🛠 Open Source / Software

I value practical and reproducible research. This page highlights open-source libraries, benchmarks, and system tools built by my group and collaborators. Many of these projects make AI systems auditable across the deployment stack: anomaly and out-of-distribution detection in data, trust and robustness evaluation of foundation models, and auditability and control of agent systems across their lifecycle, with applications in science and high-stakes domains. Several of these projects have been accepted into the Anthropic Claude for Open Source Program and OpenAI's Codex for Open Source. For all repositories, see my GitHub profile.

External adoption: PyOD is named by OpenAI as expected operational tooling in its Technical Intelligence Analyst role, shipped as a first-class ModelHandler in Apache Beam (Apache Software Foundation), runs the live-traffic alerting subsystem in PostHog (34K+ stars), is the canonical anomaly-detection flavor in MLflow's community-flavor docs, and is embedded in Genentech (Roche) drug-discovery validators. Across the wider ecosystem, 5,493 public repositories and 139 packages depend on PyOD (May 2026 snapshot). DoD CDAO lists PyOD and TrustLLM; ESA OPS-SAT uses PyOD; NIST AI 100-2e2025 and the FLI AI Safety Index cite TrustLLM. On the agent-auditing line, Aegis is used as a quantitative baseline (Praetor, arXiv:2604.26274) and Auditable Agents has drawn third-party citations and editorial coverage within weeks of release.

Adoption & Recognition

PyOD — Python Library for Outlier Detection

9,800+ stars · 50M+ downloads · 5,493 dependent repos + 139 packages (GitHub Dependents, May 2026) · GitHub
TypeEvidenceSource
AI Lab OpenAI Careers names PyOD as expected operational tooling in the Technical Intelligence Analyst job posting Qualifications block: "Have experience with anomaly detection tools, such as PyOD, and discovery processes for surfacing novel or low-prevalence patterns." OpenAI · snapshot
UK Gov UK Government Algorithmic Transparency Record (London Borough of Sutton, "Access Assure" technology-enabled care) links PyOD KNN documentation in its Model Specification (section 4.2.6); a production deployment record GOV.UK
Space Agency Selected by ESA for OPS-SAT spacecraft telemetry benchmark (all 30 algorithms) Nature Sci. Data
Saudi Gov Saudi Data & AI Authority (SDAIA) Deepfakes Guidelines names PyOD (p.10) among recommended tools for detecting suspicious activity patterns SDAIA PDF
U.S. DoD CDAO Generative AI Responsible AI Toolkit lists PyOD as a Production / High-maturity OOD-detection tool (entry p.49, embedded in Stage 3.1.10 assessment workflow) ai.mil PDF
EU Project SEDIMARK Horizon Europe D3.1 (p.18) names PyOD and TODS in the outlier-detection module of the EU data-space toolbox SEDIMARK D3.1
Gov / Labs Cited in research papers by authors affiliated with Deutsche Bundesbank, NIH, CDC, RAND, NASA JPL, German DLR and DESY, and the Sandia, Brookhaven, and Argonne national labs, plus multiple Fraunhofer institutes (citing ECOD, COPOD, PyOD, ADBench, TODS, and LSCP) Audit details
Platform Apache Software Foundation / Apache Beam (8.5K+ stars) ships a first-class PyOD ModelHandler at sdks/python/apache_beam/ml/anomaly/detectors/pyod_adapter.py; Apache Beam underlies Google Cloud Dataflow apache/beam
Enterprise PostHog (34K+ stars, YC unicorn product analytics) runs a multi-detector PyOD subsystem at posthog/tasks/alerts/detectors/pyod_detectors/ for live-traffic alerting (eight algorithm wrappers: KNN, IForest, COPOD, ECOD, OCSVM, LOF, PCA, HBOS) PostHog/posthog
Platform MLflow (25.8K+ stars) official community-flavor docs list PyOD as the canonical anomaly-detection flavor with worked KNN-detector example via mlflavors mlflow/mlflow
Pharma Genentech (Roche) Data Detective embeds PyOD/ADBench in its drug-discovery validator factories (adbench_validator_method_factory.py, adbench_multimodal, adbench_ood_inference) Genentech/data-detective
Enterprise Walmart real-time pricing anomaly detection (1M+ daily updates) KDD 2019
Enterprise Databricks Kakapo framework for unsupervised outlier detection Databricks Blog
Enterprise IQVIA healthcare fraud detection (123K+ pharmacy claims) SUOD Paper
Enterprise Ericsson Anomaly Detection Framework (E-ADF) built on PyOD Ericsson Blog
Patents 15 patents cite PyOD/COPOD/ECOD/LSCP/SUOD/TODS (WIPO x2, EU x2, US x4, China x6, Slovakia x1); recent additions include Tencent CN112989338B (COPOD), DBAPPSecurity CN117216660B (TODS), and Chongqing University CN117648656A (ADBench) Ericsson · Actimize · Tencent
Journal Two 2026 Nature Scientific Reports papers implement anomaly detection via PyOD in their Methods (eight PyOD detectors in one; COPOD/ECOD/IForest in the other), each citing the PyOD JMLR paper s41598-026-45091-2
Encyclopedia Wikipedia "Anomaly detection" Software section names PyOD; reference list cites Zhao, Nasrullah, Li 2019 JMLR Wikipedia
Education Featured in 5 books (Manning, O'Reilly, Apress, Routledge, IntechOpen) Manning
Education DataCamp course with dedicated chapter (19M+ platform learners) DataCamp

Agent Auditability Line — auditable, Aegis, agent-audit, Auditable Agents, GRADE, Implicit Execution Tracing

Early external validation of the lab's current frontier (2026 releases)
TypeEvidenceSource
Benchmark Baseline Praetor (arXiv:2604.26274) deploys Aegis as a quantitative baseline behavioral firewall, reporting 12.8% attack success for Aegis versus 2.2% for their method, and cites the Aegis paper by title and arXiv ID arXiv
Media The Agent Times feature names Yue Zhao and explains all five auditability dimensions of the Auditable Agents framework Agent Times
Industry Blog WIDTH applies the Auditable Agents five dimensions and the 8.3 ms firewall-overhead result to compliance infrastructure WIDTH
Media Machine Brief explains the mechanism, results, and limitations of Implicit Execution Tracing (recovering which agent produced a harmful output after logs are stripped) Machine Brief
Downstream Citations Auditable Agents has seven third-party scholarly citations within weeks of posting, including AgentBound, DEMM, Trace2Policy, and OpenClawBench AgentBound · DEMM
Ecosystem agent-audit is indexed in the ReputAgent ecosystem directory; awesome-auditable-ai is referenced from a Hacker News thread and indexed by Ecosyste.ms ReputAgent · HN

ADBench — Anomaly Detection Benchmark

1,000+ stars · NeurIPS 2022 · GitHub
TypeEvidenceSource
Consulting Deloitte Germany cites ADBench in an AIxAML anti-money-laundering transaction-monitoring solution Deloitte PDF
Enterprise Cited in papers with authors affiliated with Microsoft Research, Tencent, Amazon, BlackRock, Visa, Bosch, Siemens, and Ericsson Audit details
Pharma Genentech (Roche) Data Detective uses ADBench in adbench_validator_method_factory.py, adbench_multimodal, and adbench_ood_inference validator factories for drug-discovery data validation Genentech/data-detective
Journal Nature Communications study links the ADBench repository in its Data availability section for the benchmark anomaly datasets used Nature Comms

TrustLLM — Trustworthiness Benchmark for LLMs

620+ stars · ICML 2024 · GitHub
TypeEvidenceSource
U.S. Senate Cited in HSGAC "Hedge Fund Use of Artificial Intelligence" report (footnote 119) Senate PDF
U.S. DoD Listed in CDAO Generative AI Responsible AI Toolkit ai.mil PDF
NIST Named in NIST AI 100-2e2025 Section 3.6 "Benchmarks for AML Vulnerabilities" NIST PDF
Policy Official benchmark in all 4 editions of the FLI AI Safety Index (2024, 2025 x2, Summer 2026) FLI Report · Summer 2026
National Lab Lawrence Livermore National Laboratory feature article; LLNL/DOE SafeAI report cites TrustLLM LLNL · SafeAI PDF
International Cited in International AI Safety Report 2026 (citation #881; led by Yoshua Bengio, 100+ experts, 30+ countries) Report
U.S. DOE Oak Ridge National Laboratory technical report ORNL/TM-2025/3935, "Scalable Workflow for Evaluating Trustworthiness of Large Language Models," discusses and cites TrustLLM OSTI
G7 / OECD Salesforce G7 Hiroshima AI Process Transparency Report (OECD-hosted) cites TrustLLM (p.4) among its trust-and-safety evaluation metrics OECD
Industry Lab NTT Technical Review names TrustLLM as an LLM-safety benchmark and links the benchmark repository (reference [4]) NTT
Media Featured by 机器之心 (Jiqizhixin) and 澎湃新闻 (The Paper) 机器之心 · 澎湃
Enterprise Editorial Samsung SDS Insights treats TrustLLM as a flagship LLM trustworthiness evaluation framework in its Korean enterprise editorial; reference list cites arXiv:2401.05561 Samsung SDS

Recent Institutional Visibility — 2026 research releases

Selected institutional pages for newer benchmarks and datasets
TypeEvidenceSource
Institute Vector Institute highlights TrustGen in its ICLR 2026 research roundup Vector
Industry Lab Adobe Research lists FigEdit ("Charts Are Not Images") and the benchmark release Adobe Research

TDC — Therapeutics Data Commons

1,200+ stars · NeurIPS 2021 · with Harvard & Stanford · GitHub
TypeEvidenceSource
Journal Published in Nature Chemical Biology (2022) Nature Chem. Bio.
University Harvard Medical School feature: "Can AI transform drug discovery?" HMS News
Science Press Phys.org syndication of Harvard article Phys.org
Industry Amazon Science feature article Amazon Science
Pharma Cited by researchers at AstraZeneca, Pfizer, Roche, Novartis, Merck, Sanofi, Eli Lilly Audit details
Labs Cited in papers by researchers at Los Alamos and Brookhaven national labs (cheminformatics, model uncertainty) and OpenAI (biomedical reasoning) Audit details

DoxBench — Geolocation Privacy Leakage Benchmark

ICLR 2026 · GitHub
TypeEvidenceSource
Policy Cited by Privacy International in "Nowhere to Hide? Privacy Risks and Policy Implications of AI Geolocation" (p.28, footnote 56) Report
Chinese Media 机器之心Pro reporting via Sina names "南加州大学教授赵越(Yue Zhao)团队", paper title "Doxing via the Lens", and the arXiv link Sina

Full Project Portfolio

Flagship project pages: agent-style, anywhere-agents, agent-audit, Aegis, PyOD 3.

Sort by: