Computational tools for drug discovery
We develop open-source machine learning models, deep learning architectures, and integrated software platforms that accelerate every stage of the early drug discovery pipeline starting from data capture and target prioritization to hit triage and portfolio management.

DocuStore
RAG for Drug Discovery
A chemistry-aware, domain-constrained RAG and document intelligence platform for drug discovery. DocuStore extracts and connects chemical structures, compounds, biological entities, bioactivity measurements, and scientific text to enable evidence-grounded search and question answering across complex research documents.

ChemCellar
Compound Registration & Management Platform
An open-source compound registration, management, and screening platform for research organizations. ChemCellar provides a centralized environment for managing molecules, assays, protocols, compound collections, screening data, and research projects.
DAIKON AI & Curator Studio
Next Gen of DAIKON
A collaborative drug discovery platform for scientific data curation, AI/ML prediction, analysis, and project management across the early discovery lifecycle, from genes and targets through screening, hit assessment, portfolio management, and post-portfolio studies. This is the next generation of DAIKON, extending a platform used by the TB Drug Accelerator consortium since 2022 to support collaborative tuberculosis drug discovery.

DAIKON
Data Acquisition, Integration, and Knowledge Capture Web Application
DAIKON is a cloud-deployable, open-source web platform that unifies the full target-based drug discovery pipeline under a single interface. It connects genes, targets, screens, validated hits, and project portfolios, giving multidisciplinary teams across academia, industry, and non-profits a shared, up-to-date source of truth at every stage – from target identification through to post-portfolio clinical tracking.

CAGE-Fusion
Co-Attention Graph Embedding Fusion For Nuisance Compound Detection
CAGE-Fusion is a multimodal deep learning model for identifying assay nuisance compounds, including PAINS, frequent hitters, aggregators, and reactive electrophiles that generate costly false positives in high-throughput screening. It fuses molecular graph (GNN) and SMILES sequence (Transformer) representations through a novel gated co-attention mechanism.

PARSNIP
Protein Assessment and Ranking System for Novel Input from Public users
PARSNIP is a publicly accessible target assessment tool for transparent target evaluation in Mycobacterium tuberculosis and other bacterial pathogens. Developed alongside the TBDA consortium’s target assessment framework, it allows researchers outside the consortium to apply the same rigorous multi-dimensional scoring methodology, integrating chemical validation, genetic essentiality, vulnerability, and small-molecule feasibility to any protein of interest.
