Walking through
chemical space.

I develop computational methods that integrate chemical priors into molecular representation learning and generative drug design. My research spans AI-driven drug discovery and protein–ligand and protein–protein interaction modeling, with emerging directions in phenotype-based and polypharmacology drug design. My latest work, Ouroboros, advances this broader program through a molecular foundation model that learns chemically informed representations from conformational-space and pharmacophore similarities for molecular property prediction and generation.

Postdoctoral Researcher
Institute of Systems Medicine
Chinese Academy of Medical Sciences

Ouroboros connects chemically informed molecular representations to both prediction and candidate design.

  1. 01 · Recognize

    Chemically meaningful similarity

    Conformational-space and pharmacophore relationships keep functionally relevant molecules connected even when their 2D scaffolds differ.

  2. 02 · Generalize

    Reusable prediction

    The same molecular space supports property and cellular-activity prediction, similarity search, and target-oriented screening.

  3. 03 · Design

    Property-guided exploration

    Prediction objectives can guide candidate design for directed optimization and targeted polypharmacology.

Advanced Science · 2026

Ouroboros

Ouroboros addresses a persistent divide in AI-driven drug discovery: representations optimized for prediction are rarely reusable for molecular design. By organizing chemical space with fingerprint, conformational-space, and pharmacophore relationships, it supports property prediction, virtual screening, and property-guided generation within one representation. This creates a common basis for hit discovery, lead optimization, and multi-target design.

From recognizing cross-scaffold similarity to designing within it.

GeminiMol showed that conformational-space and pharmacophore relationships can reveal functionally related molecules across scaffold boundaries. Ouroboros extends this concept into a shared chemical space that connects prediction with candidate design.

GeminiMol

Functional similarity across scaffolds

GeminiMol prioritizes molecules with related 3D and pharmacophore profiles even when their 2D structures are dissimilar.

Outcome cross-scaffold screening and the experimentally validated GluN1/GluN3A inhibitor GM-10.

Ouroboros

Prediction and design in one chemical space

Ouroboros organizes several molecular relationships together so prediction and generative exploration share a chemically meaningful space.

Scope property prediction, directed optimization, and targeted polypharmacology.

Composite research figure showing the progression from GeminiMol cross-scaffold similarity search to Ouroboros prediction and molecular design.
Research progression and representative discovery settings. Open the full-resolution figure ↗

01 · Similarity

Recognize similarity beyond 2D scaffolds

Conformational-space and pharmacophore relationships preserve functionally relevant neighborhoods across structurally distinct molecules.

02 · Prediction

Transfer chemical knowledge across tasks

A shared representation supports property prediction, similarity-based screening, and target-oriented prioritization across discovery settings.

03 · Design

Connect prediction with molecular design

Property objectives guide the proposal of candidates for directed optimization and targeted polypharmacology.

Ouroboros symbol showing a molecular graph enclosed by a serpent beside an ordered molecular representation vector.
The visual mark summarizes a molecular graph mapped to an ordered representation vector.

A shared chemical space turns molecular relationships into testable design hypotheses.

By preserving fingerprint, conformational-space, and pharmacophore relationships, Ouroboros supports cross-scaffold search, transferable property models, hit optimization, and multi-target design in a common representation. Proposed molecules remain candidates for computational and experimental evaluation.

  1. 01

    Cross-scaffold similarity

    Functionally relevant neighborhoods can remain visible even when molecules share little two-dimensional structural similarity.

  2. 02

    Property prediction

    One representation supports diverse molecular-property and cellular-activity objectives.

  3. 03

    Directed optimization

    Candidate design can pursue selected property profiles while retaining relevant features of a starting molecule.

  4. 04

    Polypharmacology hypotheses

    Multiple references or objectives can define candidates intended to balance several target profiles.

The shared representation supports complementary scientific problems across early drug discovery.

Cross-scaffold screening

Prioritize molecules whose conformational and pharmacophore relationships suggest shared function despite low 2D similarity.

Hit-to-lead optimization

Explore property improvements while retaining chemically meaningful relationships to a starting hit.

Multi-target design

Propose candidates intended to balance several target or phenotype objectives for targeted polypharmacology.

Ouroboros research figure connecting chemically informed molecular representations with cross-scaffold screening, hit discovery, and hit-to-lead optimization.
Applications evaluated in the study: chemical-space exploration and optimization from a known hit. Open the full-resolution figure ↗

Selected research

This selection includes methods I led or co-developed and broader collaborative studies to which I contributed. Together they span molecular representation, phenotypic screening, protein structure and interactions, and synthetic accessibility.

2024

Molecular representation

GeminiMol

Finding functional similarity beyond shared 2D scaffolds.

GeminiMol addresses the problem that functionally similar molecules can appear unrelated in two-dimensional structure. It supports cross-scaffold property prediction, virtual screening, target identification, and scaffold hopping; screening 18 million compounds led to GM-10, an experimentally validated GluN1/GluN3A inhibitor.

2025

Multimodal phenotype

PhenoModel

Aligning molecular graphs with cellular morphology.

Cell morphology offers a complementary view of compound activity. PhenoModel is a collaborative study that aligns molecular graphs with cellular images in a multimodal model for phenotypic virtual screening. The work explores how molecular structure and cellular responses can inform the prioritization of compounds for further study.

2026

Protein conformations

ProteinConformers

Evaluating simulated protein conformational landscapes.

Protein simulations generate many structures, making the assessment of entire conformational landscapes an important challenge. ProteinConformers is a collaborative benchmark for evaluating the diversity and plausibility of simulated protein conformations. It provides a framework for assessing the ensembles produced by simulation methods, alongside research on protein structure and molecular recognition.

2022

PPI and molecule glue

PPI-Miner

Finding receptor-compatible backbone motifs beyond sequence similarity.

PPI-Miner reveals interaction partners that sequence-motif searches can miss because unrelated proteins may preserve the local backbone geometry required for receptor recognition. In the CRBN case, it produced 1,739 candidates and recovered 16 previously reported cases; at least 12 additional candidates in the released library later received compound-dependent recruitment support.

2023

Synthetic accessibility

DeepSA

Predicting synthetic accessibility for molecular design.

Molecular design also needs to consider whether proposed compounds may be difficult to synthesize. DeepSA is a co-developed deep-learning predictor of small-molecule synthetic accessibility. It provides an estimate of synthesis difficulty that can inform candidate prioritization and complement other computational assessments during molecular design.

Open tools for everyday research.

I also maintain or contribute to practical tools for database searching, molecular simulation, and reproducible computer-aided drug-design workflows.

01 · Database navigation

Biodb-Search

A browser-based navigator that sends one query to multiple biomedical databases, designed to make routine literature and data lookup more direct.

02 · Molecular dynamics

AutoMD / AutoTRJ

Automation for Desmond system setup and molecular-dynamics simulation, paired with a configurable pipeline for trajectory processing and analysis.

03 · CADD workflows

CADD-Scripts

Shell utilities for virtual screening, cross-docking, and protein modeling with Schrödinger and Rosetta, shared to support reproducible computational workflows.

Recent publications

The full record contains 35 publications and book chapters from 2020–2026 across molecular representation, structural biology, protein interactions, and drug discovery.

Loading publication record…

View all publications

News

Selected milestones from collaborative research, software releases, competitions, and publications, including the archive from the earlier site.

  1. The CoCoBind study connects RNA–compound interaction prediction with nucleotide-level binding-site localization, supporting site-aware RNA–ligand recognition under distribution shift. It was published in Journal of Medicinal Chemistry.

  2. The Ouroboros study, Learned Conformational Space and Pharmacophore Into Molecular Foundational Model, was published in Advanced Science.

  3. The PhenoModel study on multimodal phenotypic drug design was published in Acta Pharmaceutica Sinica B.

  4. The PhenoScreen project introduced a dual-space contrastive learning framework combining conformational-space and molecular-image similarities for virtual screening.

  5. A collaborative study using GeminiMol reported a GluN1/GluN3A inhibitor (IC50 = 0.98 μM); the project received first prize in the 2023 Shanghai International Computational Biology Innovation Competition.

  6. The GeminiMol paper was published in Advanced Science.

  7. The GeminiMol project introduced molecular representation learning informed by conformational-space similarity; see the code and bioRxiv preprint.

  8. DeepSA was published in the Journal of Cheminformatics; the DeepSA web server is also available.

  9. DeepSA, co-developed with Shihang Wang, was released as a model for estimating the synthetic accessibility of small molecules.

  10. The original plmd scripts were expanded and renamed AutoMD, adding support for additional force fields and configurable trajectory-analysis pipelines.

  11. The PPI-Miner study showed how conserved backbone motifs can reveal protein-interaction candidates even when their sequences diverge. The paper appeared in the Journal of Chemical Information and Modeling, with source code on GitHub.

  12. A retrospective cross-reference of the released CRBN candidate database found at least 12 candidates beyond the paper's original set of 16 that were later supported by compound-dependent CRBN-recruitment assays; mutating the predicted G-loop glycine abolished the recruitment signal (Science, 2025).

  13. A collaborative study reported an allosteric pocket on the SARS-CoV-2 spike protein and inhibitor candidates designed to restrict its conformational change; the paper appeared in the Journal of Medicinal Chemistry.

  14. A collaborative study reported five candidate small-molecule blockers targeting the SARS-CoV-2 spike protein in Acta Pharmacologica Sinica.

  15. Automated Schrödinger workflows for cross-docking (XDock), virtual screening (GVSrun), and molecular-dynamics simulation (plmd) were released, with source code on GitHub.

  16. Biodb-Search was released as a browser-based navigator that submits one query to multiple biomedical databases.

  17. GetPDB was released as a script for downloading single-chain protein structures by UniProt ID.

Training and current research.

My training has progressed from life science and structure-based modeling to machine learning for molecules, proteins, and cellular phenotypes.

  1. 2019

    B.S. in Life Science

    Northeast Agricultural University

    Training in life science, followed by early work with computational drug-discovery workflows.

  2. 2019–2024

    Ph.D.

    ShanghaiTech University

    Research in structural biology, molecular dynamics, protein interactions, and chemically informed molecular representation.

  3. Current

    Postdoctoral Researcher

    Institute of Systems Medicine, CAMS

    Research on molecular and protein machine-learning models, phenotype-aware representation learning, and computational drug-discovery applications.

Research interests and collaboration.

I welcome discussions and collaborations that connect molecular modeling and machine learning with experimentally testable drug-discovery questions.

Wanglin1102@outlook.com