
Ligand-based virtual screening (LBVS) methods use compound information to predict activity by measuring the similarity of library compounds to reference compounds known to be active against a target of interest. 2D or 3D chemical structures and molecular descriptors of known actives are used to retrieve similar compounds from a database through similarity measures, common substructure searching, or pharmacophore matching. LBVS offers a significant advantage over other virtual screening approaches: it involves no macromolecules in the calculations and requires no prior knowledge of active ligands, making it one of the most popular strategies for drug discovery and lead optimization—especially when three-dimensional structures of potential drug targets are unavailable.
Profacgen's LBVS platform provides a comprehensive suite of ligand-based methods, including 2D molecular similarity (fingerprint-based), 3D similarity search (molecular shape, optionally colored by physicochemical or electrostatic properties), pharmacophore modeling, and machine learning–assisted 2D/3D QSAR. Our integrated in silico and experimental teams, supported by a high-performance computing cluster of 60 blades and 720 cores and a compound database exceeding 10 million purchasable molecules, deliver fast, accurate, and customizable screening solutions.
LBVS is grounded in the similarity property principle: structurally similar molecules tend to share biological activity, though activity cliffs pose notable exceptions. Molecular descriptors convert chemical structures into numerical vectors for computational comparison, spanning 0D (constitutional), 1D (fragment-based), 2D (topological), and 3D (geometric/electrostatic) representations.
Descriptor choice critically impacts LBVS outcomes. 2D fingerprints (Morgan/ECFP, MACCS, atom pair) efficiently capture connectivity and excel at identifying analogues within established series. 3D descriptors (ROCS shape overlay, EON electrostatics, pharmacophores) enable scaffold hopping by identifying structurally distinct molecules with similar spatial feature arrangements. Profacgen employs multi-descriptor consensus strategies to balance 2D sensitivity with 3D diversity-enabling capability.
Profacgen applies a range of complementary ligand-based techniques to maximize hit diversity and quality:
Figure 1. Schematic representation of virtual screening approach. (Polgár and Keseru, 2011)
Building on core similarity and QSAR methods, Profacgen offers two specialized LBVS sub-services for high-precision ligand design:

Construct robust pharmacophore models by superimposing a set of structurally diverse compounds that bind the same target and extracting the common chemical features essential for bioactivity. These models enable virtual screening of compound libraries to identify novel ligand candidates and guide the design of new chemical entities with optimized potency and selectivity.

Establish quantitative relationships between three-dimensional molecular fields (steric, electrostatic, hydrophobic) and biological activity using robust chemometric techniques such as PLS, G/PLS, and ANN. 3D-QSAR models predict the activity of untested compounds, guide the selection of promising molecules for synthesis, and drive rational lead optimization.
2D Fingerprint Similarity Searching
Rapid screening of compound libraries using Morgan (ECFP), MACCS keys, atom pair, and topological torsion fingerprints. Tanimoto, Dice, and Cosine similarity metrics enable flexible threshold tuning. Optimal for analog searching, patent landscape analysis, and chemical series expansion with sub-second database screening performance.
3D Shape and Electrostatic Matching
Shape-based virtual screening using Rapid Overlay of Chemical Structures (ROCS) for Gaussian shape comparison and Electrostatic Overlap (EON) for field-based similarity. Conformer generation with OMEGA or RDKit ETKDG ensures comprehensive conformational coverage. Enables scaffold hopping beyond 2D structural similarity.
Machine Learning-Based Activity Prediction
Supervised learning models including support vector machines (SVM), random forests, gradient boosting, and deep neural networks trained on bioactivity databases. Molecular graph convolutional networks (GCN) and transformer architectures process SMILES strings and molecular graphs for end-to-end activity prediction with uncertainty quantification.
Matched Molecular Pair Analysis (MMPA)
Systematic identification of structurally related compound pairs differing by single chemical transformations. Statistical analysis of activity shifts associated with specific structural changes guides medicinal chemistry optimization. Integration with R-group decomposition for combinatorial library design and SAR table generation.
Background:
A pharmaceutical client required novel scaffolds for selective 5-HT2C agonism to treat metabolic disorders, with the constraint of avoiding intellectual property surrounding known indazole and azaindole chemotypes.
Our Solution:
Profacgen executed a comprehensive LBVS campaign combining multiple approaches: (1) 3D pharmacophore modeling from 12 crystallographically resolved 5-HT2C ligands; (2) ROCS shape overlay screening of 3.5 million Enamine REAL compounds against a high-affinity reference agonist; (3) EON electrostatic field matching to prioritize compounds with complementary charge distributions; and (4) machine learning classification using a random forest model trained on 2,400 serotonin receptor ligands with measured binding affinities.
Final Results:
The consensus ranking identified 89 compounds from 23 distinct scaffolds for purchase and testing. Bioluminescence resonance energy transfer (BRET) assays revealed 19 compounds with EC50 < 100 nM for 5-HT2C activation, representing a 21% hit rate. Lead compound PFC-5HT-14, containing a previously unreported pyrrolo[3,2-b]pyridine core, showed EC50 = 8.3 nM with 87-fold selectivity over 5-HT2A and 134-fold selectivity over 5-HT2B, meeting the client's selectivity requirements. Subsequent optimization improved oral bioavailability from 4% to 31% in rats while maintaining sub-10 nM potency.
Background:
An academic research group sought inhibitors of the SARS-CoV-2 papain-like protease (PLpro), an essential viral enzyme with deISGylation and deubiquitination activities contributing to immune evasion.
Our Solution:
With limited crystallographic structures available at project initiation, Profacgen designed an LBVS strategy centered on machine learning. We assembled a training set of 4,200 known protease inhibitors from ChEMBL (annotated against cysteine proteases, coronavirus proteases, and viral enzymes) with pIC50 > 5.0. Molecular descriptors included 2,048-bit Morgan fingerprints, 200 physicochemical properties from RDKit, and 300 continuous data-driven descriptors from the Mordred calculator. A stacked ensemble model combining gradient boosting (XGBoost), support vector regression, and a graph attention network (GAT) achieved R2 = 0.78 on a held-out test set of 420 compounds.
Final Results:
Screening 5.6 million ZINC20 compounds with this ensemble identified 245 candidates with predicted pIC50 > 6.5. After ADMET filtering and diversity clustering, 48 compounds were procured and tested in a PLpro fluorogenic assay. Eight compounds showed IC50 < 10 μM, with PFC-PLP-03 (a 2-cyanopyrrolidine derivative) achieving 0.42 μM IC50 and antiviral activity in Calu-3 cells (EC50 = 3.8 μM, CC50 > 100 μM, SI > 26).
References:
Fill out this form and one of our experts will respond to you within one business day.