Fast QSAR, Fast Lane, Slow Endpoints, and EHS Endpoints run different endpoint scopes from the QSAR model library.
Prerequisite: Load The Molecule From ChemrytIQ
Before opening ChemrytIQ-QSAR, search the molecule in ChemrytIQ by SMILES, InChI, molecule name, or CAS number. Confirm the correct molecule on the ChemrytIQ page, then open the required Chemryt app from that same molecule context so the selected structure is loaded into the app automatically.
What It Does
ChemrytIQ-QSAR submits molecule SMILES and selected endpoints to server-side prediction APIs, runs model inference, and returns endpoint-level ADME, toxicity, physicochemical, environmental, and safety signals with interpretation context.
Select SMILES, SELFIES, or both to control which model representation endpoints are submitted and how results are separated.
Read the combined QSAR tab first, then compare SMILES Endpoints and SELFIES Endpoints tabs when both representations are available.
Quick Tutorial
- Load a valid molecule in ChemrytIQ and confirm the canonical SMILES before opening QSAR.
- In Model, select SMILES for the standard chemistry string model, SELFIES for the robust SELFIES representation, or both when you want side-by-side representation evidence.
- Run Fast QSAR first. It runs the priority 12-endpoint panel and quickly populates the Latest QSAR Run summary plus the QSAR result tab.
- Run Fast Lane when you want the broader fast library. The button may show 96 total endpoints with 67 slow endpoints deferred for later.
- Run Slow Endpoints only after Fast QSAR/Fast Lane when deeper endpoints are needed and you can wait longer.
- Run EHS Endpoints when environmental, health, safety, ecotoxicity, persistence, exposure, or physical hazard screening is the main question.
- Use Clear QSAR memory before switching molecules, changing representation choice, retrying after a stopped run, or clearing stale partial results.
Main Areas
| Area | What to review | When to use it |
|---|---|---|
| Model selector | SMILES and SELFIES checkboxes under the Model legend. | Use before running QSAR to choose which representation-specific endpoints are submitted. |
| Run buttons | Run Fast QSAR (12), Run Fast Lane (96, 67 slow later), Run Slow Endpoints (67), Run EHS Endpoints (46), and Clear QSAR memory. | Use to control endpoint scope, runtime, and cache/retry behavior. |
| Latest QSAR Run | Completion summary such as active endpoint count, representation mix, and model-library count. | Use to confirm what actually completed before interpreting tabs. |
| Result tabs | QSAR, SMILES Endpoints, and SELFIES Endpoints with badge counts. | Use to read combined output first, then representation-specific endpoint cards. |
| Endpoint cards | Endpoint value/class, model confidence, applicability-domain notes, skipped/error status, and health/chemistry grouping. | Use to decide whether an endpoint is strong enough for follow-up. |
What The Model Selector Means
The Model box controls the molecular representation used by QSAR endpoints. It does not change the molecule itself; it changes which model-representation lane is requested for the current molecule.
| Control | Meaning | When to select it |
|---|---|---|
| SMILES | Runs endpoints that use the canonical SMILES representation from the current molecule. | Use for the standard QSAR readout and for most routine screening. Confirm the SMILES is correct before running. |
| SELFIES | Runs endpoints that use SELFIES, a robust molecular string representation designed to avoid invalid molecule strings. | Use when you want representation-robust evidence, SELFIES-aware endpoint comparison, or later token-hotspot/analog work. |
| SMILES + SELFIES | Runs both available representation lanes and separates the completed endpoint cards into SMILES and SELFIES tabs. | Use when endpoint confidence matters and you want to see whether both representations point in the same direction. This may take longer and endpoint counts can differ by representation. |
Run Buttons And When To Use Them
QSAR is split into lanes so users can get a quick answer first, then expand only when the molecule deserves deeper screening.
| Button | Purpose | What to expect |
|---|---|---|
| Run Fast QSAR (12) | Runs the priority fast panel first. | Use this first for triage. It should quickly update Latest QSAR Run and produce the first QSAR endpoint cards. In the screenshot example, 12 active endpoints completed from the 209-model library. |
| Run Fast Lane (96, 67 slow later) | Expands beyond the 12 priority endpoints into the broader fast lane while deferring the slow endpoint set. | Use after Fast QSAR when the compound looks worth deeper review. Expect more endpoint cards and a broader ADME/safety/property profile, but not the full slow panel yet. |
| Run Slow Endpoints (67) | Runs endpoints marked as slower or deferred. | Use after Fast QSAR/Fast Lane when slow endpoints are needed for a final screening note or when a project decision depends on those endpoints. |
| Run EHS Endpoints (46) | Runs the environmental, health, and safety endpoint lane. | Use for biodegradation, bioaccumulation, ecotoxicity, exposure, environmental fate, physical hazard, or safety-screening questions. EHS can run separately from the normal fast/slow sequence. |
| Clear QSAR memory | Clears the latest QSAR run state, cached partial results, reused endpoint results, progress messages, and tab counts. | Use before changing molecule, changing SMILES/SELFIES choices, retrying failed/skipped endpoints, or removing stale results. It does not delete the molecule structure. |
How To Read The Result Tabs
The tab badge is a completed-result count, not always the configured endpoint count. Counts can differ because some endpoints are representation-specific, skipped, failed, filtered, or still streaming.
| Tab | What it shows | How to use it |
|---|---|---|
| QSAR | Combined QSAR result view with all completed, non-skipped endpoint cards from the current run. | Start here. It gives the broad decision-support view across active SMILES and SELFIES outputs. |
| SMILES Endpoints | Only endpoint cards whose prediction came from the SMILES representation. | Use to inspect the standard representation outputs and compare them with the SELFIES tab when both were selected. |
| SELFIES Endpoints | Only endpoint cards whose prediction came from the SELFIES representation. | Use to see representation-robust outputs and to support SELFIES hotspot or analog-perturbation interpretation. |
| Latest QSAR Run | A run summary above the tabs, such as fast panel endpoint count, representation mix, deferred endpoints, and model-library count. | Read this before the tabs so you know which lane completed and whether results are partial, fast-only, EHS-only, or expanded. |
Calculated Molecular Descriptors
The Descriptors tab is a separate molecular-characterization report calculated from the current structure with RDKit, Mordred Community, and formula-specified implementations. It is useful for interpretation, comparison, and feature review, but it does not replace endpoint predictions or alter the existing QSAR model input vector. Expand a category or search by descriptor ID, scientific name, or category.
| Display item | What it means | How to use it |
|---|---|---|
| Value | The calculated descriptor value. Most descriptors are unitless indices; dimensional descriptors state or imply their scientific unit in the name/definition. | Compare only like-for-like descriptor IDs calculated by the same method. Do not assume that a larger number is inherently better. |
| Exact match | The displayed ID exactly matches an entry in the system descriptor catalog and uses its catalog scientific description. | Use the Catalog ID and description when recording or exporting the result. |
| Family | The engine descriptor was assigned to the nearest scientific family, but its name is not an exact catalog-name match. | Treat the engine ID as authoritative and avoid claiming exact equivalence to a similarly named catalog descriptor. |
| n/a | The value could not be calculated, was non-finite, needs an unavailable atom type, or requires geometry that could not be generated. | Do not convert n/a to zero. Zero can be a valid result, especially for an absent feature or atom pair. |
| Geometry note | 3D families use the deterministic conformer method reported below the table: ETKDGv3 followed by MMFF94 when possible, otherwise UFF or an unoptimized conformer. | Treat 3D values as conformation-dependent estimates. Experimental conformers, protonation, stereochemistry, salts, and tautomer choice can change them. |
| Formula implementation notes | Records special conventions used by catalog formulas, including heavy-atom scope, population standard deviation, atomic weighting, distance units, and fallback behavior. | Open this section before reproducing a value in another package or comparing it with a publication. |
Descriptor Categories: Composition, Topology, And 2D Structure
These categories describe what the molecule contains and how atoms are connected. They are derived primarily from the molecular graph and normally do not require a 3D conformer.
| Category | What it describes | Special interpretation |
|---|---|---|
| Constitutional indices | Atom, element, bond, mass, heteroatom, hydrogen, and simple composition statistics. | Strongly size-dependent; compare normalized or average variants when molecule sizes differ greatly. Formula statistics use non-hydrogen atoms, and reported standard deviations use the population denominator. |
| Ring descriptors | Ring number, size, aromaticity, fusion, spiro, bridgehead, and ring-system composition. | Ring perception depends on the sanitized molecular graph; fused and aromatic systems should not be interpreted as a simple count of drawn circles. |
| Topological indices | Global graph shape, branching, centrality, distance, Zagreb, Wiener, Balaban, complexity, and related indices. | Unitless and often correlated with molecular size; use them comparatively rather than as pass/fail limits. |
| Walk and path counts | Counts or weighted summaries of graph walks and paths of different lengths. | Long-order values can grow rapidly with molecular size and branching. |
| Connectivity indices | Kier-Hall/Randić-style connectivity and valence-connectivity measures. | Order and suffix matter; zero-, first-, and higher-order variants are not interchangeable. |
| Information indices | Graph information content, symmetry, neighborhood diversity, and related entropy-style measures. | Values depend on the partition or neighborhood order used; compare the same descriptor ID only. |
| 2D matrix-based descriptors | Spectral and matrix operators derived from adjacency, distance, Laplacian, chi, reciprocal-distance, Barysz, or related matrices. | Suffixes identify the source matrix or atomic weighting. Eigenvector signs are mathematically arbitrary; the system fixes sign deterministically for reproducible signed VE values. |
| 2D autocorrelations | ATS, ATSC, Moran, Geary, and related correlations of atomic properties across topological lags. | The lag is graph-bond distance, not Angstrom distance. Weight suffixes can represent mass, volume, electronegativity, polarizability, ionization potential, intrinsic state, or charge. |
| Burden eigenvalues | Highest/lowest eigenvalues of modified Burden matrices with atomic-property weighting. | Names encode weighting and rank. The implementation uses carbon-normalized diagonal weights, square-root bond orders, small nonbonded entries, and zero padding when rank exceeds heavy-atom count. |
| ETA indices | Extended topochemical atom indices describing size, branching, functionality, and electronic/topological environment. | Composite and local variants have specific formulas; do not interpret ETA values as measured electronic properties. |
| Edge adjacency indices | Indices derived from relationships between bonds rather than only between atoms. | Sensitive to bond topology and bond typing; different weighting suffixes represent distinct descriptors. |
| 2D Atom Pairs | Counts or weighted occurrences of atom-type pairs separated by a stated topological distance. | Distance is in bonds. A zero usually means that the requested pair at that lag is absent. |
| MDE descriptors | Molecular distance-edge descriptors for selected atom types and graph distances. | Atom-type-specific and size-sensitive; compare the same element-pair definition. |
Descriptor Categories: Surface, Chemistry, And Medicinal-Chemistry Features
These categories summarize chemical functionality, atomic environments, surface partitions, and familiar developability heuristics.
| Category | What it describes | Special interpretation |
|---|---|---|
| P_VSA-like descriptors | Partitions molecular surface area into bins based on partial charge, LogP contribution, molar refractivity, E-state, or pharmacophore character. | A suffix/bin is part of the definition. Bin values are surface contributions, not whole-molecule LogP, charge, or PSA. |
| Functional group counts | Counts recognizable chemical groups such as carbonyls, acids, amides, amines, halides, and heterocycles. | Overlapping patterns may count different chemical concepts in the same region. A zero means no implemented pattern matched, not proof that every possible functional-group definition was tested. |
| Atom-centred fragments | Counts atoms in specific local element, valence, hydrogen, aromaticity, and neighbor environments. | Only the explicitly implemented environment rules are exact formula matches; other fragment families can remain unavailable. |
| Atom-type E-state indices | Electrotopological-state sums, extrema, and counts for atom types. | Uses intrinsic state plus graph-distance perturbation on heavy-atom connectivity. Empty atom types aggregate to zero; isolated-atom global extremes may be undefined. |
| Pharmacophore descriptors | Counts and distance distributions of donor, acceptor, hydrophobic, aromatic, positive, and negative feature pairs. | They are rule-based feature encodings, not evidence of binding to a particular protein. Pair distance may be topological or 3D depending on the descriptor family. |
| Charge descriptors | Partial-charge extrema, totals, surface-charge summaries, and charge-transfer/topological charge indices. | Partial charges are calculation-model dependent. Gasteiger-based values should not be treated as experimental charge or a quantum-mechanical observable. |
| Molecular properties | Whole-molecule properties such as molecular weight, LogP, refractivity, polar surface area, volume, and polarizability. | Check the specific algorithm: calculated LogP/PSA values can differ from experimental measurements and from other software implementations. |
| Drug-like indices | Lipinski, Veber, QED, and related rule or desirability summaries. | These are prioritization heuristics, not safety, efficacy, bioavailability, or regulatory decisions. Borderline values should be interpreted with endpoint and assay context. |
| Chirality descriptors | Counts and indices describing stereocenters, chiral topology, and stereochemical complexity. | Results require correctly specified stereochemistry; unspecified centers and mixtures can make the descriptor incomplete or ambiguous. |
| SASA descriptors | Solvent-accessible surface area and surface partitions by atom/property type. | Probe radius, atomic radii, protonation, and conformation affect SASA. Do not equate SASA directly with TPSA. |
Descriptor Categories: Geometry And 3D Shape
These categories require or benefit from a generated 3D conformer. They describe one standardized computational geometry, not the full solution-state conformational ensemble.
| Category | What it describes | Special interpretation |
|---|---|---|
| Geometrical descriptors | Molecular dimensions, radius of gyration, eccentricity, shape, aromatic bond geometry, displacement, and quadrupole-style indices. | Distance-based values use the generated conformer and can change with stereochemistry and conformational sampling. |
| 3D matrix-based descriptors | Spectral operators calculated from geometry, reciprocal geometry, geometry/topology ratio, and Coulomb matrices. | Matrix suffixes G, RG, G/D, and Coulomb define different mathematics and scales; never compare their raw values as if they were the same property. |
| 3D autocorrelations | Atomic-property autocorrelations combining topological lags with Cartesian separation. | TDB descriptors use lags 1-10, Angstrom distances, carbon-scaled weights, and RDKit covalent radii for the r suffix. |
| RDF descriptors | Radial distribution function signals describing weighted pair-distance distributions. | Uses Angstrom distances and the documented smoothing convention (beta 100 A^-2, normalization f=1). Signal position and weight suffix are essential parts of the ID. |
| 3D-MoRSE descriptors | Electron-diffraction-inspired transforms of weighted interatomic distances over signal values 01-32. | Signal 01 uses the sinc limit; suffixes encode atomic weighting, including Gasteiger charge. Values are conformer dependent. |
| WHIM descriptors | Weighted Holistic Invariant Molecular shape, symmetry, size, and directional distribution indices. | Computed from centered 3D coordinates with atomic-property weighting; degenerate or highly symmetric geometries can yield zero or undefined components. |
| GETAWAY descriptors | Geometry, topology, and atom-weight assembly descriptors using leverage and influence/distance relationships. | Sensitive to the generated coordinates and weighting scheme; leverage-based terms are not direct measures of biological influence. |
| Randic molecular profiles | Shape profiles derived from interatomic-distance distributions and molecular size. | Profile positions and summary indices must be compared at matching definitions and conformer methods. |
| 3D Atom Pairs | Euclidean distances or summaries for selected element pairs in 3D. | G(X..Y) sums heavy-atom pair distances in Angstroms and returns zero when the requested pair is absent. |
| CATS 3D descriptors | Three-dimensional pharmacophore atom-pair counts across distance bins. | Feature assignment and bin boundaries matter; they encode spatial pattern similarity, not receptor-specific activity. |
| WHALES descriptors | Weighted Holistic Atom Localization and Entity Shape descriptors based on 3D molecular landmarks. | Conformer and atomic weighting affect the result; use the full descriptor ID and the same geometry protocol for comparisons. |
Formula Reference Used By The Descriptor Engine
The table below summarizes the principal equations implemented directly by the QSAR descriptor service. A is the number of heavy atoms, E is the molecular-graph edge set, d_i is vertex degree, d_ij is topological distance, r_ij is Cartesian distance in Angstroms, w_i is an atomic-property weight, and Delta_k is the number of atom pairs at lag k. RDKit and Mordred descriptors that are not formula-specified here use their library implementations; the engine name beside each result identifies that source.
| Descriptor family | Formula used | Implementation details |
|---|---|---|
| Atomic-property sum and mean | S_X = sum_i w_i; M_X = (1/A) sum_i w_i. | Uses non-hydrogen atoms. For catalog-weighted properties, atomic values are normally scaled to carbon, so w_C = 1. |
| Population standard deviation | stdX = sqrt[(1/A) sum_i (x_i - x_bar)^2]. | Used for atomic mass, van der Waals volume, Sanderson electronegativity, polarizability, and ionization-potential statistics. The denominator is A, not A-1. |
| Zagreb indices | ZM1 = sum_i d_i^2; ZM2 = sum_(i,j in E) d_i d_j. | Degrees count heavy-atom graph neighbors. |
| Randić/Kier connectivity | X0 = sum_i d_i^(-1/2); X1 = sum_(i,j in E) (d_i d_j)^(-1/2). Valence variants replace d_i with the valence vertex degree. | Higher-order connectivity descriptors sum the corresponding path products. Zero-degree terms are excluded where division would be undefined. |
| Broto-Moreau autocorrelation | ATS_k = ln[1 + sum_(d_ij=k) w_i w_j]. | k is topological lag. The value is unavailable if the logarithm argument is not positive. |
| Centred autocorrelation | ATSC_k = sum_(d_ij=k) (w_i - w_bar)(w_j - w_bar). | Centered on the molecule-wide mean atomic weight. |
| Moran autocorrelation | MATS_k = {Delta_k^(-1) sum_(d_ij=k)(w_i-w_bar)(w_j-w_bar)} / {A^(-1) sum_i(w_i-w_bar)^2}. | Undefined when the atomic property has zero variance or no pairs exist at the requested lag. |
| Geary autocorrelation | GATS_k = {[2 Delta_k]^(-1) sum_(d_ij=k)(w_i-w_j)^2} / {[(A-1)^(-1)] sum_i(w_i-w_bar)^2}. | Undefined for zero variance or an empty lag. |
| 2D atom pairs | B_k[X-Y] = 1 when at least one X/Y pair exists at d_ij=k, otherwise 0; F_k[X-Y] = the number of matching pairs. | X may mean any halogen in catalog definitions. Topological distance is measured in bonds. |
| 3D atom pairs | G(X..Y) = sum r_ij over matching heavy-atom element pairs. | Distances are in Angstroms; returns zero when the requested pair is absent. |
| RDF | RDF(R,w) = f sum_(i<j) w_i w_j exp[-beta(R-r_ij)^2]. | The system uses beta = 100 A^-2 and f = 1. R and the weighting suffix are part of the descriptor ID. |
| 3D-MoRSE | Mor_s(w) = sum_(i<j) w_i w_j sin(s r_ij)/(s r_ij). | Signals 01-32 use s = 0...31 A^-1. At s=0 the sinc factor is defined as 1. |
| Matrix Wiener/average operators | Wi_M = (1/2) sum_i sum_j M_ij; WiA_M = 2 Wi_M/[A(A-1)]; AVS_M = (1/A) sum_i sum_j M_ij. | M can be adjacency, distance, Laplacian, chi, reciprocal squared distance, geometry, reciprocal geometry, geometry/distance, Coulomb, or another named matrix. |
| Matrix spectral operators | SpAbs = sum |lambda_i|; SpPos = sum_(lambda_i>0) lambda_i; SpMax = max lambda_i; SM_q = sum lambda_i^q; EE = ln(sum exp(lambda_i)). | Eigenvalues lambda_i come from the matrix named in the suffix. Normalized A and logarithmic variants apply the division or log stated in their catalog names. |
| Burden eigenvalues | Eigenvalues of a modified Burden matrix with weighted diagonal atoms, square-root bond-order off-diagonals, and 0.001 for nonbonded pairs. | High/low ranked eigenvalues are selected by the descriptor ID; missing ranks are zero-padded. Atomic-property weights are carbon normalized. |
| E-state | I_i = intrinsic atomic state; S_i = I_i + sum_(j != i)(I_i-I_j)/(d_ij+1)^2. | Atom-type descriptors report counts, sums, minima, or maxima of S_i for matching heavy-atom environments. Empty types sum to zero. |
| P_VSA and SASA partitions | P_VSA_bin = sum_i area_i * 1[property_i is in the named bin]. | The atomic area model and property bins follow the named RDKit/Mordred or catalog implementation. Bin sums are not whole-molecule property measurements. |
| Feature and fragment counts | N_pattern = sum of matches to the descriptor-specific atom, bond, functional-group, ring, or pharmacophore rule. | Rules can overlap unless a descriptor definition explicitly makes them exclusive. CATS/SHED variants additionally bin feature-pair graph distances; CATS3D bins Cartesian distances. |
| 3D weighted covariance families | C = [sum_i w_i (r_i-r_bar)(r_i-r_bar)^T] / sum_i w_i; eigenvalues and eigenvectors of C generate WHIM-style size, shape, symmetry, and directional indices. | Coordinates are centered and come from the reported ETKDGv3 geometry. WHIM, GETAWAY, WHALES, and related families then apply their catalog-specific leverage, landmark, or normalization operators. |
| Drug-likeness rules | Lipinski violations count MW>500, LogP>5, HBD>5, and HBA>10; Veber-style checks use rotatable bonds and TPSA; QED is the weighted geometric mean of property desirabilities. | These equations are screening heuristics. The precise QED desirability functions use the RDKit implementation. |
How To Reproduce A Descriptor Value
A formula alone is not enough to reproduce many descriptors. The molecular form, atom typing, weighting table, geometry, bin boundaries, and software implementation are also part of the calculation.
| Record this | Why it matters | QSAR report location |
|---|---|---|
| Canonical structure and molecular form | Salt fragments, charge, tautomer, isotope, hydrogen state, and stereochemistry can change graph and property values. | Resolved molecule and canonical SMILES. |
| Exact descriptor ID | Lag, rank, bin, matrix, atom pair, and weight suffix select the actual equation. | Descriptor ID and Catalog ID columns. |
| Calculation engine | RDKit, Mordred Community, and Formula specification can implement different descriptor families or conventions. | Mapping source/engine column and report footer. |
| Geometry method | All coordinate-based descriptors depend on the conformer and optimization method. | Geometry field in the descriptor report footer. |
| Formula notes and missing-value state | Documents carbon scaling, charge model, heavy-atom scope, distance unit, fallback, and whether the value was undefined. | Formula implementation notes and calculated/failed summary. |
What Users Should Expect From Each Lane
Each lane answers a different question. Users should avoid reading the first fast panel as the final answer when slow or EHS endpoints are relevant.
| Lane | Best for | User expectation |
|---|---|---|
| Fast QSAR | Immediate first-pass triage. | Expect a small set of high-priority endpoint cards. Use it to decide whether the molecule is worth deeper screening. |
| Fast Lane | Broader but still practical endpoint coverage. | Expect more endpoint groups and a better overall profile while slow endpoints remain deferred. |
| Slow Endpoints | Deeper endpoint completion after the fast lanes. | Expect longer runtime and additional cards that may be important for borderline compounds. |
| EHS Endpoints | Environmental and safety-specific screening. | Expect endpoints focused on environmental fate, ecotoxicity, biodegradation, bioaccumulation, and hazard-style interpretation rather than general drug-like triage. |
Advanced QSAR Views
Advanced QSAR Views are optional interpretation and optimization panels. Load them only after the main endpoint run is complete and only when you need deeper explanation, chemistry utilities, or decision support.
| Advanced view | What it opens | When the user should use it |
|---|---|---|
| Visual Analytics | Opens heavier visual interpretation tools such as endpoint profile summaries, SAR contribution hints, molecular clipping/heatmap-style views, property radar, and endpoint-focused analytics. | Use after reading the endpoint cards when you need to explain why a model may be sensitive to a structure region or when you want a visual summary for one endpoint. It is interpretive support, not experimental attribution. |
| Chemistry Utility | Opens chemistry helper panels such as logD/pH-style utility views and drug-likeness radar summaries derived from the current molecule and QSAR profile. | Use when the question is about chemical behavior behind the predictions, such as ionization/pH context, drug-likeness balance, or descriptor-driven interpretation. |
| Decision Support | Opens comparison and optimization panels, including original-vs-edited variant SMILES, variant QSAR, endpoint shift comparison, and MPO/Chemryt Score criteria tools. | Use when you want to compare a proposed analog or edited SMILES against the original molecule and tune target-product-profile style criteria. |
| Hide advanced views | Closes the currently loaded advanced panel while keeping the completed QSAR run and main endpoint tabs available. | Use when you are finished with deeper analysis, want a cleaner page, or want to reduce visual load before reading endpoint cards or switching tabs. |
How To Use Advanced Views Safely
Advanced views are designed for interpretation after the main result exists. They should not replace the endpoint cards, confidence notes, applicability-domain notes, or experimental follow-up.
| User action | Reason | Expected result |
|---|---|---|
| Run QSAR first | Advanced panels need a completed profile. During streaming or partial runs, the module keeps advanced content hidden so the user does not interpret incomplete data. | After the run completes, the Advanced QSAR Views block becomes useful for deeper reading. |
| Start with endpoint cards | The cards contain the direct model result, confidence, applicability-domain status, and skipped/error notes. | The user understands the model output before opening secondary visual or optimization tools. |
| Open one advanced view at a time | Each button loads a heavier panel with a different purpose. | The page stays focused and the user can separate visual explanation, chemistry utility, and decision support. |
| Hide advanced views before changing task | The advanced panel may be tied to the current selected endpoint, current molecule, or edited variant. | The user avoids confusing an old advanced-panel context with a new molecule, representation, or endpoint lane. |
SELFIES Token Hotspots
ChemrytIQ - SELFIES Token Hotspots is a model-agnostic perturbation view. It edits valid SELFIES tokens, decodes candidate molecules, runs selected QSAR endpoints, and reports which token changes move endpoint predictions.
| Control or output | What it means | How the user should read it |
|---|---|---|
| Run SELFIES Hotspots | Starts a SELFIES token perturbation scan for the current molecule and selected endpoint context. | Use after a QSAR run when you want to know which local token edits may improve or worsen a prediction. Expect the page to show progress while candidates are generated and scored. |
| Filtered SELFIES tokens | Default perturbation mode that focuses on practical token edits instead of every possible edit. | Use for a manageable first pass. It is faster and easier to interpret than a full token-wide scan. |
| Bioisosteric replacements | Perturbation mode for chemistry-aware replacement ideas. | Use when you want analog-like edits that may preserve medicinal chemistry intent while shifting endpoint risk. |
| Full molecule - token-wide | Broader coverage mode that scans more SELFIES token positions across the molecule. | Use when the first filtered scan is not enough. Expect more candidates, more runtime, and more review work. |
| Endpoint movement table | Shows candidate edits and how selected endpoint predictions changed. | Look for edits that consistently improve the endpoint without creating new high-risk signals. Do not act on a single favorable delta without checking validity and applicability. |
| Map | Maps a selected SELFIES hotspot back onto approximate atoms in the structure. | Use to see where the token edit is likely acting. Atom mapping uses RDKit/MCS attribution and should be treated as approximate. |
| Map heatmap | Aggregates token perturbation effects into a structure heatmap. | Blue regions summarize beneficial endpoint shifts and red regions summarize harmful shifts. Use it as a design clue, not proof of causal atom attribution. |
| Clear token map | Clears the current hotspot results and selected token/heatmap context. | Use before changing endpoint focus, rerunning with a different mode, or moving to another molecule. |
Similar Compounds With Known Data
Similar Compounds With Known Data is a ChEMBL neighbor reality check. It searches for near-neighbor compounds, computes local Morgan Tanimoto similarity, and summarizes available experimental activity records.
| Control or output | What it means | How the user should read it |
|---|---|---|
| Similar compounds with known data | Expands the known-data panel for the current molecule. | Use when a QSAR-only call is important enough to check against experimental neighbors. |
| Run ChEMBL check | Searches ChEMBL similarity neighbors and computes local RDKit.js Morgan similarity. | Use after a valid QSAR result. The check can support or weaken practical confidence depending on how close and how consistent the neighbors are. |
| Refresh ChEMBL check | Repeats the known-data lookup for the current molecule. | Use if the previous lookup failed, if source data changed, or if you want to reload neighbor evidence after a page-state change. |
| ChEMBL neighbor table | Shows nearest known-data neighbors, names or ChEMBL IDs, similarity, links, and summarized activity examples where returned. | A close neighbor with conflicting measured activity, unfavorable properties, or concerning known data should lower confidence in a QSAR-only decision. |
| Known-data source note | Explains whether ChEMBL or PubChem neighbor evidence was returned and whether structural alignment was available. | Use this note to understand limitations. If full substructure alignment is unavailable, rely on table-level evidence and Morgan similarity rather than visual alignment. |
| Hide known-data check | Collapses the neighbor-evidence panel. | Use once you have captured the evidence or when you want to return to endpoint cards without extra context on screen. |
How Hotspots And Known Data Work Together
SELFIES hotspots suggest possible edits; known-data neighbors provide experimental context. Strong decisions should consider both when available.
| Situation | Recommended reading | Decision cue |
|---|---|---|
| Hotspot suggests an improvement and close ChEMBL neighbors agree | Check whether the improved candidate stays chemically plausible and whether neighbor data supports the same endpoint direction. | This is a stronger design clue, but still needs experimental confirmation. |
| Hotspot suggests an improvement but known neighbors disagree | Treat the edit as uncertain. Review applicability-domain notes, endpoint confidence, and neighbor assay context. | Do not prioritize the edit without additional evidence. |
| Known-data neighbor has concerning measured activity | Use the neighbor as a caution even if QSAR output looks favorable. | Lower practical confidence and consider follow-up assays or safer analog directions. |
| No ChEMBL neighbors returned | Use QSAR cards, applicability-domain notes, and SELFIES hotspot trends, but note the lack of experimental neighbor support. | The result remains model-led and should be validated more carefully. |
ML Model / Computation Used
| Model or method | What it predicts | Implementation details |
|---|---|---|
| Server-side QSAR endpoint models | ADME, toxicity, physicochemical, environmental, and safety endpoint predictions. | The module submits canonical SMILES to QSAR APIs and returns endpoint cards with confidence/applicability context. The inspected package includes ONNX/joblib runtime dependencies; endpoint-specific model artifacts are managed by the QSAR server package. |
| SELFIES hotspot and analog perturbation workflow | Valid token edits and virtual analog scans for endpoint movement. | SELFIES is used as a safe molecular representation layer; canonical SMILES remain the authority sent into prediction models. |
Good Practice
QSAR output is computational decision support, not experimental proof. Confirm key safety, potency, and developability decisions experimentally.
Reference Used
This Tutorial page mirrors the ChemrytIQ reference module: ChemrytIQ-QSAR.