Modern drug discovery no longer fits neatly into the old categories.
For decades, informatics systems were built around a relatively clean separation. Small molecules belonged to cheminformatics. Proteins and nucleic acids belonged to bioinformatics. Chemists worked with atoms and bonds. Biologists worked with sequences. Each world had its own conventions, tools, file formats, databases, and mental models.
That separation worked well when therapeutic entities were either unambiguously small molecules or unambiguously natural biopolymers. Today’s discovery programs, however, are increasingly filled with hybrid modalities: synthetic peptides, cyclic peptides, modified oligonucleotides, antibody-drug conjugates, antibody-oligonucleotide conjugates, peptide-drug conjugates, chimeras, and other engineered biomolecules.
These entities are not just molecules, nor are they just sequences. They are both, and that complexity creates a difficult challenge for record keeping.
When a scientist creates a modified peptide, an oligonucleotide analog, or an ADC, a registration system is often forced to answer a question that the science itself refuses to answer cleanly:
Should this be registered as a molecule with atom-level format or sequence format?
For many modern therapeutic modalities, either answer loses something important.
A conventional small molecule is typically represented as a graph of atoms and bonds. That representation captures chemical composition, connectivity, stereochemistry, and the information needed for structure search, molecular weight calculation, property prediction, and other cheminformatics workflows.
A conventional protein or nucleic acid, on the other hand, is often represented as a sequence. For natural amino acids or nucleotides, sequence notation is compact, familiar, and efficient. A biologist does not want to draw thousands of atoms to describe a protein chain. A sequence is usually the right level of abstraction.
An example of an entity represented as a sequence vs a chemical structure
The problem begins when a therapeutic entity needs both levels of meaning.
Consider a synthetic peptide with non-natural amino acids. As a sequence, it is easy to see the order of monomers and the relationship to the corresponding natural peptide. But unless the registration system also understands the chemical structures of the modified monomers, critical chemical information may be hidden in a dictionary, an attachment, a footnote, or a scientist’s memory.
Represent the same peptide only as a fully drawn chemical structure, and the reverse problem appears. The chemistry may be present, but the biological meaning becomes hard to visualize. The amino acid composition, sequence order, residue identity, and relationship to related analogs are no longer obvious.
Modified oligonucleotides present a similar challenge. A simple sequence may capture the bases, but it may not fully express modified sugars, phosphate replacements, conjugated groups, backbone changes, duplexes, hairpins, strand breaks, or other structural features that matter to the project.
ADCs pose an even more significant problem. An antibody has chains, domains, disulfide bridges, sequence regions, and higher-order biological meaning. A payload or linker has a chemical structure. The conjugation may occur at defined sites, variable sites, or sites that are not yet fully known. Registering the ADC as only an antibody sequence loses the chemistry. Registering it as only a chemical representation loses the biology.
As a consequence, many organizations end up using workarounds.
A peptide sequence might live in one field, custom monomer definitions in a separate table, chemical drawings in an attachment, and assay results somewhere else. An oligo might be captured as a text string plus footnotes. An ADC might be represented as an antibody record, a payload record, and a spreadsheet describing how they are connected.
These workarounds may be manageable for a small number of compounds. They become fragile as programs scale.
Registration is not just a filing exercise. It is the foundation for everything that happens downstream: searching, comparing, analyzing, reporting, decision-making, collaboration, and eventually regulatory traceability.
When complex biologics are represented incompletely, the consequences show up in many ways.
Scientists may struggle to find related analogs because the system cannot search across both sequence similarity and chemical structure. Data managers may have to manually reconcile custom monomers, aliases, and attachment points. Chemists may not trust calculated properties if the full structure is not represented. Biologists may not recognize a biologically meaningful entity if it is shown only as a dense atom-and-bond diagram. Informatics teams may be forced to maintain fragile connections among multiple systems, file formats, and spreadsheets.
The root cause is usually the same: the representation has thrown away information.
A sequence without chemistry cannot fully describe a chemically modified biomolecule. A chemical structure without sequence context cannot fully describe a biomolecule analog. A note field cannot support reliable structure search. A drawing attachment cannot support calculation. A global monomer dictionary can become a bottleneck when teams are constantly inventing new building blocks.
The more innovative the modality, the more painful the compromise becomes.
Chemically aware biologics registration means that complex biomolecules can be represented at the right level of abstraction without losing the underlying chemistry.
A scientist should be able to view a peptide, oligo, antibody, or conjugate in a form that makes sense for the modality. That may be a sequence, a monomer-based sketch, an antibody glyph, a cartoon-like representation, or a full atom-and-bond structure.
These views should not be disconnected illustrations. They should be different points of view of the same registered entity.
The representation should know what the monomers are. It should know how they connect. It should know where the attachment points are. It should know which pieces are natural, which are modified, and which are custom. It should be able to expand into an all-atom structure when needed. It should support calculations and searching. It should allow biological information to sit on top of biochemical and chemical information rather than replacing it.
In other words, the system should not force a false choice between molecule and sequence. It should support both.
The benefits of chemically aware biologics registration are practical.
For scientists, it means less ambiguity. The registered entity more closely matches what was actually made, tested, and discussed.
For chemists, it means that modified monomers, linkers, payloads, stereochemistry, and attachment points are not lost in shorthand.
For biologists, it means that sequence, chain, domain, and antibody-level views remain accessible and recognizable.
For informatics teams, it means that the data can support search, calculation, comparison, analysis, and integration rather than serving only as a static record.
For project teams, it means fewer spreadsheet workarounds, fewer duplicate interpretations, and fewer disconnected fragments of truth.
This becomes especially important as portfolios diversify. Many organizations are no longer running only small-molecule programs or only biologics programs. They are working across modalities, often in the same discovery organization, with teams that need shared data infrastructure.
A platform that can handle small molecules and chemically complex biologics in one environment can reduce the friction between disciplines.
The central challenge of modern biologics registration is not just size or complexity. It is meaning.
A fully drawn structure may contain the atoms but obscure the sequence. A sequence may communicate the biology but omit the chemistry. A cartoon may be easy to read but useless for calculation. A spreadsheet may be flexible but unreliable as a scientific system of record.
Modern therapeutic entities need a representation that preserves meaning across levels.
That is the promise of chemically aware biologics registration: to describe peptides, oligos, antibodies, ADCs, and other complex biomolecules in a way that chemists, biologists, and computers can all understand.
The awkward choice between molecule and sequence is no longer good enough. Modern drug discovery needs both.
In subsequent articles, we will address specific challenges and solutions of registering oligonucleotides, peptides and ADCs as chemically aware biologics, as well as technical file format considerations.
To learn more about CDD Vault's Chemically-Aware Biologics Registration capabilities, contact us to schedule a demo.