Discover HLA Antigens Beyond the Canonical Proteome
Many antigen-discovery workflows begin with annotated protein-coding sequences. That approach is useful,
but it can miss peptides produced from alternative translation events, aberrant transcripts, retained
introns, untranslated regions, non-coding RNAs, transposable elements, endogenous retroviral sequences,
alternative reading frames, and other sequence sources outside conventional protein databases. When such
peptides are processed and loaded onto HLA molecules, they become part of the experimentally observable
immunopeptidome.
Cryptic antigen discovery extends the search space beyond canonical proteins and asks
whether these unconventional sequence sources generate HLA-presented peptides in the biological system of
interest. Our project design can combine HLA peptide enrichment and LC-MS/MS with custom or
sample-specific sequence databases, proteogenomic evidence, de novo sequencing support, HLA annotation,
normal-reference comparison, and downstream validation modules. For broad HLA ligand profiling, see our HLA Peptidomics
platform.
Non-Canonical Antigen Sources We Can Investigate
The sequence space can be tailored to the project rather than restricted to one predefined class of
cryptic peptide. Depending on the available biological data and supplier-supported workflow, discovery
modules may include the following sources.
Alternative and Upstream ORFs
Search peptides translated from uORFs, overlapping ORFs,
alternative initiation sites, and alternative reading frames that are absent from standard protein
annotations.
UTR-Derived Translation Products
Investigate peptides originating from translated 5′ or 3′
untranslated regions when transcript and sequence evidence supports an expanded search space.
Intronic and Intergenic Sources
Identify candidate HLA ligands associated with retained introns,
cryptic transcription, intergenic translation, or other non-canonical genomic regions.
lncRNA and ncRNA-Derived Peptides
Evaluate translated products from long non-coding RNAs and other
transcripts that may contribute peptides to the HLA-presented repertoire.
Repeat and Retroelement-Associated Peptides
Explore peptides linked to endogenous retroviruses, transposable
elements, LINE/LTR-related sequences, and repeat-associated transcription when relevant to the
biological system.
Aberrant Splicing and Transcript Processing
Build search candidates from intron retention, alternative splice
junctions, transcript rearrangements, or other atypical RNA-processing events.
Variant and Non-Coding Mutation Sources
Extend discovery to non-coding or unconventional sequence changes
when matched genomic or transcriptomic data are available for personalized database construction.
Custom Cryptic Antigen Hypotheses
Project-specific search spaces can incorporate additional
unconventional peptide-generation mechanisms, including selected peptide-splicing or processing
hypotheses, when appropriate evidence and validation controls can be defined.
Integrated Discovery Strategy
Cryptic antigen discovery is most informative when direct HLA presentation evidence is connected to the
sequence source that could generate each peptide. A project can therefore combine immunopeptidomics with
one or more genomic, transcriptomic, translation-aware, or computational layers. The exact combination is
selected according to the research question and available material.
HLA Immunopeptidomics
Enrich HLA-associated peptides and acquire tandem MS evidence for
sequences that are naturally presented in the sample.
Custom Sequence Database Construction
Build expanded search databases from annotated and non-canonical
ORFs, sample-specific variants, transcript structures, or other project-defined sequence sources.
RNA-Seq and Transcript Evidence
Use transcript expression, splice events, retained introns, or
tumor-associated transcription to support the plausibility and specificity of candidate peptide
sources.
Translation-Aware Evidence
Ribosome-profiling or other translation-related data can be
incorporated when available to support candidate ORFs and distinguish plausible translated regions
from sequence-only hypotheses.
De Novo and Open-Search Support
Selected projects can use de novo peptide sequencing or
complementary search strategies to broaden discovery beyond a single fixed reference database and
generate candidates for subsequent validation.
Normal-Reference and Specificity Analysis
Compare candidate sources and HLA ligands against matched controls,
public benign-tissue resources, or study-specific reference sets to distinguish condition-enriched
from broadly presented peptides.
Projects focused on a broader mutation-driven or personalized tumor antigen pipeline can also be
integrated with our Proteogenomic
Neoantigen Analysis service. When the central goal is direct MS confirmation across a wider
neoantigen space, Mass
Spectrometry Neoantigen Discovery can be incorporated into the same project plan.
Cryptic Antigen Discovery Workflow
Study Design and HLA Context
Define the biological comparison, HLA background, candidate source
classes, and available multi-omics data
HLA Peptide Enrichment and MS Acquisition
Recover naturally presented HLA ligands and acquire LC-MS/MS data for
discovery
Expanded Sequence-Space Construction
Build project-specific search space from canonical and non-canonical
sequence sources
Identification and Evidence Integration
Integrate MS evidence with HLA, genomic, transcriptomic, and
translation-aware context
Candidate Prioritization and Validation
Prioritize cryptic antigens for confirmation and downstream functional
studies
Confidence Assessment for Non-Canonical Peptides
Expanding a database increases the number and diversity of candidate peptide sequences, which can also
increase ambiguous or false-positive spectrum assignments if evidence is interpreted too loosely. For
cryptic antigen discovery, confidence should therefore be built from multiple independent evidence layers
rather than from a single peptide-spectrum match.
MS/MS Evidence Quality
Assess fragmentation support, precursor behavior, spectrum
quality, and reproducibility across technical or biological observations when available.
Canonical-Proteome Exclusion
Evaluate whether the sequence can be explained by an annotated
protein, isoform, common variant, or other conventional source before assigning a non-canonical
origin.
Genomic and Transcript Context
Map candidate sequences to the underlying genomic or transcript
region and assess whether the proposed source is supported by sample-specific sequence
information.
HLA Presentation Context
Use HLA typing, peptide motifs, binding predictions, and
enrichment design to assess whether the candidate is consistent with the measured HLA repertoire.
Specificity and Reference Screening
Compare candidates against control samples or benign-tissue
references to distinguish broadly presented non-canonical peptides from context-enriched or
tumor-associated candidates.
Orthogonal Confirmation
Higher-priority candidates can be advanced to synthetic-peptide
MS confirmation, targeted MS, peptide-HLA assays, pMHC/TCR studies, or functional T-cell testing
according to the evidence required.
Applications of Cryptic Antigen Discovery
Tumor-Specific and Tumor-Associated Antigen Discovery
Expand the antigen search beyond coding-region mutations to
reveal non-mutational or non-canonical HLA ligands that may be enriched in tumor material.
Cancer Vaccine and T-Cell Target Research
Generate experimentally observed candidate peptides for
downstream prioritization, pMHC validation, TCR screening, or functional immune studies.
Treatment-Responsive Antigen Landscapes
Investigate whether epigenetic, splicing, translation, stress, or
pathway perturbations alter the repertoire of canonical and non-canonical presented peptides.
Viral and Pathogen Antigen Research
Examine alternative ORFs or unconventional translation products
from pathogen genomes when standard annotations do not capture the full presented antigen
repertoire.
Autoimmunity and Aberrant Antigen Presentation
Explore whether unusual transcription, translation, or processing
events generate self-derived HLA ligands that warrant mechanistic or immune-recognition follow-up.
Mechanism and Target-Generation Studies
Connect antigen presentation with splicing defects, epigenetic
derepression, nonsense-mediated decay, transposable-element activation, or alternative translation
mechanisms.
Sample and Multi-Omics Study Design
The most informative design depends on the claim the project needs to support. Direct immunopeptidomics
is central when natural HLA presentation must be demonstrated, while genomic, transcriptomic, or
translation-aware datasets improve interpretation of where a cryptic peptide came from and how selectively
it is expressed.
| Input or Evidence Layer |
Role in the Project |
Typical Use |
| HLA immunopeptidomics sample |
Provides direct evidence of naturally presented HLA-bound peptides |
Core discovery layer for experimentally observed cryptic antigens |
| HLA typing |
Supports allele-aware interpretation and candidate presentation plausibility |
Useful for donor-derived, primary, or multiallelic material |
| RNA-seq |
Supports expression, aberrant splicing, intron retention, transcript structure, and condition
specificity |
Recommended when transcript-derived candidate origin is important |
| WES or WGS |
Provides sample-specific sequence variants and personalized reference information |
Useful for mutation-containing or individualized cryptic antigen searches |
| Ribosome profiling or translation-aware data |
Provides direct or indirect evidence that unconventional ORFs are translated |
Valuable for difficult ncORF or alternative-translation hypotheses |
| Matched control or benign reference |
Helps evaluate recurrence, enrichment, and off-target normal-tissue presentation |
Important when tumor or condition specificity is a prioritization criterion |
| Follow-up validation material |
Supports targeted MS, synthetic-peptide confirmation, binding, pMHC/TCR, or cellular testing
|
Used to increase evidence depth for prioritized candidates |
Representative Results
The examples below illustrate common analytical outputs for cryptic antigen discovery. Actual plots and
evidence layers depend on the search space, available multi-omics data, HLA context, and validation
strategy.
Origin of Non-Canonical HLA Peptides
Proteogenomic Evidence for a Cryptic Antigen
MS/MS Confirmation of a Candidate Peptide
Target-Context Specificity Across Reference Samples
Typical Deliverables
Deliverables are assembled according to the experimental design and may include:
- HLA-Bound Peptide Identification Table
Detected peptide sequences with MS evidence, source mapping, and project-specific annotation fields.
- Cryptic Source Annotation
Classification of candidates by alternative ORF, UTR, intronic, intergenic, lncRNA, repeat, aberrant transcript, variant, or other custom source categories.
- Genomic and Transcript Mapping
Coordinates and sequence context linking candidate peptides to their proposed genomic, transcriptomic, or translated origin.
- HLA Presentation Annotation
HLA typing context, motif or binding-prediction support, and other project-appropriate evidence for peptide-HLA compatibility.
- Multi-Omics Evidence Matrix
Integrated view of MS, RNA, DNA, translation-aware, recurrence, and reference-sample evidence when these data layers are included.
- Specificity and Prioritization Results
Candidate ranking by presentation evidence, source plausibility, context enrichment, normal-reference occurrence, recurrence, and downstream validation relevance.
- Representative Visualizations
Source-category summaries, spectrum evidence, genomic/transcript context, HLA annotations, specificity plots, and project-specific comparison figures.
- Validation Candidate List
Shortlisted peptides suitable for synthetic-peptide confirmation, targeted MS, binding assays, pMHC/TCR studies, or functional T-cell testing.
- Analytical Report and Data Package
Methods summary, quality-control notes, interpretation, data tables, visualizations, and project-specific files required for downstream analysis.
Related Antigen Discovery and Validation Options
Cryptic antigen discovery can be used as a focused service or combined with broader antigen-discovery and
validation modules. The appropriate configuration depends on whether the project begins with unexplored
sequence space, known genomic variants, predefined peptides, or candidate immune targets.
| Project Need |
Recommended Module |
Role in an Integrated Program |
| Expand discovery beyond annotated proteins |
Cryptic Antigen Discovery |
Build and interrogate non-canonical antigen space with direct presentation evidence |
| Combine multiple tumor antigen sources in one program |
Neoantigen
Discovery |
Integrate mutation-derived, non-canonical, viral, shared, or other antigen classes into a
broader discovery strategy |
| Rank candidates computationally before or after MS |
Neoantigen
Prediction & Prioritization |
Add HLA presentation, binding, processing, or immunogenicity-oriented prioritization |
| Test receptor-level recognition after candidate selection |
TCR-pMHC
Validation |
Advance selected peptide-HLA candidates into receptor-binding or functional recognition studies
|
References
For research use only. Not for use in diagnostic or therapeutic procedures.
FAQ for Cryptic Antigen Discovery
What is a cryptic or non-canonical antigen?
+
A cryptic or non-canonical antigen is an HLA-presented peptide that originates
from a sequence source or peptide-generation mechanism not captured well by standard annotated protein
databases. Examples include alternative ORFs, untranslated regions, intronic or intergenic translation,
lncRNAs, transposable elements, aberrant splicing, non-coding variants, and other unconventional
translation or processing events.
Is cryptic antigen discovery limited to cancer?
+
No. Cancer is a major application because tumors frequently alter transcription,
splicing, epigenetic regulation, and translation, but the same discovery logic can be applied to
infection, vaccine research, autoimmune mechanisms, stress responses, engineered systems, and other
biological contexts that may generate unconventional HLA ligands.
Which non-canonical sequence sources can be included?
+
Projects can be configured to investigate upstream or alternative ORFs, UTRs,
introns, intergenic regions, lncRNAs, transposable elements, endogenous retroviral sequences,
alternative splice products, non-coding variants, and other project-specific sequence hypotheses. The
final search space depends on the available data and the evidence threshold required.
Do I need RNA-seq or ribosome profiling?
+
Not for every project. HLA immunopeptidomics can identify candidate non-canonical
peptides directly, while RNA-seq strengthens transcript-origin and specificity interpretation. Ribosome
profiling or other translation-aware evidence can add support for unconventional ORFs. The most useful
combination depends on whether the goal is broad discovery, mechanistic source assignment, or
high-confidence target nomination.
Can WES or WGS be integrated with cryptic antigen discovery?
+
Yes. Sample-specific DNA sequencing can contribute mutations and personalized
sequence information, including unconventional sequence changes that fall outside standard coding exons.
These data can be integrated with RNA evidence and HLA immunopeptidomics in a project-specific
proteogenomic workflow.
How are false-positive non-canonical peptide identifications controlled?
+
Expanded databases require stricter interpretation because more candidate
sequences create more opportunities for ambiguous matches. Confidence can be strengthened by spectrum
quality, sequence uniqueness, canonical-proteome exclusion, HLA compatibility, transcript or translation
evidence, recurrence, reference-sample screening, and targeted or synthetic-peptide confirmation for
high-priority candidates.
Can de novo sequencing be used for cryptic antigen discovery?
+
Yes, as a complementary strategy. De novo sequencing can nominate peptide
sequences that are not readily recovered by a fixed database search. Candidate de novo sequences should
then be mapped back to plausible genomic or transcript sources and evaluated with additional MS and
biological evidence before being treated as high-confidence cryptic antigens.
Can both HLA class I and HLA class II cryptic antigens be investigated?
+
Yes, depending on the HLA expression, sample type, enrichment strategy, and
project goal. HLA-I is the most established context for many non-canonical tumor-antigen studies, while
HLA-II can also present peptides from unconventional protein sources. Class-specific enrichment and
bioinformatics should be planned separately when both compartments are important.
Does a non-canonical peptide automatically mean tumor-specific?
+
No. Healthy tissues can also present peptides from non-canonical genomic regions.
Tumor or condition specificity must be evaluated independently using matched controls, benign-tissue
references, expression evidence, recurrence patterns, and the biological context of the source
transcript or translated region.
Does MS detection prove that a cryptic antigen is immunogenic?
+
No. MS-based immunopeptidomics provides evidence that a peptide is present in the
enriched HLA ligand pool. Immunogenicity additionally depends on peptide-HLA stability, abundance,
T-cell repertoire, receptor recognition, and biological context. Functional immune assays are required
when the project needs evidence of T-cell recognition or activity.
How can prioritized cryptic antigens be validated?
+
Validation can be staged according to the claim required. Options include
synthetic-peptide spectrum matching, targeted MS, peptide-HLA binding assays, pMHC reagent studies, TCR
screening, and cellular T-cell assays. Higher-priority development programs can combine several evidence
layers rather than relying on a single validation method.
Can cryptic antigen discovery be combined with conventional neoantigen analysis?
+
Yes. Mutation-derived neoantigens, non-canonical antigens, viral peptides, shared
tumor antigens, and other candidate classes can be analyzed within an integrated antigen-discovery
program. This is often useful when the goal is to maximize target space rather than restrict the project
to one antigen origin.