Immunopeptidomics Experimental Design for HLA Ligand and Neoantigen Discovery
Immunopeptidomics fails most often at the junction between a scarce sample and an ambitious claim. A tumor specimen may be sufficient to identify a broad HLA ligand repertoire but insufficient for reproducible detection of a specific mutant peptide. A peptide-spectrum match may pass a global 1% false discovery rate while remaining unreliable within the much smaller neoantigen subset. A sequence may be predicted to bind the patient's HLA allele without having been demonstrated on the cell surface.
These are not interchangeable limitations. Sample input determines how much HLA-peptide material reaches the instrument. Immunoaffinity controls determine whether the observed peptides are associated with HLA capture rather than beads, antibodies, or proteolytic background. Search-space and FDR design determine how many sequence assignments are credible. Targeted MS and functional assays address different layers of validation.
The study should therefore begin with the intended evidence claim: repertoire profiling, differential presentation, direct detection of a candidate antigen, absolute or relative copy-number estimation, or demonstration of T-cell recognition. The same discovery workflow cannot support all of these claims at the same evidentiary level.
Evidence Categories in Immunopeptidomics: MS detection supports physical presentation of a peptide in an HLA-enriched sample. HLA-binding prediction supports biological plausibility. A T-cell assay addresses recognition and function. None of these measurements can substitute automatically for the others.
Study Design for HLA Ligand Profiling and Neoantigen Discovery
An HLA ligandome survey asks which peptides can be detected after immunoaffinity enrichment. A quantitative perturbation study asks which presented peptides change between conditions. A neoantigen discovery project asks whether a variant sequence is directly presented by a specified HLA context. A therapeutic target program may also require selectivity across normal tissues, abundance, reproducibility, and functional recognition.
Each question changes the design.
| Intended claim | Minimum evidence structure | Common overinterpretation to avoid |
|---|---|---|
| Broad HLA-I or HLA-II repertoire | HLA enrichment, appropriate controls, peptide-level identification, allele-aware annotation, and transparent FDR | Treating every identified peptide as a unique therapeutic target |
| Differential presentation after treatment | Biological replication, balanced processing, quantitative normalization, and condition-level statistics | Interpreting presence/absence in one run as regulated presentation |
| Mutant or pathogen-derived peptide detection | Sample-specific database, subset-aware error control, spectrum review, and targeted or synthetic-peptide confirmation | Relying on a permissive custom-database hit alone |
| HLA restriction assignment | High-resolution HLA typing plus motif or binding analysis, monoallelic evidence, or orthogonal assignment | Treating a NetMHCpan prediction as direct assignment proof |
| Immunogenic antigen | Credible presentation evidence plus a fit-for-purpose T-cell assay | Assuming that abundance or HLA binding guarantees immunogenicity |
The project should reserve material for the highest-priority downstream evidence. If all HLA-enriched peptide is consumed in a discovery run, an HLA-matched primary specimen may not be available for targeted confirmation. When material cannot be regenerated, allocating an aliquot for a secondary screen can be more valuable than marginally increasing discovery depth.
Sample Requirements for HLA-I and HLA-II Immunopeptidomics
HLA-bound peptides are low-abundance analytes recovered through several loss-prone steps: tissue disruption or cell lysis, solubilization of peptide–HLA complexes, immunoaffinity capture, washing, acid elution, peptide cleanup, concentration, and LC-MS injection. Starting material therefore affects both the number of ligands detected and the probability of observing a specific low-copy target.
Traditional tissue immunopeptidomics workflows have often used 500–1,000 mg of wet tissue or up to approximately one billion cells for deep neoantigen discovery. More sensitive integrated workflows have reported HLA-I and HLA-II analyses from about 50 mg of tissue, but reduced input should be understood as a workflow-specific feasibility achievement, not a universal guarantee of equal depth [1]. Other recent studies have used approximately 60 mg of tissue or 10⁸ cultured cells per enrichment [2].
Sample Input by Specimen Type
| Sample context | Practical planning range reported across workflows | Factors that increase input need | Interpretation at the lower end |
|---|---|---|---|
| Established cell line with high HLA-I expression | Often 10⁷–10⁸ cells for optimized studies; deeper discovery may use substantially more | Weak HLA expression, multiple alleles, class-II analysis, treatment-induced cell loss, fractionation | Feasibility depends strongly on HLA density and recovery; nondetection is not proof of absence |
| Primary cells or patient-derived cultures | Commonly closer to 10⁸ cells where available | Donor heterogeneity, limited viability, variable HLA induction, mixed cell populations | Pilot enrichment and HLA-expression assessment are especially important |
| Fresh-frozen tumor tissue | Approximately 50–500 mg in optimized or conventional workflows; some deep studies use more | Low tumor content, necrosis, stromal admixture, HLA-II objective, low-abundance neoantigens | Lower inputs may support repertoire profiling but reduce confidence for rare target nondetection |
| Small biopsy or archival tissue | Workflow-specific low-input feasibility | Ischemia, fixation, peptide loss, limited HLA material, inability to repeat extraction | Prioritize the primary claim and reserve material only when analytically feasible |
These ranges should be reported as preferred, feasible, and elevated-risk input—not as a single minimum. A laboratory's validated recovery, antibody format, LC-MS sensitivity, fractionation strategy, and expected peptide yield matter more than a generic mass threshold.
The starting amount should also be linked to HLA expression. Flow cytometry, immunohistochemistry, or another fit-for-purpose assessment can reveal whether the sample contains sufficient HLA-I or HLA-II-bearing cells. Tumor percentage, necrosis, immune infiltration, and tissue composition should be documented because wet mass alone does not describe the amount of recoverable peptide–HLA complex.
Pre-Analytical Variables in Immunopeptidomics
Cold ischemia, freeze-thaw history, proteolysis, cell dissociation, interferon exposure, culture medium, serum, drug treatment, and harvest timing can all alter the peptides recovered from HLA molecules. Samples compared biologically should be collected and processed under matched conditions. Protease inhibitors and cold handling protect complexes after harvest, but they do not reverse biological changes that occurred before stabilization.
For treatment studies, the harvest time must reflect the expected kinetics of protein turnover, antigen processing, HLA loading, and surface presentation. A change in source-protein abundance does not necessarily produce an immediate or proportional change in the corresponding HLA peptide.
An immunopeptidome profiling feasibility review should therefore consider HLA class, sample type, input, tissue quality, intended target class, and the amount that must remain for HLA typing, DNA/RNA sequencing, total proteome analysis, or targeted confirmation.
Figure 1. Input-planning framework for immunopeptidomics based on sample type, HLA expression, analytical objective, and acceptable nondetection risk.
HLA Typing and Customized Databases for Neoantigen Discovery
High-resolution HLA typing is not optional metadata when the interpretation depends on allele assignment. The typing method, nomenclature, resolution, and sample identity should be recorded according to accepted HLA conventions. For mixed or clinical samples, identity checks should connect the HLA typing, genomic data, transcriptomic data, and immunopeptidomics specimen.
MIAIPE reporting guidance recommends documenting the biological source, tissue or cell type, HLA/MHC allotypes, typing method, starting material, replicates, isolation conditions, LC-MS method, database, search parameters, FDR strategy, and quantification procedure [3]. These fields are not administrative details: they determine whether another team can evaluate the credibility of the ligand assignments.
Reference Proteome and Customized Neoantigen Databases
A standard-proteome HLA ligand search can use a curated reference proteome with relevant isoforms and contaminants. A neoantigen or noncanonical search may additionally include sample-specific somatic variants, fusions, alternative open reading frames, pathogen sequences, or other translated events supported by genomic or transcriptomic evidence.
Every added sequence expands the search space. Including all theoretical variants from an unfiltered pipeline increases the chance that an unrelated spectrum will be assigned to a biologically attractive sequence. Customized databases should therefore be traceable to variant calls, expression evidence, reading-frame logic, and sample identity. Separate search or FDR strata may be required for rare candidate classes.
Companion total-proteome data can support source-protein context but should not be used as a strict requirement for HLA presentation. HLA-I and HLA-II pathways sample proteins through different degradation and trafficking routes; some presented peptides arise from proteins not detected in a conventional tryptic proteome [1]. RNA expression and source-protein abundance are contextual evidence, not proof of presentation.
Quality Controls for HLA Immunoaffinity Purification
The immunoaffinity step concentrates HLA complexes but also creates opportunities for non-specific adsorption, antibody-derived peptides, proteolytic fragments, and environmental contaminants. Controls should be selected to answer the likely failure mode rather than added as a ritual.
Process Blank
A process blank contains buffers, resin, cleanup materials, and handling steps without biological lysate. It identifies reagent and environmental contaminants introduced after sample receipt. It is particularly useful for keratins, antibody fragments, plastics-related background, and persistent LC-MS carryover.
Blank-Bead or Mock-IP Control
Blank beads processed with representative lysate reveal peptides and proteins that bind the solid support or workflow independently of the HLA-capture antibody. Some published workflows remove peptides observed in blank-bead negative-control IPs before final HLA ligand reporting [1]. The control should be processed at a scale that produces interpretable background without consuming material needed for the biological comparison.
Irrelevant Antibody or Isotype Control
An isotype-matched or irrelevant antibody can assess non-specific interactions attributable to the antibody and coupling chemistry. It is not universally mandatory. Its value depends on the capture reagent, resin, antibody amount, species, coupling method, and availability of a relevant control antibody. A poorly matched isotype can create a different background rather than a realistic negative model.
Biological Negative Control
An HLA-null or beta-2-microglobulin-deficient model, an HLA-mismatched sample, an uninfected control, or a wild-type sequence control can provide strong evidence for particular claims. The correct negative depends on the target. For a pathogen-derived peptide, uninfected cells may be appropriate. For a mutation-derived peptide, a matched wild-type or variant-negative sample helps test sequence specificity.
Recovery and System-Suitability Controls
Defined peptide–MHC complexes, synthetic peptides, retention-time standards, or reference cell material can monitor enrichment recovery and LC-MS performance. The point at which a standard is added determines what it corrects. A peptide spiked after elution cannot measure immunoprecipitation recovery; a peptide–MHC standard added before capture can interrogate more of the workflow.
| Control | Failure mode addressed | What it cannot establish alone |
|---|---|---|
| Process blank | Reagent, handling, cleanup, and LC-MS background | Non-specific binding from a complex biological lysate |
| Blank-bead or mock IP | Resin-associated and non-antibody capture background | Allele specificity or biological presentation of a candidate |
| Isotype or irrelevant antibody | Antibody- and coupling-dependent nonspecific enrichment | A universal background for every capture antibody |
| HLA-null, mismatched, uninfected, or variant-negative sample | Target-specific biological background | General recovery and instrument stability |
| Exogenous pMHC or peptide standard | Recovery, retention, and targeted quantitative performance, depending on spike point | Endogenous immunogenicity |
Figure 2. Control architecture for separating true HLA-associated peptide evidence from process, resin, antibody, and biological background.
Non-Tryptic Database Search and FDR Control in Immunopeptidomics
Conventional bottom-up proteomics benefits from enzymatic specificity: trypsin restricts the candidate sequences assigned to each spectrum. HLA peptides are naturally processed and generally searched without tryptic specificity. The theoretical search space is therefore much larger, with many sequence candidates sharing similar precursor masses. Customized neoantigen and noncanonical databases expand it further.
Immunopeptidomics is consequently vulnerable to attractive false positives. A low global FDR can coexist with a high error rate in a small candidate subset if that subset occupies a much larger fraction of the database than it does in the true ligandome [4–6].
PSM-, Peptide-, and Subset-Level FDR
The analysis plan should distinguish:
- PSM-level FDR: the estimated error among spectrum-to-sequence assignments;
- peptide-level FDR: the estimated error among unique peptide sequences after combining PSM evidence;
- subset-specific FDR: the estimated error within a category such as mutant, pathogen-derived, noncanonical, spliced, or PTM-containing peptides;
- candidate-level confirmation: direct review and orthogonal validation for the short list used to support therapeutic or mechanistic claims.
The decoy strategy, peptide-length range, allowed charge states, modifications, database composition, search engine, rescoring method, and grouping of HLA-I versus HLA-II data should be reported. FDR estimated for a large set of canonical self-peptides should not be assumed to transfer unchanged to a handful of variant candidates.
Search-Space Control and Spectrum Rescoring
Database construction is part of error control. Include sample-supported sequences and remove impossible or duplicate entries where justified. Search canonical and expanded sequence classes in a way that permits transparent assessment of each class. Retention-time and fragmentation prediction can improve scoring, but these models are not a substitute for a defensible target-decoy framework or empirical validation.
Deep-learning spectral prediction has improved identification of non-tryptic HLA peptides and exposed incorrect assignments in previously proposed peptide classes [7]. This is useful evidence that advanced rescoring can improve sensitivity and specificity. It is also a warning that an algorithmically appealing sequence should not bypass candidate-level confirmation.
For studies that require specialized search and quantitative workflows, DIA data analysis or other immunopeptidomics-specific processing must retain peptide-level evidence, database provenance, and category-aware error estimates. A protein-centric report alone is insufficient because the therapeutic entity is the peptide–HLA complex.
HLA Binding Prediction and MS-Detected Peptide Presentation
NetMHCpan and related tools can estimate whether a peptide is compatible with one or more HLA alleles. Motif deconvolution can help assign peptides in multi-allelic samples. These tools are valuable for annotation, quality assessment, and candidate prioritization, but the predicted binding score is not physical evidence that the peptide was presented in the sample.
Prediction also has uneven performance across alleles and peptide classes. Well-represented alleles benefit from richer training data, whereas rare alleles and noncanonical peptides may be less certain. Filtering every discovery result through a binding predictor can remove genuine ligands that fall outside learned motifs; accepting every predicted binder can retain false sequence assignments.
The strongest interpretation uses concordant evidence:
- credible MS/MS sequence assignment;
- retention-time and fragmentation behavior consistent with the sequence;
- HLA allele and motif compatibility;
- recurrence or quantitative consistency across biological replicates where relevant;
- absence or strong reduction in appropriate negative controls;
- targeted confirmation for high-value candidates.
Validation of HLA Ligands and Neoantigens
Candidate validation should escalate according to the consequence of being wrong.
MS/MS Spectrum and Retention-Time Validation
Review fragment-ion coverage, mass accuracy, co-isolation, chromatographic peak shape, and competing sequence assignments. Predicted spectra and retention times can identify inconsistencies, but agreement with a model remains computational evidence.
Synthetic Peptide Validation
Analyze a stable-isotope-labeled synthetic peptide with the biological sample where possible. Coelution and matching fragment-ion patterns reduce ambiguity more effectively than comparing spectra acquired in separate runs. Take precautions against carryover from high-concentration synthetic standards.
Targeted PRM and SureQuant Confirmation
PRM or internal-standard-triggered methods such as SureQuant can improve sensitivity for a defined peptide list and support relative or absolute quantification. SureQuant MHC uses heavy peptide standards to trigger sensitive acquisition of the endogenous target and can incorporate peptide–MHC standards for recovery or calibration [8]. The method confirms the selected targets; it does not rescue an unsupported sequence or demonstrate T-cell activity.
Functional Validation of Candidate Antigens
Detection of a peptide–HLA complex does not guarantee immunogenicity. Depending on the program, the next evidence may include HLA stabilization, peptide–MHC multimer binding, T-cell activation, cytokine release, cytotoxicity, or target-cell recognition. Conversely, T-cell reactivity to exogenously loaded peptide does not prove that the peptide is naturally processed and presented at the relevant density.
An integrated program can use discovery proteomics for source-protein context and targeted proteomics for focused peptide confirmation, while keeping physical presentation and immune function as separate claims.
Figure 3. Evidence ladder from untargeted HLA peptide discovery through synthetic-standard confirmation, targeted quantification, and functional immune testing.
Evidence Requirements by Immunopeptidomics Application
| Project type | Discovery design | Confirmation needed before a strong claim | Key limitation to state |
|---|---|---|---|
| Cell-line perturbation | Independent cultures, matched treatment timing, HLA typing, mock-IP background, quantitative comparison | Replication plus targeted confirmation for selected regulated ligands | Cell-line presentation may not represent patient tissue |
| Patient-tumor neoantigen discovery | Tumor-rich tissue, matched DNA/RNA, high-resolution HLA typing, custom database, category-specific FDR | Synthetic standard and targeted MS; functional testing for immunogenicity | Nondetection may reflect input or sensitivity rather than true absence |
| Pathogen-derived ligand discovery | Infected and uninfected controls, matched HLA background, pathogen-aware database | Targeted confirmation and infection-specific biological control | Rare pathogen peptides can have a higher subset error rate than the global ligandome |
| Normal-tissue selectivity assessment | Relevant HLA alleles, tissue metadata, comparable processing, sufficient cohort representation | Cross-tissue targeted measurement and orthogonal expression context | Absence in a limited tissue set is not proof of universal tumor specificity |
| Quantitative pMHC assay development | Discovery-supported target list and proteotypic peptide behavior | Stable-isotope standards, calibration, recovery, precision, selectivity, and matrix assessment | Peptide copies per cell do not alone predict T-cell response |
Immunopeptidomics Reporting Requirements
The final ligand list should remain connected to the evidence needed for the next decision. Candidate-level records should retain the peptide sequence, charge, precursor mass error, retention time, spectrum identifier, fragment-ion evidence, identification score, PSM- and peptide-level q-values, database category, source protein or genomic event, sample-specific HLA type, predicted allele assignment, control-sample behavior, and validation status.
The study-level package should document specimen provenance, tissue composition, storage, replicate structure, antibody and resin, coupling chemistry, lysate-to-antibody ratio, wash and elution conditions, cleanup, fractionation, LC-MS method, database version, decoy strategy, modification list, peptide-length filters, quantitative normalization, missing-value handling, and public data accession where appropriate. For variant and other rare candidate classes, the report should also preserve the customized-database provenance and the subset-specific error-control procedure.
Candidate Evidence Levels
- Discovery-supported HLA ligand: credible sequence assignment in an HLA-enriched sample, with the relevant controls and identification statistics retained.
- Orthogonally confirmed presented peptide: agreement with a synthetic standard or targeted MS measurement under documented acceptance criteria.
- Quantitatively characterized target: a fit-for-purpose assay with appropriate standards, calibration, recovery, precision, and matrix assessment.
- Functionally supported antigen: presentation evidence paired with a defined T-cell recognition or activity assay in the relevant HLA context.
This tiered handoff prevents physical presentation, quantitative abundance, allele assignment, and immunogenicity from being collapsed into a single label. It also allows downstream vaccine, TCR, or biomarker teams to see which evidence has been generated and which experiments remain necessary. For sample-limited studies, an immunopeptidome profiling service feasibility assessment can align input allocation, HLA typing, discovery acquisition, and targeted confirmation before irreplaceable material is consumed.
References:
- Abelin JG, Bergstrom EJ, Rivera KD, et al. Workflow enabling deepscale immunopeptidome, proteome, ubiquitylome, phosphoproteome, and acetylome analyses of sample-limited tissues. Nature Communications. 2023;14:1851. DOI: 10.1038/s41467-023-37547-0.
- Shapiro IE, Huber F, Michaux J, Bassani-Sternberg M. Sensitive neoantigen discovery by real-time mutanome-guided immunopeptidomics. Nature Communications. 2025;16:7269. DOI: 10.1038/s41467-025-62647-4.
- Lill JR, van Veelen PA, Tenzer S, et al. Minimal Information About an Immuno-Peptidomics Experiment (MIAIPE). Proteomics. 2018;18(12):1800110. DOI: 10.1002/pmic.201800110.
- Fritsche J, Kowalewski DJ, Backert L, et al. Pitfalls in HLA Ligandomics—How to Catch a Li(e)gand. Molecular & Cellular Proteomics. 2021;20:100110. DOI: 10.1016/j.mcpro.2021.100110.
- Nesvizhskii AI. Proteogenomics: concepts, applications and computational strategies. Nature Methods. 2014;11:1114–1125.
- Leddy O, Cui Y, Ahn R, et al. Validation and quantification of peptide antigens presented on MHCs using SureQuant. Nature Protocols. 2025;20:1196–1222. DOI: 10.1038/s41596-024-01076-x.
- Wilhelm M, et al. Deep learning boosts sensitivity of mass spectrometry-based immunopeptidomics. Nature Communications. 2021;12:3346. DOI: 10.1038/s41467-021-23713-9.
- Stopfer LE, et al. Absolute quantification of tumor antigens using embedded MHC-I isotopologue calibrants. Proceedings of the National Academy of Sciences. 2021;118:e2111173118. DOI: 10.1073/pnas.2111173118.
4D Proteomics with Data-Independent Acquisition (DIA)