Meta Intent: A practical guide for selecting, building, and validating protein sequence databases for LC-MS/MS studies in non-model organisms, with emphasis on controlling search-space inflation and producing defensible biological evidence.
Proteomics can be technically successful and still answer the wrong biological question. A tissue digest may yield high-quality tandem mass spectra, stable chromatography, and a well-controlled instrument run—yet the final report can contain few useful protein identifications if the sequence database does not represent the organism, tissue, strain, or microbial community that produced the peptides.
That distinction matters most outside well-curated model systems. In human, mouse, yeast, or commonly studied microbial strains, researchers usually begin with a maintained reference proteome and a relatively predictable annotation landscape. In a non-model organism, the reference may be incomplete, derived from a distant relative, fragmented across de novo transcript contigs, or entirely absent. The database is therefore not a passive lookup table. It is a design choice that determines what the mass-spectrometry experiment can detect, how false matches are controlled, and how confidently a peptide can be assigned to a protein or biological pathway.
The productive question is not, "Which database contains the most sequences?" It is, "Which sequence space is sufficiently complete for this biological question, while remaining constrained enough to support reliable peptide-spectrum matching?" A defensible project connects the biological sample to the most relevant source of sequence evidence, expands the search space only when needed, and validates the novel portion of the result separately from familiar identifications.
This guide provides that framework. It is designed for investigators studying orphan species, agricultural and aquatic organisms, wild populations, host-associated samples, unusual cell systems, and complex environmental communities. The same principles also apply when an existing reference proteome is present but does not adequately capture tissue-specific isoforms, strain variation, non-canonical translation products, or proteins contributed by associated organisms.
Figure 1. Decision routes for non-model organism proteomics should begin with the available sequence evidence—not with a default search engine or database download.
The Reference Proteome Is Often the First Experimental Bottleneck
In routine bottom-up proteomics, proteins are enzymatically digested, peptide precursors are measured by LC-MS/MS, and fragmentation spectra are compared with theoretical spectra derived from a protein sequence database. The process feels computational, but every identification is conditional: a peptide can only be reported if an appropriate candidate sequence was included in the search space.
For a poorly characterized species, several biological realities complicate that assumption:
- The closest public proteome may represent a related species rather than the sample under study.
- Gene models may be incomplete, particularly in lineage-specific regions, long genes, alternative isoforms, and recently duplicated paralogs.
- A sample may include more than one biological source, such as host tissue plus microbiota, parasite material, diet-derived proteins, or environmental contaminants.
- The transcriptome supporting a database may come from a different tissue, developmental stage, sex, season, or experimental condition.
- A broadly translated genome can generate an enormous number of sequences that are technically searchable but biologically implausible.
These problems are not solved by simply appending every available FASTA file. A large database can reduce identification sensitivity because more candidate sequences compete for each spectrum. It can also complicate false discovery rate (FDR) estimation and protein inference. The project must therefore separate two goals that are often conflated: increasing the chance that a true sequence is present, and increasing the chance that a reported match is correct.
A useful starting point is to classify the project by its source of sequence evidence.
| Starting condition | Most defensible first database route | Key limitation to acknowledge |
|---|---|---|
| A high-quality proteome exists for the same species | Curated reference proteome with project-specific additions | May not contain isoforms, strain variants, or unannotated proteins |
| A close relative is well annotated | Homology-guided database | Peptide assignments can lose species and paralog specificity |
| Matched tissue RNA is obtainable | RNA-seq-derived custom database | Transcript presence is not proof of protein expression |
| A draft genome or metagenome exists | Predicted proteome or carefully filtered translation database | Gene prediction and search-space inflation require rigorous control |
| No usable sequence resource is available | De novo peptide sequencing plus homology-supported interpretation | Sequence calls require stringent downstream validation |
The table is a decision aid, not a ranking. A close-relative database may be entirely appropriate for a conserved enzyme survey, whereas it is a weak foundation for a species-specific isoform or evolutionary adaptation study. Similarly, a matched transcriptome may be essential for a tissue-resolved discovery project but unnecessary when the goal is to quantify a small set of well-conserved proteins.
Before database construction begins, define the inferential unit. Is the study trying to establish peptide evidence, protein-group abundance, a full-length proteoform, a taxon-resolved function, or a lineage-specific candidate sequence? The more specific the intended conclusion, the more direct and independent the supporting evidence must be.
That early definition also changes the analytical plan. A broad discovery project may reasonably begin with whole-proteome shotgun protein identification, while a project centered on sequence gaps should reserve space for protein sequencing by mass spectrometry and subsequent targeted confirmation. In either case, sample provenance, extraction conditions, and tissue identity must be recorded before sequence information is added. A database cannot correct a sample mix-up or compensate for proteins that were never efficiently extracted.
Route 1: Use a Close-Relative Database When Conservation Supports the Question
Homology-guided searching is often the fastest entry point for a non-model species. Investigators select one or more annotated relatives, search the spectra against their protein sequences, and interpret identified peptides through orthology or protein-family relationships. This route can work well when the study asks about conserved pathways, abundant structural proteins, broadly preserved metabolic enzymes, or high-level functional shifts.
The strengths are practical. Public reference proteomes are already formatted, annotated, and linked to established functional resources. Search engines can operate within a manageable sequence space, and the result can produce a rapid first view of proteome composition. For exploratory work, the approach can reveal whether sample processing is suitable for protein identification, whether protein yield supports the intended depth, and which functional areas merit deeper sequencing support.
Its central limitation is specificity. A peptide that matches a protein from a related species may provide valid family-level evidence but not necessarily prove that the exact ortholog, isoform, or paralog is present in the sampled organism. This is especially important in lineages with gene-family expansions, rapid adaptive evolution, whole-genome duplication, or weakly resolved taxonomy.
A close-relative search is therefore strongest when the report is framed at the level actually supported by the data. "Peptides support the presence of a cytochrome P450 family member" may be appropriate. "The sample expresses this exact species-specific P450 isoform" may not be justified if all identified peptides are shared among homologs. Likewise, a pathway-level shift may be interpretable when functional mapping is transparent, while a claim of lineage-specific adaptation needs direct sequence and validation evidence.
The practical remedy is not to abandon homology searching. It is to make the database and the claims proportionate to one another. Start with the closest available taxa, record their evolutionary distance and annotation quality, and retain accession-level provenance for every sequence. Where several related proteomes are combined, remove exact duplicate proteins but preserve source annotations so that shared peptides do not appear more species-specific than they are.
A carefully constructed close-relative database may also include common contaminants, experimentally relevant host proteins, and a restricted set of sequences expected from associated organisms. What it should not include by default is every taxonomically adjacent proteome available in a public repository. Excessive taxonomic breadth inflates the search space and can create attractive but weak species assignments.
Figure 2. A close-relative database and a sample-matched RNA-seq database can both support valid discovery, but they encode different types of biological evidence and uncertainty.
When a Close-Relative Database Is a Good First Choice
Choose this route first when most of the following conditions are true:
- The question concerns conserved pathways, proteins, or protein families rather than exact isoforms.
- A reasonably close relative has a curated and well-annotated proteome.
- The sample is limited, degraded, archived, or otherwise unsuitable for matched RNA collection.
- The project needs a rapid feasibility assessment before committing to sequencing-based database construction.
- The final interpretation can transparently remain at the ortholog, protein-group, or pathway level.
It is less suitable when the objective depends on species-specific peptides, recent duplication events, poorly conserved secreted proteins, unusual toxins, novel short proteins, or exact host-versus-symbiont attribution. Those are the situations in which direct sequence evidence becomes more valuable.
Route 2: Build a Sample-Matched RNA-Seq-Derived Protein Database
When fresh or appropriately preserved material is available, a matched transcriptome is often the most informative way to define the protein search space. The RNA does not need to come from the same aliquot as the proteomics sample, but it should represent the same organism, tissue context, developmental state, and experimental condition as closely as practical.
The workflow is conceptually simple:
- Collect or select biologically matched material.
- Generate RNA-seq data and assess read quality, contamination, and sample identity.
- Assemble transcripts against a reference or de novo, depending on genome availability.
- Predict coding sequences and translate plausible open reading frames.
- Remove exact redundancy, flag short or low-support predictions, and append appropriate contaminant sequences.
- Search LC-MS/MS data against this versioned custom database.
- Validate the candidate sequences that are absent from established references.
The difficulty lies in the word "plausible." A de novo transcriptome is not a direct protein catalog. Assemblies can fragment long transcripts, collapse related paralogs, preserve incomplete coding regions, or create chimeric contigs. Short-read assemblies are particularly vulnerable when alternative isoforms and repeat-rich gene families are biologically important. For that reason, the database should preserve a clear link from each predicted protein back to its transcript or contig, sequence source, sample, and filtering status.
Matched RNA-seq also changes what is possible analytically. Instead of asking whether a spectrum matches a distant canonical protein, the search can ask whether it matches a coding sequence observed in the actual biological context. This is particularly valuable for unusual tissues, seasonal phenotypes, non-standard developmental stages, and environmentally responsive systems in which a public reference may capture only a small fraction of expressed proteins.
The approach works best as part of an integrated design rather than as an add-on after the mass-spectrometry run. Sample preparation should be coordinated across the nucleic-acid and proteomics arms. Tissue dissection, storage conditions, lysis chemistry, and abundance range all influence whether transcript and peptide evidence can later be compared meaningfully. Robust protein sample preparation is therefore not merely a wet-lab prelude; it protects the link between the biological material and the database that will be used to interpret it.
What a Useful Custom Database Should Contain
A project-specific FASTA file should be smaller and more traceable than a naive translation of every assembled contig. At minimum, it should distinguish predicted proteins supported by matched transcripts; sequences inherited from a same-species or close-relative reference; candidate novel or alternative sequences; likely contaminants and laboratory background proteins; and host, microbiome, parasite, or dietary components when the sample source requires them.
Each entry should carry a stable identifier and source tag. That simple discipline enables later class-specific review. Known reference proteins, transcript-supported additions, and highly speculative sequences should not be treated as one homogeneous category when calculating confidence or interpreting biological significance.
The custom database must also be versioned. A result from "database_final.fasta" is difficult to reproduce six months later. A result from "SpeciesX_liver_RNAseq_v1.2_filteredORF_2026-08-29.fasta," accompanied by the filtering logic and source metadata, is auditable. This is especially important when the proteomics data will be reanalyzed as assemblies improve.
The key insight is that transcript-supported sequences narrow uncertainty without eliminating it. RNA evidence increases the biological plausibility of a candidate sequence, but peptide evidence is still required to show that the corresponding protein was observed in the proteomics experiment. A mature workflow uses each data layer for what it can establish, rather than allowing one omics layer to overclaim for another.
Route 3: Use Genome-Derived Translation Databases as Discovery Space, Not as a Default Endpoint
A draft genome, metagenome, or six-frame translation can expose protein-coding possibilities that a reference proteome misses. It is particularly useful when the biological question concerns novel open reading frames, strain-specific sequence variants, incomplete annotation, or organisms for which RNA is unavailable. The attraction is obvious: translating genomic sequence makes almost every possible coding region searchable.
That same feature makes the approach statistically dangerous. Most sequences generated by an unrestricted six-frame translation are not translated proteins in the sampled biology. They may be short accidental open reading frames, incorrect frame choices, untranslated regions, partial fragments, or redundant variants of the same locus. If those sequences are placed directly beside a conventional reference proteome, they expand the candidate pool without contributing equivalent biological prior evidence.
Search-space inflation has two practical consequences. First, peptide-spectrum matches must compete against many more theoretical candidates, which can reduce sensitivity at a fixed nominal FDR. Second, conventional target-decoy assumptions can become less reliable when a very large fraction of target entries are implausible sequences. A match to a speculative translation may be counted as a target hit even when it is functionally similar to a random false assignment.
The preferred response is staged analysis. Run an initial broad search designed to discover plausible regions of the expanded sequence space. Use that result, together with genomic context, transcript support, tryptic plausibility, peptide uniqueness, and recurrence across replicates, to build a reduced candidate database. Then re-search the spectra against the compact, evidence-filtered database and control FDR separately for conventional and novel sequence classes.
Figure 3. Broad discovery searches should be followed by evidence-based database reduction and focused re-searching, not interpreted as a single final identification pass.
Practical Filters Before and After Translation
Database reduction should start before the MS search whenever supporting information exists. Gene-prediction models, transcript evidence, long-read isoforms, ribosome profiling, homology, coding-potential scores, and minimum peptide-length rules can all remove sequences unlikely to produce observable tryptic peptides. These filters should be documented rather than silently applied, because they determine which biological possibilities were excluded.
After the initial search, the second-stage candidate list should be based on more than a single peptide-spectrum match. Useful criteria include peptide uniqueness within the complete database, fragment-ion coverage, retention-time plausibility, independent detection in technical or biological replicates, and absence of a better explanation among modified known peptides. The purpose is not to force every candidate through a rigid universal cutoff. It is to require increasingly strong evidence as the biological novelty of a claim increases.
For projects that must investigate unannotated proteins or difficult peptide classes, de novo peptide and protein sequencing can provide an independent candidate-generation route. It should be integrated with database evidence, however, rather than used to bypass statistical review.
Route 4: Use De Novo Peptide Sequencing to Generate Candidates, Then Test Them
De novo peptide sequencing infers an amino-acid sequence directly from tandem mass-spectrometry fragmentation patterns without requiring a matching sequence to be present in a database. This makes it attractive for non-model organisms, natural products, unusual venom or secretion studies, poorly annotated microbiomes, and projects in which a database search leaves a substantial fraction of high-quality spectra unexplained.
Its role is best understood as candidate discovery. A predicted sequence can reveal a peptide that is absent from the current database, direct a homology search, identify a missed open reading frame, or suggest that a transcriptome assembly should be revisited. It does not automatically establish the existence of a full-length protein, a specific isoform, or an exact taxonomic origin.
This distinction is increasingly important as deep-learning methods improve de novo reconstruction. Better prediction scores can make sequence hypotheses easier to obtain, but confidence scores alone are not independent biological validation. Isobaric residues such as leucine and isoleucine remain intrinsically difficult to distinguish by conventional MS/MS. Co-isolation, unexpected modifications, incomplete fragmentation, and peptides from homologous proteins can also yield apparently persuasive but incorrect candidates.
A practical validation loop has three gates. First, assess the spectrum itself: precursor mass accuracy, fragment coverage, charge state, diagnostic ions, and whether a modified known peptide offers a better explanation. Second, reconcile the candidate with orthogonal sequence evidence by searching transcript assemblies, genomic translations, close-relative proteomes, and appropriate taxonomic databases. Third, where the claim matters biologically, confirm the peptide using targeted MS, a synthetic standard, or independent replicate observations.
Figure 4. De novo sequence predictions are most useful when treated as testable candidates that pass through spectral, sequence-context, and targeted-validation gates.
Protein Inference Is a Separate Problem From Peptide Identification
A confidently assigned peptide does not always identify one complete protein. In non-model organism studies, the gap between peptide-level and protein-level certainty can be especially wide because transcript contigs are fragmented, paralogs are closely related, and predicted proteins may overlap or differ only at unobserved regions.
This is where many reports become overconfident. Several peptides may support a protein group without resolving the exact member of a paralogous family. A peptide may map to one short contig and another peptide to a second contig, even though the two fragments were never demonstrated to derive from the same full-length transcript. Conversely, a genuine expressed protein may be underreported because its unique regions were not observed in the experiment.
Report the evidence at the level it supports. Use peptide-level reporting for novel sequences, protein-group reporting for shared evidence, and isoform-level calls only when unique peptides or direct transcript structure resolve the ambiguity. Preserve the mapping between every reported peptide, every inferred protein group, and the database version used in the search.
Protein inference becomes more robust when a project explicitly distinguishes three evidence layers:
- Peptide evidence: the spectrum supports this amino-acid sequence under stated FDR and quality criteria.
- Protein-group evidence: the observed peptide set is consistent with one or more related protein sequences.
- Exact protein or isoform evidence: unique peptides, transcript structure, or other orthogonal information distinguish one candidate from its alternatives.
Figure 5. Fragmented contigs and shared peptides can support a protein group while leaving the exact full-length isoform unresolved.
Apply FDR and Validation Rules According to Evidence Class
A single global FDR threshold is not enough when a database combines mature reference proteins with speculative translated sequences. The prior probability that a match is correct differs across these classes. A peptide matching a well-supported reference protein and a peptide matching a hypothetical six-frame ORF may carry the same nominal score while requiring different levels of scrutiny.
Class-specific FDR is a useful starting point. Separate known reference entries, transcript-supported additions, and novel genomic or de novo candidates before evaluating their identification behavior. Do not assume that a global 1% peptide-spectrum match FDR automatically provides 1% error control in each subgroup. Entrapment sequences, alternative database formulations, and targeted re-searches can provide additional checks when claims depend on new proteins or rare sequences.
Validation should also be matched to the intended conclusion. If the output is pathway-level exploration, transparent protein-group assignments and cautious annotation may be enough. If the output is a new gene model, species-specific peptide, functional isoform, or candidate sequence of biological interest, require stronger orthogonal evidence. A sensible escalation ladder is: high-quality spectrum; unique peptide; recurrence across samples; transcript or genomic support; targeted MS confirmation; and, where relevant, independent functional testing.
Choose DIA, DDA, or a Hybrid Acquisition Strategy With the Database in Mind
Database design and acquisition mode are linked. Data-dependent acquisition (DDA) can yield clean individual fragmentation spectra that are useful for candidate discovery, manual inspection, and sequence validation. Data-independent acquisition (DIA) can offer more consistent sampling across samples, but its multiplexed spectra can make unusual or poorly represented sequences more difficult to resolve without an appropriate library or database-guided strategy.
For non-model organism discovery projects, a hybrid design is often practical: use representative DDA data to build or refine a sample-specific library and identify high-confidence candidates, then use DIA for broader comparative quantification. The appropriate choice depends on sample number, expected proteome complexity, available sequence resources, and whether the central deliverable is novel sequence discovery or reproducible abundance comparison.
When quantitative consistency across many samples is the primary need, a DIA quantitative proteomics workflow can be paired with a versioned custom database and explicit peptide-selection rules. When the sequence search space is still unsettled, DDA discovery data and targeted follow-up may provide more transparent candidate review. Neither mode solves a weak database on its own.
The same constraint applies when the discovery target is unusually short or non-canonical: microprotein and sORF-encoded peptide discovery in tissue depends on a search library that makes the proposed sequence testable without uncontrolled search-space expansion. When the central challenge is instead a weak target signal in a high-dynamic-range fluid, the relevant decision shifts from sequence coverage to matrix depth and missingness, as discussed in this plasma DIA detectability guide.
Functional Annotation Should Preserve Uncertainty Rather Than Hide It
Once proteins have been identified, functional annotation can make an unfamiliar proteome biologically interpretable. But annotation is another inference layer. A BLAST-like homology match, conserved domain call, ortholog assignment, or protein-structure prediction can suggest a function; it does not prove the role of a specific sequence in the sampled organism.
For non-model systems, annotation should retain the evidence chain. Report whether a term is transferred from a close ortholog, supported by a conserved domain, inferred from a protein family, or experimentally characterized in the studied organism. This is especially important for rapidly evolving proteins, secreted peptides, immune-related genes, and gene families with lineage-specific expansion.
A transparent functional workflow typically includes sequence similarity searching, domain and motif analysis, orthology-aware mapping, enrichment testing with an explicit background set, and evidence-ranked biological interpretation. The background set should reflect the proteins that could realistically have been detected in the project—not an unrelated whole-genome catalog. This makes functional annotation and enrichment analysis more meaningful and reduces the risk of treating database incompleteness as biology.
Where evolutionary context is central to the question, protein evolution analysis can help separate conserved family membership from evidence of a lineage-specific sequence feature. It should be interpreted alongside peptide uniqueness and sequence quality, not used to rescue an ambiguous protein assignment.
Figure 6. The best strategy depends on the available evidence, required sequence specificity, tolerance for search-space expansion, and capacity for validation.
A Practical Selection Matrix
| Strategy | Best use case | Main strength | Main risk | Recommended validation focus |
|---|---|---|---|---|
| Close-relative homology database | Conserved pathways and rapid feasibility studies | Compact, annotated, fast to deploy | Overstated species or isoform specificity | Unique peptide review and orthology provenance |
| Matched RNA-seq-derived database | Tissue- and condition-specific discovery | Direct biological relevance to the sample | Assembly fragmentation and incomplete expression coverage | Contig provenance, peptide uniqueness, database versioning |
| Genome-derived translation database | Annotation refinement and novel coding-region discovery | Broad potential sequence coverage | Search-space inflation and miscalibrated FDR | Two-stage searches and class-specific FDR |
| De novo peptide sequencing | Unexplained high-quality spectra and minimal sequence resources | Can reveal database-absent candidates | Sequence ambiguity and overinterpretation | Spectrum quality, homology checks, targeted confirmation |
| Hybrid proteogenomics | Mixed evidence sources or high-value discovery questions | Balances completeness with constraint | Requires explicit integration rules | Versioned database, separate evidence classes, orthogonal validation |
For many projects, the hybrid route is the most realistic. Begin with the best existing reference, add sample-matched transcript-supported sequences where possible, restrict expanded genomic candidates with explicit rules, use de novo sequencing for unresolved spectra, and validate the novel tier separately. A hybrid database is not a maximal database. It is an evidence-weighted database.
Four Phases for an Audit-Friendly Non-Model Organism Proteomics Project
- Define the biological claim. Specify whether the project needs pathway trends, protein-family evidence, species-specific peptides, full-length isoforms, or novel coding sequences. Record specimen provenance, biological replicates, associated organisms, and the intended inference level.
- Audit the available sequence evidence. Evaluate public proteomes, close relatives, draft genomes, transcriptomes, taxonomic context, and the feasibility of matched RNA or DNA sequencing. Choose a starting database route before LC-MS/MS data are interpreted.
- Build and search a versioned, evidence-tagged database. Keep reference, transcript-supported, translated, and contaminant sequences distinguishable. Use staged searching when the database contains speculative sequences. Preserve every filtering rule and accession source.
- Validate, annotate, and report at the supported level. Review novel candidates separately, use targeted confirmation where appropriate, apply transparent protein inference, and state whether functional assignments are direct, orthology-based, or domain-based.
These phases are iterative: a high-confidence novel peptide may expose a missing transcript, a weak protein group may flag an assembly problem, and a revised database may justify a focused re-search. An integrated custom proteomics service and proteomics bioinformatics workflow can keep these decisions connected from planning through evidence review.
Figure 7. A versioned, four-phase workflow creates a traceable path from specimen provenance to validated protein and functional conclusions.
Conclusion: Treat the Database as Part of the Experiment
For non-model organisms, the database strategy sets the upper limit of what proteomics can credibly discover. The right choice is not the largest possible FASTA file or the most sophisticated algorithm in isolation. It is the smallest evidence-appropriate search space that still contains the biological sequences required to answer the question.
Close-relative databases are valuable for conserved biology. Matched RNA-seq databases improve relevance and specificity. Genome-derived translations open discovery but need staged searching and tailored FDR control. De novo sequencing can reveal missing candidates but requires independent validation. When these routes are combined deliberately, proteomics becomes a tool not only for measuring an unfamiliar proteome, but also for improving the sequence resources that make future studies possible.
Frequently Asked Questions
Can I perform non-model organism proteomics without RNA-seq?
Yes. A well-chosen close-relative database or de novo peptide sequencing route can support useful results. RNA-seq becomes more important when exact sequence identity, tissue-specific isoforms, or lineage-specific proteins are central to the study.
How close must a reference species be for homology-based searching?
There is no universal sequence-identity threshold. Suitability depends on the conservation of the target proteins, the evolutionary history of relevant gene families, and the claim being made. Assess multiple close relatives when possible and avoid interpreting shared peptides as species-specific evidence.
Does a transcript-supported protein require less proteomic validation?
No. RNA support increases plausibility, but it does not prove that the protein was present in the analyzed proteome. Novel or high-impact findings still require peptide-level quality review and, where appropriate, orthogonal confirmation.
Can DIA be used when no spectral library exists for the species?
Yes, but the analysis should use an appropriate sequence database or project-built library and should retain transparent peptide-selection criteria. Representative DDA acquisition can be useful for building a sample-specific discovery foundation.
Why not search every possible six-frame translation?
An unrestricted translation introduces many implausible sequences, reducing sensitivity and making error control more difficult. A broad first-pass discovery search followed by evidence-based filtering and focused re-searching is usually more defensible.
How should I report ambiguous protein assignments?
Report them as protein groups or homologous families unless unique peptides, transcript structure, or other direct evidence resolve a single sequence. Do not convert peptide-level evidence into unsupported isoform-level claims.
References
- Li H, Joh YS, Kim H, Paek E, Lee SW, Hwang KB. Evaluating the effect of database inflation in proteogenomic search on sensitive and reliable peptide identification. BMC Genomics. 2016;17(Suppl 13):1031. doi: 10.1186/s12864-016-3327-5.
- Zhang X, Ning Z, Mayne J, et al. MetaPro-IQ: a universal metaproteomic approach to studying human and mouse gut microbiota. Microbiome. 2016;4:31. doi: 10.1186/s40168-016-0176-z.
- Cheng K, Ning Z, Zhang X, et al. MetaLab: an automated pipeline for metaproteomic data analysis. Microbiome. 2017;5:157. doi: 10.1186/s40168-017-0375-2.
- Werner J, Géron A, Kerssemakers J, Matallana-Surget S. mPies: a novel metaproteomics tool for the creation of relevant protein databases and automatized protein annotation. Biology Direct. 2019;14:21. doi: 10.1186/s13062-019-0253-x.
- Chen TW, Gan RC, Fang YK, et al. FunctionAnnotator, a versatile and efficient web tool for non-model organism annotation. Scientific Reports. 2017;7:10430. doi: 10.1038/s41598-017-10952-4.
- Yilmaz M, Fondrie WE, Bittremieux W, et al. Sequence-to-sequence translation from mass spectra to peptides with a transformer model. Nature Communications. 2024;15:6427. doi: 10.1038/s41467-024-49731-x.





