
Quantitative proteomics has become essential for understanding protein expression across biological conditions, with data-independent acquisition (DIA) methods now enabling analysis of hundreds of samples daily. However, DIA data contains substantial noise and interference from co-eluting and co-fragmenting peptide species, making accurate quantification challenging. Traditional fragment selection methods rely on single metrics like extracted ion chromatogram (XIC) correlation, which can exclude valid quantitative information whilst failing to capture the complex factors determining fragment reliability. The critical barrier has been the absence of ground truth data for protein abundances, preventing conventional machine learning approaches to optimise fragment selection.
Vu et al. developed QuantSelect, a deep learning framework addressing this limitation through self-supervised learning. The system integrates multiple quality features, including XIC correlation, retention time deviation, mass accuracy, and precursor intensities, to identify high-quality fragment ions for quantification. A regularised weighted variance loss function enables training without ground truth by assuming that median intensities of high-quality proteins approximate true abundance values.
For experimental validation, the team from the Mann lab utilised an Orbitrap Astral interfaced with an Evosep One. Peptides were separated on an Aurora® Rapid™ 5×75 C18 UHPLC column using the Whisper Zoom 80SPD method with a 16.3-minute gradient. Ionisation was performed via an EASY-Spray source equipped with a FAIMS Pro interface.
QuantSelect achieved 12-15% improved precision compared to single-metric filtering across mixed-species benchmark datasets, with up to 4-fold enhanced sensitivity in differential expression analysis. Applied to challenging single-cell embryonic stem cell data (169 cells, 6,329 proteins), QuantSelect demonstrated superior reproducibility and an 18% increase in sensitivity.
This advancement enables more reliable quantification across the entire protein abundance range, particularly benefiting low-abundance proteins where traditional methods struggle, ultimately improving biological discoveries in high-throughput DIA proteomics applications.
Publication
bioRxiv
Authors
Duc Tung Vu, Georg Wallmann, Marvin Thielert, Enes Ugur, Marc Oeller, Maximilian Zwiebel, Constantin Ammar, & Matthias Mann;
Title
Deep Learning-Driven Fragment Ion Selection for Improved Quantification in MS based Proteomics


