Targeting lightweight spectrogram library construction system for high-sensitivity single-cell proteome research and application of targeting lightweight spectrogram library construction system

The targeted lightweight spectral library construction system solves the problems of high sample dependence, poor generalization and limited recognition depth in single-cell proteomics. It achieves efficient recognition of low-abundance regulatory proteins, improves recognition accuracy and platform compatibility, and is suitable for basic scientific research, clinical research and high-throughput proteomics platforms.

CN121768474AActive Publication Date: 2026-03-31ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-03
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies in single-cell proteomics analysis suffer from problems such as high sample dependence, poor generalization, severe coverage bias, limited identification depth, poor result reproducibility, and fragmented analysis process, making it difficult to effectively identify low-abundance regulatory proteins and improve identification accuracy.

Method used

A targeted, lightweight spectral library construction system is used to construct a lightweight spectral library suitable for single-cell proteomics through steps such as high-confidence peptide extraction, priority import of functional proteins, expression score-driven compression, structural noise reduction and spectrum optimization, and pseudo-spectrum generation. It supports multi-platform compatibility and personalized construction.

Benefits of technology

It improves the functional protein recognition rate, reduces the redundancy of the spectral library, enhances the accuracy and repeatability of results, and strengthens the platform's versatility and workflow integration, making it suitable for practical application scenarios with multiple samples and multiple conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768474A_ABST
    Figure CN121768474A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of bioinformatics and proteomics, and particularly relates to a targeting lightweight spectrogram library construction system for high-sensitivity single-cell proteome research and application of the targeting lightweight spectrogram library construction system. According to the system, on the basis of single cell DIA data, accurate construction and redundancy control of a spectrogram library are achieved through high-confidence-coefficient peptide fragment extraction, functional protein priority import, target-oriented spectrogram deduction, expression score-driven entry compression and structure noise reduction optimization. According to the method, the protein recognition rate and repeatability of single-cell DIA data are remarkably improved, and particularly, the method is outstanding in recognition of functional proteins such as membrane receptors, kinases, transcription factors and ubiquitin regulatory proteins. The method can also be applied to key functional protein identification of various immune cells and tumor-related macrophage subpopulation identification, provides efficient and reliable technical support for functional analysis and subpopulation classification of a single-cell proteome level, and has good universality and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and proteomics, and in particular to a targeted lightweight spectral library construction system for high-sensitivity single-cell proteomics research and its application, which can achieve efficient identification and specificity enhancement of functional proteins, thereby improving the detection sensitivity and biological interpretability of rare regulatory proteins. Background Technology

[0002] In recent years, single-cell proteomics based on mass spectrometry has become a frontier direction in proteomics research as an important means of revealing cellular functional heterogeneity. Mass spectrometry data acquisition-independent (DIA) technology has become a core means of single-cell proteomics analysis due to its good reproducibility and recognition depth. To improve the protein recognition efficiency of DIA data, the quality of the spectral library has a decisive impact on the overall analytical performance. The current spectral library generation methods mainly include three categories: (1) Reference libraries constructed through experimental means, which usually require a large number of samples and preprocessing steps such as fractionation. (2) Self-learning methods based on the sample's own data, i.e., the "Library-free" strategy, which extracts pseudo-spectrals from the sample's DIA data and constructs a database. (3) Deep learning model prediction, which generates virtual spectral libraries (in silico spectral libraries) with peptide retention time and fragment ion characteristics through neural network models. Its performance is highly dependent on the training data and the generalization ability of the model.

[0003] Currently, all three methods have significant limitations in single-cell proteomics analysis: 1. Limitations of the experimental reference library method: (1) High starting sample size dependence. The reference library constructed in the experiment usually depends on a large number of samples and requires strategies such as multidimensional fractionation, long gradient separation, and repeated injection. However, the protein content of single cells is extremely low (usually 200 pg, about one-thousandth of that of a normal sample), which is far lower than the input amount required for data acquisition.

[0004] (2) Poor generalization. Experimental libraries rely on specific experimental platforms, mass spectrometer models, and tissue sources, and their content features lack broad applicability across platforms and samples. Directly using experimental libraries from different sources for single-cell data analysis will lead to a significant decrease in identification efficiency and accuracy due to problems such as incomplete peptide coverage, mismatched fragment ion distribution, and retention time shift.

[0005] (3) Severe coverage bias. Spectroscopic libraries tend to capture high-abundance structural and metabolic enzyme proteins, making it difficult to cover low-abundance regulatory proteins, such as membrane receptors, kinases, and transcription factors. Although these proteins have low copy numbers, they play a decisive role in cell signal transduction, immune regulation, and disease development. This bias is relatively less pronounced in constant-volume samples due to high starting quantities and fractionation strategies; however, at the single-cell level, these key proteins are often completely absent, severely limiting the depth and explanatory power of single-cell proteomics applications from "descriptive research" to "mechanism analysis."

[0006] 2. Limitations of the self-learning library-free method: (1) The library construction stage is prone to introducing pseudo-identification. Self-learning libraries rely on extracting pseudo-spectra directly from sample DIA data to construct the database. However, raw single-cell DIA data generally face challenges such as weak ion signals, significant noise background, sparse quantification, and unstable fragment ion distribution, which greatly limits the accurate extraction of pseudo-spectra. Single-cell self-learning libraries often contain a large number of low-confidence or noise-driven entries, which can easily lead to problems such as incorrect matching, affecting the accuracy of protein identification and the reliability of downstream analysis.

[0007] (2) Limited recognition depth. Since it does not rely on external knowledge or prior information, its recognition ability is entirely limited to the signals that can be extracted from the current data. For single-cell data with weak original signals and complex backgrounds, it is very easy to miss regulatory proteins with low expression levels and insignificant fragment spectral features, such as membrane receptors, kinases, and transcription factors. If they are not captured in the library construction stage, irreversible information loss will occur in the early stages of analysis, severely limiting the functional analysis and annotation capabilities of single-cell proteomes.

[0008] (3) Poor reproducibility of results. Single-cell data itself is highly volatile and subject to technical noise, such as mass spectrometry signal drift, ion loss, and random detection. These problems are common, which makes the pseudo-spectrums extracted by this method vary greatly between different single-cell samples. This leads to a lack of consistency in protein identification results between samples, which weakens the comparability and statistical power of downstream quantitative analysis and limits its application in systematic single-cell studies.

[0009] 3. Limitations of deep learning-based methods for predicting spectral libraries: (1) Prediction distortion caused by training data bias. Existing models are mostly trained on high-input, high-quality conventional sample data, and their prediction results are based on the assumption of high signal-to-noise ratio, high coverage and low randomness of spectral distribution. However, in single-cell samples, the mass spectrometry signal intensity is low, the ion statistics are unstable and the fragment spectrum is highly variable, which deviates significantly from the characteristic distribution of the training data, resulting in a significant decrease in the adaptability of the model output in single-cell applications.

[0010] (2) Lack of customized screening and feedback mechanisms. Existing spectral library construction strategies generate "full coverage" data by targeting all enzymatically digested peptides, resulting in extremely large scales, typically exceeding the actual size of single-cell spectra by more than a hundred times. This leads to high redundancy and low effective information density. The lack of quality assessment and targeted screening mechanisms for single-cell data features significantly increases the computational complexity of the search process and exacerbates spectral interference. In analytical scenarios where single-cell signals are sparse and background ion interference is limited, this problem weakens search efficiency and identification accuracy, amplifying the risk of false positives.

[0011] (3) Fragmented analysis workflow. Currently, spectral library prediction methods are often nested as independent modules within complex analysis workflows. Researchers need to use different platforms (such as DIA-NN, Spectronaut, AlphaPeptDeep, Prosit, etc.) to complete multiple steps such as model training, peptide prediction, spectral library generation, search matching, and functional annotation. This cross-platform, multi-module analysis method relies on a large amount of manual data transfer, which not only increases the complexity of the operation but also easily introduces implicit errors such as file format conversion errors, version compatibility issues, and parameter inconsistencies, which seriously restricts its widespread application in large-scale single-cell sequencing. Summary of the Invention

[0012] To address the aforementioned problems in existing technologies, this invention provides a targeted, lightweight spectral library construction system for highly sensitive single-cell proteomics research and its applications. The objective of this invention can be achieved through the following technical solutions, including: Firstly, a spectral library construction system is provided, encompassing peptide selection, functional enhancement, expression screening, noise reduction and compression, platform compatibility, quality control, and personalized construction. The system includes: S1: High-confidence peptide extraction. Based on representative single-cell DIA data, a set of high-confidence peptides with good reproducibility and stable signal intensity was extracted as the basis for construction; The peptide set should cover basic characteristic parameters such as the parent ion mass-to-nucleus ratio, major fragment types, post-translational modification information, relative intensity distribution, and retention time interval, and can be used as a reference for the structural regulation and matching of subsequent entries.

[0013] The single-cell DIA data were obtained from single-cell sorting, lysis, enzyme digestion, and detection using liquid chromatography-high resolution mass spectrometry. S2: Prioritize the import of functional proteins and construct a target-guided spectral deduction library. Combine functional protein category annotation information to prioritize the inclusion of peptide information of functional proteins. A high-priority retention mechanism is configured for this type of protein during the construction process; During the construction process, a spectral deduction library was built for functional proteins not included in S1.

[0014] S3: Expression Score-Driven Item Compression. Based on single-cell protein expression data, all candidate items are scored for expression propensity, and content compression and item filtering are performed according to the scores. The data expressed can be derived from public databases and experimentally constructed micro and constant samples. Low-scoring items will be downgraded or removed according to the rules.

[0015] S4: Structural Denoising and Spectral Optimization. To address the characteristics of few ion fragments and strong background signals in single-cell mass spectrometry data, structural denoising was performed by setting thresholds for peptide length, fragment ion quantity, and m / z region density. In this process, redundant entries in highly overlapping areas are removed, significantly reducing redundancy and improving matching efficiency.

[0016] S5: The output spectral library includes fields such as standardized retention time, fragment structure, and functional classification, and the file format is compatible with mainstream analysis platforms such as DIA-NN and Spectronaut.

[0017] S6: After the spectral library is built, in order to meet the DIA search engine's requirements for false discovery rate (FDR) control, automatically construct a spectrum matching the target spectrum. Figure 1 A corresponding pseudo-spectral (decoy) is generated and used for subsequent FDR estimation and result back-screening to improve the reliability and accuracy of the overall identification.

[0018] It provides a configurable spectrum library construction interface, allowing users to specify target cell types, species, target protein sets, or research platforms as needed to configure spectrum library generation parameters. It automatically generates target-oriented lightweight spectrum libraries in batches, with good scalability and reusability.

[0019] Specifically, the extraction of high-confidence peptides from S1 includes the following steps: The target cells were sorted, lysed, and enzymatically digested. Raw DIA mass spectrometry data of single cells were obtained by liquid chromatography-high resolution mass spectrometry. Based on more than 20 raw DIA mass spectrometry data of each cell type, the raw mass spectrometry data files were subjected to characteristic peak extraction, retention time alignment, and theoretical sequence comparison through a public search platform. Based on the feature extraction module, peptides that appeared in most samples and had stable signals were generated. The mass-to-nucleus ratio, retention time, fragment distribution, and relative intensity characteristics of the parent ion were extracted to construct a high-confidence seed set.

[0020] Specifically, the list of functional proteins in S2 includes protein classification results from public annotation databases, and the proteins include at least membrane receptors, ligands, kinases, ubiquitin ligases (including E1, E2, and E3 ubiquitin ligases), and transcription factors.

[0021] During the construction process, the spectra of the peptides corresponding to the above categories are retained first, and they are not removed even if they do not appear in S1.

[0022] Specifically, the construction of the target-oriented spectrogram inference library is based on the AlphaPeptDeep architecture and incorporates a transfer learning strategy.

[0023] Specifically, the expression scoring mechanism in S3 includes the following steps: Based on public or experimental expression data for the target cell type, assign probability values ​​to each candidate peptide spectrum. If the value is lower than the set threshold and the structural score is also low, it will be rejected.

[0024] Specifically, the structural optimization measures in S4 include: The peptide length is set to 7-25 amino acids, and the number of fragment ions does not exceed 12. Remove frequently occurring repetitive peptides in the m / z set region to reduce false matches.

[0025] Specifically, the spectrum library output in S5 includes the following fields: Precursor m / z, Charge, Fragment type, Fragment ion m / z, Relative Intensity, iRT calibration value, Protein name; The output format is .tsv, so no manual format conversion is required.

[0026] Furthermore, the spectral library of the present invention can be directly used for single-cell DIA data search after being loaded by the Spectronaut and DIA-NN search modules, and the recognition rate is improved by about 30-40% compared with the traditional self-learning library.

[0027] It performs particularly well in functional protein recognition, with a 42.7% increase in the number of membrane protein recognitions and a 58.3% increase in kinases.

[0028] Specifically, the pseudo-spectrum (decoy) generation method in S6 is as follows: While keeping the m / z value of the parent ion unchanged, the amino acid sequence is reversed to construct random but structurally valid peptide entries, and the corresponding fragment spectra are generated by the same spectrogram calculation framework. The obtained pseudo-spectrum and the target spectrum are used together for search analysis to construct a target-non-target distribution model, thereby completing the FDR curve fitting and confidence threshold setting.

[0029] In a second aspect, the application of the method described in the first aspect in the prediction of key functional proteins of various immune cells is provided, characterized in that the immune cells include lymphocytes, neutrophils and macrophages; and the key functional proteins include single-cell level identification of membrane receptor ligands, kinases, transcription factors and ubiquitin-regulated proteins.

[0030] Thirdly, the method described in the first aspect is provided for the application of tumor-associated macrophage (TAM) subset identification, characterized in that the tumor-associated macrophages are obtained from the mouse melanoma microenvironment; the application is used to predict functional proteins that regulate biological processes such as phagocytosis, inflammation, metabolic reprogramming, and self-renewal.

[0031] Furthermore, single-cell proteomics results based on a customized functional spectrum library were used to construct a TAM subgroup classification model to distinguish functional subgroups such as TAM.1, TAM.2, TAM.3, and TAM.4.

[0032] Furthermore, at least one type of biological characteristic in the TAM subgroup is predicted by single-cell functional protein feature vectors, including but not limited to: phagocytosis-enhanced, metabolically active, self-renewing, or inhibited subgroups.

[0033] Fourthly, an electronic device is provided for performing the method of the first aspect, the electronic device comprising: a processor, and a memory coupled to the processor, the memory for storing a computer program; the processor being configured to execute the computer program stored in the memory such that the electronic device performs the method as described in any possible implementation of the first aspect.

[0034] Fifthly, a computer-readable storage medium is provided, comprising a computer program or instructions that, when executed on a computer, cause the computer to perform a method as described in any of the possible implementations of the first aspect.

[0035] In a sixth aspect, a computer program product for customized analysis of single-cell proteomics is provided, comprising a computer program or instructions that, when run on a computer, cause the computer to perform the method as described in any possible implementation of the first aspect.

[0036] The beneficial effects and advantages of the present invention include, but are not limited to: 1) Improve the recognition rate of functional proteins, especially by enhancing the treatment of regulatory proteins that are easily lost in traditional methods; 2) Reduce the redundancy of the spectrum library by using a dual screening mechanism of expression probability and structure scoring to achieve a dynamic balance between the number of spectra and analysis efficiency; 3) Improve the accuracy and repeatability of results; multi-dimensional quality control and targeted optimization ensure the reliability of the search process. 4) Enhance the platform's versatility and process integration, support user customization and be compatible with mainstream analysis software, and have promotional application value; 5) Applicable to real-world application scenarios with multiple samples and conditions, supporting flexible deployment in different tissue types, disease models, or functional studies.

[0037] In summary, this invention provides a spectrum library construction strategy that integrates targeting, adaptability, and lightweight design, breaking through the applicability limitations of traditional construction methods at the single-cell level. It can be widely applied in basic research, clinical research, and high-throughput proteomics platforms, and is of great significance for promoting the functional analysis of single-cell proteomics technology in rare cell populations. Attached Figure Description

[0038] Figure 1 This is a schematic diagram of the workflow of a single-cell proteomics analysis system for enhancing functional protein recognition according to the present invention.

[0039] Figure 2 The size statistics of immune cells were selected for this invention. Figure 2 a is a single-cell image acquired by a single-cell sorting system (Cellenone™ technology platform) with image recognition capabilities; Figure 2 b is a single-cell diameter distribution map generated statistically based on the single-cell sorting system. The value marked in the map is the mode of the diameter distribution, which is used to characterize the representative cell diameter of the corresponding cell type.

[0040] Figure 3 This is the result of the application of the present invention in the recognition of low-abundance functional proteins in immune cells. Figure 3 a represents the statistical results of total protein identification for different immune cell types; Figure 3 b shows the identification results of four key functional proteins: membrane proteins, transcription factors, protein kinases, and ubiquitin ligases, and demonstrates the percentage improvement in identification compared to the control method using the method of this invention.

[0041] Figure 4 This presents the application results of the present invention in the identification of tumor-associated macrophage (TAM) subsets. Among them, Figure 4 a is the UMAP distribution map of the four TAM functional subgroups (TAM.1–TAM.4) identified after dimensionality reduction and clustering of the data obtained based on the customized spectral library of this invention; Figure 4 b is a distribution diagram of the number of proteins identified at the single-cell level in each TAM subpopulation; Figure 4 c is a comparison of the distribution characteristics of four key functional proteins—membrane proteins, transcription factors, protein kinases, and ubiquitin ligases—in various TAM subgroups and in data not constructed using the spectral library construction method of this invention. Detailed Implementation

[0042] The following detailed embodiments further illustrate the concept and technical effects of the present invention to fully understand its purpose, features, and effects. Unless otherwise specified, all methods described are conventional methods. Unless otherwise specified, all materials are available from publicly available commercial sources. The illustrative embodiments and descriptions of the present invention are used to explain the invention and do not constitute an undue limitation thereof. It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0043] It should be noted that although functional modules are divided in the schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than those shown in the schematic diagram or the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0044] Example 1: Construction of a Single-Cell Proteomics Data Atlas Library System

[0045] Please see Figure 1 This application provides a targeted, lightweight spectral library construction system for single-cell proteomics data. The system includes, but is not limited to, the following six core steps: S1. High-confidence peptide spectrum extraction. First, the target single cells were sorted, lysed, and enzyme-digested. Then, raw DIA mass spectrometry data of the single cells were acquired using liquid chromatography-high resolution mass spectrometry (timsTOF or Orbitrap platform).

[0046] Specifically, the single-cell sorting is achieved using a single-cell sorting system (Cellenone™ technology platform) with image recognition capabilities.

[0047] Specifically, the single-cell lysis product is obtained using one or more of the nonionic surfactants dodecyl maltodextrin, n-octyl β-D-glucopyranoside, sodium dodecyl sarcosinate, and sodium dodecyl sulfate at a concentration of 0.1%, 0.2%, or 0.5%.

[0048] Specifically, in the single-cell enzymatic digestion process, peptides are obtained from proteins by digestion with one or more enzymes, including trypsin and Lys-C enzyme, at 37°C for 1 to 4 hours.

[0049] Specifically, the single-cell sorting, lysis, and enzyme digestion processes are performed using non-transfer reaction vessels, such as 96-well plates, 384-well plates, or liquid chromatography vials.

[0050] Specifically, based on DIA data from at least N single cells for each cell type, a public search engine (DIA-NN or Spectronaut) is used to complete feature peak extraction, retention time alignment, and peptide identification. Here, N represents the number of single-cell DIA data points, preferably 10-50, and more preferably 20.

[0051] Define the candidate peptide set as:

[0052] After feature extraction of all DIA data, for any candidate peptide p∈P, the proportion f(p) of its occurrence in the sample set is defined as:

[0053] Where II(·) is an indicator function, which is 1 if peptide p is detected in the i-th single-cell sample, and 0 otherwise.

[0054] The signal stability evaluation index is defined as the intensity variation coefficient CV(p) of the first-order mass spectrometer (MS1):

[0055] Among them, I p Let be the set of MS1 intensities of peptide p across all single-cell samples, where μ and σ are the mean and standard deviation, respectively.

[0056] Peptides that meet the following criteria are selected to construct a high-confidence seed set P. High :

[0057] For each retained peptide p, the following characteristic parameters were extracted: precursor m / z, retention time (RT), charge, major fragment ion type and m / z, and relative intensity.

[0058] S2. Construction of a Functional Protein Priority Import and Target-Guided Spectral Inference Library. To enhance the identification of key regulatory proteins in single-cell proteomics data, this invention further introduces a functional protein priority import mechanism to construct a target-guided functional enhancement spectral library.

[0059] Specifically, a functional protein set G is constructed based on a public functional annotation database. funcThe functional protein set includes at least membrane receptors, ligands, protein kinases, transcription factors, and ubiquitin ligases (E1, E2, E3). The functional protein set is constructed as follows: Membrane protein set: Based on subcellular localization annotation information in the Gene Ontology database, the set of membrane proteins was obtained by filtering according to the annotation entries related to "cellmembrane" or "cell surface". It includes 3,353 mice-derived proteins and 4,288 human-derived proteins.

[0060] Transcription factor set: A set of transcription factor proteins was obtained by screening the AnimalTFDB database, including 1,171 mouse transcription factors and 1,617 human transcription factors.

[0061] Protein kinase collection: A collection of protein kinases was constructed based on the kinase.com database, which includes 508 mouse protein kinases and 509 human protein kinases.

[0062] Ubiquitin ligase set: Based on entries labeled "ubiquitin enzyme" in the BRENDA database, including 300 mouse ubiquitin ligases and 317 human ubiquitin ligases.

[0063] All functional protein entries are limited to UniProt reviewed (Swiss-Prot) entries to ensure the reliability and consistency of annotation information, and are uniformly converted to UniProt protein IDs.

[0064] For any candidate peptide p, its corresponding protein is denoted as g(p). Define a function R to prioritize and preserve performance. func (p) is as follows:

[0065] When R func When (p)=1, the peptide is identified as a functional protein-related peptide. During the spectral library construction process, functional protein-related peptides are retained and included in the candidate spectral library even if they do not meet the screening conditions in step S1.

[0066] In addition, when the following conditions are met: and

[0067] If the peptide belongs to a functional protein-related peptide but was not detected in the high-confidence peptide set in step S1, then the present invention further automatically constructs its deduced spectrum entry based on the peptide sequence, fragmentation rules and instrument parameters to complete its structural information in the spectrum library.

[0068] Specifically, based on the amino acid sequence information of this peptide, the following information is simulated, including but not limited to the theoretical precursor ion m / z, the relative intensity distribution of the precursor ion, neutral loss modification, the theoretical fragment ion type and m / z, and the relative intensity distribution of fragment ions. The simulated spectra are then compared with the single-cell data source spectra. Figure 1 And included in the candidate spectral library.

[0069] Specifically, the construction of the inferred spectral entries is based on the AlphaPeptDeep architecture, incorporating a transfer learning strategy. The core idea of ​​this method is to utilize the potential physicochemical correlations between the amino acid sequence of a single-cell peptide and its retention time (RT) and secondary spectra (MS2) to infer the most likely RT and MS2 features of a given sequence. In implementation, a transfer learning algorithm is used to allow the model to learn the single-cell data distribution characteristics of the corresponding cell type through specified single-cell DIA data. The weights are then adjusted to more accurately model the influencing factors of single-cell data. Ultimately, this method can effectively capture the RT and MS2 features of peptides in single-cell samples, thereby achieving more accurate spectral library inference.

[0070] Specifically, the preferred parameters for the transfer learning are: epoch_ms2 = 100; warmup_epoch_ms2 = 10; batch_size_ms2 = 512; epoch_rt_ccs = 40; warmup_epoch_rt_ccs = 10; default_instrument = Astral; default_nce = 25.0.

[0071] Specifically, the preferred parameters for spectral deduction are: enzyme cleavage specificity of trypsin, allowing a maximum of one missed cleavage site; no fixed modification; a maximum of one variable modification allowed; peptide length of 7–35 amino acids; parent ion charge of 2–4; and fragment ion charge up to 2.

[0072] Specifically, the learning effect of the model after transfer learning is judged by quality control, and the indicators include the coefficient of determination (R²), Pearson correlation coefficient (PCC), and angle similarity (SA).

[0073] Specifically, R² is defined as follows:

[0074] in, and These are the RT and extrapolated RT for peptide i, respectively, derived from single-cell data. This represents the mean RT for all single-cell data.

[0075] Specifically, PCC is defined as follows:

[0076] in, and The values ​​represent the extrapolation of the same peptide segment and the fragment ion intensity from single-cell data, respectively. and Let n be its mean, and n be the number of fragment ions.

[0077] Specifically, SA is defined as follows:

[0078] in, and To deduce the L2 normalized vector of fragment ion intensity from single-cell data.

[0079] S3. Expression score-driven item compression. Based on single-cell protein expression data, a peptide expression scoring mechanism is introduced, and content compression and item filtering are performed according to the score.

[0080] Specifically, the single-cell protein expression data comes from public databases, including SPDB (https: / / scproteomicsdb.com / ) and Slavov Lab (https: / / scp.slavovlab.net / data) data.

[0081] Specifically, define peptide Expression probability score:

[0082] Among them, E j (g(p)) represents the expression level of protein g(p) in the j-th public database sample, T is the set expression threshold, M is the number of samples in the public database, and II(·) is the indicator function.

[0083] Furthermore, combined with the structural score S struct (p):

[0084] in, The fragment ion quantity score measures the actual number of fragment ions observed and normalizes it to the theoretically generated number of fragment ions. To score the coverage of fragment ion m / z distribution, i.e. to calculate the coverage ratio of fragments in different m / z ranges; To score the consistency of isotopic structures, the cosine similarity between the experimental isotopic distribution and the theoretical isotopic pattern is calculated; w1-3 are the corresponding weights.

[0085] An entry is removed when the following conditions are met:

[0086] S4. Structural Denoising and Spectral Optimization. To address the characteristics of few ion fragments and strong background signals in single-cell mass spectrometry data, structural denoising is performed by setting thresholds for peptide length, fragment ion quantity, and m / z region density.

[0087] Specifically, structural noise reduction follows the following filtering conditions: Peptide length restrictions: 7 ≤ Length(p) ≤ 25; Number of fragment ions ≤ 12; m / z density determination: A maximum of 2 parent ions are allowed within a 5 Da window, and the density function is defined as follows:

[0088] Among them, #Precursors in [m / z a , m / z b [m / z] a to m / z b Number of precursor ions within the range, m / z a and m / z b These are the lower and upper limits of the 5Da window, respectively. If D(p) > 0.4, it is marked as a high-density redundant entry and removed.

[0089] S5. Standardized output format compatible with multiple proteomics analysis platforms. The output spectrum library is a .tsv file, with fields including: Precursor m / z, Charge, Fragment type, Fragment ion m / z, Relative Intensity, iRT calibration value (obtained from standard peptide regression), and Protein name.

[0090]

[0091] Where α and β are parameters obtained by regression fitting of standard peptides.

[0092] S6, Pseudospectral Library (P) decoy Library generation and false discovery rate (FDR) control. To meet the FDR control requirements, a pseudo-spectrum set P is constructed. decoy For each target peptide p, a pseudo peptide p is generated by sequence inversion. With the parent ion m / z remaining unchanged, the corresponding RT and MS2 spectra are derived using the single-cell type model after S2 transfer learning.

[0093] During the search process, the following FDR definition is used:

[0094] in, To match the number of peptides in P_decoy, The number of peptides matched to the template library (i.e., the spectral library constructed in this invention).

[0095] To enhance user experience and construction flexibility, this invention provides an intuitive and user-friendly visual spectral library construction platform. The platform is deployed via a web interface and features a graphical navigation bar divided into five modules: Home, Services, Documentation, Help, and About us. The Home module showcases the platform's overall framework and design philosophy, including methodological advantages (such as targeted functional protein optimization, lightweight compression, and high compatibility with the DIA engine), and provides an overview of the operation process to help users quickly understand the platform's purpose and operational logic. The Services module provides users with personalized spectral library construction functions, including the following input items: cell type selection, such as "macrophage" or "T cell"; species selection, supporting "Human" and "Mouse" species; target protein list upload, supporting user-uploaded custom protein IDs, such as UniProtID or Gene Symbol; peptide sequence input, supporting FASTA sequence pasting or txt file upload; after the user submits the parameters, the system automatically executes the species, cell type, and peptide retrieval, feature extraction, and spectral library generation process in the background. The generated spectrogram library structure supports interactive browsing and downloading, and the results page provides additional information for each entry, including protein information, source basis and construction method (experimental / deductive), expression score, and functional classification. The Documentation module details the principles of spectrogram library construction, data sources, scoring and filtering algorithm implementation, model training strategies, and parameter setting basis. The documentation supports English / Chinese bilingual switching. The Help module provides a concise user guide covering typical use cases (such as building a "mouse macrophage library"), and illustrates input steps and output formats with diagrams. The Aboutus module showcases the platform's development background, participating team members, major publications, and future expansion directions, enhancing academic transparency and the potential for sustainable development.

[0096] The embodiments of this application described above will be specifically described below through a combination of multiple examples.

[0097] Example 2: Application in enhancing the recognition of low-abundance functional proteins in immune cells

[0098] To verify the practicality and effectiveness of the spectral library construction system described in this invention, this embodiment applies the system to the recognition of key functional proteins of various immune cells. The target proteins include regulatory molecules such as membrane proteins, protein kinases, transcription factors, and ubiquitin regulatory proteins.

[0099] (1) Experimental background and target cell type

[0100] Three representative immune cell types were selected, including: lymphocytes: including CD4+ cells. T cells are primarily responsible for antigen recognition and immune memory; neutrophils are important innate immune cells that participate in chemotaxis, phagocytosis, and NET formation; macrophages perform functions such as phagocytosis, inflammation regulation, and tissue repair, and are key immune cells in the tumor microenvironment.

[0101] All the cells were primary cells derived from mice. Single-cell sorting was performed using an image recognition single-cell sorting platform, followed by lysis, enzyme digestion, and DIA mass spectrometry data acquisition.

[0102] (2) Experimental results

[0103] ①Please refer to Figure 2 According to measurements and statistics, the average cell diameters of the three cell types are approximately: CD T cells ~10μm, neutrophils ~13μm, macrophages ~17μm.

[0104] ② Customized spectrogram libraries were constructed for the three types of cells, and the inference spectrogram scoring results after transfer learning are shown in Table 1: Table 1 Spectrum Scoring Results

[0105] All three indicators take values ​​in the range of 0–1. The closer the value is to 1, the higher the accuracy of the inferred spectrum and the closer the structural features are to the real spectrum.

[0106] The results showed that the model after transfer learning exhibited high consistency and stability in all three types of immune cells, indicating that the model had successfully learned and mastered the typical spectral features in single-cell mass spectrometry data. This capability provides a solid foundation for the deduction of functional protein peptide spectral entries in this invention.

[0107] ③ Finally, three customized spectral libraries were generated: Lib_CD8T, Lib_Neu, and Lib_M, and their basic information is shown in Table 2: Table 2 Basic Information of the Three Customized Spectral Libraries

[0108] ④ Please refer to Figure 3 The DIA data of the three cell types were searched and compared using both a traditional self-learning spectronaut library (library-free, constructed by Spectronaut) and the customized spectronaut library of this invention. The method of this invention showed a significant improvement in each cell type, with the total protein identification count being CD8. + T cells 2595, neutrophils 3102, macrophages 5002. Detailed functional protein enhancement performance is shown in Table 3. Table 3 Comparison of DIA data for three types of cells

[0109] ⑤ False discovery rate of the customized spectral library of this invention:

[0110] The results are far below the 1% FDR control standard generally accepted in the field of proteomics, which fully demonstrates that the present invention improves the recognition ability of key functional proteins without introducing additional false positive burden, thus ensuring high accuracy and high reliability of the recognition results.

[0111] In summary, this embodiment verifies the significant advantages of the customized spectral library system of the present invention in the task of identifying functional proteins in immune cells. Through mechanisms such as priority retention of functional proteins, expression scoring-driven compression, and structure deduction, the system successfully improves the breadth and depth of identification of key functional proteins without relying on experimental construction of a reference library, demonstrating its practical application value.

[0112] Example 3: Application in the identification of tumor-associated macrophage (TAM) subsets

[0113] This embodiment aims to verify the applicability and performance advantages of the spectral library construction system described in this invention in the analysis of complex immune cell heterogeneity, with a particular focus on the identification of single-cell subsets of tumor-associated macrophages (TAMs) in a mouse melanoma model.

[0114] (1) Experimental background and target cell type

[0115] Using the B16-OVA mouse melanoma model, CD11 was isolated from the tumor tissue. / F4 / 8 / CD20 The TAM population was sorted using an image recognition-based single-cell sorting system (CellenONE platform). The total number of single-cell samples was:

[0116] All cells underwent sequential lysis, trypsin digestion, and DIA mass spectrometry data acquisition.

[0117] (2) Experimental results

[0118] ① The basic information of the customized spectral library Lib_TAM is shown in Table 4: Table 4 Basic Information on Customized Spectra

[0119] ②Please refer to Figure 4 a. Based on the protein expression matrix of the above 282 TAM single cells, PCA dimensionality reduction and UMAP clustering analysis were performed using the Seurat package to identify four functionally distinct TAM subgroups, which were named TAM.1, TAM.2, TAM.3, and TAM.4, respectively.

[0120] ③ Please refer to Figure 4 Figure b shows the distribution of protein identification counts for the four TAM subpopulations at the single-cell level. The results show that the overall distribution of protein identification counts for each subpopulation from TAM.1 to TAM.4 is similar, without any systematic differences, indicating that the formation of different subpopulations is not due to deviations in protein identification depth.

[0121] ③ Please refer to Figure 4c. To further evaluate the ability of the method of this invention to resolve cellular functional heterogeneity, a systematic comparison was conducted on the expression distribution of membrane proteins, transcription factors, kinases, and ubiquitin ligases in four TAM subpopulations. The results showed that, without the application of the customized spectral library of this invention (i.e., the traditional library-free method), the aforementioned key regulatory proteins often exhibited low detection rates and discontinuous expression signals at the single-cell level, making it difficult to demonstrate clear subpopulation specificity. However, using the method of this invention significantly enhanced the overall coverage of functional proteins, not only increasing the detection frequency but also enhancing the expression differentiation among different subpopulations, highlighting the sensitivity and stability of this invention in identifying low-abundance, function-guided proteins. Although the TAM subpopulations share common basic phagocytic characteristics, each subpopulation exhibits clear molecular specialization. For example, TAM.1 is enriched in proteins related to membrane transport, such as Ehd1, Ap2b1, Ap2a2, Ap3b1, and Rab5a, which are involved in vesicle transport. Factors regulating cytoskeleton dynamics (Cdc42, Iqgap2) and immune recognition-related molecules (Tlr2, Itgb2, Irak4, Mapk14, Ripk1) are also significantly upregulated, highlighting their role in antigen uptake and pro-inflammatory signaling. TAM.2 is characterized by high expression of chemokine receptors and adhesion-related molecules (Ccr5, Gnaq, Itgal, Hpse), while the upregulated multiple kinases and ubiquitin ligases (Map2k4, Hck, Stub1, Usp8, Dnmt1) are closely related to signal regulation and protein homeostasis. TAM.3 is extensively enriched in immune regulation and metabolic reprogramming, including Apoa1, Tfrc, and Igf2r involved in lipid and iron metabolism, kinases Gsk3b and Cdk4, and transcriptional regulators Stat5b and Stat6. Multiple energy metabolism and intracellular transport-related enzymes are also upregulated, consistent with its metabolically active functional state. In contrast, the proteomic signature of TAM.4 is relatively sparse, with only a few marker molecules (such as Asap2 and Prox2), suggesting a low degree of functional specialization or activation. In summary, the method of this invention, by enhancing the recognition ability of low-abundance regulatory proteins, successfully revealed the molecular differences of TAM subsets in immune activation, metabolic reprogramming, and self-renewal potential, laying a proteomic foundation for a deeper understanding of the functional state of macrophages in the tumor microenvironment, and has broad application prospects and translational value.

[0122] Example 4: Construction of a Customized Spectral Library System Based on an Intelligent Computing Platform

[0123] This embodiment provides a scalable intelligent computing platform that integrates multiple algorithms. This platform is used to build a customized spectral library construction system and supports peptide selection, functional enhancement, expression screening, noise reduction and compression, quality control, personalized construction, and visualization analysis.

[0124] The system can be deployed in a cloud platform environment, possessing excellent scalability and computing performance. The platform adopts a modular, layered architecture design, decoupling the front-end user interface from the back-end business processing logic, enabling flexible invocation of functional components and efficient system operation and maintenance management. The front-end interface is based on a dynamic responsive design, supporting interactive parameter input and result display, significantly improving the intuitiveness of user operation and system response speed. The back-end service integrates a high-performance task scheduling mechanism and concurrent request processing module, coupled with automated load balancing strategies, ensuring stable system operation under high access intensity. The platform's data layer is built on a highly reliable database management system, supporting structured storage, rapid indexing, and cross-module invocation of large-scale spectral data and analysis results. Users can access the system's core functional modules through a unified entry point, conveniently completing parameter configuration, data submission, and result acquisition, meeting the application needs of multi-scenario single-cell proteomics analysis.

[0125] This application also provides an electronic device, including: a processor, and a memory coupled to the processor, the memory being used to store a computer program; the processor being used to execute the computer program stored in the memory, so that the electronic device performs the method as described in any of the above embodiments.

[0126] Electronic devices can be computing devices such as desktop computers, laptops, handheld computers, and cloud servers. These electronic devices may include, but are not limited to, processors and memory.

[0127] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting various parts of the device via various interfaces and lines.

[0128] The memory can be used to store the computer program, and the processor implements various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.

[0129] The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc.; the data storage area may store data created based on the use of the mobile phone, etc. Furthermore, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0130] This application embodiment also provides a storage medium, which is a computer-readable storage medium. The computer program is stored in the computer-readable storage medium, and when executed by a processor, the computer program can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0131] This application also provides a computer program product, including: a computer program or instructions that, when the computer program or instructions are run on a computer, cause the computer to perform any of the above possible implementation methods.

[0132] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.

[0133] Although the present invention has been described in detail with general description and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention are within the scope of protection claimed by the present invention.

Claims

1. A method for constructing a targeted, lightweight spectral library for single-cell proteomics data, characterized in that, The method includes: S1, extraction of high-confidence peptide spectra; S2. Construction of a library for preferential importation of functional proteins and target-guided spectral deduction; S3, Expressing score-driven item compression; S4. Structural noise reduction and spectral optimization; S5, standardized output format compatible with multiple proteomics analysis platforms; S6, Decoy library generation and false discovery rate (FDR) control.

2. The method according to claim 1, characterized in that, Step S1 involves single-cell sorting, lysis, and enzymatic digestion of the target cells. Raw DIA mass spectrometry data of single cells are obtained using liquid chromatography-high resolution mass spectrometry (LC-HDMS). Based on more than 20 raw DIA mass spectrometry data of each cell type, characteristic peaks are extracted, retention times are aligned, and theoretical sequences are compared using a public search platform. Peptides that appear in most samples and have stable signals are generated based on the feature extraction module. Their parent ion mass-to-nucleus ratio, retention time, fragment distribution, and relative intensity characteristics are extracted to construct a high-confidence seed set.

3. The method according to claim 2, characterized in that, Step S2 involves constructing a functional protein set G_func based on a public functional annotation database. The functional protein set includes at least membrane receptors, ligands, protein kinases, transcription factors, and ubiquitin ligases. The entries for the functional proteins are limited to UniProt reviewed entries. During the construction process, the spectra of the peptides corresponding to the above categories are preferentially retained, even if they do not appear in S1.

4. The method according to claim 3, characterized in that, Step S3 involves introducing a peptide expression scoring mechanism based on single-cell protein expression data, and performing content compression and item screening according to the score. The single-cell protein expression data comes from public databases, including SPDB and Slavov laboratory data. Based on public or experimental expression data of the target cell type, a probability value is assigned to each candidate peptide spectrum. If the value is lower than a set threshold and the structure score is also low, it is rejected.

5. The method according to claim 4, characterized in that, Step S4 involves setting peptide length, fragment ion quantity, and m / z region density thresholds to perform structural noise reduction processing, taking into account the characteristics of few ion fragments and strong background signals in single-cell mass spectrometry data. The optimization of the structure includes: setting the peptide length to 7-25 amino acids and the number of fragment ions not exceeding 12; removing repetitive peptides with excessive frequency in the m / z concentration region to reduce false matching.

6. The method according to claim 5, characterized in that, Step S5 outputs a spectral library as a .tsv file, with fields including: Precursor m / z, Charge, Fragment type, Fragment ion m / z, RelativeIntensity, iRT calibration value, and Protein name.

7. The method according to claim 6, characterized in that, Step S6 involves constructing a pseudospectral set P. decoy For each target peptide p, a pseudo peptide p is generated by sequence inversion. With the parent ion m / z remaining constant, the corresponding RT and MS2 spectra are derived from the single-cell type model after S2 transfer learning. The resulting pseudo-spectrum and target spectrum are used together to search and analyze, and are used to construct a target-non-target distribution model, thereby completing the FDR curve fitting and confidence threshold setting.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it can implement the steps of the method according to any one of claims 1 to 7.

9. An electronic device, comprising: A processor, and a memory coupled to the processor, the memory being used to store computer programs; The processor is configured to execute the computer program stored in the memory, so that the electronic device performs the method as described in any one of claims 1 to 7.

10. The application of the method according to any one of claims 1 to 7 in the prediction of key functional proteins of immune cells, characterized in that, The immune cells include lymphocytes, neutrophils, and macrophages; the key functional proteins include single-cell level identification of membrane receptor ligands, kinases, transcription factors, and ubiquitin-regulated proteins.

Citation Information

Patent Citations

  • Method for improving single-cell proteome identification coverage rate based on deep learning

    CN114639444A

  • Novel proteome identification method based on spectrogram merging strategy

    CN117037900A