Raw spectrum processing method and device, computer device and storage medium

By using a deep learning model to automatically process chromatographic peaks in proteomics, generate spectral feature matrices, and extract peptide feature vectors, the inaccurate and inefficient analytical results caused by manual intervention in existing technologies are solved, achieving efficient and accurate spectral analysis.

CN114283884BActive Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-08-17
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing LC-MS/MS-based proteomics analysis methods rely on human intervention, resulting in highly subjective, inaccurate, and inefficient analysis results.

Method used

An automated method based on a deep learning model is adopted to extract chromatographic peaks from a target peptide library, generate a spectral feature matrix, and perform feature extraction to determine the classification information of the original spectrum, thereby achieving automated feature extraction.

Benefits of technology

It improves the accuracy and efficiency of proteomics analysis, reduces the impact of human intervention, and significantly enhances the reliability of analytical results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283884B_ABST
    Figure CN114283884B_ABST
Patent Text Reader

Abstract

The application provides a raw spectrum processing method and device, computer equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: extracting a plurality of chromatographic peaks from a raw spectrum of proteomics based on a target peptide library; determining a plurality of spectral line feature matrices based on the plurality of chromatographic peaks; performing feature extraction on the plurality of spectral line feature matrices respectively based on a deep learning model to obtain a plurality of peptide feature vectors; and determining typing information of the raw spectrum based on the plurality of peptide feature vectors, wherein the typing information is used to indicate the probability of the raw spectrum corresponding to a target disease type. The above technical solution can automatically extract features in the raw spectrum of proteomics without human intervention, the extracted features are more accurate and the extraction efficiency is high, thereby significantly improving the accuracy and analysis efficiency of the analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, and storage medium for processing raw spectra. Background Technology

[0002] In the field of proteomics research, LC-MS / MS (chromatography-mass spectrometry-mass spectrometry) based measurement methods have become the mainstream measurement methods due to their high precision, high stability, and high throughput, and play a vital role in clinical diagnosis, drug development, and disease treatment.

[0003] Currently, peptide search engine-based analysis methods are commonly used to analyze the raw spectra acquired by LC-MS / MS acquisition methods, and then the determination is completed based on the analysis results.

[0004] However, the above-mentioned analytical methods mainly rely on manual labeling of chromatographic peaks and manually designed feature extraction algorithms, which require a lot of human intervention, resulting in subjective and unstable analytical results, affecting the accuracy of the analytical results and causing low analytical efficiency. Summary of the Invention

[0005] This application provides a method, apparatus, computer device, and storage medium for processing raw proteomics spectra. It automatically extracts features from raw proteomics spectra without human intervention, resulting in more accurate and efficient feature extraction, thus significantly improving the accuracy and efficiency of the analysis results. The technical solution is as follows:

[0006] On the one hand, a method for processing raw spectra is provided, the method comprising:

[0007] Multiple chromatographic peaks are extracted from the original proteomics spectrum based on the target peptide library, wherein the target peptide library is associated with the original spectrum, and one peptide in the target peptide library corresponds to at least one chromatographic peak.

[0008] Based on the multiple chromatographic peaks, multiple spectral feature matrices are determined, and each spectral feature matrix represents the chromatographic peak characteristics of a peptide.

[0009] Based on a deep learning model, feature extraction is performed on the multiple spectral feature matrices to obtain multiple peptide feature vectors, which are used to represent the features of peptide segments.

[0010] Based on the multiple peptide feature vectors, the genotyping information of the original spectrum is determined, and the genotyping information is used to indicate the probability of the original spectrum corresponding to the target disease type.

[0011] On the other hand, a raw spectrum processing apparatus is provided, the apparatus comprising:

[0012] A chromatographic peak extraction module is used to extract multiple chromatographic peaks from the original spectrum of proteomics based on a target peptide library, wherein the target peptide library is associated with the original spectrum, and one peptide in the target peptide library corresponds to at least one chromatographic peak.

[0013] The matrix determination module is used to determine multiple spectral feature matrices based on the multiple chromatographic peaks, where each spectral feature matrix represents the chromatographic peak characteristics of a peptide.

[0014] The feature extraction module is used to extract features from the multiple spectral feature matrices based on a deep learning model to obtain multiple peptide feature vectors, which are used to represent the features of peptide segments.

[0015] The information determination module is used to determine the genotyping information of the original spectrum based on the multiple peptide feature vectors, wherein the genotyping information is used to indicate the probability of the original spectrum corresponding to a target disease type.

[0016] In some embodiments, the matrix determination module includes:

[0017] A spectral line adjustment unit is used to adjust the spectral line lengths of the multiple chromatographic peaks to a target length;

[0018] The spectral line stacking unit is used to stack the spectral lines of the chromatographic peaks corresponding to any peptide segment to obtain the spectral line feature matrix corresponding to the peptide segment.

[0019] In some embodiments, the spectral stacking unit is used to stack the spectral lines of the chromatographic peaks corresponding to the peptide to obtain an intermediate matrix; and to stack at least two intermediate matrices in the channel dimension to obtain the spectral feature matrix corresponding to the peptide, wherein one channel of the spectral feature matrix corresponds to one intermediate matrix.

[0020] In some embodiments, the apparatus further includes:

[0021] The sorting module is used to sort the chromatographic peaks corresponding to library ions based on the Pearson correlation coefficient, wherein the library ions are the peptide ions corresponding to peptide segments in the target peptide library;

[0022] The sorting module is also used to sort the chromatographic peaks corresponding to theoretical ions based on the extraction order, wherein the theoretical ions are peptide ions corresponding to peptides not included in the target peptide library.

[0023] In some embodiments, the apparatus further includes:

[0024] A smoothing module is used to smooth the multiple chromatographic peaks.

[0025] In some embodiments, the step of determining the genotyping information of the original spectrum based on the plurality of peptide feature vectors is implemented based on a genotyping model, which is used to determine the probability of the target disease type based on the peptide feature vectors.

[0026] In some embodiments, the classification model is trained based on the following method:

[0027] Multiple sample chromatographic peaks are extracted from the original sample chromatograms based on a sample peptide library, wherein the sample peptide library is associated with the original sample chromatograms, and one sample peptide in the sample peptide library corresponds to at least one sample chromatographic peak.

[0028] Based on the multiple sample chromatographic peaks, multiple sample spectral matrixes are determined, and each sample spectral matrix represents the chromatographic peak characteristics of a sample peptide.

[0029] Based on the deep learning model, feature extraction is performed on the multiple sample spectral matrixes to obtain multiple sample peptide feature vectors, which are used to represent the features of sample peptide segments.

[0030] The typing model is trained based on the peptide feature vectors of the multiple samples.

[0031] In some embodiments, the step of training the model based on the plurality of sample peptide feature vectors to obtain the trained typing model includes:

[0032] The multiple sample peptide feature vectors are concatenated to obtain a sample feature vector, wherein the number of elements in the sample feature vector is the sum of the number of elements in the multiple sample peptide feature vectors;

[0033] The sample feature vector is divided into five equal parts, and the hyperparameters of the fractal model are adjusted based on five-fold cross-validation.

[0034] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the operations performed in the original spectrum processing method in the embodiments of this application.

[0035] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to perform the operations performed as in the original spectrum processing method in the embodiments of this application.

[0036] On the other hand, a computer program product is provided, comprising computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the raw spectrum processing method provided in various alternative implementations of the above aspects.

[0037] The beneficial effects of the technical solutions provided in this application are:

[0038] In this application embodiment, a novel raw spectrum processing method is provided, which can automatically extract features from raw proteomics spectra without human intervention. The extracted features are more accurate and efficient, thereby significantly improving the accuracy and efficiency of the analysis results. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a schematic diagram of the implementation environment of a raw spectrum processing method provided in the embodiments of this application;

[0041] Figure 2 This is a flowchart of a raw spectrum processing method provided according to an embodiment of this application;

[0042] Figure 3 This is a flowchart of another raw spectrum processing method provided according to an embodiment of this application;

[0043] Figure 4 This is a schematic diagram of chromatographic peak extraction from an original spectrum according to an embodiment of this application;

[0044] Figure 5 This is a flowchart illustrating another method for processing raw spectra according to an embodiment of this application;

[0045] Figure 6 This is a block diagram of a raw spectrum processing apparatus provided according to an embodiment of this application;

[0046] Figure 7 This is a block diagram of another raw spectrum processing apparatus provided according to an embodiment of this application;

[0047] Figure 8This is a schematic diagram of the structure of a server according to an embodiment of this application. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0049] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0050] In this application, the term "at least one" means one or more, and "multiple" means two or more.

[0051] The following is an explanation of the terms used in this application.

[0052] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0053] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation. The solutions provided in this application's embodiments involve machine learning and deep learning technologies in artificial intelligence; details are provided in the embodiments.

[0054] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0055] Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer. In this application's embodiments, the original spectra, peptide libraries, extracted chromatographic peaks, peptide feature vectors, deep learning models, and classification models can all be stored in the blockchain system.

[0056] Dwell time is the time each ion spends in the scan. The sum of the dwell times of all particles is the time taken for one scan, which is the time interval between two scan points on the chromatogram.

[0057] Biomarkers are biochemical indicators that can mark changes or potential changes in the structure or function of systems, organs, tissues, cells, and subcellular structures. They have a very wide range of applications. Biomarkers can be used for disease diagnosis, disease staging, or to evaluate the safety and efficacy of new drugs or therapies in target populations. Functionally, they are generally classified into: exposure biomarkers; effect biomarkers; and susceptibility biomarkers.

[0058] The original spectrum processing method provided in this application embodiment can be executed by a computer device. Figure 1 This is a schematic diagram illustrating the implementation environment of a raw spectrum processing method provided in an embodiment of this application. See also... Figure 1 The implementation environment includes terminal 101 and server 102.

[0059] Terminal 101 and server 102 can be connected directly or indirectly via wired or wireless communication, and this application does not impose any restrictions on this.

[0060] In some embodiments, server 102 is a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 102 is used to process the raw proteomics spectra to determine whether the raw spectra correspond to the target disease type.

[0061] In some embodiments, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.

[0062] In this implementation environment, the computer device is server 102. Server 102 acquires the raw proteomics spectrum uploaded by terminal 101, and then, based on the raw spectrum processing method provided in this application embodiment, extracts chromatographic peaks from the raw spectrum, converts the extracted chromatographic peaks into a structured spectral feature matrix, and then extracts features from the spectral feature matrix. Based on the extracted peptide feature vectors, the probability of the raw spectrum corresponding to a target disease type is determined. Server 102 can then return this probability to terminal 101, thereby assisting medical personnel in classifying the raw proteomics spectrum for disease.

[0063] In some embodiments, terminal 101 may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited thereto. Those skilled in the art will understand that the number of terminals 101 may be more or less. For example, there may be only one terminal 101, or there may be dozens or hundreds of terminals 101, or even more. This application does not limit the number or type of terminals in its embodiments.

[0064] Figure 2 This is a flowchart of a raw spectrum processing method provided according to an embodiment of this application, such as... Figure 2 As shown in the embodiment of this application, the method is described using an example executed by a server. The original spectrum processing method includes the following steps:

[0065] 201. The server extracts multiple chromatographic peaks from the original proteomics spectrum based on the target peptide library, which is associated with the original spectrum. One peptide in the target peptide library corresponds to at least one chromatographic peak.

[0066] In this embodiment, the server acquires the raw spectrum to be processed and the target peptide library associated with the raw spectrum. The raw spectrum is a raw spectrum from proteomics. The target peptide library is one of two types: a pre-constructed prior peptide library or a peptide library constructed based on deconvolution algorithms and search engine peptide libraries. The peptide library includes information such as the mass-to-charge ratio, charge, sequence, daughter ion composition, and standard residence time of peptides. Based on the information in the target peptide library, the server can extract multiple chromatographic peaks from the raw spectrum, with one peptide corresponding to one or more chromatographic peaks.

[0067] 202. Based on these multiple chromatographic peaks, the server determines multiple spectral feature matrices. Each spectral feature matrix represents the chromatographic peak characteristics of a peptide.

[0068] In this embodiment, the server converts multiple chromatographic peaks into multiple structured spectral feature matrices. The number of spectral feature matrices is no greater than the number of chromatographic peaks; that is, a spectral feature matrix is ​​obtained based on one or more chromatographic peaks, so that a single spectral feature matrix can represent all the chromatographic peak features of a peptide.

[0069] 203. The server uses a deep learning model to extract features from the multiple spectral feature matrices to obtain multiple peptide feature vectors, which are used to represent the features of peptide segments.

[0070] In this embodiment, the server can perform feature extraction based on a deep learning model pre-trained using a deep learning algorithm. The peptide feature vector extracted by the server based on the deep learning model can more effectively represent spectral features, improving the sensitivity of peptide detection. In some embodiments, the deep learning model is a ResNet18 model pre-trained on the ImageNet dataset, or other deep learning models, including but not limited to: AlexNet, VGG (Visual Geometry Group Network), Inception, SqueezeNet, EfficientNet, PReLU-nets (Parametric Rectified Linear Unit Nets), PreActResNet, Stochastic Depth ResNet, WRN (Wide Residual Networks), GeNet (GPU-Efficient Networks), MetaQNN (Meta Q-Leaning Neural Networks), PyramidNet, DenseNet, FractalNet, Residual Attention Network, Xception, PolyNet, and DPN (Dual Path Network). Networks (dual-path network) model, MobileNet model, NasNet (Neural Architecture Search Networks) model, etc.

[0071] 204. Based on the multiple peptide feature vectors, the server determines the typing information of the original spectrum, which is used to indicate the probability of the original spectrum corresponding to the target disease type.

[0072] In this embodiment, the server can concatenate multiple peptide feature vectors into a single vector, and then determine the probability of the original spectrum corresponding to a target disease type based on the concatenated vector, thus obtaining typing information. The concatenation order of the peptide feature vectors is determined based on the order of the chromatographic peaks corresponding to the peptide segments. Since each peptide feature vector has the same number of elements, the number of elements in the concatenated vector is M×N, where M represents the number of elements in a peptide feature vector, N represents the number of peptide segments, and M and N are positive integers. The target disease type is the disease type to be analyzed. This typing information can be used to assist medical personnel in performing disease typing based on the original spectrum.

[0073] For example, let's illustrate this by concatenating the first and last features of three peptides. One peptide feature vector is [1,1,1,1,1,1], another is [2,2,2,2,2,2], and the third is [3,3,3,3,3,3]. The server concatenates these three peptide feature vectors to obtain [1,1,1,1,1,1,1,2,2,2,2,2,2,3,3,3,3,3,3].

[0074] The solution provided in this application provides a novel method for processing raw spectra. It can automatically extract features from raw proteomics spectra without human intervention. The extracted features are more accurate and efficient, thereby significantly improving the accuracy and efficiency of the analysis results.

[0075] The above Figure 2 A flowchart of the raw spectrum processing method provided in an embodiment of this application is illustrated. The following description further illustrates this raw spectrum processing method based on an application scenario. See also... Figure 3 As shown, Figure 3 This is a flowchart of another raw spectrum processing method provided according to an embodiment of this application, such as... Figure 3 As shown in the embodiment of this application, the method of performing disease classification on the original spectrum by the server is used as an example for illustration. The original spectrum processing method includes the following steps:

[0076] 301. The server retrieves the raw spectra of proteomics.

[0077] In this embodiment of the application, the original proteomics spectrum is obtained by the server acquiring the collected original spectrum, or by the server acquiring it.

[0078] In some embodiments, the acquisition methods of the original spectrum include the following: SRM (Selected Reaction Monitoring), PRM (Parallel Reaction Monitoring), TMT-SRM (Tandem Mass Tag-Selected Reaction Monitoring), TMT-PRM (Tandem Mass Tag-Parallel Reaction Monitoring), and DIA (Data-Independent Acquisition), etc., and the embodiments of this application do not limit the methods.

[0079] In some embodiments, the server is a node in a blockchain system. If the original spectrum of the proteomics is stored in the blockchain system, the server can retrieve the original spectrum from local storage. If the original spectrum of the proteomics is not stored in the blockchain system, the server can upload the original spectrum to the blockchain after any device uploads the original spectrum, thereby obtaining the original spectrum.

[0080] 302. The server retrieves the target peptide library associated with the original spectrum, which includes multiple peptide segments.

[0081] In this embodiment, after obtaining the original spectrum, the server can acquire the target peptide library associated with the original spectrum based on the associated peptide library information of the original spectrum. This associated peptide library information indicates the peptide library associated with the original spectrum. If the associated peptide library information of the original spectrum does not indicate a peptide library associated with the original spectrum, the server can construct the target peptide library associated with the original spectrum based on the original spectrum. The target peptide library includes information such as the mass-to-charge ratio, charge, sequence, daughter ion composition, and standard residence time of the peptide segments.

[0082] In some embodiments, for raw spectra acquired using SRM or PRM methods, the associated peptide library information of the raw spectra indicates a prior peptide library provided by medical personnel. After acquiring the raw spectra, medical personnel select this prior peptide library for the raw spectra. The associated peptide library information stores the association between the prior peptide library and the raw spectra, and the server retrieves the prior peptide library based on this information. For raw spectra acquired using DIA methods, if the associated peptide library information of the raw spectra indicates a prior peptide library provided by medical personnel, the server retrieves the prior peptide library based on this information. If the raw spectra is not associated with a peptide library, the server can construct a peptide library for the raw spectra based on a deconvolution algorithm combined with a search engine algorithm, and use the newly constructed peptide library as the target peptide library associated with the raw spectra.

[0083] In some embodiments, the server is a node in a blockchain system. If the target peptide library is already stored in the blockchain system, the server can retrieve the target peptide library from local storage. If the target peptide library is not stored in the blockchain system, the server can upload the target peptide library to the blockchain after any device uploads it, thereby obtaining the target peptide library.

[0084] 303. The server extracts multiple chromatographic peaks from the original spectrum based on the target peptide library. One peptide in the target peptide library corresponds to at least one chromatographic peak.

[0085] In this embodiment, the server extracts chromatographic peaks corresponding to each peptide segment in the target peptide library from the original spectrum based on information in the target peptide library, resulting in multiple chromatographic peaks. The server extracts different chromatographic peaks depending on the acquisition method: for original spectra acquired using SRM or PRM methods, the server extracts chromatographic peaks of the library precursor ion, library daughter ion, theoretical daughter ion, and precursor ion isotope; for original spectra acquired using TMT-SRM or TMT-PRM methods, the server extracts chromatographic peaks of the precursor ion and TMT reporter ion; for original spectra acquired using DIA methods, the server extracts chromatographic peaks of the library precursor ion, library daughter ion, theoretical daughter ion, precursor ion isotope, daughter ion isotope, precursor ion decoy light isotope, daughter ion decoy light isotope, and unfractured precursor ion.

[0086] In some embodiments, after extracting chromatographic peaks, the server can also determine whether the residence time corresponding to each peptide is accurate, thereby determining the actual residence time of the peptide. Accordingly, for raw spectra acquired using SRM, PRM, TMT-SRM, or TMT-PRM methods, the server determines the actual residence time of each peptide based on the chromatographic peaks of the primary mass spectrometry signal; for raw spectra acquired using DIA, the server randomly samples peptides from the target peptide library, and fits a linear regression model based on the standardized residence time in the sampled peptide library and the actual residence time obtained by the chromatographic peak-finding algorithm. Based on this linear regression model, the actual residence time of all peptides in the target peptide library is determined. This peak-finding algorithm identifies the portion of the curve above the target value as a chromatographic peak. Of course, the server can also use other peak-finding algorithms, such as peak-finding methods based on the second derivative or peak-finding methods that find the maximum value among adjacent points; this application does not limit this approach.

[0087] 304. The server sorts the multiple chromatographic peaks.

[0088] In this embodiment, the multiple chromatographic peaks include both those corresponding to library ions and those corresponding to theoretical ions. The library ions are peptide ions corresponding to peptides in the target peptide library, while the theoretical ions are peptide ions corresponding to peptides not included in the target peptide library. It should be noted that, since ions are also a form of peptide, library ions can also represent ions in the target peptide library, and theoretical ions represent ions not in the target peptide library.

[0089] In some embodiments, the server first sorts the chromatographic peaks corresponding to the library ions based on the Pearson correlation coefficient. Then, it sorts the chromatographic peaks corresponding to the theoretical ions based on the extraction order. That is, the chromatographic peaks corresponding to the library ions are in the order of the theoretical ions.

[0090] For example, for any chromatographic peak corresponding to a library ion, if the Pearson correlation coefficient between the peak and other chromatographic peaks is the largest, it means that the peak has the highest correlation with other chromatographic peaks, and thus the peak has the highest accuracy. The server will then prioritize this peak.

[0091] 305. The server smooths multiple chromatographic peaks.

[0092] In this embodiment, the server smooths the sorted chromatographic peaks using a Gaussian smoothing algorithm. Of course, the server can also smooth chromatographic peaks using other algorithms, such as neighborhood averaging and median filtering; this embodiment does not limit this approach. Smoothing significantly filters noise from chromatographic peaks, improving the accuracy of feature extraction.

[0093] It should be noted that the server can also execute step 305 directly after executing step 303, that is, skip the sorting step, in order to shorten the overall processing time; or the server can execute step 306 directly after executing step 304, that is, skip the smoothing step, in order to shorten the overall processing time; or the server can execute step 306 directly after executing step 303, that is, skip both the sorting and smoothing steps, in order to shorten the overall processing time.

[0094] 306. The server adjusts the spectral lengths of multiple chromatographic peaks to the target length.

[0095] In this embodiment, the server can fill in shorter spectral lines and truncate longer spectral lines to ensure that the spectral length of each chromatographic peak is the target length. For chromatographic peaks with a spectral length shorter than the target length, the server fills in zero values ​​on both sides of the peak's spectral line, using the midpoint of the residence time as a reference, so that the adjusted spectral length of the peak is the target length. For chromatographic peaks with a spectral length longer than the target length, the server truncates the spectral lines on both sides of the peak, using the midpoint of the residence time as a reference, so that the adjusted spectral length of the peak is the target length. This target length can be set according to actual needs; this embodiment does not limit the target length.

[0096] 307. For any peptide segment, the server stacks the spectral lines of the corresponding chromatographic peaks to obtain the spectral feature matrix of the peptide segment.

[0097] In this embodiment, the server stacks the adjusted chromatographic peaks into a spectral feature matrix. Taking any peptide as an example, the peptide corresponds to one or more chromatographic peaks. The server stacks the spectral lines of the chromatographic peaks corresponding to the peptide into a matrix to obtain the spectral feature matrix corresponding to the peptide. Based on the original spectrum, the server can obtain N spectral feature matrices, where N represents the number of peptides contained in the original spectrum, and N is a positive integer.

[0098] In some embodiments, a chromatographic peak includes one or more channels. When a chromatographic peak includes two or more channels, the server stacks the spectral lines of the chromatographic peak corresponding to the peptide to obtain an intermediate matrix. Then, the server stacks at least two intermediate matrices along the channel dimension to obtain a spectral feature matrix corresponding to the peptide. One channel of the spectral feature matrix corresponds to one intermediate matrix, and the at least two intermediate matrices are identical. The number of rows in the intermediate matrix is ​​the number of spectral lines of the chromatographic peak corresponding to the peptide.

[0099] For example, the original spectrum includes three channels: R, G, and B. For any peptide, the server stacks the spectral lines of the corresponding chromatographic peaks to obtain an intermediate matrix. Then, the server copies the intermediate matrix to obtain three intermediate matrices, one intermediate matrix for each channel. Finally, the server stacks the three intermediate matrices along the channel dimension to obtain the spectral feature matrix corresponding to the peptide.

[0100] 308. The server extracts features from multiple spectral feature matrices based on a deep learning model to obtain multiple peptide feature vectors, which are used to represent the features of peptide segments.

[0101] In this embodiment of the application, the server can extract features from multiple spectral feature matrices obtained based on a pre-trained deep learning model, thereby obtaining the peptide feature vector corresponding to each peptide segment.

[0102] In some embodiments, the deep learning model has fully connected layers with fixed weights, causing the deep model to output a vector of fixed dimension. This application does not limit the value of the fixed weights; in this application, the fixed weight is described as 1. The fixed dimension is 1-dimensional by default, so the deep learning model outputs a 1-dimensional vector, i.e., a scalar.

[0103] It should be noted that the peptide feature vector extracted by the server based on the deep learning model can also be called the protein feature vector.

[0104] It should also be noted that the deep learning model is any of the deep learning models shown in step 203 above, and will not be described again here.

[0105] 309. The server determines the typing information of the original spectrum based on multiple peptide feature vectors. This typing information is used to indicate the probability of the original spectrum corresponding to the target disease type.

[0106] In this embodiment of the application, this step is the same as step 204 above, and will not be repeated here.

[0107] In some embodiments, this step can be implemented based on a typing model, which is used to determine the probability of a target disease type based on peptide feature vectors. The typing model is a classification model. In this embodiment, the typing model is a binary classification model. The typing model can be obtained through multiple iterations of training. Taking one iteration as an example, the one-step iteration process of the typing model includes steps (1) to (4).

[0108] (1) The server extracts multiple sample chromatographic peaks from the original sample spectrum of proteomics based on the sample peptide library. The sample peptide library is associated with the original sample spectrum, and one sample peptide in the sample peptide library corresponds to at least one sample chromatographic peak.

[0109] In this embodiment, the method for obtaining the original sample spectrum is the same as the method for obtaining the original spectrum by the server in step 301 above, and will not be repeated here. The method for obtaining the sample peptide library is the same as the method for obtaining the target peptide library by the server in step 302 above, and will not be repeated here. In addition, when obtaining the original sample spectrum, the server will also obtain the corresponding sample category label. The sample category label is used to indicate the disease type corresponding to the original sample spectrum. The sample category label is marked by medical personnel or domain experts.

[0110] The process of extracting multiple sample chromatographic peaks in this step is the same as the process of extracting chromatographic peaks in step 303 above, and will not be repeated here.

[0111] (2) Based on the multiple sample chromatographic peaks, the server determines multiple sample spectral matrixes. Each sample spectral matrix represents the chromatographic peak characteristics of a sample peptide.

[0112] In the embodiments of this application, the process of determining the sample spectral line matrix in this step is described in steps 304-307, and will not be repeated here.

[0113] To make steps (1) and (2) above easier to understand, see [link to relevant documentation]. Figure 4 As shown, Figure 4 This is a schematic diagram of chromatographic peak extraction from an original spectrum according to an embodiment of this application. For example... Figure 4 Using the original sample spectrum and a sample peptide library as training samples, the server extracts chromatographic peaks from the original sample spectrum based on the sample peptide library, obtaining multiple sample chromatographic peaks. These peaks are then stacked according to peptide dimensions to obtain a sample spectral matrix. Multiple sample spectral matrices originating from the same original sample spectrum constitute a spectral data matrix. One spectral data matrix corresponds to one original sample spectrum. Rows in the spectral data matrix correspond to the peptide dimension (one row represents one peptide), and columns correspond to the residence time dimension. Figure 5 In this context, a plus sign indicates a positive sample, and a minus sign indicates a negative sample.

[0114] (3) The server extracts features from the multiple sample spectral matrixes based on a deep learning model to obtain multiple sample peptide feature vectors, which are used to represent the features of the sample peptide segments.

[0115] In this embodiment of the application, the process of feature extraction of the sample spectral matrix in this step is the same as the process of feature extraction of the spectral feature matrix in step 308, and will not be repeated here.

[0116] (4) The server trains the typing model based on the peptide feature vectors of multiple samples.

[0117] In this embodiment, the server constructs a genotyping model based on the XGBoost model, and then trains the model based on the aforementioned multiple sample peptide feature vectors, adjusting the hyperparameters of the genotyping model. This genotyping model constructed based on the XGBoost model is a classification model; in this embodiment, a binary classification model is used as an example for illustration.

[0118] In some embodiments, the step of the server training a model based on the multiple sample peptide feature vectors to obtain a trained genotyping model includes: the server concatenating the multiple sample peptide feature vectors to obtain a sample feature vector, the number of elements of which is the sum of the number of elements of the multiple sample peptide feature vectors; then the server dividing the sample feature vector into five equal parts and adjusting the hyperparameters of the genotyping model based on five-fold cross-validation. The above-mentioned concatenation method of the sample peptide feature vectors is end-to-end concatenation, and the server determines the concatenation order according to the order of the sample chromatographic peaks corresponding to the sample peptide segments.

[0119] In some embodiments, the server can filter out biomarkers that contribute significantly to the model output based on the feature importance of the XGBoost model (i.e., the genotyping model), thereby making the output of the genotyping model interpretable.

[0120] It should be noted that steps 301-309 above are exemplified by processing the first-order mass spectrometry signal in the original spectrum. Correspondingly, the server also processes the second-order and third-order mass spectrometry signals simultaneously. See [link to documentation]. Figure 5 As shown, Figure 5 This is a flowchart illustrating another method for processing raw spectra according to an embodiment of this application. The steps before the server obtains the spectral feature matrix are the same as those for primary mass spectrometry signals, and will not be repeated here. During the training of the genotyping model, the server simultaneously extracts chromatographic peaks from the raw spectra of multiple samples. Each sample's raw spectra corresponds to a spectral data matrix, where one row corresponds to the spectral feature matrix of a peptide in the sample's raw spectra. For secondary mass spectrometry signals, such as... Figure 5 As shown, the server performs spectral line decomposition on the spectral data matrix, extracting multiple spectral line feature matrices. Then, the server uses a deep learning model ( Figure 5 Taking a ResNet18 model pre-trained on ImageNet as an example, feature extraction is performed on the spectral feature matrix to extract peptide features, resulting in multiple peptide feature vectors. These multiple peptide feature vectors are then concatenated end-to-end to obtain a one-dimensional vector with M×N elements, where M represents the number of elements in a peptide feature vector and N represents the number of peptides (M and N are positive integers). The server processes the concatenated vector based on a genotyping model to perform a classification task. For a three-level mass spectrum, such as... Figure 5As shown, the server performs spectral line decomposition on the spectral data matrix, extracting multiple spectral line feature matrices, which also include the daughter ion dimension. Then, the server uses a deep learning model ( Figure 5 Taking a ResNet18 model pre-trained on ImageNet as an example, feature extraction is performed on the spectral feature matrix to extract peptide features, resulting in multiple peptide feature vectors. These multiple peptide feature vectors are then concatenated end-to-end to obtain a vector with M×N elements. The server processes the concatenated vector based on a classification model to perform a classification task. In this embodiment, the classification task is disease classification. This classification task can also be for various clinical tasks such as survival analysis and efficacy prediction; this embodiment does not limit the scope of the classification task.

[0121] The solution provided in this application offers a novel method for processing raw spectra. It replaces traditional heuristic spectra feature extraction algorithms with machine learning algorithms. Based on machine learning algorithms, it extracts proteomics spectra features automatically without human intervention. The extracted features are more accurate and efficient, thus significantly improving the accuracy and efficiency of the analysis results. This enables LC-MS / MS-based proteomics methods to be applied more quickly and accurately to disease classification and clinical diagnosis, ultimately achieving the goal of precision medicine.

[0122] Figure 6 This is a block diagram of a raw spectrum processing apparatus according to an embodiment of this application. The apparatus is used to perform the steps in the above-described raw spectrum processing method. (See also...) Figure 6 The device includes: a chromatographic peak extraction module 601, a matrix determination module 602, a feature extraction module 603, and an information determination module 604.

[0123] The chromatographic peak extraction module 601 is used to extract multiple chromatographic peaks from the original spectrum of proteomics based on a target peptide library associated with the original spectrum, wherein one peptide in the target peptide library corresponds to at least one chromatographic peak.

[0124] The matrix determination module 602 is used to determine multiple spectral feature matrices based on the multiple chromatographic peaks, where each spectral feature matrix represents the chromatographic peak characteristics of a peptide.

[0125] The feature extraction module 603 is used to extract features from the multiple spectral feature matrices based on a deep learning model to obtain multiple peptide feature vectors, which are used to represent the features of peptide segments.

[0126] The information determination module 604 is used to determine the typing information of the original spectrum based on the multiple peptide feature vectors. The typing information is used to indicate the probability of the original spectrum corresponding to the target disease type.

[0127] In some embodiments, see Figure 7 As shown, Figure 7 This is a block diagram of another raw spectrum processing apparatus according to an embodiment of this application. The matrix determination module 602 includes:

[0128] The spectral line adjustment unit 6021 is used to adjust the spectral line lengths of the multiple chromatographic peaks to the target length;

[0129] The spectral stacking unit 6022 is used to stack the spectral lines of the corresponding chromatographic peaks for any peptide segment to obtain the spectral feature matrix of the peptide segment.

[0130] In some embodiments, the spectral stacking unit 6022 is used to stack the spectral lines of the chromatographic peaks corresponding to the peptide to obtain an intermediate matrix; and to stack at least two intermediate matrices in the channel dimension to obtain the spectral feature matrix corresponding to the peptide, wherein one channel of the spectral feature matrix corresponds to one intermediate matrix.

[0131] In some embodiments, see Figure 7 As shown, the device also includes:

[0132] The sorting module 605 is used to sort the chromatographic peaks corresponding to the library ions based on the Pearson correlation coefficient. The library ions are the peptide ions corresponding to the peptides in the target peptide library.

[0133] The sorting module 605 is also used to sort the chromatographic peaks corresponding to theoretical ions based on the extraction order. The theoretical ions are peptide ions corresponding to peptides not included in the target peptide library.

[0134] In some embodiments, see Figure 7 As shown, the device also includes:

[0135] The smoothing module 606 is used to smooth the multiple chromatographic peaks.

[0136] In some embodiments, the step of determining the genotyping information of the original spectrum based on the plurality of peptide feature vectors is implemented based on a genotyping model, which is used to determine the probability of the target disease type based on the peptide feature vectors.

[0137] In some embodiments, the classification model is trained in the following manner:

[0138] Multiple sample chromatographic peaks are extracted from the original sample chromatogram based on the sample peptide library. The sample peptide library is associated with the original sample chromatogram, and one sample peptide in the sample peptide library corresponds to at least one sample chromatographic peak.

[0139] Based on these multiple sample chromatographic peaks, multiple sample spectral matrixes are determined, and each sample spectral matrix represents the chromatographic peak characteristics of a sample peptide.

[0140] Based on this deep learning model, feature extraction is performed on the spectral matrix of the multiple samples to obtain multiple sample peptide feature vectors, which are used to represent the features of the sample peptide segments.

[0141] The typing model is trained based on the peptide feature vectors of these multiple samples.

[0142] In some embodiments, the process of training the model based on the multiple sample peptide feature vectors to obtain the trained typing model includes:

[0143] The multiple sample peptide feature vectors are concatenated to obtain a sample feature vector, the number of elements of which is the sum of the number of elements of the multiple sample peptide feature vectors;

[0144] The sample feature vector is divided into five equal parts, and the hyperparameters of the genotyping model are adjusted based on five-fold cross-validation.

[0145] This application provides a novel raw spectrum processing device that can automatically extract features from raw proteomics spectra without human intervention. The extracted features are more accurate and efficient, thereby significantly improving the accuracy and efficiency of the analysis results.

[0146] It should be noted that the raw spectrum processing apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when processing raw spectra. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the raw spectrum processing apparatus and the raw spectrum processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0147] In this embodiment of the application, the computer device is a server. Figure 8This is a schematic diagram of a server structure according to an embodiment of this application. The server 800 can vary considerably due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 801 and one or more memories 802. The memory 802 stores at least one computer program, which is loaded and executed by the processor 801 to implement the raw spectrum processing method provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0148] This application also provides a computer-readable storage medium storing at least one computer program. This computer program is loaded and executed by a processor of a computer device to implement the operations performed by the computer device in the original spectrum processing method of the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0149] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.

[0150] This application also provides a computer program product including computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the raw spectrum processing method provided in the various optional implementations described above.

[0151] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0152] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for processing raw spectra, characterized in that, The method includes: Multiple chromatographic peaks are extracted from the original proteomics spectrum based on the target peptide library, wherein the target peptide library is associated with the original spectrum, and one peptide in the target peptide library corresponds to at least one chromatographic peak. Based on the Pearson correlation coefficient, the chromatographic peaks corresponding to the library ions are sorted, where the library ions are the peptide ions corresponding to the peptide segments in the target peptide library; Based on the extraction order, the chromatographic peaks corresponding to the theoretical ions are sorted, where the theoretical ions are peptide ions corresponding to peptides not included in the target peptide library; Based on the multiple chromatographic peaks, multiple spectral feature matrices are determined, and each spectral feature matrix represents the chromatographic peak characteristics of a peptide. Based on a deep learning model, feature extraction is performed on the multiple spectral feature matrices to obtain multiple peptide feature vectors, which are used to represent the features of peptide segments. Based on the multiple peptide feature vectors, the genotyping information of the original spectrum is determined, and the genotyping information is used to indicate the probability of the original spectrum corresponding to the target disease type.

2. The method according to claim 1, characterized in that, The determination of multiple spectral feature matrices based on the multiple chromatographic peaks includes: Adjust the spectral lengths of the multiple chromatographic peaks to the target length; For any peptide segment, the spectral lines of the corresponding chromatographic peaks are stacked to obtain the spectral feature matrix of the peptide segment.

3. The method according to claim 2, characterized in that, The step of stacking the spectral lines corresponding to the chromatographic peaks of the peptide to obtain the spectral feature matrix corresponding to the peptide includes: The spectral lines of the chromatographic peaks corresponding to the peptide segments are stacked to obtain an intermediate matrix; At least two intermediate matrices are stacked along the channel dimension to obtain the spectral feature matrix corresponding to the peptide, wherein one channel of the spectral feature matrix corresponds to one intermediate matrix.

4. The method according to claim 1, characterized in that, The method further includes: The multiple chromatographic peaks are smoothed.

5. The method according to any one of claims 1-4, characterized in that, The step of determining the genotyping information of the original spectrum based on the multiple peptide feature vectors is implemented based on a genotyping model, which is used to determine the probability of the target disease type based on the peptide feature vectors.

6. The method according to claim 5, characterized in that, The method further includes: Multiple sample chromatographic peaks are extracted from the original sample chromatograms based on a sample peptide library, wherein the sample peptide library is associated with the original sample chromatograms, and one sample peptide in the sample peptide library corresponds to at least one sample chromatographic peak. Based on the multiple sample chromatographic peaks, multiple sample spectral matrixes are determined, and each sample spectral matrix represents the chromatographic peak characteristics of a sample peptide. Based on the deep learning model, feature extraction is performed on the multiple sample spectral matrixes to obtain multiple sample peptide feature vectors, which are used to represent the features of sample peptide segments. The typing model is trained based on the peptide feature vectors of the multiple samples.

7. The method according to claim 6, characterized in that, The step of training the genotyping model based on the multiple sample peptide feature vectors includes: The multiple sample peptide feature vectors are concatenated to obtain a sample feature vector, wherein the number of elements in the sample feature vector is the sum of the number of elements in the multiple sample peptide feature vectors; The sample feature vector is divided into five equal parts, and the hyperparameters of the fractal model are adjusted based on five-fold cross-validation.

8. A raw spectrum processing device, characterized in that, The device includes: A chromatographic peak extraction module is used to extract multiple chromatographic peaks from the original spectrum of proteomics based on a target peptide library, wherein the target peptide library is associated with the original spectrum, and one peptide in the target peptide library corresponds to at least one chromatographic peak. The sorting module is used to sort the chromatographic peaks corresponding to library ions based on the Pearson correlation coefficient, wherein the library ions are the peptide ions corresponding to peptide segments in the target peptide library; The sorting module is also used to sort the chromatographic peaks corresponding to theoretical ions based on the extraction order, wherein the theoretical ions are peptide ions corresponding to peptides not included in the target peptide library; The matrix determination module is used to determine multiple spectral feature matrices based on the multiple chromatographic peaks, where each spectral feature matrix represents the chromatographic peak characteristics of a peptide. The feature extraction module is used to extract features from the multiple spectral feature matrices based on a deep learning model to obtain multiple peptide feature vectors, which are used to represent the features of peptide segments. The information determination module is used to determine the genotyping information of the original spectrum based on the multiple peptide feature vectors, wherein the genotyping information is used to indicate the probability of the original spectrum corresponding to a target disease type.

9. The apparatus according to claim 8, characterized in that, The matrix determination module includes: A spectral line adjustment unit is used to adjust the spectral line lengths of the multiple chromatographic peaks to a target length; The spectral line stacking unit is used to stack the spectral lines of the chromatographic peaks corresponding to any peptide segment to obtain the spectral line feature matrix corresponding to the peptide segment.

10. The apparatus according to claim 9, characterized in that, The spectral stacking unit is used to stack the spectral lines of the chromatographic peaks corresponding to the peptide to obtain an intermediate matrix; at least two intermediate matrices are stacked in the channel dimension to obtain the spectral feature matrix corresponding to the peptide, and one channel of the spectral feature matrix corresponds to one intermediate matrix.

11. The apparatus according to claim 8, characterized in that, The device further includes: A smoothing module is used to smooth the multiple chromatographic peaks.

12. The apparatus according to any one of claims 8-11, characterized in that, The step of determining the genotyping information of the original spectrum based on the multiple peptide feature vectors is implemented based on a genotyping model, which is used to determine the probability of the target disease type based on the peptide feature vectors.

13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as the raw spectrum processing method according to any one of claims 1 to 7.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one computer program for performing the raw spectrum processing method according to any one of claims 1 to 7.