Multispectral intelligent recognition system for protein post-translational modification sites

By using parallel scanning and deep learning models in a multispectral intelligent recognition system, the problems of cumbersome mass spectrometry processes and insufficient algorithm adaptability have been solved. This has enabled efficient and accurate identification of protein post-translational modification sites, meeting the needs of rapid screening and high-throughput analysis of large-scale samples.

CN121281642BActive Publication Date: 2026-05-05山西省汾阳医院
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
山西省汾阳医院
Filing Date
2025-10-16
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing mass spectrometry techniques are cumbersome, time-consuming, and costly, making it difficult to meet the needs of rapid screening and high-throughput analysis of large-scale samples. Existing algorithms are not adaptable to handling complex biological samples, resulting in difficulties in data analysis and insufficient accuracy and reliability, especially in the lack of ability to identify novel or low-abundance modification types.

Method used

A multispectral intelligent recognition system is adopted, which achieves efficient processing and accurate recognition of multispectral data through parallel spectral scanning, analog-to-digital conversion, standardized feature extraction and deep learning models, combined with attention mechanism and feature fusion algorithm.

Benefits of technology

It significantly improves the accuracy and throughput of identifying protein post-translational modification sites, solves the problems of low efficiency and complex sample analysis in traditional mass spectrometry, enhances the ability to discover novel and low-abundance modification types, and provides an efficient and automated analysis workflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121281642B_ABST
    Figure CN121281642B_ABST
Patent Text Reader

Abstract

This invention discloses a multispectral intelligent identification system for protein post-translational modification sites, comprising: a spectral data acquisition module, which controls different spectral analysis instruments to perform parallel spectral scanning of protein samples and performs analog-to-digital conversion and formatting to obtain various raw multispectral data of proteins; an intelligent analysis module, which preprocesses and extracts features from each type of raw multispectral data of proteins to obtain various standardized spectral feature vectors; fuses each standardized spectral feature vector to obtain a multidimensional feature vector and inputs it into a pre-trained deep learning model for modification site identification, outputting a list of identification results, which includes the location of the identified protein post-translational modification sites in the protein sequence and the modification type corresponding to each protein post-translational modification site; and a result output module, which visualizes and exports the list of identification results. This invention significantly improves the accuracy and throughput of protein post-translational modification site identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of biomedicine and artificial intelligence, and in particular to a multispectral intelligent recognition system for protein post-translational modification sites. Background Technology

[0002] Post-translational modifications of proteins play a crucial role in vital activities such as cell signal transduction and protein function regulation, making it a hot topic in biomedical research. However, current mainstream analytical techniques have significant limitations, severely restricting the development of this field. On the one hand, mass spectrometry, as a core analytical tool, while possessing high sensitivity, suffers from cumbersome and time-consuming analytical procedures, coupled with expensive equipment and high operating costs, making it difficult to meet the needs of rapid screening and high-throughput analysis of large-scale samples, greatly limiting the efficiency and scale of research. On the other hand, existing algorithms are poorly adaptable to handling complex biological samples, with different spectral signals interfering with each other, leading to difficulties in data analysis and significantly reducing accuracy and reliability. Especially when faced with novel or low-abundance modification types, the algorithms lack recognition capabilities and cannot respond promptly to the constantly emerging unknown modification sites, severely hindering the depth and breadth of research on protein post-translational modifications.

[0003] Therefore, there is an urgent need to provide a technical solution to address the above problems. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a multispectral intelligent identification system for protein post-translational modification sites. The technical solution of this system is as follows:

[0005] The spectral data acquisition module is used to control different spectral analysis instruments to perform spectral scanning of the protein sample to be tested in parallel, and to perform analog-to-digital conversion and preliminary formatting on the analog spectral signals acquired by each spectral analysis instrument to obtain multiple protein multispectral raw data of the protein sample.

[0006] The intelligent analysis module is used to preprocess and extract features from the raw multispectral data of each protein to obtain a standardized spectral feature vector corresponding to the raw multispectral data of each protein; to fuse each standardized spectral feature vector to obtain a multidimensional feature vector; to input the multidimensional feature vector into a pre-trained deep learning model for modification site recognition for analysis and calculation, and to output a list of recognition results. The list of recognition results includes the location information of at least one identified protein post-translational modification site in the protein sequence and the modification type corresponding to each protein post-translational modification site.

[0007] The results output module is used to visualize and export the list of recognition results in a structured report format.

[0008] Furthermore, the spectral data acquisition module is specifically used for:

[0009] Send synchronous control commands to different spectral analyzers to start parallel scanning, and receive the analog spectral signals collected by each spectral analyzer in real time;

[0010] The analog spectral signals acquired by each spectral analyzer are converted from analog to digital to obtain the corresponding digital spectral data;

[0011] Each digital spectral data is initially formatted according to a predefined unified data format to generate multiple protein multispectral raw data of the protein sample; wherein, each spectral analysis instrument corresponds to one protein multispectral raw data.

[0012] Furthermore, the intelligent analysis module is specifically used for:

[0013] Noise filtering and baseline correction were performed on the raw multispectral data of each protein to obtain the preprocessed raw multispectral data of each protein.

[0014] Using a pre-defined feature extraction algorithm, spectral features of each type of preprocessed protein multispectral raw data are extracted and normalized to generate a standardized spectral feature vector corresponding to each type of protein multispectral raw data.

[0015] Furthermore, the intelligent analysis module is specifically used for:

[0016] The wavelet transform algorithm is used to extract the frequency domain features of any preprocessed protein multispectral raw data and construct an initial feature vector. The principal component analysis algorithm is then used to reduce the dimensionality of the initial feature vector to obtain a dimensionality-reduced feature vector. Finally, the maximum-minimum normalization algorithm is used to perform a scaling transformation to generate the standardized spectral feature vector corresponding to the preprocessed protein multispectral raw data.

[0017] Repeat the above steps until you obtain the standardized spectral feature vector corresponding to the raw multispectral data of each preprocessed protein.

[0018] Furthermore, the intelligent analysis module is specifically used for:

[0019] Each standardized spectral feature vector is grouped and arranged according to the type of the corresponding spectral analyzer to form a grouped feature set;

[0020] Based on the grouped feature set, a feature fusion algorithm with an attention mechanism is used to calculate the weight coefficient of each normalized spectral feature vector;

[0021] The weighted feature representation is obtained by summing the corresponding standardized spectral feature vectors according to the calculated weight coefficients.

[0022] The weighted feature representation is concatenated with all standardized spectral feature vectors to generate the multidimensional feature vector used to characterize the multispectral features of the protein sample to be tested.

[0023] Furthermore, the intelligent analysis module is specifically used for:

[0024] For any standardized spectral feature vector in the grouped feature set, calculate the correlation score between the standardized spectral feature vector and other standardized spectral feature vectors, convert the correlation score into attention weights through the softmax function, and then perform weighted fusion with the instrument-specific parameters corresponding to the standardized spectral feature vector to obtain the weight coefficient of the standardized spectral feature vector.

[0025] Repeat the above steps until the weight coefficients of each standardized spectral eigenvector are obtained.

[0026] Furthermore, the attention weights are calculated using the following formula:

[0027]

[0028] in, The weight coefficients represent the i-th type of standardized spectral eigenvector. This represents the sigmoid activation function. The dimension of the feature vector. Let represent the attention score between the i-th normalized spectral feature vector and the j-th normalized spectral feature vector. Represents the j-th standardized spectral eigenvector. The weighting coefficients represent the instrument-specific parameters. This represents the instrument-specific parameter of the spectral analyzer corresponding to the i-th standardized spectral eigenvector. This represents the total number of types of standardized spectral eigenvectors.

[0029] Furthermore, the intelligent analysis module is specifically used for:

[0030] The multidimensional feature vector is input into the deep neural network layer of the deep learning model for modification site recognition to extract high-order features and obtain a deep feature representation.

[0031] The deep feature representation is input into the modification type classifier to calculate the probability distribution of each amino acid site in the sequence belonging to different modification types. At the same time, the deep feature representation is input into the position regressor to predict the precise position coordinates of each amino acid site in the protein sequence.

[0032] Based on the probability distribution and the precise location coordinates, amino acid sites with probability values ​​exceeding a preset threshold and their corresponding most likely modification types are identified as candidate protein post-translational modification sites, and a set of candidate modification sites is generated. Each candidate protein post-translational modification site in the set of candidate modification sites is sorted and filtered based on probability confidence to generate the identification result list.

[0033] Furthermore, the modification type classifier calculates the modification type probability distribution for each amino acid site using the following formula:

[0034]

[0035] in, This represents the probability distribution that the k-th amino acid site belongs to modification type c. This represents the trainable weight parameter matrix in the modification type classification layer. This represents the depth feature representation of the k-th amino acid site. Higher-order feature representation after feature transformation represents the trainable bias vector in the modifier type classification layer, and Softmax represents the multi-class activation function that transforms the input vector into a probability distribution.

[0036] Furthermore, the result output module is specifically used for:

[0037] A structured data table is generated based on the protein post-translational modification site location information and modification type in the identification result list, and the structured data table is converted into an interactive visualization chart.

[0038] The structured data table and the interactive visualization chart are displayed simultaneously in the graphical user interface;

[0039] When a user operation instruction for exporting the report is received, the currently displayed structured data table and interactive visualization chart are exported as an analysis report in the specified format.

[0040] The technical solution of this invention can solve the problems of cumbersome process and low analysis efficiency of existing mass spectrometry technology, and overcome the limitations of existing algorithms in analyzing complex samples, significantly improving the accuracy and throughput of protein post-translational modification site identification.

[0041] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0043] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0044] Figure 1 This is a schematic diagram of an embodiment of a multispectral intelligent identification system for protein post-translational modification sites according to the present invention. Detailed Implementation

[0045] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0046] Figure 1 This diagram illustrates a structural schematic of an embodiment of a multispectral intelligent recognition system for protein post-translational modification sites provided by the present invention. Figure 1 As shown, the system includes:

[0047] The spectral data acquisition module 110 is used to control different spectral analysis instruments to perform spectral scanning of the protein sample to be tested in parallel, and to perform analog-to-digital conversion and preliminary formatting on the analog spectral signals acquired by each spectral analysis instrument to obtain multiple protein multispectral raw data of the protein sample.

[0048] In this context, "spectral analysis instrument" refers to a device used to scan protein samples and generate spectral signals; for example, in analyzing phosphorylation modifications of human serum samples, an infrared spectrometer, a Raman spectrometer, and a fluorescence spectrometer are used to scan the sample in parallel. "Protein sample to be analyzed" refers to a protein sample that needs to be analyzed to identify post-translational modification sites; for example, a specific serum protein sample in the analysis of phosphorylation modifications in human serum samples. "Analog spectral signal" refers to a continuous electrical signal acquired by a spectroscopic analysis instrument; for example, the analog electrical signal waveform generated by an infrared spectrometer scanning a human serum sample. "Raw protein multispectral data" refers to the set of spectral data obtained after digitizing and formatting analog spectral signals; for example, multiple sets of digital spectral data obtained by converting the infrared, Raman, and fluorescence analog signals of a human serum sample.

[0049] The intelligent analysis module 120 is used to preprocess and extract features from the raw multispectral data of each protein to obtain a standardized spectral feature vector corresponding to the raw multispectral data of each protein; to fuse each standardized spectral feature vector to obtain a multidimensional feature vector; to input the multidimensional feature vector into a pre-trained deep learning model for modification site recognition for analysis and calculation, and to output a list of recognition results. The list of recognition results includes the location information of at least one identified protein post-translational modification site in the protein sequence and the modification type corresponding to each protein post-translational modification site.

[0050] The standardized spectral feature vector refers to the standardized feature data formed after preprocessing and feature extraction; for example, the feature vector obtained after wavelet transform and normalization of infrared spectral data of human serum samples. The multidimensional feature vector refers to the comprehensive feature representation formed by fusing multiple standardized spectral feature vectors; for example, the multidimensional vector formed by fusing infrared, Raman, and fluorescence feature vectors of human serum samples. The pre-trained deep learning model for modification site recognition refers to the modification site recognition algorithm model trained on a large amount of protein spectral data; for example, a neural network model trained using protein spectral data with known phosphorylation modification sites. The recognition result list refers to the set of modification site recognition results output by the model analysis; for example, a list containing phosphorylation sites identified in human serum samples and their location information. The post-translational modification site of a protein refers to the specific amino acid position in the protein sequence where chemical modification occurs; for example, the phosphorylation site of serine at position 32 in human serum albumin. The protein sequence refers to the order of amino acids in a protein; for example, the amino acid sequence structure of human serum albumin. Location information refers to the specific position of the modification site in the protein sequence; for example, serine residue 32 in human serum albumin. Modification type refers to the type of chemical modification that occurs after protein translation; for example, phosphorylation, acetylation, or methylation.

[0051] The result output module 130 is used to visualize and export the list of recognition results in the form of a structured report.

[0052] Structured report format refers to a report format that organizes data according to a specific format; for example, a standard analytical report format that includes tables, charts, and text descriptions.

[0053] The technical solution of this embodiment can solve the problems of cumbersome process and low analysis efficiency of existing mass spectrometry technology, and overcome the limitations of existing algorithms in analyzing complex samples, significantly improving the accuracy and throughput of protein post-translational modification site identification.

[0054] In an alternative embodiment, the spectral data acquisition module 110 is specifically used for:

[0055] Synchronization control commands are sent to different spectral analyzers to initiate parallel scanning, and the analog spectral signals acquired by each spectral analyzer are received in real time.

[0056] Among them, the synchronous control command refers to the command signal that controls multiple instruments to start scanning simultaneously; for example, the electronic command to start infrared, Raman and fluorescence spectrometer scanning simultaneously.

[0057] The analog spectral signals acquired by each spectral analyzer are converted from analog to digital to obtain the corresponding digital spectral data.

[0058] Digital spectral data refers to digital signals obtained by analog-to-digital conversion; for example, the digital signal sequence obtained by converting the infrared analog signal of a human serum sample.

[0059] Each digital spectral data is initially formatted according to a predefined unified data format to generate multiple protein multispectral raw data of the protein sample; wherein, each spectral analysis instrument corresponds to one protein multispectral raw data.

[0060] Among the above-mentioned optional methods, efficient acquisition and format unification of multispectral data are further realized. Through parallel scanning and real-time analog-to-digital conversion, data acquisition efficiency is improved, providing standardized input for subsequent intelligent analysis and ensuring smooth connection of the analysis process.

[0061] In one alternative embodiment, the intelligent analysis module 120 is specifically used for:

[0062] Noise filtering and baseline correction were performed on the raw multispectral data of each protein to obtain the preprocessed raw multispectral data of each protein.

[0063] Among them, the preprocessed protein multispectral raw data refers to the spectral data after noise filtering and baseline correction; for example, the Raman spectral data of human serum samples after noise removal.

[0064] Using a pre-defined feature extraction algorithm, spectral features of each type of preprocessed protein multispectral raw data are extracted and normalized to generate a standardized spectral feature vector corresponding to each type of protein multispectral raw data.

[0065] The preset feature extraction algorithm refers to a pre-defined method for extracting features from spectral data; for example, a wavelet transform algorithm used to extract features from human serum spectral data.

[0066] Among the above-mentioned optional methods, data quality is further optimized through noise filtering and feature extraction. Pre-set algorithms are used to remove interference information and extract key spectral features. After normalization processing, standardized feature vectors are generated, which significantly improves the usability and recognition accuracy of the feature vectors.

[0067] In one alternative embodiment, the intelligent analysis module 120 is specifically used for:

[0068] The wavelet transform algorithm is used to extract the frequency domain features of any preprocessed protein multispectral raw data and construct an initial feature vector. The principal component analysis algorithm is then used to reduce the dimensionality of the initial feature vector to obtain a dimensionality-reduced feature vector. Finally, the minimax normalization algorithm is used to perform a scaling transformation to generate the standardized spectral feature vector corresponding to the preprocessed protein multispectral raw data.

[0069] Among them, wavelet transform algorithm refers to an algorithm that uses wavelet basis functions for time-frequency analysis; for example, an algorithm for extracting frequency domain features from the fluorescence spectrum of human serum samples. Initial feature vector refers to the preliminary extracted feature representation without dimensionality reduction processing; for example, the initial feature set obtained after wavelet transforming the infrared spectrum of human serum samples. Principal component analysis algorithm refers to a statistical method that reduces the dimensionality of data through linear transformation; for example, an algorithm for reducing the dimensionality of the initial features of human serum sample spectra. Dimensionality-reduced feature vector refers to the feature representation after dimensionality reduction processing; for example, the low-dimensional feature vector obtained after principal component analysis of human serum sample spectral data.

[0070] Repeat the above steps until you obtain the standardized spectral feature vector corresponding to the raw multispectral data of each preprocessed protein.

[0071] Among the above-mentioned optional methods, features are further extracted and normalized through wavelet transform and principal component analysis, which effectively reduces the data dimensionality while retaining core information. The maximum-minimum normalization algorithm is used to unify the feature scale, ensuring the uniformity and comparability of feature vectors, thus laying a solid foundation for subsequent fusion analysis.

[0072] In one alternative embodiment, the intelligent analysis module 120 is specifically used for:

[0073] Each standardized spectral feature vector is grouped and arranged according to the type of the corresponding spectral analyzer to form a grouped feature set.

[0074] The type of spectroscopic instrument refers to the category of spectroscopic instruments classified according to their working principle; for example, different types such as infrared spectrometers, Raman spectrometers, and fluorescence spectrometers. The grouped feature set refers to the set of feature vectors grouped and arranged according to instrument type; for example, a set formed by grouping the infrared, Raman, and fluorescence feature vectors of human serum samples.

[0075] Based on the grouped feature set, a feature fusion algorithm using an attention mechanism is employed to calculate the weight coefficients of each standardized spectral feature vector.

[0076] Among them, the feature fusion algorithm based on the attention mechanism refers to an algorithm that uses attention weights for feature fusion; for example, a fusion algorithm that calculates the importance weights of different spectral features of human serum samples. The weight coefficient refers to a numerical parameter representing the importance of a feature; for example, the importance weight of infrared spectral features of human serum samples in fusion.

[0077] The weighted feature representation is obtained by summing the corresponding standardized spectral feature vectors based on the calculated weight coefficients.

[0078] Weighted feature representation refers to the feature representation after weighting according to weight coefficients; for example, the feature representation after weighted summation of different spectral features of human serum samples.

[0079] The weighted feature representation is concatenated with all standardized spectral feature vectors to generate the multidimensional feature vector used to characterize the multispectral features of the protein sample to be tested.

[0080] In the above-mentioned optional methods, the weights are further dynamically calculated through grouping and attention mechanisms, and the weights are accurately allocated based on feature correlation and instrument-specific parameters to achieve optimal fusion of cross-spectral features, highlight important feature information, improve the recognition model's ability to focus on key features, and enhance recognition accuracy.

[0081] In one alternative embodiment, the intelligent analysis module 120 is specifically used for:

[0082] For any standardized spectral feature vector in the grouped feature set, calculate the correlation score between the standardized spectral feature vector and other standardized spectral feature vectors. Convert the correlation score into attention weights using a softmax function. Then, weight and fuse the attention weights with the instrument-specific parameters corresponding to the standardized spectral feature vector to obtain the weight coefficients of the standardized spectral feature vector.

[0083] The correlation score refers to a numerical value representing the degree of correlation between feature vectors; for example, the correlation score between infrared and Raman spectral features of a human serum sample. The attention weight refers to the importance weight calculated through an attention mechanism; for example, the attention weight value of fluorescence spectral features of a human serum sample in the fusion process.

[0084] Repeat the above steps until the weight coefficients of each standardized spectral eigenvector are obtained.

[0085] In the above-mentioned optional methods, further correlation calculation and instrument-specific parameters are fused to accurately measure the correlation strength between different feature vectors. The correlation score is converted into attention weight and combined with the instrument parameters to effectively compensate for the differences between different spectral data and enhance the consistency and fusion effect of cross-instrument data.

[0086] In an alternative approach, the attention weights are calculated using the following formula:

[0087]

[0088] in, The weight coefficients represent the i-th type of standardized spectral eigenvector. This represents the sigmoid activation function. The dimension of the feature vector. Let represent the attention score between the i-th normalized spectral feature vector and the j-th normalized spectral feature vector. Represents the j-th standardized spectral eigenvector. The weighting coefficients represent the instrument-specific parameters. This represents the instrument-specific parameter of the spectral analyzer corresponding to the i-th standardized spectral eigenvector. This represents the total number of types of standardized spectral eigenvectors.

[0089] Among them, instrument-specific parameters refer to parameters that reflect the system error of a specific instrument; for example, the baseline drift parameter unique to a certain infrared spectrometer.

[0090] In the above-mentioned optional methods, feature fusion is further automated and standardized by formulaic weight calculation. Based on feature dimensions, correlation scores and instrument parameters, the weight values ​​are standardized by using the sigmoid activation function to ensure that the weight coefficient calculation process is reasonable, stable and reliable, and improve the transparency and repeatability of the fusion process.

[0091] In one alternative embodiment, the intelligent analysis module 120 is specifically used for:

[0092] The multidimensional feature vector is input into the deep neural network layer of the deep learning model for modification site recognition to extract high-order features and obtain a deep feature representation.

[0093] Here, a deep neural network layer refers to a network layer used for feature extraction in a deep learning model; for example, a convolutional neural network layer in a modified site recognition model. Deep feature representation refers to high-order features extracted through a deep neural network; for example, deep features extracted from human serum sample spectral data via a neural network.

[0094] The deep feature representation is input into the modification type classifier to calculate the probability distribution of each amino acid site in the sequence belonging to different modification types. At the same time, the deep feature representation is input into the position regressor to predict the precise position coordinates of each amino acid site in the protein sequence.

[0095] Here, a modification type classifier refers to a classification algorithm used to determine the type of modification; for example, a classification model that distinguishes between phosphorylation, acetylation, and other modification types. An amino acid site refers to a single amino acid position in a protein sequence; for example, serine position 32 in human serum albumin. A probability distribution refers to a set of probability values ​​belonging to different categories; for example, the probability set of a site in a human serum sample belonging to phosphorylation, acetylation, or other modification types. A position regressor refers to a regression algorithm that predicts the position of a modification site; for example, a regression model that predicts the specific position of a phosphorylation site in a human serum sample. Precise position coordinates refer to the precise location identifier of the modification site in the sequence; for example, the precise sequence position of serine position 32 in human serum albumin.

[0096] Based on the probability distribution and the precise location coordinates, amino acid sites with probability values ​​exceeding a preset threshold and their corresponding most likely modification types are identified as candidate protein post-translational modification sites, and a set of candidate modification sites is generated. Each candidate protein post-translational modification site in the set of candidate modification sites is sorted and filtered based on probability confidence to generate the identification result list.

[0097] Here, the preset threshold refers to a pre-defined probability threshold; for example, a probability threshold of 0.85 for determining a valid phosphorylation modification site. Candidate post-translational modification sites refer to sites that are preliminarily screened for potential modifications; for example, potential phosphorylation sites in human serum samples with a probability value exceeding 0.85. The candidate modification site set refers to the set of candidate modification sites; for example, the set of all potential phosphorylation sites in a human serum sample. The probability confidence level refers to the degree of reliability of the modification site identification results; for example, the confidence score for determining a site in a human serum sample as a phosphorylation modification.

[0098] Among the above-mentioned optional methods, the location and type of modification sites can be accurately predicted by using deep learning models, high-order features can be extracted by using deep neural networks, and classifiers and regressors can be combined to achieve accurate localization and comprehensive identification of sites, significantly improving the accuracy and comprehensiveness of the analysis and meeting the needs of efficient identification of multiple modification types in complex samples.

[0099] In an alternative approach, the modification type classifier calculates the modification type probability distribution for each amino acid site using the following formula:

[0100]

[0101] in, This represents the probability distribution that the k-th amino acid site belongs to modification type c. This represents the trainable weight parameter matrix in the modification type classification layer. This represents the depth feature representation of the k-th amino acid site. Higher-order feature representation after feature transformation represents the trainable bias vector in the modifier type classification layer, and Softmax represents the multi-class activation function that transforms the input vector into a probability distribution.

[0102] Among them, higher-order feature representation refers to the deep feature representation obtained through nonlinear transformation; for example, the feature representation of human serum sample spectral data after nonlinear transformation by a neural network.

[0103] In the above-mentioned optional methods, the features are further mapped to a probability distribution through the Softmax function, which efficiently completes the classification of modification types. The deep feature representation is mapped to the probability space of different modification types, which intuitively reflects the modification tendency of each amino acid site, provides a quantitative basis for subsequent screening, and enhances the intelligence level and decision support capability of the analysis.

[0104] In an alternative embodiment, the result output module 130 is specifically used for:

[0105] A structured data table is generated based on the protein post-translational modification site location information and modification type in the identification result list, and the structured data table is converted into an interactive visualization chart.

[0106] Structured data tables refer to data tables organized according to a specific structure; for example, a data table containing the location and type of phosphorylation sites in human serum samples. Interactive visualization charts refer to graphical data displays that support user interaction; for example, a zoomable and filterable distribution map of modification sites in human serum samples.

[0107] The structured data table and the interactive visualization chart are displayed simultaneously in the graphical user interface.

[0108] The graphical user interface refers to a software interface that provides a graphical operating interface; for example, a software interface that displays the analysis results and operating controls of human serum samples.

[0109] When a user operation instruction for exporting the report is received, the currently displayed structured data table and interactive visualization chart are exported as an analysis report in the specified format.

[0110] User operation instructions refer to the operation commands issued by the user through the interface; for example, the instruction signal generated when the user clicks the export report button. Analysis report refers to the final generated formatted analysis result document; for example, a PDF report document containing the phosphorylation analysis results of human serum samples.

[0111] Among the above optional methods, the user experience is further enhanced by interactive visualization and export functions. The recognition results are transformed into intuitive charts and data tables, which are displayed synchronously in the graphical interface, making it convenient for users to quickly browse key information. It also supports the export of reports in multiple formats, meeting users' needs for result sharing and archiving, and improving the system's practicality and ease of use.

[0112] To better illustrate the technical solution of this embodiment, the following example is used for complete explanation:

[0113] S10: Control three different spectroscopic analysis instruments—infrared spectrometer, Raman spectrometer, and fluorescence spectrometer—to perform spectral scanning on the human serum sample, which is the protein sample to be tested, in parallel. Perform analog-to-digital conversion and preliminary formatting on the analog spectral signals acquired by each spectroscopic analysis instrument to obtain the raw multispectral data of the three proteins in the human serum sample.

[0114] S20: Noise filtering and baseline correction are performed on the three types of protein multispectral raw data to obtain preprocessed protein multispectral raw data. Wavelet transform algorithm is used to extract the frequency domain features of each type of preprocessed data and form an initial feature vector. Principal component analysis algorithm is used to reduce the dimensionality of the initial feature vector to obtain a dimensionality-reduced feature vector. Max-min normalization algorithm is used to perform scaling transformation to generate three standardized spectral feature vectors.

[0115] S30: The three standardized spectral feature vectors are grouped and arranged according to the corresponding spectral analysis instrument type to form a grouped feature set. The feature fusion algorithm with attention mechanism is used to calculate the weight coefficient of each standardized spectral feature vector. The weighted sum of each standardized spectral feature vector is obtained according to the weight coefficient. The weighted feature representation is concatenated with all standardized spectral feature vectors to generate a multidimensional feature vector.

[0116] S40: Input the multidimensional feature vector into the deep neural network layer of the pre-trained deep learning model for modification site recognition to extract high-order features and obtain deep feature representation. Input the deep feature representation into the modification type classifier to calculate the probability distribution of each amino acid site in the sequence belonging to different modification types. At the same time, input the deep feature representation into the position regressor to predict the precise position coordinates of each amino acid site in the protein sequence.

[0117] S50: Based on probability distribution and precise location coordinates, amino acid sites with probability values ​​exceeding a preset threshold of 0.85 and their corresponding most likely modification types are identified as candidate protein post-translational modification sites, generating a set of candidate modification sites. Each candidate protein post-translational modification site in the set of candidate modification sites is sorted and filtered based on probability confidence, generating a list of identification results.

[0118] S60: Generate a structured data table based on the protein post-translational modification site location information and modification type in the identification result list, convert the structured data table into an interactive visualization chart, and simultaneously display the structured data table and the interactive visualization chart in the graphical user interface. When a user operation command is received, export the currently displayed structured data table and interactive visualization chart into a PDF analysis report.

[0119] The overall technical solution of this embodiment has the following technical effects:

[0120] First, the system controls different spectral analysis instruments to acquire data in parallel through the spectral data acquisition module, realizing the synchronous acquisition and standardized processing of multi-source spectral data. This effectively solves the problems of cumbersome and time-consuming processes in traditional mass spectrometry technology, significantly improving data acquisition efficiency and analysis throughput, and can meet the actual needs of rapid screening and high-throughput analysis of large-scale protein samples.

[0121] Second, the system employs a multi-level feature processing and fusion strategy through an intelligent analysis module. First, it standardizes the raw multispectral data of each protein to generate a standardized spectral feature vector. Then, it dynamically calculates the weight coefficients through an attention mechanism feature fusion algorithm, effectively integrating feature information from different spectral sources. This overcomes the problem of mutual interference between different spectral signals in the analysis of complex biological samples, and significantly improves the accuracy and reliability of data analysis.

[0122] Third, the system performs in-depth analysis of the fused multidimensional feature vectors through a deep learning model for modification site identification, extracts high-order features using deep neural network layers, and combines a modification type classifier and a position regressor to achieve accurate identification and localization of modification sites. This significantly enhances the ability to discover novel and low-abundance modification types, and effectively expands the depth and breadth of research on protein post-translational modifications.

[0123] Fourth, the system transforms the identification results into a visual analysis report through a structured result output module, providing a complete solution from data acquisition to result display. This greatly reduces the technical requirements of operators and equipment dependence on traditional analysis methods, making it highly practical and easy to use, and conducive to promoting the widespread application of protein post-translational modification analysis technology.

[0124] Fifth, the system has a reasonable overall architecture and close connections between modules, realizing fully automated processing from raw data to the final analysis report. This not only improves analysis efficiency but also ensures the consistency and reproducibility of analysis results, providing reliable technical support for protein post-translational modification research.

[0125] Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above.

[0126] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0127] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0128] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A multispectral intelligent recognition system for protein post-translational modification sites, characterized in that, The system includes: The spectral data acquisition module is used to control different spectral analysis instruments to perform spectral scanning of the protein sample to be tested in parallel, and to perform analog-to-digital conversion and preliminary formatting on the analog spectral signals acquired by each spectral analysis instrument to obtain multiple protein multispectral raw data of the protein sample. The intelligent analysis module is used to preprocess and extract features from the raw multispectral data of each protein to obtain a standardized spectral feature vector corresponding to the raw multispectral data of each protein; to fuse each standardized spectral feature vector to obtain a multidimensional feature vector; to input the multidimensional feature vector into a pre-trained deep learning model for modification site recognition for analysis and calculation, and to output a list of recognition results. The list of recognition results includes the location information of at least one identified protein post-translational modification site in the protein sequence and the modification type corresponding to each protein post-translational modification site. The results output module is used to visualize and export the list of recognition results in the form of a structured report. The intelligent analysis module is specifically used for: Each standardized spectral feature vector is grouped and arranged according to the type of the corresponding spectral analyzer to form a grouped feature set; Based on the grouped feature set, a feature fusion algorithm with an attention mechanism is used to calculate the weight coefficient of each normalized spectral feature vector; The weighted feature representation is obtained by summing the corresponding standardized spectral feature vectors according to the calculated weight coefficients. The weighted feature representation is concatenated with all standardized spectral feature vectors to generate the multidimensional feature vector used to characterize the multispectral features of the protein sample to be tested. For any standardized spectral feature vector in the grouped feature set, calculate the correlation score between the standardized spectral feature vector and other standardized spectral feature vectors, convert the correlation score into attention weights through the softmax function, and then perform weighted fusion with the instrument-specific parameters corresponding to the standardized spectral feature vector to obtain the weight coefficient of the standardized spectral feature vector. Repeat the above steps until the weight coefficients of each standardized spectral eigenvector are obtained; The weighting coefficients are calculated using the following formula: in, The weight coefficients represent the i-th type of standardized spectral eigenvector. This represents the sigmoid activation function. The dimension of the feature vector. The attention weights are obtained by transforming the correlation scores between the i-th and j-th standardized spectral feature vectors using the softmax function. Represents the j-th standardized spectral eigenvector. The weighting coefficients represent the instrument-specific parameters. This represents the instrument-specific parameter of the spectral analyzer corresponding to the i-th standardized spectral eigenvector. This represents the total number of types of standardized spectral eigenvectors.

2. The multispectral intelligent identification system for protein post-translational modification sites according to claim 1, characterized in that, The spectral data acquisition module is specifically used for: Send synchronous control commands to different spectral analyzers to start parallel scanning, and receive the analog spectral signals collected by each spectral analyzer in real time; The analog spectral signals acquired by each spectral analyzer are converted from analog to digital to obtain the corresponding digital spectral data; Each digital spectral data is initially formatted according to a predefined unified data format to generate multiple protein multispectral raw data of the protein sample; wherein, each spectral analysis instrument corresponds to one protein multispectral raw data.

3. The multispectral intelligent identification system for protein post-translational modification sites according to claim 1, characterized in that, The intelligent analysis module is specifically used for: Noise filtering and baseline correction were performed on the raw multispectral data of each protein to obtain the preprocessed raw multispectral data of each protein. Using a pre-defined feature extraction algorithm, spectral features of each type of preprocessed protein multispectral raw data are extracted and normalized to generate a standardized spectral feature vector corresponding to each type of protein multispectral raw data.

4. The multispectral intelligent identification system for protein post-translational modification sites according to claim 3, characterized in that, The intelligent analysis module is specifically used for: The wavelet transform algorithm is used to extract the frequency domain features of any preprocessed protein multispectral raw data and construct an initial feature vector. The principal component analysis algorithm is then used to reduce the dimensionality of the initial feature vector to obtain a dimensionality-reduced feature vector. Finally, the maximum-minimum normalization algorithm is used to perform a scaling transformation to generate the standardized spectral feature vector corresponding to the preprocessed protein multispectral raw data. Repeat the above steps until you obtain the standardized spectral feature vector corresponding to the raw multispectral data of each preprocessed protein.

5. The multispectral intelligent identification system for protein post-translational modification sites according to claim 1, characterized in that, The intelligent analysis module is specifically used for: The multidimensional feature vector is input into the deep neural network layer of the deep learning model for modification site recognition to extract high-order features and obtain a deep feature representation. The deep feature representation is input into the modification type classifier to calculate the probability distribution of each amino acid site in the sequence belonging to different modification types. At the same time, the deep feature representation is input into the position regressor to predict the precise position coordinates of each amino acid site in the protein sequence. Based on the probability distribution and the precise location coordinates, amino acid sites with probability values ​​exceeding a preset threshold and their corresponding most likely modification types are identified as candidate protein post-translational modification sites, and a set of candidate modification sites is generated. The candidate protein post-translational modification sites in the candidate modification site set are sorted and filtered based on probability confidence to generate the identification result list.

6. The multispectral intelligent identification system for protein post-translational modification sites according to claim 5, characterized in that, The modification type classifier calculates the modification type probability distribution for each amino acid site using the following formula: in, This represents the probability distribution that the k-th amino acid site belongs to modification type c. This represents the trainable weight parameter matrix in the modification type classification layer. This represents the depth feature representation of the k-th amino acid site. Higher-order feature representation after feature transformation represents the trainable bias vector in the modifier type classification layer, and Softmax represents the multi-class activation function that transforms the input vector into a probability distribution.

7. The multispectral intelligent recognition system for protein post-translational modification sites according to claim 6, characterized in that, The result output module is specifically used for: A structured data table is generated based on the protein post-translational modification site location information and modification type in the identification result list, and the structured data table is converted into an interactive visualization chart. The structured data table and the interactive visualization chart are displayed simultaneously in the graphical user interface; When a user operation instruction for exporting the report is received, the currently displayed structured data table and interactive visualization chart are exported as an analysis report in the specified format.

Citation Information

Patent Citations

  • Protein active site multi-classification identification method based on multi-modal deep learning

    CN120600125A

  • Quantum and region sensing fused protein methylation site prediction method

    CN120727109A