An intelligent proteomics data anomaly detection and analysis method
By identifying discrete points in proteomics data and adjusting environmental temperature and sampling frequency, the detection error caused by voltage fluctuations in the detection equipment and sample storage time was resolved, thereby improving the accuracy and continuity of detection.
Patent Information
- Application Number
- CN202510940610.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-07-09
AI Technical Summary
In existing technologies, voltage fluctuations and increased storage time during long-term monitoring of proteomics data detection equipment lead to voltage fluctuations caused by temperature increases in the detection equipment, resulting in inaccurate detection results and data errors caused by prolonged sample storage time.
By comparing proteomics data with standard data, discrete points are identified, and the ambient temperature and sampling frequency of the detection equipment are adjusted to reduce errors caused by voltage fluctuations and sample storage time. Virtual sampling points are inserted to supplement information in areas where the detection data changes drastically.
It improves the accuracy and continuity of detection, reduces detection errors caused by voltage fluctuations and sample storage time, and ensures the stability and integrity of detection results.
Smart Images

Figure CN120783873B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of data detection, in particular to an intelligent proteomics data anomaly detection and analysis method. BACKGROUND
[0002] In the prior art, the mean and standard deviation of each feature in the proteomics data set are calculated, a reasonable threshold range is set, and data points exceeding the threshold range are determined as abnormal values. The method is suitable for the case that the data distribution is relatively uniform and meets the normal distribution assumption, but if the data has obvious skewness, it may lead to misjudgment. PCA aims to convert the original high-dimensional proteomics data to a new low-dimensional space through linear transformation, so that the variance of the data in the new coordinate system is maximized, that is, the main feature information of the data is retained. First, the covariance matrix of the data is calculated, then the eigenvalues and eigenvectors are solved, and the principal components are selected according to the size of the eigenvalues. Normal data has a certain distribution pattern in the low-dimensional space, while abnormal data deviates from this pattern. SVM can distinguish between normal and abnormal data by finding an optimal hyperplane. When dealing with nonlinear separable data, kernel functions can be used to map data to high-dimensional space to find a suitable hyperplane. The performance of SVM is highly dependent on the selection of kernel functions and the tuning of parameters. Different kernel functions and parameter settings may result in significant differences in results. By using known protein functional annotation information such as biological processes and cellular localization in which the protein is involved, if the expression level or modification state of a protein is in obvious conflict with its functional annotation, it may be considered abnormal.
[0003] Chinese Patent Publication No. CN114298214A discloses a protein anomaly detection method based on a super-large scale evolution algorithm and hardware acceleration, which comprises the following steps: 1. collecting mass spectrometry feature data; 2. generating a mass spectrometry feature selection scheme population and an external archive, and setting parameters; 3. updating the external archive and rapidly clustering and grouping the mass spectrometry features; 4. performing mating pool selection, and simultaneously generating offspring mass spectrometry feature selection schemes in the original space and the reduced space after grouping; 5. using the offspring feature selection scheme to adaptively adjust the algorithm parameters, and then merging the offspring and parent mass spectrometry feature selection scheme populations to perform environmental selection, iteratively selecting high-quality feature selection schemes, and finally obtaining the optimal protein anomaly detection feature selection scheme. It can be seen that the protein anomaly detection method based on the super-large scale evolution algorithm and hardware acceleration has the problems of voltage fluctuation caused by the increase of the monitoring time of the monitoring equipment, which leads to the increase of the plug-in position temperature and further causes the voltage fluctuation, or the increase of the storage time of the subsequent detection samples caused by the increase of the monitoring time, which leads to the inaccuracy of the detected protein data. SUMMARY
[0004] To this end, the application provides an intelligent proteomics data anomaly detection and analysis method to overcome the problem of inaccurate protein data detected due to the increase in monitoring time, which leads to voltage fluctuations caused by the increase in the temperature of the plug-in position of the monitoring device, or the increase in the storage time of subsequent detection samples, which leads to inaccurate protein data detected.
[0005] To achieve the above-mentioned purpose, the application provides an intelligent proteomics data anomaly detection and analysis method, comprising:
[0006] detecting proteins by a mass spectrometer to obtain a plurality of sets of proteomics data;
[0007] comparing the plurality of sets of proteomics data with standard data, and determining discrete error data based on the comparison result;
[0008] determining the discrete type of proteomics data based on the discrete characteristics of the discrete error data;
[0009] determining the detection execution mode according to the discrete type, including adjusting the environmental temperature of the connection line interface of the mass spectrometer according to the Pearson correlation coefficient of the corresponding position of the discrete error data of proteomics data and the detection voltage fluctuation value on the detection time sequence of proteomics data,
[0010] or, adjusting the sampling frequency according to the deviation amplitude of the deviation trend;
[0011] obtaining the peptide segment peak width of the protein detected according to the detection execution mode, and calculating the peak width standard deviation according to the peptide segment peak width;
[0012] determining the insertion position of the virtual sampling point in the detection time sequence according to the peak width standard deviation;
[0013] comparing the proteomics data and the standard data again according to the insertion position of the virtual sampling point, and marking the real abnormal data based on the re-comparison result.
[0014] Further, determining the discrete type of proteomics data based on the discrete characteristics of the discrete error data comprises:
[0015] equally dividing the detection time sequence into a plurality of detection time periods;
[0016] obtaining the position of the discrete error data on a plurality of sets of proteomics data according to the detection time sequence;
[0017] respectively calculating the data density in the plurality of detection time periods;
[0018] fitting the change slope of a plurality of data densities;
[0019] If the slope of change is less than or equal to a preset first slope, then the first detection execution method is executed;
[0020] If the slope of change is greater than or equal to the preset second slope, then the second detection execution mode is executed.
[0021] Furthermore, the data density is the ratio of the number of discrete error data within a single detection period to the total number of data.
[0022] Furthermore, the horizontal axis of the slope of change represents the detection time, and the vertical axis of the slope of change represents the data density.
[0023] Furthermore, the ambient temperature of the mass spectrometer's connection interface is adjusted based on the Pearson correlation coefficient between the discrete error data of the proteomics data and the corresponding position of the detection voltage fluctuation value in the detection time sequence of the proteomics data, including:
[0024] The discrete error data and the detected voltage fluctuation value are acquired respectively;
[0025] Calculate the Pearson correlation coefficient between the discrete error data and the detected voltage fluctuation value;
[0026] If the Pearson correlation coefficient is greater than or equal to a preset coefficient, then the ambient temperature of the mass spectrometer's connection interface is reduced.
[0027] Furthermore, the detection voltage fluctuation value includes the peak voltage value when the detection voltage is greater than a preset second voltage and the trough voltage value when the detection voltage is less than or equal to a preset first voltage during the detection of the proteomics data.
[0028] Further, adjusting the sampling frequency based on the magnitude of the deviation trend includes:
[0029] Obtain the deviation magnitude of the deviation trend;
[0030] The deviation range is compared with the preset range;
[0031] If the deviation is greater than or equal to the preset deviation, the sampling frequency within a single detection period is reduced.
[0032] Furthermore, the deviation magnitude is the absolute value of the difference between the chromatographic column corresponding to the discrete error data and the chromatographic column corresponding to the standard data.
[0033] Further, determining the insertion position of the virtual sampling point in the detection time sequence based on the peak width standard deviation includes:
[0034] Obtain the peak width standard deviation within several of the aforementioned detection periods;
[0035] The peak width standard deviation is compared with the preset standard deviation;
[0036] During the detection period in which the peak width standard deviation is greater than or equal to the preset standard deviation, several virtual sampling points of equal duration are inserted.
[0037] Furthermore, the number of virtual sampling points inserted is positively correlated with the peak width standard deviation.
[0038] Compared with existing technologies, the beneficial effects of this invention are as follows: The method of this invention identifies discrete points in the data by comparing proteomics data with standard data. During the acquisition of several sets of proteomics data, as the detection time increases, the temperature at the wire interface of the detection device rises, leading to voltage fluctuations. These voltage fluctuations cause deviations in the acquired proteomics data. The discrete error data caused by voltage fluctuations exhibits a dispersed distribution. The first detection execution mode adjusts the ambient temperature to maintain temperature stability in the detection device, thereby reducing proteomics data deviations caused by voltage fluctuations. Simultaneously, lowering the ambient temperature helps stabilize the detected protein components, improving detection accuracy. Furthermore, as the detection time increases, the composition of subsequent protein samples changes due to storage time, resulting in an increasing number of samples distributed sequentially according to the detection time. The second detection execution mode adjusts the sampling frequency to capture the changes in proteomics data caused by sample storage... The gradual changes in proteins over time can be addressed by various factors. For example, residual proteases in the sample may slowly hydrolyze peptide bonds, or hydrophobic interactions between protein molecules may lead to the slow formation of aggregates. Solvent evaporation can also cause discrete error data to become increasingly concentrated in the period before the end of the detection. Adjusting the sampling frequency reduces high-frequency harmonic interference, effectively reducing detection errors caused by prolonged sample storage time. Because the detection execution method reduces error sources, protein signals become more stable. When discrete types show deviations, such as signal drift or abrupt changes, aliasing or signal distortion can occur, making it impossible to accurately capture the dynamic changes in protein signals. Therefore, filtering and noise reduction in data processing can amplify errors. By calculating the peak width standard deviation and determining the insertion position of virtual sampling points, additional sampling points can be added in areas of drastic data changes. Limited by hardware sampling frequency and detection cycle, it is impossible to capture transient abnormal signals at key time points of protein activity changes. By generating supplementary data points, the perception blind spots of the time series are filled, improving the continuity of detection.
[0039] Furthermore, the method of the present invention determines the discrete type. The first detection execution mode adjusts the ambient temperature to address the data deviation problem caused by voltage fluctuations, thereby reducing the impact of detection errors caused by voltage fluctuations during the operation of the detection equipment from the source. The second detection execution mode adjusts the sampling frequency to adapt to changes in data distribution characteristics, ensuring that more key information is captured in areas of drastic data changes, thus improving detection efficiency.
[0040] Furthermore, the method of the present invention reduces the ambient temperature of the mass spectrometer's connection interface by lowering the Pearson correlation coefficient, thereby reducing voltage fluctuations caused by temperature increases in the detection equipment and improving the stability of data acquisition. At the same time, lowering the ambient temperature helps to slow down the degradation rate of protein samples during the detection process, preserving the original characteristics of the samples and improving the accuracy of the detection results.
[0041] Furthermore, the method of the present invention reduces the sampling frequency according to the deviation magnitude of the deviation trend, thereby reducing the operating load of the equipment when the detection time or sample storage time is extended. This avoids hardware wear and data redundancy problems caused by long-term high-frequency sampling. Since the protein dispersion error data increases over time due to the longer storage time of proteins detected later than proteins detected earlier when the mass spectrometer detects several samples, abnormal signals can still be captured within the critical time period. This reduces harmonic interference from high-frequency sampling while ensuring detection accuracy, providing clearer time series characteristics for subsequent data analysis, reducing noise interference introduced by excessive data density, and further improving the accuracy of anomaly detection.
[0042] Furthermore, the method described in this invention solves the problem of time series perception blind spots caused by hardware limitations or excessively large sampling intervals by calculating the peak width standard deviation and inserting virtual sampling points. Since changes in the detection environment during detection execution can lead to changes in proteomics data, increasing the number of virtual sampling points can compensate for information that may be missed during actual sampling, thereby improving data integrity and enhancing the ability to identify transient abnormal signals. Attached Figure Description
[0043] Figure 1 This is an overall flowchart of the intelligent proteomics data anomaly detection and analysis method according to an embodiment of the present invention;
[0044] Figure 2 This is a flowchart illustrating the determination of discrete types in the intelligent proteomics data anomaly detection and analysis method according to an embodiment of the present invention;
[0045] Figure 3This is a flowchart illustrating the adjustment of the ambient temperature of the connection interface of the mass spectrometer in the intelligent proteomics data anomaly detection and analysis method according to an embodiment of the present invention.
[0046] Figure 4 This is a flowchart illustrating the adjustment of sampling frequency in the intelligent proteomics data anomaly detection and analysis method according to an embodiment of the present invention. Detailed Implementation
[0047] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0048] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0049] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0050] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0051] Please see Figure 1 , Figure 2 , Figure 3 as well as Figure 4 The diagrams shown are, respectively, the overall flowchart, the flowchart for determining the discrete type, the flowchart for adjusting the ambient temperature of the mass spectrometer's connection interface, and the flowchart for adjusting the sampling frequency of the intelligent proteomics data anomaly detection and analysis method according to an embodiment of the present invention. The intelligent proteomics data anomaly detection and analysis method according to an embodiment of the present invention includes:
[0052] Proteins are detected using mass spectrometry to obtain several sets of proteomics data;
[0053] Several sets of proteomics data were compared with standard data, and discrete error data were determined based on the comparison results.
[0054] The discrete type of proteomics data is determined based on the discrete characteristics of the discrete error data.
[0055] The detection execution method is determined according to the discrete type, including adjusting the ambient temperature of the mass spectrometer's connection interface based on the Pearson correlation coefficient between the discrete error data of the proteomics data and the corresponding position of the detection voltage fluctuation value in the detection time sequence of the proteomics data.
[0056] Alternatively, the sampling frequency can be adjusted according to the magnitude of the deviation trend;
[0057] Obtain the peak width of the peptides detected according to the detection execution method, and calculate the standard deviation of the peak width based on the peak width;
[0058] The insertion position of the virtual sampling point in the detection time sequence is determined based on the peak width standard deviation.
[0059] The proteomics data is compared with the standard data again according to the insertion position of the virtual sampling point, and real abnormal data are marked based on the results of the second comparison.
[0060] Specifically, the types of proteins detected include Large ribosomal subunit protein uL4, Plasma serine protease inhibitor, Bone morphogenetic protein 1, Amine oxidase [copper-containing] 3, T-complex protein 1 subunit alpha (Fragment), and Villin-1.
[0061] Specifically, proteomics data includes protein name, peptide sequence, peptide peak area, mass-to-charge ratio, peptide peak width, and chromatographic column.
[0062] Specifically, several sets of proteomics data refer to multiple independent sets of proteomics data obtained by mass spectrometry analysis of multiple different samples.
[0063] Specifically, the standard data consists of known historical protein data for the corresponding protein type.
[0064] Specifically, discrete error data involves comparing each data point in several sets of proteomics data with the corresponding data in standard data to identify data points that deviate from the standard data. Examples of this include missing peptides, phosphorylation / glycosylation site errors, and...
[0065] Specifically, the connection interface is the mass spectrometer's socket and its local circuit area, with the local circuit area not exceeding 0.04 m². 2 The ambient temperature of the connection line interface is regulated by a cooler located below the local circuit area.
[0066] Specifically, the method for calculating the peak width standard deviation is a well-known technique in the art and will not be elaborated here.
[0067] In implementation, the method of this invention identifies discrete points in the data by comparing proteomics data with standard data. During the acquisition of several sets of proteomics data, the temperature at the wire interface of the detection device rises with increasing detection time, leading to voltage fluctuations. These voltage fluctuations cause deviations in the acquired proteomics data. The discrete error data caused by voltage fluctuations exhibits a dispersed distribution. The first detection execution mode adjusts the ambient temperature to maintain temperature stability in the detection device, thereby reducing proteomics data deviations caused by voltage fluctuations. Simultaneously, lowering the ambient temperature helps stabilize the detected protein components, improving detection accuracy. Furthermore, as the detection time increases, the composition of subsequent protein samples changes due to storage time, resulting in an increasing number of samples distributed sequentially. The second detection execution mode adjusts the sampling frequency to capture protein changes in the proteomics data caused by sample storage time. The gradual changes in protein composition, such as the slow hydrolysis of peptide bonds by residual proteases in the sample, the slow formation of aggregates due to hydrophobic interactions of protein molecules, or the concentration of discrete error data in the period before the end of detection due to solvent evaporation, can be addressed by adjusting the sampling frequency to reduce high-frequency harmonic interference, thereby effectively reducing detection errors caused by prolonged sample storage time. Because the detection execution method reduces error sources, the protein signal becomes more stable. When discrete types show deviations, such as signal drift or abrupt changes, aliasing or signal distortion occurs, making it impossible to accurately capture the dynamic changes of the protein signal. Therefore, filtering and noise reduction in data processing can amplify errors. By calculating the peak width standard deviation and determining the insertion position of virtual sampling points, sampling points can be added in areas of drastic data changes. Limited by hardware sampling frequency and detection cycle, it is impossible to capture transient abnormal signals at key time points of protein activity changes. By generating supplementary data points, the perception blind spots of the time series are filled, improving the continuity of detection.
[0068] Specifically, determining the discrete type of proteomics data based on the discrete characteristics of the discrete error data includes:
[0069] The detection sequence is divided into several detection time periods;
[0070] Obtain the position of the discrete error data in several sets of proteomics data according to the detection time sequence;
[0071] Calculate the data density within each of the several detection time periods;
[0072] Fit the slope of the change in several of the aforementioned data densities;
[0073] If the slope of change is less than or equal to a preset first slope, then the first detection execution method is executed;
[0074] If the slope of change is greater than or equal to the preset second slope, then the second detection execution mode is executed.
[0075] Specifically, under the conditions of a trypsin to protein mass ratio of 1:100 and a mass spectrometer resolution of 80000 FWHM, the general range of the preset first slope is (0, 0.1], the general range of the preset second slope is [0.4, 1], the preferred embodiment of the preset first slope is 0.05, and the preferred embodiment of the preset second slope is 0.6.
[0076] Those skilled in the art will understand that the range of preset first slope and preset second slope provided in this embodiment, as well as the preferred embodiment, are the values that best address the technical problem solved by the technical solution of the present invention under the conditions that the mass ratio of trypsin to protein is 1:100 and the mass spectrometer resolution is 80000 FWHM. In actual applications or experiments, those skilled in the art can make adaptive adjustments to the preset first slope and preset second slope according to the actual application environment and application scenario.
[0077] In practice, the method of the present invention determines the discrete type. The first detection execution mode adjusts the ambient temperature to deal with the data deviation problem caused by voltage fluctuations, thereby reducing the impact of detection errors caused by voltage fluctuations during the operation of the detection equipment from the source. The second detection execution mode adjusts the sampling frequency to adapt to changes in data distribution characteristics, ensuring that more key information is captured in areas of drastic data changes, thus improving detection efficiency.
[0078] Specifically, the data density is the ratio of the number of discrete error data within a single detection period to the total number of data.
[0079] Specifically, the horizontal axis of the slope change represents the detection time, and the vertical axis of the slope change represents the data density.
[0080] Specifically, the ambient temperature of the mass spectrometer's connection interface is adjusted based on the Pearson correlation coefficient between the discrete error data of the proteomics data and the corresponding position of the detection voltage fluctuation value in the detection time sequence of the proteomics data, including:
[0081] The discrete error data and the detected voltage fluctuation value are acquired respectively;
[0082] Calculate the Pearson correlation coefficient between the discrete error data and the detected voltage fluctuation value;
[0083] If the Pearson correlation coefficient is greater than or equal to a preset coefficient, then the ambient temperature of the mass spectrometer's connection interface is reduced.
[0084] Specifically, the detection voltage fluctuation value includes the peak voltage value where the detection voltage is greater than a preset second voltage and the trough voltage value where the detection voltage is less than or equal to a preset first voltage during the detection of the proteomics data.
[0085] Specifically, under the conditions that the mass ratio of trypsin to protein is 1:100 and the external voltage of the mass spectrometer is 220V, the general range of the preset coefficient is [0.4, 1], and the preferred embodiment of the preset coefficient is 0.6.
[0086] Those skilled in the art will understand that the range of preset coefficients and preferred embodiments provided in this embodiment are the values that best address the technical problem solved by the present invention under the conditions of a trypsin-to-protein mass ratio of 1:100 and an external voltage of 220V for the mass spectrometer. In actual applications or experiments, those skilled in the art can make adaptive adjustments to the preset coefficients according to the actual application environment and application scenario.
[0087] In practice, the method of the present invention reduces the ambient temperature of the mass spectrometer's connection interface based on the Pearson correlation coefficient, thereby reducing voltage fluctuations caused by temperature increases in the detection equipment and improving the stability of data acquisition. At the same time, lowering the ambient temperature helps to slow down the degradation rate of protein samples during the detection process, preserving the original characteristics of the samples and improving the accuracy of the detection results.
[0088] Specifically, adjusting the sampling frequency based on the magnitude of the deviation from the stated deviation trend includes:
[0089] Obtain the deviation magnitude of the deviation trend;
[0090] The deviation range is compared with the preset range;
[0091] If the deviation is greater than or equal to the preset deviation, the sampling frequency within a single detection period is reduced.
[0092] Specifically, the deviation magnitude is the absolute value of the difference between the chromatographic column corresponding to the discrete error data and the chromatographic column corresponding to the standard data.
[0093] In practice, the method of this invention reduces the sampling frequency based on the deviation magnitude of the deviation trend. This reduces the operating load of the equipment when the detection time or sample storage time is extended, avoiding hardware wear and data redundancy problems caused by long-term high-frequency sampling. Since the storage time of proteins detected later is longer than that of proteins detected earlier when the mass spectrometer detects several samples, the discrete error data of proteins increases over time. Abnormal signals can still be captured in the critical time period, thereby reducing harmonic interference from high-frequency sampling while ensuring detection accuracy. This provides clearer time series characteristics for subsequent data analysis, reduces noise interference introduced by excessive data density, and further improves the accuracy of anomaly detection.
[0094] Specifically, determining the insertion position of the virtual sampling point in the detection time series based on the peak width standard deviation includes:
[0095] Obtain the peak width standard deviation within several of the aforementioned detection periods;
[0096] The peak width standard deviation is compared with the preset standard deviation;
[0097] During the detection period in which the peak width standard deviation is greater than or equal to the preset standard deviation, several virtual sampling points of equal duration are inserted.
[0098] Specifically, the number of virtual sampling points inserted is positively correlated with the peak width standard deviation.
[0099] In practice, the method described in this invention solves the problem of time series perception blind spots caused by hardware limitations or excessively large sampling intervals by calculating the peak width standard deviation and inserting virtual sampling points. Since changes in the detection environment during detection execution can lead to changes in proteomics data, increasing the number of virtual sampling points can compensate for information that may be missed during actual sampling, improve data integrity, and enhance the ability to identify transient abnormal signals.
[0100] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. An intelligent proteomics data anomaly detection and analysis method, characterized in that, The method comprises: detecting proteins by a mass spectrometer to obtain a plurality of sets of proteomic data; comparing the plurality of sets of proteomic data with standard data, and determining discrete error data based on the comparison result; determining a discrete type of the proteomic data based on discrete characteristics of the discrete error data, comprising dividing a detection time sequence into a plurality of detection time periods, obtaining positions of the discrete error data on the plurality of sets of proteomic data according to the detection time sequence, respectively calculating data densities in the plurality of detection time periods, fitting variation slopes of the plurality of data densities, and if the variation slope is less than or equal to a preset first slope, performing a first detection execution mode, and if the variation slope is greater than or equal to a preset second slope, performing a second detection execution mode; determining a detection execution mode according to the discrete type, the first detection execution mode being adjusting an ambient temperature of a connection line interface of the mass spectrometer according to a Pearson correlation coefficient of the discrete error data of the proteomic data and a detection voltage fluctuation value at corresponding positions on a detection time sequence of the proteomic data, and the second detection execution mode being adjusting a sampling frequency according to a deviation amplitude of a deviation trend; obtaining a peptide segment peak width of a protein detected according to the detection execution mode, and calculating a peak width standard deviation according to the peptide segment peak width; determining an insertion position of a virtual sampling point in the detection time sequence according to the peak width standard deviation; respectively comparing the proteomic data and the standard data again according to the insertion position of the virtual sampling point, and marking real abnormal data based on the re-comparison result. 2.The intelligent proteomics data anomaly detection and analysis method of claim 1, wherein, The data density is a ratio of a number of discrete error data in a single detection time period to a number of all data. 3.The intelligent proteomics data anomaly detection and analysis method of claim 2, characterized in that, The horizontal coordinate of the variation slope is a detection time, and the vertical coordinate of the variation slope is a data density. 4.The intelligent proteomics data anomaly detection and analysis method of claim 3, wherein, The adjusting of the ambient temperature of the connection line interface of the mass spectrometer according to the Pearson correlation coefficient of the discrete error data of the proteomic data and the detection voltage fluctuation value at the corresponding positions on the detection time sequence of the proteomic data comprises: respectively obtaining the discrete error data and the detection voltage fluctuation value; calculating the Pearson correlation coefficient of the discrete error data and the detection voltage fluctuation value; if the Pearson correlation coefficient is greater than or equal to a preset coefficient, reducing the ambient temperature of the connection line interface of the mass spectrometer. 5.The intelligent proteomics data anomaly detection and analysis method of claim 4, wherein, The detection voltage fluctuation value comprises a voltage peak value of a detection voltage greater than a preset second voltage and a voltage valley value of a detection voltage less than or equal to a preset first voltage in a process of detecting the proteomic data. 6.The intelligent proteomics data anomaly detection and analysis method of claim 5, wherein, The adjusting of the sampling frequency according to the deviation amplitude of the deviation trend comprises: obtaining the deviation amplitude of the deviation trend; comparing the deviation amplitude with a preset amplitude; if the deviation amplitude is greater than or equal to the preset amplitude, reducing the sampling frequency in a single detection time period.
7. The intelligent proteomics data anomaly detection and analysis method of claim 6, wherein, The deviation amplitude is an absolute value of a difference between a chromatographic column corresponding to the discrete error data and a chromatographic column corresponding to the standard data. 8.The intelligent proteomics data anomaly detection and analysis method of claim 7, wherein, The determining of the insertion position of the virtual sampling point in the detection time sequence according to the peak width standard deviation comprises: obtaining peak width standard deviations in the plurality of detection time periods; comparing the peak width standard deviation with a preset standard deviation; inserting a plurality of virtual sampling points with equal time length in the detection period when the peak width standard deviation is greater than or equal to the preset standard deviation. 9.The intelligent proteomics data anomaly detection and analysis method of claim 8, wherein, The number of the inserted virtual sampling points is positively correlated with the peak width standard deviation.
Citation Information
Patent Citations
Protein anomaly detection method based on super-large scale evolutionary algorithm and hardware acceleration
CN114298214A
Sequencing nucleic acids by barcoding in discrete entities
CN115011670A