Mass spectrometer data processing method, system and server
By building a local database and implementing a specific search process, the problem of low accuracy in commercial spectral libraries was solved, enabling high-accuracy substance identification through mass spectrometry data processing.
Patent Information
- Application Number
- CN202511604875.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-11-05
AI Technical Summary
In existing mass spectrometer data processing methods, the accuracy of data information from commercial spectral libraries is not high, the number of substances that can be accurately identified in the sample is relatively small, and different chromatographic conditions have a significant impact on the results, making it difficult to adjust the details of the search process and parameters.
A pre-built local database is constructed, and a specific search process is used to search the self-built local database. Peak correction and time alignment are performed using the chromatographic mass spectrometry information of the mass spectrometer, and feature matching is performed in combination with the local database to improve the accuracy of the search.
It improves the accuracy of database searches, ensures accurate identification of substances in samples, and reduces the impact of changes in chromatographic conditions.
Smart Images

Figure CN121070878B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of mass spectrometer data processing, in particular to a mass spectrometer data processing method, system and server. BACKGROUND
[0002] Non-targeted identification of metabolites in samples, the most important step in the process is called searching. The key two parts of searching are the technology and method of searching, and the second is the reference metabolite spectrum library. Previously, users usually use commercial searching software (such as Compound Discovery) provided by upstream instrument manufacturers and related spectrum libraries (such as mzCloud, mzVault, ChemSpider). But the data information accuracy of such spectrum library is not high, the number of substances that can be accurately identified in the sample accounts for a small proportion, and because of different chromatographic conditions, retention time and other parameters cannot be reused, etc. Factors have adverse effects on the results.
[0003] In summary, the details of the existing searching process and the parameters are difficult to adjust, which affects the processing results of searching. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a mass spectrometer data processing method, system and server, which uses a pre-constructed local database to complete the construction of the spectrum library, and realizes the searching process of the self-built local database through a specific searching process, thereby improving the accuracy of searching.
[0005] In a first aspect, the present application provides a mass spectrometer data processing method, which comprises:
[0006] Obtaining the original data corresponding to the mass spectrometer, and converting the original data into a first format file according to a first formatting rule;
[0007] Determining the chromatography mass spectrometry information corresponding to the sample in the mass spectrometer by using the first format file, and obtaining the first chromatography peak signal corresponding to the sample based on the chromatography mass spectrometry information;
[0008] Performing peak correction on the first chromatography peak signal based on the correction strategy corresponding to the sample, and obtaining the peak correlation result corresponding to the first chromatography peak signal after time alignment processing of the sample corresponding to the corrected first chromatography peak signal;
[0009] After determining the feature result of the sample according to the peak correlation result, performing feature matching between the feature result and the preset local database, and obtaining the data processing result corresponding to the sample.
[0010] Optionally, after the step of obtaining the raw data corresponding to the mass spectrometer and converting the raw data into a first format file according to a first formatting rule, the method further comprises:
[0011] obtaining a non-target library mode parameter corresponding to the mass spectrometer; wherein the non-target library mode parameter at least includes a mass priority parameter and a performance priority parameter;
[0012] If the non-target library mode parameter is the performance priority parameter, the first format file is split according to the positive and negative modes based on the parent ion charge parameter corresponding to the raw data, and the split result is used to update the first format file.
[0013] Optionally, if the non-target library mode parameter is the mass priority parameter, the first format file is read according to the positive and negative modes.
[0014] Optionally, after obtaining the raw data corresponding to the mass spectrometer, the method further comprises:
[0015] converting the raw data into a second format file according to a second formatting rule;
[0016] determining the feature result of the sample by using the second format file, and performing feature matching between the feature result and the local database to obtain the data processing result corresponding to the sample.
[0017] Optionally, the construction process of the local database comprises:
[0018] obtaining an initial database associated with the type data of the sample from a preset cloud mass spectrometry database;
[0019] obtaining a background noise library corresponding to the initial database, controlling the background noise library to be merged with the initial database to obtain an extended database corresponding to the sample, and saving the extended database in the local server;
[0020] verifying the extended database by using a preset standard sample, generating attribute information of the extended database according to the verification result, and constructing the local database by using the attribute information.
[0021] Optionally, the step of obtaining the background noise library corresponding to the initial database, controlling the background noise library to be merged with the initial database to obtain the extended database corresponding to the sample, and saving the extended database in the local server comprises:
[0022] According to the screening strategy corresponding to the sample, the database with a score exceeding a preset threshold in the initial database is determined as a to-be-screened database;
[0023] obtaining a name string of the to-be-screened database, determining a replacement string corresponding to the name string based on the type data of the sample, and updating the replacement string to the screening database.
[0024] The initial database corresponds to a background noise database, and the background noise database is merged with the screening database to obtain an extended database corresponding to the sample.
[0025] Optionally, based on the correction strategy corresponding to the sample, the peak of the primary chromatographic peak signal is corrected, and after the time alignment processing of the sample corresponding to the corrected primary chromatographic peak signal, the peak correlation result corresponding to the primary chromatographic peak signal is obtained.
[0026] The primary chromatographic peak signal corresponding to the chromatographic peak is obtained, and the correction strategy is determined based on the mass deviation of the adjacent chromatographic peak, and the primary chromatographic peak signal is corrected by using the correction strategy;
[0027] The time parameter of the sample is determined, and the reference sample is determined from the sample based on the time parameter, and the sample is time-aligned according to the reference sample;
[0028] The deviation data corresponding to the sample that has completed the time alignment processing is obtained according to the type data of the sample, and the deviation data is used to determine the peak correlation result corresponding to the primary chromatographic peak signal.
[0029] Optionally, after the feature result of the sample is determined according to the peak correlation result, the feature result is matched with the preset local database to obtain the data processing result corresponding to the sample.
[0030] The retention time feature and the nuclear-cytoplasmic ratio feature corresponding to the sample are determined based on the peak correlation result, and the feature result of the sample is determined according to the retention time feature and the nuclear-cytoplasmic ratio feature;
[0031] The local database is processed by using the feature result to perform a cyclic search database feature matching processing, and the secondary database data corresponding to the feature result is queried from the local database;
[0032] The secondary feature data of the feature result and the spectrum entropy relationship corresponding to the secondary database data are determined to obtain the secondary information, and the data processing result corresponding to the sample is determined based on the secondary information.
[0033] In a second aspect, the present application provides a mass spectrometer data processing system, which comprises:
[0034] A data conversion module is configured to obtain the original data corresponding to the mass spectrometer, and convert the original data into a first format file according to a first formatting rule;
[0035] A peak detection module is configured to determine the chromatographic mass spectrometry information corresponding to the sample in the mass spectrometer by using the first format file, and obtain the primary chromatographic peak signal corresponding to the sample based on the chromatographic mass spectrometry information;
[0036] The peak processing module is configured to perform peak correction on the primary chromatographic peak signal based on a correction strategy corresponding to the sample, perform time alignment processing on the sample corresponding to the corrected primary chromatographic peak signal, and obtain a peak correlation result corresponding to the primary chromatographic peak signal.
[0037] The feature matching module is configured to perform feature matching between the feature result and a preset local database after determining the feature result of the sample according to the peak correlation result, and obtain a data processing result corresponding to the sample.
[0038] In a third aspect, an embodiment of the present application further provides a server, comprising a processor and a memory, the memory storing computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to implement the steps of the mass spectrometer data processing method provided in the first aspect.
[0039] In a fourth aspect, an embodiment of the present application further provides a storage medium, the storage medium storing computer executable instructions, and the computer executable instructions, when invoked and executed by a processor, cause the processor to implement the steps of the mass spectrometer data processing method provided in the first aspect.
[0040] The mass spectrometer data processing method, system and server provided by the embodiment of the present application, in the process of searching the database for non-target identification of metabolites in the sample, first obtains the original data corresponding to the mass spectrometer, and converts the original data into a first format file according to a first formatting rule; then determines the chromatographic mass spectrometry information corresponding to the sample in the mass spectrometer by using the first format file, and obtains the primary chromatographic peak signal corresponding to the sample based on the chromatographic mass spectrometry information; then performs peak correction on the primary chromatographic peak signal based on the correction strategy corresponding to the sample, and performs time alignment processing on the sample corresponding to the corrected primary chromatographic peak signal, and obtains the peak correlation result corresponding to the primary chromatographic peak signal; finally, after determining the feature result of the sample according to the peak correlation result, performing feature matching between the feature result and a preset local database, and obtaining the data processing result corresponding to the sample. The method uses a pre-constructed local database to complete the construction of the spectrum database, and realizes the process of searching the self-built local database through a specific search process, thereby improving the accuracy of the search.
[0041] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and achieved by the structure particularly pointed out in the description, claims and drawings.
[0042] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the accompanying drawings needed to be used in the specific embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0044] Figure 1 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0045] Figure 2 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0046] Figure 3 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0047] Figure 4 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0048] Figure 5 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0049] Figure 6 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0050] Figure 7 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0051] Figure 8 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0052] Figure 9 The flow chart of a mass spectrometer data processing method provided by the embodiment of the present application is shown in the figure.
[0053] Figure 10 The structural schematic diagram of a mass spectrometer data processing system provided by the embodiment of the present application is shown in the figure.
[0054] Figure 11 The structural schematic diagram of a server provided by the embodiment of the present application is shown in the figure.
[0055] Icon:
[0056] 1010 - Data Conversion Module; 1020 - Peak Detection Module; 1030 - Peak Processing Module; 1040 - Feature Matching Module;
[0057] 101 - Processor; 102 - Memory; 103 - Bus; 104 - Communication interface. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0059] Non-targeted identification of metabolites in samples involves a crucial step known as library search. Library search hinges on two key components: the search technique and methodology, and the reference metabolite spectral library. Previously, users typically relied on commercial library search software (such as Compound Discovery) and related spectral libraries (such as mzCloud, mzVault, and ChemSpider) provided by upstream instrument manufacturers. However, these libraries suffer from low accuracy, accurately identifying only a small percentage of substances in the sample. Furthermore, factors such as the inability to reuse parameters like retention time (RT) due to varying chromatographic conditions negatively impact the results. It is evident that the detailed methods and parameters involved in existing library search processes are difficult to adjust, thus affecting the processing outcomes. Therefore, this invention provides a mass spectrometer data processing method, system, and server. This method utilizes a pre-built local database to construct a spectral library and employs a specific search process to search this self-built local database, thereby improving search accuracy.
[0060] To facilitate understanding of this embodiment, a mass spectrometer data processing method disclosed in this embodiment will first be described in detail, such as... Figure 1 As shown, the method includes:
[0061] Step S101: Obtain the raw data from the mass spectrometer and convert it into a first format file according to the first formatting rule.
[0062] The raw data from the mass spectrometer is usually stored in an instrument-specific format (such as Waters' RAW, Thermo's RAW, Bruker's d format, etc.), containing basic information such as ion signal intensity, mass-to-charge ratio (m / z), retention time (RT), as well as metadata such as instrument parameters, experimental conditions, etc. Data acquisition methods include direct export from the mass spectrometer control system, extraction through laboratory information management system (LIMS) interface, or batch reading from local storage path.
[0063] The first formatting rule is based on common mass spectrometry data standards (such as mzML, mzXML, mzData, etc.) to develop conversion rules, which requires defining data structures (such as spectrum list, scan event, peak list), metadata fields (instrument model, acquisition parameters, sample information), and storage formats (binary or text format). Conversion implementation can be performed through related scripts or open source tools (such as MSConvert, ProteoWizard), and data integrity needs to be checked (such as missing value filling, outlier marking) during the process, and an index file containing sample ID, acquisition time, etc. is generated to facilitate subsequent tracing.
[0064] In step S102, the chromatography-mass spectrometry information corresponding to the sample in the mass spectrometer is determined using the first format file, and the first chromatography peak signal corresponding to the sample is obtained based on the chromatography-mass spectrometry information.
[0065] The chromatography-mass spectrometry information analysis process can extract a three-dimensional data matrix from the first format file, separate overlapping peaks through chromatography peak recognition algorithms (such as XIC extraction, baseline correction, peak detection), and obtain the chromatogram (RT-intensity) and mass spectrum (m / z-intensity) of each sample. Key parameters may include chromatography peak broadening threshold, signal-to-noise ratio (SNR) filtering conditions, isotope peak cluster recognition rules, etc., to exclude noise and false positive signals.
[0066] The first chromatography peak signal extraction is mainly aimed at target compounds or full-spectrum scanning data, and by extracting ion chromatogram (XIC) or mass chromatogram (MC), the chromatography peak corresponding to the specific m / z is located, and the peak area, retention time, peak height, half-peak width, etc. Quantitative / qualitative parameters are obtained. For complex samples (such as biological fluids, environmental samples), background subtraction algorithms (such as moving window method, polynomial fitting) need to be combined to remove matrix interference and improve the accuracy of peak signals.
[0067] In step S103, the first chromatography peak signal is peak-corrected based on the correction strategy corresponding to the sample, and after time alignment processing of the sample corresponding to the corrected first chromatography peak signal, the peak correlation result corresponding to the first chromatography peak signal is obtained.
[0068] In the process of correcting the signal of the primary chromatographic peak, the internal standard method (adding a standard with a known concentration) or the external standard method can be used for intensity normalization to eliminate batch-to-batch differences in view of the sensitivity drift of the instrument. The mass-to-charge axis can also be calibrated by using calibration compounds (such as PFPP and Naformate) with accurate masses to ensure that the mass-to-charge ratio accuracy reaches the ppm level. The peak shape distortion can be corrected by mathematical fitting (such as Gaussian function and Lorentz function) of the tailing and fronting peaks. Then, the statistical model (such as principal component analysis PCA and partial least squares discriminant analysis PLS-DA) is used to identify the system error, and the machine learning algorithm (such as random forest) can be used to predict the correction parameters.
[0069] In the time alignment process and the peak association process, the chromatogram of a representative sample can be selected as a reference based on the reference spectrum, and the chromatographic peaks of other samples can be aligned to the reference time axis by using the dynamic time warping (DTW) algorithm. In the retention time prediction process, the retention time can be predicted by using the structural information of the compound (such as carbon chain length and functional group) or the machine learning model (such as neural network) to correct the actual RT deviation. In the peak association result generation process, the peak table (containing sample ID, RT, m / z, and intensity) is constructed after alignment, and the peak matching algorithm (such as MZmine and XCMS) is used to identify the peaks of the same compound in different samples to generate an association matrix (sample-peak-intensity) for subsequent feature analysis.
[0070] In step S104, after determining the feature results of the samples according to the peak association results, the feature results are matched with the preset local database to obtain the data processing results corresponding to the samples.
[0071] In the feature result determination process, the differential features (such as peaks with VIP value > 1.0 and p < 0.05) can be selected from the peak association matrix, and the m / z, RT, and isotope distribution parameters of the feature peaks can be extracted by peak intensity normalization (such as total ion flow TIC normalization and median normalization) and standardization processing. For unknown compounds, the molecular formula can be inferred by isotope peak cluster analysis (such as 13C and 15N isotope abundance ratio), or the possible structure candidates can be generated by combining the fragment ion prediction tool (such as ChemSpider and MassBank).
[0072] In the process of cyclic search and feature matching and acquisition of secondary information, the local database is used for implementation. In the construction process of the local database, the mass spectrum (primary and secondary MS / MS), retention time, molecular formula, and structural information of known compounds are included. The local database can be based on public databases (such as HMDB, Metlin, and MassBank) or self-built standard sample library.
[0073] The library searching strategy involves primary library searching and secondary library searching. Specifically, the primary library searching is matched with the database by m / z and RT to screen candidate compounds. In the secondary library searching process, if the sample contains secondary mass spectrum data (such as DDA), the characteristic peak corresponding fragment ion spectrum is compared with the MS / MS spectrum in the database, and the matching score (such as Cosine similarity, DotProduct) is calculated. For unmatched characteristic peaks, the derivatives are predicted by combining the metabolic transformation rules (such as hydroxylation, methylation), and the library searching is performed again to improve the identification coverage.
[0074] Based on the library searching matching score, isotopic matching degree, retention time deviation and other parameters, the candidate compounds are sorted by confidence (such as Level 1-4 identification level), and a data processing result report containing the compound name, structure, matching score and quantitative value is generated.
[0075] Optionally, after the step S101 of obtaining the raw data corresponding to the mass spectrometer and converting the raw data into a first format file according to the first formatting rule, as shown in the following formula (I), the method further comprises: Figure 2
[0076] In the step S201, the non-target library searching mode parameters corresponding to the mass spectrometer are obtained, wherein the non-target library searching mode parameters at least include mass priority parameters and performance priority parameters.
[0077] In the step S202, if the non-target library searching mode parameters are the performance priority parameters, the first format file is split according to the positive and negative modes based on the parent ion charge parameters corresponding to the raw data, and the first format file is updated by using the split result.
[0078] The mass priority parameters and the performance priority parameters involved in the non-target library searching mode parameters correspond to the high-quality non-target library searching process and the high-performance non-target library searching process, respectively. If the non-target library searching mode parameters are the performance priority parameters, the high-performance non-target library searching process needs to be performed, and at this time, the first format file is split according to the positive and negative modes based on the parent ion charge parameters corresponding to the raw data, and the first format file is updated by using the split result. The advantage of this process is mainly low cost, about half of the high-quality process. The biological sample is not collected in the on-line stage, and more secondary information is collected for substance identification
[0079] Optionally, if the non-target search mode parameter is a quality priority parameter, the first format file is read in positive and negative modes. Specifically, a sample preparation mass spectrometry on-machine detection file is prepared and converted into a general mzxml format, and the positive and negative modes are separated. The positive and negative modes refer to the principle of mass spectrometry detection. Mass spectrometry detection is mainly to obtain the mass-to-charge ratio of ions entering the mass spectrometer. Most endogenous substances do not carry charges by themselves and need to be charged in the ion source of the mass spectrometer. Due to the different characteristics of each substance, the charges carried by each substance are different. Therefore, each sample is tested in positive and negative modes to see what types of substances can be identified in the positive and negative modes. The positive and negative modes are converted separately, and the search library is searched separately. For details, refer to the flowchart of another mass spectrometer data processing method shown in Figure 3 .
[0080] Optionally, after obtaining the raw data of the mass spectrometer, as shown in Figure 4 , the method further comprises:
[0081] Step S401, converting the raw data into a second format file according to a second formatting rule;
[0082] Step S402, determining the characteristic results of the sample by using the second format file, and performing characteristic matching between the characteristic results and a local database to obtain a data processing result corresponding to the sample.
[0083] Referring to Figure 3 , after obtaining the raw data raw of the mass spectrometer, the raw data is converted into a second format file, i.e., an mgf file. Then, the characteristics of the raw data are annotated, the characteristic results of the sample are determined by using the second format file, and the characteristic results are subjected to a cyclic search library processing with a local database to determine secondary information corresponding to the characteristic results, and the data processing result corresponding to the sample is determined based on the secondary information.
[0084] The database used in the search library process is locally established. Optionally, the construction process of the local database includes the steps shown in Figure 5 .
[0085] Step S501, obtaining an initial database associated with the type data of the sample from a preset cloud mass spectrometry database;
[0086] Step S502, obtaining a background noise library corresponding to the initial database, controlling the background noise library to be merged with the initial database to obtain an extended database corresponding to the sample, and saving the extended database in a local server;
[0087] Step S503, verifying the extended database by using a preset standard sample, generating attribute information of the extended database according to a verification result, and constructing the local database by using the attribute information.
[0088] Specifically, according to the type data of the sample, the data is searched from the cloud mass spectrum database such as the NIST library, and the data is reserved according to specific requirements to obtain an initial database. Then, a background noise library corresponding to the initial database is obtained, and the initial database and the background noise library are merged to obtain an extended database corresponding to the sample, and the extended database is saved in the local server. The local extended database needs to be verified by a specific standard sample, so as to obtain attribute information of the extended database according to the verification result, and then the attribute information is used to perfect the extended database, so as to complete the construction of the local database.
[0089] Optionally, the step S502 of obtaining the background noise library corresponding to the initial database, controlling the background noise library and the initial database to be merged to obtain the extended database corresponding to the sample, and saving the extended database in the local server, as shown in Figure 6 , includes:
[0090] In step S601, the database with a score exceeding a preset threshold in the initial database is determined as a screening database according to a screening strategy corresponding to the sample.
[0091] In step S602, a name string of the screening database is obtained, a replacement string corresponding to the name string is determined based on the type data of the sample, and the replacement string is updated to the screening database.
[0092] In step S603, a background noise library corresponding to the initial database is determined, and the background noise library and the screening database are merged to obtain an extended database corresponding to the sample.
[0093] Specifically, the extended database is realized by expanding the background noise library based on the initial database. First, the high-score database in the initial database is determined as a screening database according to the screening strategy corresponding to the sample, and then the name string of the screening database is obtained. The replacement rule corresponding to the sample is obtained through the type data of the sample, and then the replacement string corresponding to the name string is obtained by using the replacement rule, and the replacement string is updated to the screening database to complete the replacement of the name. Then, the background noise library corresponding to the initial database is determined, and the background noise library and the screening database are merged to obtain the corresponding extended database. The above process can be referred to as another flowchart of the construction process of the local database shown in Figure 7 .
[0094] Optionally, the step S103 of obtaining the peak association result corresponding to the primary chromatographic peak signal based on the correction strategy corresponding to the sample to perform peak correction on the primary chromatographic peak signal, and performing time alignment processing on the sample corresponding to the corrected primary chromatographic peak signal, as shown in Figure 8 , includes:
[0095] Step S801, obtain the chromatographic peak corresponding to the primary chromatographic peak signal, determine a correction strategy based on the mass deviation of adjacent chromatographic peaks, and perform peak correction on the primary chromatographic peak signal using the correction strategy;
[0096] Step S802, determine the time parameter of the sample, and determine the reference sample from the sample based on the time parameter, and perform time alignment processing on the sample according to the reference sample;
[0097] Step S803, obtain the deviation data corresponding to the sample after time alignment processing according to the type data of the sample, and determine the peak correlation result corresponding to the primary chromatographic peak signal using the deviation data.
[0098] The above process involves three steps of peak correction, retention time alignment and peak correlation. Specifically, first, the peak detection process is performed. Peak detection is mainly used to extract the primary chromatographic peak signal in the sample. The signal extraction considers the influence of signal-to-noise ratio, the minimum intensity threshold of peak signal, and the filtering of peak width. In actual scenarios, a signal-to-noise ratio of more than 20 is often required, and the peak width of normal metabolites in chromatography needs to be between 10-50 seconds, otherwise it may be affected by noise.
[0099] Peak correction is necessary when the directly extracted chromatographic peak may have poor peak shape or overlapping peaks. Peak optimization or correction is needed. Generally, the mass deviation at the center of two peaks and the height difference ratio of overlapping peaks are considered. On a high-resolution mass spectrometry platform, the relative mass deviation can be set to 5-10 ppm. The calculation method is: (peak 1 mass-peak 2 mass) / peak 2 mass*10^6.
[0100] Retention time alignment is an important step to avoid false positive ions. Specifically, one or a group of samples are selected as references, and other samples are aligned to this group of samples using linear smoothing method. According to experience, while performing metabolic detection, strict technical repeat incubation is required, which is interspersed in actual biological samples as quality control samples, referred to as QC samples. At the same time, this group of samples belongs to technical repeat samples, so they can be used as reference sample group for retention time alignment.
[0101] Peak correlation mainly correlates peaks in different samples that may belong to the same substance. Since each sample is extracted separately, and the mass spectrometer has a detection tolerance, the instrument coefficient obtained for the same substance will have some deviation. Peaks with mass and retention time deviation within a certain range are correlated to obtain the characteristics of the entire sample group. Each characteristic is a potential metabolite that can be identified.
[0102] In the actual scene, peak filling can also be performed, and due to the sensitivity of the mass spectrometer, the same feature can exist in different technical repeated samples and cannot be detected. Therefore, the peak integration quantitative result is also filled by referring to the signal condition of the adjacent peak.
[0103] Optionally, after determining the feature result of the sample according to the peak correlation result, the feature result is matched with a preset local database to obtain a data processing result corresponding to the sample, as shown in step S104. Figure 9 as shown, comprising:
[0104] In step S901, the retention time feature and the nuclear mass ratio feature corresponding to the sample are determined based on the peak correlation result, and the feature result of the sample is determined according to the retention time feature and the nuclear mass ratio feature.
[0105] In step S902, the feature result is used for cyclic database searching feature matching processing of the local database, and the secondary database data corresponding to the feature result is queried from the local database.
[0106] In step S903, the secondary feature data of the feature result and the spectrum entropy relationship corresponding to the secondary database data are determined to obtain secondary information, and the data processing result corresponding to the sample is determined based on the secondary information.
[0107] Step S104 mainly realizes the feature annotation process in Figure 3 The feature annotation is to combine the feature and the database to identify which metabolite the feature may be. The identified features are cyclically searched in the database, and the main logic is: the rt (retention time) and mz (nuclear mass ratio) of the feature are obtained, the secondary information of the matched sample is queried, the substances in the first mass deviation range and the secondary information are obtained by searching the database, the secondary information of the feature and the secondary information in the database are matched by spectrum entropy method, and finally the data processing result is output.
[0108] From the mass spectrometer data processing method mentioned in the above embodiment, it can be seen that the method uses the pre-constructed local database to complete the construction of the spectrum library, and realizes the searching process of the self-built local database through a specific searching process, thereby improving the searching accuracy.
[0109] Corresponding to the mass spectrometer data processing method provided in the foregoing embodiment, an embodiment of the present application provides a mass spectrometer data processing system, as shown in Figure 10 The system comprises:
[0110] The data conversion module 1010 is configured to obtain the original data corresponding to the mass spectrometer, and convert the original data into a first format file according to a first formatting rule.
[0111] The peak detection module 1020 is configured to determine chromatography mass spectrum information corresponding to the sample in the mass spectrometer by using the file in the first format, and obtain a first chromatographic peak signal corresponding to the sample based on the chromatography mass spectrum information.
[0112] The peak processing module 1030 is configured to perform peak correction on the first chromatographic peak signal based on a correction strategy corresponding to the sample, and perform time alignment processing on the sample corresponding to the corrected first chromatographic peak signal, and obtain a peak correlation result corresponding to the first chromatographic peak signal.
[0113] The feature matching module 1040 is configured to determine a feature result of the sample according to the peak correlation result, perform feature matching between the feature result and a preset local database, and obtain a data processing result corresponding to the sample.
[0114] It can be known from the mass spectrometer data processing system mentioned in the above embodiment that the system completes the construction of the spectrum library by using the pre-constructed local database, and realizes the process of searching the self-constructed local database through a specific search database process, thereby improving the accuracy of the search database.
[0115] The mass spectrometer data processing system provided in the embodiment has the same implementation principle and technical effects as the foregoing mass spectrometer data processing method embodiments. For brevity of description, the part of the system embodiment not mentioned can be referred to the corresponding content in the foregoing mass spectrometer data processing method embodiments.
[0116] The embodiment also provides a server, and a structure diagram of the server is shown in Figure 11 The device includes a processor 101 and a memory 102; wherein the memory 102 is configured to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the steps of the foregoing mass spectrometer data processing method.
[0117] Figure 11 The server shown in the figure also includes a bus 103 and a communication interface 104, and the processor 101, the communication interface 104 and the memory 102 are connected through the bus 103.
[0118] The memory 102 can include a high-speed random access memory (RAM), and can also include a non-volatile memory, for example, at least one disk memory. The bus 103 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 11 only one bidirectional arrow is used to represent one bus or one type of bus.
[0119] The communication interface 104 is configured to connect with at least one user terminal and other network units through a network interface, and transmit the encapsulated IPv4 packet or the IPv4 packet to the user terminal through the network interface.
[0120] The processor 101 can be an integrated circuit chip with processing capability. In implementation process, each step of the above method can be completed by integrated logic circuit or software form instruction in the processor 101. The processor 101 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; or can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component. Each method, step and logic block in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as a hardware code processor to execute, or be executed by a combination of hardware and software modules in the code processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage medium in the art. The storage medium is located in the memory 102, and the processor 101 reads the information in the memory 102 and combines the hardware to complete the steps of the method of the above embodiments.
[0121] The embodiments of the present application also provide a storage medium, and the storage medium stores a computer program. When the computer program is run by a processor, the steps of the mass spectrometer data processing method in the above embodiments are executed.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, devices, and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0123] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0124] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0125] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0126] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, but not to limit the same. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that any person skilled in the art can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features, within the technical scope disclosed by the present application. The modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of mass spectrometer data processing, characterized by, The method comprises: acquiring raw data corresponding to a mass spectrometer and converting the raw data into a first format file according to a first formatting rule; determining chromatography-mass spectrometry information corresponding to a sample in the mass spectrometer based on the first format file, and acquiring a primary chromatography peak signal corresponding to the sample based on the chromatography-mass spectrometry information; performing peak correction on the primary chromatography peak signal based on a correction strategy corresponding to the sample, and performing time alignment processing on the sample corresponding to the primary chromatography peak signal after the peak correction, to obtain a peak correlation result corresponding to the primary chromatography peak signal; determining a feature result of the sample according to the peak correlation result, and performing feature matching between the feature result and a preset local database to obtain a data processing result corresponding to the sample; determining a feature result of the sample according to the peak correlation result, and performing feature matching between the feature result and a preset local database to obtain a data processing result corresponding to the sample; determining a retention time feature and a nucleo-cytoplasmic ratio feature corresponding to the sample based on the peak correlation result, and determining the feature result of the sample according to the retention time feature and the nucleo-cytoplasmic ratio feature; performing cyclic database search feature matching processing on the local database by using the feature result, and querying a secondary database data corresponding to the feature result from the local database; determining secondary information of the secondary feature data of the feature result and the spectral entropy relationship corresponding to the secondary database data, and determining the data processing result corresponding to the sample based on the secondary information.
2. The mass spectrometer data processing method of claim 1, wherein, After the step of acquiring raw data corresponding to a mass spectrometer and converting the raw data into a first format file according to a first formatting rule, the method further comprises: acquiring non-target database search mode parameters corresponding to the mass spectrometer; wherein the non-target database search mode parameters at least include: mass priority parameters and performance priority parameters; if the non-target database search mode parameters are the performance priority parameters, then splitting the first format file according to a positive-negative mode based on parent ion charge parameters corresponding to the raw data, and updating the first format file by using the splitting result.
3. The mass spectrometer data processing method of claim 2, wherein, if the non-target database search mode parameters are the mass priority parameters, then reading the first format file according to the positive-negative mode.
4. The mass spectrometer data processing method of claim 1, wherein, After acquiring raw data corresponding to a mass spectrometer, the method further comprises: converting the raw data into a second format file according to a second formatting rule; determining the feature result of the sample by using the second format file, and performing feature matching between the feature result and the local database to obtain a data processing result corresponding to the sample.
5. The mass spectrometer data processing method of claim 1, wherein, The construction process of the local database comprises: acquiring an initial database associated with type data of the sample from a preset cloud mass spectrometry database; acquiring background noise libraries corresponding to the initial database, controlling the background noise libraries to be merged with the initial database to obtain an expanded database corresponding to the sample, and saving the expanded database in a local server; After the extended database is verified by using the preset standard sample, attribute information of the extended database is generated according to a verification result, and the local database is constructed by using the attribute information.
6. The mass spectrometer data processing method of claim 5, wherein, The step of obtaining the background noise library corresponding to the initial database, controlling the background noise library to be merged with the initial database to obtain the extended database corresponding to the sample, and saving the extended database in the local server comprises: According to the screening strategy corresponding to the sample, a database with a score exceeding a preset threshold in the initial database is determined as a database to be screened. A name string of the database to be screened is obtained, a replacement string corresponding to the name string is determined based on the type data of the sample, and the replacement string is updated to the database to be screened. A background noise library corresponding to the initial database is determined, and the background noise library is merged with the database to be screened to obtain the extended database corresponding to the sample.
7. The mass spectrometer data processing method of claim 1, wherein, The step of obtaining the peak correlation result corresponding to the primary chromatographic peak signal based on the correction strategy corresponding to the sample for peak correction of the primary chromatographic peak signal and time alignment processing of the sample corresponding to the primary chromatographic peak signal after the primary chromatographic peak signal is corrected comprises: The chromatographic peak corresponding to the primary chromatographic peak signal is obtained, and the correction strategy is determined based on the mass deviation of adjacent chromatographic peaks, and the primary chromatographic peak signal is corrected by using the correction strategy; The time parameter of the sample is determined, and the reference sample is determined from the sample based on the time parameter, and the sample is time-aligned according to the reference sample; The deviation data corresponding to the sample that has completed time alignment processing is obtained according to the type data of the sample, and the peak correlation result corresponding to the primary chromatographic peak signal is determined by using the deviation data.
8. A mass spectrometer data processing system characterized by, The system comprises: A data conversion module is configured to obtain raw data from a mass spectrometer and convert the raw data into a first format file according to a first formatting rule; A peak detection module is configured to determine chromatography mass spectrometry information corresponding to a sample in the mass spectrometer by using the first format file, and obtain a primary chromatographic peak signal corresponding to the sample based on the chromatography mass spectrometry information; A peak processing module is configured to obtain a peak correlation result corresponding to the primary chromatographic peak signal by performing peak correction on the primary chromatographic peak signal based on a correction strategy corresponding to the sample and performing time alignment processing on the sample corresponding to the primary chromatographic peak signal after the primary chromatographic peak signal is corrected; A feature matching module is configured to determine a feature result of the sample according to the peak correlation result, perform feature matching between the feature result and a preset local database, and obtain a data processing result corresponding to the sample. The feature matching module is further configured to: determine the retention time feature and the nucleus-to-cytoplasm ratio feature corresponding to the sample based on the peak correlation result, and determine the feature result of the sample according to the retention time feature and the nucleus-to-cytoplasm ratio feature; perform cyclic database search feature matching processing on the local database by using the feature result, and query the secondary database data corresponding to the feature result from the local database; determine secondary information by correlating the secondary feature data of the feature result with the spectrum entropy of the secondary database data, and determine the data processing result corresponding to the sample based on the secondary information.
9. A server, characterized by A processor and a memory are included, the memory stores computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to implement the steps of the mass spectrometer data processing method in any one of claims 1 to 7. A processor and a memory are included, the memory stores computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to implement the steps of the mass spectrometer data processing method in any one of claims 1 to 7.
Citation Information
Patent Citations
Biomass spectrometry data standardization method, molecular map database and sharing method
CN115394362A