Automatic quantitative analysis method, system and equipment for mass spectrum detection data, medium and program product
By automating the construction and screening of standard curves for mass spectrometry detection data, the problem of instability in quantitative results caused by reliance on manual operation in existing technologies has been solved. This achieves fully automated processing of mass spectrometry detection data and improves the stability and consistency of quantitative results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI BIOTREE
- Filing Date
- 2026-04-13
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the construction of standard curves for mass spectrometry detection data relies on manual operation and lacks an adaptive optimization mechanism, resulting in poor stability and consistency of quantitative results.
Mass spectrometry detection data is obtained through automated quantitative analysis methods, at least two candidate standard curves are constructed, a target standard curve is screened using multidimensional evaluation indicators, and internal standard data is automatically switched when preset conditions are not met, and candidate standard curves are reconstructed until a target standard curve that meets the conditions is obtained.
It has achieved full automation of the quantitative analysis process of mass spectrometry detection data, improved the rationality and adaptability of standard curve selection, reduced the intensity of manual operation, and enhanced the stability and consistency of quantitative results.
Smart Images

Figure CN122016992A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mass spectrometry data processing technology, and in particular to an automated quantitative analysis method, system, device, medium and program product for mass spectrometry detection data. Background Technology
[0002] In analytical applications based on gas / liquid chromatography-tandem mass spectrometry (GC-MS), mass spectrometry data is typically used for the quantitative analysis of target analytes in samples. During data processing, a standard curve is usually constructed to establish the correlation between the target analyte's response signal and its concentration, and quantitative calculations are then performed on the sample based on this standard curve.
[0003] In existing technologies, standard curve construction is typically based on preset internal standards and fitting methods. When the constructed standard curve does not meet the preset requirements, manual adjustments to the selection of internal standards or fitting conditions are usually necessary to obtain a suitable standard curve model. However, this process heavily relies on operator experience, and the selected standard curves may differ between different operators or batches, affecting the stability of quantitative results. Furthermore, quantitative calculations are usually performed based on a single standard curve during standard curve construction, lacking a mechanism for comparing and filtering results from multiple standard curves. When the constructed standard curve does not meet the conditions, there is a lack of effective alternative strategies, making adaptive selection of the standard curve difficult and resulting in poor consistency of quantitative results.
[0004] Therefore, how to automatically construct and adaptively select standard curves during the quantitative analysis of mass spectrometry detection data, thereby improving the stability and consistency of quantitative results, has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides an automated quantitative analysis method, system, device, medium, and program product for mass spectrometry detection data, which at least solves the technical problems in existing technologies where the construction and selection of standard curves rely on manual methods and lack an adaptive optimization mechanism, resulting in poor stability and consistency of quantitative results.
[0006] To achieve the above objectives and other advantages, some embodiments of this application provide the following aspects:
[0007] In a first aspect, some embodiments of this application provide an automated quantitative analysis method for mass spectrometry detection data, including:
[0008] Acquire mass spectrometry detection data of mixed standard samples, and determine the response signal of the target analyte based on the mass spectrometry detection data;
[0009] Based on the response signal, at least two candidate standard curves are constructed to describe the relationship between the response signal and the concentration of the target analyte;
[0010] The candidate standard curves are screened based on at least one evaluation index to determine the target standard curve that meets the preset conditions.
[0011] When the candidate standard curve does not meet the preset conditions, the internal standard data is automatically switched, the candidate standard curve is reconstructed and the filtering is performed until the target standard curve is obtained.
[0012] Based on the target standard curve, the quantitative concentration of the sample to be analyzed is calculated to obtain the quantitative analysis results of the target analyte in the sample to be analyzed.
[0013] Secondly, some embodiments of this application also provide an automated quantitative analysis system for mass spectrometry detection data, including:
[0014] The data acquisition module is used to acquire mass spectrometry detection data of mixed standard samples;
[0015] The signal processing module is used to determine the response signal of the target analyte based on the mass spectrometry detection data;
[0016] The standard curve construction module is used to construct at least two candidate standard curves based on the response signal to describe the relationship between the response signal and the concentration of the target analyte.
[0017] The curve screening module is used to screen the candidate standard curves based on at least one evaluation index to determine the target standard curve that meets the preset conditions.
[0018] The internal standard switching module is used to automatically switch the internal standard data, reconstruct the candidate standard curve and perform filtering when the candidate standard curve does not meet the preset conditions, until the target standard curve is obtained.
[0019] The quantitative analysis module is used to calculate the quantitative concentration of the sample to be analyzed based on the target standard curve, so as to obtain the quantitative analysis results of the target analyte in the sample to be analyzed.
[0020] Thirdly, some embodiments of this application also provide an electronic device, the electronic device comprising:
[0021] One or more processors; and a memory storing computer program instructions that, when executed, cause the processors to perform an automated quantitative analysis method for mass spectrometry detection data as described above.
[0022] Fourthly, some embodiments of this application also provide a computer-readable storage medium having a computer program and / or instructions stored thereon, which, when executed by a processor, implement an automated quantitative analysis method for mass spectrometry detection data as described above.
[0023] Fifthly, some embodiments of this application also provide a computer program product, including a computer program and / or instructions, which, when executed by a processor, implement an automated quantitative analysis method for mass spectrometry detection data as described above.
[0024] Compared with existing technologies, the solution provided in this application automatically determines the response signal of the target analyte in the mixed standard sample based on mass spectrometry detection data, constructs at least two candidate standard curves, and filters the candidate standard curves using evaluation indicators. When a candidate standard curve does not meet preset conditions, the internal standard data is automatically switched, and a new candidate standard curve is constructed and filtered again until the target standard curve is obtained. The quantitative concentration of the sample to be analyzed is then calculated based on the target standard curve, thus forming a closed-loop optimization mechanism based on adaptive switching of multiple candidate standard curves and internal standards. This achieves fully automated processing of the quantitative analysis process of mass spectrometry detection data. Through this method, a standard curve meeting preset conditions can be automatically obtained without manual intervention, avoiding quantitative deviations caused by relying on a single standard curve or a fixed internal standard. This improves the rationality and adaptability of standard curve selection, while significantly reducing manual operation intensity and improving data processing efficiency. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other implementation methods can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is one of the flowcharts illustrating an automated quantitative analysis method for mass spectrometry detection data provided in this application embodiment;
[0027] Figure 2 This is a second schematic flowchart of an automated quantitative analysis method for mass spectrometry detection data provided in the embodiments of this application;
[0028] Figure 3 This diagram illustrates the data loading and target analyte selection interface in the mass spectrometry detection data processing workflow.
[0029] Figure 4 A schematic diagram of the chromatographic peak overlay of the target analyte is shown;
[0030] Figure 5 A schematic diagram of chromatographic peaks used for peak screening and anomaly identification is shown;
[0031] Figure 6 This is a schematic diagram of the structure of an automated quantitative analysis system for mass spectrometry detection data provided in an embodiment of this application;
[0032] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0034] Some embodiments of this application relate to an automated quantitative analysis method for mass spectrometry detection data, see below. Figure 1 As shown, the method may include the following steps:
[0035] Step S1: Obtain mass spectrometry detection data of the mixed standard sample, and determine the response signal of the target analyte in the mixed standard sample based on the mass spectrometry detection data.
[0036] A mixed standard sample refers to a standard sample prepared by combining multiple target analytes of known concentrations in a preset ratio. It is used to characterize the response characteristics of each target analyte under uniform detection conditions. A mixed standard sample can be prepared by preparing standard solutions of multiple target analytes according to a set concentration gradient, with each target analyte corresponding to a known theoretical concentration value. In practical applications, multiple concentration gradient points can be set for different target analytes, forming a set of standard point data for fitting, thereby providing basic data for subsequently establishing the correspondence between response signals and concentrations. In some embodiments, the mixed standard sample may also include an internal standard substance for correcting detection errors. This internal standard substance has a preset matching relationship with the target analytes in terms of retention time, ionization characteristics, or response stability, to improve the stability and comparability of the response signal.
[0037] Mass spectrometry (MS) data can be obtained by detecting mixed standard samples using MS equipment. MS equipment can include liquid chromatography-mass spectrometry (LC-MS), gas chromatography-mass spectrometry (GC-MS), or ion mobility spectrometry (IMS). MS data can be represented as a time-varying signal data stream, containing signal intensity information related to the target analyte. MS data acquisition modes can utilize multiple reaction monitoring (MRM), parallel reaction monitoring (PARM), data-independent acquisition, and data-dependent acquisition methods.
[0038] After acquiring mass spectrometry detection data, signal processing can be performed to determine the response signal of the target analyte in the mixed standard sample. The signal processing procedure may include preprocessing the mass spectrometry detection data, signal separation, feature extraction, and signal modeling, to progressively extract the effective signal components corresponding to the target analyte from the original signal and reduce the impact of noise interference and signal overlap on the results.
[0039] In a preferred embodiment, step S1, determining the response signal of the target analyte in the mixed standard sample based on mass spectrometry detection data, specifically includes:
[0040] Step S101: Based on the characteristics of data acquisition frequency and local scanning density, the original mass spectrometry signal of the mixed standard sample is adaptively smoothed, and the signal is separated by dynamic background noise modeling in order to extract the target signal components.
[0041] Data acquisition frequency characterizes the sampling interval of the signal in the time dimension; local scan density characterizes the data distribution of the signal in different time intervals. Adaptive smoothing parameters can be constructed based on the acquisition frequency and local scan density of mass spectrometry data. Specifically, the local scan density of different time intervals can be determined by calculating the local sampling interval or the number of data points per unit time of the mass spectrometry signal in the time dimension, and the size of the smoothing window can be dynamically adjusted accordingly. In regions with high local scan density, a smaller smoothing scale is used to preserve the detailed structure of the signal; in regions with low local scan density, a larger smoothing scale is used to enhance the denoising effect, thereby achieving adaptive smoothing processing of non-uniformly sampled signals. Adaptive smoothing processing can be implemented using a smoothing method based on a Gaussian kernel function or other smoothing methods, and the Gaussian kernel parameter or smoothing window is adaptively adjusted according to the local scan density to balance noise suppression and signal structure preservation.
[0042] In some implementations, the smoothing process can be constrained and adjusted by incorporating local variation characteristics of the signal to avoid over-smoothing of the feature structure in areas of drastic signal change, thereby improving the accuracy of signal processing.
[0043] After adaptive smoothing, a dynamic background noise model can be constructed to model and separate background noise from the raw mass spectrometry signal. The dynamic background noise model can adaptively estimate and dynamically update the background signal based on its temporal variation trend, characterizing baseline drift and low-frequency variation features. By separating the background noise from the smoothed signal, the target signal component containing information about the target analyte can be obtained. Dynamic background noise modeling can be achieved through baseline estimation or background subtraction, and the baseline can be dynamically adjusted according to the signal characteristics in different time intervals, thereby improving the accuracy of background signal separation. By separating the background noise from the smoothed signal, the target signal component containing information about the target analyte can be obtained.
[0044] In some implementations, the parameters of the adaptive smoothing process and dynamic background noise modeling process can be adjusted according to the degree of signal change or signal-to-noise ratio in different time intervals, so as to improve the adaptability of signals to different sample conditions or different detection environments, thereby further improving the accuracy and stability of target signal extraction.
[0045] Step S102: Perform multi-scale feature analysis on the target signal components in the time-scale joint domain, and locate the candidate peaks of the target analyte by identifying the extreme value propagation trajectories at different scales.
[0046] The time-scale joint domain refers to the domain that characterizes the response characteristics of a signal at different scales within the joint analysis space composed of the time and scale dimensions. The time dimension characterizes the temporal location of the signal, while the scale dimension characterizes the resolution or smoothness of the signal analysis. At smaller scales, the detailed structure of the signal can be reflected; at larger scales, the overall trend of the signal's change can be reflected. By analyzing the same signal at multiple scales, the signal response characteristics corresponding to different scales can be obtained. Multi-scale feature analysis can be achieved using wavelet transform methods. By mapping the target signal components to wavelet coefficient spaces at different scales, the response characteristics of the signal at different scales can be obtained, thereby achieving multi-resolution characterization of the signal in the time-scale joint domain.
[0047] Local extremum information can be extracted based on the distribution characteristics of wavelet coefficients at different scales, and correlation analysis can be performed on extremum points at different scales. Specifically, extremum points at different scales can be matched and connected based on their proximity in the time dimension and their continuous change characteristics in the scale dimension, thereby establishing the propagation relationship of extremum points in the process of scale change and forming the corresponding extremum propagation trajectory. Furthermore, candidate peaks can be located based on the extremum propagation trajectory. When an extremum point exists continuously at multiple scales and its position change in the time dimension is within a preset range, the propagation trajectory corresponding to the extremum point can be determined as a stable trajectory, and its corresponding time position can be determined as the candidate peak position. Conversely, extremum points that appear only at a single scale or a few scales and do not have continuity in the process of scale change can be determined as noise or spurious peaks and eliminated.
[0048] In some implementations, candidate peaks can be further screened based on characteristics such as the number of consecutive scales, cross-scale stability, or scale coverage of extreme value propagation trajectories to improve the accuracy and robustness of candidate peak localization. For example, a minimum continuous scale threshold can be constrained to determine the degree of continuous existence of extreme value propagation trajectories in the scale dimension. When the number of consecutive scales of a certain extreme value propagation trajectory in the scale dimension is greater than or equal to the minimum continuous scale threshold, the extreme value propagation trajectory is determined as a valid trajectory, and its corresponding time position is determined as the candidate peak position; when the number of consecutive scales is less than the minimum continuous scale threshold, the corresponding extreme value propagation trajectory is determined as an unstable trajectory and is eliminated.
[0049] In some implementations, the cross-scale stability of the extreme value propagation trajectory can be evaluated by considering the degree of positional shift of the trajectory at different scales. When the temporal position change of the extreme point at different scales is less than a preset shift threshold, the extreme value propagation trajectory can be determined to have high stability, thereby further improving the reliability of candidate peak identification.
[0050] Step S103: Based on the preset time axis drift correction model and combined with the response signal of the internal standard data, perform multidimensional verification on the candidate peaks to determine the characteristic peaks corresponding to the target analyte.
[0051] A time axis drift correction model can be established based on the correspondence between a reference time and the actual detection time. Specifically, the actual detection time of the internal standard in the current sample can be obtained and compared with its preset reference time to determine the time offset; a time mapping relationship can be constructed based on the time offset to correct the time axis drift. In some implementations, the time mapping relationship can be fitted based on the time offset information of multiple internal standards to characterize the nonlinear drift characteristics in different time intervals; and the time axis drift correction model can be dynamically updated based on data from different samples or different detection batches to adapt to the time offset caused by changes in detection conditions.
[0052] After obtaining the time axis drift correction model, the time position of the candidate peak on the original time axis can be input into the time axis drift correction model to obtain the corresponding corrected time position. A matching judgment is then made based on the corrected time position and the reference time of the target analyte. When the corrected time position of the candidate peak falls within the preset time tolerance range, the candidate peak is determined to meet the time consistency constraint verification.
[0053] Based on satisfying the time consistency constraint, the response consistency of candidate peaks can be verified by combining the response signals of internal standard data. Specifically, based on the correlation between the response signals of candidate peaks and the response signals of internal standard data, it can be determined whether the response signals of candidate peaks meet preset response constraints, such as whether they are within a preset ratio range or trend range, thereby eliminating spurious peaks caused by noise or abnormal fluctuations.
[0054] Furthermore, cross-sample consistency verification of candidate peaks can be performed based on their response changes across multiple samples. Specifically, the response change trend of a candidate peak in different samples can be analyzed and compared with the change trend of the internal standard signal. When the response change trend of a candidate peak meets a preset consistency condition, the candidate peak is determined to have stability characteristics.
[0055] In some implementations, the time axis drift correction model can be adaptively modified based on changes in the internal standard signal, ensuring that the time correction result remains consistent with the current detection environment, thereby further improving the accuracy and stability of characteristic peak identification. Through the aforementioned multi-dimensional verification mechanism based on time correction and internal standard constraints, reliable identification of characteristic peaks of the target analyte can be achieved even in the presence of time drift and signal fluctuations.
[0056] Step S104: Model the characteristic peaks and decompose the overlapping peaks based on iterative optimization separation to obtain the response signal of the target analyte in the mixed standard sample.
[0057] Modeling is used to mathematically represent the morphological characteristics of characteristic peaks. It can fit the characteristic peaks based on a preset peak shape function to obtain the model parameters corresponding to the characteristic peaks. The model parameters include parameters such as peak area, peak height, peak width, and peak position. Among them, the peak area can be used as a basic quantity to characterize the response signal of the target analyte.
[0058] When a signal corresponding to a feature peak has multiple overlapping components, the feature peak can be separated. This separation can be achieved using deconvolution methods. By constructing a superposition model of multiple peak shape functions, the overlapping signals are inverted and decomposed to separate the composite peak into multiple sub-peaks corresponding to different components. In some implementations, the deconvolution process can be solved based on an iterative optimization strategy. By minimizing the residual between the model fitting result and the actual signal, parameters such as the peak position, peak width, and peak area of each sub-peak are updated. Specifically, the difference between the current model fitting result and the actual signal can be used as the residual signal, and the model parameters can be gradually corrected based on the residual signal to gradually approximate the true signal distribution. When the residual signal is less than a preset threshold, the deconvolution process is considered to have reached convergence, thus obtaining the final parameters of each sub-peak.
[0059] In some implementations, constraints can be introduced during the deconvolution process to limit the relative positional relationship between sub-peaks or the range of peak shape parameter variations, thereby avoiding decomposition results that do not conform to actual detection characteristics and improving the stability and reliability of peak segmentation results.
[0060] Through the above-mentioned deconvolution-based separation process, effective separation of multiple component signals can be achieved even in the case of peak overlap or signal interference, obtaining the independent response signal corresponding to each target analyte, thereby improving the accuracy of subsequent quantitative analysis.
[0061] The solution provided in this application constructs a mass spectrometry signal processing mechanism combining multi-scale identification, multi-dimensional verification, and deconvolution separation through steps S101 to S104, achieving multi-level constraint processing of mass spectrometry data from the signal layer to the structural layer. Specifically, multi-scale feature analysis improves the completeness of candidate peak detection, multi-dimensional consistency verification enhances the reliability of feature peak determination, and deconvolution-based separation further solves the problem of inaccurate analysis of overlapping peaks. Through the aforementioned multi-layer constraint and collaborative processing mechanism, the technical defects of existing technologies, such as susceptibility to noise interference, unstable peak identification, and difficulty in separating overlapping peaks, are effectively overcome, thereby significantly improving the accuracy and stability of quantitative mass spectrometry analysis.
[0062] In actual mass spectrometry data processing, factors such as low target analyte concentration, signal noise interference, or unclear peak shape may cause some standard samples to fail to identify characteristic peaks corresponding to the target analyte, resulting in missing response signals for some standard samples and affecting the completeness and accuracy of subsequent standard curve construction. To address this, and avoid incomplete data or deviations during standard curve construction due to missing characteristic peaks, in an optional embodiment, a compensation mechanism is implemented for samples in the mixed standard samples where the characteristic peak corresponding to the target analyte is not identified. This compensation mechanism includes:
[0063] Step S1101: Construct a reference sample set based on the standard samples that belong to the same detection batch as the mixed standard samples and have identified the characteristic peaks of the target analyte, and obtain the retention time distribution sequence corresponding to each characteristic peak in the reference sample set.
[0064] The peak identification results of standard samples belonging to the same detection batch as the mixed standard samples can be traversed to select standard samples that meet preset identification conditions as a reference sample set. The same detection batch can refer to a group of standard sample data acquired continuously under the same detection conditions to ensure consistency among samples in chromatographic separation conditions and mass spectrometry response environment. Identification conditions may include: consistency constraints on the characteristic peak passage time of the corresponding target analyte and multidimensional verification constraints. During the screening process, samples that fail verification or whose quality evaluation is below a threshold can be removed, retaining only valid standard samples that successfully identify the characteristic peak of the target analyte to construct the reference sample set. Subsequently, the retention time of each sample is extracted from the reference sample set to form a retention time distribution sequence according to sample order or time order. Furthermore, the retention time distribution sequence can be preliminarily cleaned; for example, based on statistical dispersion indicators or preset deviation thresholds, abnormal retention times can be identified and removed to improve the stability of subsequent statistical calculations.
[0065] Step S1102: Calculate the median retention time based on the retention time distribution sequence.
[0066] The retention time distribution sequence is sorted, and the median value is selected as the median retention time. When the sample size is odd, the median retention time is the value in the middle position after sorting; when the sample size is even, the average of the two middle values can be used as the median retention time. In some implementations, to improve robustness, the retention time distribution sequence can be constrained before calculating the median, for example, only data falling within a preset time range can be retained, or outliers can be screened out based on quantile intervals, thereby reducing the impact of outlier samples on the statistical results. Through the above processing, a representative time position estimate of the target analyte under the current batch conditions can be obtained.
[0067] Step S1103: Construct an integration window based on the median retention time and a preset integration bandwidth.
[0068] In some implementations, the median retention time can be used as the center position, and an integration window can be constructed by combining it with a preset integration bandwidth, the expression of which is:
[0069]
[0070]
[0071] in, This represents the median retention time of the target analyte in the reference sample set; Indicates the integral bandwidth; Indicates the start time position of the integration window; Indicates the end time position of the integration window.
[0072] The above expression is used to construct a symmetrical time interval centered on the median retention time, thereby defining the extraction range of the target signal on the time axis.
[0073] Step S1104: Map the integration window to the corresponding position on the time axis of the sample where no characteristic peak was identified, and perform integration processing within the integration window to obtain the compensation response signal of the target analyte in the sample where no characteristic peak was identified.
[0074] In some implementations, the integration window can be mapped to the time axis interval of samples where no characteristic peaks were identified, and the signal sequence within that interval can be extracted, as expressed by:
[0075]
[0076] in, Indicates the first The set of signals within the integration window for samples in which no characteristic peaks of the target analyte were identified; Indicates at a point in time Signal strength at the location; Represents discrete sampling time points; This indicates the time interval corresponding to the integration window.
[0077] Subsequently, the signal set can be integrated to obtain the compensated response signal of the target analyte, the expression of which is:
[0078]
[0079] in, Indicates the first The compensation response signal for a sample where no characteristic peak of the target analyte was identified; the summation operation is used to accumulate the signal intensity within the integration window to approximate the peak area of the target analyte in that sample.
[0080] In the solution provided in this application embodiment, a compensation processing mechanism based on batch statistical characteristics is introduced through steps S1101 to S1104, transforming the problem of unidentified characteristic peaks into a signal extraction problem based on time constraints. By constructing a reference sample set and modeling the retention time distribution, statistical estimation of the target analyte's time position is achieved; the introduction of the median retention time improves the anti-interference capability of time positioning; and by combining integral window constraints and signal integration processing, effective response signals can still be extracted even in cases of weak signals or the absence of obvious peaks. Through these mechanisms, the applicability of mass spectrometry quantitative analysis in abnormal samples or low-abundance scenarios is significantly improved, enhancing data utilization and result stability.
[0081] Step S2: Construct at least two candidate standard curves based on the response signals to describe the relationship between the response signals of the target analyte and its concentration.
[0082] In some implementations, candidate standard curves can be constructed based on the theoretical concentration of a standard sample and the corresponding response signal, where the response signal can be the ratio of the peak area of the target analyte to the peak area of the internal standard. Candidate standard curves can be constructed in the following form:
[0083]
[0084] in, Indicates the first The response values corresponding to the candidate standard curves; Indicates the theoretical concentration of the target analyte; Indicates the first The slope of the candidate standard curve; Indicates the first The intercepts of the candidate standard curves.
[0085] In some implementations, multiple candidate standard curves can be constructed based on different standard sample selection strategies. For example, multiple candidate curves can be constructed based on the sorting method or distribution characteristics of standard sample concentrations, such as: constructing candidate standard curves based on selecting standard points from high to low concentration; constructing candidate standard curves based on selecting standard points from low to high concentration; or constructing candidate standard curves based on the breadth of the concentration range covered.
[0086] In some implementations, candidate standard curves can be constructed based on different combinations of standard points. For example, under the condition of satisfying the preset minimum number of points constraint, the combination containing the most standard points can be selected, or the combination of standard points covering a large concentration range can be selected for fitting, thereby forming multiple candidate standard curves.
[0087] In this way, at least two candidate standard curves can be generated. Each candidate standard curve is used to characterize the correspondence between the response signal of the target analyte and its concentration, providing a diverse candidate basis for subsequent screening based on evaluation indicators.
[0088] Step S3: Screen the candidate standard curves based on at least one evaluation index to determine the target standard curve that meets the preset conditions.
[0089] In a preferred embodiment, step S3 specifically includes:
[0090] Step S301: Establish a multidimensional evaluation matrix that includes goodness of fit, quantitative accuracy, number of standard points, and concentration coverage.
[0091] In some implementations, for each candidate standard curve constructed in step S2, its corresponding evaluation feature parameters can be extracted, and a multi-dimensional evaluation matrix can be constructed. Specifically, for the first... Given several candidate standard curves, the following evaluation vector can be constructed:
[0092]
[0093] in, Indicates the first The goodness of fit of the candidate standard curves is used to measure the degree of linear correlation between the response signal and the concentration; Indicates the first The quantitative accuracy of the candidate standard curve is used to characterize the deviation between the concentration calculated based on the curve and the theoretical concentration. Indicates the number of standard points involved in the fitting; This indicates the concentration range covered by the standard point, used to characterize the applicable range of the curve.
[0094] Step S302: According to the preset priority logic, the candidate standard curves are graded and screened based on the multi-dimensional evaluation matrix, including:
[0095] Step S3021: Among the candidate standard curves that meet the first evaluation constraint, screen the candidate standard curves that simultaneously meet the requirements of goodness of fit and quantitative accuracy, and select the candidate standard curves whose number of standard points reaches the maximum value or whose concentration coverage reaches the maximum value.
[0096] Step S3022: If no candidate standard curves that meet the first evaluation constraint are found, the candidate standard curves are screened again based on the second evaluation constraint which is lower than the first evaluation constraint, and the candidate standard curves that meet the quantitative accuracy requirements and whose concentration coverage reaches the maximum value are selected.
[0097] Candidate standard curves can be categorized and screened based on preset evaluation constraints. These constraints may include goodness-of-fit thresholds, quantitative accuracy ranges, and minimum number of standard points.
[0098] In some implementations, a first set of evaluation constraints can be set first, such as a high goodness-of-fit threshold and a strict quantitative accuracy range. Then, a set of candidate standard curves that simultaneously satisfy both the goodness-of-fit and quantitative accuracy constraints is selected from the multidimensional evaluation matrix. After obtaining the set of candidate curves that meet the first evaluation constraints, they can be further optimized based on the number of standard points and the concentration coverage range. Specifically, candidate standard curves with a larger number of standard points can be preferentially selected to improve fitting stability; when the number of standard points is the same or indistinguishable, candidate standard curves with a larger concentration coverage range can be selected to improve the applicability of the curve.
[0099] When no candidate standard curve meets the first evaluation constraint, the evaluation constraint can be relaxed to form a second evaluation constraint. For example, the goodness-of-fit threshold can be lowered or some constraints can be relaxed. Based on this, candidate standard curves are re-screened, and priority is given to candidate standard curves that meet the quantitative accuracy requirements and have a large concentration coverage, so as to expand the applicability of the curve while ensuring quantitative reliability.
[0100] In a specific example, the process of screening candidate standard curves can be illustrated using a particular target analyte.
[0101] Under the first evaluation constraint, a goodness-of-fit threshold of 0.99 can be set, the quantitative accuracy range can be 80% to 120%, and the number of standard points can be no less than 5. Under this constraint, candidate standard curves can be constructed and screened based on different standard point selection strategies, including:
[0102] Standard points were selected from high to low according to the concentration of the calibrated lines for fitting, and the candidate standard curve that meets the above evaluation constraints and contains the most standard points was selected.
[0103] Alternatively, standard points can be selected based on the concentration coverage range for fitting, and a candidate standard curve that meets the above evaluation constraints, covers the widest concentration range, and contains the most standard points can be selected.
[0104] Alternatively, standard points can be selected from low to high concentrations for fitting, and the candidate standard curve that meets the above evaluation constraints and contains the most standard points can be selected.
[0105] In the further screening process, when there are no candidate standard curves that meet the goodness-of-fit requirement greater than 0.99, the goodness-of-fit threshold can be appropriately lowered, for example, to 0.98 or 0.95. Under the relaxed evaluation constraints, candidate standard curves are re-screened, and priority is given to the candidate standard curves that meet the quantitative accuracy requirements, have the widest concentration coverage, and contain the most standard points.
[0106] The above example process allows you to select the target standard curve that meets the criteria from multiple candidate standard curves.
[0107] Step S303: Determine the candidate standard curves obtained from the screening as the target standard curve.
[0108] The candidate standard curves obtained through screening are used as the final target standard curves for subsequent quantitative calculations of the samples to be analyzed. The evaluation parameters and screening path information corresponding to the target standard curve are recorded, such as the level of evaluation constraints used and the set of standard points involved in the fitting, to facilitate subsequent result traceability or quality control.
[0109] In the solution provided in this application embodiment, through steps S301 to S303, candidate standard curves are uniformly modeled based on a multidimensional evaluation matrix and then hierarchically screened according to a preset priority logic. When higher evaluation constraints are met, candidate standard curves with high goodness of fit, reliable quantitative accuracy, and reasonable standard point distribution are preferentially selected. When the conditions are not met, adaptive screening is achieved by relaxing the evaluation constraints, thereby stably obtaining the optimal standard curve under different data quality conditions. By combining multidimensional evaluation with hierarchical screening, the instability caused by relying solely on a single evaluation index for standard curve selection is avoided, improving the robustness of standard curve screening. Simultaneously, it considers both fitting accuracy and applicability, effectively enhancing the accuracy and reliability of mass spectrometry quantitative analysis results.
[0110] Step S4: When the candidate standard curve does not meet the preset conditions, automatically switch the internal standard data, reconstruct the candidate standard curve and perform filtering until the target standard curve is obtained.
[0111] In a preferred embodiment, step S4 specifically includes:
[0112] Step S401: Extract a set of candidate internal standards that have a mapping relationship with the target analyte from the preset internal standard database.
[0113] An internal standard database is pre-built to store the association information between multiple internal standards and target analytes. This association information can be established based on similarity in physicochemical properties, proximity of retention times, consistency of ion response characteristics, or historical experimental calibration relationships. For example, a mapping table can be established by analyzing the degree of matching between the target analyte and each internal standard in terms of chromatographic retention behavior, ionization efficiency, or response stability. Upon receiving the target analyte identifier, a set of internal standards with a mapping relationship can be retrieved from the internal standard database to obtain a candidate internal standard set. In some embodiments, the candidate internal standard set can also be pre-screened, for example, by removing internal standard data with abnormal or missing responses in the current batch, to improve the stability of subsequent processing.
[0114] Step S402: Select internal standard data from the candidate internal standard set as the current reference internal standard according to the preset matching priority.
[0115] For multiple internal standards in the candidate internal standard set, they can be sorted based on a preset matching priority. The matching priority can include at least one of the following indicators: retention time difference with the target analyte, response signal stability, signal-to-noise ratio (SNR), and historical quantitative accuracy performance. Specifically, retention time difference characterizes the temporal proximity of the candidate internal standard and the target analyte during chromatographic separation; response signal stability can be measured by the coefficient of variation in the same batch or historical batches; SNR reflects the detectability and anti-interference ability of the internal standard signal; and historical quantitative accuracy performance characterizes the accuracy and consistency of the internal standard in previous quantitative analyses. Based on these indicators, a comprehensive score can be calculated for each internal standard, and the candidate internal standard set can be sorted according to the score results. Subsequently, internal standards are selected sequentially according to the sorting order as the current reference internal standard for subsequent response signal normalization processing.
[0116] Step S403: Normalize the response signal of the target analyte based on the current reference internal standard to obtain an updated response signal.
[0117] For each sample, the response signal of the target analyte and the response signal of the current reference internal standard can be obtained, and a normalized ratio can be constructed based on the two to eliminate systematic errors between samples and the influence of instrument fluctuations.
[0118] Normalization can be expressed as:
[0119]
[0120] in: Indicates the first The normalized response signal of each sample; Indicates the first The original response signal of the target analyte in each sample; Indicates the first The response signal of the current reference internal standard in each sample.
[0121] The above normalization process can effectively reduce the impact of factors such as sample matrix effect, injection volume fluctuation and instrument drift on the response signal, thereby improving the stability and consistency of subsequent standard curve construction.
[0122] Step S404: Based on the updated response signal, re-execute the steps of constructing at least two candidate standard curves and screening the candidate standard curves.
[0123] The normalized response signal is used as new input data to reconstruct candidate standard curves. Specifically, multiple candidate standard curves can be constructed based on the correspondence between the theoretical concentration of the standard sample and the normalized response signal. For example, different standard point combination strategies, different fitting intervals, or different data screening methods can be used to construct multiple candidate curves. Subsequently, the constructed candidate standard curves can be screened based on preset evaluation indicators (such as goodness of fit, quantitative accuracy, number of standard points, and concentration coverage) to determine the optimal candidate standard curve under the current internal standard conditions.
[0124] Step S405: If a target standard curve that meets the preset conditions is not obtained, switch to the next internal standard data in the candidate internal standard set, and re-execute the normalization processing of the response signal and the construction and screening steps of the candidate standard curve until the target standard curve is obtained.
[0125] When none of the candidate standard curves corresponding to the current reference internal standard meet the preset conditions (e.g., insufficient goodness of fit or substandard quantitative accuracy), the internal standard switching mechanism can be triggered. The next-ranked internal standard data is selected from the candidate internal standard set as the new reference internal standard, and steps S403 and S404 are repeated. During each internal standard switching process, the corresponding standard curve evaluation results can be recorded and compared with historical results to determine whether the preset conditions have been met. When a candidate standard curve that meets the preset conditions exists, it is determined as the target standard curve, and the internal standard switching process is terminated.
[0126] In the solution provided in this application embodiment, through steps S401 to S405, a candidate internal standard set is constructed based on a preset internal standard database. Internal standards are then selected sequentially according to matching priority for response signal normalization. When a candidate standard curve does not meet preset conditions, the internal standard is automatically switched, and the candidate standard curve is reconstructed and screened, forming an internal standard-driven iterative optimization mechanism. This achieves adaptive adjustment of the standard curve construction process. This technical solution avoids the system bias problem caused by relying on a single internal standard. Multiple rounds of normalization and modeling screening of the response signal are performed under different internal standard conditions, improving the fitting quality and quantitative accuracy of the standard curve. Simultaneously, by automatically traversing the candidate internal standard set and executing a closed-loop screening process, the need for manual intervention is reduced, enhancing the system's robustness and adaptability in complex sample environments.
[0127] Step S5: Calculate the quantitative concentration of the sample to be analyzed based on the target standard curve to obtain the quantitative analysis results of the target analyte in the sample to be analyzed.
[0128] A target standard curve is used to describe the functional relationship between the response signal of a target analyte and its concentration. In practical applications, the target standard curve can be represented using either a linear or nonlinear model. In the case of linear fitting, the target standard curve can be expressed as:
[0129]
[0130] in, Indicates the first Quantitative concentration of the target analyte in a sample to be analyzed; Indicates the first The response signal of the target analyte in the sample to be analyzed; Indicates the slope of the target standard curve; This represents the intercept of the target standard curve.
[0131] In practical implementation, the response signal of the sample to be analyzed can be input into the target standard curve model to calculate the corresponding concentration value of the target analyte. If the response signal is a signal that has been normalized by an internal standard, the calculated concentration value is a corrected quantitative result, thereby effectively reducing the impact of systematic errors. In some implementations, the validity of the calculated concentration result can also be verified. For example, it can be determined whether the response signal falls within the effective concentration range of the target standard curve; when the response signal exceeds the coverage range of the standard curve, it can be marked as an extrapolation result or an anomaly handling mechanism can be triggered to avoid outputting unreliable quantitative results.
[0132] In a preferred embodiment, the method further includes: correcting the quantitative analysis results of the target analyte in the sample to be analyzed based on a preset subtraction rule, specifically including:
[0133] Step S5A1: Obtain the response signal of the blank sample belonging to the same detection batch as the sample to be analyzed, and determine the detection threshold according to the preset multiple relationship;
[0134] Step S5A2: Compare the response signal of the target analyte in the sample to be analyzed with the detection threshold;
[0135] Step S5A3: If the response signal is less than or equal to the detection threshold, the target analyte in the sample to be analyzed is determined to be in an undetected state, and the corresponding quantitative analysis result is corrected to the preset undetected label;
[0136] Step S5A4: If the response signal is greater than the detection threshold, the original quantitative value of the target analyte in the sample to be analyzed is determined based on the target standard curve, and the original quantitative value is corrected according to the response signal of the blank sample to obtain the corrected quantitative analysis result.
[0137] After determining the response signal of the target analyte in the sample to be analyzed, the quantitative analysis results can be further corrected based on blank samples. Specifically, blank samples can be identified from the sample set belonging to the same detection batch as the sample to be analyzed, and the response signal of the blank sample under the corresponding detection conditions of the target analyte can be obtained. The response signal of the blank sample is used to characterize the background noise level or matrix interference intensity in the current detection batch.
[0138] In some implementations, the response signals of multiple blank samples can be statistically processed to obtain reference values characterizing the background level. For example, the mean, median, or quantile of the blank sample response signals can be calculated as background reference values. Based on this, the background reference values can be amplified according to a preset multiple relationship to determine the detection threshold used to determine whether the target analyte is detected.
[0139] For each sample to be analyzed, the corresponding target analyte response signal (e.g., peak area or normalized response value) can be obtained and compared with a detection threshold. When the response signal of the sample to be analyzed is less than or equal to the detection threshold, it can be considered that the response signal mainly originates from background noise or system interference and does not form a valid target analyte signal. In this case, the target analyte in the sample to be analyzed can be determined as undetectable, and its quantitative analysis results can be corrected. For example, the corresponding concentration value can be set to zero or marked with a preset undetectable flag. Through the above processing, background noise can be avoided from being misjudged as a valid signal, thereby reducing the probability of false positive results and improving the reliability of quantitative analysis results.
[0140] When the response signal of the sample to be analyzed is higher than the detection threshold, it can be determined that the target analyte in the sample has been effectively detected. At this time, the response signal can be substituted into the target standard curve model to calculate the corresponding raw quantitative concentration value.
[0141] Considering that even in the detection state, the response signal may still contain a certain degree of background signal, the original quantitative concentration value can be corrected based on the response signal of the blank sample. For example, the corresponding background signal can be subtracted from the original response signal, and the concentration value can be recalculated based on the corrected response signal, or the original quantitative concentration value can be directly compensated and corrected to obtain the corrected quantitative analysis result. In some embodiments, the corrected concentration can be expressed as:
[0142]
[0143] in, Indicates the first Corrected concentrations of the target analyte in the sample to be analyzed; This represents the response signal of the target analyte in the sample to be analyzed; This represents the background signal of the blank sample; This represents the inverse function relationship corresponding to the target standard curve.
[0144] In the solution provided in this application embodiment, through steps S5A1 to S5A4, a background reference can be established based on a blank sample within the same detection batch, and the quantitative results of the sample to be analyzed can be adaptively corrected. While suppressing false positive determination of low signal samples, background subtraction is performed on the validly detected samples, thereby improving the accuracy and reliability of the quantitative analysis results.
[0145] In one specific embodiment, such as Figure 2 As shown, the automated quantitative analysis method for mass spectrometry detection data provided in this application may include the following process:
[0146] First, acquire the mass spectrometry detection data of the sample to be analyzed. Mass spectrometry detection data can be acquired using a gas / liquid chromatography-mass spectrometry (GC-MS) system and is presented as a signal data stream that varies over time. Simultaneously with acquiring the raw mass spectrometry detection data, further analysis can be performed using methods such as... Figure 3The analysis interface shown loads the molecular and sample information to be processed. Area ① of this interface displays the selection screen for the candidate molecule list, which lists several labeled compounds. Users can select the target analyte or corresponding internal standard by checking boxes, thus determining the analytical object for subsequent data processing. Area ② displays the structural information of the sample data, including fields such as sample type, sample name, and collection order. Sample types can include pooled standard samples (QC), blank samples (BLK), and samples to be analyzed (SPL). Different types of samples serve different functions in subsequent data processing: quality control samples are used to assess data stability and method reliability; blank samples are used for background signal estimation and blank subtraction; and samples to be analyzed are used for the quantitative calculation of the target analyte.
[0147] After data loading is completed, the mass spectrometry detection data is preprocessed, including adaptive smoothing, baseline subtraction, and dynamic background noise modeling of the raw signal to reduce noise interference and eliminate the influence of baseline drift, thereby obtaining the pre-cleaned signal data.
[0148] After data preprocessing, characteristic peak identification and response signal extraction are performed on the signal. Multi-scale analysis can be conducted within the joint time-scale domain, extracting local extremum information at different scales through wavelet transform, and identifying candidate peaks based on extremum propagation trajectories. Simultaneously, by combining a time-axis drift correction model and internal standard data, multi-dimensional verification is performed on the candidate peaks to determine the characteristic peaks corresponding to the target analyte. For overlapping peaks, peak separation can be performed using deconvolution to obtain the response signal of the target analyte. Figure 4 As shown, a peak overlay plot visualizes the response signal of the target analyte in a mixed standard sample. The horizontal axis represents retention time, and the vertical axis represents signal intensity. Different colored curves correspond to the changes in the response signal of different samples under the corresponding detection channel of the target analyte. In this plot, it can be observed that the mixed standard samples form an overlay peak distribution within the same retention time region, indicating that the target analyte responds at similar time points in the mixed standard samples. By limiting this region to a window, the response signal of the target analyte can be extracted. Specifically, an integration window can be constructed based on a preset retention time range. The integration window is defined by left and right boundaries to constrain the integration interval of the target peak. The vertical lines in the plot are used to mark the boundary positions of the integration window, which can include a start boundary and an end boundary to define the effective integration range of the peak. Within the integration window, the target peak can be integrated by area integration to obtain the response signal of the target analyte in the mixed standard sample.
[0149] Furthermore, such as Figure 5As shown, the peak position can be identified within the integration window, and the peak shape can be analyzed with the peak as the center. When there is a slight shift in the retention time of the peaks of multiple samples, the peak position can be corrected by combining the time axis drift correction model, so that the peaks of different samples are aligned under a unified reference time, thereby improving the consistency of response signal extraction.
[0150] In some implementations, anomalous peaks can also be identified based on peak shape characteristics. For example, when the peaks of certain samples deviate significantly from those of other samples in terms of height, width, or symmetry, they can be identified as anomalous peaks and corrected or removed in subsequent processing, thereby preventing anomalous data from affecting the construction of the standard curve.
[0151] After obtaining the response signal of the target analyte, at least two candidate standard curves are constructed based on the response signal. Specifically, multiple candidate standard curves can be constructed by using different combinations of standard points or different fitting intervals, based on the correspondence between the theoretical concentration of the standard sample and the response signal.
[0152] Subsequently, the candidate standard curves are screened. A multi-dimensional evaluation matrix can be established, including goodness of fit, quantitative accuracy, number of standard points, and concentration coverage. The candidate standard curves are then categorized and screened according to a preset priority logic to determine whether a target standard curve that meets the preset conditions exists.
[0153] When a candidate standard curve does not meet the preset conditions, an internal standard switching mechanism is triggered. Different internal standards can be selected sequentially from the candidate internal standard set as references to normalize the response signal of the target analyte, and the candidate standard curve construction and screening steps are re-executed until a target standard curve that meets the preset conditions is obtained, thus forming an optimization process based on multiple internal standard switching.
[0154] After obtaining the target standard curve, the quantitative concentration of the target analyte in the sample to be analyzed is calculated based on the target standard curve. Furthermore, detection determination and background subtraction can be performed based on the response signal of the blank sample to obtain corrected quantitative analysis results.
[0155] Finally, the quantitative analysis results can be summarized and a standardized analysis report can be generated. The analysis report data structure can include at least: sample identification information, target analyte identification, response signal, quantitative concentration, parameters of the target standard curve used, and corresponding evaluation index information. The report data structure can be converted into a standardized analysis report file according to a preset report template, and batch generation and export of analysis results for multiple samples are supported. The report file can be presented in tabular or structured document format to facilitate result traceability and batch output.
[0156] In summary, the automated quantitative analysis method for mass spectrometry detection data provided in this application automatically determines the response signal of the target analyte in a mixed standard sample based on the mass spectrometry detection data, constructs at least two candidate standard curves, and filters the candidate standard curves using evaluation indicators. When a candidate standard curve does not meet preset conditions, the method automatically switches the internal standard data, reconstructs the candidate standard curve, and performs the filtering process again until the target standard curve is obtained. Based on the target standard curve, the quantitative concentration of the analyte is calculated, thus forming a closed-loop optimization mechanism based on adaptive switching of multiple candidate standard curves and internal standards. This achieves fully automated processing of the quantitative analysis process for mass spectrometry detection data. Through this method, a standard curve meeting preset conditions can be automatically obtained without manual intervention, avoiding quantitative bias caused by relying on a single standard curve or a fixed internal standard. This improves the rationality and adaptability of standard curve selection, while significantly reducing manual operation intensity and improving data processing efficiency.
[0157] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0158] Some embodiments of this application also relate to an automated quantitative analysis system for mass spectrometry detection data, see reference Figure 6 As shown, it includes:
[0159] The data acquisition module is used to acquire mass spectrometry detection data of mixed standard samples;
[0160] The signal processing module is used to determine the response signal of the target analyte based on mass spectrometry detection data;
[0161] The standard curve construction module is used to construct at least two candidate standard curves based on the response signals to describe the relationship between the response signals of the target analyte and its concentration.
[0162] The curve screening module is used to screen candidate standard curves based on at least one evaluation index in order to determine the target standard curve that meets the preset conditions.
[0163] The internal standard switching module is used to automatically switch the internal standard data, reconstruct the candidate standard curve and perform filtering when the candidate standard curve does not meet the preset conditions, until the target standard curve is obtained.
[0164] The quantitative analysis module is used to calculate the quantitative concentration of the sample to be analyzed based on the target standard curve, so as to obtain the quantitative analysis results of the target analyte in the sample to be analyzed.
[0165] The data acquisition module is used to acquire mass spectrometry detection data of mixed standard samples. Mixed standard samples are standard samples containing multiple target analytes at known concentrations, used to construct standard curves. Mass spectrometry detection data can be acquired by liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, or ion mobility mass spectrometry, and is presented as a time-varying signal data sequence. In some embodiments, the data acquisition module can also be used to load sample information and analyte information, including sample type, sample identifier, and selection information for target analytes and internal standards, providing a data foundation for subsequent processing.
[0166] The signal processing module is used to determine the response signal of the target analyte based on mass spectrometry detection data. Specifically, the signal processing module can perform data preprocessing operations on the raw signal, including adaptive smoothing, baseline subtraction, and dynamic background noise modeling, to reduce noise interference and eliminate the influence of baseline drift. Peak detection can be performed on the signal using multi-scale analysis methods, and candidate peaks can be verified by combining a time axis drift correction model and internal standard data to finally determine the characteristic peak corresponding to the target analyte, and the response signal is obtained through integration calculation.
[0167] The standard curve construction module is used to construct at least two candidate standard curves based on the response signals of a mixed standard sample. Multiple candidate standard curves can be constructed using different standard point combination strategies or different fitting methods, based on the correspondence between the theoretical concentrations of each target analyte in the mixed standard sample and the corresponding response signals, thus providing multiple alternative models for subsequent standard curve screening.
[0168] The curve screening module is used to screen candidate standard curves based on at least one evaluation metric to determine the target standard curve that meets preset conditions. Evaluation metrics may include goodness of fit, quantitative accuracy, number of standard points, and concentration coverage. By performing hierarchical screening of candidate standard curves, the optimal standard curve that meets the preset conditions can be obtained.
[0169] The internal standard switching module automatically switches internal standard data, reconstructs candidate standard curves, and performs filtering when candidate standard curves do not meet preset conditions. It can extract a set of candidate internal standards from a preset internal standard database and select internal standards sequentially according to matching priority, normalizing the response signals of the mixed standard samples. Based on the normalized response signals, the standard curve construction and filtering steps are re-executed. If the candidate standard curve corresponding to the current internal standard still does not meet the preset conditions, the module continues to switch to the next internal standard and repeats the above process until the target standard curve is obtained, thus forming a closed-loop optimization mechanism based on multi-internal standard switching.
[0170] The quantitative analysis module is used to calculate the quantitative concentration of the analyte in the sample based on the target standard curve, thereby obtaining the quantitative analysis results of the target analyte in the sample. It can acquire the response signal of the sample and input the response signal into the target standard curve model to calculate the corresponding target analyte concentration value.
[0171] Through the synergistic effect of the above modules, the entire process of constructing and screening standard curves based on mixed standard samples, as well as the quantitative calculation of samples to be analyzed based on the target standard curve, can be automated. Furthermore, the adaptability and stability of standard curve construction can be improved through the internal standard switching mechanism, thereby enhancing the accuracy and reliability of quantitative analysis results.
[0172] The content of the above-described automated quantitative analysis method embodiments for mass spectrometry detection data is applicable to this device embodiment. The specific functions implemented in this device embodiment are the same as those in the above-described automated quantitative analysis method embodiments for mass spectrometry detection data, and the beneficial effects achieved are also the same as those achieved in the above-described automated quantitative analysis method embodiments for mass spectrometry detection data. To avoid repetition, further details are omitted here.
[0173] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as cellular phones, smartphones, wearable devices, and other similar computing devices.
[0174] The electronic device includes: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform an automated quantitative analysis method for mass spectrometry detection data as provided in any one or more of the above embodiments. Figure 7An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0175] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103, and output device 1104 may be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0176] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0177] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).
[0178] In this embodiment, a computer-readable medium stores a computer program / instruction, which, when executed by a processor, implements an automated quantitative analysis method for mass spectrometry detection data provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.
[0179] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.
[0180] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0181] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0182] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0183] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, R, Python, PyQt, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0184] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0185] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0186] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0187] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. Terms such as "first," "second," etc., are used only to distinguish descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.
[0188] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. An automated quantitative analysis method for mass spectrometry detection data, characterized in that, include: Acquire mass spectrometry detection data of mixed standard samples, and determine the response signal of the target analyte in the mixed standard samples based on the mass spectrometry detection data; Based on the response signal, at least two candidate standard curves are constructed to describe the relationship between the response signal and the concentration of the target analyte; The candidate standard curves are screened based on at least one evaluation index to determine the target standard curve that meets the preset conditions. When the candidate standard curve does not meet the preset conditions, the internal standard data is automatically switched, the candidate standard curve is reconstructed and the filtering is performed until the target standard curve is obtained. Based on the target standard curve, the quantitative concentration of the sample to be analyzed is calculated to obtain the quantitative analysis results of the target analyte in the sample to be analyzed.
2. The automated quantitative analysis method for mass spectrometry detection data according to claim 1, characterized in that, The step of determining the response signal of the target analyte in the mixed standard sample based on the mass spectrometry detection data includes: Based on the characteristics of data acquisition frequency and local scanning density, the raw mass spectrometry signal of the mixed standard sample is adaptively smoothed, and the signal is separated by dynamic background noise modeling in order to extract the target signal components. Multi-scale feature analysis is performed on the target signal components in the joint time-scale domain, and candidate peaks of the target analyte are located by identifying extreme value propagation trajectories at different scales. Based on a preset time axis drift correction model, and combined with the response signal of the internal standard data, the candidate peaks are verified in multiple dimensions to determine the characteristic peaks corresponding to the target analyte. The characteristic peaks are modeled and the overlapping peaks are decomposed based on iterative optimization separation to obtain the response signal of the target analyte in the mixed standard sample.
3. The automated quantitative analysis method for mass spectrometry detection data according to claim 2, characterized in that, Also includes: A compensation mechanism is implemented for samples in the mixed standard samples where no characteristic peak corresponding to the target analyte is identified. The compensation mechanism includes: A reference sample set is constructed based on the standard samples that belong to the same detection batch as the mixed standard samples and have identified the characteristic peaks of the target analyte, and the retention time distribution sequence corresponding to each characteristic peak in the reference sample set is obtained; Based on the retention time distribution sequence, calculate the corresponding median retention time; An integration window is constructed based on the median retention time and a preset integration bandwidth; The integration window is mapped to the corresponding position on the time axis of the sample where no characteristic peak was identified, and integration processing is performed within the integration window to obtain the compensation response signal of the target analyte in the sample where no characteristic peak was identified.
4. The automated quantitative analysis method for mass spectrometry detection data according to claim 1, characterized in that, The step of screening the candidate standard curves based on at least one evaluation index to determine the target standard curve that meets preset conditions includes: Establish a multidimensional evaluation matrix that includes goodness of fit, quantitative accuracy, number of standard points, and concentration coverage. According to a preset priority logic, the candidate standard curves are subjected to hierarchical screening based on the multidimensional evaluation matrix, including: Among the candidate standard curves that meet the first evaluation constraint, the candidate standard curves that simultaneously meet the requirements of goodness of fit and quantitative accuracy are selected, and the candidate standard curves whose number of standard points reaches the maximum value or whose concentration coverage reaches the maximum value are selected from the selected candidate standard curves. If no candidate standard curves that meet the first evaluation constraint are found, the candidate standard curves are screened again based on the second evaluation constraint that is lower than the first evaluation constraint, and the candidate standard curves that meet the quantitative accuracy requirements and whose concentration coverage reaches the maximum value are selected. The candidate standard curves obtained from the screening are determined as the target standard curve.
5. The automated quantitative analysis method for mass spectrometry detection data according to claim 1, characterized in that, The step of automatically switching internal standard data, reconstructing candidate standard curves, and performing filtering when the candidate standard curve does not meet the preset conditions, until the target standard curve is obtained, includes: Extract a set of candidate internal standards that have a mapping relationship with the target analytes of the mixed standard samples from a pre-set internal standard database; According to the preset matching priority, internal standard data are selected sequentially from the candidate internal standard set as the current reference internal standard; The response signal of the target analyte in the mixed standard sample is normalized based on the current reference internal standard to obtain an updated response signal. Based on the updated response signal, the steps of constructing at least two candidate standard curves and screening candidate standard curves are re-executed; If a target standard curve that meets the preset conditions is not obtained, the system switches to the next internal standard data in the candidate internal standard set and re-executes the normalization processing of the response signal and the construction and screening steps of the candidate standard curve until the target standard curve is obtained.
6. The automated quantitative analysis method for mass spectrometry detection data according to claim 1, characterized in that, Also includes: The quantitative analysis results of the target analyte in the sample to be analyzed are corrected based on a preset subtraction rule, including: Obtain the response signal of a blank sample belonging to the same detection batch as the sample to be analyzed, and determine the detection threshold according to a preset multiple relationship; The response signal of the target analyte in the sample to be analyzed is compared with the detection threshold. If the response signal is less than or equal to the detection threshold, the target analyte in the sample to be analyzed is determined to be in an undetected state, and the corresponding quantitative analysis result is corrected to the preset undetected label. If the response signal is greater than the detection threshold, the original quantitative value of the target analyte in the sample to be analyzed is determined based on the target standard curve, and the original quantitative value is corrected according to the response signal of the blank sample to obtain the corrected quantitative analysis result.
7. An automated quantitative analysis system for mass spectrometry detection data, characterized in that, include: The data acquisition module is used to acquire mass spectrometry detection data of mixed standard samples; The signal processing module is used to determine the response signal of the target analyte based on the mass spectrometry detection data; The standard curve construction module is used to construct at least two candidate standard curves based on the response signal to describe the relationship between the response signal and the concentration of the target analyte. The curve screening module is used to screen the candidate standard curves based on at least one evaluation index to determine the target standard curve that meets the preset conditions. The internal standard switching module is used to automatically switch the internal standard data, reconstruct the candidate standard curve and perform filtering when the candidate standard curve does not meet the preset conditions, until the target standard curve is obtained. The quantitative analysis module is used to calculate the quantitative concentration of the sample to be analyzed based on the target standard curve, so as to obtain the quantitative analysis results of the target analyte in the sample to be analyzed.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the automated quantitative analysis method for mass spectrometry detection data as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that, When the computer program and / or instructions are executed by the processor, they implement the automated quantitative analysis method for mass spectrometry detection data as described in any one of claims 1-6.
10. A computer program product comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by the processor, they implement the automated quantitative analysis method for mass spectrometry detection data as described in any one of claims 1-6.