Liquid chromatography mass spectrometry data peak shape identification processing method and related equipment
By dynamically adjusting peak shape recognition parameters and constructing an optimized database, the problem of automatic identification and optimization of abnormal peak shapes in liquid chromatography-mass spectrometry (LC-MS) technology was solved, achieving an efficient and reliable analytical process and reducing manual intervention and misjudgment rate.
Patent Information
- Application Number
- CN202511537062.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing liquid chromatography-mass spectrometry techniques suffer from undesirable peak shapes such as peak tailing, peak broadening, and peak overlap when processing complex samples, resulting in poor analytical accuracy. Furthermore, existing software has low recognition accuracy, relies heavily on manual verification, and lacks systematic optimization methods.
By acquiring liquid chromatography-mass spectrometry (LC-MS) data, calculating signal complexity indices, dynamically adjusting peak shape recognition parameters, and combining a pre-set optimization database, the system automatically identifies abnormal peaks and recommends optimization schemes for liquid chromatography methods, thus constructing a closed-loop system with self-learning and evolution capabilities.
It achieves automatic and accurate identification of abnormal peaks, reduces false positive and false negative rates, significantly improves analysis efficiency and reliability, reduces human intervention, and provides systematic optimization suggestions.
Smart Images

Figure CN120992831B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular, to a liquid chromatography-mass spectrometry data peak shape identification processing method and related equipment. BACKGROUND
[0002] Liquid chromatography-mass spectrometry (LC-MS) technology is a key technology in modern analytical laboratories, and LC-MS data directly depends on the shape of the chromatographic peak. In complex sample analysis, due to poor chromatographic conditions or matrix interference, often produce abnormal peak shapes such as peak tailing, peak broadening and peak overlap, which seriously affect the accuracy of the analysis results.
[0003] Currently, the evaluation of the chromatographic peak shape in LC-MS data mainly depends on two ways. One is manual visual inspection by analysts, which is intuitive, but when dealing with high-throughput samples, not only is the efficiency extremely low, but the evaluation standard is also easily affected by subjective factors, resulting in poor repeatability of the results. The other is assisted identification by automatic data processing software, and existing software can usually perform preliminary screening of the peak shape according to preset fixed parameters (such as symmetry, half-peak width, etc.). However, when dealing with complex or variable signal-to-noise ratio samples, these software often result in a decrease in identification accuracy due to the fixed parameters, resulting in a large number of false positive or false negative results, which still requires a large amount of manual review.
[0004] More importantly, whether it is manual inspection or software-assisted identification, the workflow of the existing technical means stops at the level of "finding problems". Once it is determined that a certain chromatographic peak is an abnormal peak, how to systematically adjust and optimize the upstream liquid phase method to solve the problem depends on the professional knowledge and trial-and-error experience of the analysts. However, when faced with a large amount of complex data, analysts will face a huge workload in the peak shape evaluation link, and the reliability and accuracy of the identification cannot be guaranteed. SUMMARY
[0005] In view of the above technical problems and defects, the purpose of the present application is to provide a liquid chromatography-mass spectrometry data peak shape identification processing method and related equipment, which can automatically and accurately identify various abnormal peak shapes and guarantee the reliability and accuracy of the identification.
[0006] To achieve the above object, in a first aspect, the application provides a liquid chromatography-mass spectrometry data peak shape identification processing method, comprising: acquiring liquid chromatography-mass spectrometry data, the liquid chromatography-mass spectrometry data comprising extracted ion current chromatogram time series and liquid phase method condition information; based on the extracted ion current chromatogram time series, extracting the morphological feature data of the chromatographic peak, the morphological feature data comprising peak points, valley points, peak width, peak area and peak symmetry; calculating the signal complexity index of the extracted ion current chromatogram time series, the signal complexity index comprising signal-to-noise ratio and peak density; adjusting the peak shape identification parameters according to the signal complexity index, the peak shape identification parameters comprising noise threshold and peak width range; using the adjusted peak shape identification parameters, distinguishing the chromatographic peak into normal peak or abnormal peak; searching the liquid phase method optimization scheme matched with the morphological feature data of the abnormal peak in the preset optimization database, the liquid phase method optimization scheme comprising gradient elution time adjustment, mobile phase proportion optimization or chromatographic column replacement suggestion; sending the liquid phase method optimization scheme to the user.
[0007] The application adopts the above method, introduces the signal complexity index, enables the system to first "understand" whether the entire chromatogram is "clean" or "complex", and then dynamically adjusts the identification parameters accordingly, which is like giving the system a pair of automatic zoom glasses, ensuring accurate observation under various conditions. Moreover, the application not only "finds problems", but also builds an expert knowledge base, which can automatically match and recommend verified and operable liquid phase method optimization schemes according to the specific morphological features of the abnormal peak. In this way, a link that originally requires a lot of manual intervention and is full of uncertainty is transformed into an automatic, accurate and solution-providing intelligent workflow, fundamentally improving the efficiency and reliability of the analysis work.
[0008] Optionally, in some embodiments, after sending the liquid phase method optimization scheme to the designated user, it further comprises: receiving the final optimization scheme confirmed by the user as effective in response to the abnormal peak; associating the morphological feature data of the abnormal peak with the final optimization scheme to obtain a feature-optimization mapping data pair; updating the feature-optimization mapping data pair to the optimization database to optimize the accuracy of subsequent search matching.
[0009] The technical scheme of the above embodiment introduces the user feedback and database updating mechanism to build a closed-loop system with self-learning and evolution capabilities. This design can continuously and structurally integrate the tacit knowledge of experts into the optimization database, so that the knowledge base of the scheme matching model grows with application, thereby continuously improving the accuracy, coverage and reliability of subsequent search matching, realizing the intelligent iteration of the system and the long-term optimization of the performance.
[0010] Optionally, in some embodiments, the signal complexity index of the extracted ion chromatogram time sequence is calculated, including: according to the preset target substance information, determining the chromatographic peak detected in the retention time window in the extracted ion chromatogram time sequence as a sample peak; determining a time period corresponding to the peak width of the sample peak as a baseline time period in the adjacent area of the baseline of the sample peak; calculating the signal integral area in the baseline time period and determining it as a baseline noise area; calculating the ratio of the peak area of the sample peak to the baseline noise area to obtain a signal-to-noise ratio value; determining a peak density value based on the number of intersections of the chromatographic peak at a set peak height and the baseline; and integrating the signal-to-noise ratio value and the peak density value to obtain the signal complexity index.
[0011] The technical scheme of the above embodiment provides an objective, accurate and statistically significant signal complexity index calculation method. By calculating the noise area in the baseline area corresponding to the peak width of the sample peak, the fairness and representativeness of the signal-to-noise ratio calculation are ensured; in combination with the quantification of the peak density, the crowdedness of the chromatogram, a stable and reliable quantification input is provided for the adaptive adjustment of the subsequent peak shape recognition parameters, thereby ensuring the scientificity and accuracy of the entire dynamic recognition strategy.
[0012] Optionally, in some embodiments, the morphological feature data of the chromatographic peak is extracted based on the extracted ion chromatogram time sequence, including: generating an initial peak shape feature set based on the extracted ion chromatogram time sequence; inputting the initial peak shape feature set into a pre-trained feature optimization model to obtain an optimized feature weight set; and filtering low-weight morphological features according to the feature weight set to obtain the morphological feature data.
[0013] The technical scheme of the above embodiment upgrades the extraction process of the peak shape feature from the traditional manual selection or full use to an intelligent importance evaluation and screening mechanism by introducing a pre-trained feature optimization model. This method can automatically and data-drivenly quantify the actual contribution of each feature in the initial feature set to distinguishing the advantages and disadvantages of the peak shape. This not only solves the problem of high dimensionality of the initial feature set and the inclusion of a large amount of redundant and noise information, but more importantly, it can discover and preferentially retain the core features with the highest discriminability, laying a solid foundation for subsequent construction of an efficient and robust classification model, thereby improving the accuracy and computational efficiency of the entire recognition system.
[0014] Optionally, in some embodiments, the morphological feature data is obtained by filtering low-weight morphological features according to the feature weight set, including: determining a weight threshold based on the confirmed qualified peak shape feature weight distribution in the historical data; comparing each feature weight in the feature weight set with the weight threshold and filtering the morphological features with a weight lower than the weight threshold to obtain filtered high-weight features; and generating the morphological feature data according to the filtered high-weight features.
[0015] The technical scheme of the above embodiment provides an objective, robust and statistically significant decision benchmark for feature filtering. By determining the weight threshold based on the weight distribution of historical qualified peak shapes, the method ensures that the feature screening criteria are not set by humans or fixed, but derived from a deep understanding of the importance of the inherent characteristics of "ideal peak shapes". This makes the filtering process more scientific and effectively avoids the risk of deleting key features or retaining irrelevant features due to improper threshold setting, thereby maximizing the generation of morphological feature data that is both concise and highly representative and reliable.
[0016] Optionally, in some embodiments, the chromatographic peaks are distinguished into normal peaks or abnormal peaks using the adjusted peak shape identification parameters, including: generating a multi-dimensional peak shape feature vector based on the morphological feature data and the adjusted peak shape identification parameters; performing dimension reduction processing on the multi-dimensional peak shape feature vector to obtain a two-dimensional spatial coordinate set; calculating the Euclidean distance between each coordinate point in the two-dimensional spatial coordinate set to obtain a spatial distance matrix; classifying the chromatographic peaks into normal clusters or abnormal clusters based on the spatial distance matrix through a clustering algorithm; and marking the chromatographic peaks in the normal clusters as normal peaks and marking the chromatographic peaks in the abnormal clusters as abnormal peaks.
[0017] The technical scheme of the above embodiment proposes a robust peak shape classification method based on unsupervised learning. By reducing the multi-dimensional feature vector and applying the density clustering algorithm, the method can automatically classify the data into normal clusters and abnormal clusters according to the inherent similarity of peak shape features. This method does not require pre-set complex classification rules, is particularly good at discovering unknown or atypical abnormal peak shapes, has high adaptability and robustness, and ensures the objectivity and comprehensiveness of peak shape classification.
[0018] Optionally, in some embodiments, the liquid phase method optimization scheme matching the morphological feature data of the abnormal peak is retrieved in the pre-set optimization database, including: encoding the morphological feature data of the abnormal peak into a feature vector; inputting the feature vector into a pre-trained scheme matching model to calculate the similarity between the feature vector and the historical abnormal peak feature vector through the scheme matching model, obtaining a similarity value and a matching optimization scheme; and when the similarity value exceeds a pre-set confidence threshold, outputting the matching liquid phase optimization scheme as the retrieval result.
[0019] The technical scheme of the above embodiment constructs an intelligent and high-precision scheme matching and retrieval system based on a machine learning model. This method realizes quantitative matching of optimization schemes by calculating the similarity between the current abnormal peak and the historical case feature vector, which is much better than traditional rule or keyword retrieval. At the same time, the introduction of the pre-set confidence threshold adds a reliability verification stage for the output of the retrieval result, ensuring that the system only recommends highly reliable optimization schemes, effectively avoiding errors, and ensuring the scientificity and practical value of the recommendations.
[0020] In a second aspect, an electronic device is provided, including one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is configured to store computer program codes including computer instructions, and the one or more processors are configured to invoke the computer instructions to cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0021] In a third aspect, a computer-readable storage medium is provided, including instructions, and when the instructions are executed on the electronic device, the electronic device is caused to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0022] In a fourth aspect, a computer program product is provided, including instructions, and when the computer program product is executed on the electronic device, the electronic device is caused to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0023] It can be understood that the electronic device provided in the second aspect, the storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the method provided in the first aspect of the present application. Therefore, the beneficial effects that can be achieved are referred to the beneficial effects in the corresponding method, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 is a flowchart of a liquid chromatography mass spectrometry data peak shape identification processing method according to an embodiment of the present application;
[0025] Figure 2 is a multi-color spectrum peak distribution example diagram according to an embodiment of the present application; wherein the horizontal axis represents the data point scanning sequence number without unit; and the vertical axis represents the signal intensity level without unit;
[0026] Figure 3 is a hardware architecture schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] To clearly illustrate the technical effects of the present application, a high-throughput quantitative analysis of drug metabolites in plasma samples is taken as a typical application scenario for illustration. In this scenario, the analysis task requires LC-MS analysis on hundreds of pretreated plasma samples to achieve accurate quantification of low-concentration basic drug metabolites. Due to the high complexity of the plasma matrix and the serious interference of endogenous substances, the peak shape of the target metabolite in the chromatogram generally presents a significant tailing phenomenon, accompanied by low signal-to-noise ratio and a large number of matrix interference peaks.
[0028] In this scenario, analysts will face a heavy workload of data review if using the prior art. If relying on manual visual inspection, hundreds of chromatograms need to be reviewed one by one, and the analyst needs to judge whether the tailing degree of each target peak meets the quality control standard according to personal experience. This method not only takes hours or even days, but also leads to poor reproducibility due to the lack of uniformity in the judgment standard. The same peak shape may lead to contradictory conclusions at different time points or by different operators. If using automated data processing software, a fixed symmetry threshold (e.g. 0.9) is usually preset as the basis for judgment.
[0029] However, due to the inherent low signal-to-noise ratio of plasma samples, the software with fixed parameters is prone to misjudging a large number of sharp peaks generated by matrix noise as unqualified peaks (producing false positives); at the same time, for real target peaks whose tailing degree is just in the critical range of the preset threshold, they may be missed (producing false negatives) because they do not meet the trigger standard. The more critical problem is that even after a large amount of manpower is invested to complete the screening of all tailing abnormal data points, the workflow of the prior art is terminated. Analysts still need to rely on their own professional knowledge to infer the cause of the abnormality, such as judging whether the tailing phenomenon is caused by improper pH value of the mobile phase or by column contamination. For inexperienced analysts, the subsequent process will inevitably fall into a trial-and-error process without theoretical guidance to seek effective methods to optimize strategies.
[0030] In comparison, when processing the same plasma sample data using the peak shape recognition and processing method for liquid chromatography-mass spectrometry provided by the embodiments of the present application, the entire analysis process is optimized. The system first acquires all the chromatogram data and the corresponding liquid chromatography method condition information (including key parameters such as column specifications, mobile phase component ratio, etc.). Then, the system finds that the overall data of this batch presents the characteristics of low signal-to-noise ratio and high peak density through calculation, and accordingly determines that it is a complex matrix sample. Based on this judgment, the system automatically performs a dynamic adjustment strategy for peak shape recognition parameters: appropriately increasing the noise judgment threshold to effectively filter out matrix interference signals, and using a more stringent peak boundary recognition algorithm to accurately define the integration area of the target peak, so as to accurately recognize the peak of the target drug metabolite as an abnormal peak with tailing characteristics, and effectively eliminate a large number of potential false positive results.
[0031] After the abnormal peak recognition is completed, the system further extracts key morphological feature parameters of the abnormal peak, such as “the peak symmetry value is significantly lower than 1.0, and an obvious tailing morphology is presented”, and matches it with known liquid phase method condition information, specifically including “the mobile phase system is a gradient elution combination of formic acid aqueous solution and acetonitrile, and the target analyte belongs to an alkaline compound”. The system performs retrieval matching in the preset optimization database based on the above information, obtains the most possible cause explanation and the corresponding solution.
[0032] Finally, the system presents the diagnostic results and optimization suggestions to the user in the form of a structured report: “a serious tailing phenomenon is detected in the target peak. The inferred cause is that the secondary interaction occurs between the alkaline compound and the residual silanol group on the surface of the chromatographic column packing. The recommended liquid phase method optimization scheme includes: (1) increasing the formic acid concentration in the mobile phase to 0.2% to enhance the ion suppression effect on the basic site; (2) replacing the C18 chromatographic column with end group tailing treatment optimized for alkaline compounds.” This method compresses the original manual review and fault diagnosis process of several hours or even several days to several minutes, significantly improving the analysis efficiency. The dynamic adaptive recognition algorithm based on signal complexity is superior to traditional fixed parameter software in terms of accuracy and robustness, greatly reducing the misjudgment rate and omission rate, ensuring the reliability and consistency of data quality control. At the same time, this method realizes the whole process closed loop from abnormal discovery to root cause diagnosis to optimization scheme recommendation, not only pointing out the problem, but also providing specific improvement measures based on expert knowledge base, effectively reducing the dependence on the professional experience of the operator, speeding up the development and troubleshooting process.
[0033] The following will be described in combination with Figure 1 to illustrate a liquid chromatography mass spectrometry data peak shape recognition processing method provided by the embodiment, which specifically includes the following steps:
[0034] Step 101, acquiring liquid chromatography mass spectrometry data, the liquid chromatography mass spectrometry data including an extracted ion flow chromatogram time sequence and liquid phase method condition information.
[0035] Among them, the extracted ion flow chromatogram (Extracted Ion Chromatogram, EIC) is a graph showing only the signal intensity of a specific mass-to-charge ratio ion changing with the chromatographic retention time, which is used to track target compounds from complex mass spectrometry data.
[0036] The extracted ion flow chromatogram time sequence is a series of data point sets arranged in time sequence constituting the extracted ion flow chromatogram, each data point consisting of a specific time value and an ion intensity value corresponding to the time point.
[0037] Liquid method condition information refers to the structured information corresponding to the chromatographic data file, recording all the instrument parameters and experimental settings used in this analysis, such as column specifications, mobile phase components, gradient program, and column temperature, etc.
[0038] Specifically, the data processing system (hereinafter referred to as the system) accesses and reads the raw data files generated by the liquid chromatography-mass spectrometry instrument through a pre-set interface protocol or file parsing module. These files are usually in instrument-specific binary format (such as.raw,.d,.wiff) or open standard format (such as mzML). The "acquisition" behavior covers the programmed parsing of these complex data structures.
[0039] Secondly, for the key operation of "extracting ion flow chromatogram time series", according to the specific mass-to-charge ratio (m / z) of the target analyte, all ion intensity signals within the pre-set mass tolerance range (for example, ±5 ppm) are filtered out from the full mass spectrum data (three-dimensional data: retention time, mass-to-charge ratio, ion intensity) recorded in each scan period, and these intensity values are mapped with their corresponding time points, thereby generating a two-dimensional time series data point set with retention time as the horizontal coordinate and ion intensity as the vertical coordinate. This extracted ion flow chromatogram time series constitutes the response signal stream of a single target compound during the entire chromatographic analysis process and is the basis for all subsequent peak shape feature extraction.
[0040] At the same time, the system also needs to obtain the liquid method condition information strictly corresponding to the data file. This part of information is usually stored in the metadata area of the data file or in an independent experimental method file. Its acquisition process involves parsing these files and structurally extracting all parameters related to chromatographic separation behavior, including but not limited to: the physical and chemical properties of the chromatographic column (such as brand, model, stationary phase chemical properties, particle size, length, and inner diameter), the specific components of the mobile phase (such as the solvents and additives of A and B phases and their concentrations), the gradient elution program (including the mobile phase ratio and flow rate changes at each time point), and the column temperature, etc.
[0041] The synchronous acquisition of the extracted ion flow chromatogram time series and the liquid method condition information is crucial. The time series data provides the "phenomenon", while the method condition information provides the "background" for explaining and optimizing the phenomenon. Together, they constitute the complete input for subsequent intelligent diagnosis and optimization recommendations.
[0042] Step 102, based on the extracted ion flow chromatogram time series, extract the morphological feature data of the chromatographic peak.
[0043] Among them, the morphological feature data includes peak points, valley points, peak width, peak area, and peak symmetry.
[0044] This step aims to convert the extracted ion chromatogram time series signal into a set of numerical feature vectors that can quantitatively describe the geometric shape of the chromatographic peaks.
[0045] The specific process begins with applying a series of signal preprocessing algorithms to the input time series data, such as using Gaussian smoothing or Savitzky-Golay filtering algorithm to remove high-frequency noise, and using baseline fitting algorithms such as asymmetric least squares to perform baseline correction, to obtain pure peak signals. On this basis, the system identifies the start, end and vertex position of the chromatographic peaks through peak detection algorithms (such as zero judgment method based on first and second derivatives).
[0046] Specifically, "peak point" is defined as the data point where the ion intensity reaches the maximum value in the identified peak area, and its coordinates (retention time and peak height) are accurately recorded. "Valley point" refers to the local minimum point formed by the intersection with the baseline at the beginning and end of the peak signal or between overlapping peaks, which defines the boundary of peak integration. "Peak width" is a key indicator to measure the efficiency of the chromatographic column. In this embodiment, multiple dimensions of peak width are usually calculated, among which the most core is the half peak width (Full Width at Half Maximum, FWHM), which directly reflects the efficiency of chromatographic separation.
[0047] In addition, the peak width at 5% of the peak height defined by the United States Pharmacopoeia (USP) can also be calculated for subsequent symmetry calculation. "Peak area" as the basis of quantitative analysis is obtained by numerical integration (such as trapezoidal rule) of the baseline-corrected peak signal between the starting and ending valley points, which represents the total response amount of the substance. "Peak symmetry" is a key parameter to evaluate abnormal shapes such as peak tailing or front stretching, usually quantified by asymmetry factor (Asymmetry Factor) or tailing factor (Tailing Factor, Tf). For example, the USP tailing factor is defined as the ratio of the width of the latter half to the width of the former half of the peak (measured at 5% of the peak height), and the Tf value of an ideal Gaussian peak is 1. Greater than 1 indicates tailing, less than 1 indicates front stretching.
[0048] Through systematic calculation of the above features, each chromatographic peak is converted into a multi-dimensional digital fingerprint, providing an objective and quantitative basis for subsequent classification, identification and optimization scheme retrieval.
[0049] Step 103, calculate the signal complexity index of the extracted ion chromatogram time series, including signal-to-noise ratio and peak density.
[0050] The step is to quantitatively evaluate the overall quality and crowding degree of the extracted ion chromatogram (EIC) time series, so as to provide an objective basis for subsequent adaptive peak shape recognition.
[0051] Firstly, the EIC data as a one-dimensional time series, i.e. a series of (time, intensity) data point pairs, is obtained. Then, the calculation of the signal complexity index is performed.
[0052] Among them, for the signal-to-noise ratio (S / N), it is defined as the ratio of signal response intensity to noise response intensity, which is a key indicator to measure signal clarity. In this embodiment, the calculation of signal-to-noise ratio is not for a single chromatographic peak, but for a global evaluation of the entire chromatogram.
[0053] In specific implementation, the system first automatically identifies one or more baseline regions where there are no obvious chromatographic peaks. By calculating the standard deviation (SD) or root mean square (RMS) of signal intensity in these regions, a value representing the global noise level (N) is obtained. Subsequently, the system determines the maximum signal intensity value in the entire chromatogram, or takes the average value of several data points with the highest signal intensity as the representative signal level (S). Finally, the global signal-to-noise ratio is S / N.
[0054] For peak density, it is defined as the number of chromatographic peaks detected per unit time, which is used to represent the crowding degree of chromatographic separation. To calculate this index, the system will perform a preliminary, low threshold peak search algorithm on the entire EIC time series, such as based on first derivative zero crossing or local maximum search, aiming to quickly identify all potential peak events without accurate integration or bounding. The total number of identified potential peaks is divided by the total time of the entire chromatographic analysis or the effective analysis time window, and the peak density value is obtained.
[0055] Finally, the calculated signal-to-noise ratio and peak density, two independent values, together constitute a two-dimensional signal complexity index, which can be represented in the form of a vector, for example, the signal complexity index vector. This vector can comprehensively and quantitatively describe the complexity of the current EIC data from two dimensions of signal clarity and component separation.
[0056] Step 104, adjusting the peak shape recognition parameters according to the signal complexity index, the peak shape recognition parameters including noise threshold and peak width range.
[0057] This step is the core link to realize the peak shape recognition adaptability, its purpose is to convert the above signal complexity index into specific and optimized algorithm parameters for micro-peak recognition, so as to get rid of the limitations of traditional fixed parameter method in processing diversified samples. The adjustment process is based on pre-set logical rules or mathematical models.
[0058] Firstly, for the adjustment of the noise threshold, the noise threshold defines the intensity limit to distinguish the true signal from the baseline noise. The system will dynamically set this threshold according to the signal-to-noise ratio (S / N) index. When a low signal-to-noise ratio is detected, it indicates that the baseline noise is large, and the system will automatically increase the noise threshold, for example, set it to a higher multiple of the baseline noise standard deviation (such as from the default 3 times SD to 5 times or higher), to enhance the robustness, effectively avoid misjudging noise spikes as effective chromatographic peaks, and thus reduce the false positive rate. Conversely, when a high signal-to-noise ratio is detected, it indicates that the baseline is flat and clean, and the system can appropriately reduce the noise threshold to improve the detection sensitivity, ensure that real low-abundance component peaks are not missed, and avoid false negatives.
[0059] Secondly, for the adjustment of the peak width range, the peak width range defines the reasonable time width interval ([minimum width, maximum width]) that should be possessed by the identified effective chromatographic peaks. The system mainly adjusts according to the peak density index. When a high peak density is detected, it means that the chromatogram is very crowded, and the chromatographic peaks may become narrow and sharp due to rapid elution or adjacent peak interference. At this time, the system will automatically narrow the acceptable peak width range, especially reduce the maximum width threshold, to more accurately separate closely overlapping peaks and prevent multiple incompletely separated peaks from being incorrectly identified and combined into an abnormally wide peak. Conversely, when the peak density is low, it may correspond to good separation but wide peak chromatographic conditions, and the system will widen the peak width range, especially increase the maximum width threshold, to ensure that these normally shaped but wide chromatographic peaks can be correctly identified, avoiding being incorrectly marked as abnormal due to exceeding the narrow default range.
[0060] By establishing a functional mapping or rule set between signal complexity indicators and peak shape recognition parameters, this method can tailor the optimal recognition criteria for each EIC data, achieving intelligent and high-precision peak shape judgment.
[0061] Step 105, using the adjusted peak shape recognition parameters, the chromatographic peaks are divided into normal peaks or abnormal peaks.
[0062] This step is the decision-making stage of peak shape quality evaluation, and the core is to apply the adaptive recognition parameters dynamically generated in the previous step to perform multi-dimensional morphological evaluation and classification on each chromatographic peak that has undergone preliminary detection and integration.
[0063] Firstly, the system processes the extracted ion chromatogram with the adjusted noise threshold, only identifying the signal area above this threshold as potential peak events, and then determining the precise start and end points of the peaks according to the principle of signal falling back to the baseline or valley.
[0064] Subsequently, the system immediately applies the adjusted peak width range to the delimited peaks for preliminary screening. The actual width of each peak (peak tail time minus peak start time) is calculated, and if its value is not within the reasonable interval set dynamically, it is directly classified as an abnormal peak, such as a "too narrow peak" (usually a noise spike) or a "too wide peak" (possibly indicating a decrease in column efficiency or an inappropriate gradient).
[0065] For chromatographic peaks that pass the width screening, the system will enter a more detailed morphological feature calculation and evaluation phase. This phase calculates a series of key morphological parameters, with the most core being the symmetry indicators such as Tailing Factor (Tf) or Asymmetry Factor (As). The calculation method of these factors follows the pharmacopoeia specifications (such as USP or EP), which quantifies the degree of peak tailing or fronting by measuring the ratio of the width of the latter half to the former half of the peak at a set height (such as 5% or 10% of the peak height). The system compares the calculated symmetry value with the pre-set, but user-configurable, qualified standard (for example, Tf should be between 0.9 and 1.5). Any peak that exceeds this range will be marked as a "tailing peak" or "fronting peak" or other specific abnormal types. In addition, the system will also evaluate the single-peak nature of the peak by detecting whether there are "shoulder peaks" or "bifurcations" at the peak top to identify "overlapping peaks".
[0066] Finally, a chromatographic peak is only considered a "normal peak" if it meets the dynamically adjusted width requirement, symmetry standard, and other morphological indicators (such as single-peak nature); otherwise, any unqualified indicator will be classified as a corresponding "abnormal peak", and its detailed morphological feature data (such as peak shape classification, tailing factor specific value, etc.) will be recorded to provide accurate input for subsequent optimization suggestion retrieval.
[0067] Step 106, retrieve the liquid phase method optimization scheme that matches the morphological feature data of the abnormal peak in the pre-set optimization database, including gradient elution time adjustment, mobile phase proportion optimization, or chromatographic column replacement suggestion.
[0068] This step aims to combine automated fault diagnosis with intelligent solution suggestion, building a closed-loop workflow from problem identification to solution generation. The core of this process is a pre-set optimization database, which is essentially a structured expert knowledge base that stores a large number of "problem-reason-solution" triple correspondence relationships.
[0069] The "key" or query index of the optimization database is the morphological feature data of the abnormal peak identified in the previous step. It is not just a single "abnormal" label, but a multi-dimensional feature vector that can include: the specific classification of the abnormal peak (such as "tailing peak", "fronting peak", "broadening peak"), the severity of the abnormality (such as the tailing factor value), the physicochemical properties of the analyte (such as acidity and alkalinity, if known), and the current liquid phase method conditions (such as column type, mobile phase composition, etc.).
[0070] When the system receives the morphological feature data of an abnormal peak, it will use it as a query request to perform a matching search in the optimization database. The search algorithm can range from simple exact matching (e.g., searching for all solutions for "alkaline compound tailing peak") to more complex fuzzy matching or rule-based reasoning. For example, when a severe tailing peak (Tf>2.0) is detected and the mobile phase is neutral, the database will preferentially match high-priority solutions such as "adjust the mobile phase pH to acidic to suppress the effect of silanol groups" or "add a competitive alkaline additive to the mobile phase". The "liquid phase method optimization solutions" stored in the database are specific and actionable recommendations, covering aspects such as gradient elution time adjustment (e.g., extending the elution time to improve resolution), mobile phase proportion optimization (e.g., changing the initial or final proportion of the organic phase), mobile phase component modification (e.g., replacing buffer salts or adjusting pH), and column replacement suggestions (e.g., suggesting a different column with different packing material or particle size to improve column efficiency or selectivity).
[0071] The search results are not a single recommendation, but a list of recommendations prioritized according to historical success rate, ease of implementation, or universality, providing users with clear and logical guidance.
[0072] In this embodiment, the construction of the optimization database is a process of structuring, digitizing, and systematizing the implicit knowledge of chromatography experts and the deterministic principles of chromatographic separation science. It begins with a multi-dimensional knowledge framework that aims to establish an exact mapping relationship between quantifiable "abnormal peak morphological features" (problems) and specific "liquid phase method adjustment strategies" (solutions).
[0073] In the specific construction, first, the troubleshooting experience of experienced chromatography analysts, authoritative chromatography theory monographs, application notes from instrument manufacturers, and a large number of proven data from published scientific literature need to be collected and organized. Then, these unstructured knowledge is converted into structured data entries. Each entry contains a problem vector and a solution set.
[0074] The problem vector is multi-dimensional, including not only the type of abnormal peak (e.g. tailing, fronting, splitting), but also the quantified indicator of its severity (e.g. tailing factor > 2.0), the key physicochemical properties of the target analyte (e.g. pKa value, hydrophobicity), and the core parameters of the current method (e.g. column packing type, mobile phase pH, organic phase type).
[0075] The solution set corresponds to one or more prioritized optimization schemes, each containing four core elements: the adjustment instruction (e.g. "lower mobile phase A pH to 3.0"), the theoretical basis (e.g. "suppress secondary interactions of residual silanol groups"), the expected effect (e.g. "improve peak symmetry"), and the potential risks or considerations (e.g. "confirm whether the analyte is stable under acidic conditions").
[0076] By establishing such a large and detailed rule base or relational model, the database becomes a "smart brain" that can simulate the fault diagnosis and method optimization of chromatography experts, providing a solid foundation for subsequent automated retrieval.
[0077] Step 107, send the liquid phase method optimization scheme to the user.
[0078] Specifically, it can be delivered through a software-integrated graphical user interface (GUI) in the form of an interactive, visual diagnostic report, the core meaning of which is to convert the results of automated data analysis into actionable intelligence that directly drives experimental decisions.
[0079] Specific implementation, when the user reviews the analysis results, the system will intuitively point out the problematic chromatographic peaks on the interface through highlighting, marking, etc., and display their quantified morphological parameters (e.g. tailing factor = 2.3) and abnormal classification (e.g. "severe tailing") side by side.
[0080] Next, the system will display the solutions matched from the optimization database in a clear list according to the recommended priority. Each solution is presented in a structured way of "what to do (adjustment suggestion) - why to do it (principle explanation) - how to do it (specific operation guide)".
[0081] The significance of this design lies in the construction of a seamless cognitive bridge from "finding problems" to "understanding problems" to "solving problems", greatly reducing the dependence on individual experience of analysts. This not only provides expert-level guidance for inexperienced technical personnel, avoiding blind trial and error and significantly shortening the method development and optimization cycle; at the same time, it also provides an efficient decision-making aid tool for experienced experts, enabling them to quickly verify and judge and conduct systematic optimization.
[0082] Ultimately, this delivery method will upgrade software from a simple data processing tool to an intelligent partner that can deliver knowledge, guide practice, and ensure the scientific, standardized, and efficient development of the entire laboratory method development process.
[0083] The embodiment first realizes objective and quantitative evaluation of data quality of different samples by calculating signal complexity indicators such as signal-to-noise ratio and peak density, and dynamically adjusts peak shape identification parameters based on this. This adaptive mechanism greatly improves the accuracy and robustness of abnormal peak identification in complex matrix samples, overcoming the defect of high misjudgment rate under low signal-to-noise ratio caused by traditional fixed parameter method.
[0084] Secondly, and most innovatively, the embodiment breaks through the limitation of existing technology that can only "find problems". By establishing a correlation database between abnormal peak morphology characteristics and liquid phase method optimization schemes, it realizes intelligent recommendation from automatic identification of abnormal peaks to root cause optimization suggestions. This design not only externalizes and systematizes the tacit knowledge of experienced analysts, greatly reducing the dependence on the experience of operating personnel, but also significantly shortens the cycle of method development and troubleshooting. In summary, the technical effect of the present application is reflected in that it upgrades peak shape identification from an independent quality control link to an automated system integrating accurate diagnosis and intelligent decision support, thereby fundamentally improving analysis efficiency, data reliability, and the intelligent level of the entire LC-MS workflow.
[0085] Figure 2 The distribution of multiple chromatographic peaks (Peak I, Peak II, Peak III) is described, where Peak I and Peak II are relatively close and may have some overlap risk. The chart clearly shows the distribution and separation of three adjacent chromatographic peaks (Peak I, Peak II, Peak III) in liquid chromatography-mass spectrometry (LC-MS) analysis. The horizontal axis (X-axis) is the data point (0-1000), representing the time or scan point sequence, and the vertical axis (Y-axis) is the intensity (0-5), representing the response value of the mass spectrometer detector.
[0086] Figure 2Peak I is located at about 400 data points, with an intensity close to 5, showing a high and sharp shape, indicating that the component concentration is high and the chromatographic behavior is ideal; Peak II is located at about 550 data points, with a significantly lower intensity (about 1), a wide and low peak shape, suggesting that it may be a low-abundance component or there is integration difficulty; Peak III is located at about 750 data points, with an intensity of about 2.5, and a shape between the first two. Most importantly, the peak valley between Peak I and Peak II does not completely drop to the baseline, and the distance between the two peaks is relatively close (about 150 data points), which directly reveals the risk of incomplete chromatographic separation (i.e. peak sticking phenomenon), and this partially overlapping peak shape will seriously affect the qualitative and quantitative accuracy of Peak II.
[0087] And the method of the embodiment solves the co-elution problem caused by insufficient peak separation degree shown in the prior art through its adaptive peak recognition algorithm. Figure 2
[0088] Specifically, first, the peak density index of the region is calculated (based on the number and distribution of peaks per unit data point), and it is identified that there is an overlap risk between Peak I (about 380 data points) and Peak II (about 500-650 data points) - the characteristic is that the valley between the two peaks does not drop to the baseline (intensity about 0.8), and the intensity of Peak II is low (about 1) and the peak shape is wide, which is easily covered by the tailing of the high intensity of Peak I.
[0089] Subsequently, the method triggers dynamic parameter adjustment: narrows the peak width tolerance range (for example, reduces the maximum allowed peak width from 100 data points to 60), and applies a clustering algorithm (such as DBSCAN) to finely segment the overlapping region, forcing the distinction between the falling edge of Peak I and the rising edge of Peak II, so as to ensure that Peak II can be independently identified and accurately integrated.
[0090] At the same time, the optimization suggestion module based on the historical knowledge base matches similar cases (such as "low-intensity wide peak and adjacent high peak sticking"), outputs targeted liquid phase condition optimization scheme (for example, "extend the gradient elution time to expand the retention time window of the 500-700 data point region" or "adjust the organic phase ratio in the mobile phase to improve selectivity"), effectively improving the chromatographic separation efficiency and avoiding quantitative errors. This process does not require human intervention, realizing the closed-loop automation from problem diagnosis to solution generation.
[0091] In some embodiments, the embodiment also provides another liquid chromatography mass spectrometry data peak shape recognition processing method, specifically comprising the following steps:
[0092] S201. Obtain liquid chromatography mass spectrometry data.
[0093] This step can refer to the foregoing embodiments, which will not be repeated here.
[0094] S202. Based on the extracted ion flow chromatogram time series, an initial peak shape feature set is generated.
[0095] The purpose of this step is to comprehensively and unbiasedly characterize each preliminary detected chromatographic peak, and the core purpose is to convert one-dimensional time series signal into a high-dimensional digital fingerprint, i.e. the initial peak shape feature set.
[0096] First, a series of standard signal processing and peak recognition algorithms are applied to the input extracted ion flow chromatogram time series, such as baseline correction, smoothing and noise reduction, and peak boundary determination.
[0097] After the precise start and end times are determined for each potential peak, the system will perform an exhaustive feature calculation program. This program not only calculates conventional morphological parameters such as peak height, retention time, peak area, half-peak width, and peak width and asymmetry factor at different peak height percentages (e.g. 5%, 10%), but also deeply mines deep features that can describe subtle changes in peak shape. These features can include: statistical moments that describe the skewness and sharpness of the peak shape (such as the third moment skewness and the fourth moment kurtosis); model parameters obtained after fitting the peak shape to an ideal mathematical model (such as Gaussian function, exponential modified Gaussian function) (such as Gaussian standard deviation, exponential decay constant); and other shape factors that describe the irregularity of the peak shape.
[0098] Finally, each chromatographic peak is mapped to a high-dimensional vector containing dozens or even hundreds of numerical features, and this vector set constitutes the initial peak shape feature set, which serves as the original and complete data basis for subsequent feature optimization. The design concept is "better safe than sorry", ensuring that all potential information related to peak shape quality is initially captured.
[0099] S203. Input the initial peak shape feature set into the pre-trained feature optimization model, and output the optimized feature weight set.
[0100] The core of this step is to use machine learning techniques to intelligently evaluate the importance of the large and redundant initial feature set generated in the previous step, in order to distinguish the key features that have high discrimination for peak shape quality.
[0101] Among them, the pre-trained feature optimization model is a complex mathematical model based on supervised learning algorithms, such as Gradient Boosting Decision Trees (GBDT), Random Forest, or specific structure neural network.
[0102] The pre-training process of this feature optimization model is done on a large-scale, expert-annotated historical LC-MS dataset. This dataset contains thousands of chromatographic peaks, each of which is explicitly labeled as a "normal peak" or a specific abnormal class (such as a tailing peak, a fronting peak, a split peak, etc.), and is accompanied by its complete initial peak shape feature set. The model learns the correspondence between these "features-labels" to automatically identify and quantify the contribution of each feature to correct classification.
[0103] In actual application, the system takes the "initial peak shape feature set" of the current chromatographic peak to be analyzed as input and delivers it to this already trained model. The model internally scores each feature of the input through its complex decision logic (such as tree split gain or feature permutation importance calculation), and finally outputs an "optimized feature weight set".
[0104] This optimized feature weight set is a numerical vector with the same dimension as the initial feature set, where each value (i.e., weight) accurately quantifies the relative importance of the corresponding feature in distinguishing normal peaks from abnormal peaks in this specific task. The higher the weight value, the stronger the discriminant ability of the feature.
[0105] S204. Filter low-weight morphological features according to the feature weight set to obtain morphological feature data.
[0106] Specifically, first, a fixed numerical value is set as the weight threshold. Then, each feature weight in the feature weight set output in the previous step is compared with the threshold. If the weight value of the feature is greater than or equal to the threshold, the feature is determined to be an important feature and is retained; if it is less than the threshold, its contribution is considered to be low and it is removed.
[0107] After comparing and screening all features, all high-weight features retained are recombined into a new feature set, which is the final optimized morphological feature data used for peak shape identification.
[0108] In some embodiments, another implementation is also provided.
[0109] The system first normalizes the optimized feature weight set output in the previous step so that the sum of the weights of all features is equal to 1, thereby converting each weight value to the contribution percentage of the feature to the overall performance of the model.
[0110] Subsequently, all the morphological features and their corresponding normalized weights are sorted in descending order of their weight values. Then, starting from the feature with the highest weight, the system accumulates the weight values one by one in descending order and monitors the cumulative sum in real time. When the cumulative sum first reaches or exceeds a pre-set cumulative contribution threshold (e.g., 95%), the accumulation process stops immediately. At this point, all the features that have been included in the cumulative calculation, i.e., from the feature with the highest weight to the last feature that meets the threshold condition, collectively constitute the high-weight feature subset that is filtered out. All the features that are outside this sorting and are not included in the cumulative sum are considered as low-weight features and are discarded.
[0111] Finally, the set composed of this filtered high-weight feature subset is the morphological feature data that is finally used for peak shape classification.
[0112] This method is adaptive, as it does not care about the absolute weight of individual features, but focuses on the overall performance of feature combinations, and can dynamically determine the number of features to be retained according to the importance distribution of different data sets, thereby achieving a better balance between maximizing the retention of key information and minimizing the feature dimension.
[0113] In some embodiments, step S204 can specifically include:
[0114] S2041. Determine the weight threshold based on the weight distribution of the confirmed qualified peak shape features in the historical data.
[0115] This step aims to establish an objective and data-driven decision-making benchmark for the feature filtering link.
[0116] The system first accesses a pre-constructed and expert-confirmed historical database that stores a large number of labeled "qualified" or "ideal" chromatographic peak samples. For each qualified peak in the library, the system performs the same feature extraction and weight calculation process as the sample to be tested, thereby obtaining a set of feature weights that are specific to the qualified peaks.
[0117] By statistically analyzing the weight sets of a large number of qualified samples, the system can depict the probability distribution of the importance of each morphological feature in an ideal peak shape. Based on this distribution, the system will use a robust statistical quantity to determine the weight threshold, for example, the weight value of a certain low percentile (such as the 10th percentile) or the mean minus several times the standard deviation (μ-kσ) of all features in the qualified peak sample set can be set as the weight threshold. The core significance of this weight threshold is that it represents the lowest effective contribution that a feature can exhibit in a recognized qualified peak shape, and any feature below this benchmark can be considered statistically insignificant in terms of importance.
[0118] S2042. Compare each feature weight in the feature weight set with the weight threshold, and filter the morphological features with weight lower than the weight threshold, to obtain the filtered high-weight features.
[0119] This step is a deterministic filtering operation, aiming to refine the initial feature set of the current chromatographic peak according to the objective criteria established in the previous step.
[0120] The system receives the feature weight set of the current peak and the weight threshold determined in the previous step as inputs. This process is implemented through an iterative comparison algorithm: the system will traverse each element in the feature weight set, i.e., each morphological feature and its corresponding weight value.
[0121] In each iteration, the system directly compares the weight value of the current feature with the pre-set weight threshold. If the weight value of the feature is greater than or equal to the weight threshold, it is determined that this feature is a high-weight feature that significantly contributes to the quality assessment of the peak shape, and it is retained; otherwise, if its weight value is less than the weight threshold, it is determined that this feature is redundant or noise information, and it is filtered out from the candidate list. This process is repeated until all features are evaluated.
[0122] Finally, all the features that are retained through comparison together form a new set, i.e., the filtered high-weight features, which is an optimized subset of the original feature space with more information content.
[0123] S2043. Generate morphological feature data according to the filtered high-weight features.
[0124] This step is the final link of the feature selection process, which aims to reorganize the selected key feature information into a structured and standardized data entity for use by subsequent classification models.
[0125] Specifically, first, use the filtered high-weight feature set as an index. According to this index list, the system returns to the initial peak shape feature set calculated for the current chromatographic peak, which contains all features and their original values, to search. For each high-weight feature identifier (such as feature name) in the index list, the system will accurately locate the feature in the initial feature set and extract its corresponding value.
[0126] Subsequently, the system assembles these extracted (feature identifier, feature value) data pairs into a dimension-determined and fixed-order feature vector or similar data structure.
[0127] The final generated feature vector, namely the "morphological feature data", represents a highly condensed and optimized digital portrait of the original chromatographic peak morphology. Not only is it of lower dimensionality and more computationally efficient, but also its effectiveness and robustness as input to subsequent classification algorithms are significantly enhanced due to the elimination of irrelevant variable interference.
[0128] S205. According to the preset target substance information, the chromatographic peak detected in the retention time window in the extracted ion chromatogram time sequence is determined as the sample peak.
[0129] This step aims to accurately locate the target signal for signal-to-noise ratio calculation from the entire chromatogram. The system first retrieves the key identification information of the target analyte from the preset analysis method or database, mainly including its specific mass-to-charge ratio (m / z) and expected theoretical retention time.
[0130] Subsequently, the system sets a retention time window with a specific tolerance range around the theoretical retention time, for example, expected retention time ± 0.2 minutes, to cope with the chromatographic drift caused by instrument or experimental condition changes. The retention time window is a time interval with a small floating range set around the expected retention time of the target analyte, used to accurately identify and capture the target chromatographic peak within the range.
[0131] Finally, the system only performs peak detection algorithm on the extracted ion chromatogram time sequence within this specified time window, and the chromatographic peak with the strongest response or the largest area detected within the window is explicitly anchored and marked as the "sample peak". This ensures that the subsequent signal-to-noise ratio calculation is for the true target compound signal, not for other interferents or noise.
[0132] S206. In the baseline adjacent area of the sample peak, a time period equivalent to the peak width of the sample peak is determined as the baseline time period.
[0133] The core of this step is to select a representative and fair baseline interval for noise quantification. After determining the sample peak, the system first accurately calculates the width of the peak, which is usually the difference between the start and end time points at the baseline.
[0134] Then, the system automatically searches for the baseline adjacent area adjacent to the sample peak. The baseline adjacent area refers to a flat area with no significant signal response before or after the sample peak elutes, only showing instrument background fluctuations.
[0135] To ensure the comparability of noise and signal in evaluation, the system will extract a time length equivalent to the measured peak width of the sample peak in this baseline adjacent area, forming a specific "time period", which is the "baseline time period".
[0136] This operation guarantees the noise area for the subsequent calculation, which is measured on the same scale as the signal duration, so that the comparison between signal and noise is more statistically meaningful and reasonable.
[0137] S207. Calculate the signal integral area of the baseline period, and determine the baseline noise area.
[0138] This step aims to convert the random fluctuations in the baseline period selected in the previous step into a quantitative value.
[0139] The system first extracts all (time, intensity) data points contained in the baseline period to form a baseline signal. Before integration, the baseline signal is usually corrected to zero, that is, the average intensity value of the signal is subtracted, to eliminate the effect of baseline drift, so that the integral result can better reflect the true noise fluctuation amplitude.
[0140] Then, the system applies standard numerical integration algorithms such as the trapezoidal rule or Simpson's rule to perform point-by-point integration on the corrected baseline signal, and calculates the sum of the absolute areas enclosed by the signal intensity curve and the zero baseline in the baseline period. The final integral value calculated, which is defined as the "baseline noise area", represents the total amount of instrument background noise in the same time span as the sample peak.
[0141] S208. Calculate the ratio of the peak area of the sample peak to the baseline noise area to obtain the signal-to-noise ratio value.
[0142] This step is the final quantification step of the signal-to-noise ratio index, which uses a strategy based on the comparison of the total signal response and the total noise response.
[0143] The system first obtains the peak area of the sample peak, which is obtained by numerically integrating the sample peak signal between its start and end points in the early peak detection and integration process, representing the total signal amount of the target substance.
[0144] At the same time, the system retrieves the baseline noise area calculated in the previous step. Then, the system performs an arithmetic division operation, that is, the peak area of the sample peak is divided by the baseline noise area. This calculation result is a dimensionless ratio, which directly reflects the size relationship between the cumulative intensity of the target signal and the cumulative intensity of the background noise. This ratio is determined as the "signal-to-noise ratio" defined by this method, which is used to measure the clarity and reliability of the signal.
[0145] S209. Determine the peak density value based on the number of intersections between the chromatographic peak at the set peak height and the baseline.
[0146] The purpose of this step is to quantify the crowdedness of the whole chromatogram. First, the system will perform a global peak search operation on the whole extracted ion current chromatogram time series to identify all potential chromatographic peak events. In this process, the system will only confirm the signal events with peak top intensity exceeding a pre-set intensity threshold, i.e. the“set peak height”, as valid chromatographic peaks to filter out low-level baseline noise. The system will count the total number of valid chromatographic peaks identified within the whole analysis time window. This total number, i.e. the actual quantification result corresponding to the concept of“number of intersections with baseline at set peak height” (each peak corresponds to two baseline intersections, so the number of peaks is half of the number of intersections, and they are proportional to each other). Finally, divide this peak total number by the total analysis time, and the peak number per unit time, i.e. the“peak density value” is obtained. The higher this value, the more complex the chromatographic separation, and the greater the possibility of peak overlap.
[0147] S210. Integrate the signal-to-noise ratio value and the peak density value to obtain the signal complexity index.
[0148] This step is to integrate the data of the two independent evaluation dimensions calculated in the previous steps into a multi-dimensional comprehensive index. The system combines the scalar values“signal-to-noise ratio” and“peak density” calculated in the previous steps as an ordered data pair. In technical implementation, this is usually constructed as a two-dimensional feature vector, such as [signal-to-noise ratio, peak density], or a data structure object containing two named attributes. This integrated data entity, i.e. the“signal complexity index”, can give a comprehensive and quantitative description of the overall complexity of the chromatogram from the vertical quality (clarity, reflected by signal-to-noise ratio) and horizontal quality (crowdedness, reflected by peak density) of the signal. Two orthogonal dimensions provide a direct and rich decision basis for subsequent adaptive adjustment of identification parameters.
[0149] S211. Adjust the peak shape identification parameters according to the signal complexity index, including the noise threshold and the peak width range.
[0150] This step can refer to the previous embodiments, which will not be repeated here.
[0151] S212. Generate a multi-dimensional peak shape feature vector based on the morphological feature data and the adjusted peak shape identification parameters.
[0152] The core task of this step is to convert each chromatographic peak to be analyzed into a standardized mathematical representation containing all its key information.
[0153] The system first locates to each chromatographic peak identified by dynamic parameter. For each peak, the system extracts and integrates two types of information: the first type is the previously calculated and screened morphological feature data, that is, the set of high-weight feature values that can most effectively describe the peak shape; the second type is the adjusted peak shape identification parameters used in this analysis, such as the actual noise threshold and peak width, which are also considered as environmental or contextual information.
[0154] The system combines and arranges these two types of numerical information in a predefined order to form an ordered numerical list, and this fixed-length list is the "multi-dimensional peak shape feature vector" described above. It constructs a comprehensive, high-dimensional digital fingerprint for each chromatographic peak, which is the direct data basis for subsequent pattern recognition and clustering analysis.
[0155] S213. Perform dimensionality reduction on the multi-dimensional peak shape feature vector to obtain a two-dimensional spatial coordinate set.
[0156] This step aims to convert the high-dimensional, abstract feature vector into a coordinate point that can be intuitively represented and calculated on a two-dimensional plane, facilitating subsequent clustering analysis.
[0157] Since direct distance calculation in high-dimensional space has the "curse of dimensionality" problem, that is, the effectiveness of distance measurement decreases with increasing dimension, dimensionality reduction must be performed. The system will use a nonlinear manifold learning algorithm such as t-distributed stochastic neighbor embedding (t-SNE) or uniform manifold approximation and projection (UMAP) to process the "multi-dimensional peak shape feature vector" set of all peaks. The core idea of this algorithm is to project the data into a low-dimensional (here, two-dimensional) space while maintaining the local neighborhood structure of the data points in the high-dimensional space.
[0158] After the algorithm is executed, each original high-dimensional feature vector is mapped to a unique (x, y) coordinate pair. The set of all these coordinate pairs, that is, the "two-dimensional spatial coordinate set" described above, in which chromatographic peaks with similar morphologies are close to each other in the two-dimensional space.
[0159] S214. Calculate the Euclidean distance between each coordinate point in the two-dimensional spatial coordinate set to obtain a spatial distance matrix.
[0160] This step aims to quantify the pairwise similarity of all chromatographic peaks in the feature space after dimensionality reduction.
[0161] The system takes the "two-dimensional spatial coordinate set" as input, and calculates the Euclidean distance between any two coordinate points P i (x i ,y i ) and P j (x j ,yj ) apply the Euclidean distance formula, i.e. This formula calculates the straight-line distance between two points in a two-dimensional plane. The system will systematically and exhaustively calculate the distance between every pair of coordinate points in the dataset. All these distance values are finally organized into an N x N symmetric square matrix, where N is the total number of chromatographic peaks, which is the so-called “spatial distance matrix”.
[0162] The value of the element located at the i-th row and j-th column in the spatial distance matrix precisely represents the degree of similarity (the smaller the distance, the higher the similarity) between the i-th chromatographic peak and the j-th chromatographic peak in terms of morphological characteristics, providing the necessary input for the subsequent clustering algorithm.
[0163] S215. Based on the spatial distance matrix, classify the chromatographic peaks into normal clusters or abnormal clusters through a clustering algorithm.
[0164] This step is an unsupervised pattern discovery process, aiming to automatically group chromatographic peaks according to peak shape similarity.
[0165] The system takes the “spatial distance matrix” generated in the previous step as input and applies a density-based clustering algorithm, such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise). The DBSCAN algorithm does not require the number of clusters to be specified in advance. It identifies core points by checking the density of points in their neighborhood and continuously expands from core points to merge density-reachable points into a cluster. Based on the general assumption that well-shaped normal peaks dominate in number and have high morphological consistency in a single analysis, they will form a core area with the largest size and highest density in the two-dimensional feature space.
[0166] This algorithm automatically identifies this dominant high-density area as a cluster, which the system immediately defines as the “normal cluster”. All other scattered points or small-scale clusters with lower density and far from the core area are classified as “abnormal clusters”.
[0167] S216. Label the chromatographic peaks in the normal cluster as normal peaks and the chromatographic peaks in the abnormal cluster as abnormal peaks.
[0168] This step is the last step of the classification process, and its function is to convert the abstract grouping results output by the clustering algorithm into explicit, business-meaningful final labels for each chromatographic peak.
[0169] The system will check the attribution information of each chromatographic peak generated in the previous clustering process one by one. For any chromatographic peak, the system will query whether the cluster it belongs to is a normal cluster or an abnormal cluster. If the peak is assigned to a normal cluster defined as dominant, the system will add a classification label of normal peak to the metadata of the peak. Conversely, if the peak is assigned to any abnormal cluster, whether it is a noise point or belongs to a small abnormal group, the system will add an abnormal peak label to it.
[0170] After this operation, all chromatographic peaks in the entire dataset are assigned a clear quality evaluation conclusion, thus completing the entire process of automatic peak shape recognition and classification.
[0171] S217. Encode the morphological feature data of the abnormal peak into a feature vector.
[0172] This step aims to convert the descriptive information of the identified abnormal chromatographic peak into a standardized mathematical format that can be processed by machine learning models.
[0173] The system first extracts the morphological feature data corresponding to the abnormal peak, i.e. the set of high-weight features and their numerical values after screening. Subsequently, the system arranges these numerical values in order according to a pre-set, globally unified feature sequence, forming a numerical array with fixed dimensions. To eliminate the calculation bias caused by the dimension difference between different features, the system usually normalizes or standardizes the array, such as scaling each element to the [0, 1] interval or converting it to a standard normal distribution with mean 0 and standard deviation 1.
[0174] This ordered and scaled numerical array is the feature vector of the abnormal peak, which is like a "digital fingerprint" of the peak's abnormal morphology, and can be accurately used by subsequent models for similarity measurement and pattern matching.
[0175] S218. Input the feature vector into the pre-trained solution matching model to calculate the similarity between the feature vector and the historical abnormal peak feature vector through the solution matching model, and obtain the similarity value and the matched optimization solution.
[0176] This step is the core of abnormal diagnosis and solution retrieval. The system submits the feature vector of the current abnormal peak to a solution matching model that has been pre-trained with historical data. The model maintains a large knowledge base inside, which stores a large number of historical abnormal peak feature vectors, and each historical vector is bound to an expert-verified, explicit fault cause and corresponding "optimization solution".
[0177] After receiving a new feature vector, the model calculates the "distance" or "angle" between the input vector and all historical vectors in the knowledge base in the multi-dimensional feature space using an efficient similarity measurement algorithm such as cosine similarity or radial basis function kernel. The model finds one or more historical vectors that are most similar to the input vector and outputs two key results: a quantitative "similarity value" representing the matching degree of the current anomaly with historical cases, and the "optimization scheme" associated with the most similar historical case.
[0178] In this embodiment, the construction and training of the scheme matching model is a supervised learning process based on historical experience data. First, a detailed knowledge base needs to be constructed, which is composed of a large number of historical abnormal peak samples, each containing two key pieces of information: one is the complete morphological feature vector of the abnormal peak as the input data (X) of the model, and the other is the root cause and the corresponding solution text determined by expert artificial diagnosis after the abnormal peak occurs as the label (Y) of the model.
[0179] In the training phase, the model (which can use K-Nearest Neighbors (KNN), Support Vector Machine (SVM), or more complex deep learning-based Siamese Network) learns how to measure the similarity between different feature vectors and establishes an accurate mapping relationship from a specific abnormal feature pattern to the correct solution label. For example, for a Siamese network, the model learns an embedding space by optimizing the loss function, so that feature vectors of abnormal peaks caused by the same cause are close to each other in the space, while those caused by different causes are far apart. After training, the model has the ability to accurately match new unknown abnormal peaks and automatically recommend solutions.
[0180] S219. When the similarity value exceeds the pre-set confidence threshold, output the matched liquid phase optimization scheme as the retrieval result.
[0181] This step is a decision and filtering link to ensure the reliability and accuracy of the output suggestions.
[0182] The pre-set confidence threshold is a key parameter (e.g., 0.95) determined by domain experts based on experience or statistical analysis of historical data, which represents the minimum similarity standard that the system considers a matching result to be acceptable. The system directly compares the "similarity value" calculated by the model in the previous step with this threshold. Only when the similarity value is greater than or equal to the confidence threshold, the system will determine that the matching is highly reliable, i.e., the cause of the current abnormal peak is likely to be the same as the matched historical case. In this case, the system will present the corresponding liquid phase optimization scheme (such as "suggest cleaning the sample inlet" or "check the gas path tightness") as the final and effective retrieval result to the user.
[0183] If the similarity value is lower than the threshold value, the system will determine that the matching degree is insufficient and will not provide any suggestions to avoid giving false guidance.
[0184] S220. Send the liquid phase method optimization scheme to the user.
[0185] This step can refer to the foregoing embodiments, which will not be repeated here.
[0186] In some embodiments, it is also necessary to consider improving the depth and efficiency of human-computer interaction. In this step, an "interactive diagnosis and virtual optimization workbench" integrating causal inference explainable AI (Causal-XAI) and a physics-informed neural network digital twin (Physics-Informed Neural Network Digital Twin, PINN-DT) can be designed. This design aims to free analysts from the traditional troubleshooting mode of "guessing-trying" and enter a new paradigm of "insight-verification-optimization". The specific implementation includes:
[0187] First, when the system identifies an abnormal peak, its built-in Causal-XAI engine starts. Instead of simply showing feature importance, the engine reversely infers the observed abnormal peak morphology features (effects) to the most likely physical and chemical root causes (reasons) based on a pre-trained causal graph containing chromatographic mechanisms, and presents them to the user in the form of a probabilistic "root cause hypothesis list" (for example: "column contamination causes secondary interaction" with a confidence of 75%, "improper mobile phase pH" with a confidence of 20%). At the same time, the system highlights the key morphology areas supporting each hypothesis in the form of a "heat map" on the chromatogram, achieving complete transparency in the diagnosis process.
[0188] Next, the user can select one or more high-confidence root cause hypotheses to activate the PINN-DT module. Instead of a traditional empirical model, this digital twin is a deep learning network whose loss function not only contains a data fitting term but also forcibly embeds partial differential equations that describe the liquid chromatography process (such as the general rate theory model), enabling it to not only accurately match the "individuality" of a specific system but also strictly adhere to the physical laws of chromatographic separation.
[0189] In the activated "virtual optimization workbench" interface, the Causal-XAI module automatically maps the selected hypothesis to adjustable key experimental parameters (such as selecting "column contamination", which automatically emerges options such as "column regeneration parameters", "virtual replacement of new column", etc.). The user can adjust these parameters in the virtual environment as if on a real instrument, and the PINN-DT can render the adjusted predicted chromatographic peak shape in real time within seconds.
[0190] Further, the user can define optimization goals (such as "symmetry factor > 0.95 and retention time drift < 2%"), and the system can start a built-in Bayesian optimization algorithm to automatically explore the parameter combinations in the digital twin space and recommend a "best solution" that has been virtually verified.
[0191] This seamless closed loop from causal diagnosis to virtual experiment upgrades the software's function from "post hoc advice" to "proactive prediction and active optimization", greatly shortening the method development and troubleshooting cycle, and providing analysts with unprecedented mechanistic insights.
[0192] S221. Receive the user's feedback on the confirmed effective final optimization scheme for the abnormal peak.
[0193] This step is the starting point of building the system's self-learning and evolution ability, aiming to make the tacit knowledge of human experts explicit and integrate it into the system knowledge base.
[0194] The system provides a human-computer interaction interface. When a chromatographic peak is marked as "abnormal", the operator or domain expert will troubleshoot according to the actual situation and take appropriate corrective measures. Once the problem is solved, the expert can submit a clear feedback through this interface for this specific abnormal peak instance. The core content of this feedback is the "confirmed effective final optimization scheme", which may be confirmed in the system recommended scheme, may be selected from a pre-set standardized scheme library, or may be a new, practice-tested solution description input by the user in free text form.
[0195] The system background will accurately record this feedback information and strictly bind it with the unique identifier of the abnormal peak that triggered this feedback (such as analysis time, peak ID, etc.), ensuring the uniqueness and traceability of the feedback information source, and laying a solid foundation for subsequent data association.
[0196] S222. Associate the morphological feature data of the abnormal peak with the final optimization scheme to obtain a feature-optimization mapping data pair.
[0197] This step is the core link of data processing and integration, and its goal is to create a complete, model-learnable supervised sample.
[0198] After receiving the user's submitted effective feedback record, the system first retrieves the complete "morphological feature data" (i.e. high-dimensional feature vector) corresponding to the abnormal peak when it was first detected from the historical database according to the unique identifier in the feedback information.
[0199] Subsequently, the system normalizes the user feedback "final optimization scheme" text. This process may involve natural language processing (NLP) techniques such as keyword extraction, synonym merging, and intent recognition, aiming to map the user's free text or options into a predefined, standardized solution category label (for example, both "I cleaned the ion source" and "performed ion source maintenance" are classified as "standard operation-ion source cleaning").
[0200] Finally, the system binds the feature vector of this abnormal peak as input features (X) and the normalized solution label as target output (Y) together, forming a new, formatted "feature-optimization mapping data pair". This data pair constitutes the basic unit of the system's knowledge base growth, accurately recording the causal relationship between a "specific abnormal phenomenon" and its "effective solution".
[0201] S223. Update the feature-optimization mapping data pair to the optimization database to optimize the accuracy of subsequent retrieval matching.
[0202] This step is the closed-loop endpoint of the system's self-evolution and performance improvement. The system persistently stores the "feature-optimization mapping data pair" generated in the previous step and adds it to the "optimization database" as the "scheme matching model" knowledge source. This operation is not just a simple data addition, but more importantly, it triggers the model's learning and updating mechanism.
[0203] According to the design of the system architecture, this update can take two main strategies: one is batch retraining (Batch Retraining), which periodically (for example, during the night low peak period) integrates the new data into the entire training set, re-trains the scheme matching model, generates a new version of the model with better performance and more comprehensive knowledge, and deploys it online; the other is incremental learning (Incremental Learning), for models that support this feature (such as some types of neural networks or K-nearest neighbor algorithms), the system can integrate new data pairs into the existing model in real time or quasi-real time, dynamically adjust the model parameters, so that it can immediately benefit from new knowledge.
[0204] Regardless of the strategy adopted, this step ensures that the system can continuously learn from users' actual operation experience, and its knowledge base becomes increasingly rich over time, significantly improving the accuracy, coverage, and reliability of its automatic diagnosis and solution recommendation when facing similar or new abnormal peaks in the future.
[0205] In some embodiments, considering the rare chromatographic problems in some complex biological matrix analysis and new chemical entity (NCE) development, this embodiment proposes a distributed knowledge base enhancement scheme based on hierarchical federated transfer learning (HFTL) and proof of contribution (PoC). This scheme can be named as HFTL-PoC framework in this embodiment.
[0206] The HFTL-PoC framework in this embodiment is designed for the collaborative environment of multi-center and multi-instrument platforms (such as UHPLC-QTOF systems of different manufacturers). Under the premise of strictly protecting the privacy of proprietary compound structures and original data of each participant (such as pharmaceutical companies and contract research organizations CRO), efficient knowledge aggregation and personalized model optimization are realized. The specific implementation mode of the HFTL-PoC framework includes:
[0207] Firstly, a three-layer "client-cluster-server" federated learning architecture is established. The software client of each laboratory is the bottom layer, which is responsible for model training on the locally confirmed "feature-optimization mapping data pair" by the user. The middle layer is the "cluster", and the server automatically groups the clients according to the non-identifiable metadata (such as instrument model, chromatographic column type, and application field such as "metabolomics" or "antibody drug analysis") uploaded by the client. The cluster server is responsible for aggregating the model updates of similar backgrounds within the cluster to form a "domain expert model" for a specific application scenario. The top central server aggregates the "domain expert models" to generate a "global basic model" with wide applicability.
[0208] Secondly, the transfer learning mechanism is introduced. The central server does not issue the final model, but a high-performance "global basic model". Each client fine-tunes the model on a small amount of proprietary data locally, thereby enjoying the wisdom of the global while obtaining a personalized solution matching model that is highly consistent with the specific analysis task.
[0209] Moreover, the system integrates an incentive and quality control mechanism based on proof of contribution (PoC). When the client uploads the model update each time, the central server will use an independent and standardized "challenge data set" (containing various classic and difficult chromatographic problems) to evaluate the degree of improvement of the global model performance by the update, and generate a quantifiable "contribution score". This score not only serves as the weight in federated aggregation, but also can be converted into a "knowledge token" through the combination of technologies such as alliance chain, to motivate high-quality data contribution and build a fair and self-consistent knowledge sharing economy.
[0210] Through the HFTL-PoC framework, the training difficulty of standard federated learning under highly heterogeneous (Non-IID) chromatographic data is solved, and through the inherent quality evaluation and incentive mechanism, the professionalism and evolution of the knowledge base are ensured, so that the application is upgraded from a tool software to an industry-level chromatographic method development collaborative platform.
[0211] The method provided by the above embodiment can also be specifically executed by an electronic device. Next, the electronic device in the embodiment of the application is described from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic diagram of an entity device structure of the electronic device in the embodiment of the application.
[0212] It should be noted that Figure 3 The structure of the electronic device shown is only an example, and should not impose any limitation on the functions and use range of the embodiment of the application.
[0213] As Figure 3 shown, the electronic device includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes, for example, execute the method described in the above embodiment, according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage portion 408 to a random access memory (RAM) 403. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0214] The following components are connected to the I / O interface 405: an input portion 406 including an audio input device, a button switch, and the like; an output portion 407 including a liquid crystal display (LCD), an audio output device, an indicator, and the like; a storage portion 408 including a hard disk and the like; and a communication portion 409 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication portion 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 410 as needed, so that a computer program read therefrom is installed in the storage portion 408 as needed.
[0215] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program in accordance with embodiments of the present application. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising computer instructions for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 409 and / or installed from the removable media 411. When the computer program is executed by the central processing unit (CPU) 401, various functions defined in the present application are performed.
[0216] Note that specific examples of computer readable storage media can include without limitation: an electrical connection having one or more wires; a portable computer diskette; a hard disk; a random access memory (RAM); a read-only memory (ROM); an erasable programmable read only memory (EPROM); a flash memory; an optical fiber; a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing. In the present application, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0217] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved.
[0218] In particular, the electronic device of the present embodiment includes a processor and a memory coupled to the one or more processors, the memory storing computer program code comprising computer instructions to be invoked by the one or more processors to cause the electronic device to perform the method provided by the above-described embodiments.
[0219] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The storage medium carries one or more computer programs, which, when executed by a processor of the electronic device, enable the electronic device to implement the methods provided in the above embodiments.
[0220] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still make modifications to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0221] Those skilled in the art can understand that all or part of the flow of the above-mentioned embodiment method can be implemented by a computer program to instruct the relevant hardware to complete, and the program can be stored in a computer readable storage medium. The program, when executed, can include the flow of each method embodiment as described above. The foregoing storage medium includes ROM, random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.
Claims
1. A liquid chromatography mass spectrometry data peak shape recognition processing method, characterized by, The method comprises the following steps: acquiring liquid chromatography mass spectrometry data, wherein the liquid chromatography mass spectrometry data comprises extracted ion current chromatogram time series and liquid phase method condition information; extracting morphological feature data of chromatographic peaks based on the extracted ion current chromatogram time series, wherein the morphological feature data comprises peak points, valley points, peak width, peak area and peak symmetry; calculating a signal complexity index of the extracted ion current chromatogram time series, wherein the signal complexity index comprises signal-to-noise ratio and peak density; adjusting peak shape identification parameters according to the signal complexity index, wherein the peak shape identification parameters comprise noise threshold and peak width range; distinguishing the chromatographic peaks into normal peaks or abnormal peaks by using the adjusted peak shape identification parameters, comprising: processing the extracted ion current chromatogram by using the adjusted noise threshold, identifying signal regions higher than the threshold as potential peak events, and completing accurate start and end point delimitation of the peaks according to the principle that the signal falls back to the baseline or the valley point; applying the adjusted peak width range to the delimited peaks for preliminary screening, calculating the actual width of each peak, and if the actual width is not within the reasonable interval dynamically set, the peak is classified as an abnormal peak; calculating the symmetry index of the chromatographic peaks passing the width preliminary screening, comparing the symmetry index with a preset qualified standard which can be configured by a user, and the peaks exceeding the range will be marked as specific abnormal types; a chromatographic peak is determined as a normal peak only when it meets the dynamically adjusted width requirement, symmetry standard and other morphological indexes at the same time, and if any index is unqualified, the peak will be classified as a corresponding abnormal peak; searching for a liquid phase method optimization scheme matching the morphological feature data of the abnormal peak in a preset optimization database, wherein the liquid phase method optimization scheme comprises gradient elution time adjustment, mobile phase proportion optimization or chromatographic column replacement suggestion; sending the liquid phase method optimization scheme to a user.
2. The method of claim 1, wherein, After the liquid phase method optimization scheme is sent to the designated user, the method further comprises the following steps: receiving a final optimization scheme confirmed by the user as effective in response to the abnormal peak; associating the morphological feature data of the abnormal peak with the final optimization scheme to obtain a feature-optimization mapping data pair; updating the feature-optimization mapping data pair to the optimization database to optimize the accuracy of subsequent search matching.
3. The method of claim 1, wherein, The method of calculating the signal complexity index of the extracted ion current chromatogram time series comprises the following steps: determining a chromatographic peak detected in a retention time window in the extracted ion current chromatogram time series as a sample peak according to preset target substance information; determining a time period corresponding to the peak width of the sample peak as a baseline time period in the adjacent region of the baseline of the sample peak; calculating the signal integral area in the baseline time period and determining it as a baseline noise area; calculating the ratio of the peak area of the sample peak to the baseline noise area to obtain a signal-to-noise ratio value; determining a peak density value based on the number of intersection points of the chromatographic peak at a set peak height and the baseline; integrating the signal-to-noise ratio value and the peak density value to obtain the signal complexity index.
4. The method according to any one of claims 1 to 3, characterized in that, The method of extracting morphological feature data of chromatographic peaks based on the extracted ion current chromatogram time series comprises the following steps: generating an initial peak shape feature set based on the extracted ion chromatogram time series; inputting the initial peak shape feature set into a pre-trained feature optimization model to obtain an optimized feature weight set; filtering low-weight morphological features according to the feature weight set to obtain the morphological feature data.
5. The method of claim 4, wherein, The filtering of low-weight morphological features according to the feature weight set to obtain the morphological feature data comprises: determining a weight threshold value based on the confirmed qualified peak shape feature weight distribution in historical data; comparing each feature weight in the feature weight set with the weight threshold value, and filtering morphological features with weights lower than the weight threshold value to obtain filtered high-weight features; generating the morphological feature data according to the filtered high-weight features.
6. The method of claim 1, wherein, The use of the adjusted peak shape identification parameters to distinguish the chromatographic peaks into normal peaks or abnormal peaks comprises: generating a multi-dimensional peak shape feature vector based on the morphological feature data and the adjusted peak shape identification parameters; performing dimension reduction processing on the multi-dimensional peak shape feature vector to obtain a two-dimensional spatial coordinate set; calculating the Euclidean distance between each coordinate point in the two-dimensional spatial coordinate set to obtain a spatial distance matrix; classifying the chromatographic peaks into normal clusters or abnormal clusters based on the spatial distance matrix through a clustering algorithm; labeling the chromatographic peaks in the normal clusters as the normal peaks, and labeling the chromatographic peaks in the abnormal clusters as the abnormal peaks.
7. The method of claim 1, wherein, The retrieval of a liquid phase method optimization scheme matching the morphological feature data of the abnormal peak in a pre-set optimization database comprises: encoding the morphological feature data of the abnormal peak into a feature vector; inputting the feature vector into a pre-trained scheme matching model to calculate the similarity between the feature vector and historical abnormal peak feature vectors through the scheme matching model to obtain a similarity value and a matched optimization scheme; when the similarity value exceeds a pre-set threshold value, outputting the matched liquid phase optimization scheme as a retrieval result.
8. An electronic device, comprising: comprises one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is configured to store computer program code comprising computer instructions, and the one or more processors are configured to invoke the computer instructions to cause the electronic device to perform the method of any one of claims 1-7.
9. A computer readable storage medium storing computer instructions, characterized in that, When the computer instructions run on the electronic device, the electronic device is caused to perform the method of any one of claims 1-7.
10. A computer program product, characterised in that, When the computer program product runs on the electronic device, the electronic device is caused to perform the method of any one of claims 1-7.
Citation Information
Patent Citations
Automation-based high-throughput analysis method and system, medium and product
CN119375390A
High resolution ms1 based quantification
US20180224406A1