Intelligent liquid chromatography peak analysis method and system based on deep learning

By constructing an intelligent analysis network that includes a feature extractor, a physical parameter probability prediction head, and a physical signal generator, the problems of low accuracy and limited efficiency in liquid chromatography analysis are solved. This enables accurate analysis and reliability assessment of complex liquid chromatography signals, improving both analysis accuracy and efficiency.

CN121208237BActive Publication Date: 2026-04-24RELAIS (HANGZHOU) MEDICAL TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RELAIS (HANGZHOU) MEDICAL TECH CO LTD
Filing Date
2025-11-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing liquid chromatography analysis methods suffer from low resolution and limited efficiency when dealing with complex samples, especially when faced with severe overlapping peaks, baseline drift, and noise interference. They rely on manually setting parameters and thresholds, resulting in low resolution and limited efficiency.

Method used

A deep learning-based intelligent peak analysis method for liquid chromatography is adopted, and an intelligent peak analysis network is constructed, including a feature extractor, a physical parameter probability prediction output head, and a physical signal generator. By dynamically adjusting the sample chromatogram and its composition, the network is continuously updated, and a chromatographic peak heatmap is generated to improve the analysis accuracy and efficiency.

Benefits of technology

It achieves accurate analysis and reliability assessment of complex liquid chromatography signals, and improves the accuracy and efficiency of analysis by outputting results through probability distribution, thereby enhancing the interpretability and credibility of the analysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121208237B_ABST
    Figure CN121208237B_ABST
Patent Text Reader

Abstract

The application discloses a liquid chromatography peak intelligent analysis method and system based on deep learning, and relates to the technical field of data processing.The method comprises the following steps: constructing a chromatography peak intelligent analysis network based on deep learning, training the network using sample chromatograms until convergence; inputting a chromatography signal to be analyzed into the chromatography peak intelligent analysis network, and evaluating the reliability of the prediction result based on the probability distribution of the key physical parameters output by the network to obtain a result confidence; dynamically adjusting the number and composition of the sample chromatograms based on the result confidence and a pre-set confidence threshold, and continuously updating the chromatography peak intelligent analysis network; recording the weight score assigned to each data point by the attention mechanism of the chromatography peak intelligent analysis network, mapping the weight score back to the time sequence of the original chromatography signal, and visualizing the weight score through a color gradient, which is superimposed on the original chromatography curve in the form of a semi-transparent layer to generate a chromatography peak heat map.The application effectively improves the efficiency and accuracy of liquid chromatography analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method and system for intelligent peak analysis in liquid chromatography based on deep learning. Background Technology

[0002] Chromatographic analysis, as a core method of separation and detection, is widely used in chemical, environmental and pharmaceutical fields. It is necessary to analyze key physical parameters such as peak position and peak area from chromatographic signals to achieve quantitative and qualitative analysis of components. However, as the complexity of samples increases, the demand for analysis efficiency and accuracy increases significantly.

[0003] In liquid chromatography analysis, traditional peak resolution methods have the advantages of high computational efficiency and clear principles when dealing with simple samples with good separation. However, when faced with severe overlapping peaks, baseline drift, and noise interference commonly found in complex samples, they usually rely on manually setting parameters and thresholds, resulting in low resolution accuracy and limited efficiency. Summary of the Invention

[0004] This application provides a method and system for intelligent peak analysis in liquid chromatography based on deep learning, aiming to solve the technical problems of low accuracy and limited efficiency in liquid chromatography analysis in the prior art.

[0005] In view of the above problems, this application provides a method and system for intelligent peak analysis in liquid chromatography based on deep learning.

[0006] Firstly, this application provides a deep learning-based intelligent peak analysis method for liquid chromatography, including:

[0007] Based on deep learning, a chromatographic peak intelligent analysis network is constructed, and the chromatographic peak intelligent analysis network is trained using sample chromatograms until convergence. The chromatographic peak intelligent analysis network includes a feature extractor, a physical parameter probability prediction output head, and a physical signal generator.

[0008] The chromatographic signal to be analyzed is input into the intelligent chromatographic peak analysis network, and the reliability of the prediction results is evaluated based on the probability distribution of the key physical parameters output by the intelligent chromatographic peak analysis network, so as to obtain the confidence level of the results.

[0009] Based on the confidence level of the results and the preset confidence threshold, the number and composition of sample chromatograms are dynamically adjusted, and the intelligent peak analysis network is continuously updated.

[0010] The attention mechanism of the chromatographic peak intelligent analysis network records the weight score assigned to each data point, maps the weight score back to the time series of the original chromatographic signal, and visualizes it through color gradient. The weight score is then superimposed on the original chromatographic curve in the form of a semi-transparent layer to generate a chromatographic peak heatmap.

[0011] Secondly, this application provides a deep learning-based intelligent peak analysis system for liquid chromatography, comprising:

[0012] The network construction and training module is used to construct a chromatographic peak intelligent analysis network based on deep learning, and to train the chromatographic peak intelligent analysis network using sample chromatograms until convergence. The chromatographic peak intelligent analysis network includes a feature extractor, a physical parameter probability prediction output head, and a physical signal generator.

[0013] The result confidence assessment module is used to input the chromatographic signal to be analyzed into the intelligent chromatographic peak analysis network, and to assess the reliability of the prediction results based on the probability distribution of the key physical parameters output by the intelligent chromatographic peak analysis network, thereby obtaining the result confidence level.

[0014] The dynamic adaptive update module is used to dynamically adjust the number and composition of sample chromatograms based on the confidence level of the results and a preset confidence threshold, and to continuously update the intelligent analysis network of chromatographic peaks.

[0015] The heatmap visualization module is used to record the weight score assigned to each data point by the attention mechanism of the chromatographic peak intelligent analysis network. The weight score is mapped back to the time series of the original chromatographic signal and visualized through color gradient. It is superimposed on the original chromatographic curve in the form of a semi-transparent layer to generate a chromatographic peak heatmap.

[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0017] This application provides a deep learning-based intelligent peak analysis method and system for liquid chromatography, achieving accurate analysis and reliability assessment of complex liquid chromatography signals. By constructing an intelligent analysis network comprising a feature extractor, a physical parameter probability prediction head, and a physical signal generator, key physical parameters can be robustly extracted from overlapping peaks and noise interference, and the results are output in the form of a probability distribution. A dynamic sample optimization mechanism based on prediction confidence continuously improves network performance and adaptability. The heatmap visualization technology of attention weights provides an intuitive basis for decision-making in the analysis process, enhancing interpretability and credibility, and effectively improving the accuracy and efficiency of the liquid chromatography analysis method. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1A flowchart illustrating the intelligent peak analysis method for liquid chromatography based on deep learning provided in this application embodiment;

[0020] Figure 2 A schematic diagram of the structure of the intelligent peak analysis system for liquid chromatography based on deep learning provided in the embodiments of this application;

[0021] Figure 3 A chromatogram containing one symmetrical peak is provided for an embodiment of this application;

[0022] Figure 4 The chromatographic peak thermogram provided for the embodiments of this application;

[0023] The components represented by each number in the attached diagram are explained below:

[0024] The network construction and training module 11, the result confidence evaluation module 12, the dynamic adaptive update module 13, and the heatmap visualization module 14 are all included. Detailed Implementation

[0025] This application provides a deep learning-based intelligent peak analysis method and system for liquid chromatography, which addresses the technical problems of low accuracy and limited efficiency in existing liquid chromatography analysis techniques.

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0027] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.

[0028] Example 1, as Figure 1 As shown, this application provides a deep learning-based intelligent peak analysis method for liquid chromatography, the method comprising:

[0029] S100: Based on deep learning, a chromatographic peak intelligent analysis network is constructed, and the chromatographic peak intelligent analysis network is trained using sample chromatograms until convergence. The chromatographic peak intelligent analysis network includes a feature extractor, a physical parameter probability prediction output head, and a physical signal generator.

[0030] In this embodiment, a chromatographic peak intelligent analysis network is constructed based on deep learning, and the network is trained using sample chromatograms until convergence. This network includes a feature extractor, a physical parameter probability prediction output head, and a physical signal generator. A one-dimensional CNN automatically extracts complex chromatographic features, avoiding the limitations of manually designed features; the probability prediction branch quantifies parameter uncertainty, improving analysis reliability; and the differentiable generator constrains predictions through a physical model, ensuring the output conforms to the physical laws of chromatographic peaks. Labeled samples are used to simultaneously optimize parameter prediction accuracy and signal reconstruction effect through composite loss, enabling the network to learn the correlation between peak shape features and physical parameters from the data, achieving end-to-end optimization through backpropagation.

[0031] Step S100 in the method provided in this application embodiment includes:

[0032] The method provided in this application embodiment constructs a chromatographic peak intelligent analysis network based on deep learning. This network includes at least a feature extractor, a physical parameter probability prediction output head, and a physical signal generator.

[0033] The feature extractor is constructed based on a one-dimensional CNN. When a chromatogram is input, the feature tensor of the chromatographic peak can be identified and output. The feature tensor includes at least the rising edge, falling edge, apex, and baseline drift trend of the chromatographic peak.

[0034] Following the feature extractor, multiple independent probability prediction branches are connected in parallel, wherein each probability prediction branch corresponds to the parameter prediction result of a potential chromatographic peak, the parameter prediction result including the probability distribution of key physical parameters, the probability distribution including the prediction mean and prediction variance;

[0035] The physical signal generator is constructed based on a differentiable function.

[0036] First, a feature extractor is constructed based on a one-dimensional CNN. Inputting a chromatogram, it identifies and outputs a feature tensor of chromatographic peaks. This feature tensor includes at least the rising edge, falling edge, vertex, and baseline drift trend of the chromatographic peak. A one-dimensional CNN is a convolutional neural network suitable for processing one-dimensional signals such as time series, extracting features from local continuous data through sliding convolutional kernels. The feature extractor is constructed using a one-dimensional convolutional neural network, with a network structure containing multiple convolutional layers, batch normalization layers, and pooling layers. The input is the original chromatogram, i.e., a one-dimensional sequence of time-response intensity, such as response values ​​at 1000 time points. Local features are extracted through convolutional layers, such as the rate of change of response between adjacent time points. Pooling layers reduce dimensionality and retain key features, ultimately outputting a feature tensor. This feature tensor includes: rising edge: the increasing trend of the response over time, such as slope changes; falling edge: the decreasing trend of the response over time; vertex: the time point and intensity corresponding to the maximum response; baseline drift: the slow changing trend of the background response over time. For example, ... Figure 3 As shown, the input chromatogram contains one symmetrical peak: time 0-100s, peak at 50s, rising edge response from 100 to 5000 from 0-50s, falling edge response from 50-100s to 100, baseline drift 50; feature extractor output feature tensor: rising edge: 4900 / 50=98; falling edge: 98; peak: 50s, 5000; baseline drift: 0.5.

[0037] Secondly, following the feature extractor, multiple independent probability prediction branches are connected in parallel. Each probability prediction branch corresponds to the parameter prediction result of a potential chromatographic peak. The parameter prediction result includes the probability distribution of key physical parameters, which includes the prediction mean and prediction variance. Key physical parameters include: peak height, retention time, peak width, and asymmetry factor. At the output of the feature extractor, N independent probability prediction branches are set in parallel, where N is the preset maximum number of potential chromatographic peaks, such as 5. Each branch consists of a fully connected layer and a probability output layer. After the feature tensor is mapped by the fully connected layer, the probability output layer is parameterized through a normal distribution, outputting the probability distribution of the key physical parameters of each potential chromatographic peak, including the probability distribution of peak height, retention time, peak width, and asymmetry factor, describing the statistical distribution of possible parameter values. The mean μ represents the most likely value, and the variance... This represents uncertainty. Each branch corresponds to a potential peak; peaks that do not appear are filtered by variance approaching 0 or a probability threshold. For example, for a chromatogram with two overlapping peaks, three probability prediction branches are set. Branch 1 outputs the retention time. Peak height Half-peak width asymmetric factor Branch 2 output retention time Peak height Half-peak width asymmetric factor Branch 3 outputs have a maximum variance, such as , indicating no corresponding peak.

[0038] Finally, the physical signal generator is constructed based on a differentiable function. The physical signal generator is constructed using a differentiable function, with the functional form based on the chromatographic peak physical model. Taking the Gaussian peak model as an example, the signal function of a single peak is: Where H is the peak height, T is the retention time, W is the half-peak width correlation parameter, and B(t) is the baseline function, such as a linear function. The generator takes as input the physical parameters of each branch sample: H, T, W, a, b, etc., and calculates and outputs the predicted chromatographic signal, i.e., the time series response value, through function superposition. Because the function is differentiable, it supports collaborative backpropagation with other parts of the network. For example, input: Branch 1 sampling parameters: H1=3000, T1=10.2s, W1=0.5s, baseline parameters: a=0.1, b=50; Branch 2 sampling parameters: H2=2500, T2=10.8s, W2=0.4s, baseline parameters: a=0.1, b=50. The generator calculates and superimposes two Gaussian peaks and the baseline, outputting a predicted response sequence from 0 to 20 seconds.

[0039] The method provided in this application embodiment involves obtaining a chromatogram labeled with chromatographic peak information as a sample chromatogram, and training the intelligent chromatographic peak analysis network using multiple sample chromatograms until convergence, including:

[0040] Obtain the chromatogram labeled with chromatographic peak information as the sample chromatogram;

[0041] The preprocessed sample chromatogram is input into the feature extractor for feature tensor extraction.

[0042] The extracted feature tensor is input into the physical parameter probability prediction output head to obtain the probability distribution of key physical parameters;

[0043] Based on this probability distribution, sampling is performed, and the sampled values ​​are input into the physical signal generator to generate a predictive chromatographic signal;

[0044] The composite loss between the predicted chromatographic signal and the actual chromatographic signal is calculated, and the network weights are updated through backpropagation based on the composite loss until the intelligent chromatographic peak analysis network converges.

[0045] First, obtain chromatograms labeled with chromatographic peak information, serving as sample chromatograms. Collect chromatograms from actual detection scenarios, such as gas chromatography signals detecting mixed solvents. Record the key physical parameters and actual chromatographic signals of each chromatographic peak through manual annotation or high-precision instrument-assisted annotation, forming a sample set. The samples should cover scenarios with simple peaks and complex peaks to improve network generalization. For example, the sample set contains 1000 chromatograms, of which 300 are single-peak, 200 are double-peak, 300 are multi-peak, and 200 are complex peaks with strong noise. Each chromatogram is labeled with parameters such as retention time (e.g., 10.2s ± 0.05s) and peak height (e.g., 3000 ± 50).

[0046] Next, the preprocessed sample chromatogram is input into the feature extractor for feature tensor extraction. Preprocessing of the sample chromatogram includes techniques such as wavelet thresholding to remove high-frequency noise and polynomial fitting to correct baseline drift. The preprocessed signal is input into the feature extractor, which performs operations such as convolution and pooling to output a feature tensor. For example, for noisy sample chromatograms, such as those with random noise superimposed on the original signal and a response fluctuation of ±200, wavelet thresholding reduces the noise fluctuation to ±50. For samples with baseline drift, such as those where the baseline rises from 100 to 300 within 0-100 seconds, a third-order polynomial is used to fit the baseline and subtract the corrected signal. This corrected signal is then input into the feature extractor, which outputs a tensor containing the denoised peak shape features.

[0047] Furthermore, the extracted feature tensor is input into the physical parameter probability prediction output head to obtain the probability distribution of key physical parameters. The feature tensor output by the feature extractor is simultaneously input into each probability prediction branch. After mapping through a fully connected layer, each branch outputs the probability distribution of the physical parameters corresponding to the potential peak. For example, for the retention time parameter, the branch output... This indicates that the retention time of the peak is most likely 10.2 s, with relatively low uncertainty.

[0048] Then, sampling is performed based on this probability distribution, and the sampled values ​​are input into the physical signal generator to generate a predicted chromatographic signal. Random sampling is performed from the probability distributions of each branch to obtain specific values ​​of the physical parameters, such as sampling 10.18 s from the retention time distribution N(10.2, 0.01). The sampled parameters are input into the physical signal generator, which calculates and superimposes each peak signal and baseline using a differentiable function to generate a predicted chromatographic signal, i.e., a time series of the same length as the sample chromatogram. For example, 2980 s is sampled from the peak height distribution N(3000, 100) of branch 1, and 10.19 s is sampled from the retention time distribution N(10.2, 0.01). After being input into the generator, these samples, together with the sampling parameters from branch 2, generate a predicted signal sequence containing two overlapping peaks.

[0049] Finally, the composite loss between the predicted and actual chromatographic signals is calculated, and the network weights are updated via backpropagation based on this composite loss until the intelligent peak analysis network converges. The composite loss between the predicted and actual chromatographic signals is calculated, including reconstruction loss and probability loss. Reconstruction loss, such as mean squared error (MSE), measures the signal fit; probability loss, such as KL divergence, measures the deviation between the predicted distribution and the labeled parameters. Backpropagation refers to updating the network weights along the direction of loss reduction using a gradient descent algorithm. The composite loss is optimized through backpropagation, and the iterations are repeated until the loss value stabilizes, indicating network convergence. For example, if the MSE between the actual and predicted signals for a sample is 5000, the KL divergence is 0.1, and the composite loss is set to 0.8 × MSE + 0.2 × KL = 4000.02, backpropagation is used to adjust the convolutional kernel weights of the feature extractor and the weights of the fully connected layers in the prediction branch, reducing the loss in the next iteration to below 4000.02. After 800 iterations, the loss stabilizes at around 100, indicating that the intelligent peak analysis network has converged.

[0050] In this embodiment, by constructing a chromatographic peak intelligent analysis network, simple and complex chromatographic peaks can be automatically analyzed, and the output probability distribution can quantify the analysis uncertainty; by constraining the physical generator, the physical rationality of parameter prediction is improved; after sufficient training, the analysis accuracy and efficiency are significantly improved, and it has the ability to generalize to multi-sample scenarios.

[0051] S200: Input the chromatographic signal to be analyzed into the intelligent chromatographic peak analysis network, and evaluate the reliability of the prediction results based on the probability distribution of the key physical parameters output by the intelligent chromatographic peak analysis network to obtain the confidence level of the results.

[0052] In this embodiment, the chromatographic signal to be analyzed is input into a chromatographic peak intelligent analysis network. Based on the probability distribution of key physical parameters output by the network, the reliability of the prediction results is evaluated, and the confidence level is obtained. This achieves a quantitative assessment of the reliability of the analysis results, with the confidence level numerically reflecting the degree of credibility. The peak area confidence interval provides an error range reference for quantitative analysis. By dynamically combining parameter uncertainty and signal matching degree, the limitations of a single indicator are avoided, ultimately improving the credibility and application value of the chromatographic peak analysis results.

[0053] Step S200 in the method provided in this application embodiment includes:

[0054] Extract the prediction variance of each physical parameter output by the intelligent analysis network for chromatographic peaks, and calculate the normalized weighted average of the prediction variance of each physical parameter.

[0055] Based on the error propagation principle, the confidence interval of the peak area is derived from the normalized weighted average of the prediction variances of each physical parameter, and the interval width is obtained.

[0056] The signal matching degree is obtained by comparing the degree of matching between the predicted chromatographic signal reconstructed by the intelligent peak analysis network and the chromatographic signal to be analyzed.

[0057] The confidence level of the obtained result is evaluated based on the interval width and the signal matching degree.

[0058] First, the predicted variances of each physical parameter output by the intelligent peak analysis network are extracted, and the normalized weighted average of the predicted variances of each physical parameter is calculated. Each variance is then divided by the maximum value of all parameter variances to obtain a normalized value between 0 and 1. This is to eliminate differences in the units of different parameters, such as the unit of retention time variance. The peak height variance is expressed in response values. 2 Direct comparison is not possible. Weights are assigned to each parameter based on their importance to peak resolution. For example, peak width and retention time have a greater impact on peak identification, so their weight is set to 0.3. Peak height and asymmetry factor are weighted at 0.2. The weighted sum of the normalized variances is then calculated, resulting in the normalized weighted average. For instance, the predicted variance of a certain chromatographic peak to be resolved is: peak height variance... Response value 2 Retention time variance Peak width variance Asymmetric factor variance The maximum variance is 2500, after normalization. Weighted calculation: The normalized weighted average is mainly contributed by the normalized variance of the peak height.

[0059] Secondly, based on the error propagation principle, the confidence interval of the peak area is derived from the normalized weighted average of the prediction variances of each physical parameter, thus obtaining the interval width.

[0060] The method provided in this application embodiment, based on the error propagation principle, derives the confidence interval of the peak area from the normalized weighted average of the prediction variances of each physical parameter, and obtains the interval width, including:

[0061] Obtain the probability distribution of the prediction of each physical parameter by the intelligent analysis network of chromatographic peaks, and use the prediction variance of the probability distribution as a quantitative indicator of the prediction uncertainty of that parameter.

[0062] Based on the prediction variance of the probability distribution, the optimal parameter estimate of the physical parameter is obtained, and the sensitivity of the peak area to the change of each physical parameter near the optimal parameter estimate is calculated as the sensitivity coefficient of the physical parameter.

[0063] Multiply the prediction variance of each physical parameter by the square of its corresponding sensitivity coefficient to obtain multiple uncertainty contribution values;

[0064] The summation of multiple uncertainty contributions and the square root are used as the standard uncertainty of the peak area.

[0065] Configure a multiplier factor, and determine the interval width based on the multiplier factor and the standard uncertainty.

[0066] First, the probability distribution of the chromatographic peak intelligent analysis network's predictions for each physical parameter is obtained. The prediction variance of the probability distribution is used as a quantitative indicator of the prediction uncertainty of that parameter. The prediction variance of each parameter is extracted from the probability distribution output by the network and directly used as a quantitative indicator of the parameter's prediction uncertainty. The larger the variance, the higher the probability that the parameter's true value deviates from the predicted value. Continuing the previous example, the variance of each parameter is: peak height... Retention time Peak width Asymmetric factors .

[0067] Secondly, based on the prediction variance of the probability distribution, the optimal parameter estimates of the physical parameters are obtained. The sensitivity of the peak area to changes in each physical parameter near the optimal parameter estimate is then calculated, serving as the sensitivity coefficient of the physical parameter. Peak area is a core quantitative indicator in chromatographic analysis and is a function of key physical parameters, denoted as [equation missing]. Where H is the peak height, T is the retention time, W is the peak width, and S is the asymmetry factor. The sensitivity coefficient is defined as the partial derivative of the peak area with respect to each parameter, i.e. The sensitivity coefficient is used to measure the impact of small changes in parameters on the peak area. The larger the absolute value of the sensitivity coefficient, the more significant the impact of parameter changes on the peak area. For example, assuming the peak area function is... k is a constant related to the asymmetry factor S, calculated from S=1.2, and the optimal parameter estimate is... Sensitivity coefficient at peak height: Sensitivity coefficient for peak width: Sensitivity coefficient for retention time: Sensitivity coefficient of the asymmetry factor: .

[0068] Furthermore, the prediction variance of each physical parameter is multiplied by the square of its corresponding sensitivity coefficient to obtain multiple uncertainty contribution values. The contribution of each parameter to the peak area uncertainty is the prediction variance of that parameter multiplied by the square of its sensitivity coefficient, i.e. This is used to quantify the impact of the uncertainty of a single parameter on the total uncertainty of the peak area. For example, the contribution of peak height: Peak width contribution value: Asymmetric factor contribution value: Retention time contribution value: .

[0069] Then, the summation of multiple uncertainty contributions and the square root are taken as the standard uncertainty of the peak area. It is the square root of the sum of the uncertainties of each parameter, i.e. This is used to quantify the overall dispersion of peak area predictions, when the peak area follows a normal distribution. This is the standard deviation. For example, .

[0070] Finally, a multiplier factor is configured, and the interval width is determined based on the multiplier factor and the standard uncertainty.

[0071] The method provided in this application embodiment, which configures a multiplier factor and determines the interval width based on the multiplier factor and the standard uncertainty, includes:

[0072] Based on the risk profile of the application scenario, a pre-set confidence level is established, wherein the confidence level is proportional to the risk level.

[0073] Based on the confidence level, the corresponding Z value is retrieved from the standard normal distribution table and used as the multiple factor;

[0074] Based on the multiplication factor, the optimal parameter estimate, and the standard uncertainty, the confidence interval of the peak area is obtained;

[0075] The width of the interval is obtained based on the confidence interval of the peak area.

[0076] First, a confidence level is preset based on the risk profile of the application scenario, where the confidence level is directly proportional to the risk level. The confidence level refers to the probability that the true peak area falls within a confidence interval; it is set according to the risk level of the application scenario, with higher risks requiring higher confidence levels. For example, environmental pollutant detection requires strict control over the risk of false positives, so a preset confidence level of 95% is used.

[0077] Secondly, based on the confidence level, the corresponding Z-value is retrieved from the standard normal distribution table and used as the multiple factor. The Z-value is a standard normal distribution. The quantiles corresponding to the confidence level, such as Z=1.96 for a 95% confidence level, indicate that 95% of the probability density in a normal distribution is concentrated in the interval [-1.96, 1.96]. These quantiles are used as a multiplier to expand the standard uncertainty. For example, the Z value corresponding to a 95% confidence level is 1.96, hence the multiplier... .

[0078] Then, based on the multiplier factor, the optimal parameter estimate, and the standard uncertainty, the confidence interval for the peak area is obtained. The confidence interval is: optimal parameter estimate ± (multiplier factor × standard uncertainty). The optimal estimate A of the peak area is calculated from the optimal parameters, and its confidence interval is... This indicates that the true peak area has a specified probability of falling within this interval. For example, the optimal estimate of the peak area... The confidence interval is [5000-1.96×255.15, 5000+1.96×255.15]≈[5000-500, 5000+500]=[4500, 5500].

[0079] Finally, the interval width is obtained based on the confidence interval of the peak area. The interval width is the difference between the upper and lower limits of the confidence interval; the smaller the width, the lower the uncertainty in peak area prediction. For example, the interval width above = 5500 - 4500 = 1000.

[0080] Furthermore, the matching degree between the predicted chromatographic signal reconstructed by the intelligent peak analysis network and the chromatographic signal to be analyzed is compared to obtain the signal matching degree. The signal matching degree is used to measure the overall similarity between the predicted chromatographic signal reconstructed by the network and the chromatographic signal to be analyzed. It is obtained by calculating the similarity index between the two, with a value ranging from 0 to 1. The closer to 1, the higher the matching degree. For example, if the normalized mean square error between the signal to be analyzed and the predicted signal is 0.05, then the signal matching degree = 1 - 0.05 = 0.95, which is a relatively high matching degree.

[0081] Finally, based on the interval width and the signal matching degree, the confidence level of the result is evaluated. The result confidence level is a comprehensive index of the interval width and the signal matching degree, calculated using a weighted formula: Confidence Level = α × Signal Matching Degree + (1 - α) × (1 - Interval Width / Maximum Possible Width), where α is the weight, such as 0.6, ranging from 0 to 1. The closer to 1, the higher the reliability. For example, if the interval width is 1000 and the maximum possible width is 2000, then the interval width score = 1 - 1000 / 2000 = 0.5; the signal matching degree is 0.95, α = 0.6, and the confidence level = 0.6 × 0.95 + 0.4 × 0.5 = 0.57 + 0.2 = 0.77, indicating a moderately high reliability.

[0082] In this embodiment, the uncertainty of each parameter is calculated by normalizing the weighted average variance to eliminate dimensional differences and provide a basis for subsequent peak area uncertainty analysis. Based on the error propagation principle, the peak area confidence interval is derived: the uncertainty of key parameters is quantitatively propagated to the core quantitative indicators, and its reliability is quantified by the confidence interval. The network reconstruction effect is verified from the overall signal level to supplement the limitations of parameter-level analysis. Combining uncertainty and signal consistency, the reliability of the analysis results is comprehensively quantified.

[0083] S300: Based on the confidence level of the results and the preset confidence level threshold, dynamically adjust the number and composition of sample chromatograms and continuously update the intelligent peak analysis network.

[0084] In this embodiment, based on the result confidence level and a preset confidence threshold, the number and composition of sample chromatograms are dynamically adjusted to continuously update the intelligent peak analysis network. Weak chromatograms are screened using the confidence threshold to accurately pinpoint the model's analytical shortcomings; common features are extracted to identify the model's knowledge blind spots, and targeted synthetic samples are generated to supplement training data, addressing the problem of scarce data in weak scenarios in real samples; the training endpoint is controlled based on the average confidence level of the validation set, ensuring a significant improvement in the generalization ability of the updated intelligent peak analysis network.

[0085] Step S300 in the method provided in this application embodiment includes:

[0086] Preset reliability threshold;

[0087] Chromatograms that do not meet the confidence threshold are identified as weak chromatograms;

[0088] Identify common features of weak chromatograms, including specific types of severe peak overlap and extreme baseline drift;

[0089] Based on the common features of weak chromatograms, a physical model is used to generate synthetic chromatograms with similar features as targeted samples;

[0090] The newly generated targeted samples are added to the sample chromatogram, and the intelligent chromatographic peak analysis network is retrained until the average confidence level of the model on the validation set exceeds the preset high confidence threshold. The validation set is labeled with accurate chromatographic peak information.

[0091] First, a reliability threshold is preset. The confidence threshold is the critical value used to determine whether the chromatographic peak analysis results meet reliability requirements. It is preset based on the accuracy requirements of the application scenario. For example, the threshold is higher in pharmaceutical testing scenarios with high reliability requirements, while it can be slightly lower in general chemical analysis scenarios. The threshold range is consistent with the result confidence level, being 0-1, and needs to be determined by considering historical analysis errors and business needs. For example, in the pharmaceutical industry, to ensure the accuracy of component quantification, a preset reliability threshold of 0.8 is used, meaning a result confidence level ≥ 0.8 is considered reliable; in general environmental monitoring scenarios, a preset reliability threshold of 0.7 is used.

[0092] Furthermore, chromatograms that do not meet the confidence threshold are identified as weak chromatograms. The confidence scores of all chromatographic signals to be analyzed output in step S200 are iterated, and chromatograms with confidence scores below a preset threshold are selected and defined as weak chromatograms. Weak chromatograms indicate insufficient reliability of the analysis results, reflecting that the model has not sufficiently learned its features. For example, if 80 out of 500 chromatograms to be analyzed have confidence scores below the preset confidence threshold of 0.8, these 80 chromatograms are marked as weak chromatograms.

[0093] Secondly, common features of weak chromatograms are identified, including specific types of severe peak overlap and extreme baseline drift. Feature analysis is performed on weak chromatograms to extract common signal features that lead to resolution difficulties. Common features include: specific types of severe peak overlap (e.g., retention time interval of 3 or more peaks < 0.3 s, overlap > 80%, peak height difference < 20%); extreme baseline drift (e.g., baseline drift within 100 s > 50% of the main peak height, or exhibiting non-linear drift); strong noise interference (e.g., signal-to-noise ratio < 5:1, noise exhibiting periodic fluctuations). Statistical analysis, such as calculating the frequency of feature occurrence, determines the dominant common features. For example, analysis of 80 weak chromatograms revealed that 65 exhibited severe overlap: 3 peaks with retention time intervals of 0.2-0.3 s and overlap of 85%-90%; 15 exhibited extreme baseline drift: non-linear baseline drift within 100 s, increasing from a response value of 100 to 2000.

[0094] Furthermore, based on the common features of weak chromatograms, physical models are used to generate synthetic chromatograms with similar characteristics, serving as targeted samples. Based on the identified common features, physical models of chromatographic peaks, such as Gaussian peak models, Lorentz peak models, or mixed models, can precisely control parameters such as peak position, peak shape, and baseline to generate synthetic chromatograms with similar features. During generation, it is necessary to accurately reproduce the weak features, such as controlling the retention time difference of overlapping peaks and the magnitude and trend of baseline drift. Simultaneously, accurate physical parameters are labeled as tags to form targeted samples to supplement the weak scenario data lacking in the training set. For example, for three severely overlapping peak features, a Gaussian peak model is used to generate 50 synthetic chromatograms: retention times are set to 10.0s, 10.2s, and 10.4s, with an interval of 0.2s; peak heights are all 5000±500 response values, with differences <20%; half-peak width is 0.3s; and overlap is 88%. For extremely nonlinear baseline drift features, 30 synthetic chromatograms are generated: the baseline is set to exponential drift. t represents time, one main peak is superimposed, retention time is 50s, peak height is 3000, and the physical parameters of all peaks are labeled.

[0095] Finally, the newly generated targeted samples are added to the sample chromatograms, and the intelligent peak analysis network is retrained until the model's average confidence score on the validation set exceeds a preset high confidence threshold. The validation set contains accurate peak information. The newly generated targeted samples are then added to the original sample chromatogram set to expand the training set, ensuring a reasonable proportion of new samples (e.g., 10%-20%) to avoid data imbalance. The validation set, independent of the training set, is used to objectively evaluate the model's generalization ability. It must cover diverse scenarios, including common and complex peaks, and each sample must be accurately labeled with peak information: peak start and end points: the time from which the peak rises from the baseline and the time from which it falls back to the baseline; vertex: the time point and response intensity value corresponding to the maximum peak response intensity; peak area: the area enclosed by the peak and the baseline. The expanded training set is then used to retrain the intelligent peak analysis network, following the original training process, such as calculating the composite loss and updating weights via backpropagation. After each training round, the model's average confidence score is evaluated using the validation set. Training stops when this score exceeds a preset high confidence threshold, and the network is updated. For example, the preset high confidence threshold is 0.9; the original training set contains 1000 samples, which are expanded to 1080 after adding 80 targeted samples; the validation set contains 200 samples with an initial average confidence score of 0.75. After 50 retraining rounds, when the validation set's average confidence score rises to 0.92, exceeding the high confidence threshold of 0.9, training stops, and the intelligent chromatographic peak analysis network update is complete.

[0096] In this embodiment, by selectively supplementing weak scenario samples, the intelligent chromatographic peak analysis network is able to specifically enhance its learning of complex features, avoiding the inefficiency caused by blindly expanding data. The supervision of the validation set ensures that the updated model not only performs well on the training data, but also can stably adapt to complex signals in real-world scenarios, ultimately continuously improving the reliability and generalization ability of chromatographic peak analysis, and realizing the dynamic iterative optimization of the intelligent chromatographic peak analysis network.

[0097] S400: The attention mechanism of the recording chromatographic peak intelligent analysis network assigns a weight score to each data point, maps the weight score back to the time series of the original chromatographic signal, and visualizes it through color gradient. It is then superimposed on the original chromatographic curve in the form of a semi-transparent layer to generate a chromatographic peak heatmap.

[0098] In this embodiment, firstly, the attention mechanism of the intelligent chromatographic peak analysis network is used to assign weight scores to each data point. The attention mechanism is a deep learning module that simulates human attention allocation. It assigns weight scores to each input data point by calculating its contribution to the model's decision-making; a higher weight indicates a greater impact on peak identification and parameter prediction. During network inference, the weight scores output by the attention mechanism are recorded in real time. For an input one-dimensional chromatographic signal [time series containing N data points, each corresponding to a response intensity at a specific time], the attention mechanism generates a weight score sequence of length N, ranging from 0 to 1, or mapped to this range after normalization. For example, when inputting a chromatographic signal containing a single peak [time 0-100s, a total of 1000 data points, with the peak located at 40-60s], the attention mechanism assigns a high weight of 0.7-0.9 to the data points at 40-60s and a low weight of 0.1-0.3 to the baseline regions at 0-30s and 70-100s. The recorded weight score sequence corresponds one-to-one with the 1000 time points.

[0099] Secondly, the weight scores are mapped back to the time series of the original chromatographic signal. The original chromatographic signal has time as the horizontal axis, with each time point corresponding to a unique index. Although the weight score sequence output by the attention mechanism may undergo dimensionality adjustments within the network, through a preset index mapping relationship, each weight score can be precisely associated with a time point of the original signal, ensuring that the weight scores are completely aligned with the original time series in the time dimension. For example, during internal network processing, 1000 original data points are temporarily reduced to 500 feature points through pooling with a step size of 2. After the attention mechanism assigns weights to these 500 feature points, the weight score of each feature point is copied to the corresponding two original data points according to the mapping rule: original point index = feature point index × 2. For example, the weight of feature point 1 is assigned to original points 1 and 2, ultimately resulting in a weight sequence that is completely aligned with the original 1000 time points.

[0100] Finally, as Figure 4As shown, visualization is achieved through a color gradient, overlaid as a semi-transparent layer on the original chromatographic curve to generate a chromatographic peak heatmap. A color gradient scheme is used, such as from cool to warm tones, to visualize the weighted scores: low weights (0-0.3) correspond to blue; medium weights (0.3-0.7) correspond to yellow; and high weights (0.7-1.0) correspond to red. The color depth increases with the weight, forming a continuous gradient. The weighted scores must first be normalized to 0-1 to ensure they match the color gradient range. Above the time axis of the original chromatographic curve, a color heatmap layer of the weighted scores is plotted with the same horizontal axis range. The height of the heatmap layer can be set to 10%-20% of the vertical axis range of the original curve to avoid obscuring the curve, and the transparency should be set to 50%-70% to ensure that the thermal color distribution is clearly visible without obscuring the original peak shape. Integrating the original curve and the heatmap layer forms the final chromatographic peak heatmap, visually displaying the key time points of interest in the model. For example, a certain overlapping peak chromatographic signal [time 10-20s, containing two overlapping peaks, located at 12-14s and 13-15s respectively], is assigned a high weight of 0.8-0.9 to the overlapping region at 12-15s, displayed as red, with the deepest red near 13s, i.e., the core of the overlapping two peaks; the baseline or non-overlapping regions at 10-12s and 15-20s are assigned a low weight of 0.2-0.4, displayed as light blue. In the heatmap, the red band is concentrated at 12-15s, perfectly matching the overlapping peak region of the original curve, and the translucent layer clearly covers the curve below, showing both the peak shape and highlighting the model's focus.

[0101] In this embodiment, the interpretability of the chromatographic peak analysis process is achieved. The heatmap clearly shows the peak regions that the model focuses on, such as peak values, overlapping boundaries, and ignored baselines or noise regions, verifying the rationality of the analysis logic. At the same time, if the heatmap shows that the weights are abnormally concentrated in noise points or baseline drift regions, it can indicate that there is a learning bias in the intelligent chromatographic peak analysis network, providing intuitive clues for the optimization of the intelligent chromatographic peak analysis network, and ultimately enhancing the user's trust in the analysis results and the practicality of the intelligent chromatographic peak analysis network.

[0102] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects: Accurate extraction of key physical parameters of complex chromatographic peaks is achieved through deep learning networks, and the reliability of the analysis results is quantified by probability distribution, providing an objective basis for subsequent decision-making. Targeted samples are dynamically supplemented with confidence feedback, continuously optimizing the model's adaptability to complex scenarios such as severely overlapping peaks and extreme baseline drift, significantly improving generalization performance. Simultaneously, the model's decision dependence on key regions such as peak shoulders and inflection points is visually presented through heatmaps, breaking the black-box nature of the analysis process. This ensures both the accuracy and efficiency of chromatographic analysis and enhances the verifiability of the analysis logic, providing efficient and reliable technical support for qualitative and quantitative analysis of chromatographic detection.

[0103] Example 2, as Figure 2 As shown, this application provides a deep learning-based intelligent peak analysis system for liquid chromatography, the system comprising:

[0104] The network construction and training module 11 is used to construct a chromatographic peak intelligent analysis network based on deep learning, and to train the chromatographic peak intelligent analysis network using sample chromatograms until convergence. The chromatographic peak intelligent analysis network includes a feature extractor, a physical parameter probability prediction output head, and a physical signal generator.

[0105] The result confidence assessment module 12 is used to input the chromatographic signal to be analyzed into the intelligent chromatographic peak analysis network, and to assess the reliability of the prediction results based on the probability distribution of the key physical parameters output by the intelligent chromatographic peak analysis network, and obtain the result confidence.

[0106] The dynamic adaptive update module 13 is used to dynamically adjust the number and composition of sample chromatograms based on the confidence level of the results and the preset confidence level threshold, and continuously update the intelligent analysis network of chromatographic peaks.

[0107] The heatmap visualization module 14 is used to record the weight score assigned to each data point by the attention mechanism of the chromatographic peak intelligent analysis network, map the weight score back to the time series of the original chromatographic signal, and visualize it through color gradient. It is superimposed on the original chromatographic curve in the form of a semi-transparent layer to generate a chromatographic peak heatmap.

[0108] In one embodiment, the network construction and training module 11 is further configured to:

[0109] Specifically, a chromatographic peak intelligent analysis network is constructed based on deep learning. This network includes at least a feature extractor, a physical parameter probability prediction output head, and a physical signal generator, comprising:

[0110] The feature extractor is constructed based on a one-dimensional CNN. When a chromatogram is input, the feature tensor of the chromatographic peak can be identified and output. The feature tensor includes at least the rising edge, falling edge, apex, and baseline drift trend of the chromatographic peak.

[0111] Following the feature extractor, multiple independent probability prediction branches are connected in parallel, wherein each probability prediction branch corresponds to the parameter prediction result of a potential chromatographic peak, the parameter prediction result including the probability distribution of key physical parameters, the probability distribution including the prediction mean and prediction variance;

[0112] The physical signal generator is constructed based on a differentiable function.

[0113] The process includes obtaining chromatograms labeled with chromatographic peak information as sample chromatograms, and training the intelligent peak analysis network using multiple sample chromatograms until convergence, including:

[0114] Obtain the chromatogram labeled with chromatographic peak information as the sample chromatogram;

[0115] The preprocessed sample chromatogram is input into the feature extractor for feature tensor extraction.

[0116] The extracted feature tensor is input into the physical parameter probability prediction output head to obtain the probability distribution of key physical parameters;

[0117] Based on this probability distribution, sampling is performed, and the sampled values ​​are input into the physical signal generator to generate a predictive chromatographic signal;

[0118] The composite loss between the predicted chromatographic signal and the actual chromatographic signal is calculated, and the network weights are updated through backpropagation based on the composite loss until the intelligent chromatographic peak analysis network converges.

[0119] In one embodiment, the result confidence assessment module 12 is further configured to:

[0120] Extract the prediction variance of each physical parameter output by the intelligent analysis network for chromatographic peaks, and calculate the normalized weighted average of the prediction variance of each physical parameter.

[0121] Based on the error propagation principle, the confidence interval of the peak area is derived from the normalized weighted average of the prediction variances of each physical parameter, and the interval width is obtained.

[0122] The signal matching degree is obtained by comparing the degree of matching between the predicted chromatographic signal reconstructed by the intelligent peak analysis network and the chromatographic signal to be analyzed.

[0123] The confidence level of the obtained result is evaluated based on the interval width and the signal matching degree.

[0124] Based on the error propagation principle, the confidence interval of the peak area is derived from the normalized weighted average of the prediction variances of each physical parameter, thus obtaining the interval width, including:

[0125] Obtain the probability distribution of the prediction of each physical parameter by the intelligent analysis network of chromatographic peaks, and use the prediction variance of the probability distribution as a quantitative indicator of the prediction uncertainty of that parameter.

[0126] Based on the prediction variance of the probability distribution, the optimal parameter estimate of the physical parameter is obtained, and the sensitivity of the peak area to the change of each physical parameter near the optimal parameter estimate is calculated as the sensitivity coefficient of the physical parameter.

[0127] Multiply the prediction variance of each physical parameter by the square of its corresponding sensitivity coefficient to obtain multiple uncertainty contribution values;

[0128] The summation of multiple uncertainty contributions and the square root are used as the standard uncertainty of the peak area.

[0129] Configure a multiplier factor, and determine the interval width based on the multiplier factor and the standard uncertainty.

[0130] The configuration of the multiplier factor, and the determination of the interval width based on the multiplier factor and the standard uncertainty, includes:

[0131] Based on the risk profile of the application scenario, a pre-set confidence level is established, wherein the confidence level is proportional to the risk level.

[0132] Based on the confidence level, the corresponding Z value is retrieved from the standard normal distribution table and used as the multiple factor;

[0133] Based on the multiplication factor, the optimal parameter estimate, and the standard uncertainty, the confidence interval of the peak area is obtained;

[0134] The width of the interval is obtained based on the confidence interval of the peak area.

[0135] In one embodiment, the dynamic adaptive update module 13 is further configured to:

[0136] Preset reliability threshold;

[0137] Chromatograms that do not meet the confidence threshold are identified as weak chromatograms;

[0138] Identify common features of weak chromatograms, including specific types of severe peak overlap and extreme baseline drift;

[0139] Based on the common features of weak chromatograms, a physical model is used to generate synthetic chromatograms with similar features as targeted samples;

[0140] The newly generated targeted samples are added to the sample chromatogram, and the intelligent chromatographic peak analysis network is retrained until the average confidence level of the model on the validation set exceeds the preset high confidence threshold. The validation set is labeled with accurate chromatographic peak information.

[0141] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0142] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0143] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A deep learning-based intelligent peak analysis method for liquid chromatography, characterized in that, The method includes: Based on deep learning, a chromatographic peak intelligent analysis network is constructed, and the chromatographic peak intelligent analysis network is trained using sample chromatograms until convergence. The chromatographic peak intelligent analysis network includes a feature extractor, a physical parameter probability prediction output head, and a physical signal generator. The chromatographic signal to be analyzed is input into the intelligent chromatographic peak analysis network, and the reliability of the prediction results is evaluated based on the probability distribution of the key physical parameters output by the intelligent chromatographic peak analysis network, so as to obtain the confidence level of the results. Based on the confidence level of the results and the preset confidence threshold, the number and composition of sample chromatograms are dynamically adjusted, and the intelligent peak analysis network is continuously updated. The attention mechanism of the chromatographic peak intelligent analysis network records the weight score assigned to each data point, maps the weight score back to the time series of the original chromatographic signal, and visualizes it through color gradient. It is then superimposed on the original chromatographic curve in the form of a semi-transparent layer to generate a chromatographic peak heatmap. The chromatographic signal to be analyzed is input into the intelligent peak analysis network, and the reliability of the prediction results is evaluated based on the probability distribution of key physical parameters output by the network, thus obtaining the confidence level of the results, including: Extract the prediction variance of each physical parameter output by the intelligent analysis network for chromatographic peaks, and calculate the normalized weighted average of the prediction variance of each physical parameter. Based on the error propagation principle, the confidence interval of the peak area is derived from the normalized weighted average of the prediction variances of each physical parameter, and the interval width is obtained. The signal matching degree is obtained by comparing the degree of matching between the predicted chromatographic signal reconstructed by the intelligent peak analysis network and the chromatographic signal to be analyzed. Based on the interval width and the signal matching degree, the confidence level of the obtained result is evaluated; Based on the confidence level of the results and a preset confidence threshold, the number and composition of sample chromatograms are dynamically adjusted, and the intelligent peak analysis network is continuously updated, including: Preset reliability threshold; Chromatograms that do not meet the confidence threshold are identified as weak chromatograms; Identify common features of weak chromatograms, including specific types of severe peak overlap and extreme baseline drift; Based on the common features of weak chromatograms, a physical model is used to generate synthetic chromatograms with similar features as targeted samples; The newly generated targeted samples are added to the sample chromatogram, and the intelligent chromatographic peak analysis network is retrained until the average confidence level of the model on the validation set exceeds the preset high confidence threshold. The validation set is labeled with accurate chromatographic peak information.

2. The intelligent peak analysis method for liquid chromatography based on deep learning according to claim 1, characterized in that, Based on deep learning, a chromatographic peak intelligent analysis network is constructed. This network includes at least a feature extractor, a physical parameter probability prediction output head, and a physical signal generator, comprising: The feature extractor is constructed based on a one-dimensional CNN. When a chromatogram is input, the feature tensor of the chromatographic peak can be identified and output. The feature tensor includes at least the rising edge, falling edge, apex, and baseline drift trend of the chromatographic peak. Following the feature extractor, multiple independent probability prediction branches are connected in parallel, wherein each probability prediction branch corresponds to the parameter prediction result of a potential chromatographic peak, the parameter prediction result including the probability distribution of key physical parameters, the probability distribution including the prediction mean and prediction variance; The physical signal generator is constructed based on a differentiable function.

3. The intelligent peak analysis method for liquid chromatography based on deep learning according to claim 1, characterized in that, Obtain chromatograms labeled with chromatographic peak information as sample chromatograms, and train the intelligent peak analysis network using multiple sample chromatograms until convergence, including: Obtain the chromatogram labeled with chromatographic peak information as the sample chromatogram; The preprocessed sample chromatogram is input into the feature extractor for feature tensor extraction. The extracted feature tensor is input into the physical parameter probability prediction output head to obtain the probability distribution of key physical parameters; Based on this probability distribution, sampling is performed, and the sampled values ​​are input into the physical signal generator to generate a predictive chromatographic signal; The composite loss between the predicted chromatographic signal and the actual chromatographic signal is calculated, and the network weights are updated through backpropagation based on the composite loss until the intelligent chromatographic peak analysis network converges.

4. The intelligent peak analysis method for liquid chromatography based on deep learning according to claim 1, characterized in that, Based on the error propagation principle, the confidence interval for the peak area is derived from the normalized weighted average of the prediction variances of each physical parameter, thus obtaining the interval width, including: Obtain the probability distribution of the prediction of each physical parameter by the intelligent analysis network of chromatographic peaks, and use the prediction variance of the probability distribution as a quantitative indicator of the prediction uncertainty of that parameter. Based on the prediction variance of the probability distribution, the optimal parameter estimate of the physical parameter is obtained, and the sensitivity of the peak area to the change of each physical parameter near the optimal parameter estimate is calculated as the sensitivity coefficient of the physical parameter. Multiply the prediction variance of each physical parameter by the square of its corresponding sensitivity coefficient to obtain multiple uncertainty contribution values; The summation of multiple uncertainty contributions and the square root are used as the standard uncertainty of the peak area. Configure a multiplier factor, and determine the interval width based on the multiplier factor and the standard uncertainty.

5. The intelligent peak analysis method for liquid chromatography based on deep learning according to claim 4, characterized in that, Configure the multiplier factor, and determine the interval width based on the multiplier factor and the standard uncertainty, including: Based on the risk profile of the application scenario, a pre-set confidence level is established, wherein the confidence level is proportional to the risk level. Based on the confidence level, the corresponding Z value is retrieved from the standard normal distribution table and used as the multiple factor; Based on the multiplication factor, the optimal parameter estimate, and the standard uncertainty, the confidence interval of the peak area is obtained; The width of the interval is obtained based on the confidence interval of the peak area.

6. A deep learning-based intelligent peak analysis system for liquid chromatography, characterized in that, The system is used to implement the deep learning-based intelligent peak analysis method for liquid chromatography according to any one of claims 1-5, the system comprising: The network construction and training module is used to construct a chromatographic peak intelligent analysis network based on deep learning, and to train the chromatographic peak intelligent analysis network using sample chromatograms until convergence. The chromatographic peak intelligent analysis network includes a feature extractor, a physical parameter probability prediction output head, and a physical signal generator. The result confidence assessment module is used to input the chromatographic signal to be analyzed into the intelligent chromatographic peak analysis network, and to assess the reliability of the prediction results based on the probability distribution of the key physical parameters output by the intelligent chromatographic peak analysis network, thereby obtaining the result confidence level. The dynamic adaptive update module is used to dynamically adjust the number and composition of sample chromatograms based on the confidence level of the results and a preset confidence threshold, and to continuously update the intelligent peak analysis network. The heatmap visualization module is used to record the weight score assigned to each data point by the attention mechanism of the chromatographic peak intelligent analysis network. The weight score is mapped back to the time series of the original chromatographic signal and visualized through color gradient. It is superimposed on the original chromatographic curve in the form of a semi-transparent layer to generate a chromatographic peak heatmap.

Citation Information

Patent Citations

  • Splicing site prediction and interpretation method based on attention mechanism

    CN114566216A

  • Chromatogram analysis method and device based on artificial intelligence and computer equipment

    CN119295714A

  • Petroleum product chemical composition automatic analysis system and method based on deep learning

    CN120847319A