Dynamic detection system and method for optimizing molecular weight distribution of cubilose peptide

By constructing a dynamic molecular weight distribution matrix using liquid chromatography-mass spectrometry (LC-MS), and combining it with multi-scale decomposition and adaptive clustering algorithms, the problems of dynamic changes and insufficient resolution in the detection of bird's nest peptide molecular weight were solved, enabling precise control of bird's nest peptide product quality and optimization of the production process.

CN122017055APending Publication Date: 2026-05-12QINGDAO CANON BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO CANON BIOTECHNOLOGY CO LTD
Filing Date
2025-11-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for detecting the molecular weight of bird's nest peptides are insufficient to accurately obtain information on the change of molecular weight distribution over time, and their resolution is inadequate. They cannot accurately analyze complex samples or eliminate interference from impurities, resulting in low accuracy and reliability of detection, and thus failing to meet market demands for quality control of bird's nest peptide products.

Method used

A dynamic molecular weight distribution matrix was constructed using liquid chromatography-mass spectrometry (LC-MS). Molecular weight distribution features were extracted through multi-scale decomposition and adaptive clustering algorithms to generate molecular weight distribution clusters. A dynamic evolution model was then used to predict the trend of molecular weight distribution changes, and detection parameters were optimized to improve resolution.

Benefits of technology

It enables dynamic monitoring and precise detection of the molecular weight distribution of bird's nest peptides, providing detailed molecular weight distribution information, supporting product quality assessment and production process control, and ensuring that product quality meets high standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122017055A_ABST
    Figure CN122017055A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cubilose peptide detection, and discloses a dynamic detection system and method for optimizing cubilose peptide molecular weight distribution. The method comprises the following steps: acquiring liquid chromatography-mass spectrometry data of a cubilose peptide sample; constructing a molecular weight dynamic distribution matrix based on the data, wherein the dimension of the molecular weight dynamic distribution matrix covers a time point index, a mass-to-charge ratio index and a signal intensity value; performing multi-scale decomposition on the matrix, and extracting molecular weight distribution characteristics under different time scales, such as a main peak position, a peak width and a peak area ratio; dynamically grouping the features by adopting a self-adaptive clustering algorithm, and generating a molecular weight distribution cluster comprising a cluster center, a cluster boundary and intra-cluster dispersion; on the basis, a dynamic evolution model is constructed and used for simulating the variation trend of molecular weight distribution along with time and predicting the merging or splitting behavior of molecular weight distribution clusters; adjusting acquisition parameters of liquid chromatography-mass spectrometry data according to a prediction result, and optimizing the molecular weight distribution resolution of subsequent detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bird's nest peptide detection technology, specifically to an optimized dynamic detection system and method for the molecular weight distribution of bird's nest peptides. Background Technology

[0002] With the continuous advancement of technology, bird's nest peptides have gradually come into people's view. Bird's nest peptides are hydrolyzed products of bird's nest and have been proven to possess superior bioactivity in some areas, surpassing traditional bird's nest. Currently, research on bird's nest peptides mainly focuses on preparation methods and bioactivity studies. Regarding preparation methods, researchers have tried various enzymatic hydrolysis techniques to obtain bird's nest peptides with different structures and functions. In the field of bioactivity research, the antioxidant, anti-aging, whitening, immune-enhancing, and anti-inflammatory effects of bird's nest peptides have been widely reported. However, research on methods for detecting the molecular weight distribution of bird's nest peptides is relatively limited.

[0003] Traditional methods for determining the molecular weight of bird's nest peptides, such as gel filtration chromatography and polyacrylamide gel electrophoresis, while capable of analyzing the molecular weight of bird's nest peptides to some extent, have several limitations. Firstly, they struggle to accurately obtain information on how the molecular weight distribution of bird's nest peptides changes over time. In actual production and research processes, the molecular weight distribution of bird's nest peptides is affected by various factors, such as enzymatic hydrolysis time, temperature, and pH value, and traditional methods cannot track these changes in real time.

[0004] Secondly, existing methods have limitations in detection resolution. Bird's nest peptides are a complex mixture containing peptides of various molecular weights, making it difficult for traditional detection methods to perform precise separation and analysis, thus hindering the accurate acquisition of detailed information on the molecular weight distribution of bird's nest peptides. Furthermore, for some complex bird's nest peptide samples, existing methods have limited analytical capabilities and are easily affected by impurities and interfering substances, thereby reducing the accuracy and reliability of detection.

[0005] In recent years, with increasing attention to health and beauty, the market demand for bird's nest peptide products has shown a rapid growth trend. From nutritional supplements to skincare products, bird's nest peptides are becoming increasingly common. In the field of nutritional supplements, bird's nest peptides are formulated into oral liquids, capsules, and other forms, providing consumers with convenient nourishing options; in the field of skincare products, bird's nest peptides are added to products such as face masks and serums to achieve whitening, moisturizing, and anti-wrinkle effects.

[0006] With market expansion, higher demands are being placed on the quality control and testing accuracy of bird's nest peptide products. Accurately determining the molecular weight distribution of bird's nest peptides is crucial for assessing product quality, optimizing production processes, and protecting consumer rights. Traditional testing methods can no longer meet market demands, necessitating a new technology to achieve dynamic and precise detection of the molecular weight distribution of bird's nest peptides. This patented technology was developed against this backdrop, aiming to address the shortcomings of existing testing methods and provide strong support for the quality control and research and development of bird's nest peptide products. Summary of the Invention

[0007] The purpose of this invention is to provide an optimized dynamic detection system and method for the molecular weight distribution of bird's nest peptides, so as to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides a method for optimizing the dynamic detection of molecular weight distribution of bird's nest peptides, the method comprising:

[0009] Acquire liquid chromatography-mass spectrometry data of bird's nest peptide samples, wherein the liquid chromatography-mass spectrometry data includes time series signals and mass-to-charge ratio distribution information;

[0010] A molecular weight dynamic distribution matrix is ​​constructed based on the liquid chromatography-mass spectrometry data. The dimensions of the molecular weight dynamic distribution matrix include time point index, mass-to-charge ratio index, and signal intensity value.

[0011] The molecular weight dynamic distribution matrix is ​​decomposed into multi-scale components to extract molecular weight distribution features at different time scales. The molecular weight distribution features include the position of the main peak, the peak width, and the peak area ratio.

[0012] An adaptive clustering algorithm is used to dynamically group the molecular weight distribution features to generate molecular weight distribution clusters, each of which includes a cluster center, a cluster boundary, and intra-cluster dispersion.

[0013] A dynamic evolution model is constructed based on the molecular weight distribution clusters. The dynamic evolution model is used to simulate the changing trend of molecular weight distribution over time and predict the merging or splitting behavior of molecular weight distribution clusters.

[0014] Based on the prediction results of the dynamic evolution model, the acquisition parameters of the liquid chromatography-mass spectrometry data are adjusted to optimize the molecular weight distribution resolution for subsequent detection.

[0015] Preferably, the step of constructing a dynamic molecular weight distribution matrix based on the liquid chromatography-mass spectrometry data includes:

[0016] The liquid chromatography-mass spectrometry data are subjected to time alignment and noise suppression processing to generate a standardized time series signal;

[0017] The signal intensity values ​​of the standardized time series signal in different mass-to-charge ratio intervals are extracted to construct a three-dimensional matrix. The row index of the three-dimensional matrix corresponds to the time point, the column index corresponds to the mass-to-charge ratio interval, and the matrix element value corresponds to the signal intensity.

[0018] The three-dimensional matrix is ​​normalized to eliminate signal intensity differences at different time points, thereby generating a dynamic molecular weight distribution matrix.

[0019] Preferably, the multi-scale decomposition of the molecular weight dynamic distribution matrix includes:

[0020] The sliding window algorithm is used to traverse the molecular weight dynamic distribution matrix and calculate the statistical characteristics of molecular weight distribution within different time windows.

[0021] A multi-scale feature vector is constructed based on the statistical characteristics of the molecular weight distribution. The multi-scale feature vector includes short-time peak shape characteristics, medium-time trend characteristics, and long-time stability characteristics.

[0022] Key molecular weight distribution features are extracted by reducing the dimensionality of the multi-scale feature vectors through principal component analysis.

[0023] Preferably, the step of dynamically grouping the molecular weight distribution features using an adaptive clustering algorithm includes:

[0024] The initial cluster centers are calculated based on the similarity measure of the key molecular weight distribution characteristics.

[0025] The number of clusters is adjusted based on dynamic density peak detection, which is achieved through local density and minimum distance threshold.

[0026] The boundaries of the clusters are iteratively optimized until the intra-cluster dispersion converges to a preset range.

[0027] Preferably, the construction of the dynamic evolution model based on the molecular weight distribution cluster includes:

[0028] Calculate the similarity matrix of molecular weight distribution clusters at adjacent time points, and use the similarity matrix to quantify the migration probability between clusters;

[0029] The evolution path of molecular weight distribution clusters is simulated based on the Markov chain model, and the evolution path includes cluster merging, cluster splitting and cluster stable state;

[0030] The evolutionary path is used to predict the trend of molecular weight distribution cluster changes at future time points.

[0031] Preferably, adjusting the acquisition parameters of the liquid chromatography-mass spectrometry data based on the prediction results of the dynamic evolution model includes:

[0032] If the prediction results indicate that molecular weight distribution clusters are about to merge, increase the mass spectrometry resolution to distinguish overlapping peaks;

[0033] If the prediction results indicate that the molecular weight distribution clusters are about to split, then extend the chromatographic separation time to improve peak resolution;

[0034] If the prediction results indicate that the molecular weight distribution cluster remains stable, then the current acquisition parameters are maintained to reduce redundant data.

[0035] Preferably, the method further includes:

[0036] Real-time monitoring of abnormal shifts in molecular weight distribution clusters, including abrupt changes in cluster centers or abnormal expansion of cluster boundaries;

[0037] When an abnormal offset is detected, the liquid chromatography-mass spectrometry data is reacquired, and the molecular weight dynamic distribution matrix is ​​updated.

[0038] Preferably, the method further includes:

[0039] A deep learning model is trained based on historical molecular weight distribution data, and the deep learning model is used to assist in the prediction of dynamic evolution models.

[0040] The output of the deep learning model is weighted and fused with the prediction results of the dynamic evolution model to generate the final molecular weight distribution prediction result.

[0041] Preferably, the method further includes:

[0042] Establish quality assessment indicators for molecular weight distribution, including peak symmetry, baseline drift, and signal-to-noise ratio;

[0043] The parameters of the clustering algorithm and evolutionary model are dynamically adjusted based on the quality assessment indicators to optimize detection accuracy.

[0044] Preferably, the present invention also includes a bird's nest peptide detection system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, implements the steps of the above-described bird's nest peptide detection method.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] In the research and production of bird's nest peptides, a comprehensive understanding of their molecular weight distribution information is crucial. This invention achieves comprehensive capture of the molecular weight distribution information of bird's nest peptides by acquiring liquid chromatography-mass spectrometry data of bird's nest peptide samples and constructing a dynamic molecular weight distribution matrix based on this data.

[0047] Traditional detection methods struggle to capture information on the temporal changes in the molecular weight distribution of bird's nest peptides. However, the method of this invention records time-series signals and mass-to-charge ratio distribution information. The constructed dynamic molecular weight distribution matrix encompasses time point indices, mass-to-charge ratio indices, and signal intensity values, making the molecular weight distribution of bird's nest peptides at different time points readily apparent. This is akin to providing researchers with a "dynamic documentary" of the molecular weight distribution of bird's nest peptides, clearly showing its changes at different stages and providing a rich data foundation for in-depth research on the properties of bird's nest peptides. This comprehensive data plays a crucial role whether exploring the formation mechanism of bird's nest peptides or studying their interactions with other substances.

[0048] Multi-scale decomposition and adaptive clustering algorithms are key technologies in this invention, playing a crucial role in accurately extracting the molecular weight distribution characteristics of bird's nest peptides. By performing multi-scale decomposition on the dynamic molecular weight distribution matrix, molecular weight distribution characteristics at different time scales can be extracted, including the position of the main peak, peak width, and peak area ratio. These characteristics, like the "fingerprint" of bird's nest peptides, reflect their quality and properties.

[0049] For example, the position of the main peak can indicate the molecular weight of the main components in bird's nest peptides, the peak width reflects the dispersion of molecular weight distribution, and the peak area ratio reflects the relative content of peptides with different molecular weights. An adaptive clustering algorithm is used to dynamically group these features, generating molecular weight distribution clusters and clearly defining the cluster center, cluster boundaries, and intra-cluster dispersion. This provides a more accurate basis for the quality assessment of bird's nest peptides, effectively distinguishing bird's nest peptide products of different quality grades. During the production process, companies can strictly control product quality based on these characteristics and cluster information, ensuring that bird's nest peptide products released to the market meet high-quality standards.

[0050] The dynamic evolution model is another innovation of this invention. It can predict the changing trend of the molecular weight distribution of bird's nest peptides over time and anticipate the merging or splitting behavior of molecular weight distribution clusters. In actual production and research, the molecular weight distribution of bird's nest peptides is affected by a variety of factors and changes. Traditional methods cannot know these changes in advance, resulting in delayed detection results and an inability to adjust production or research strategies in a timely manner. Attached Figure Description

[0051] Figure 1 This is a schematic diagram illustrating the working principle of the optimized dynamic detection method for the molecular weight distribution of bird's nest peptides described in this invention.

[0052] Figure 2 A flowchart for constructing a dynamic molecular weight distribution matrix based on liquid chromatography-mass spectrometry data;

[0053] Figure 3 This is a flowchart for multi-scale decomposition of the molecular weight dynamic distribution matrix. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] Please see Figure 1 This invention provides a system and method for dynamically detecting the molecular weight distribution of bird's nest peptides, the method comprising:

[0056] By integrating liquid chromatography-mass spectrometry (LC-MS) with time-series data analysis, real-time monitoring and parameter optimization of molecular weight distribution are achieved. LC-MS data of bird's nest peptide samples were acquired. This data includes time-series signals and mass-to-charge ratio (MTR) distribution information, sourced from the output of a standard LC-MS spectrometer. The time-series signals record the mass spectrometry scan results at different time points, while the MTR distribution information reflects the signal intensity corresponding to the molecular weight. A dynamic molecular weight distribution matrix was constructed based on the LC-MS data. The dimensions of the dynamic molecular weight distribution matrix include time point indices, MTR indices, and signal intensity values. The matrix construction process involved data alignment and normalization to ensure temporal consistency. Multi-scale decomposition of the dynamic molecular weight distribution matrix was performed to extract molecular weight distribution features at different time scales, including peak position, peak width, and peak area percentage. The multi-scale decomposition employed sliding window and statistical analysis methods to capture short-, medium-, and long-term distribution changes. An adaptive clustering algorithm is used to dynamically group molecular weight distribution features, generating molecular weight distribution clusters. Each cluster includes a cluster center, cluster boundary, and intra-cluster dispersion. The clustering process automatically adjusts the number of clusters based on feature similarity. A dynamic evolution model is constructed based on these molecular weight distribution clusters to simulate the changing trend of molecular weight distribution over time and predict the merging or splitting behavior of molecular weight distribution clusters. The model uses probabilistic methods to describe the evolution path. The acquisition parameters of the liquid chromatography-mass spectrometry (LC-MS) data are adjusted based on the prediction results of the dynamic evolution model to optimize the resolution of subsequent molecular weight distribution detection. Parameter adjustment involves the dynamic configuration of mass spectrometry resolution and chromatographic separation time.

[0057] Example 1: See Figure 2In constructing the molecular weight dynamic distribution matrix, time alignment and noise suppression were performed on the liquid chromatography-mass spectrometry (LC-MS) data to generate standardized time series signals. Time alignment was achieved through internal standard retention time calibration. An internal standard compound stably present in the bird's nest peptide sample was selected as the time reference point, and its retention time remained constant during LC separation. The signal at each detection time point was aligned and corrected with the peak position of the internal standard. A dynamic time warping algorithm was used to calculate the offset between the actual and theoretical time points. The signal was resampled using cubic spline interpolation to eliminate time drift caused by flow rate fluctuations or column efficiency changes. Noise suppression employed a combination of wavelet transform and moving average filtering to perform multi-resolution analysis on the original signal. The wavelet basis function decomposes the signal into five levels, extracts wavelet coefficients at each scale, eliminates high-frequency noise components through soft thresholding, retains low-frequency signals that reflect the true molecular weight distribution, and dynamically determines the window width of the moving average filter based on the signal sampling frequency. The window width is set to 11 data points, and the filtered signal is reconstructed to obtain a smooth time series.

[0058] A three-dimensional matrix was constructed by extracting signal intensity values ​​from standardized time-series signals across different mass-to-charge ratio (M / C) intervals. The M / C interval division was determined based on the mass spectrometer's mass accuracy, dividing the entire detection range into intervals of 0.1 units in width. Each interval corresponds to a specific mass number range. The row indices of the three-dimensional matrix strictly correspond to the sampling time points of the liquid chromatography (LC-MS), while the column indices correspond to the M / C interval numbers. The matrix element values ​​store the maximum signal intensity value within the corresponding M / C interval at that time point. The three-dimensional matrix was constructed using a dynamic memory allocation method, expanding the matrix dimensions in real time as LC-MS data acquisition progressed. Newly acquired data was validated and appended to the end of the matrix. The matrix data structure adopted a sparse matrix storage format, storing only non-zero signal intensity values ​​to optimize memory usage.

[0059] The three-dimensional matrix is ​​normalized to eliminate signal strength differences at different time points. Normalization is performed independently for each time point, calculating the arithmetic mean and standard deviation of signal strength across all mass-to-charge ratio intervals within a single time point. Z-score normalization is used to convert each signal strength value into a standard score; the conversion formula is the current signal strength value minus the mean, divided by the standard deviation. The normalized signal strength values ​​conform to a standard normal distribution. After generating the molecular weight dynamic distribution matrix, data integrity verification is performed to check for outliers or missing values. Outlier detection uses a box plot method, considering values ​​exceeding 1.5 times the quartile range as outliers and replacing them with linear interpolation. Missing values ​​are filled using a weighted average of neighboring time points. The dynamic time warping algorithm used in time alignment employs symmetric constraints to ensure the monotonicity and continuity of the time axis. Dynamic programming is used to search for warped paths, and Euclidean distance is used to fill the cumulative distance matrix. Finally, the minimum cost path is found as the optimal alignment scheme. The wavelet threshold selection in noise suppression is based on the principle of unbiased risk estimation. Adaptive thresholds are set for different decomposition levels, and the threshold size is proportional to the noise level, ensuring that effective signal features are preserved while denoising. An index lookup table is established for the mass-to-charge ratio interval mapping of the three-dimensional matrix. Each interval number corresponds to a specific upper and lower limit of the mass-to-charge ratio. The interval boundary processing adopts the rounding principle to avoid the signal value being segmented into different intervals.

[0060] The normalized molecular weight dynamic distribution matrix is ​​visualized and verified, generating a heatmap to show the signal intensity trend over time. A human-computer interface allows operators to check the matrix quality and confirm the uniformity and continuity of the signal distribution. Matrix data is stored in HDF5 format, retaining complete metadata information including sampling time points, mass-to-charge ratio range parameters, and normalization parameters, supporting fast read / write operations and subsequent processing module access. The time-aligned reference internal standard is a compound with stable chromatographic behavior, exhibiting a constant concentration in the bird's nest peptide sample and a retention time variation coefficient of less than 0.5%, ensuring alignment accuracy meets analytical requirements.

[0061] Noise suppression parameters are dynamically adjusted based on real-time signal-to-noise ratio (SNR) monitoring results. The SNR is calculated using the ratio of peak signal to baseline noise. When the SNR falls below a preset threshold, the filtering intensity is automatically increased. The wavelet denoising threshold is selected based on the Stan unbiased risk estimation method to minimize the bias introduced during the denoising process. The construction of the three-dimensional matrix implements data compression, using run-length encoding to compress continuous zero-value regions, reducing storage space usage. The matrix access interface provides the function of extracting sub-matrices by slicing according to time range or mass-to-charge ratio range. The Z-score method for normalization ensures that data at each time point are of the same magnitude, avoiding feature extraction bias caused by signal strength fluctuations. Normalization parameters are saved separately for the same processing of new data. The construction process of the molecular weight dynamic distribution matrix includes a quality control feedback mechanism. Each processing step generates a quality assessment report, recording processing parameters and execution results. When data quality indicators fail to meet the standards, a reprocessing process is triggered. The accuracy of time alignment is evaluated by calculating the standard deviation of the retention time of the aligned internal standard. The standard deviation is controlled within 0.1 seconds to ensure temporal consistency. The noise suppression effect is quantitatively evaluated by calculating signal smoothness and noise variance. The memory management of the three-dimensional matrix adopts a block storage strategy, which divides the large matrix into multiple sub-blocks to reduce memory fragmentation, and the matrix operation utilizes multi-threaded parallel computing to accelerate the processing speed.

[0062] The normalized matrix is ​​then subjected to a distribution uniformity test using... The signal intensity distribution at different time points was examined and compared to ensure consistency in the distribution shape. Time points that failed the test were marked as suspicious and manually reviewed. The entire matrix construction process was automated and streamlined. After the raw liquid chromatography-mass spectrometry data was input, the time alignment, noise suppression, 3D matrix construction, and normalization processing sequences were automatically executed. The processing log recorded the input, output, and anomalies of each step in detail. The time alignment algorithm included an anomaly handling mechanism. When the internal standard peak was missing, it automatically switched to retention time prediction mode and used a linear regression model to predict the current time point based on historical data.

[0063] Noise suppression processing offers a variety of filter options, including The system employs filters and Kalman filters, selecting appropriate denoising methods based on different signal characteristics. Filter parameters are optimized through cross-validation. The 3D matrix data structure supports fast query and update operations, providing an application programming interface for reading and writing matrix elements. Matrix dimension information is stored separately in header and data files. The reversible design of normalization processing allows for restoration of the original signal intensity when needed. Normalization parameters are stored separately in configuration files for easy tracing of processing history. The persistent storage of the molecular weight dynamic distribution matrix uses a columnar storage format to improve the efficiency of querying by mass-to-charge ratio intervals. The data index is built based on a B+ tree structure to support fast range queries. The selection of reference points for time alignment considers the case of multiple internal standards. When multiple internal standards are used, a weighted average method is used to calculate the comprehensive alignment scheme, with weights allocated according to the stability of the internal standards. The signal after noise suppression undergoes baseline correction. An asymmetric weighted least squares algorithm is used to estimate the baseline profile, and the baseline component is subtracted from the original signal to obtain the pure compound signal. The construction of the 3D matrix supports an incremental update mode. Newly acquired data is appended to the existing matrix after the same processing. Matrix version management records the timestamp and changes for each update.

[0064] The robust design of the normalization process handles abnormal signal conditions. When a signal is completely missing at a certain time point, it automatically skips that point to prevent calculation errors, and marks all time points as invalid when the signal strength is zero. Metadata management for the molecular weight dynamic distribution matrix adopts a standardized format, containing complete information such as instrument parameters, sampling conditions, and processing methods to ensure data reproducibility. The accuracy verification of time alignment processing is achieved by calculating the coefficient of variation of the retention time of the internal standard before and after alignment; a coefficient of variation of less than 1% is required for effective alignment. The noise suppression effect is evaluated using a signal distortion index, striking a balance between denoising degree and signal fidelity, and achieving optimal denoising effect through parameter tuning. Storage optimization of the three-dimensional matrix includes data partitioning and compression strategies. Large matrices are divided into appropriately sized blocks, each independently compressed and stored, and decompressed on demand during access to reduce memory pressure. The normalization method for normalization processing is selected based on data distribution characteristics. For skewed data, logarithmic transformation preprocessing is used to approximate a normal distribution before Z-score normalization. The construction process of the molecular weight dynamic distribution matrix is ​​implemented with real-time monitoring, displaying processing progress and quality on a dashboard.

[0065] Example 2: See Figure 3When performing multi-scale decomposition of the molecular weight dynamic distribution matrix, a sliding window algorithm is used to traverse the matrix. The sliding window algorithm sets three different window widths: a short-term window with 5 consecutive time points, a medium-term window with 20 consecutive time points, and a long-term window with 100 consecutive time points. The window sliding step size is fixed at 1 time point. Each window covers all mass-to-charge ratio intervals within the corresponding time range of the matrix. The statistical characteristics of the molecular weight distribution within different time windows are calculated. These characteristics include the mean, median, standard deviation, skewness, and kurtosis of the signal intensity within the window. These statistics are calculated independently for each mass-to-charge ratio interval, forming a complete description of the window characteristics. The sliding window algorithm uses a circular buffer to achieve efficient data access. When the window moves, the oldest time point data is removed and the latest time point data is added, maintaining the temporal continuity of the data within the window.

[0066] A multi-scale feature vector is constructed based on the statistical characteristics of molecular weight distribution. This multi-scale feature vector consists of three sub-vectors: a short-time kurtosis feature vector extracts the position, height, and width of signal peaks within a window; a medium-time trend feature vector calculates the linear regression slope and coefficient of determination of signal changes within the window; and a long-time stability feature vector evaluates the autocorrelation function and variance rate of change of the signal within the window. Each sub-vector contains feature values ​​across multiple dimensions: the short-time kurtosis feature vector has a dimension of 10, the medium-time trend feature vector has a dimension of 8, and the long-time stability feature vector has a dimension of 6. These three sub-vectors are concatenated to form a 24-dimensional multi-scale feature vector. The feature vector values ​​are standardized to ensure that the feature values ​​of each dimension are of the same order of magnitude, preventing certain features from dominating subsequent analysis due to excessively large numerical ranges.

[0067] Principal Component Analysis (PCA) is used to reduce the dimensionality of multi-scale eigenvectors. The PCA algorithm calculates the covariance matrix of the 24-dimensional eigenvectors, solves for the eigenvalues ​​and eigenvectors of the covariance matrix, and selects principal components by sorting the eigenvalues ​​from largest to smallest. The top k principal components with a cumulative variance contribution rate of 95% are retained. The value of k is determined based on the actual data distribution, typically between 8 and 12. The original 24-dimensional eigenvectors are projected onto the k-dimensional principal component space to form the key molecular weight distribution features. Before PCA, the eigenvectors are Z-score standardized to eliminate the influence of differences in dimensions on the covariance calculation. An adaptive clustering algorithm is used to dynamically group the molecular weight distribution features. This adaptive clustering algorithm is based on an improved k-means algorithm framework. Initial cluster centers are selected using the k-means algorithm, which selects points farther from the previously selected centers using a probability distribution method to avoid initial centers clustering in local areas. Initial cluster centers are calculated based on the similarity measure of key molecular weight distribution characteristics. The similarity measure uses a weighted Euclidean distance formula, and the weight of each feature dimension is set according to its variance contribution rate. Dimensions with higher variance contribution rates are given larger weights.

[0068] The number of clusters is adjusted based on dynamic density peak detection. The dynamic density peak detection algorithm calculates the local density of each feature point within its neighborhood radius, where local density is defined as the number of other points in the neighborhood. The neighborhood radius is adaptively determined based on the data distribution. Simultaneously, the distance from each feature point to its nearest high-density point is calculated, called the minimum distance. The product of the local density and the minimum distance is used as a measure of cluster centrality, and the point with the highest centrality is selected as the candidate cluster center. The number of clusters is automatically determined by setting local density thresholds and minimum distance thresholds. The local density threshold is the median of the local densities of all points, and the minimum distance threshold is the average of the minimum distances of all points. The boundaries of the clusters are iteratively optimized until the intra-cluster dispersion converges to a preset range. The iteration process uses the expectation-maximization algorithm. The expectation step calculates the probability that each feature point belongs to each cluster, and the maximization step updates the cluster center positions and cluster boundaries. Intra-cluster dispersion is represented by the average distance from all feature points within the cluster to the cluster center. The preset convergence range is set to a change in intra-cluster dispersion of less than 0.01 between two adjacent iterations. The cluster separation is monitored in real time during clustering, and the neighborhood radius is automatically adjusted and the density peak is recalculated when cluster overlap occurs.

[0069] The sliding window algorithm employs time series segmentation, with each window processed independently to generate feature vectors. Overlapping areas between windows are weighted using a Hamming window function to reduce boundary effects. Short-term peak shape extraction of multi-scale feature vectors utilizes Gaussian fitting, applying a Gaussian function to the signal profile within each mass-to-charge ratio interval of the window and extracting the fitting parameters as feature values. Mid-term trend features are calculated using least squares linear regression, with the regression model including intercept and slope terms to assess the strength and direction of signal change trends. Long-term stability features are calculated using the autocorrelation function with a fast Fourier transform algorithm, improving computational efficiency through frequency domain transformation. The principal component analysis algorithm includes an eigenvalue decomposition step, using the QR algorithm to iteratively solve for the eigenvalues ​​of the covariance matrix. Eigenvectors are orthogonalized using the Schmitt orthogonalization method. The dimensionality-reduced key molecular weight distribution features are stored as a low-dimensional matrix, preserving the main variation information of the original features. The k-means++ initialization process for the adaptive clustering algorithm employs multiple rounds of random sampling to ensure the representativeness of the initial centroids, and matrix operations are used to optimize distance calculation and accelerate the clustering process.

[0070] The neighborhood radius for dynamic density peak detection is determined using the k-nearest neighbor algorithm. For each point, the distance to its k-th nearest neighbor is calculated as the radius for local density calculation, with k set to the square root of the total number of points based on the data size. Local density calculation employs a kernel density estimation method, using a Gaussian kernel function to weight points within the neighborhood, assigning larger weights to closer points. Minimum distance calculation is accelerated by constructing a kd-tree spatial index structure to improve nearest neighbor search efficiency. The expectation-maximization algorithm in the iterative optimization process includes a regularization term to prevent cluster centers from overfitting to outliers. Cluster boundary updates use a soft assignment strategy, allowing each point to belong to multiple clusters with a certain probability. Monitoring intra-cluster dispersion uses a statistical process control method, triggering cluster re-initialization when dispersion abnormally increases. Clustering results are evaluated using the silhouette coefficient, which measures the balance between intra-cluster compactness and inter-cluster separation, selecting the clustering scheme with the largest silhouette coefficient as the final result. The time complexity of the sliding window algorithm is optimized through incremental computation, reusing some calculation results from the previous window as the window moves, reducing redundant computation. Multi-scale eigenvectors are stored using a sparse representation method, compressing and storing eigenvalues ​​close to zero. Principal component analysis (PCA) eigenvalue decomposition uses a block-based algorithm to process large-scale data, dividing the matrix into blocks and computing eigenvalues ​​in parallel. The parallel implementation of the adaptive clustering algorithm utilizes multi-threading technology, splitting the data and performing clustering operations simultaneously in different threads.

[0071] The candidate center point selection for dynamic density peak detection employs a hierarchical strategy, first screening potential center regions at a coarse-grained level, and then precisely locating center points at a fine-grained level. The convergence criterion for the iterative optimization process combines absolute and relative changes; convergence is considered achieved when the dispersion change in three consecutive iterations is less than a threshold. The visualization of the clustering results uses... Dimensionality reduction techniques project high-dimensional features into a two-dimensional space, facilitating manual inspection of clustering quality. The entire multi-scale decomposition and clustering process integrates a quality monitoring mechanism, outputting quality evaluation metrics at each step. When metrics are abnormal, algorithm parameters are automatically adjusted or previous steps are re-executed. Algorithm parameter settings are managed through configuration files, allowing for flexible adjustments based on different sample characteristics. Intermediate results during processing are persistently stored, facilitating fault recovery and result traceability. Temporal alignment of multi-scale feature vectors ensures the comparability of features generated by different windows, and window boundary processing employs mirror expansion to avoid data truncation.

[0072] Principal component analysis (PCA) employs cross-validation to evaluate the impact of different principal component numbers on subsequent clustering results, selecting the number of principal components that maximizes the cluster silhouette coefficient. The adaptive clustering algorithm supports multiple distance metrics, including Mahalanobis distance and cosine distance, automatically selecting the most suitable metric based on data distribution characteristics. The threshold setting for dynamic density peak detection utilizes statistical learning, training a threshold prediction model using clustering results from historical data to improve threshold accuracy. The iterative optimization process incorporates cluster merging and splitting mechanisms to detect unstable cluster assignments. When the dispersion of a cluster continuously increases, a cluster splitting operation is automatically triggered; when the center distance between two clusters is too close, they are automatically merged into a larger cluster. The stability of the clustering results is evaluated through multiple runs of the clustering algorithm, selecting the most frequently occurring clustering pattern as the final result. The execution time of the entire process is controlled within a reasonable range through algorithm optimization, supporting real-time processing of liquid chromatography-mass spectrometry (LC-MS) data streams.

[0073] Example 3: When constructing a dynamic evolution model based on molecular weight distribution clusters, the similarity matrix of molecular weight distribution clusters at adjacent time points is calculated. The similarity matrix is ​​used to quantify the migration probability between clusters. The construction of the similarity matrix adopts... The exponential method measures the degree of overlap between two clusters in the mass-to-charge ratio space. For cluster i at time t and cluster j at time t+1, the similarity is... Calculated using the following formula:

[0074]

[0075] in: Cluster with cluster The similarity value between them ranges from 0 to 1; Cluster The set of regions occupied by the mass-to-charge ratio in space; Cluster The set of regions occupied by the mass-to-charge ratio in space; This represents the size of the intersection of two sets of regions, i.e., the number of overlapping mass-to-charge ratio intervals; This represents the size of the union of two region sets, i.e., the total number of all involved mass-to-charge ratio intervals. The dimension of the similarity matrix is ​​equal to the product of the number of clusters of adjacent time points, and the row indices of the matrix correspond to the time points. Cluster number, column index corresponding to the time point The clusters are numbered, and each matrix element stores the corresponding similarity value. Before similarity calculation, the cluster regions need to be discretized, and the mass-to-charge ratio space is divided into intervals of fixed width, with each interval width set to 0.01 units, to ensure the accuracy of intersection and union calculations.

[0076] A Markov chain model is constructed based on the similarity matrix to simulate the evolution path of molecular weight distribution clusters. The state space of the Markov chain model consists of all possible cluster identifiers, and the state transition probability matrix is ​​directly derived from the similarity matrix. Each element in the similarity matrix... After row normalization, it is converted into transition probabilities. Row normalization divides each row element by the sum of its elements, ensuring that the total transition probability from any state is 1. The Markov chain model assumes the Markov property, meaning future states depend only on the current state and are independent of historical states. The evolutionary path simulation uses the Viterbi algorithm to find the most probable state sequence. The Viterbi algorithm calculates the maximum probability path at each time point through dynamic programming. The path results include cluster merging events, cluster splitting events, and cluster stable states. A cluster merging event corresponds to multiple states transitioning to the same state, a cluster splitting event corresponds to one state transitioning to multiple states, and a cluster stable state indicates that the state remains unchanged.

[0077] This method predicts the molecular weight distribution cluster change trend at future time points based on the evolutionary path. The prediction process is based on the n-step transition probability matrix of a Markov chain, which is calculated by raising the nth power of the single-step transition probability matrix. The prediction time range is set according to the detection requirements, typically predicting cluster state changes at the next 5-10 time points. For each future time point, the distribution probability of each state is calculated. The trend analysis includes the probability of cluster disappearance, the probability of new cluster appearance, and the probability of cluster stability. When the predicted probability of a state is below a threshold, it is marked as possibly disappearing; when the predicted probability of a new state is above the threshold, it is marked as possibly appearing. The prediction results are output in probability distribution form, containing uncertainty information for easy risk assessment.

[0078] Based on the predictions of the dynamic evolution model, the acquisition parameters of the liquid chromatography-mass spectrometry (LC-MS) data were adjusted. When the prediction indicated that molecular weight distribution clusters were about to merge, the mass spectrometry resolution was increased to distinguish overlapping peaks. This adjustment was achieved by reducing the scan step size of the mass analyzer from the default 0.1 units to 0.05 units. Increasing the resolution increased the mass spectrometry scan time but enhanced peak discrimination ability. The distinction of overlapping peaks was verified by checking the changes in peak sharpness along the mass-to-charge ratio axis. When the prediction indicated that molecular weight distribution clusters were about to split, the chromatographic separation time was extended to improve peak resolution. This extension was achieved by reducing the mobile phase flow rate from 1.0 mL / min to 0.5 mL / min. Extending the separation time improved the separation efficiency of the chromatographic column. Peak resolution was evaluated by calculating the depth ratio of adjacent peaks. When the prediction indicated that molecular weight distribution clusters remained stable, the current acquisition parameters were maintained to reduce redundant data. The current acquisition parameters included a mass spectrometry resolution step size of 0.1 units and a chromatographic flow rate of 1.0 mL / min. Parameter maintenance is based on stability probability assessment. When the stability probability exceeds 95%, no adjustment is considered necessary, reducing unnecessary data acquisition and saving storage space and processing time. Parameter adjustment decisions are executed in real time through a control algorithm. Adjustment commands are sent to the liquid chromatography-mass spectrometry instrument via a digital interface, with an instrument response time of less than 100 milliseconds.

[0079] The calculation of the similarity matrix involves the measurement of the overlap of cluster regions, which are defined as the convex hull or polygonal boundary in the mass-to-charge ratio space. Region discretization maps the boundary onto the mass-to-charge ratio interval grid. The exponential calculation is accelerated using a bitmap method, where each cluster region is represented as a binary bitmap, and bit operations quickly calculate the intersection and union sizes. The similarity matrix is ​​stored in a sparse matrix format, storing only non-zero similarity values ​​to reduce memory usage. The state transition probability matrix of the Markov chain model is dynamically adjusted as new data arrives, using an exponentially weighted moving average update strategy, where new transition probabilities are weighted and fused with historical probabilities. The Viterbi algorithm implementation includes initialization and recursive steps. Initialization sets the initial state probabilities, and the recursive step calculates the maximum probability path at each time point. Path backtracking extracts the optimal state sequence, with a sequence length equal to the number of detection time points. Multi-step predictions in the prediction model use matrix exponentiation, which optimizes computational efficiency through eigenvalue decomposition. The parameter adjustment logic integrates a feedback control mechanism, comparing prediction results with actual observations in real time, and adaptively optimizing the adjustment strategy based on error signals.

[0080] Normalization of the similarity matrix ensures the reasonableness of the transition probabilities, while row normalization prevents probability leakage. The steady-state distribution calculation of the Markov chain model assesses the long-term evolution trend; the steady-state distribution is obtained by solving for the eigenvectors. Confidence interval estimation of the prediction results uses Monte Carlo simulation to generate multiple possible paths and calculate probability distributions. Parameter adjustment thresholds are set based on experimental data: the merging probability threshold is set to 0.7, the splitting probability threshold to 0.6, and the stability probability threshold to 0.95. Specific parameters for mass spectrometry resolution adjustment include the mass analyzer's RF voltage and scan speed; increasing resolution requires increasing the RF voltage and decreasing the scan speed. Chromatographic separation time adjustment involves modifying the gradient elution program, extending the duration of the linear gradient segment. Parameter maintenance monitors instrument stability, and periodic calibration prevents drift. The event triggering mechanism of the control algorithm detects probability threshold crossovers, and the triggering conditions include hysteresis to prevent oscillation.

[0081] The granularity of region discretization in similarity calculation affects accuracy; too coarse granularity leads to overestimation of similarity, while too fine granularity increases computational burden. Higher-order extensions of Markov chain models consider historical states, increasing the dimensionality of the state space. Rolling time-domain optimization of the prediction model executes only the next adjustment after each prediction, reducing the impact of long-term prediction uncertainty. Prioritization of parameter adjustments handles multiple events, prioritizing merging events over splitting events. Symmetry handling of the similarity matrix is ​​important for bidirectional evolution, but Markov chain models typically employ unidirectional transitions. Path pruning in the Viterbi algorithm reduces computational complexity; the pruning threshold is set to one percent of the probability value. Visualization of prediction results helps operators understand trends; the probability distribution is presented as a heatmap. Historical logs of parameter adjustments record the reasons and effects of each adjustment, used for model improvement.

[0082] The limits of mass spectrometry resolution adjustment are constrained by instrument performance, with the maximum resolution corresponding to a minimum step size of 0.01 units. Chromatographic separation time extension is limited by the analysis cycle, with the maximum extension not exceeding 20% ​​of the total cycle. Parameter maintenance duration is monitored to prevent over-conservatism; a reassessment is performed after 10 time points of stable operation. The control algorithm robustly handles sensor noise, employing Kalman filtering to smooth probability inputs. The similarity matrix update frequency is synchronized with data acquisition, recalculating the most recent matrix at each new time point. Markov chain model state reduction merges similar states to reduce model complexity. An ensemble method for prediction models combines multiple Markov chains to improve prediction accuracy. Coordinated control of parameter adjustments coordinates mass spectrometry and chromatographic parameters to avoid mutual interference. The region definition in the exponential calculation considers the cluster density distribution, and the weighted overlap area emphasizes high-density regions. The non-homogeneous expansion of the Markov chain model adapts to time-varying transition probabilities, which are functions of time. Uncertainty propagation in the prediction results is quantified through error variance, which is used to adjust the confidence level. Gradient descent optimization for parameter adjustment finds the optimal parameter set to minimize the prediction error.

[0083] Example 4: Real-time monitoring of abnormal migration of molecular weight distribution clusters is achieved by continuously tracking changes in cluster center coordinates and cluster boundary contours. Abnormal migration includes two types: cluster center abrupt changes and abnormal cluster boundary expansion. The monitoring system acquires the current state of the molecular weight distribution clusters every 30 seconds and compares it with historical states. Historical state data is stored in the detection records of the past 24 hours. Cluster center abrupt changes are detected by calculating the Euclidean distance between the current cluster center position and the average center position of the previous 10 time points. When the distance exceeds a threshold, it is marked as a abrupt event. Abnormal cluster boundary expansion is detected by comparing the ratio of the current cluster boundary area to the historical average area. When the ratio exceeds a set range, an alarm is triggered. The threshold parameters for abnormal migration monitoring are determined based on statistical analysis of historical data, as shown in Table 1.

[0084] Table 1: Threshold Table for Abnormal Offset Monitoring Parameters

[0085] Monitoring parameters Calculation method Normal range Warning threshold Abnormal threshold Data source Cluster center movement distance Euclidean distance calculation 0-0.5 units 0.5-1.0 units >1.0 unit 10 consecutive time points Cluster boundary expansion rate Area change rate ±5% / minute ±5-10% / minute >±10% / minute Slide window 20 minutes Intra-cluster dispersion variation Standard deviation ratio 0.8-1.2 0.6-0.8 or 1.2-1.5 <0.6 or >1.5 Comparison of adjacent time points Signal strength fluctuation coefficient of variation <0.15 0.15-0.25 >0.25 All points within a single cluster

[0086] Upon detecting an abnormal offset, the system immediately triggers a reacquisition of LC-MS data, prioritizing the routine acquisition schedule. After reacquisition begins, the system pauses the current data stream processing and sends an interrupt command to the LC-MS instrument. The instrument then immediately begins a new acquisition sequence after completing the current scan cycle. Reacquisition parameter settings are adjusted based on the type and severity of the abnormal offset; mass spectrometry scan density is increased for abrupt cluster center changes, and chromatographic separation time is extended for abnormal cluster boundary expansion. The newly acquired data, after validation, replaces the original abnormal data segments, and the molecular weight dynamic distribution matrix is ​​updated accordingly, maintaining the temporal continuity of the matrix during the update process.

[0087] A deep learning model was trained using historical molecular weight distribution data. The model employed a Long Short-Term Memory (LSTM) network architecture. The network input consisted of molecular weight distribution feature sequences from the past 60 time points, and the output was the cluster state probability distribution for the next 10 time points. Training data comprised 5000 complete detection sequences from the past three months, each covering a 120-minute detection period. Data preprocessing included normalization and sequence alignment. Normalization scaled each feature dimension to the 0-1 range, while sequence alignment ensured consistent input sequence length. The LTM network contained three hidden layers with 128 neurons each. A dropout rate of 0.2 was used to prevent overfitting. The Adam algorithm was used as the optimizer, with an initial learning rate of 0.001. Five-fold cross-validation was employed during training. The dataset was randomly divided into five subsets, with four subsets used for training and the remaining subset for validation. The training cycle was set to 100 epochs, with an early stopping mechanism terminating training if the validation set loss did not improve for 10 consecutive epochs. Model evaluation used mean absolute error and root mean square error (RMSE) metrics, and the optimal model parameters were saved as a weight file. The forward propagation computation of Long Short-Term Memory (LSTM) networks involves activation function operations for the input gate, forget gate, and output gate. Cell state updates are controlled by a gating mechanism to manage information flow. Backpropagation uses the time-varying backpropagation algorithm, and gradient clipping limits the gradient norm to no more than 5.0.

[0088] The output of the deep learning model and the prediction results of the dynamic evolution model are weighted and fused. The weights are dynamically calculated based on the model's performance on the validation set. The weights of the deep learning model are set based on its prediction accuracy, which is calculated as the ratio of the predicted probability to the actual state. The weights of the dynamic evolution model consider the reliability of its state transition probabilities, which is evaluated through the calibration of historical predictions. The weighted fusion adopts a linear combination method, and the fusion formula is: final probability = w1 × deep learning model probability + w2 × dynamic evolution model probability, with the weights w1 and w2 summing to 1. The fusion result generates the final molecular weight distribution prediction result, which includes the probability distribution and confidence interval of each cluster state at each future time point. The specific implementation of abnormal offset monitoring includes multi-indicator collaborative judgment. The calculation of the cluster center movement distance uses the Euclidean distance formula in mass-to-charge ratio space, and the distance value is standardized to eliminate the influence of dimensions. The calculation of the cluster boundary expansion rate is based on the change in the area of ​​the boundary polygon, which is reconstructed from scattered data using the Alpha shape algorithm. Monitoring of intra-cluster dispersion changes employs a sliding window statistical method, with a window width set to 15 time points. The dispersion ratio is calculated as the ratio of the standard deviation of the current window to that of historical windows. Signal strength fluctuations are assessed using the coefficient of variation (CV), calculated as the ratio of the standard deviation to the mean, thus eliminating the influence of absolute signal strength.

[0089] The re-acquisition process is triggered by multiple threshold settings: a warning threshold triggers system logging, while an anomaly threshold initiates actual re-acquisition. Data quality verification during re-acquisition includes signal-to-noise ratio checks, baseline stability testing, and peak shape symmetry assessment; failure to pass any of these verifications necessitates re-acquisition. The molecular weight dynamic distribution matrix is ​​updated using a version control mechanism, retaining historical versions for backtracking analysis, while the new version matrix undergoes multi-scale decomposition and cluster analysis.

[0090] Training data augmentation for deep learning models employs a time-series warping method, generating new training samples by slightly stretching or compressing the time axis. Data standardization is performed separately for each feature dimension, with the mean and standard deviation for each dimension calculated from the training set. Hyperparameter tuning for the Long Short-Term Memory network utilizes a grid search method, experimenting with different combinations of layers, neurons, and dropout rates. Model deployment leverages the TensorFlow framework, while inference optimization utilizes a GPU for acceleration. Weighted fusion updates weights daily, recalculating them using validation results from the most recent 7 days. Weight calculation considers both model stability and consistency; stability measures the volatility of model performance, while consistency assesses the degree of agreement between model predictions and actual conditions. Confidence intervals for fusion results are estimated using the Bootstrap method, calculating quantiles by repeatedly sampling 1000 times from the model output distribution. The system implementation for anomaly offset monitoring includes a real-time data pipeline, with latency from instrument acquisition to monitoring result output controlled within 100 milliseconds. The monitoring algorithm employs a parallel computing architecture, simultaneously calculating multiple monitoring metrics to improve efficiency. Historical data retrieval uses a time-series database, supporting fast range queries and aggregation calculations. The alarm management module takes different measures based on the anomaly level; low-level alarms are only logged, while high-level alarms trigger email and SMS notifications. Instrument control during the re-acquisition process is implemented via the OPCUA protocol, and parameter settings are transmitted in a structured data format. The data verification algorithm includes multiple checkpoints: raw data verification checks file integrity and format correctness, and quality checks evaluate signal quality indicators. Matrix update operations ensure atomicity, avoiding data inconsistencies during the update process. Version management records the timestamp, changes, and operator for each update.

[0091] The online learning function of the deep learning model supports incremental updates. Newly collected data is added to the training set after labeling, and the model is retrained weekly. Model performance monitoring tracks prediction bias and loss function values, triggering emergency updates when performance deteriorates. Model interpretability analysis uses the SHAP method to identify the feature dimensions that have the greatest impact on prediction results. The model deployment environment is containerized to ensure consistency of the operating environment.

[0092] The weighted fusion system features a redundant design with multiple backup models, automatically switching to a backup model when one fails. A smoothing factor is incorporated into weight calculation to prevent drastic weight fluctuations from affecting fusion stability. The visualization of fusion results includes probability distribution plots and trend curves, supporting interactive exploration. System integration testing covers various abnormal scenarios to ensure the robustness of the fusion logic. Anomaly offset monitoring calibration is performed regularly, using known standard samples to verify the accuracy of the monitoring algorithm. Threshold parameters are dynamically adjusted based on seasonal changes and instrument status, using control chart methods. Statistical analysis of monitoring results generates daily reports, including the number, type distribution, and duration of abnormal events. The system maintenance module regularly checks the performance indicators of the monitoring algorithm and updates algorithm parameters promptly. Optimization of the re-acquisition logic considers instrument load balancing to avoid frequent re-acquisitions affecting other detection tasks. Data validation standards are set differently based on sample type, with more lenient validation standards for complex samples. Matrix update conflict resolution adopts a timestamp-first principle, with later data overwriting earlier data. The system recovery mechanism automatically resumes processing after power outages or network interruptions, ensuring data integrity.

[0093] Distributed training of deep learning models utilizes multiple GPU servers, with training data shards stored in a distributed file system. Model version management supports rapid rollback, allowing restoration of older models when new models underperform. Model inference services are provided through... The system exposes and supports concurrent request processing. A model monitoring panel displays request volume, response time, and error rate in real time. A weighted fusion anomaly handling mechanism detects the reasonableness of input data, marking models as suspicious when their output significantly deviates from historical ranges. Adaptive weight adjustment considers the correlation between models, reducing the weights of highly correlated models. Post-processing of the fusion results includes smoothing filtering and outlier removal to improve output stability. System performance benchmark tests are conducted regularly to ensure real-time requirements are met.

[0094] Example 5: Establishing a molecular weight distribution quality assessment system, including three core indicators: peak symmetry, baseline drift, and signal-to-noise ratio. Peak symmetry was calculated using Gaussian fitting to determine the half-width-to-height ratio of the chromatographic peak, with a goodness-of-fit requirement of R² ≥ 0.98. Baseline drift was assessed by fitting the baseline slope using a cubic polynomial, with the normal range limited to [value missing]. The signal-to-noise ratio (SNR) is calculated as the ratio of peak height to noise standard deviation, with a threshold of 20:1. These indicators drive dynamic optimization of detection parameters: when peak symmetry is below 0.9, the weight of the mass-to-charge ratio dimension in the clustering algorithm is increased from 1.0 to 1.5; when baseline drift exceeds the limit, the evolutionary model introduces a drift compensation factor and shortens the prediction step size; when the SNR decreases, the mass spectrometry parameters (ion source temperature +5-10℃, collision energy -2-3eV) are adjusted and the acquisition time is extended by 10%-20%. The system establishes a mapping relationship between quality indicators and algorithm parameters, achieving adaptive optimization through closed-loop feedback. The clustering algorithm adjusts the neighborhood radius based on the SNR and optimizes the number of clusters based on peak symmetry; the evolutionary model adjusts the learning rate (0.01-0.001) and regularization intensity based on baseline drift. Parameter adjustment adopts grid search and incremental optimization, with a single adjustment amplitude ≤10%, while a parameter knowledge base is established to record the optimal configuration. The real-time monitoring dashboard displays indicator trends, and a multi-level alarm mechanism triggers automatic adjustment. The system periodically verifies the accuracy of indicators using standard samples, employs joint parameter tuning and cross-validation to ensure optimization, and ultimately uses quality control charts to continuously track improvements in detection accuracy.

[0095] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0096] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for optimizing the dynamic detection of molecular weight distribution of bird's nest peptides, characterized in that, include: Acquire liquid chromatography-mass spectrometry data of bird's nest peptide samples, wherein the liquid chromatography-mass spectrometry data includes time series signals and mass-to-charge ratio distribution information; A molecular weight dynamic distribution matrix is ​​constructed based on the liquid chromatography-mass spectrometry data. The dimensions of the molecular weight dynamic distribution matrix include time point index, mass-to-charge ratio index, and signal intensity value. The molecular weight dynamic distribution matrix is ​​decomposed into multi-scale components to extract molecular weight distribution features at different time scales. The molecular weight distribution features include the position of the main peak, the peak width, and the peak area ratio. An adaptive clustering algorithm is used to dynamically group the molecular weight distribution features to generate molecular weight distribution clusters, each of which includes a cluster center, a cluster boundary, and intra-cluster dispersion. A dynamic evolution model is constructed based on the molecular weight distribution clusters. The dynamic evolution model is used to simulate the changing trend of molecular weight distribution over time and predict the merging or splitting behavior of molecular weight distribution clusters. Based on the prediction results of the dynamic evolution model, the acquisition parameters of the liquid chromatography-mass spectrometry data are adjusted to optimize the molecular weight distribution resolution for subsequent detection.

2. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 1, characterized in that, The construction of the molecular weight dynamic distribution matrix based on the liquid chromatography-mass spectrometry data includes: The liquid chromatography-mass spectrometry data are subjected to time alignment and noise suppression processing to generate a standardized time series signal; The signal intensity values ​​of the standardized time series signal in different mass-to-charge ratio intervals are extracted to construct a three-dimensional matrix. The row index of the three-dimensional matrix corresponds to the time point, the column index corresponds to the mass-to-charge ratio interval, and the matrix element value corresponds to the signal intensity. The three-dimensional matrix is ​​normalized to eliminate signal intensity differences at different time points, thereby generating a dynamic molecular weight distribution matrix.

3. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 2, characterized in that, The multi-scale decomposition of the molecular weight dynamic distribution matrix includes: The sliding window algorithm is used to traverse the molecular weight dynamic distribution matrix and calculate the statistical characteristics of molecular weight distribution within different time windows. A multi-scale feature vector is constructed based on the statistical characteristics of the molecular weight distribution. The multi-scale feature vector includes short-time peak shape characteristics, medium-time trend characteristics, and long-time stability characteristics. Key molecular weight distribution features are extracted by reducing the dimensionality of the multi-scale feature vectors through principal component analysis.

4. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 3, characterized in that, The step of dynamically grouping the molecular weight distribution features using an adaptive clustering algorithm includes: The initial cluster centers are calculated based on the similarity measure of the key molecular weight distribution characteristics. The number of clusters is adjusted based on dynamic density peak detection, which is achieved through local density and minimum distance threshold. The boundaries of the clusters are iteratively optimized until the intra-cluster dispersion converges to a preset range.

5. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 4, characterized in that, The construction of a dynamic evolution model based on the molecular weight distribution cluster includes: Calculate the similarity matrix of molecular weight distribution clusters at adjacent time points, and use the similarity matrix to quantify the migration probability between clusters; The evolution path of molecular weight distribution clusters is simulated based on the Markov chain model, and the evolution path includes cluster merging, cluster splitting and cluster stable state; The evolutionary path is used to predict the trend of molecular weight distribution cluster changes at future time points.

6. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 5, characterized in that, The step of adjusting the acquisition parameters of liquid chromatography-mass spectrometry data based on the prediction results of the dynamic evolution model includes: If the prediction results indicate that molecular weight distribution clusters are about to merge, increase the mass spectrometry resolution to distinguish overlapping peaks; If the prediction results indicate that the molecular weight distribution clusters are about to split, then extend the chromatographic separation time to improve peak resolution; If the prediction results indicate that the molecular weight distribution cluster remains stable, then the current acquisition parameters are maintained to reduce redundant data.

7. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 6, characterized in that, The method further includes: Real-time monitoring of abnormal shifts in molecular weight distribution clusters, including abrupt changes in cluster centers or abnormal expansion of cluster boundaries; When an abnormal offset is detected, the liquid chromatography-mass spectrometry data is reacquired, and the molecular weight dynamic distribution matrix is ​​updated.

8. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 7, characterized in that, The method further includes: A deep learning model is trained based on historical molecular weight distribution data, and the deep learning model is used to assist in the prediction of dynamic evolution models. The output of the deep learning model is weighted and fused with the prediction results of the dynamic evolution model to generate the final molecular weight distribution prediction result.

9. The method for dynamically detecting the molecular weight distribution of bird's nest peptides according to claim 8, characterized in that, The method further includes: Establish quality assessment indicators for molecular weight distribution, including peak symmetry, baseline drift, and signal-to-noise ratio; The parameters of the clustering algorithm and evolutionary model are dynamically adjusted based on the quality assessment indicators to optimize detection accuracy.

10. A bird's nest peptide detection system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the bird's nest peptide detection method according to any one of claims 1 to 9.