AI-assisted protein purification result analysis method and system
By using AI-assisted methods to analyze protein purification results and combining multimodal data fusion technology, the problem of insufficient adaptability of existing systems is solved, and the protein purification process is automated and the results are standardized for analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGZHOU SMART LIFESCI CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-28
AI Technical Summary
Existing protein purification systems lack adaptive capabilities and cannot achieve cross-modal data fusion, resulting in scattered evaluation of purification effects, long analysis cycles, and results that rely on subjective human judgment.
AI-assisted methods are used to collect and preprocess chromatograms, electrophoresis images, and mass spectrometry peak tables. Electrophoresis bands and chromatographic peaks are identified through convolutional neural networks and autoregressive models. Multimodal data fusion is performed using deep learning to generate a comprehensive analysis report.
It enables automated result analysis of the protein purification process, improves analytical accuracy and result standardization, and supports batch comparison and process trend analysis.
Smart Images

Figure CN121938468A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biopharmaceutical analysis and intelligent data processing technology, and in particular to an AI-assisted method and system for analyzing protein purification results. Background Technology
[0002] Current mainstream protein purification systems are equipped with automatic integration and peak detection functions, but they still mainly rely on threshold settings for simple peak identification, and the algorithms lack adaptive capabilities. Some research institutions have attempted to use Python or MATLAB scripts for data post-processing, including baseline subtraction, peak integration, and electrophoretic grayscale statistics, but these methods are still based on traditional signal processing algorithms and have limited effectiveness in identifying complex peak shapes and overlapping signals.
[0003] Furthermore, existing systems often only analyze single data types, failing to achieve cross-modal fusion. Mass spectrometry data analysis is typically performed independently, lacking correlation with chromatographic and electrophoretic results. This leads to a fragmented purification effect evaluation process, long analysis cycles, and reliance on subjective human judgment for result interpretation, making it difficult to establish unified standards. Therefore, this proposal suggests an AI-assisted protein purification result analysis method and system to address these issues. Summary of the Invention
[0004] The purpose of this invention is to provide an AI-assisted protein purification result analysis method and system to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an AI-assisted protein purification result analysis method, the analysis method comprising the following steps: S1, to collect chromatographic data, electrophoretic image data and mass spectrometry peak table data; S2, after obtaining chromatogram data, electrophoresis image data and mass spectrometry peak table data, data preprocessing is performed; S3 uses AI to identify and analyze preprocessed data, enabling electrophoretic band recognition, chromatographic peak recognition, and mass spectrometry deconvolution. S4 achieves a comprehensive determination of protein purity and co-eluted impurities by multimodal fusion of chromatographic, gel electrophoretic, and mass spectrometry information sources. S5 automates the generation of reports and produces visual charts.
[0006] Preferably, the preprocessing of chromatogram data is chromatographic signal preprocessing, which involves using a sliding window polynomial fitting to smooth the signal points within a window, followed by baseline fitting using least squares or an asymmetric least squares algorithm to remove drift and retain the true peak area, correcting the integration error caused by baseline rise; the preprocessing of electrophoresis image data is electrophoresis image preprocessing, which involves normalizing the electrophoresis image and using rolling ball filtering and top-cap filtering algorithms to remove gel background, followed by gamma correction and contrast stretching to enhance the band areas; the preprocessing of mass spectrometry peak table data is mass spectrometry peak data preprocessing, which involves denoising the mass spectrometry peak data, using dynamic time warping or mass-to-charge ratio matching algorithms to correct instrument drift, and using local extremum methods or peak identification based on Gaussian fitting to extract peak positions and peak areas.
[0007] Preferably, AI recognition and analysis includes the following steps: S301, electrophoretic band recognition, uses a convolutional neural network combined with a sequence model to achieve automatic detection, segmentation and quantification of electrophoretic bands; S302, Chromatographic Peak Identification, treats the chromatographic signal as a time series and fits the smooth trend through an autoregressive model or an autoregressive integral moving average model, thereby automatically detecting peak regions in the residual signal; S303, mass spectrometry deconvolution, performs automatic clustering based on peak shape parameters to achieve objective classification.
[0008] Preferably, the electrophoretic band recognition first inputs a preprocessed electrophoretic grayscale image, then performs feature extraction and sequence modeling on the preprocessed electrophoretic grayscale image, and finally outputs the grayscale integral value of each lane band.
[0009] Preferably, the chromatographic peak identification first performs baseline modeling and fits the overall trend, then extracts the residual signal and performs derivative analysis and threshold screening, calculates the integral for each peak region after derivative analysis and threshold screening, and performs parameter optimization processing, and outputs the results after optimization.
[0010] Preferably, mass spectrometry deconvolution first extracts peak features, then standardizes the features and performs calculations using a clustering algorithm. After the calculations are completed, classification judgments are made and the results are optimized. Finally, the optimized results are output.
[0011] Preferably, multimodal fusion includes feature layer fusion and decision layer fusion. Feature layer fusion maps data from different modalities to the same latent feature space and performs correlation modeling through a deep learning model architecture. Decision layer fusion fuses different judgment results output by different analysis modules.
[0012] Preferably, the automatically generated analysis report includes chromatographic peak identification and integration results, electrophoretic band grayscale analysis, mass spectrometry protein matching table, protein purity and impurity ratio, and batch-to-batch consistency index. The report is exported in PDF or HTML format and includes visual charts.
[0013] An AI-assisted protein purification result analysis system is used to implement the aforementioned AI-assisted protein purification result analysis method. The analysis system consists of a data acquisition module, a preprocessing module, an AI analysis module, a multimodal data fusion module, and a result visualization and reporting module. The data acquisition module is used for acquiring chromatogram data, electrophoresis images, and mass spectrometry peak tables. The preprocessing module is used for chromatographic signal preprocessing, electrophoretic image preprocessing, and mass spectrometry peak data preprocessing. The AI analysis module is used for chromatographic peak identification, electrophoretic band identification, and mass spectrometry deconvolution. The multimodal data fusion module is used for purity assessment and impurity source analysis; The results visualization report module is used for generating analysis reports and outputting charts.
[0014] The technical effects and advantages of this invention are as follows: This invention introduces deep learning and multimodal data fusion technology to achieve comprehensive and automatic analysis of chromatographic data, electrophoresis images and mass spectrometry results, thereby realizing automated result analysis of the protein purification process. It uses artificial intelligence algorithms to identify peak structures and electrophoretic bands, realizes multimodal data fusion, improves analysis accuracy, and generates standardized, traceable, and visualized analysis reports of the data results, while supporting batch comparison and process trend analysis. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the operational implementation of the analytical method of the present invention.
[0016] Figure 2 This is a block diagram of the system configuration of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1: The present invention provides as follows Figure 1 The AI-assisted protein purification result analysis method shown includes the following steps: S1, to collect chromatographic data, electrophoretic image data and mass spectrometry peak table data; It should be noted that the chromatographic input data comes from the ultraviolet absorption or conductivity signals of the rapid protein liquid chromatography system, the electrophoresis image input data is the grayscale image of the sample protein bands acquired by a gel imaging system, and the mass spectrometry peak table input data is the peak value, intensity, m / z ratio, etc., obtained by liquid chromatography-mass spectrometry or matrix-assisted laser desorption / ionization time-of-flight mass spectrometry analysis. These different detection methods provide complementary information on the purity and composition of the protein sample. Chromatography reflects elution behavior and protein concentration, electrophoresis reflects molecular weight and impurity bands, and mass spectrometry provides molecular identification and quantification. The combined input of these three data sources provides the foundational data for the subsequent multimodal learning of the AI model. S2, after obtaining chromatogram data, electrophoresis image data and mass spectrometry peak table data, data preprocessing is performed; The preprocessing of chromatogram data is the preprocessing of chromatographic signals. After smoothing the signal points by performing polynomial fitting within the window range using a sliding window polynomial fitting, the least squares baseline fitting or the asymmetric least squares algorithm is used to remove drift and retain the true peak area, correcting the integral error caused by baseline rise. It should be noted that the Savitzky-Golay filter is used for smoothing. A sliding window polynomial fitting is employed to smooth the signal points within a window, essentially a local low-pass filter. The calculation method is as follows: Assuming a window length w = 11 and a polynomial order p = 3, the smoothed signal is... After the calculation is completed, the baseline is corrected. The least squares baseline fitting or the asymmetric least squares algorithm is used to remove the drift, and the true peak area is preserved and the integral error caused by the baseline rise is corrected. The electrophoresis image data preprocessing is to normalize the electrophoresis image and then use rolling ball filtering and top cap filtering algorithms to remove the gel background, and perform gamma correction and contrast stretching to enhance the band region. It should be noted that during the preprocessing of electrophoresis image data, the image needs to be normalized, the gray range of the electrophoresis image is standardized to [0,1] to eliminate exposure differences, the gel background is removed by rolling ball filtering or top cap filtering algorithm, and gamma correction and contrast stretching are performed to enhance the strip area. Preprocessing of mass spectrometry peak data involves denoising the mass spectrometry peak data, correcting instrument drift using dynamic time warping or mass-to-charge ratio matching algorithms, and extracting peak position and peak area using local extremum method or Gaussian fitting-based peak identification. It should be noted that the preprocessing of mass spectrometry peak table data requires noise reduction first, using wavelet transform or local polynomial regression filtering to reduce random noise, then peak alignment, using dynamic time warping or mass-to-charge ratio matching algorithms to correct instrument drift, and finally peak detection, using local extremum method or Gaussian fitting-based peak identification to extract peak position and peak area. S3 uses AI to identify and analyze preprocessed data, enabling electrophoretic band recognition, chromatographic peak recognition, and mass spectrometry deconvolution. AI recognition and analysis includes the following steps: S301, electrophoretic band recognition, uses a convolutional neural network combined with a sequence model to achieve automatic detection, segmentation and quantification of electrophoretic bands; Electrophoretic band recognition first inputs a preprocessed electrophoretic grayscale image, then extracts features from the preprocessed electrophoretic grayscale image and performs sequence modeling, and finally outputs the grayscale integral value of each lane band. It should be noted that the method uses a convolutional neural network (CNN) combined with a sequence model (LSTM) to identify and quantify the grayscale of electrophoretic bands, achieving automatic detection, segmentation, and quantification of electrophoretic bands. The first step involves inputting pre-processed electrophoretic grayscale images with uniform dimensions. The second step, a feature extraction module, employs a lightweight CNN structure to extract local grayscale gradients, edges, and band morphological features through convolutional kernels, using ReLU or Swish activation functions to avoid gradient vanishing. The third step, a sequence modeling module, unfolds the features output by the CNN along the lane direction into a sequence input and learns the grayscale variation patterns of the bands through a sequence model to identify the band position and intensity within each lane. The fourth step, the output layer, uses a sigmoid channel to generate a band segmentation mask and simultaneously outputs the grayscale integral value of each lane band. S302, Chromatographic Peak Identification, treats the chromatographic signal as a time series and fits the smooth trend through an autoregressive model or an autoregressive integral moving average model, thereby automatically detecting peak regions in the residual signal; Chromatographic peak identification first involves baseline modeling and fitting the overall trend, then extracting the residual signal and performing derivative analysis and threshold screening. After derivative analysis and threshold screening, integral calculation is performed for each peak region, and parameter optimization is carried out. After optimization, the results are output. It should be noted that the chromatographic signal is treated as a time series, and a smooth trend is fitted using an autoregressive model or an autoregressive integral moving average model, thereby automatically detecting peak regions in the residual signal. The first step is baseline modeling, assuming the signal is: (1); in These are the autoregressive coefficients. For time The observed values, Let be the order of the autoregressive model. In time The observed values, The error value is used to train the AR model using the least squares method to fit the overall trend; the second step is to extract the residual signal, calculated using the following formula: (2); That is, the difference between the observed value and the model prediction, where For time Time residuals For time The observed value at time, In time The model prediction value at time 1 is represented by the positive deviation in the residual, which indicates the peak region. The third step, derivative analysis and threshold selection, involves calculating the first derivative of the residual signal using the following formula: (3); The peak starting point is where the derivative changes from positive to negative, and the peak ending point is where the derivative changes from negative to positive; the fourth step is peak integral calculation, which calculates the integral for each peak region using the following formula: (4); Among them, peak area Corresponding protein elution amount, For signal value, The baseline value is used as the parameter; the fifth step is parameter optimization, where the AR order p is automatically selected using the AIC criterion, and non-stationary sequences are automatically switched to the ARIMA model; the sixth step is result output, which outputs the start and end times, peak height, and peak area of each peak, and automatically determines the main peak and impurity peaks, while also outputting the smoothed chromatographic curve after baseline correction and the integral result of the target protein peak. S303, mass spectrometry deconvolution, automatic clustering based on peak shape parameters to achieve objective classification; Mass spectrometry deconvolution first extracts peak features, then standardizes the features and performs calculations using a clustering algorithm. After the calculations are completed, classification judgments are made and the results are optimized. Finally, the optimized results are output. It should be noted that purified chromatograms often contain the target peak and several impurity peaks. Unsupervised learning can be used to automatically cluster peaks based on peak shape parameters, achieving objective classification. The first step is to extract peak features, where the peak symmetry formula is: (5); Peak height is Half-peak width is The peak area is The retention period is The second step is feature standardization, which uses Z-score standardization to eliminate the influence of dimensions, transforming the original feature distribution into a standard normal distribution with a mean of 0 and a standard deviation of 1. The calculation formula is as follows: (6); In the formula, For the original data points, The mean of this feature. The standard deviation of this feature. The values are standardized; after transformation, all feature data follow a distribution with a mean of 0 and a standard deviation of 1; this effectively eliminates the dimensional differences between different data sources, enabling deep models to learn feature relationships on a uniform scale when fusing multimodal signals (electrophoresis, chromatography, mass spectrometry), thereby improving the accuracy and generalization performance of protein purity analysis; the third step, when the number of peaks is large and the distribution is regular, uses K-means (partitioning) clustering, the goal of which is to divide the sample set into groups by minimizing the within-class variance. There are several clusters, aiming to make data points within the same cluster as similar as possible, and to minimize the differences between data points in different clusters. The objective function for optimization is: (7); In the formula, For the first One cluster, For the first The centroid of each cluster, For sample points, Euclidean distance is used to measure intra-class variance. When the number of peaks is small and the peak shapes vary greatly, DBSCAN (density-based) clustering algorithm is used. Its core idea is to define the cluster structure through density reachability, so that high-density regions of arbitrary shapes are identified as a class, while low-density regions are regarded as noise. Let the sample set be: (8); and define The field is: (9); like ,but As the core point, it can be defined by "density reachability" and "density connectivity," with all reachable points forming a cluster; the two key parameters in the formula have the following meanings: (Epsilon) is the neighborhood radius, and MinPts is the minimum number of points in the neighborhood; for any data point Perform calculations, if in Points within a radius greater than or equal to MinPts are called "Core Points". Points reachable from a Core Point via a neighborhood connectivity path belong to the same cluster. Points that are neither Core Points nor within the neighborhood of any Core Point are marked as "Noise". The fourth step involves classification and judgment, defining the peak with the largest area at the cluster center as the target peak, and other peaks as impurity peaks or co-eluting peaks. The fifth step is result optimization; if the chromatographic overlap between the main peak and impurity peaks is greater than 30%, it is marked as a "co-eluting risk peak". The sixth step is result output, which includes the target peak number and parameters, a list of impurity peaks, and a visualization of the peak clustering results (two-dimensional scatter plot: retention time vs. peak area). S4 achieves a comprehensive determination of protein purity and co-eluted impurities by multimodal fusion of chromatographic, gel electrophoretic, and mass spectrometry information sources. Multimodal fusion includes feature layer fusion and decision layer fusion. Feature layer fusion maps data from different modalities to the same latent feature space and performs correlation modeling through a deep learning model architecture. Decision layer fusion fuses the different judgment results output by different analysis modules. It should be noted that the first step, feature layer fusion, in the deep learning framework, allows data from different modalities (image features, signal features, spectral features) to be mapped to the same latent feature space. Correlation modeling is performed using a Transformer structure. Features are first extracted from each modality: from chromatographic features, peak area, retention time, peak width, and symmetry; from electrophoretic features, band grayscale, number of bands, and band distribution location; and from mass spectrometry features, major protein ID, peptide coverage, and relative intensity. Then, the features from each modality are concatenated into a unified vector, calculated using the following formula: (10); The data is then input into the Transformer encoder, where the multi-head attention layer learns the inter-modal dependencies and outputs fused features. The overall purification quality is characterized; by using purity labels (obtained experimentally) or expert annotations as supervision signals, the model parameters are optimized, and finally the predicted purity value of the target protein, the contribution weight of each modality, and the model self-explanation map (attention map) are output. The second step, decision-level fusion, works by addressing the fact that different analysis modules may output different judgments. Decision fusion improves the overall confidence level of the judgment. First, the prediction results from the CNN module, electrophoresis grayscale module, AR peak integral module, and clustering module are uniformly encoded. Then, one of the fusion strategies, weighted voting, is employed, setting the output probability of each model... Corresponding weight Its formula is: (11); The Bayesian fusion method calculates the posterior probability based on the assumption of independence among the modules. The formula is as follows: (12); The third step is confidence assessment. If the confidence of the fusion result is <0.7, it is marked as "awaiting manual review"; otherwise, the final judgment result is output. The fourth step involves outputting results and calculating metrics. First, the purity of the target protein is calculated using the following formula: (13); in The system calculates the integral area of the target peak; then it identifies co-eluting impurities. If multiple mass spectrometry protein signals are detected within the retention time range of the target peak, and these signals do not match the target protein sequence, or multiple bands appear in the electrophoresis bands, the system automatically identifies them as "potential co-eluting impurities." Finally, it outputs the protein purity percentage, co-eluting risk markers, various modal analysis charts (chromatographic peak diagram, electrophoresis bands, mass spectrometry peak matching diagram), and fusion confidence score. S5 automates the generation of reports and produces visual charts.
[0019] The automatically generated analysis report includes chromatographic peak identification and integration results, electrophoresis band grayscale analysis, mass spectrometry protein matching table, protein purity and impurity ratio, and batch-to-batch consistency index. The report is exported in PDF or HTML format and includes visual charts. It should be noted that the system automatically generates analysis reports including chromatographic peak identification and integration results, electrophoresis band grayscale analysis, mass spectrometry protein matching tables, protein purity and impurity ratio, and batch-to-batch consistency index. The formula for calculating the batch-to-batch consistency index, used to evaluate the consistency of purification between batches, is as follows: (14); in For the sample size, For the first Measurement values for each batch The automatically generated report, which is the average value for all batches, is exported in PDF or HTML format and includes visual charts (peak plots, band distributions, heat maps, etc.). Example 2: The present invention provides as follows Figure 2 The AI-assisted protein purification result analysis system shown is used for the AI-assisted protein purification result analysis method in Example 1. The system is characterized by comprising a data acquisition module, a preprocessing module, an AI analysis module, a multimodal data fusion module, and a result visualization and reporting module. The data acquisition module is used for acquiring chromatogram data, electrophoresis images, and mass spectrometry peak tables; The preprocessing module is used for chromatographic signal preprocessing, electrophoresis image preprocessing, and mass spectrometry peak data preprocessing; The preprocessing module is used for chromatographic signal preprocessing, electrophoresis image preprocessing, and mass spectrometry peak data preprocessing. Chromatographic signal preprocessing involves: using a sliding window polynomial fitting to smooth the signal points within a window; then using least squares baseline fitting or an asymmetric least squares algorithm to remove drift and retain the true peak area, correcting the integration error caused by baseline rise. Electrophoresis image preprocessing involves: normalizing the electrophoresis image and then using rolling ball filtering and top-cap filtering algorithms to remove the gel background, followed by gamma correction and contrast stretching to enhance the band regions. Mass spectrometry peak data preprocessing involves denoising the mass spectrometry peak data, using dynamic time warping or mass-to-charge ratio matching algorithms to correct instrument drift, and using local extremum methods or Gaussian fitting-based peak identification to extract peak positions and peak areas. The AI analysis module is used for chromatographic peak identification, electrophoretic band identification, and mass spectrometry deconvolution. The AI analysis module is used for chromatographic peak identification, electrophoretic band identification, and mass spectrometry deconvolution. It uses a convolutional neural network (CNN) combined with a sequence model (LSTM) to achieve automatic detection, segmentation, and quantification of electrophoretic bands. Then, it uses an autoregressive model or an autoregressive integral moving average model to fit the smooth trend, thereby automatically detecting peak regions in the residual signal. Through unsupervised learning, it can automatically cluster based on peak shape parameters to achieve objective classification. The multimodal data fusion module is used for purity assessment and impurity source analysis; The multimodal data fusion module enables collaborative analysis of chromatographic data, electrophoresis images, and mass spectrometry information. Through the multimodal fusion mechanism, it enhances the accuracy of target protein identification and purity assessment, achieving a more comprehensive purification quality evaluation. The results visualization report module is used for generating analysis reports and outputting charts; The results visualization report module supports the output of standardized analysis reports, including chromatographic peak tables, electrophoresis grayscale analysis results, purity assessment charts, etc., and supports audit trail functionality, which is suitable for the document management requirements of the biopharmaceutical industry.
[0020] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An AI-assisted method for analyzing protein purification results, characterized in that, The analytical method includes the following steps: S1, to collect chromatographic data, electrophoretic image data and mass spectrometry peak table data; S2, after obtaining chromatogram data, electrophoresis image data and mass spectrometry peak table data, data preprocessing is performed; S3 uses AI to identify and analyze preprocessed data, enabling electrophoretic band recognition, chromatographic peak recognition, and mass spectrometry deconvolution. S4 achieves a comprehensive determination of protein purity and co-eluted impurities by multimodal fusion of chromatographic, gel electrophoretic, and mass spectrometry information sources. S5 automates the generation of reports and produces visual charts.
2. The AI-assisted protein purification result analysis method according to claim 1, characterized in that, The preprocessing of the chromatogram data is chromatographic signal preprocessing. After smoothing the signal points by performing polynomial fitting within the window range using a sliding window polynomial fitting, the least squares baseline fitting or the asymmetric least squares algorithm is used to remove drift and retain the true peak area, correcting the integral error caused by baseline rise. The preprocessing of the electrophoresis image data involves normalizing the electrophoresis image and then using rolling ball filtering and top-cap filtering algorithms to remove the gel background, followed by gamma correction and contrast stretching to enhance the band regions. The preprocessing of the mass spectrometry peak data involves denoising the mass spectrometry peak data, correcting instrument drift using dynamic time warping or mass-to-charge ratio matching algorithms, and extracting peak positions and peak areas using local extremum methods or peak identification based on Gaussian fitting.
3. The AI-assisted protein purification result analysis method according to claim 1, characterized in that, The AI recognition and analysis includes the following steps: S301, electrophoretic band recognition, uses a convolutional neural network combined with a sequence model to achieve automatic detection, segmentation and quantification of electrophoretic bands; S302, Chromatographic Peak Identification, treats the chromatographic signal as a time series and fits the smooth trend through an autoregressive model or an autoregressive integral moving average model, thereby automatically detecting peak regions in the residual signal; S303, mass spectrometry deconvolution, performs automatic clustering based on peak shape parameters to achieve objective classification.
4. The AI-assisted protein purification result analysis method according to claim 3, characterized in that, The electrophoretic band recognition process first inputs a preprocessed electrophoretic grayscale image, then extracts features from the preprocessed electrophoretic grayscale image and performs sequence modeling, and finally outputs the grayscale integral value of each lane band.
5. The AI-assisted protein purification result analysis method according to claim 3, characterized in that, The chromatographic peak identification process first involves baseline modeling and fitting the overall trend, then extracting the residual signal and performing derivative analysis and threshold screening. After derivative analysis and threshold screening, integral calculation is performed for each peak region, and parameter optimization is carried out. After optimization, the results are output.
6. The AI-assisted protein purification result analysis method according to claim 3, characterized in that, The mass spectrometry deconvolution first extracts peak features, then standardizes the features and performs calculations using a clustering algorithm. After the calculations are completed, classification judgments are made and the results are optimized. Finally, the optimized results are output.
7. The AI-assisted protein purification result analysis method according to claim 1, characterized in that, The multimodal fusion includes feature layer fusion and decision layer fusion. Feature layer fusion maps data from different modalities to the same latent feature space and performs correlation modeling through a deep learning model architecture. Decision layer fusion fuses different judgment results output by different analysis modules.
8. The AI-assisted protein purification result analysis method according to claim 1, characterized in that, The automatically generated analysis report includes chromatographic peak identification and integration results, electrophoretic band grayscale analysis, mass spectrometry protein matching table, protein purity and impurity ratio, and batch-to-batch consistency index. The report is exported in PDF or HTML format and includes visual charts.
9. An AI-assisted protein purification result analysis system, used to implement the AI-assisted protein purification result analysis method according to any one of claims 1-9, characterized in that, The analysis system consists of a data acquisition module, a preprocessing module, an AI analysis module, a multimodal data fusion module, and a result visualization and reporting module. The data acquisition module is used for acquiring chromatogram data, electrophoresis images, and mass spectrometry peak tables. The preprocessing module is used for chromatographic signal preprocessing, electrophoretic image preprocessing, and mass spectrometry peak data preprocessing. The AI analysis module is used for chromatographic peak identification, electrophoretic band identification, and mass spectrometry deconvolution. The multimodal data fusion module is used for purity assessment and impurity source analysis; The results visualization report module is used for generating analysis reports and outputting charts.
Citation Information
Cited By
Method and system for recognizing batch-to-batch impurity difference of OLED material
CN122150477A