Method for predicting optimal extraction condition of snail bioactive peptide by utilizing machine learning

By using machine learning to predict the optimal extraction conditions for snail active peptides, the high cost and long cycle problems caused by reliance on manual experience in traditional methods were solved, and a steady increase in the yield of snail active peptides and intelligent production were achieved.

CN120673913APending Publication Date: 2025-09-19WORLD UNION BIOENGINEERING WUXI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510750051.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies rely on manual experience in the extraction process of snail active peptides, resulting in high trial and error costs, long cycles, and difficulty in global optimization. They are unable to handle the nonlinear interactions of multiple factors and lack the ability to dynamically adjust driven by real-time data.

Method used

Using machine learning methods, a standardized database was constructed by collecting and processing snail mucus pretreatment, enzymatic hydrolysis and post-processing data. A hybrid machine learning model was used to predict the yield and molecular weight distribution of active peptides. The optimal parameter combination was generated through a multi-objective optimization algorithm, and real-time verification and dynamic updates were carried out in combination with a digital twin system.

Benefits of technology

The yield of snail active peptides has been steadily increased to over 92%, supporting real-time adjustment and monitoring of process parameters and improving the level of intelligent production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673913A_ABST
    Figure CN120673913A_ABST
Patent Text Reader

Abstract

The invention provides a method for predicting an optimal extraction condition of snail bioactive peptide by utilizing machine learning, and belongs to the technical field of biology. Comprising the following steps: S1, collecting snail mucus pretreatment parameters, enzymolysis process parameters and post-treatment data, and constructing a standardized database; s2, performing missing value filling, standardization processing and feature interaction item construction on the data to generate a high-value feature set; s3, predicting the active peptide yield and molecular weight distribution by adopting a mixed machine learning model, and calculating a comprehensive quality score; and S4, generating a Pareto optimal parameter combination based on a multi-objective optimization algorithm, and outputting recommended values of the temperature, the pH value and the reaction time. According to the method, the multi-parameter nonlinear relation is disclosed through machine learning modeling, the number of experiments is reduced, the extraction condition is dynamically optimized, the yield of the bioactive peptide is stabilized to be 92% or above, real-time adjustment and remote monitoring of process parameters are supported, and the production intelligence level is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to a method for predicting optimal extraction conditions of snail active peptides by using machine learning. Background Art

[0002] The extraction efficiency and quality of snail active peptides are affected by multiple parameters such as pretreatment, enzymatic hydrolysis and post-treatment. Traditional methods rely on manual experience to adjust process conditions, which has problems such as high trial and error costs, long cycles, and difficulty in global optimization. In the prior art, although some studies have used the response surface method to optimize a single parameter, it is unable to handle the nonlinear interactions of multiple factors and lacks the ability to dynamically adjust based on real-time data. Therefore, this application provides a method for using machine learning to predict the optimal extraction conditions for snail active peptides to meet the needs. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method for predicting the optimal extraction conditions of snail active peptides using machine learning to solve the existing problems.

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0005] A method for predicting optimal extraction conditions of snail bioactive peptides using machine learning, comprising the following steps:

[0006] S1. Collect snail mucus pretreatment parameters, enzymatic hydrolysis process parameters and post-processing data to build a standardized database;

[0007] S2. Fill missing values, standardize data, and construct feature interaction terms to generate a high-value feature set;

[0008] S3. Use a hybrid machine learning model to predict the yield and molecular weight distribution of active peptides and calculate the comprehensive quality score;

[0009] S4. Generate Pareto optimal parameter combinations based on a multi-objective optimization algorithm and output recommended values ​​for temperature, pH value, and reaction time;

[0010] S5. Verify the optimization results through the digital twin system and dynamically update the model parameters.

[0011] Preferably, the S1 specifically includes:

[0012] S101, collecting mucus pretreatment parameters in real time through sensors, including ethanol concentration, ultrasonic power, and degreasing time;

[0013] S102, recording enzymatic hydrolysis process parameters, including enzyme type, temperature, pH value, and reaction time;

[0014] S103. Post-test processing data, including ultrafiltration membrane molecular weight cut-off, freeze-drying temperature and active peptide purity.

[0015] Preferably, the feature engineering in S2 includes:

[0016] S201, missing value processing: use K nearest neighbor algorithm to fill in abnormal data;

[0017] S202, characteristic structure:

[0018] Calculate the temperature-time integral effect: TTI = ∫(T(t)-T0)dt, where T0 = 50°C;

[0019] Construct enzyme concentration-pH interaction factor: CPH = [E] × |pH-8.0|;

[0020] S203, Feature Selection: Using XGBoost feature importance scoring, retain features with importance ≥ 0.85.

[0021] Preferably, the model training in S3 includes:

[0022] S301, Random Forest Model: Predicting discrete variables;

[0023] S302, LSTM neural network: modeling enzymatic reaction kinetics time series data (time window = 10 minutes);

[0024] S303, XGBoost regression model: predict the yield of active peptides, the objective function is:

[0025]

[0026] Where λ = 0.1 and tree depth = 8.

[0027] Preferably, the multi-objective optimization algorithm in S4 is:

[0028] S401. Establishing the objective function:

[0029]

[0030] S402, use NSGA-II algorithm to search Pareto frontier;

[0031] S403. Output the optimal solution set: enzymatic hydrolysis temperature 52±0.5°C, pH 8.2±0.1, reaction time 4.5±0.2h.

[0032] Preferably, the digital twin system in S5 includes:

[0033] Virtual mapping of process parameters: real-time synchronization of sensor data from the enzymatic reactor;

[0034] Abnormal working condition simulation: inject noise data to test the robustness of the model;

[0035] Dynamic calibration module: Compare the predicted value with the measured value every 24 hours, and trigger model retraining when the error is greater than 5%.

[0036] Compared with the prior art, the present invention has at least the following beneficial effects:

[0037] In the above scheme, machine learning modeling is used to reveal the nonlinear relationship between multiple parameters, reduce the number of experiments, and dynamically optimize the extraction conditions to stabilize the yield of active peptides at above 92%. It also supports real-time adjustment and remote monitoring of process parameters, thereby improving the level of intelligent production. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings, which are incorporated herein and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and to enable one skilled in the relevant art to make and use the present disclosure.

[0039] Figure 1 Flow chart of the method of the present invention.

[0040] As shown in the figure, in order to clearly implement the structure of the embodiment of the present invention, specific structures and devices are marked in the figure, but this is only for illustrative purposes and is not intended to limit the present invention to the specific structure, device and environment. According to specific needs, ordinary technicians in this field can adjust or modify these devices and environments, and the adjustments or modifications made are still included in the scope of the appended claims. DETAILED DESCRIPTION

[0041] The following is a detailed description of a method for predicting the optimal extraction conditions of snail active peptides using machine learning, provided by the present invention, in conjunction with the accompanying drawings and specific examples. It is also noted that, in order to make the examples more detailed, the following embodiments are listed as best and preferred embodiments, and those skilled in the art may also adopt other alternative methods for implementation; and the accompanying drawings are only for the purpose of describing the embodiments in more detail, and are not intended to specifically limit the present invention.

[0042] It should be noted that references in the specification to "one embodiment," "an embodiment," "an exemplary embodiment," "some embodiments," etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not every embodiment necessarily includes such specific features, structures, or characteristics. In addition, when specific features, structures, or characteristics are described in conjunction with an embodiment, it is within the knowledge of those skilled in the relevant art to implement such features, structures, or characteristics in conjunction with other embodiments (whether or not explicitly described).

[0043] In general, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending at least in part on the context, allow for the presence of other factors that are not necessarily explicitly described.

[0044] It will be understood that the meanings of “on,” “over,” and “above” in this disclosure should be interpreted in the broadest manner, such that “on” means not only “directly on” something, but also includes being “on” something with intervening features or layers, and “on” or “over” means not only “on” or “above” something, but also includes being “on” or “above” something with no intervening features or layers.

[0045] Additionally, spatially relative terms such as "below," "beneath," "lower," "above," "upper," and the like may be used herein for descriptive convenience to describe the relationship of one element or feature to another element or features, as illustrated in the accompanying drawings. Spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the accompanying drawings. The device may be oriented in other ways, and the spatially relative descriptors used herein should be similarly interpreted accordingly.

[0046] like Figure 1 As shown, an embodiment of the present invention provides a method for predicting the optimal extraction conditions of snail active peptides using machine learning, comprising the following steps:

[0047] S1. Collect snail mucus pretreatment parameters, enzymatic hydrolysis process parameters and post-processing data to build a standardized database;

[0048] S2. Fill missing values, standardize data, and construct feature interaction terms to generate a high-value feature set;

[0049] S3. Use a hybrid machine learning model to predict the yield and molecular weight distribution of active peptides and calculate the comprehensive quality score;

[0050] S4. Generate Pareto optimal parameter combinations based on a multi-objective optimization algorithm and output recommended values ​​for temperature, pH value, and reaction time;

[0051] S5. Verify the optimization results through the digital twin system and dynamically update the model parameters.

[0052] It should be further explained that, in this embodiment, S1 specifically includes:

[0053] S101, collecting mucus pretreatment parameters in real time through sensors, including ethanol concentration (60%-90%), ultrasonic power (200W-500W), and degreasing time (1h-3h);

[0054] S102, recording the enzymatic hydrolysis process parameters, including enzyme type (alkaline protease, trypsin), temperature (40°C-60°C), pH value (7.0-9.0), and reaction time (2h-6h);

[0055] S103. Post-test processing data, including ultrafiltration membrane molecular weight cut-off (1 kDa-10 kDa), freeze-drying temperature (-40°C to -50°C) and active peptide purity (≥95%).

[0056] It should be further explained in this embodiment that the feature engineering in S2 includes:

[0057] S201, missing value processing: use K nearest neighbor algorithm (K=5) to fill in abnormal data;

[0058] S202, characteristic structure:

[0059] Calculate the temperature-time integral effect: TTI = ∫(T(t)-T0)dt, where T0 = 50°C;

[0060] Construct enzyme concentration-pH interaction factor: CPH = [E] × |pH-8.0|;

[0061] S203, Feature Selection: Using XGBoost feature importance scoring, retain features with importance ≥ 0.85.

[0062] It should be further explained in this embodiment that the model training in S3 includes:

[0063] S301, Random Forest Model: Predict discrete variables (enzyme type selection, membrane material adaptability);

[0064] S302, LSTM neural network: modeling enzymatic reaction kinetics time series data (time window = 10 minutes);

[0065] S303, XGBoost regression model: predict the yield of active peptides (output range 0%-100%), the objective function is:

[0066]

[0067] Where λ = 0.1 and tree depth = 8.

[0068] It should be further explained in this embodiment that the multi-objective optimization algorithm in S4 is:

[0069] S401. Establishing the objective function:

[0070]

[0071] S402, using the NSGA-II algorithm (population size = 200, crossover probability = 0.9) to search the Pareto front;

[0072] S403. Output the optimal solution set: enzymatic hydrolysis temperature 52±0.5°C, pH 8.2±0.1, reaction time 4.5±0.2h.

[0073] It should be further explained in this embodiment that the digital twin system in S5 includes:

[0074] Virtual mapping of process parameters: real-time synchronization of sensor data from the enzymatic reactor;

[0075] Abnormal working condition simulation: inject noise data to test the robustness of the model;

[0076] Dynamic calibration module: Compare the predicted value with the measured value every 24 hours, and trigger model retraining when the error is greater than 5%.

[0077] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.

[0078] Those skilled in the art will understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc.

[0079] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for predicting the optimal extraction conditions of snail active peptides using machine learning, characterized in that: The following steps are involved: S1. Collect snail mucus pretreatment parameters, enzymatic hydrolysis process parameters and post-processing data to build a standardized database; S2. Fill missing values, standardize data, and construct feature interaction terms to generate a high-value feature set; S3. Use a hybrid machine learning model to predict the yield and molecular weight distribution of active peptides and calculate the comprehensive quality score; S4. Generate Pareto optimal parameter combinations based on a multi-objective optimization algorithm and output recommended values ​​for temperature, pH value, and reaction time; S5. Verify the optimization results through the digital twin system and dynamically update the model parameters.

2. A method for predicting the optimal extraction conditions of snail active peptides using machine learning according to claim 1, characterized in that: Said S1 specifically includes: S101, collecting mucus pretreatment parameters in real time through sensors, including ethanol concentration, ultrasonic power, and degreasing time; S102, recording enzymatic hydrolysis process parameters, including enzyme type, temperature, pH value, and reaction time; S103. Post-test processing data, including ultrafiltration membrane molecular weight cut-off, freeze-drying temperature and active peptide purity.

3. A method for predicting the optimal extraction conditions of snail active peptides using machine learning according to claim 1, characterized in that: The feature engineering in S2 includes: S201, missing value processing: use K nearest neighbor algorithm to fill in abnormal data; S202, characteristic structure: Calculate the temperature-time integral effect: TTI = ∫(T(t)-T0)dt, where T0 = 50°C; Construct enzyme concentration-pH interaction factor: CPH = [E] × |pH-8.0|; S203, Feature Selection: Using XGBoost feature importance scoring, retain features with importance ≥ 0.

85.

4. A method for predicting the optimal extraction conditions of snail active peptides using machine learning according to claim 1, characterized in that: The model training in S3 includes: S301, Random Forest Model: Predicting discrete variables; S302, LSTM neural network: modeling enzymatic reaction kinetics time series data (time window = 10 minutes); S303, XGBoost regression model: predict the yield of active peptides, the objective function is: Where λ = 0.1 and tree depth = 8.

5. A method for predicting the optimal extraction conditions of snail active peptides using machine learning according to claim 1, characterized in that: The multi-objective optimization algorithm in S4 is: S401. Establishing the objective function: S402, use NSGA-II algorithm to search Pareto frontier; S403. Output the optimal solution set: enzymatic hydrolysis temperature 52±0.5°C, pH 8.2±0.1, reaction time 4.5±0.2h.

6. A method for predicting the optimal extraction conditions of snail active peptides using machine learning according to claim 1, characterized in that: The digital twin system in S5 includes: Virtual mapping of process parameters: real-time synchronization of sensor data from the enzymatic reactor; Abnormal working condition simulation: inject noise data to test the robustness of the model; Dynamic calibration module: Compare the predicted value with the measured value every 24 hours, and trigger model retraining when the error is greater than 5%.