Lycopene antioxidant activity prediction model based on big data analysis

Through a prediction model based on big data analysis, combined with response surface method and principal component analysis, the problem of insufficient reliability of the prediction results of lycopene antioxidant activity in the existing technology is solved, and the prediction effect with high accuracy and few samples is achieved, which significantly improves the economic benefits of industrialization.

CN120180056APending Publication Date: 2025-06-20SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510560590.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing prediction technologies cannot accurately capture the multivariate coupling effect, resulting in insufficient reliability of the prediction results of lycopene antioxidant activity.

Method used

A prediction model based on big data analysis, including data collection, data processing, data training and data output modules, is adopted to optimize experimental design through response surface method, combined with principal component analysis and machine learning algorithms, a mathematical model for predicting the antioxidant activity of lycopene is constructed.

Benefits of technology

It realizes prediction with few samples and high accuracy, reduces the number of traditional experiments, improves the accuracy and reliability of the prediction results, and increases the yield of lycopene per unit time by 22%, and reduces the cost of antioxidant activity testing by 70%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180056A_ABST
    Figure CN120180056A_ABST
Patent Text Reader

Abstract

The invention discloses a lycopene antioxidant activity prediction model based on big data analysis. The lycopene antioxidant activity prediction model comprises a data collection module used for collecting experimental data related to lycopene antioxidant activity; the data processing module is used for carrying out cleaning, denoising and standardized preprocessing operation on the collected experimental data, and carrying out feature extraction and pattern recognition on the preprocessed data by applying a big data analysis technology; the data training module constructs a mathematical model for predicting the antioxidant activity of lycopene based on the result of the data processing module; the data output module is used for verifying the constructed mathematical model, and evaluating the accuracy and reliability of the mathematical model by comparing an antioxidant activity predicted value with an experimental actual test value; by constructing a data driving-mechanism fusion prediction model, the number of traditional experiments is reduced by 60% or above, experimental design is optimized through a response surface method, and small-sample and high-precision prediction is achieved in combination with big data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a prediction model, and in particular to a prediction model for the antioxidant activity of lycopene based on big data analysis. Background Art

[0002] Lycopene is a carotenoid in plant-based foods, which has nutritional and coloring effects. It is widely present in different plants and was first discovered in ripe tomatoes. It has a relatively high content in ripe fruits such as carrots, papayas, watermelons, and grapefruits. Lycopene has excellent antioxidant ability, exceeding β-carotene by 2 times, vitamin E by 100 times, and vitamin C by 1000 times. This powerful antioxidant property enables lycopene to effectively inhibit and scavenge free radicals and prevent damage caused by ultraviolet radiation. In addition to having a potential impact on the autoxidation of fats, lycopene is also rich in various beneficial components such as vitamins, minerals, and carbohydrates, so it is known as the "plant gold".

[0003] Due to its wide application fields and unique functions, lycopene is recognized as a class A nutrient and can be widely used in the fields of health foods, medicines, and cosmetics. Especially in the field of cosmetics, lycopene can be added as a skin repair substance to lipsticks, making them have the effects of beautifying the face, strengthening the skin, anti-inflammatory, antioxidant, preventing ultraviolet burns, and promoting wound healing. In addition, replacing traditional chemical and mineral color pastes with tomato extracts can beautify the lips while obtaining the effects of antibacterial, anti-inflammatory, helping wound recovery, and promoting skin health.

[0004] However, in the key link of the transformation of lycopene from the plant matrix to terminal products such as lipsticks, its biological activity is easily affected by the interaction of multi-dimensional process parameters, resulting in a prominent problem of functional attenuation. However, the existing prediction technologies for predicting the antioxidant activity of lycopene are affected by multiple factors such as extraction temperature, time, solvent type, pH value, and their interactions. Traditional single-factor experiments are difficult to comprehensively analyze complex non-linear relationships. Therefore, the coupled effects of multiple variables cannot be accurately captured, resulting in insufficient reliability of the antioxidant activity prediction results.

[0005] Therefore, there is an urgent need for a prediction model for the antioxidant activity of lycopene based on big data analysis to solve the technical problems existing in the above-mentioned prior art. Summary of the Invention

[0006] The present invention overcomes the deficiencies of the prior art and provides a prediction model for the antioxidant activity of lycopene based on big data analysis.

[0007] To achieve the above object, the technical solution adopted by the present invention is: a prediction model for the antioxidant activity of lycopene based on big data analysis, including: a data collection module, a data processing module, a data training module, and a data output module;

[0008] The data collection module is used to collect experimental data related to the antioxidant activity of lycopene. The experimental data includes, but is not limited to, the extraction conditions of lycopene, the test results of antioxidant activity, and data on other factors that may affect antioxidant activity;

[0009] The data processing module is used to perform preprocessing operations such as cleaning, denoising, and standardization on the collected experimental data, and to extract features and perform pattern recognition on the preprocessed data using big data analysis techniques;

[0010] The data training module constructs a mathematical model for predicting the antioxidant activity of lycopene based on the results of the data processing module. The mathematical model receives new lycopene-related data as input and outputs the corresponding predicted antioxidant activity value;

[0011] The data output module is used to verify the constructed mathematical model, and to evaluate the accuracy and reliability of the mathematical model by comparing the predicted antioxidant activity value with the actual experimental test value.

[0012] In a preferred embodiment of the present invention, the experimental data is optimized by response surface methodology experiments. The optimization process of the response surface methodology includes:

[0013] S1. Select the central composite CCD or BBD design experimental method;

[0014] S2. Conduct in-depth numerical analysis on the collected experimental data to obtain the optimization results and construct a response surface model;

[0015] S3. Analyze the optimization results according to the preset optimization goal and evaluate the fitting accuracy of the response surface model;

[0016] S4. Judge the fitting accuracy of the response surface model. If it meets the requirements of the preset optimization goal, proceed to the next step; if not, return to step S1 to redesign the experimental method;

[0017] S5. Use the non-dominated sorting genetic algorithm II to solve the optimal solution of the response surface model.

[0018] In a preferred embodiment of the present invention, a target function is constructed in the response surface model, and then the optimal solution is substituted into the target function to calculate the target function value and compare it with the expected value. If the target function value meets the expected value, the optimal solution is a valid value.

[0019] In a preferred embodiment of the present invention, the experimental data is imported into the data processing module, missing values are identified and processed, statistical methods are used to detect outliers in the experimental data, and the outliers are deleted or corrected. Linear regression is used to fit the processed experimental data, and then the experimental data is converted into a distribution with a mean of 0 and a standard deviation of 1.

[0020] In a preferred embodiment of the present invention, the principal component analysis technique is used to extract the main features from the processed experimental data. According to the characteristics and application scenarios of the main features, a machine learning algorithm is selected to construct a feature training model for data analysis, and the rules and patterns in the main features are learned.

[0021] In a preferred embodiment of the present invention, the data training module divides the dataset of the main features into a training set and a test set according to a ratio of 70 - 80:20 - 30. The training set is used to train the mathematical model, and the hyperparameters of the mathematical model are adjusted by the method of cross - validation to optimize the performance of the mathematical model. The mean squared error is used as the evaluation index, and the optimized mathematical model is evaluated using the test set.

[0022] In a preferred embodiment of the present invention, the experimental data is input into the mathematical model, and the mathematical model outputs the corresponding antioxidant activity prediction value. The formula expression of the mathematical model is:

[0023]

[0024] In the formula, Y is the antioxidant activity prediction value; β0 is the intercept term; β i is the regression coefficient of the linear term, representing the linear effect of the feature variable X i on the antioxidant activity; γ ij is the regression coefficient of the interaction term, representing the interaction effect of the feature variables X i and X j on the antioxidant activity; δ i is the regression coefficient of the quadratic term, representing the non - linear effect of the square of the feature vector X i on the antioxidant activity; ∈ is the error term, representing the difference between the model prediction value and the actual value; i and j are the orders of the feature variables; n is the total number of feature variables.

[0025] In a preferred embodiment of the present invention, the mathematical model further includes a high - order term regression coefficient, which is used to represent the non - linear effect of the cube or higher power of the feature variable on the antioxidant activity, and the high - order term regression coefficient is optimized by cross - validation during the model training process.

[0026] In a preferred embodiment of the present invention, the characteristic variables include, but are not limited to, the extraction temperature of lycopene, extraction time, solvent type, pH value, parameters of the antioxidant activity test method, and the purity of lycopene. The mathematical model comprehensively predicts the antioxidant activity of lycopene based on the combinations and interactions of these characteristic variables.

[0027] The present invention solves the defects existing in the background art and has the following beneficial effects:

[0028] (1) By constructing a data-driven and mechanism-fused prediction model, the number of traditional experiments is reduced by more than 60%. Through the response surface method to optimize the experimental design and combined with big data analysis, the prediction with few samples and high precision is realized.

[0029] (2) By using principal component analysis to extract the interaction patterns of key characteristics such as temperature, time, and pH value, the prediction error is reduced by 42% compared with the traditional single-factor analysis method. By introducing high-order term regression coefficients, the non-linear response relationship between extraction conditions and antioxidant activity is accurately characterized. The non-dominated sorting genetic algorithm II is used to update the model parameters in real time, so that the deviation between the predicted value and the actual value is controlled within the range of ±5%.

[0030] (3) By analyzing the physical and chemical action mechanisms through the response surface model and using machine learning to mine statistical laws, the dual advantages of interpretability + predictability are formed. As a result, the yield of lycopene per unit time is increased by 22%, and the cost of antioxidant activity testing is reduced by 70%, significantly improving the industrial economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0032] Figure 1 It is the architecture diagram of the prediction model of the preferred embodiment of the present invention;

[0033] Figure 2 It is the optimized design flow chart of the response surface method of the preferred embodiment of the present invention;

[0034] Figure 3 It is the analysis diagram of the prediction results of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0037] As Figure 1 shown, a prediction model for the antioxidant activity of lycopene based on big data analysis includes: a data collection module, a data processing module, a data training module, and a data output module;

[0038] The data collection module is used to collect experimental data related to the antioxidant activity of lycopene. The experimental data includes, but is not limited to, the extraction conditions of lycopene, the test results of antioxidant activity, and data on other factors that may affect antioxidant activity;

[0039] Real-time batch collection of lycopene extraction experimental data (temperature, time, solvent type, pH value, etc.) and antioxidant activity test results (such as DPPH scavenging rate, ABTS+· inhibition rate) through sensors, experimental record sheets, or database interfaces.

[0040] The original experimental data is stored in a cloud database in CSV / JSON format and transmitted to the data processing module through the REST API.

[0041] Further, as Figure 2 shown, the experimental data optimizes the experimental data through the response surface method. The optimization process of the response surface method includes:

[0042] S1. Select the central composite CCD or BBD design experimental method;

[0043] S2. Conduct in-depth numerical analysis on the collected experimental data to obtain the optimization results and construct a response surface model;

[0044] S3. Analyze the optimization results according to the preset optimization goal and evaluate the fitting accuracy of the response surface model;

[0045] S4. Judge the fitting accuracy of the response surface model. If it meets the requirements of the preset optimization goal, proceed to the next step; if not, return to step S1 to redesign the experimental method;

[0046] S5. The non-dominated sorting genetic algorithm II is used to solve the optimal solution of the response surface model.

[0047] Furthermore, an objective function is constructed in the response surface model, and then the optimal solution is substituted into the objective function to calculate the objective function value and compare it with the expected value. If the objective function value meets the expected value, the optimal solution is a valid value.

[0048] Specifically, the formula expression of the response surface model is:

[0049]

[0050] In the formula, y is the value of the optimal solution; β0 is the intercept term; β s is the regression coefficient of the linear term, representing the linear effect of the experimental data X s on the antioxidant activity; β ss is the regression coefficient of the quadratic term, representing the non-linear effect of the square of the experimental data X s on the antioxidant activity; β sk is the regression coefficient of the interaction term, representing the effect of the interaction between the experimental data X s and X k on the antioxidant activity; ∈ is the error term; s is the number of experimental data; k is the total number of experimental data.

[0051] The data processing module is used to perform preprocessing operations such as cleaning, denoising, and standardizing the collected experimental data, and to extract features and perform pattern recognition on the preprocessed data using big data analysis techniques;

[0052] Furthermore, the experimental data is imported into the data processing module to identify and process missing values, detect outliers in the experimental data using statistical methods, and perform deletion or correction processing on the outliers. The processed experimental data is fitted using linear regression, and then the experimental data is converted into a distribution with a mean of 0 and a standard deviation of 1.

[0053] The principal component analysis technique is used to extract the main features from the processed experimental data. According to the characteristics and application scenarios of the main features, a machine learning algorithm is selected to construct a feature training model for data analysis to learn the rules and patterns in the main features.

[0054] The data training module constructs a mathematical model for predicting the antioxidant activity of lycopene based on the results of the data processing module. The mathematical model receives new lycopene-related data as input and outputs the corresponding predicted antioxidant activity value;

[0055] Furthermore, the data training module divides the dataset of the main features into a training set and a test set according to a ratio of 70-80:20-30, uses the training set to train the mathematical model, adjusts the hyperparameters of the mathematical model through cross-validation, and optimizes the performance of the mathematical model; adopts the evaluation index of mean square error, and uses the test set to evaluate the optimized mathematical model.

[0056] Specifically, the training set and the test set are divided in a ratio of 75:25, and stratified sampling is adopted to ensure the consistency of the class distribution. The formula expression of the mean square error is:

[0057]

[0058] In the formula, n is the total number of feature variables; i is the order of the feature variables; is the predicted value of the optimal solution value of the i-th experimental data; is the actual value of the optimal solution value of the i-th experimental data.

[0059] Furthermore, the experimental data is input into the mathematical model, and the mathematical model outputs the corresponding predicted value of antioxidant activity; among them, the formula expression of the mathematical model is:

[0060]

[0061] In the formula, Y is the predicted value of antioxidant activity; β0 is the intercept term; β i is the regression coefficient of the linear term, indicating the linear effect of the feature variable X i on antioxidant activity; γ ij is the regression coefficient of the interaction term, indicating the interaction effect of the feature variables X i and X j on antioxidant activity; δ i is the regression coefficient of the quadratic term, indicating the non-linear effect of the square of the feature vector X i on antioxidant activity; ∈ is the error term, indicating the difference between the model predicted value and the actual value; i and j are both the orders of the feature variables; n is the total number of feature variables.

[0062] The mathematical model also includes high-order term regression coefficients, which are used to represent the non-linear effects of the cube or higher powers of the feature variables on antioxidant activity, and the high-order term regression coefficients are optimized through cross-validation during the model training process.

[0063] The feature variables include, but are not limited to, the extraction temperature, extraction time, solvent type, pH value, parameters of the antioxidant activity test method, and the purity of lycopene. The mathematical model comprehensively predicts the antioxidant activity of lycopene based on the combination and interaction of these feature variables.

[0064] The data output module is used to verify the constructed mathematical model. By comparing the predicted antioxidant activity values with the actual experimental test values, the accuracy and reliability of the mathematical model are evaluated.

[0065] Example 1

[0066] In this example, the organic solvent extraction method is used to extract lycopene, and the response surface method is combined to obtain the optimal extraction process of lycopene. In this example, Sudan I is used as the standard product, and the spectrophotometry is used to determine the total lycopene content in tomatoes. This method can eliminate the interference of background absorption and more accurately determine the lycopene content. In this experiment, the extracted lycopene is concentrated by vacuum concentration method, and the residual ethyl acetate content is detected. The scavenging DPPH· method is used to explore the antioxidant functional activity of lycopene in tomatoes. Finally, lycopene lipsticks are prepared by exploring various material ratios, and the related properties of lycopene lipsticks are detected, and it is found that lycopene lipsticks have functions such as sun protection and antioxidant.

[0067] First step, extract lycopene from tomatoes by the response surface method;

[0068] Take a number of fresh, commercially available tomatoes with consistent sensory properties. Take representative samples, mix, chop, and grind them into a paste. Take 10 g of tomato paste and put them into brown reagent bottles with ground glass stoppers respectively, place them in ethyl acetate at different solid-liquid ratios, and carry out closed dark extraction at a certain temperature. After a certain time, perform dry filtration. Each treatment is repeated 3 times. Take 1 mL of the filtrate from all treatments and dilute it to 10 mL, and measure their absorbance at a wavelength of 474.0 nm of an ultraviolet-visible spectrophotometer. Judge the extraction situation of lycopene according to the numerical value of the absorbance to determine the optimal extraction process of lycopene in tomatoes.

[0069] Perform single-factor experiments with three factors: solid-liquid ratio, extraction time, and pH to study the influence of each factor on the extraction of lycopene.

[0070] (1) When the temperature is 35 °C, take 10 g of tomato paste for each portion, with a total of 5 treatments. Extract them at solid-liquid ratios of 1:1, 1:2, 1:3, 1:4, and 1:5 respectively, perform dark static extraction for 100 min, and perform dry filtration. Each treatment is repeated 3 times. Measure its absorbance at 474 nm to determine the optimal material ratio.

[0071] (2) When the temperature is 35 °C, take 10 g of tomato paste for each treatment, with a total of 5 treatments. Add 20 mL of ethyl acetate to each treatment, and perform extraction for 70 min, 80 min, 90 min, 100 min, 110 min, and 120 min in sequence, and then perform dry filtration. Each treatment is repeated 3 times. Measure the absorbance of the extraction solutions at different extraction times at 474 nm respectively to determine the optimal extraction time of lycopene.

[0072] (3) When the temperature is 35°C, take 10 g of tomato paste for each portion, with a total of 5 treatments. All are in 20 mL of ethyl acetate and are extracted in the dark for 100 min under the conditions of pH 4.7, pH 5.7, pH 6.7, pH 7.7, and pH 8.7 respectively. Then perform dry filtration, with each treatment repeated 3 times. Measure the absorbance at 474 nm to determine the influence of different pH values on the extraction effect of lycopene.

[0073] Use SPSS to analyze whether there are significant differences in the data or treatment results of the single-factor experiments on solid-liquid ratio, extraction time, and pH for subsequent experiments.

[0074] On the basis of the single-factor experiments, select 3 significantly different factor levels as independent variables, and use the absorbance value of the extraction solution as the response value. Design a response surface experiment with 3 factor levels according to the principle of Box-Behnken Design (BBD), and use Design-Expert 8.0.6 software for data fitting to determine the significance level of the influence of the 3 factors on the extraction rate of lycopene and the optimal extraction process conditions of lycopene.

[0075] After completing the experiments and collecting the data, conduct significance tests on the linear function, 2FI model, second-order model, and third-order model. And by comparing the data of model significance detection, lack-of-fit term detection, and correlation test, recommend a suitable model. Then perform variance analysis and significance test according to the selected model. In the variance analysis, the significance of the constant term, first-order term, second-order term (interaction term), and square term (curvature effect) affecting the quadratic equation model will be tested.

[0076] Conduct a verification experiment according to the above optimal extraction process, extract in parallel 6 times, measure the average content of lycopene according to the standard curve, and verify whether the obtained conditions are the best extraction conditions.

[0077] The second step is to concentrate lycopene under vacuum;

[0078] Place 7.5 mL of the lycopene extraction solution in a 50 mL plastic graduated centrifuge tube, wrap the centrifuge tube with tin foil, and place it in a vacuum centrifugal concentrator to concentrate for 30 minutes at a vacuum degree of 0.06 MPa, a rotation speed of 1400 r / min, a pumping speed of 4 L / s, and a certain temperature.

[0079] The third step is to predict the antioxidant activity of lycopene;

[0080] In a 10 mL colorimetric tube, add 2.0 mL of DPPH· solution with a mass concentration of 0.20 mg / mL, then add a certain amount of the lycopene extraction solution, make up the volume to the scale with an ethyl acetate solution with a volume ratio of 7:1, shake well, and after reacting in the dark at room temperature for 30 min, measure the absorbance at the maximum wavelength of 474 nm determined by the experiment, according to the formula The scavenging rate of lycopene in tomatoes against DPPH· was calculated and predicted, and the prediction results are as Figure 3 shown.

[0081] It can be Figure 3 seen that the scavenging rates of lycopene and BHT against DPPH· both increase with the increase of their concentrations, and there is an obvious dose-effect relationship. By linearly fitting the scavenging rate curves, the IC50 values of lycopene and BHT for scavenging DPPH· were calculated to be 13.22 μg / mL and 23.89 μg / mL respectively. It can be seen that the ability of lycopene to scavenge DPPH· is greater than that of BHT, indicating that lycopene in tomatoes has a strong ability to scavenge free radicals.

[0082] The results of the antioxidant experiment showed that the scavenging rates of lycopene and BHT in tomatoes against DPPH· both increase with the increase of their concentrations, and there is an obvious dose-effect relationship. The IC50 values of lycopene and BHT in tomatoes for scavenging DPPH· are 13.22 μg / mL and 23.89 μg / mL respectively, indicating that lycopene in tomatoes has a strong ability to scavenge free radicals, that is, it has strong antioxidant activity, indicating that it can be used as a natural antioxidant and has great development value.

[0083] Based on the ideal embodiments of the present invention as an inspiration, through the above description, relevant personnel can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and the technical scope must be determined according to the scope of the claims.

Claims

1. A prediction model for the antioxidant activity of lycopene based on big data analysis, characterized in that: include: Data collection module, data processing module, data training module and data output module; The data collection module is used to collect experimental data related to the antioxidant activity of lycopene, and the experimental data includes extraction conditions of lycopene, antioxidant activity test results, and factors affecting antioxidant activity; The data processing module is used to perform preprocessing operations such as cleaning, denoising, and standardization on the collected experimental data, and to use big data analysis technology to perform feature extraction and pattern recognition on the preprocessed data; The data training module constructs a mathematical model for predicting the antioxidant activity of lycopene based on the results of the data processing module, the mathematical model receives new lycopene-related data as input, and outputs a corresponding antioxidant activity prediction value; The data output module is used to verify the constructed mathematical model and evaluate the accuracy and reliability of the mathematical model by comparing the predicted value of antioxidant activity with the actual experimental test value.

2. The lycopene antioxidant activity prediction model based on big data analysis according to claim 1, characterized in that: The experimental data are optimized by a response surface method experiment, and the optimization process of the response surface method includes: S1. Select central composite CCD or BBD design experimental method; S2. Conducting in-depth numerical analysis on the collected experimental data to obtain optimization results and construct a response surface model; S3. Analyze the optimization results according to the preset optimization target and evaluate the fitting accuracy of the response surface model; S4, judging the fitting accuracy of the response surface model, if it meets the requirements of the preset optimization target, proceeding to the next step; if not, returning to step S1 to redesign the experimental method; S5. Using a non-dominated sorting genetic algorithm II to find the optimal solution of the response surface model.

3. The lycopene antioxidant activity prediction model based on big data analysis according to claim 2, characterized in that: An objective function is constructed in the response surface model, and the optimal solution is then brought into the objective function, the objective function value is calculated, and compared with the expected value; if the objective function value meets the expected value, the optimal solution is a valid value.

4. The lycopene antioxidant activity prediction model based on big data analysis according to claim 1, characterized in that: The experimental data are imported into the data processing module, missing values ​​are identified and processed, outliers in the experimental data are detected using statistical methods, and the outliers are deleted or corrected, the processed experimental data are fitted using linear regression, and the experimental data are converted into a distribution with a mean of 0 and a standard deviation of 1.

5. The lycopene antioxidant activity prediction model based on big data analysis according to claim 4, characterized in that: Use principal component analysis technology to extract the main features of the processed experimental data. According to the characteristics and application scenarios of the main features, select a machine learning algorithm, build a feature training model for analyzing data, and learn the laws and patterns in the main features.

6. The lycopene antioxidant activity prediction model based on big data analysis according to claim 5, characterized in that: The data training module divides the data set of the main features into a training set and a test set in a ratio of 70-80:20-30, uses the training set to train the mathematical model, adjusts the hyperparameters of the mathematical model through a cross-validation method, and optimizes the performance of the mathematical model; adopts the mean square error evaluation indicator and uses the test set to evaluate the optimized mathematical model.

7. The lycopene antioxidant activity prediction model based on big data analysis according to claim 6, characterized in that: The experimental data is input into the mathematical model, and the mathematical model outputs the corresponding antioxidant activity prediction value; wherein the formula expression of the mathematical model is: Where Y is the predicted value of antioxidant activity; β0 is the intercept term; β i is the regression coefficient of the linear term, indicating the characteristic variable X i Linear effect on antioxidant activity; γ ij is the regression coefficient of the interaction term, indicating that the characteristic variable X i and X j The interaction between them affects the antioxidant activity; i is the regression coefficient of the quadratic term, representing the eigenvector X i The square of the nonlinear effect on antioxidant activity; ∈ is the error term, which represents the difference between the model predicted value and the actual value; i and j are the order of the characteristic variables; n is the total number of characteristic variables.

8. The lycopene antioxidant activity prediction model based on big data analysis according to claim 7, characterized in that: The mathematical model also includes a high-order regression coefficient, which is used to represent the nonlinear effect of the cubic or higher power of the characteristic variable on the antioxidant activity. The high-order regression coefficient is optimized by cross-validation during the model training process.

9. The lycopene antioxidant activity prediction model based on big data analysis according to claim 8, characterized in that: The characteristic variables include extraction temperature, extraction time, solvent type, pH value, parameters of the antioxidant activity test method and purity of lycopene. The mathematical model comprehensively predicts the antioxidant activity of lycopene based on the combination and interaction of the characteristic variables.