Method and system for purifying neuraminic acid based on multivariate linear regression analysis
The nervous acid purification process was optimized through the multivariate linear regression analysis method, and the optimal purification process combination was predicted, which solved the problem of low nervous acid purification purity and achieved high purity nervous acid extraction.
Patent Information
- Application Number
- CN202510307605.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The nervous acid purification method in the prior art results in low purity of neuric acid.
Using a method based on multivariate linear regression analysis, standard neuric acid purity data under different conditions were obtained, normality test, descriptive statistics and correlation analysis were performed to obtain the maximum correlation value and target correlation coefficient, and then the verification neuric acid purity data was preprocessed, the initial purity model was established and model training was carried out to predict the optimal purification process combination to achieve high purity extraction.
The purification purity of neuric acid is improved and the problem of low purity in the prior art is solved.
Smart Images

Figure CN119808028B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nervonic acid purification, and particularly to a nervonic acid purification method and system based on multiple linear regression analysis. Background Art
[0002] Multiple linear regression analysis is a statistical analysis method used to study the relationship between multiple independent variables and a continuous dependent variable, aiming to establish a linear model to examine the influence degree of multiple independent variables on the dependent variable and their correlation. The basic idea of multiple linear regression analysis is to use a linear equation to describe the relationship between the independent variable and the dependent variable. Among them, the dependent variable is assumed to be a linear combination of multiple independent variables plus an error term, and this error term represents the random difference that the model cannot explain.
[0003] Currently, the purification of nervonic acid is mainly achieved through technologies such as solvent extraction, chromatographic separation, crystallization purification, and reverse osmosis membranes. However, there is a problem that the purity of nervonic acid is not high due to poor purification conditions. Therefore, the current methods for purifying nervonic acid through technologies such as solvent extraction, chromatographic separation, crystallization purification, and reverse osmosis membranes have the problem of low extraction purity of nervonic acid. Summary of the Invention
[0004] The present invention provides a nervonic acid purification method and system based on multiple linear regression analysis, and its main purpose is to solve the problem of low extraction purity of nervonic acid.
[0005] To achieve the above object, a nervonic acid purification method based on multiple linear regression analysis provided by the present invention includes: obtaining standard nervonic acid purity data at different temperatures, pressures, solvent types, and reaction times, performing a normality test on the standard nervonic acid purity data to obtain the maximum correlation value; selecting a correlation analysis method based on the maximum correlation value, and using the correlation analysis method to perform a correlation analysis on the standard nervonic acid purity data to obtain the target correlation coefficient; performing descriptive statistics on the standard nervonic acid purity data to obtain the descriptive statistical results; obtaining verification nervonic acid purity data, and preprocessing the verification nervonic acid purity data based on the descriptive statistical results and the target correlation coefficient to obtain the target purity data; using a preset multiple linear regression analysis method to establish an initial purity model, and training the initial purity model according to the target purity data to obtain the target purity model; predicting the optimal purification process combination based on the target purity model, and achieving high-purity extraction of nervonic acid according to the optimal purification process combination.
[0006] Optionally, the normal distribution test on the standard nervonic acid purity data to obtain the maximum correlation value includes: obtaining the number of nervonic acid purity samples of the standard nervonic acid purity data, and comparing the size relationship between the number of nervonic acid purity samples and a preset sample threshold; if the number of nervonic acid purity samples is less than or equal to the sample threshold, then use the preset Shapiro-Wilk test to perform a normal distribution test on the standard nervonic acid purity data to obtain the maximum correlation value; if the number of nervonic acid purity samples is greater than the sample threshold, then use the preset Kolmogorov-Smirnov test to perform a normal distribution test on the standard nervonic acid purity data to obtain the maximum correlation value.
[0007] Optionally, the method of selecting a correlation analysis method based on the maximum correlation value and using the correlation analysis method to perform a correlation analysis on the standard nervonic acid purity data to obtain a target correlation coefficient includes: determining whether the maximum correlation value is greater than 0.05; if the maximum correlation value is greater than 0.05, then calculate the Pearson correlation coefficient of the standard nervonic acid purity data, and use the Pearson correlation coefficient as the target correlation coefficient; if the maximum correlation is not greater than 0.05, then calculate the Spearman correlation coefficient of the standard nervonic acid purity data, and use the Spearman correlation coefficient as the target correlation coefficient.
[0008] Optionally, the descriptive statistics on the standard nervonic acid purity data to obtain descriptive statistical results includes: identifying the mean and median of the standard nervonic acid purity data, calculating the central tendency difference of the standard nervonic acid purity data according to the mean and median; calculating the median absolute deviation of the nervonic acid purity data using the median; drawing a box plot based on the standard nervonic acid purity data; and combining the mean, central tendency difference, median absolute deviation and box plot to obtain descriptive statistical results.
[0009] Optionally, preprocessing the verified nervonic acid purity data based on the descriptive statistical results and the target correlation coefficient to obtain target purity data, including: sequentially extracting a set of verified nervonic acid purity parameters from the verified nervonic acid purity data; calculating a verification correlation coefficient and descriptive verification results according to the standard nervonic acid purity data and the set of verified nervonic acid purity parameters; respectively calculating the correlation error between the verification correlation coefficient and the target correlation coefficient and the statistical error between the descriptive verification results and the descriptive statistical results; calculating a comprehensive error according to the correlation error and the statistical error; determining whether the comprehensive error is greater than a preset error threshold; if the comprehensive error is greater than the error threshold, then return to the step of sequentially extracting the set of verified nervonic acid purity parameters from the verified nervonic acid purity data; if the comprehensive error is not greater than the error threshold, then incorporate the set of verified nervonic acid purity parameters into the standard nervonic acid purity data to obtain merged nervonic acid purity data; determining whether the verified nervonic acid purity data has completed the extraction of the set of verified nervonic acid purity parameters; if the verified nervonic acid purity data has not completed the extraction of the set of verified nervonic acid purity parameters, then return to the step of sequentially extracting the set of verified nervonic acid purity parameters from the verified nervonic acid purity data; if the verified nervonic acid purity data has completed the extraction of the set of verified nervonic acid purity parameters, then use the merged nervonic acid purity data as the target purity data.
[0010] Optionally, establishing an initial purity model using a preset multiple linear regression analysis method, including: taking temperature, pressure, solvent type, and reaction time as independent variables, and taking a preset predicted nervonic acid purity as the dependent variable; establishing an initial purity model using the multiple linear regression analysis method according to the independent variables and the dependent variable, where the initial purity model is:
[0011] , where represents the predicted nervonic acid purity, represents the intercept, represents the temperature regression coefficient, represents the temperature, represents the pressure regression coefficient, represents the pressure, represents the solvent regression coefficient, represents the solvent type, represents the time regression coefficient, represents the reaction time.
[0012] Optionally, the model training of the initial purity model according to the target purity data to obtain a target purity model includes: dividing the target purity data into k non-overlapping sub-datasets; sequentially extracting a sub-dataset from the k non-overlapping sub-datasets, using the sub-dataset as a validation set, and removing the validation set from the k non-overlapping sub-datasets to obtain a training set; substituting the training set into the initial purity model to obtain a training nervonic acid purity set; calculating the sum of squared regression and the total sum of squares using the training nervonic acid purity set and the training set; calculating the coefficient of determination according to the sum of squared regression and the total sum of squares; determining whether the coefficient of determination is less than a preset prediction error threshold; if the coefficient of determination is not less than the prediction error threshold, adjusting the parameters of the initial purity model using the coefficient of determination to obtain an iteratively adjusted purity model; updating the initial purity model using the iteratively adjusted purity model, and returning to the step of sequentially extracting a sub-dataset from the k non-overlapping sub-datasets; if the coefficient of determination is less than the prediction error threshold, calculating the mean squared error of the validation set according to the iteratively adjusted purity model; determining whether the mean squared error is less than a preset mean squared error threshold; if the mean squared error is not less than the mean squared error threshold, returning to the step of sequentially extracting a sub-dataset from the k non-overlapping sub-datasets; if the mean squared error is less than the mean squared error threshold, using the iteratively adjusted purity model as the target purity model.
[0013] Optionally, the adjusting the parameters of the initial purity model using the coefficient of determination to obtain an iteratively adjusted purity model includes: sequentially extracting preset initial coefficients in the initial purity model; setting a learning rate according to the coefficient of determination, and adjusting the initial coefficients using the learning rate to obtain adjusted coefficients; updating the initial coefficients using the adjusted coefficients to obtain an iteratively adjusted purity model.
[0014] Optionally, the predicting the optimal purification process combination based on the target purity model includes: obtaining a purification process combination set, and sequentially extracting a purification process combination from the purification process combination set; inputting the purification process combination into the target purity model to obtain a predicted product purity value set; extracting the maximum product purity value from the predicted product purity value set, and identifying the optimal purification process combination corresponding to the maximum product purity value.
[0015] To achieve the above object, the present invention also provides a nervonic acid purification system based on multiple linear regression analysis, comprising: a correlation analysis and descriptive statistics module, configured to obtain standard nervonic acid purity data at different temperatures, pressures, solvent types and reaction times, perform a normality test on the standard nervonic acid purity data to obtain a maximum correlation value; select a correlation analysis method based on the maximum correlation value, perform a correlation analysis on the standard nervonic acid purity data by using the correlation analysis method to obtain a target correlation coefficient; perform descriptive statistics on the standard nervonic acid purity data to obtain a descriptive statistics result; a verification nervonic acid purity data preprocessing module, configured to obtain verification nervonic acid purity data, preprocess the verification nervonic acid purity data based on the descriptive statistics result and the target correlation coefficient to obtain target purity data; an initial purity model training module, configured to establish an initial purity model by using a preset multiple linear regression analysis method, and perform model training on the initial purity model according to the target purity data to obtain a target purity model; an optimal purification process combination prediction module, configured to predict an optimal purification process combination based on the target purity model, and achieve high-purity extraction of nervonic acid according to the optimal purification process combination.
[0016] To solve the problems described in the background art, the present invention obtains standard nervonic acid purity data at different temperatures, pressures, solvent types and reaction times, performs a normality test, descriptive statistics and correlation analysis on the standard nervonic acid purity data to obtain a maximum correlation value, a descriptive statistics result and a target correlation coefficient. Since the standard nervonic acid purity data has the problem of small data volume, it is necessary to use the verification nervonic acid purity data to supplement the data. During the data supplement process, the verification nervonic acid purity data can be preprocessed based on the descriptive statistics result and the target correlation coefficient, and finally target purity data is obtained. An initial purity model is established by using a preset multiple linear regression analysis method. At this time, the target purity data can be used to perform model training on the initial purity model to obtain a target purity model. Finally, an optimal purification process combination is predicted based on the target purity model to achieve high-purity extraction of nervonic acid. Therefore, the present invention can solve the problem of low purity in the current extraction of nervonic acid. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flow chart of a nervonic acid purification method based on multiple linear regression analysis provided by an embodiment of the present invention.
[0018] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] An embodiment of the present application provides a method for purifying nervonic acid based on multiple linear regression analysis. The execution subject of the method for purifying nervonic acid based on multiple linear regression analysis includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for purifying nervonic acid based on multiple linear regression analysis can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.
[0021] Referring to Figure 1 As shown, it is a schematic flowchart of a method for purifying nervonic acid based on multiple linear regression analysis provided by an embodiment of the present invention. In this embodiment, the method for purifying nervonic acid based on multiple linear regression analysis includes: S1. Obtain standard nervonic acid purity data at different temperatures, pressures, solvent types, and reaction times, and perform a normality test on the standard nervonic acid purity data to obtain the maximum correlation value.
[0022] It can be explained that the solvent type refers to the type of the purification solvent during the process of purifying nervonic acid, such as: n-hexane, ethyl acetate, petroleum ether, acetone, methanol, ethanol, and (supercritical fluid solvent), etc. The temperature, pressure, and reaction time respectively refer to the ultrasonic temperature, ultrasonic pressure, and ultrasonic reaction time during the process of purifying nervonic acid. The standard nervonic acid purity data refers to the standard data of the corresponding relationship between different ultrasonic temperatures, ultrasonic pressures, solvent types, and ultrasonic reaction times and the nervonic acid purity. The standard data is nervonic acid purification data that has been manually screened and has no abnormal, missing, or incorrect values. For example: when the ultrasonic temperature is 90 , the ultrasonic pressure is 27 Mpa, the solvent type is ethyl acetate, and the ultrasonic reaction time is 6 h, the nervonic acid purity can be 23%; when the ultrasonic temperature is 79 , the ultrasonic pressure is 25 Mpa, the solvent type is ethyl acetate, and the ultrasonic reaction time is 7 h, the nervonic acid purity can be 25%, and so on.
[0023] Furthermore, the normality test refers to testing whether the standard nervonic acid purity data conforms to a normal distribution, and the maximum correlation value refers to the value for determining whether the standard nervonic acid purity data conforms to a normal distribution.
[0024] In the embodiment of the present invention, the normal distribution test on the standard nervonic acid purity data to obtain the maximum correlation value includes: obtaining the number of nervonic acid purity samples of the standard nervonic acid purity data, and comparing the size relationship between the number of nervonic acid purity samples and a preset sample threshold; if the number of nervonic acid purity samples is less than or equal to the sample threshold, then use the preset Shapiro-Wilk test to perform a normal distribution test on the standard nervonic acid purity data to obtain the maximum correlation value; if the number of nervonic acid purity samples is greater than the sample threshold, then use the preset Kolmogorov-Smirnov test to perform a normal distribution test on the standard nervonic acid purity data to obtain the maximum correlation value.
[0025] It can be understood that the number of nervonic acid purity samples refers to the number of samples in the standard nervonic acid purity data. The Shapiro-Wilk test is applicable to small sample data, and the Kolmogorov-Smirnov test is applicable to large sample data. Therefore, the size of the sample data of the standard nervonic acid purity data can be judged by the sample threshold. When the number of nervonic acid purity samples is greater than the sample threshold, it indicates that the standard nervonic acid purity data is large sample data, and the Kolmogorov-Smirnov test can be used to perform a normal distribution test on the standard nervonic acid purity data. When the number of nervonic acid purity samples is less than or equal to the sample threshold, it indicates that the standard nervonic acid purity data is small sample data, and the Shapiro-Wilk test can be used to perform a normal distribution test on the standard nervonic acid purity data. The sample threshold can be 50. The Shapiro-Wilk test and the Kolmogorov-Smirnov test are prior arts and will not be elaborated here.
[0026] It can be understood that in order to perform a normal distribution test, the standard nervonic acid purity data can be serially arranged according to the ultrasonic temperature, ultrasonic pressure, solvent type, and ultrasonic reaction time, so as to obtain a serial arrangement combination. The solvent type can be represented by numbers. For example, n-hexane is represented as 1, ethyl acetate is represented as 2, petroleum ether is represented as 3, etc. When the interval of the ultrasonic reaction time is [6h, 8h], the serial arrangement combination can be 80 、20Mpa、1、6h - 20%; 80 、20Mpa、1、6.1h - 21%; 80 、20Mpa、1、6.3h - 22%; 80 、20Mpa、1、6.4h - 23%; 80 、20Mpa、1、6.5h - 24%; 80 , 20 Mpa, 1, 6.9 h - 25%, etc., and the corresponding permutation and combination numbers are 1, 2, 3, 4, 5, 6, etc. Among them, 80 , 20 Mpa, 1, 6 h - 20% means that when the ultrasonic temperature is 80 , the ultrasonic pressure is 20 Mpa, the solvent type is n - hexane, and the ultrasonic reaction time is 6 h, the purity of nervonic acid is 20%. The permutation and combination can be determined in the form of a tree diagram.
[0027] S2. Select the correlation analysis method based on the maximum correlation value, and use the correlation analysis method to perform correlation analysis on the standard nervonic acid purity data to obtain the target correlation coefficient.
[0028] It can be understood that the correlation analysis method refers to the analysis method for analyzing the degree of association between the nervonic acid purity and the permutation and combination number. The target correlation coefficient refers to the coefficient of correlation between the nervonic acid purity and the permutation and combination number.
[0029] In the embodiment of the present invention, the method of selecting the correlation analysis method based on the maximum correlation value, using the correlation analysis method to perform correlation analysis on the standard nervonic acid purity data to obtain the target correlation coefficient includes: judging whether the maximum correlation value is greater than 0.05; if the maximum correlation value is greater than 0.05, then calculate the Pearson correlation coefficient of the standard nervonic acid purity data, and use the Pearson correlation coefficient as the target correlation coefficient; if the maximum correlation is not greater than 0.05, then calculate the Spearman correlation coefficient of the standard nervonic acid purity data, and use the Spearman correlation coefficient as the target correlation coefficient.
[0030] It can be understood that when the maximum correlation value is greater than the significance level (0.05), it means that the standard nervonic acid purity data conforms to the normal distribution, and when the maximum correlation value is not greater than 0.05, it means that the standard nervonic acid purity data does not conform to the normal distribution.
[0031] Further, the calculation process of the Pearson correlation coefficient requires that the standard nervonic acid purity data meet prerequisite conditions such as linear relationship or normal distribution. Therefore, the Pearson correlation coefficient is applicable to processing standard nervonic acid purity data under normal distribution. The calculation process of the Spearman correlation coefficient does not make assumptions about the distribution of the standard nervonic acid purity data. Therefore, the Spearman correlation coefficient is applicable to various types of standard nervonic acid purity data. Therefore, when the maximum correlation value is greater than 0.05 (i.e., the standard nervonic acid purity data conforms to normal distribution), the Pearson correlation coefficient of the standard nervonic acid purity data can be calculated, and the Pearson correlation coefficient is used as the target correlation coefficient; when the maximum correlation value is not greater than 0.05 (i.e., the standard nervonic acid purity data does not conform to normal distribution), the Spearman correlation coefficient of the standard nervonic acid purity data can be calculated, and the Spearman correlation coefficient is used as the target correlation coefficient.
[0032] Further, the calculation formula of the Pearson correlation coefficient is as follows:
[0033] , where represents the Pearson correlation coefficient, represents the purity of the j-th nervonic acid in the standard nervonic acid purity data, represents the average purity of nervonic acid in the standard nervonic acid purity data, represents the j-th permutation and combination serial number in the standard nervonic acid purity data, represents the average of the permutation and combination serial numbers in the standard nervonic acid purity data.
[0034] Further, the calculation formulas of the Pearson correlation coefficient and the Spearman correlation coefficient are prior arts and will not be elaborated here.
[0035] S3. Conduct descriptive statistics on the standard nervonic acid purity data to obtain descriptive statistical results.
[0036] It can be understood that the descriptive statistics refer to the statistics of the basic characteristics of the standard nervonic acid purity data, and the basic characteristics of the data include: mean, central tendency difference, median absolute deviation, and box plot. The descriptive statistical results refer to the statistical results of the basic characteristics of the standard nervonic acid purity data.
[0037] In the embodiment of the present invention, the conducting descriptive statistics on the standard nervonic acid purity data to obtain descriptive statistical results includes: identifying the mean and median of the standard nervonic acid purity data, calculating the central tendency difference of the standard nervonic acid purity data according to the mean and median; calculating the median absolute deviation of the nervonic acid purity data by using the median; drawing a box plot according to the standard nervonic acid purity data; and combining the mean, central tendency difference, median absolute deviation and box plot to obtain descriptive statistical results.
[0038] It is understandable that the central tendency difference refers to the difference between the mean and the median. When the mean is greater than the median, it indicates a higher degree of right skew in the standard nervonic acid purity data. When the mean is less than the median, it indicates a higher degree of left skew in the standard nervonic acid purity data. The mean refers to the mean of the nervonic acid purity, and the median refers to the median of the nervonic acid purity.
[0039] Furthermore, the median absolute deviation refers to the absolute deviation of the nervonic acid purity, and the calculation formula is as follows:
[0040] , where represents the median absolute deviation, refers to the median value function, represents the i-th nervonic acid purity in the standard nervonic acid purity data, represents the median of the standard nervonic acid purity data, represents the symbol of the absolute value. The median value function refers to the function that extracts the median from the set of differences between all nervonic acid purities and the median.
[0041] It is understandable that the box plot (Box-plot), also known as the box-and-whisker plot, is a statistical chart that describes the dispersion of the standard nervonic acid purity data. The box plot is a prior art and will not be elaborated here.
[0042] S4. Obtain the verified nervonic acid purity data, and preprocess the verified nervonic acid purity data based on the descriptive statistical results and the target correlation coefficient to obtain the target purity data.
[0043] Furthermore, the verified nervonic acid purity data refers to the nervonic acid purification data used to expand the standard nervonic acid purity data and may have abnormal, missing, incorrect and other values. The target purity data refers to the nervonic acid purity data obtained by supplementing the standard nervonic acid purity data with the verified nervonic acid purity data.
[0044] In the embodiments of the present invention, the preprocessing of the verified nervonic acid purity data based on the descriptive statistical results and the target correlation coefficient to obtain the target purity data includes: sequentially extracting a set of verified nervonic acid purity parameters from the verified nervonic acid purity data; calculating a verification correlation coefficient and a descriptive verification result according to the standard nervonic acid purity data and the set of verified nervonic acid purity parameters; respectively calculating the correlation error between the verification correlation coefficient and the target correlation coefficient and the statistical error between the descriptive verification result and the descriptive statistical results; calculating a comprehensive error according to the correlation error and the statistical error; determining whether the comprehensive error is greater than a preset error threshold; if the comprehensive error is greater than the error threshold, then returning to the step of sequentially extracting the set of verified nervonic acid purity parameters from the verified nervonic acid purity data; if the comprehensive error is not greater than the error threshold, then incorporating the set of verified nervonic acid purity parameters into the standard nervonic acid purity data to obtain merged nervonic acid purity data; determining whether the extraction of the set of verified nervonic acid purity parameters from the verified nervonic acid purity data is completed; if the extraction of the set of verified nervonic acid purity parameters from the verified nervonic acid purity data is not completed, then returning to the step of sequentially extracting the set of verified nervonic acid purity parameters from the verified nervonic acid purity data; if the extraction of the set of verified nervonic acid purity parameters from the verified nervonic acid purity data is completed, then using the merged nervonic acid purity data as the target purity data.
[0045] It can be understood that the set of verified nervonic acid purity parameters refers to the data of a single purification test in the verified nervonic acid purity data, for example: 80 , 20 Mpa, 1, 6 h - 20%. The verification correlation coefficient refers to the target correlation coefficient of the overall data after incorporating the set of verified nervonic acid purity parameters into the standard nervonic acid purity data, and the descriptive verification result refers to the descriptive statistical result of the overall data after incorporating the set of verified nervonic acid purity parameters into the standard nervonic acid purity data. The correlation error refers to the absolute value of the difference between the verification correlation coefficient and the target correlation coefficient, and the statistical error refers to the sum of the absolute values of the differences of the corresponding data characteristics between the descriptive verification result and the descriptive statistical results, such as the absolute value of the difference in mean, the absolute value of the difference in central tendency difference, the absolute value of the difference in median absolute deviation, and the absolute values of the differences of the quartiles in the box plot and other data characteristic absolute value sums. The comprehensive error refers to the weighted sum of the correlation error and the statistical error, and the weights of the correlation error and the statistical error can be set as needed.
[0046] Further, when the comprehensive error is not greater than the error threshold, it indicates that the verified nervonic acid purity parameter set meets the data standard of the standard nervonic acid purity data. Therefore, the verified nervonic acid purity parameter set can be incorporated into the standard nervonic acid purity data to obtain the merged nervonic acid purity data after the amplification of the standard nervonic acid purity data. If the comprehensive error is greater than the error threshold, it indicates that the verified nervonic acid purity parameter set does not meet the data standard of the standard nervonic acid purity data. Therefore, the verified nervonic acid purity parameter set can be directly removed.
[0047] S5. Establish an initial purity model using a preset multiple linear regression analysis method, and perform model training on the initial purity model according to the target purity data to obtain the target purity model.
[0048] It is understandable that the multiple linear regression analysis method refers to a method for studying the linear relationship between the predicted nervonic acid purity (one dependent variable) and temperature, pressure, solvent type, reaction time (multiple independent variables). The target purity model refers to the initial purity model after model training. The initial purity model is referred to in the following embodiments.
[0049] In the embodiments of the present invention, the establishment of the initial purity model using a preset multiple linear regression analysis method includes: using temperature, pressure, solvent type, and reaction time as independent variables, and using a preset predicted nervonic acid purity as the dependent variable; establishing an initial purity model using the multiple linear regression analysis method according to the independent variables and the dependent variable, where the initial purity model is:
[0050] , where represents the predicted nervonic acid purity, represents the intercept, represents the temperature regression coefficient, represents the temperature, represents the pressure regression coefficient, represents the pressure, represents the solvent regression coefficient, represents the solvent type, represents the time regression coefficient, represents the reaction time.
[0051] In an embodiment of the present invention, the training of the initial purity model based on the target purity data to obtain the target purity model includes: dividing the target purity data into k non-overlapping sub-datasets; sequentially extracting a sub-dataset from the k non-overlapping sub-datasets, using the sub-dataset as a validation set, and removing the validation set from the k non-overlapping sub-datasets to obtain a training set; substituting the training set into the initial purity model to obtain a training nervonic acid purity set; calculating the regression sum of squares and the total sum of squares using the training nervonic acid purity set and the training set; calculating the determination coefficient according to the regression sum of squares and the total sum of squares; determining whether the determination coefficient is less than a preset prediction error threshold; if the determination coefficient is not less than the prediction error threshold, adjusting the parameters of the initial purity model using the determination coefficient to obtain an iterative adjustment purity model; updating the initial purity model using the iterative adjustment purity model, and returning to the step of sequentially extracting a sub-dataset from the k non-overlapping sub-datasets; if the determination coefficient is less than the prediction error threshold, calculating the mean square error of the validation set according to the iterative adjustment purity model; determining whether the mean square error is less than a preset mean square error threshold; if the mean square error is not less than the mean square error threshold, returning to the step of sequentially extracting a sub-dataset from the k non-overlapping sub-datasets; if the mean square error is less than the mean square error threshold, using the iterative adjustment purity model as the target purity model.
[0052] It is understandable that the non-overlapping sub-datasets refer to data sets that do not have duplicate data with each other. The partitioning method of the target purity data can refer to the k-fold cross-validation method. The training nervonic acid purity set refers to the set of predicted nervonic acid purities obtained by inputting the multiple independent variables of each test in the training set into the initial purity model. For example, when the training set is 80 , 20 Mpa, 1, 6 h - 20%; 80 , 20 Mpa, 1, 6.1 h - 21%; 80 , 20 Mpa, 1, 6.3 h - 22%; 80 , 20 Mpa, 1, 6.4 h - 23%; 80 , 20 Mpa, 1, 6.5 h - 24%; 80 , 20 Mpa, 1, 6.9 h - 25%, the multiple independent variables of each test are: 80 , 20 Mpa, 1, 6 h; 80 , 20 Mpa, 1, 6.1 h; 80 , 20 Mpa, 1, 6.3 h; 80 , 20 Mpa, 1, 6.4 h; 80 , 20 Mpa, 1, 6.5 h; 80 , 20 Mpa, 1, 6.9 h. After sequentially inputting the above multiple independent variables into the initial purity model, the obtained training nervonic acid purity set may be: 18%, 21%, 23%, 25%, 27%, 28%. The regression sum of squares refers to the sum of the squares of the differences between the regression values of the dependent variable (the predicted nervonic acid purity in the training nervonic acid purity set) and the average value of the dependent variable (the mean of the actual nervonic acid purity in the training set). The total sum of squares refers to the sum of the squared differences between the actual nervonic acid purity in the training set and the mean of the actual nervonic acid purity. For example, when the training nervonic acid purity set is 18%, 21%, 23%, 25% and the corresponding actual nervonic acid purity set in the training set is 20%, 21%, 22%, 23%, at this time the mean of the actual nervonic acid purity in the training set is 21.5%, and the regression sum of squares is
[0053] , and the total sum of squares is
[0054] .
[0055] Furthermore, the coefficient of determination refers to the absolute value of the difference between the regression sum of squares and the total sum of squares. The larger the coefficient of determination, the worse the prediction effect of the initial purity model; the smaller the coefficient of determination, the better the prediction effect of the initial purity model. The mean squared error refers to the average of the squares of the differences between the actual nervonic acid purity and the corresponding predicted nervonic acid purity in the validation set. For example, when the actual nervonic acid purity in the validation set is 18%, 21%, 23%, 25% and the predicted nervonic acid purity is 19%, 20%, 22%, 24%, then the mean squared error is .
[0056] In the embodiments of the present invention, adjusting the parameters of the initial purity model by using the coefficient of determination to obtain an iteratively adjusted purity model includes: sequentially extracting the preset initial coefficients in the initial purity model; setting a learning rate according to the coefficient of determination, and adjusting the initial coefficients by using the learning rate to obtain adjusted coefficients; updating the initial coefficients by using the adjusted coefficients to obtain an iteratively adjusted purity model.
[0057] It can be understood that the initial coefficients refer to the coefficients in the initial purity model, such as: intercept, temperature regression coefficient, pressure regression coefficient, solvent regression coefficient, time regression coefficient. The learning rate refers to the speed and direction of updating the initial coefficients during the parameter adjustment process of the initial purity model. If the learning rate is set too large, the optimal solution may be skipped in each parameter adjustment process. If the learning rate is too small, the speed at which the initial coefficients converge to the optimal solution will be very slow. Therefore, when the coefficient of determination is large, the learning rate should also increase accordingly; when the coefficient of determination is small, the learning rate should also decrease accordingly.
[0058] Further, the adjustment process of the initial coefficients can be adjusted according to the gradient descent algorithm in combination with the learning rate, which can be a key hyperparameter for controlling the step size of each update of the initial coefficients in the gradient descent algorithm.
[0059] S6. Predict the optimal purification process combination based on the target purity model, and realize the high-purity extraction of nervonic acid according to the optimal purification process combination.
[0060] Further, the optimal purification process combination refers to the optimal combination of independent variables (i.e., the optimal combination of temperature, pressure, solvent type, and reaction time) predicted based on the target purity model.
[0061] In the embodiment of the present invention, predicting the optimal purification process combination based on the target purity model includes: obtaining a set of purification process combinations, and sequentially extracting purification process combinations from the set of purification process combinations; inputting the purification process combinations into the target purity model to obtain a set of predicted product purity values; extracting the maximum product purity value from the set of predicted product purity values, and identifying the optimal purification process combination corresponding to the maximum product purity value.
[0062] It is understandable that the set of predicted product purity values refers to the set of nervonic acid purities calculated after the set of purification process combinations is input into the target purity model.
[0063] To solve the problems described in the background art, the present invention obtains standard nervonic acid purity data at different temperatures, pressures, solvent types, and reaction times, performs normality tests, descriptive statistics, and correlation analysis on the standard nervonic acid purity data to obtain the maximum correlation value, descriptive statistical results, and target correlation coefficients. Since there is a problem of small data volume in the standard nervonic acid purity data, it is necessary to use the verified nervonic acid purity data to supplement the data. During the data supplementation process, the verified nervonic acid purity data can be preprocessed based on the descriptive statistical results and target correlation coefficients to finally obtain the target purity data. The initial purity model is established using the preset multiple linear regression analysis method. At this time, the target purity data can be used to train the initial purity model to obtain the target purity model. Finally, the optimal purification process combination is predicted based on the target purity model to achieve the high-purity extraction of nervonic acid. Therefore, the present invention can solve the problem of low purity in the current extraction of nervonic acid.
[0064] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for purifying nervonic acid based on multiple linear regression analysis, characterized in that: The method includes: obtaining standard neuraminic acid purity data under different temperatures, pressures, solvent types and reaction times, performing a normality test on the standard neuraminic acid purity data to obtain a maximum correlation value; selecting a correlation analysis method based on the maximum correlation value, performing a correlation analysis on the standard neuraminic acid purity data using the correlation analysis method to obtain a target correlation coefficient; performing descriptive statistics on the standard neuraminic acid purity data to obtain a descriptive statistical result; obtaining verification neuraminic acid purity data, pre-processing the verification neuraminic acid purity data based on the descriptive statistical result and the target correlation coefficient to obtain a target purity data; establishing an initial purity model using a preset multivariate linear regression analysis method, performing model training on the initial purity model according to the target purity data to obtain a target purity model; predicting an optimal purification process combination based on the target purity model, and achieving high-purity extraction of neuraminic acid according to the optimal purification process combination; the standard neuraminic acid purity data Performing a normality test to obtain a maximum correlation value includes: obtaining the number of neuraminic acid purity samples of standard neuraminic acid purity data, and comparing the size relationship between the number of neuraminic acid purity samples and a preset sample threshold; if the number of neuraminic acid purity samples is less than or equal to the sample threshold, performing a normality test on the standard neuraminic acid purity data using a preset Shapiro-Wilk test to obtain a maximum correlation value; if the number of neuraminic acid purity samples is greater than the sample threshold, performing a normality test on the standard neuraminic acid purity data using a preset Kolmogorov-Smirnov test to obtain a maximum correlation value; establishing an initial purity model using a preset multiple linear regression analysis method includes: using temperature, pressure, solvent type, and reaction time as independent variables, and using a preset predicted neuraminic acid purity as a dependent variable; establishing an initial purity model using a multiple linear regression analysis method based on the independent variables and the dependent variable, wherein the initial purity model is: ,in, Indicates the predicted purity of nervonic acid, represents the intercept, represents the temperature regression coefficient, Indicates temperature, represents the pressure regression coefficient, Indicates pressure, represents the solvent regression coefficient, Indicates the solvent type, represents the time regression coefficient, Indicates reaction time.
2. The method for purifying neuraminic acid based on multiple linear regression analysis as claimed in claim 1, characterized in that: The correlation analysis method is selected based on the maximum correlation value, and the correlation analysis method is used to perform correlation analysis on the standard neuraminic acid purity data to obtain a target correlation coefficient, including: determining whether the maximum correlation value is greater than 0.05; if the maximum correlation value is greater than 0.05, calculating the Pearson correlation coefficient of the standard neuraminic acid purity data, and using the Pearson correlation coefficient as the target correlation coefficient; if the maximum correlation is not greater than 0.05, calculating the Spearman correlation coefficient of the standard neuraminic acid purity data, and using the Spearman correlation coefficient as the target correlation coefficient.
3. The method for purifying neuraminic acid based on multiple linear regression analysis as claimed in claim 2, characterized in that: The descriptive statistics of the standard neuraminic acid purity data are performed to obtain descriptive statistical results, including: identifying the mean and median of the standard neuraminic acid purity data, and calculating the central trend difference of the standard neuraminic acid purity data based on the mean and median; using the median to calculate the median absolute deviation of the neuraminic acid purity data; drawing a box plot based on the standard neuraminic acid purity data; combining the mean, central trend difference, median absolute deviation and box plot to obtain descriptive statistical results.
4. The method for purifying neuraminic acid based on multiple linear regression analysis as claimed in claim 3, characterized in that: The method of preprocessing the verification neuraminic acid purity data based on the descriptive statistical results and the target correlation coefficient to obtain the target purity data includes: extracting the verification neuraminic acid purity parameter set in sequence from the verification neuraminic acid purity data; calculating the verification correlation coefficient and the descriptive verification result according to the standard neuraminic acid purity data and the verification neuraminic acid purity parameter set; respectively calculating the correlation error between the verification correlation coefficient and the target correlation coefficient and the statistical error between the descriptive verification result and the descriptive statistical result; calculating the comprehensive error according to the correlation error and the statistical error; judging whether the comprehensive error is greater than a preset error threshold; if the comprehensive error is greater than the error threshold, returning to the above-mentioned The steps of extracting the verification neuraminic acid purity parameter set in sequence from the verification neuraminic acid purity data; if the comprehensive error is not greater than the error threshold, merging the verification neuraminic acid purity parameter set into the standard neuraminic acid purity data to obtain merged neuraminic acid purity data; judging whether the verification neuraminic acid purity data has completed the extraction of the verification neuraminic acid purity parameter set; if the verification neuraminic acid purity data has not completed the extraction of the verification neuraminic acid purity parameter set, returning to the above-mentioned steps of extracting the verification neuraminic acid purity parameter set in sequence from the verification neuraminic acid purity data; if the verification neuraminic acid purity data has completed the extraction of the verification neuraminic acid purity parameter set, taking the merged neuraminic acid purity data as the target purity data.
5. The method for purifying neuraminic acid based on multiple linear regression analysis as claimed in claim 4, characterized in that: The method of training the initial purity model according to the target purity data to obtain the target purity model includes: dividing the target purity data into k non-overlapping sub-datasets; extracting sub-datasets in the k non-overlapping sub-datasets in turn, using the sub-datasets as validation sets, and removing the validation sets from the k non-overlapping sub-datasets to obtain a training set; substituting the training set into the initial purity model to obtain a training neuraminic acid purity set; calculating the regression sum of squares and the total sum of squares using the training neuraminic acid purity set and the training set; calculating the determination coefficient based on the regression sum of squares and the total sum of squares; judging whether the determination coefficient is less than a preset prediction error threshold; if the determination coefficient is not less than the prediction error threshold, then The determination coefficient is used to adjust the parameters of the initial purity model to obtain an iteratively adjusted purity model; the iteratively adjusted purity model is used to update the initial purity model, and the step of sequentially extracting sub-datasets from the k non-overlapping sub-datasets is returned to; if the determination coefficient is less than the prediction error threshold, the mean square error of the validation set is calculated according to the iteratively adjusted purity model; it is determined whether the mean square error is less than a preset mean square error threshold; if the mean square error is not less than the mean square error threshold, the step of sequentially extracting sub-datasets from the k non-overlapping sub-datasets is returned to; if the mean square error is less than the mean square error threshold, the iteratively adjusted purity model is used as the target purity model.
6. The method for purifying neuraminic acid based on multiple linear regression analysis according to claim 5, characterized in that: The method of using the determined coefficient to adjust the parameters of the initial purity model to obtain the iteratively adjusted purity model includes: extracting the initial coefficients preset in the initial purity model in sequence; setting a learning rate according to the determined coefficient, adjusting the initial coefficient using the learning rate to obtain the adjusted coefficient; and updating the initial coefficient using the adjusted coefficient to obtain the iteratively adjusted purity model.
7. The method for purifying neuraminic acid based on multiple linear regression analysis according to claim 6, characterized in that: The method of predicting the best purification process combination based on the target purity model includes: obtaining a purification process combination set, and extracting purification process combinations in the purification process combination set in sequence; inputting the purification process combination into the target purity model to obtain a predicted product purity value set; extracting the maximum product purity value in the predicted product purity value set, and identifying the best purification process combination corresponding to the maximum product purity value.
8. A neuraminic acid purification system based on multivariate linear regression analysis, characterized in that: The system includes: a correlation analysis and descriptive statistics module, which is used to obtain standard neuraminic acid purity data under different temperatures, pressures, solvent types and reaction times, perform a normality test on the standard neuraminic acid purity data to obtain a maximum correlation value; select a correlation analysis method based on the maximum correlation value, and use the correlation analysis method to perform a correlation analysis on the standard neuraminic acid purity data to obtain a target correlation coefficient; perform descriptive statistics on the standard neuraminic acid purity data to obtain a descriptive statistical result; the normality test on the standard neuraminic acid purity data to obtain a maximum correlation value includes: obtaining the number of neuraminic acid purity samples of the standard neuraminic acid purity data, and comparing the size relationship between the number of neuraminic acid purity samples and a preset sample threshold; if the number of neuraminic acid purity samples is less than or equal to the sample threshold, then use a preset Shapiro-Wilk test to perform a normality test on the standard neuraminic acid purity data to obtain a maximum correlation value; if the neuraminic acid purity If the number of samples is greater than the sample threshold, the preset Kolmogorov-Smirnov test is used to perform a normality test on the standard neuraminic acid purity data to obtain a maximum correlation value; a neuraminic acid purity data preprocessing module is used to obtain the neuraminic acid purity data, and preprocess the neuraminic acid purity data based on the descriptive statistical results and the target correlation coefficient to obtain the target purity data; an initial purity model training module is used to establish an initial purity model using a preset multiple linear regression analysis method, and train the initial purity model according to the target purity data to obtain a target purity model; the initial purity model is established using the preset multiple linear regression analysis method, including: taking temperature, pressure, solvent type, and reaction time as independent variables, and taking the preset predicted neuraminic acid purity as the dependent variable; according to the independent variables and the dependent variable, the initial purity model is established using the multiple linear regression analysis method, wherein the initial purity model is: ,in, Indicates the predicted purity of nervonic acid, represents the intercept, represents the temperature regression coefficient, Indicates temperature, represents the pressure regression coefficient, Indicates pressure, represents the solvent regression coefficient, Indicates the solvent type, represents the time regression coefficient, Represents reaction time; an optimal purification process combination prediction module is used to predict the optimal purification process combination based on the target purity model, and achieve high-purity extraction of neuraminic acid according to the optimal purification process combination.
Citation Information
Patent Citations
Multiple Linear Regression-Artificial Neural Network Hybrid Model Predicting Octanol-Water Partition Coefficient of Pure Organic Compound
KR1020120085143A