Risk assessment method and system for harmful substances in food based on big data

Through a risk assessment method based on big data, genetic optimization algorithms and risk prediction models are used, combined with historical quality inspection data, the risk of harmful substances in food is evaluated, and the problem of difficulty in risk assessment of harmful substances in food in the prior art is solved, and rapid and accurate risk assessment and reduced detection costs are achieved.

CN120164546AInactive Publication Date: 2025-06-17CHONGQING ACAD OF METROLOGY & QUALITY INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510230911.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the risks of harmful substances in food, especially in the food storage process. The detection process is cumbersome and costly, making it difficult to predict the risk of existing foods.

Method used

A risk assessment method based on big data is adopted, and a risk assessment model is constructed through genetic optimization algorithms and risk prediction models, combined with historical quality inspection data, and a risk assessment model is constructed to evaluate the risks of harmful substances in food under different storage conditions.

Benefits of technology

It improves the risk analysis ability of specific food types and additives, achieves rapid and accurate assessment of the risks of harmful substances in food, and reduces the detection cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164546A_ABST
    Figure CN120164546A_ABST
Patent Text Reader

Abstract

The invention discloses a big data-based risk assessment method and system for harmful substances in food, and the method comprises the steps: collecting to-be-detected food data and quality inspection standard data, and constructing a standard data set; time sequence features are extracted based on the standard data set and the to-be-detected food data, a risk change rate is obtained, a food quality index is determined in combination with a standard threshold interval, and an effect estimation value is obtained through a fixed effect model; and training a neural network model by using the effect estimation value, optimizing the model based on a genetic algorithm, constructing a risk assessment model, outputting an assessment result, obtaining a benchmark assessment result in combination with a standard data set, calculating a risk deviation degree, and outputting a risk level. Through the fixed effect model and the neural network model, the storage days and the time-varying coupling effect of the additive on the harmful substances are quantified, the risk change of the harmful substances in the food and the influence mode of the risk change can be effectively analyzed, and a more explanatory analysis tool is provided for food risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of risk assessment, and particularly to a method and system for risk assessment of harmful substances in food based on big data. Background Art

[0002] Food hygiene is an important foundation related to national development and people's livelihood. In recent years, with the booming development of the food industry, the content and types of various additives in food have been continuously increasing, and the maturity of the food supply chain has also enabled a large number of foods to enter the storage stage for preservation. Additives and food storage methods themselves exist to ensure food safety and freshness. However, due to the large number of near-expired foods and marginal foods that have not been sampled in the market, the actual risks of harmful substances in these foods cannot be effectively evaluated. Since the food detection process is long and the detection cost is high, the types and quantities of actually sampled foods are restricted, making it difficult to predict the risks of harmful substances in existing foods.

[0003] A method and system for risk assessment of harmful substances in food based on big data, through a genetic optimization algorithm and a risk prediction model, realizes the risk assessment of harmful substances in existing foods under different storage conditions based on historical quality inspection data, and improves the risk analysis ability for specific food types and additives. Summary of the Invention

[0004] The object of the present invention is to provide a method and system for risk assessment of harmful substances in food based on big data.

[0005] To achieve the above object, the present invention is implemented according to the following technical solutions:

[0006] The first aspect of the present invention provides a method for risk assessment of harmful substances in food based on big data, including the following steps:

[0007] Step S1, collect the data of the food to be detected and the quality inspection standard data, construct a standard data set through the quality inspection standard data, and preprocess the data of the food to be detected. The data of the food to be detected includes ingredient data, additive data, heavy metal data, and storage data, and the storage data includes storage time and storage temperature and humidity data;

[0008] Step S2, obtain the detection period through the standard data set, extract time series features based on the data of the food to be detected combined with the detection period, obtain the risk change rate by comparing the time series features with the standard data set, obtain the standard threshold interval based on the standard data set, obtain the food quality index based on the standard threshold interval, and input the data of the food to be detected and the food quality index into a fixed effect model to obtain the effect estimate value of the influence of additives and storage environment on the change of heavy metal content;

[0009] Step S3: Train a BP neural network model using the risk change rate, effect estimate value, and storage temperature and humidity data. Optimize the BP neural network model using a genetic algorithm to construct a risk assessment model, output an evaluation result based on the storage time data, and obtain a benchmark evaluation result based on the standard data set;

[0010] Step S4: Calculate the risk deviation using the evaluation result and the benchmark evaluation result, and output the risk level according to the risk deviation.

[0011] Furthermore, the method for preprocessing the food data to be detected includes:

[0012] Deduplicate and normalize the food data to be detected, calculate the MIV value of the normalized food data to be detected using the mean impact value algorithm, arrange the food data to be detected in descending order of the absolute value of MIV, and retain the first 90% of the data.

[0013] Furthermore, the method for obtaining the risk change rate by extracting time series features based on the food data to be detected in combination with the detection period and comparing with the standard data set includes:

[0014] Obtain the detection period through the standard data set. The detection period is divided according to the size of the natural quarter. Based on the food data to be detected in combination with the detection period, use the sliding window algorithm to extract time series features with a window size of 3 days. The time series features include sliding average concentration data and change rate data. Query the mean and standard deviation in the same quarter and the same temperature and humidity range in the standard data set according to the time series feature data, calculate the difference between the time series feature value of the food data to be detected and the standard mean, and divide by the standard deviation to obtain the normalized risk change rate.

[0015] Furthermore, the method for obtaining the standard threshold interval based on the standard data set and obtaining the food quality index through the standard threshold interval includes:

[0016] Based on the standard data set, compare each index of the food data to be detected to obtain the corresponding standard component data, standard additive data, standard heavy metal data, and standard storage data, which constitute the standard threshold interval corresponding to the food. Generate a food quality index based on the standard threshold interval. Use the food data to be detected to locate the abnormal time points, sort the continuous detection periods in combination with the food quality index and the abnormal time points, and obtain the weight of the risk change rate according to the sorting of the detection periods.

[0017] Furthermore, the method for using the food data to be detected to locate the abnormal time points includes: performing outlier detection on the food data to be detected, and setting the data points with an outlier score greater than 0.7 as the set threshold PT i , The formula is as follows:

[0018] The formula is as follows:

[0019]

[0020] where OS i is the outlier score of the i-th food, N is the number of isolation trees, the preset value is 100, h j is the path length of the data point x i in the j-th isolation tree, n i is the number of data points of the i-th food, H(n i ) represents ln(n i ) + γ, where γ is the Euler-Mascheroni constant.

[0021] Calculate the EWMA index at the corresponding time point using the data points above the threshold PT i , and calculate the dynamic threshold DT of the data points using the EWMA index i :

[0022]

[0023] where λ is the smoothing parameter, x i is the observed value of the i-th sample, EWMA i is the exponentially weighted moving average of the i-th sample, EWMA i-1 is the exponentially weighted moving average of the (i - 1)-th sample, k is the threshold adjustment coefficient, V i-1 is the variance of the (i - 1)-th sample;

[0024] Mark the data points in the food data to be detected that exceed the dynamic threshold DT i as outliers, and locate the abnormal time points through the outliers.

[0025] Furthermore, input the food data to be detected and the food quality indicators into the fixed effects model to obtain the effect estimate of the influence of additives and storage environment on the change of heavy metal content. The calculation method of the effect estimate includes:

[0026] Taking the food type and time effect as fixed effects, establish a two-way fixed effects regression model, and the formula is:

[0027] Y it = α + β1X it + β2T 1it + β3T 2it + … + μ i + λ t + ∈ it

[0028] where Y it represents the harmful substance content of the i-th food in the t-th period, α represents the baseline level of harmful substances, X itDenote the additive data of the \(i\)-th food in the \(t\)-th period, \(T\) 1it is the storage temperature of the \(i\)-th food in the \(t\)-th period, \(T\) 2it is the storage humidity of the \(i\)-th food in the \(t\)-th period, \(\beta_1\), \(\beta_2\), \(\beta_3\) represent the weights of each influencing data, \(\mu\) i represents the fixed effect of the \(i\)-th food, \(\lambda\) t represents the time fixed effect, \(\epsilon\) it is the random error term;

[0029] Convert the food quality indicators and the data of the food to be detected into a panel data structure, input it into the two-way fixed effect regression model, use the dummy variable method to introduce dummy variables for the food to be detected, convert the two-way fixed effect regression model into a multiple linear regression model, use the least squares method to obtain the maximum likelihood estimate value, and based on the maximum likelihood estimate value, obtain the effect estimate value of the influence of additives and storage environment on the change of heavy metal content. The formula is as follows:

[0030]

[0031] where \(\ln L\) represents the value of the likelihood function, \(Y\) it represents the observed value of the \(i\)-th food in the \(t\)-th period, represents the predicted value of the two-way fixed effect regression model, \(\sigma\) 2 is the variance of the random error term \(\epsilon\) it of \(n\) is the number of individuals, and \(T\) is the number of time points.

[0032] Furthermore, use the genetic algorithm to optimize the BP data network model. The training method of the neural network model includes:

[0033] Unfold the weights from the input layer to the hidden layer, the weights from the hidden layer to the output layer, the biases of the hidden layer, and the biases of the output layer in the BP neural network into one-dimensional vectors as the individual chromosomes of the genetic algorithm. Initialize the population and set parameters. Obtain the network parameters by decoding the chromosomes, train the network and calculate the fitness value. Use roulette wheel selection to select excellent individuals, perform single-point crossover and Gaussian mutation to increase the population diversity. The Gaussian mutation calculation formula is:

[0034]

[0035] where is the gene value of individual \(i\) in the next generation, is the gene value of individual \(i\) in the current generation, \(\sigma\) i is the standard deviation of individual \(i\), \(a\) represents the amplitude of adjusting the mutation, \(D\) is the current distance metric, \(D\) max is the maximum value of the distance, \(b\) is the amplitude of adjusting the additional mutation, \(R\) is the normalization constant, is the change rate of the current point, \(Pop\)t , For updating gene values, τ 2 Control the update amplitude, where exp is the exponential function;

[0036] Perform random perturbation on the mutated optimal individual. The random perturbation number is 10% of the population size, and the initial value of the perturbation range is 6% - 12% of the parameter range, which decays with iteration. The perturbation formula is:

[0037]

[0038] where is the gene value of individual i in the next generation, is the gene value of individual i in the current generation, Δg i is the perturbation step size of individual i, c is the simulated annealing cooling coefficient, c = 0.99t, t is the current iteration number, d is the simulated annealing influence factor, T is the current simulated annealing temperature, T0 is the initial simulated annealing temperature, ω d is the dimension weight factor, C(r) is the value of the chaotic mapping function;

[0039] Generate a new individual after random perturbation, and use the offspring to replace the individual with the lowest fitness value in the parent generation to complete population update. If the maximum iteration number is reached or the change in the optimal fitness is less than 1% for 10 consecutive generations, terminate early and output the optimal solution, where the optimal solution includes the optimal weight and the optimal bias;

[0040] Set up a BP neural network model based on the optimal weight and the optimal bias, input the risk change rate, the effect estimate value, and the storage temperature and humidity data into the BP neural network model for training to obtain a risk assessment model, output the evaluation result based on the storage time data, and obtain the benchmark evaluation result based on the standard data set.

[0041] Furthermore, calculate the risk deviation degree using the evaluation result and the benchmark evaluation result. According to the risk deviation degree, output the method for obtaining the risk level, including:

[0042] Calculate the risk deviation degree using the evaluation result and the benchmark evaluation result, and establish a risk level model based on the risk deviation degree: Low risk: when R < 0.1; Medium risk: when 0.1 ≤ R < 0.3; High risk: when R ≥ 0.3. The formula for the risk deviation degree is:

[0043]

[0044] where ω i is the index weight, n is the total number of risk factors, is the model prediction value of the i-th risk factor, r iis the historical data value of the i-th risk factor, T is the number of historical time periods reviewed, t is the current time point, k represents the k-th time period reviewed from the current time point t, α is the time decay factor, and ∈ prevents division by zero errors;

[0045] Based on the risk deviation degree, compare the risk level model and output the risk level.

[0046] The second aspect of the present invention provides a risk assessment system for harmful substances in food based on big data, including:

[0047] Data acquisition module: Collect the data of the food to be detected and the quality inspection standard data, construct a standard data set through the quality inspection standard data, and preprocess the data of the food to be detected. The data of the food to be detected includes ingredient data, additive data, heavy metal data, and storage data. The storage data includes storage time and storage temperature and humidity data;

[0048] Data analysis module: Obtain the detection cycle through the standard data set, extract time series features based on the data of the food to be detected combined with the detection cycle, obtain the risk change rate by comparing the time series features with the standard data set, obtain the standard threshold interval based on the standard data set, obtain the food quality index based on the standard threshold interval, and input the data of the food to be detected and the food quality index into the fixed effect model to obtain the effect estimate value of the influence of additives and storage environment on the change of heavy metal content;

[0049] Risk modeling module: Use the risk change rate, effect estimate value, and storage temperature and humidity data to train the BP neural network model, optimize the BP data network model using the genetic algorithm, construct a risk assessment model, output the assessment result based on the storage time data, and obtain the benchmark assessment result based on the standard data set;

[0050] Risk assessment module: Calculate the risk deviation degree using the assessment result and the benchmark assessment result, and output the risk level according to the risk deviation degree.

[0051] In the third aspect, an embodiment of the present application also provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the method steps described in the first aspect.

[0052] In the third aspect, an embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute the method steps described in the first aspect.

[0053] The beneficial effects of the present invention are:

[0054] The present invention relates to a method and system for risk assessment of harmful substances in food based on big data. Compared with the prior art, the present invention has the following technical effects:

[0055] (1) By establishing food quality indicators according to quality inspection standards, the present invention quantifies food quality data. At the same time, in combination with the time-varying characteristics of harmful substances in food, the risk changes of different factors on heavy metal content are obtained.

[0056] (2) Through a two-way fixed effects regression model, the present invention quantifies the time-varying coupling effect of storage environment and additives on the content of harmful substances.

[0057] (3) By using the effect estimates of different types of food to train a BP neural network model, a risk assessment model is constructed. Based on a genetic algorithm to optimize the global performance of the assessment model, the assessment efficiency and accuracy of the assessment model can be improved. Description of the Drawings

[0058] Figure 1 It is a flowchart of the steps of a method for risk assessment of harmful substances in food based on big data according to the present invention.

[0059] Figure 2 It is a schematic structural diagram of an electronic device in an embodiment of this specification. Detailed Embodiments

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0061] Referring to Figure 1 as shown, the first method for risk assessment of harmful substances in food based on big data according to the present invention includes:

[0062] Step S1: Collect data of the food to be detected and quality inspection standard data, construct a standard data set through the quality inspection standard data, and preprocess the data of the food to be detected. The data of the food to be detected includes ingredient data, additive data, heavy metal data, and storage data. The storage data includes storage time and storage temperature and humidity data;

[0063] In an actual evaluation, in the case of risk assessment of a packaged biscuit, food data of 100 samples to be tested was collected. Among them, the ingredient data includes the contents of various raw materials such as flour, sugar, and oil; the additive data records the addition amounts of preservatives, leavening agents, etc.; the heavy metal data detects the residual amounts of heavy metals such as lead, mercury, and cadmium in the biscuits; the storage data includes storage temperature and humidity and storage duration, and a standard data set is constructed through the biscuit quality inspection standards. By checking the 100 sample data collected, it is found that the data of 5 samples are completely repeated, and these 5 repeated sample data are deleted, leaving 95 valid sample data. The food data to be tested in different dimensions is unified to the [0,1] interval, and the MIV value is calculated for the 95 normalized sample data using the mean impact value algorithm. The sample data is sorted in descending order according to the absolute value of MIV, and the first 90% of the sample data is retained. Finally, 86 sample data are obtained. The final data has a relatively large impact on food quality in terms of ingredients, additives, heavy metals, and storage conditions, and can more effectively reflect the key factors of food quality.

[0064] Step S2: Obtain the detection period through the standard data set, extract time series features based on the food data to be tested combined with the detection period, obtain the risk change rate by comparing the time series features with the standard data set, obtain the standard threshold interval based on the standard data set, obtain the food quality index based on the standard threshold interval, and input the food data to be tested and the food quality index into the fixed effect model to obtain the effect estimation value of the influence of additives and storage environment on the change of heavy metal content;

[0065] It should be explained that obtaining the detection period through the standard data set, extracting time series features based on the food data to be tested combined with the detection period, and obtaining the risk change rate by comparing the time series features with the standard data set is because the time series features of the standard data set under the same season and storage conditions can show the change trend and change rate of the harmful substance data in the food data to be tested within the same time length in advance.

[0066] In an actual evaluation, for the detection data of the content of heavy metal lead, within the time range of the first quarter, every 3 days is used as a window. Within the first 30 days, a total of 10 sets of time series feature data of sliding windows are obtained to calculate the sliding average concentration data and change rate data of lead within this window. The sliding average concentration of lead in a certain sliding window is 0.06 mg / kg, the mean value of lead content in the standard data set is 0.05 mg / kg, and the standard deviation is 0.01 mg / kg. Calculate the difference 0.06 - 0.05 = 0.01 mg / kg, and then divide it by the standard deviation 0.01 mg / kg to obtain a normalized risk change rate of 1

[0067] It should be noted that by comparing the various indicators of the food to be detected with the standard data set, the corresponding standard ingredient data, standard additive data, standard heavy metal data, and standard storage data are obtained, constituting the standard threshold range corresponding to the food. Generating the food quality indicators based on the standard threshold range is because the standard data set contains the ingredient data of each food. By comparing specific foods, the corresponding standard data indicators can be screened out for detecting the food to be detected and generating the quality indicators of the corresponding food for evaluating the quality grade of a single food.

[0068] In actual evaluation, for the flour content in the ingredient data, the standard data set stipulates that the flour content in every 100 grams of biscuits should be between 40 - 70 grams; for the preservative potassium sorbate in the additives, the standard stipulates that the maximum usage amount is 0.5 grams per 100 grams of biscuits; for the heavy metal lead, the standard limit is below 0.2 mg / kg; the standard stipulates that the suitable storage temperature is 10 - 22 °C and the relative humidity is 30% - 50%; the standard storage time is 12 - 18 months. These data constitute the standard threshold range of biscuits and generate the quality indicators of biscuits. The detected value of the lead content of a certain sample is 0.06 mg / kg, which does not exceed the standard threshold range, so the food quality indicator of this sample is determined to be qualified.

[0069] It should be noted that the abnormal time points are located using the food data to be detected, and the continuous detection cycles are sorted by combining the food quality indicators and the abnormal time points. The weight of the risk change rate is obtained according to the sorting of the detection cycles. Outlier detection can screen out the corresponding abnormal time points in advance. Based on this time point, the data in the time periods before and after it can be preferentially detected to obtain a high-priority detection range in advance, further improving the efficiency and accuracy of the evaluation.

[0070] In actual evaluation, the number N of isolation trees is set to 100, and the outlier scores of the food data within 30 days are calculated. The data points with outlier scores greater than 0.7 are set as the threshold. After calculation, the outlier score of the lead content data point on the 30th day is 0.8, which is higher than 0.7 and is selected; the EWMA index corresponding to the time point is calculated, the smoothing parameter is set to 0.2, and the threshold adjustment coefficient k is 3. The dynamic threshold DT on the 30th day is calculated according to the formula 30 to be 0.07 mg / kg, while the actually detected lead content on the 30th day is 0.08 mg / kg, exceeding the dynamic threshold. Then the 30th day is marked as an abnormal time point, and the 30 days before and after the 30th day are used as the calculation range of the risk change rate of the abnormal time point to obtain the first high-risk time period. Then multiple high-risk time periods are obtained in turn, the corresponding risk change rates are calculated and sorted, and the weights are obtained according to the sorting;

[0071] In the actual evaluation, using the 86 sample data collected and the quality indicators of the biscuits, the least squares method is used to estimate the model. After calculation, it is obtained that: α = 0.05, β1 = 0.002, β2 = 0.001, β3 = 0.0005. It means that under the condition that other conditions remain unchanged, when the preservative addition amount increases by 1 unit, the lead content increases by an average of 0.002 units; when the storage temperature rises by 1 °C, the lead content increases by an average of 0.001 units; when the storage humidity increases by 1%, the lead content increases by an average of 0.0005 units.

[0072] Step S3: Input the effect estimate value and environmental data into the neural network model for training, optimize the model using the genetic algorithm, and construct a risk assessment model;

[0073] In the actual evaluation, a BP neural network model is set: the input layer has 5 nodes, corresponding to the risk change rate, effect estimate value, storage temperature, storage humidity, and storage time respectively. The hidden layer is set with 8 nodes, and the activation function uses LeakyReLU. The output layer has 3 nodes, representing three evaluation results of low risk, medium risk, and high risk, and the activation function is Softmax. The initial value of the learning rate is set to 0.3, and it decays by 10% every 100 iterations. Expand the weights from the input layer to the hidden layer, the weights from the hidden layer to the output layer, the biases of the hidden layer, and the biases of the output layer in the BP neural network into one-dimensional vectors as the individual chromosomes of the genetic algorithm. Initialize the population size to 50, set the maximum number of iterations to 200, and set the genetic algorithm parameters. The crossover probability is 0.8, the mutation probability is 0.1. Obtain the network parameters by decoding the chromosomes, use the mean square error as the fitness function to calculate the fitness value. The sum of the individual fitness values in the current population is 50, and the fitness value of individual A is 2, so the probability of individual A being selected is 0.04. Perform a single-point crossover operation, randomly select the crossover point, exchange part of the bases of the two individuals, and at the same time perform Gaussian mutation. Then the mutated gene value is 0.3015. Perform random perturbation on the 5 optimal individuals after mutation. The initial perturbation range is [0.06, 0.12]. In a certain optimal individual, the weight value from the first node of the input layer to the third node of the hidden layer is 0.35, Δg i The initial value is 0.01, and the current is the 50th iteration, then the perturbation value is 0.354158. After performing such perturbation operations on each gene value of the 5 optimal individuals to obtain new individuals, find the individual with the lowest fitness value, replace it with the new individual of the offspring to complete the population update. The optimal fitness value of the 180th generation is 25.5, and the change rate is 0.75%. Prematurely terminate the algorithm and output the finally obtained optimal weight matrix. Based on the optimal weights and optimal biases, set the BP neural network model to obtain a risk assessment model:

[0074] Step S4. Calculate the risk deviation degree using the evaluation result and the benchmark evaluation result, and output the risk level according to the risk deviation degree.

[0075] It should be noted that the evaluation result output by the model is based on the prediction result obtained through training, and it is also necessary to compare and analyze it with the benchmark evaluation result obtained from the standard data set to improve the credibility of the final result.

[0076] In the actual evaluation, it is determined that there are 3 risk factors, the number of historical time periods reviewed is 5, and the time decay α is 0.9. Then the finally obtained risk deviation degree is 0.12, which is greater than 0.1 and less than 0.3, and it is determined as medium risk.

[0077] In this embodiment, the method for using the food data to be detected to locate the abnormal time point includes:

[0078] Use the food data to be detected for outlier detection, and set the data points with an outlier score greater than 0.7 as the set threshold PT i , and the formula is as follows:

[0079]

[0080] Where OS i is the outlier score of the i-th type of food, N is the number of isolation trees, the preset value is 100, h j is the path length of the data point x i in the j-th isolation tree, n i is the number of data points of the i-th type of food, H(n i ) represents ln(n i ) + γ, and γ is the Euler-Mascheroni constant.

[0081] Use the data points higher than the threshold PT i to calculate the EWMA index of the corresponding time point, and use the EWMA index to calculate the dynamic threshold DT of the data point i :

[0082]

[0083] Where λ is the smoothing parameter, x i is the observed value of the i-th sample, EWMA i is the exponentially weighted moving average of the i-th sample, EWMA i-1 is the exponentially weighted moving average of the i-1-th sample, k is the threshold adjustment coefficient, and V i-1 is the variance of the i-1-th sample;

[0084] Mark the data points in the food data to be detected that exceed the dynamic threshold DT i as outliers, and locate the abnormal time point through the outliers.

[0085] In this embodiment, the food data to be detected and the food quality indicators are input into the fixed effect model to obtain the effect estimation values of the influence of additives and storage environment on the change of heavy metal content. The calculation method of the effect estimation values includes:

[0086] Taking food type and time effect as fixed effects, a two-way fixed effect regression model is established. The formula is:

[0087] Y it =α+β1X it +β2T 1it +β3T 2it +…+μ i +λ t +∈ it

[0088] Where Y it represents the harmful substance content of the i-th food in the t-th period, α represents the baseline level of harmful substances, X it represents the additive data of the i-th food in the t-th period, T 1it is the storage temperature of the i-th food in the t-th period, T 2it is the storage humidity of the i-th food in the t-th period, β1, β2, β3 represent the weights of each influencing data, μ i represents the fixed effect of the i-th food, λ t represents the time fixed effect, ∈ it is the random error term;

[0089] The food quality indicators and the food data to be detected are converted into a panel data structure, input into the two-way fixed effect regression model, and dummy variables are introduced for the food to be detected using the dummy variable method to transform the two-way fixed effect regression model into a multiple linear regression model. The maximum likelihood estimation value is obtained using the least squares method. Based on the maximum likelihood estimation value, the effect estimation values of additives and storage data on the harmful substance content in food are obtained. The formula is as follows:

[0090]

[0091] Where ln L represents the value of the likelihood function, Y it represents the observed value of the i-th food in the t-th period, represents the predicted value of the two-way fixed effect regression model, σ 2 is the variance of the random error term ∈ it , n is the number of individuals, and T is the number of time points.

[0092] In this embodiment, a genetic algorithm is used to optimize the BP data network model. The training method of the neural network model includes:

[0093] Set the input layer of the BP neural network: 5 nodes, hidden layer: 8 nodes, activation function LeakyReLU, output layer: 3 nodes, activation function Softmax, learning rate: initially 0.3, decaying by 10% every 100 iterations;

[0094] Unfold the weights from the input layer to the hidden layer, the weights from the hidden layer to the output layer, the biases of the hidden layer, and the biases of the output layer in the BP neural network into one-dimensional vectors as the individual chromosomes of the genetic algorithm. Initialize the population and set parameters. Obtain the network parameters by decoding the chromosomes, train the network and calculate the fitness value. Use roulette wheel selection to select excellent individuals, perform single-point crossover and Gaussian mutation to increase population diversity. The Gaussian mutation calculation formula is:

[0095]

[0096] where is the gene value of individual i in the next generation, is the gene value of individual i in the current generation, σ i is the standard deviation of individual i, a represents the amplitude of adjusting mutation, D is the current distance metric, D max is the maximum value of the distance, b is the amplitude of adjusting additional mutation, R is the normalization constant, is the change rate of the current point, Pop t , is used to update the gene value, τ 2 controls the update amplitude, exp is the exponential function;

[0097] Perform random perturbation on the mutated optimal individual. The random perturbation number is 10% of the population size. The initial value of the perturbation range is 6% - 12% of the parameter range, and it decays with iterations. The perturbation formula is:

[0098]

[0099] where is the gene value of individual i in the next generation, is the gene value of individual i in the current generation, Δg i is the perturbation step size of individual i, c is the simulated annealing cooling coefficient, c = 0.99t, t is the current iteration number, d is the simulated annealing influence factor, T is the current simulated annealing temperature, T0 is the initial simulated annealing temperature, ω d is the dimension weight factor, C(r) is the value of the chaotic mapping function;

[0100] Generate new individuals after random perturbation, and use the offspring to replace the individual with the lowest fitness value in the parent generation to complete the population update. If the maximum iteration number is reached or the change in the optimal fitness is less than 1% for 10 consecutive generations, terminate early and output the optimal solution, where the optimal solution includes the optimal weights and optimal biases;

[0101] Set up a BP neural network model based on the optimal weights and optimal biases, input the risk change rate, effect estimate value, and storage temperature and humidity data into the BP neural network model for training to obtain a risk assessment model, output the assessment result based on the storage time data, and obtain the benchmark assessment result based on the standard data set.

[0102] In this embodiment, the method for obtaining the risk level includes:

[0103] Calculate the risk deviation using the assessment result and the benchmark assessment result, and establish a risk level model based on the risk deviation: low risk: when R < 0.1; medium risk: when 0.1 ≤ R < 0.3; high risk: when R ≥ 0.3. The formula for the risk deviation is:

[0104]

[0105] where ω i is the index weight, n is the total number of risk factors, is the model prediction value of the i-th risk factor, r i is the historical data value of the i-th risk factor, T is the number of historical time periods reviewed, t is the current time point, k represents the k-th time period reviewed from the current time point t, α is the time decay factor, and ∈ prevents division by zero errors;

[0106] Compare the risk level model based on the risk deviation and output the risk level.

[0107] Figure 2 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 2 . At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include memory, such as high-speed random access memory (Random-Access Memory, RAM), and may also include non-volatile memory (non-volati le memory), such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0108] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 2 only a bidirectional arrow is used in

[0109] but it does not mean that there is only one bus or one type of bus.

[0110] Memory, which is used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0111] The above is as in the present application Figure 1The embodiment shown discloses a method for risk assessment of harmful substances in food based on big data, which can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or instructions in the form of software. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0112] The electronic device may also perform Figure 1 A risk assessment method for harmful substances in food based on big data is proposed Figure 1 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0113] An embodiment of the present application also proposes a computer-readable storage medium, which stores one or more programs, and the one or more programs include instructions. When the instructions are executed by an electronic device including multiple applications, any one of the aforementioned methods for risk assessment of harmful substances in food based on big data is executed.

[0114] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0115] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a system for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0116] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction system that implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0118] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0119] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0120] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0121] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0122] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0123] The above contents are merely examples and explanations of the structure of the present invention. The technicians in this technical field may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the claims, they should all fall within the protection scope of the present invention.

Claims

1. A risk assessment method for harmful substances in food based on big data, characterized in that: The following steps are involved: Step S1, collecting food data to be tested and quality inspection standard data, constructing a standard data set through the quality inspection standard data, and preprocessing the food data to be tested, wherein the food data to be tested includes ingredient data, additive data, heavy metal data and storage data, and the storage data includes storage time and storage temperature and humidity data; Step S2, obtaining the detection period through the standard data set, extracting the time series characteristics based on the food data to be detected in combination with the detection period, obtaining the risk change rate by comparing the standard data set according to the time series characteristics, obtaining the standard threshold interval based on the standard data set, obtaining the food quality index based on the standard threshold interval, inputting the food data to be detected and the food quality index into the fixed effect model, and obtaining the effect estimate of the influence of additives and storage environment on the change of heavy metal content; Step S3, using the risk change rate, effect estimate and storage temperature and humidity data to train the BP neural network model, using the genetic algorithm to optimize the BP neural network model, constructing a risk assessment model, outputting the assessment results based on the storage time data, and obtaining the benchmark assessment results based on the standard data set; Step S4: Calculate the risk deviation using the assessment result and the benchmark assessment result, and output the risk level based on the risk deviation.

2. According to claim 1, a method for risk assessment of harmful substances in food based on big data is characterized in that: Methods for preprocessing food data to be tested include: Based on the food data to be tested, deduplication and normalization processing are performed, and the MIV value of the normalized food data to be tested is calculated using the average influence value algorithm. The food data to be tested are arranged in descending order according to the absolute value of MIV, and the top 90% of the data are retained.

3. The method for risk assessment of harmful substances in food based on big data according to claim 1 is characterized in that: A method for obtaining a risk change rate by extracting time series features based on the food data to be tested and the testing cycle, and comparing the time series features with the standard data set includes: The detection cycle is obtained through the standard data set. The detection cycle is divided according to the size of the natural quarter. The sliding window algorithm is used based on the food data to be detected and the detection cycle to extract time series features with a window size of 3 days. The time series features include sliding average concentration data and change rate data. The mean and standard deviation of the same quarter and the same temperature and humidity range in the standard data set are queried according to the time series feature data, and the difference between the time series feature value of the food data to be detected and the standard mean is calculated and divided by the standard deviation to obtain the normalized risk change rate.

4. The method for risk assessment of harmful substances in food based on big data according to claim 1, characterized in that: A method for obtaining a standard threshold interval based on a standard data set and obtaining a food quality indicator through the standard threshold interval includes: Based on the standard data set, various indicators of the food data to be tested are compared to obtain the corresponding standard ingredient data, standard additive data, standard heavy metal data and standard storage data, which constitute the standard threshold interval corresponding to the food. The food quality index is generated based on the standard threshold interval, and the food data to be tested is used to locate the abnormal time point. The continuous detection cycles are sorted in combination with the food quality indicators and the abnormal time point, and the weight of the risk change rate is obtained according to the sorting of the detection cycles.

5. The method for risk assessment of harmful substances in food based on big data according to claim 4 is characterized in that: The method of using the food data to be tested to locate abnormal time points includes: Use the food data to be tested for outlier detection, and set the data points with outlier scores greater than 0.7 as the set threshold PT i , the formula is as follows: Among them, OS i is the outlier score of the i-th food, N is the number of isolation trees, the default value is 100, h j For data point x i The path length in the jth isolation tree, n i is the number of data points for the i-th food, H(n i ) is expressed as ln(n i )+γ, γ is the Euler-Mascheroni constant; Use PT above the threshold i The EWMA index of the corresponding time point is calculated using the EWMA index, and the dynamic threshold DT of the data point is calculated using the EWMA index. i : Among them, λ is the smoothing parameter, x i is the observed value of the i-th sample, EWMA i is the exponentially weighted moving average of the i-th sample, EWMA i-1 is the exponentially weighted moving average of the i-1th sample, k is the threshold adjustment coefficient, V i-1 is the variance of the i-1th sample; The food data to be tested that exceeds the dynamic threshold DT i The data points are marked as outliers, and the abnormal time points are located through the outliers.

6. The method for risk assessment of harmful substances in food based on big data according to claim 1, characterized in that: The food data to be tested and the food quality index are input into the fixed effect model to obtain the effect estimate of the influence of additives and storage environment on the change of heavy metal content. The calculation method of the effect estimate includes: Taking food type and time effect as fixed effects, a two-way fixed effect regression model is established, and the formula is: Y it =α+β1X it +β2T 1it +β3T 2it +…+m i +λ t +∈ it where Y it represents the harmful substance content of the i-th food in the t-th period, α represents the baseline level of harmful substances, and X it represents the additive data of the i-th food in the t-th period, T 1it is the storage temperature of the i-th food in the t-th period, T 2it is the storage humidity of the i-th food in the t-th period, β1, β2, β3 represent the weights of each influencing data, μ i represents the fixed effect of the i-th food, λ t represents the time fixed effect, ∈ it is the random error term; The food quality indicators and the food data to be tested were converted into a panel data structure, input into a two-way fixed effect regression model, and virtual variables were introduced into the food to be tested using the virtual variable method. The two-way fixed effect regression model was converted into a multivariate linear regression model, and the maximum likelihood estimate was obtained using the least squares method. Based on the maximum likelihood estimate, the effect estimate of the influence of additives and storage environment on the change of heavy metal content was obtained, and the formula is as follows: Where ln L represents the likelihood function value, Y it represents the observed value of the i-th food in the t-th period, represents the predicted value of the two-way fixed effect regression model, σ 2 is the random error term ∈ it The variance of , n is the number of individuals, and T is the number of time points.

7. The method for risk assessment of harmful substances in food based on big data according to claim 1, characterized in that: A genetic algorithm is used to optimize a BP neural network model, wherein the training method of the BP neural network model comprises: The weights from the input layer to the hidden layer, the weights from the hidden layer to the output layer, the bias of the hidden layer, and the bias of the output layer in the BP neural network are expanded into one-dimensional vectors as individual chromosomes of the genetic algorithm. The population is initialized and the parameters are set. The network parameters are obtained by decoding the chromosomes, the network is trained, and the fitness value is calculated. The roulette wheel is used to select excellent individuals, and single-point crossover and Gaussian mutation are performed to increase the diversity of the population. The Gaussian mutation formula is: in is the gene value of individual i in the next generation, is the gene value of individual i in the current generation, σ i is the standard deviation of individual i, a represents the amplitude of adjustment variation, D is the current distance measure, and D max is the maximum value of the distance, b is the adjustment of the additional variation, R is the normalization constant, is the rate of change of the current point, Pop t , Used to update gene values, τ 2 Control update amplitude, exp is exponential function; The optimal individual after mutation is randomly disturbed, and the number of random disturbances is 10% of the population size. The initial value of the disturbance range is 6%-12% of the parameter range, and decays with iteration. The disturbance formula is: in is the gene value of individual i in the next generation, is the gene value of individual i in the current generation, Δg i is the perturbation step size of individual i, c is the simulated annealing cooling coefficient, c = 0.99t, t is the current iteration number, d is the simulated annealing influence factor, T is the current simulated annealing temperature, T0 is the initial simulated annealing temperature, ω d is the dimension weight factor, C(r) is the chaotic mapping function value; Generate new individuals after random disturbance, use offspring to replace the individuals with the lowest fitness value in the parent generation to complete population update, if the maximum number of iterations is reached or the change in the optimal fitness for 10 consecutive generations is less than 1%, terminate early and output the optimal solution, which includes the optimal weight and the optimal bias; A BP neural network model is set based on the optimal weight and the optimal bias, and the risk change rate, effect estimate and storage temperature and humidity data are input into the BP neural network model for training to obtain a risk assessment model. The assessment results are output based on the storage time data, and the benchmark assessment results are obtained based on the standard data set.

8. The method for risk assessment of harmful substances in food based on big data according to claim 1, characterized in that: The risk deviation is calculated using the assessment results and the benchmark assessment results. Based on the risk deviation, the method for obtaining the risk level is output, including: The risk deviation is calculated using the assessment results and the benchmark assessment results, and a risk level model is established based on the risk deviation: low risk: when R<0.1; medium risk: when 0.1≤R<0.3; high risk: when R≥0.

3. The formula for the risk deviation is: where ω i is the indicator weight, n is the total number of risk factors, is the model prediction value of the ith risk factor, r i is the historical data value of the ith risk factor, T is the number of historical time periods reviewed, t is the current time point, k represents the kth time period reviewed from the current time point t, α is the time decay factor, ∈ prevents division by zero errors; Based on the risk deviation comparison risk level model, the risk level is output.

9. A risk assessment system for harmful substances in food based on big data, used to execute the method according to any one of claims 1 to 7, characterized in that: include: Data collection module: collects the food data to be tested and the quality inspection standard data, constructs a standard data set through the quality inspection standard data, and pre-processes the food data to be tested. The food data to be tested includes ingredient data, additive data, heavy metal data and storage data. The storage data includes storage time and storage temperature and humidity data; Data analysis module: obtain the detection period through the standard data set, extract the time series characteristics based on the food data to be tested combined with the detection period, compare the standard data set according to the time series characteristics to obtain the risk change rate, obtain the standard threshold interval based on the standard data set, obtain the food quality index based on the standard threshold interval, input the food data to be tested and the food quality index into the fixed effect model, and obtain the effect estimate of the influence of additives and storage environment on the change of heavy metal content; Risk modeling module: Use risk change rate, effect estimation value and storage temperature and humidity data to train BP neural network model, use genetic algorithm to optimize BP data network model, build risk assessment model, output assessment results based on storage time data, and obtain benchmark assessment results based on standard data set; Risk assessment module: Use the assessment results and benchmark assessment results to calculate the risk deviation, and output the risk level based on the risk deviation.

10. An electronic device comprising: processor; and a memory arranged to store computer executable instructions, which, when executed, cause the processor to perform the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • FPGA-based RBF neural network PID parameter setting method

    CN120428540A