Food risk prediction method and device, electronic equipment and storage medium
By combining and processing the data to be predicted with historical data sets and using deep learning and machine learning algorithms to train food risk prediction models, the difficulty of obtaining data in imported food risk assessment and the complexity of cross-border regulations is solved, efficient food safety risk identification and evaluation is achieved, and regulatory efficiency is improved.
Patent Information
- Application Number
- CN202510422286.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-22
AI Technical Summary
The imported food risk reasoning algorithm model has problems such as difficulty in obtaining data, complexity of transnational regulations, and the difficulty in evaluation brought about by diversified food types.
By entering the data to be predicted into the food risk prediction model, the historical data set is obtained, and the first risk prediction result is combined with the historical data set to form a new training data set, the model is trained using deep learning and machine learning algorithms to generate a trained food risk prediction model for risk prediction of the target data.
It improves the accuracy of food risk identification and evaluation, improves the overall efficiency of food safety supervision, reduces the time and resource costs of manual review, and supports the food safety supervision needs in the global market.
Smart Images

Figure CN120525546A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a food risk prediction method, device, electronic device and storage medium. Background Art
[0002] In recent years, with the rapid development of technologies such as big data, machine learning, and natural language processing, imported food risk inference algorithm models have gradually become an important research direction in the field of food safety. Imported food risk inference algorithm models use computer technology to automatically assess and analyze the potential risks of imported foods and predict food risks to ensure food safety and consumer health. Unlike traditional food safety monitoring systems, this model not only relies on static data but also incorporates dynamic information and real-time risk assessment to provide more accurate and timely decision support. Currently, imported food risk inference algorithm models face challenges such as data acquisition difficulties, the complexity of cross-border regulations, and the difficulty of assessment due to the diverse types of food. Summary of the Invention
[0003] The embodiment of the present invention provides a food risk prediction method, which aims to provide an efficient and intelligent food risk prediction method to improve the accuracy of food risk prediction. By inputting the data to be predicted into the food risk prediction model for risk prediction, a first risk prediction result is obtained, a historical data set is obtained, and the first risk prediction result is merged with the historical data set to obtain a new training data set. The food risk prediction model is trained with the new data set to obtain a trained food risk prediction model, and the target data is risk predicted with the trained food risk prediction model to obtain the risk prediction result of the target data. This can solve the problems of data acquisition difficulties in the imported food risk inference algorithm model, the complexity of cross-border regulations, and the high difficulty of evaluation caused by diversified food types, and can effectively improve the identification and evaluation of potential food safety risks, thereby improving the overall supervision efficiency.
[0004] In a first aspect, an embodiment of the present invention provides a method for predicting food risk, the method comprising the following steps:
[0005] Inputting the data to be predicted into a food risk prediction model for risk prediction to obtain a first risk prediction result, wherein the food risk prediction model is trained based on a sample training data set;
[0006] Acquire the historical data set, and merge the first risk prediction result with the historical data set to obtain a new training data set;
[0007] Training the food risk prediction model using the new data set to obtain a trained food risk prediction model;
[0008] The risk prediction of the target data is performed using the trained food risk prediction model to obtain the risk prediction results of the target data.
[0009] Optionally, inputting the data to be predicted into a food risk prediction model to perform risk prediction to obtain a first risk prediction result includes:
[0010] Obtain the data to be predicted and the food risk prediction model;
[0011] The food risk prediction model is used to perform risk prediction on the data to be predicted to obtain a first risk prediction result.
[0012] Optionally, merging the first risk prediction result with the historical data set to obtain a new training data set includes:
[0013] Taking the first risk prediction result and the data to be predicted as a first training data set, and merging the first training data set with the historical data set to obtain a second training data set;
[0014] Perform risk analysis processing on the second training data set, and perform data balancing sampling processing on the second training data set after the risk analysis processing by an interpolation method to obtain a new training data set.
[0015] Optionally, the interpolation formula is as follows:
[0016] x new =x i +λ×(x nn -x i )
[0017] Among them, x i is a minority class sample; x nn is a nearest neighbor minority class sample of i; λ is a random number in the interval [0,1].
[0018] Optionally, the training of the food risk prediction model using the new training data set to obtain a trained food risk prediction model includes:
[0019] Inputting the new training data set into the food risk prediction model for training;
[0020] During the training process, the parameters of the food risk prediction model are adjusted using a cross-validation calculation method. After the training is completed, a trained food risk prediction model is obtained. The cross-validation calculation method formula is as follows:
[0021]
[0022] Where: γ is the loss function, D val is the validation set, is the k-th fold validation set, and θ is the model parameter.
[0023] Optionally, the training of the food risk prediction model using the new training data set to obtain a trained food risk prediction model includes:
[0024] For each new training data set, incremental training is used to train the food risk prediction model. The incremental training formula is as follows:
[0025]
[0026] Where f represents the mapping function, X represents the input feature vector, and θ represents the parameters of the model, which are iteratively updated using the following loss function:
[0027]
[0028] Where: t is the number of iterations, η is the learning rate, is the loss function.
[0029] Optionally, the training of the food risk prediction model using the new training data set to obtain a trained food risk prediction model includes:
[0030] During the training process, the parameters of the food risk prediction model were adjusted using the F1 Score evaluation method. The formula of the F1 Score evaluation method is as follows:
[0031]
[0032] Among them, F1 is the F1 Score evaluation, Precision is the precision rate, and Recall is the recall rate; the F1 score formula for the weighted method is as follows:
[0033]
[0034] Where: C is the total number of categories, F1 i is the F1 score of the i-th class, w i is the sample proportion of the i-th category.
[0035] In a second aspect, an embodiment of the present invention further provides a food risk prediction device, comprising:
[0036] A first prediction module is used to input the data to be predicted into a food risk prediction model to perform risk prediction and obtain a first risk prediction result, wherein the food risk prediction model is trained based on a sample training data set;
[0037] a processing module, configured to obtain the historical data set, and merge the first risk prediction result with the historical data set to obtain a new training data set;
[0038] A training module, configured to train the food risk prediction model using the new data set to obtain a trained food risk prediction model;
[0039] The second prediction module is used to perform risk prediction on the target data using the trained food risk prediction model to obtain the risk prediction results of the target data.
[0040] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the food risk prediction method provided by the embodiment of the present invention are implemented.
[0041] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the food risk prediction method provided in the embodiment of the invention are implemented.
[0042] In an embodiment of the present invention, the data to be predicted is input into a food risk prediction model for risk prediction to obtain a first risk prediction result, and the food risk prediction model is obtained by training based on a sample training data set; a historical data set is obtained, and the first risk prediction result is merged with the historical data set to obtain a new training data set; the food risk prediction model is trained with the new data set to obtain a trained food risk prediction model; and the target data is risk predicted with the trained food risk prediction model to obtain a risk prediction result for the target data. By inputting the data to be predicted into a food risk prediction model for risk prediction to obtain a first risk prediction result, obtaining a historical data set, and merging the first risk prediction result with the historical data set to obtain a new training data set, the food risk prediction model is trained with the new data set to obtain a trained food risk prediction model, and the target data is risk predicted with the trained food risk prediction model to obtain a risk prediction result for the target data, the problems of data acquisition difficulties, the complexity of transnational regulations, and the high difficulty of evaluation caused by diversified food types in the imported food risk inference algorithm model can be solved, and the identification and evaluation of potential food safety risks can be effectively improved, thereby improving the overall supervision efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 This is a flow chart of a food risk prediction method provided by an embodiment of the present invention;
[0045] Figure 2 This is a structural diagram of a food risk prediction method provided by an embodiment of the present invention;
[0046] Figure 3 is a structural diagram of another food risk prediction method provided by an embodiment of the present invention;
[0047] Figure 4 1 is a schematic structural diagram of a food risk prediction device provided by an embodiment of the present invention;
[0048] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0050] like Figure 1 As shown, Figure 1 : is a flow chart of a food risk prediction method provided by an embodiment of the present invention, the food risk prediction method comprising the steps of:
[0051] 101. Input the data to be predicted into the food risk prediction model for risk prediction to obtain a first risk prediction result.
[0052] In an embodiment of the present invention, the food risk prediction method can be applied to a food identification platform, which can be constructed by a server or distributed server. The food identification platform includes a database, a data processing application, a knowledge database, and a call interface for a large prediction model or a large language model. The data to be predicted can be obtained through the data interface or uploaded by users.
[0053] The above-mentioned data to be predicted can be understood as food data that requires risk prediction, including food ingredient data, inspection data, production data, etc.
[0054] The food risk prediction model is trained using a sample training dataset. This sample training dataset includes a large amount of historical food safety data, historical inspection records, market feedback, logistics tracking information, and corresponding food risk prediction result labels. The food risk prediction model can be built using deep learning or machine learning, such as a convolutional neural network (CNN) or recurrent neural network (RNN). The food risk prediction model combines prediction risk codes, inferred general risks, and product risks to automatically identify and assess different types of food risk factors. The prediction risk codes can be understood as using a stochastic gradient algorithm to optimize model parameters, thereby improving the ability to identify and predict potential risk factors. The inferred general risks can be understood as using a forest decision tree algorithm to construct multiple decision trees and average or otherwise combine the results to improve model stability and accuracy. The product risks can be understood as using a forest decision tree algorithm to specifically assess product-level risks, helping to identify potential risk points for specific products or services and providing a basis for risk management. The decision tree algorithm described above performs predictions by dividing a dataset into multiple subsets and constructing a decision tree on each subset. The random forest algorithm improves the model's generalization ability by reducing the risk of overfitting. The forest decision tree algorithm can be understood as combining multiple decision trees to improve prediction accuracy and reduce the risk of overfitting.
[0055] The above risk prediction can be understood as a process of evaluating the potential risks of food based on analysis of input data.
[0056] The above-mentioned first risk prediction result may be an assessment of possible safety risks, health risks or other related risks of food.
[0057] 102. Obtain a historical data set, and merge the first risk prediction result with the historical data set to obtain a new training data set.
[0058] In an embodiment of the present invention, the above-mentioned historical data set includes a large amount of historical food safety data, historical inspection records, market feedback, logistics tracking information, etc. and corresponding food risk prediction result labels.
[0059] The above-mentioned merging process can be understood as a process of combining different data sources or data sets together to enable more comprehensive or more accurate analysis, prediction or decision-making. The first risk prediction result is combined with the original historical data set to form a new training data set.
[0060] Specifically, the data to be predicted and the first risk prediction result may be combined as a training data set and a historical data set to obtain a new training data set.
[0061] 103. The food risk prediction model is trained using a new data set to obtain a trained food risk prediction model.
[0062] In an embodiment of the present invention, the above-mentioned training may be supervised training. Supervised training is a training method in machine learning. Its core is to use a set of data with known labels to train the model, and to optimize the model parameters so that the model can predict the labels of new data or make decisions based on the characteristics of existing data.
[0063] This trained food risk prediction model can more accurately predict food risks, monitor potential risks in real time, and issue early warnings. By analyzing historical data and market trends, it provides precise risk scores, helping users and businesses make informed decisions. The trained food risk prediction model also possesses self-learning capabilities, optimizing prediction accuracy with the input of new data and improving food safety management efficiency.
[0064] It should be noted that when training food risk prediction models, it is possible to focus on data from key food safety indicators, including food quality inspection results, food process traceability information, and food safety early warning information. These key food safety indicator data are easier to obtain because they focus only on specific indicators or parameters that have a significant impact on food safety.
[0065] 104. Risk prediction of target data is performed using the trained food risk prediction model to obtain risk prediction results of the target data.
[0066] In an embodiment of the present invention, the target data can be input into a trained food risk prediction model for feature extraction to obtain target data features, perform risk assessment and prediction on the target data features, and output the risk prediction results of the target data.
[0067] In an embodiment of the present invention, the present invention adopts a food risk prediction model, combined with machine learning and big data analysis technology, and automatically identifies and evaluates different types of risk factors by integrating multiple data sources, thereby reducing the time cost and resource investment of manual review, and improving the efficiency of risk assessment, thereby supporting food safety supervision needs for the global market. The present invention can identify the specific risks faced by different types of food through automated data collection and processing, combined with deep learning algorithms. The food risk prediction model has self-learning capabilities and can optimize the evaluation results as new data is continuously input, thereby enhancing the stability and adaptability of the model.
[0068] This invention can not only improve the intelligence level of imported food risk management, but also provide scientific and reliable decision-making support for relevant enterprises and regulatory agencies, reduce the economic losses and public health risks caused by food safety issues to relevant enterprises, and promote the further development of global imported food safety management.
[0069] In an embodiment of the present invention, the data to be predicted is input into a food risk prediction model for risk prediction to obtain a first risk prediction result. The food risk prediction model is trained based on a sample training dataset; a historical dataset is obtained, and the first risk prediction result is merged with the historical dataset to obtain a new training dataset; the food risk prediction model is trained using the new dataset to obtain a trained food risk prediction model; and the trained food risk prediction model is used to perform risk prediction on target data to obtain a risk prediction result for the target data. By inputting the data to be predicted into the food risk prediction model for risk prediction to obtain a first risk prediction result, obtaining a historical dataset, and merging the first risk prediction result with the historical dataset to obtain a new training dataset, the food risk prediction model is trained using the new dataset to obtain a trained food risk prediction model, and the trained food risk prediction model is used to perform risk prediction on target data to obtain a risk prediction result for the target data, this method can address the problems of data acquisition difficulties in imported food risk inference algorithm models, the complexity of transnational regulations, and the high difficulty in assessment due to the diversity of food types. This method can effectively improve the identification and assessment of potential food safety risks, thereby enhancing overall regulatory efficiency.
[0070] It is understandable that in the specific implementation of this application, text data, voice data, task data and other related data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data, as well as the training, deployment and calling of large language models, must comply with relevant laws, regulations and standards of relevant countries and regions.
[0071] Optionally, in the step of inputting the data to be predicted into the food risk prediction model for risk prediction and obtaining a first risk prediction result, the data to be predicted and the food risk prediction model can be obtained; risk prediction is performed on the data to be predicted through the food risk prediction model to obtain a first risk prediction result.
[0072] In an embodiment of the present invention, the data to be predicted includes food ingredient data, inspection data, production data, etc.
[0073] The above-mentioned food risk prediction model can be a food risk prediction model built based on deep learning or machine learning, and is obtained by training an untrained food risk prediction model based on historical sample data. The above-mentioned food risk prediction model can be a convolutional neural network (CNN), a recurrent neural network (RNN), etc. The above-mentioned food risk prediction model combines the stochastic gradient algorithm, the forest decision tree algorithm, etc., and is used to predict and evaluate different types of risk factors. The above-mentioned stochastic gradient algorithm can be understood as a method of optimizing model parameters by calculating the gradient by randomly selecting samples and gradually updating the weights to minimize the loss function. The above-mentioned decision tree is a method of making predictions by dividing the dataset into multiple subsets and constructing a decision tree on each subset. The above-mentioned random forest improves the generalization ability of the model by reducing the risk of overfitting. The above-mentioned forest decision tree algorithm can be understood as improving prediction accuracy and reducing the risk of overfitting by combining multiple decision trees.
[0074] The above risk prediction can be understood as a process of evaluating the potential risks of food based on analysis of input data.
[0075] The first risk prediction result may be an assessment of possible safety risks, health risks or other related risks of the data to be predicted.
[0076] Optionally, in the step of merging the first risk prediction result with the historical data set to obtain a new training data set, the first risk prediction result and the data to be predicted can be used as the first training data set, and the first training data set can be merged with the historical data set to obtain a second training data set; risk analysis processing is performed on the second training data set, and data balance sampling processing is performed on the second training data set after risk analysis processing through interpolation method to obtain a new training data set.
[0077] In an embodiment of the present invention, the first risk prediction result is obtained by performing risk prediction on the data to be predicted using a food risk prediction model, including an assessment of potential safety risks, health risks, or other related risks associated with the data to be predicted. The data to be predicted may include food ingredient data, inspection data, production data, and the like.
[0078] The above historical data sets include a large amount of historical food safety data, historical inspection records, market feedback, logistics tracking information, etc., as well as corresponding food risk prediction result labels.
[0079] The above-mentioned merging process can be understood as a process of combining different data sources or data sets together.
[0080] The risk analysis process described above can be understood as a computational process that essentially performs general risk analysis and product analysis on the second training dataset. First, a preliminary risk assessment is performed on the second training dataset to identify potential risk points or outliers. This is followed by a detailed product analysis of product performance, reliability, and other aspects. The purpose of the risk analysis process is to enhance subsequent training datasets. Statistical calculations can be used to perform risk analysis, including but not limited to calculating key indicators such as risk frequency and risk percentage. The risk frequency can be the number of times a particular risk occurs, and the percentage can be the relative proportion of a particular risk among all risks.
[0081] The above-mentioned data balanced sampling process is a data oversampling method for dealing with class imbalance problems. By generating synthetic samples between the minority classes, the number of minority class samples is increased. The goal is to increase the number of minority class samples by generating synthetic samples between the minority class samples. New samples can be generated by interpolation methods instead of simply copying the minority class samples. For a given minority class sample,
[0082] The above interpolation method can be understood as the process of constructing new data points through known data points. The purpose is to find a function that can accurately pass through these known points to estimate the value of the unknown point.
[0083] In one possible implementation, for example, if data for a certain risk category is found to be insufficient or unbalanced, targeted sampling or data generation methods can be used to increase the amount of data for that category, thereby achieving balanced sample data. During data processing and enhancement, it is crucial to ensure that the dataset used for model training remains balanced. This means that data from different risk categories should be evenly distributed to avoid prediction bias caused by data skew. A balanced dataset can be used to train more accurate and reliable risk prediction models, thereby improving the accuracy and effectiveness of overall data analysis.
[0084] Optionally, the interpolation formula is as follows:
[0085] x new =x i +λ×(x nn -x i )
[0086] Among them, x i is a minority class sample; x nn is a nearest neighbor minority class sample of i; λ is a random number in the interval [0,1].
[0087] In this embodiment of the present invention, the value of the minority class sample of the position can be estimated by using the known minority class samples. i is a minority class sample; xnn is a nearest neighbor minority class sample of i, that is, i The nearest minority class sample; λ is a random number or empirical coefficient in the interval [0,1], which is used to control the degree of interpolation, that is, x new How close is it to x i or x nn of.
[0088] Specifically, this formula describes how to generate a new sample from a given minority class sample and its nearest neighbor sample, where λ is a random number used to generate a new sample in x. i and x nn Interpolate between to generate a new sample point. This new sample point will be located at x i and x nn The specific position is determined by the value of λ.
[0089] Furthermore, for each minority class sample, first find its K nearest neighbor minority class samples, randomly select a nearest neighbor, and generate a new sample according to the interpolation formula. Repeat the process of generating new samples several times until enough samples are generated.
[0090] In one possible embodiment, when there is a minority class sample x i and a nearest minority class sample x nn When x i and x nn When λ=0, x new =x i ; When λ=1, x new =x nn ; By choosing a random number with a lambda value between [0,1], a i and x nn New samples x between new .
[0091] In a possible embodiment, the interpolation iteration is multiple times, and the interpolation iteration is performed once in each round of training. When the interpolation iteration is performed for the first time, the minority class sample x can be determined. i The number N1 and its nearest neighbor minority class samples x nn The number N2 is determined based on the number N1 and the number N2.
[0092] Specifically, the formula for determining λ can be λ=(N1, N2) max / (N1+N2),(N1,N2) max Indicates the larger of the number N1 and the number N2.
[0093] In subsequent interpolation iterations, λ needs to be dynamically adjusted according to the number of interpolation iterations so that the interpolated sample is closer to the minority class sample x i .
[0094] Specifically, the formula for determining λ can be λ=[(N1,N2) max / (N1+N2)] m*(|N1-N2|) , m is the number of interpolation iterations.
[0095] The main purpose of the interpolation method is to increase the number of minority class samples by generating new subsets between minority class samples, thereby balancing the class distribution in the training set and improving the performance of the model on imbalanced data.
[0096] Optionally, in the step of training the food risk prediction model using a new training data set to obtain a trained food risk prediction model, the new training data set can be input into the food risk prediction model for training; during the training process, the parameters of the food risk prediction model are adjusted using a cross-validation calculation method. After the training is completed, a trained food risk prediction model is obtained. The cross-validation calculation method formula is as follows:
[0097]
[0098] Where: γ is the loss function, D val is the validation set, is the k-th fold validation set, and θ is the model parameter.
[0099] In an embodiment of the present invention, the cross-validation calculation method is to divide the dataset into K subsets and train and validate the model on different training-validation sets. In each iteration, one subset is used as the validation set and the remaining subsets are used as the training set. This ensures that each data point is used as the validation set once, thereby obtaining a more comprehensive performance evaluation. The cross-validation calculation method can be a 3-fold cross-validation (K=3). In each iteration, the model updates the model parameters based on the training set and calculates the error on the validation machine to minimize the loss function to select the optimal parameter combination.
[0100] The above loss function is used to measure the difference between the model's prediction and the actual result. It calculates a numerical value to represent the accuracy or error of the model's prediction. The above loss function can be a loss function such as mean squared error (MSE) or absolute error (MAE).
[0101] The above D val is the validation set, a subset of data on which the model performance is evaluated. In each fold, the model is trained on the training set and evaluated on the validation set.
[0102] above It is the validation set of the kth fold. The dataset is divided into k parts, and k-1 parts are used as training sets each time, and the remaining part is used as the validation set.
[0103] The above θ is a model parameter that needs to be adjusted through an optimization algorithm to minimize the loss function.
[0104] It should be noted that by using cross-validation methods to train and evaluate food risk prediction models, new data can be more effectively utilized to improve the performance of the model and ensure that the model has better generalization ability on new data.
[0105] Optionally, in the step of training the food risk prediction model using a new training data set to obtain a trained food risk prediction model, incremental training can be used to train the food risk prediction model for each new training data set. The incremental training formula is as follows:
[0106]
[0107] Where f represents the mapping function, X represents the input feature vector, and θ represents the parameters of the model, which are iteratively updated using the following loss function:
[0108]
[0109] Where: t is the number of iterations, η is the learning rate, is the loss function.
[0110] In an embodiment of the present invention, incremental training can be used to fit the model for each new training data set. The above fitting can be understood as adjusting the model parameters through an optimization algorithm to minimize the difference between the model prediction results and the actual labels.
[0111] The above incremental training formula indicates that in each iteration, the parameters are updated along the gradient of the loss function to minimize the loss function and thus optimize the model performance.
[0112] Incremental training is a machine learning method that allows a model to gradually update itself as it receives new data, rather than retraining from scratch. In incremental training, the model receives new data in batches and is updated immediately after each batch to adapt to changes in the new data.
[0113] It should be noted that during incremental training, whenever a new training dataset is available, the model will use this new training dataset to update the parameters of the model itself, thereby gradually improving its prediction performance.
[0114] Specifically, for each new training dataset, the model first uses the current parameter settings to make predictions; then, the error between the predicted result and the actual label is calculated based on the loss function, and the model parameters are updated based on this error; finally, the updated parameters are used to re-predict the new training dataset, and this step is repeated until the predetermined expected effect is achieved.
[0115] Optionally, in the step of training the food risk prediction model using a new training data set to obtain a trained food risk prediction model, the F1 Score evaluation method can be used to adjust the parameters of the food risk prediction model during the training process. The formula of the F1 Score evaluation method is as follows:
[0116]
[0117] Among them, F1 is the F1 Score evaluation, Precision is the precision rate, and Recall is the recall rate; the F1 score formula for the weighted method is as follows:
[0118]
[0119] Where: C is the total number of categories, F1 i is the F1 score of the i-th class, w i is the sample proportion of the i-th category.
[0120] In this embodiment of the present invention, the F1 Score is an evaluation metric that comprehensively considers both precision and recall, and is used to measure the performance of a model on a classification task. By adjusting model parameters, the F1 Score can be improved, thereby enhancing the model's prediction accuracy.
[0121] The above precision rate can be understood as the proportion of samples that are actually positive among the samples predicted by the model to be positive.
[0122] The above recall rate can be understood as the proportion of positive samples that the model can correctly identify to all actual positive samples.
[0123] For the weighted F1 score, it means that the F1 score of each category is weighted and summed according to the category weight.
[0124] In one possible embodiment, the number of samples in different classes may vary significantly. For example, the number of samples in the positive class may be far fewer than the number of samples in the negative class. In this case, even if the model's predictions for the positive class are very accurate, the overall performance may be "dragged down" by the negative class due to the small number of positive samples. The Weighted F1 Score, by considering the proportion of samples in each class, can more fairly evaluate the overall performance of the model in cases of data imbalance.
[0125] In another possible embodiment, GridSearchCV can be used to perform hyperparameter search, with the goal of selecting a set of parameters that optimizes model performance. The scoring criterion is weighted F1-score:
[0126]
[0127] Among them, F1 weighted is the weighted F1-score, c is the category, F1 c is the F1-score of the c-th class, N c is the number of samples in class c.
[0128] The aforementioned GridSearchCV is a cross-validation tool for hyperparameter tuning. It uses a grid search method to enumerate all possible parameter combinations, trains a model using these parameters, and then evaluates the performance of each parameter combination. The F1-score is a metric for evaluating classifier performance, taking into account both precision and recall. The weighted F1-score considers the importance of different categories; that is, if a category has a large number of samples, its influence on the F1-score calculation is smaller. The number of cross-validation folds used during hyperparameter search is used to evaluate the model's performance under different parameter combinations.
[0129] In summary, GridSearchCV can be used for hyperparameter tuning to find the optimal model parameters so that the model performs best under the weighted F1-score evaluation criterion.
[0130] like Figure 2 As shown, Figure 2 This is a structural diagram of a food risk prediction method provided by an embodiment of the present invention. Specifically, it includes steps ①: inputting the data to be predicted; step ②: food risk prediction model; step ③: obtaining the prediction result; step ④: putting the predicted data into the model again; step ⑤: inputting the original data; step ⑥: loading historical data; step ⑦: data preprocessing; step ⑧: risk analysis; step ⑨: high-risk sample enhancement; step ⑩: data balanced sampling; step Load / initialize the model; steps Parameter optimization; steps Model training; steps Save the model; steps Generate result report; steps Training completed.
[0131] In this embodiment, the data to be predicted are data that need to be predicted, including food ingredient data, inspection data, production data, etc.
[0132] The above-mentioned food risk prediction model is trained based on a sample training data set. The above-mentioned sample training data set includes a large amount of historical food safety data, historical inspection records, market feedback, logistics tracking information, etc., as well as corresponding food risk prediction result labels. The above-mentioned food risk prediction model can be a food risk prediction model built based on deep learning or machine learning, for example, it can be a convolutional neural network (CNN), a recurrent neural network (RNN), etc. The above-mentioned food risk prediction model combines machine learning and big data analysis technology, and automatically identifies and evaluates different types of risk factors by integrating multiple data sources, thereby reducing the time cost and resource investment of manual review, and improving the efficiency of risk assessment, thereby supporting the food safety supervision needs of the global market.
[0133] Furthermore, the data to be predicted can be input into a food risk prediction model to perform food risk prediction and obtain a risk prediction result.
[0134] The above-mentioned original data can be understood as taking the data to be predicted and the risk prediction results as the original data.
[0135] The above-mentioned loading of historical data can be understood as loading historical data into the food risk prediction model.
[0136] The above data preprocessing can be understood as a complete merging process based on historical data and original data to obtain a new training data set.
[0137] The above-mentioned risk analysis can be understood as general risk analysis and product analysis of the new training dataset.
[0138] The above-mentioned high-risk sample enhancement can be understood as an enhancement process for high-risk sample data.
[0139] The above-mentioned data balancing sampling can be understood as performing data balancing sampling on the training dataset after risk analysis through interpolation to obtain a new training dataset. By generating synthetic samples between minority classes, the number of minority class samples is increased. The goal is to increase the number of minority class samples by generating synthetic samples between minority class samples.
[0140] The above-mentioned loading / initialization model can be understood as loading the food risk prediction model / initializing the food risk prediction model.
[0141] The above parameter optimization can be understood as adjusting model parameters to improve prediction results.
[0142] The above model training can be understood as using a new training dataset to train a food risk prediction model.
[0143] The above-mentioned model preservation can be understood as preserving the trained food risk prediction model for subsequent use.
[0144] The above-mentioned generated result report can be understood as generating a detailed report of the risk prediction results.
[0145] The above training is completed, indicating that the entire model training is completed.
[0146] In this embodiment, the present invention can identify the specific risks faced by different types of food through automated data collection and processing, combined with deep learning algorithms. At the same time, the model has self-learning capabilities, and can optimize the evaluation results as new data is continuously input, thereby enhancing the stability and adaptability of the model. The present invention conducts real-time risk monitoring and early warning of imported foods, and has the advantages of fast response, high accuracy and low application cost. The food risk prediction model adopted by the present invention can automatically identify risk factors and evaluation criteria, and can intelligently conduct a comprehensive analysis of the safety of imported foods, reducing the time cost and resource investment of manual review, and improving the efficiency of risk assessment, thereby supporting the food safety supervision needs of the global market.
[0147] like Figure 3 As shown, Figure 3 This is a diagram of another food risk prediction method provided by an embodiment of the present invention. Specifically, it includes steps ①: data loading and preprocessing; ②: feature extraction; ③: sample balancing and data screening; ④: model parameter optimization; ⑤: model training; ⑥: obtaining product risk analysis rates using the model; ⑦: general risk analysis; and ⑧: ending.
[0148] In this embodiment, data loading can be understood as loading a variety of data, including historical data, market trends, and consumer feedback. Data preprocessing primarily provides clean, uniformly formatted data for model training. This includes handling missing values and formatting to ensure consistent classification feature types, thereby facilitating a smoother feature encoding process.
[0149] The above feature extraction can be understood as encoding categorical variables, converting discrete text features into numerical sparse matrices. One-Hot encoding generates a unique hot vector to represent each category, allowing the model to process these features. i One-Hot encoding of , the encoding result is:
[0150] e i =[0,…,0,1,0,…,0]
[0151] Among them, 1 corresponds to category C i , and the rest are 0.
[0152] The above sample balance and data screening can be understood as follows: in classification tasks, if there are too few samples in certain categories, it will lead to model bias. Therefore, SMOTE (Synthetic Minority Over-sampling Technique) can be used to generate new samples for the minority class to achieve sample balance.
[0153] The principle of SMOTE is to generate new samples by interpolation: for a minority class sample x i , from its nearest neighbor (assuming x j ) is randomly selected from .
[0154] Generate new synthetic samples:
[0155] x new =x i +λ·(x j -x i )
[0156] Among them, x i is a minority class sample; x nn is a nearest neighbor minority class sample of i; λ is a random number in the interval [0,1].
[0157] The above formula describes how SMOTE generates a new sample based on a given minority class sample and its nearest neighbor sample. i and x nn Interpolate between to generate a new sample point. This new sample point will be located at x i and x nn The specific position of the line connecting the two classes is determined by the value of λ. For each minority class sample, its k nearest neighbor minority class samples are first found. A nearest neighbor is randomly selected and a new sample is generated according to the interpolation formula. This process is repeated several times until a sufficient number of synthetic samples are generated. The main purpose of sample balancing and data screening is to increase the number of minority class samples by generating new samples between minority class samples, thereby balancing the class distribution in the training set and improving the performance of machine learning models on imbalanced data.
[0158] The model parameter optimization described above can be understood as a hyperparameter search using GridSearchCV, aiming to select the set of parameters that optimizes model performance. GridSearchCV is a Scikit-learn tool for parameter tuning that combines grid search and cross-validation to find optimal model parameters. Cross-validation refers to the number of folds used in the hyperparameter search to evaluate model performance.
[0159] The scoring metric is the weighted F1-score, a commonly used metric for evaluating classification performance that balances precision and recall. The weighted F1-score is a weighted average of the F1-scores for different categories, with the weights typically determined based on the amount of data in each category.
[0160] The formula for weighted F1-score is as follows:
[0161]
[0162] Among them, F1 weighted is the weighted F1-score, c is the category, F1 c is the F1-score of the c-th class, N c is the number of samples in class c.
[0163] The above model training can be understood as training based on a random forest or decision tree model. A decision tree model uses a tree-like structure to perform classification or regression predictions on sample features. The path from the root node to the leaf nodes represents the decision-making process. A random forest model is composed of multiple decision trees, using an ensemble approach to improve prediction accuracy.
[0164] The above product risk analysis rate can be understood as identifying and reducing the possibility of failure by evaluating possible adverse events that may occur during product design and manufacturing.
[0165] The above general risk analysis can be understood as identifying, assessing and responding to potential risk events.
[0166] In this embodiment, the present invention integrates multiple information sources, including historical data, market trends, and consumer feedback, to generate a more accurate and contextually accurate risk assessment report, thereby enhancing the effectiveness and reliability of decision support. This multi-layered data integration ensures that the assessment results are both scientifically sound and adaptable to complex market environments.
[0167] By employing a food risk prediction model, the present invention can transform food safety data from diverse sources and types into a unified risk feature vector (the similarities between different food types under different risk indicators are mapped to the same feature representation), thereby enabling comprehensive risk analysis of imported food. When training the model, the present invention focuses solely on data on key food safety indicators, which are more readily available than comprehensive monitoring data. Furthermore, compared to conventional monitoring methods, the use of a real-time data update algorithm can promptly reflect market dynamics and potential risks.
[0168] This invention utilizes a food risk prediction model to implement an artificial intelligence-based food safety risk assessment system, overcoming the shortcomings of traditional monitoring methods while maximizing the advantages of data analysis and real-time assessment. This model not only improves the accuracy of risk assessments but also leverages dynamic data sources to enhance responsiveness to potential risks, effectively alleviating the challenges of traditional methods in data collection and processing.
[0169] It should be noted that this application can provide relevant companies with comprehensive risk assessment, compliance advice, and market dynamics analysis capabilities. It can simultaneously achieve high accuracy, real-time response, high reliability, and multi-dimensional data analysis, providing solid technical support for food safety management.
[0170] like Figure 4 As shown, an embodiment of the present invention provides a food risk prediction device, which includes:
[0171] A first prediction module 401 is configured to input the data to be predicted into a food risk prediction model for risk prediction to obtain a first risk prediction result, wherein the food risk prediction model is trained based on a sample training data set;
[0172] A processing module 402 is configured to obtain the historical data set, and merge the first risk prediction result with the historical data set to obtain a new training data set;
[0173] A training module 403 is configured to train the food risk prediction model using the new data set to obtain a trained food risk prediction model;
[0174] The second prediction module 404 is used to perform risk prediction on the target data using the trained food risk prediction model to obtain a risk prediction result for the target data.
[0175] Optionally, the first prediction module 401 is further used to obtain data to be predicted and a food risk prediction model; perform risk prediction on the data to be predicted using the food risk prediction model to obtain a first risk prediction result.
[0176] Optionally, the processing module 402 is also used to use the first risk prediction result and the data to be predicted as a first training data set, and merge the first training data set with the historical data set to obtain a second training data set; perform risk analysis processing on the second training data set, and perform data balance sampling processing on the second training data set after risk analysis processing through an interpolation method to obtain a new training data set.
[0177] Optionally, the interpolation formula is as follows:
[0178] x new =x i +λ×(x nn -x i )
[0179] Among them, x i is a minority class sample; x nm is a nearest neighbor minority class sample of i; λ is a random number in the interval [0,1].
[0180] Optionally, the training module 403 is further configured to input the new training data set into the food risk prediction model for training;
[0181] During the training process, the parameters of the food risk prediction model are adjusted using a cross-validation calculation method. After the training is completed, a trained food risk prediction model is obtained. The cross-validation calculation method formula is as follows:
[0182]
[0183] Where: γ is the loss function, D cal is the validation set, is the k-th fold validation set, and θ is the model parameter.
[0184] Optionally, the training module 403 is further configured to train the food risk prediction model using incremental training for each new training data set. The incremental training formula is as follows:
[0185]
[0186] Where f represents the mapping function, X represents the input feature vector, and θ represents the parameters of the model, which are iteratively updated using the following loss function:
[0187]
[0188] Where: t is the number of iterations, η is the learning rate, is the loss function.
[0189] Optionally, the training of the food risk prediction model using the new training data set to obtain a trained food risk prediction model includes:
[0190] During the training process, the parameters of the food risk prediction model were adjusted using the F1 Score evaluation method. The formula of the F1 Score evaluation method is as follows:
[0191]
[0192] Among them, F1 is the F1 Score evaluation, Precision is the precision rate, and Recall is the recall rate; the F1 score formula for the weighted method is as follows:
[0193]
[0194] Where: C is the total number of categories, F1 i is the F1 score of the i-th class, w i is the sample proportion of the i-th category.
[0195] like Figure 5 As shown, an embodiment of the present invention further provides an electronic device, including a processor, and the processor can execute any of the above-mentioned food risk prediction methods.
[0196] Specifically, the method includes a processor 501, a memory 502, and a computer program for executing a food risk prediction method stored in the memory 502 and capable of running on the processor 501, wherein:
[0197] The processor 501 runs the computer program of the food risk prediction method stored in the memory 502 and performs the following steps:
[0198] Inputting the data to be predicted into a food risk prediction model for risk prediction to obtain a first risk prediction result, wherein the food risk prediction model is trained based on a sample training data set;
[0199] Acquire the historical data set, and merge the first risk prediction result with the historical data set to obtain a new training data set;
[0200] Training the food risk prediction model using the new data set to obtain a trained food risk prediction model;
[0201] The risk prediction of the target data is performed using the trained food risk prediction model to obtain the risk prediction results of the target data.
[0202] Optionally, the processor 501 inputs the data to be predicted into the food risk prediction model to perform risk prediction to obtain a first risk prediction result, including:
[0203] Obtain the data to be predicted and the food risk prediction model;
[0204] The food risk prediction model is used to perform risk prediction on the data to be predicted to obtain a first risk prediction result.
[0205] Optionally, the processor 501 performs the merging of the first risk prediction result with the historical data set to obtain a new training data set, including:
[0206] Taking the first risk prediction result and the data to be predicted as a first training data set, and merging the first training data set with the historical data set to obtain a second training data set;
[0207] Perform risk analysis processing on the second training data set, and perform data balancing sampling processing on the second training data set after the risk analysis processing by an interpolation method to obtain a new training data set.
[0208] Optionally, the interpolation formula executed by the processor 501 is as follows:
[0209] x new =x i +λ×(x nn -x i )
[0210] Among them, x i is a minority class sample; x nn is a nearest neighbor minority class sample of i; λ is a random number in the interval [0,1].
[0211] Optionally, the processor 501 performs the training of the food risk prediction model using the new training data set to obtain a trained food risk prediction model, including:
[0212] Inputting the new training data set into the food risk prediction model for training;
[0213] During the training process, the parameters of the food risk prediction model are adjusted using a cross-validation calculation method. After the training is completed, a trained food risk prediction model is obtained. The cross-validation calculation method formula is as follows:
[0214]
[0215] Where: γ is the loss function, D val is the validation set, is the k-th fold validation set, and θ is the model parameter.
[0216] Optionally, the processor 501 performs the training of the food risk prediction model using the new training data set to obtain a trained food risk prediction model, including:
[0217] For each new training data set, incremental training is used to train the food risk prediction model. The incremental training formula is as follows:
[0218]
[0219] Where f represents the mapping function, X represents the input feature vector, and θ represents the parameters of the model, which are iteratively updated using the following loss function:
[0220]
[0221] Where: t is the number of iterations, η is the learning rate, is the loss function.
[0222] Optionally, the processor 501 performs the training of the food risk prediction model using the new training data set to obtain a trained food risk prediction model, including:
[0223] During the training process, the parameters of the food risk prediction model were adjusted using the F1 Score evaluation method. The formula of the F1 Score evaluation method is as follows:
[0224]
[0225] Among them, F1 is the F1 Score evaluation, Precision is the precision rate, and Recall is the recall rate; the F1 score formula for the weighted method is as follows:
[0226]
[0227] Where: C is the total number of categories, F1 i is the F1 score of the i-th class, w i is the sample proportion of the i-th category.
[0228] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the various processes of the food risk prediction method or the application-end food risk prediction method provided by the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0229] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0230] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A food risk prediction method, characterized in that: The method comprises the following steps: Inputting the data to be predicted into a food risk prediction model for risk prediction to obtain a first risk prediction result, wherein the food risk prediction model is trained based on a sample training data set; Acquire the historical data set, and merge the first risk prediction result with the historical data set to obtain a new training data set; Training the food risk prediction model using the new data set to obtain a trained food risk prediction model; The risk prediction of the target data is performed using the trained food risk prediction model to obtain the risk prediction results of the target data.
2. The food risk prediction method according to claim 1, wherein: The step of inputting the data to be predicted into the food risk prediction model to perform risk prediction and obtain a first risk prediction result includes: Obtain the data to be predicted and the food risk prediction model; The food risk prediction model is used to perform risk prediction on the data to be predicted to obtain a first risk prediction result.
3. The food risk prediction method according to claim 2, wherein: The merging of the first risk prediction result with the historical data set to obtain a new training data set includes: Taking the first risk prediction result and the data to be predicted as a first training data set, and merging the first training data set with the historical data set to obtain a second training data set; Perform risk analysis processing on the second training data set, and perform data balancing sampling processing on the second training data set after the risk analysis processing by an interpolation method to obtain a new training data set.
4. The food risk prediction method according to claim 3, wherein: The interpolation formula is as follows: x new =x i +λ×(x nn -x i ) Among them, x i is a minority class sample; x nn is a nearest neighbor minority class sample of i; λ is a random number in the interval [0,1].
5. The food risk prediction method according to claim 4, wherein: The method of training the food risk prediction model using the new training data set to obtain a trained food risk prediction model includes: Inputting the new training data set into the food risk prediction model for training; During the training process, the parameters of the food risk prediction model are adjusted using a cross-validation calculation method. After the training is completed, a trained food risk prediction model is obtained. The cross-validation calculation method formula is as follows: Where: γ is the loss function, D val is the validation set, is the k-th fold validation set, and θ is the model parameter.
6. The food risk prediction method according to claim 4, wherein: The method of training the food risk prediction model using the new training data set to obtain a trained food risk prediction model includes: For each new training data set, incremental training is used to train the food risk prediction model. The incremental training formula is as follows: Where f represents the mapping function, X represents the input feature vector, and θ represents the parameters of the model, which are iteratively updated using the following loss function: Where: t is the number of iterations, η is the learning rate, is the loss function.
7. The food risk prediction method according to claim 6, wherein: The method of training the food risk prediction model using the new training data set to obtain a trained food risk prediction model includes: During the training process, the parameters of the food risk prediction model were adjusted using the F1 Score evaluation method. The formula of the F1 Score evaluation method is as follows: Among them, F1 is the F1 Score evaluation, Precision is the precision rate, and Recall is the recall rate; the F1 score formula for the weighted method is as follows: Where: C is the total number of categories, F1 i is the F1 score of the i-th class, w i is the sample proportion of the i-th category.
8. A food risk prediction device, characterized in that: The food risk prediction device comprises: A first prediction module is used to input the data to be predicted into a food risk prediction model to perform risk prediction and obtain a first risk prediction result, wherein the food risk prediction model is trained based on a sample training data set; a processing module, configured to obtain the historical data set, and merge the first risk prediction result with the historical data set to obtain a new training data set; A training module, configured to train the food risk prediction model using the new data set to obtain a trained food risk prediction model; The second prediction module is used to perform risk prediction on the target data using the trained food risk prediction model to obtain the risk prediction results of the target data.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the food risk prediction method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the food risk prediction method according to any one of claims 1 to 7.