Bayesian network-based food adulteration prediction method and system

The food adulteration database was constructed through the Bayesian network, and the food chain characteristics were extracted and the hierarchical warning was performed, which solved the problems of low food adulteration detection efficiency and limitations in data analysis, and achieved efficient and accurate food adulteration prediction.

CN120373552AActive Publication Date: 2025-07-25CHINESE ACAD OF INSPECTION & QUARANTINE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510465904.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-25
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art is inefficient in food adulteration detection, making it difficult to fully cover complex supply chain links, and data analysis is limited to the food itself, resulting in insufficient accuracy and timeliness of the prediction results.

Method used

The food adulteration prediction method based on Bayesian network is adopted to construct a food adulteration database, extract the food chain characteristics, calculate the comprehensive similarity, perform hierarchical early warnings, and dynamically update the Bayesian network to enhance the adaptability of the model to adulteration mode.

Benefits of technology

It improves the efficiency and accuracy of food adulteration prediction, can detect potential adulteration risks in the early stage, save resources, adapt to the needs of different food adulteration prediction systems, and is universal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373552A_ABST
    Figure CN120373552A_ABST
Patent Text Reader

Abstract

The invention discloses a Bayesian network-based food adulteration prediction method and system, and the method comprises the steps: constructing a food adulteration database, processing the food information and production and marketing logs of to-be-predicted food to obtain food chain features, calculating the comprehensive similarity between the food chain features and standard food chain features, and predicting the food adulteration. Determining adulteration early warning data and adulteration risk data according to the comprehensive similarity, directly performing adulteration early warning according to the adulteration early warning data, and determining a Bayesian network for predicting food adulteration according to the adulteration risk data. And inputting the adulteration risk data into the food adulteration prediction Bayesian network to obtain an adulteration category and an adulteration probability, and updating a food adulteration database by adopting the food adulteration prediction Bayesian network and the corresponding adulteration representation vector. The method not only can improve the food adulteration prediction efficiency and accuracy, but also has good interpretability, and can be directly applied to a food adulteration prediction system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of food quality, and particularly to a food adulteration prediction method and system based on a Bayesian network. Background Art

[0002] With the increasing complexity and globalization of the food industry supply chain, food adulteration behaviors have become more concealed and diverse in form, seriously damaging consumers' rights and interests, endangering public health, and also causing economic losses and reputational damage to food enterprises. Against this background, accurate and efficient food adulteration prediction technology has become the key to ensuring food safety and maintaining market stability, and its development has received high attention from the food industry, regulatory authorities, and the general public.

[0003] Traditional food adulteration detection mostly relies on post-event sampling inspection. First of all, this method is inefficient and difficult to comprehensively cover various types of foods and complex supply chain links. Secondly, some data analysis-based prediction technologies have limitations in data integration, often only focusing on the quality inspection data of the food itself and ignoring key information such as production and sales links, resulting in a significant reduction in the accuracy and timeliness of prediction results. In response to the above problems, the rise of new technologies such as big data and artificial intelligence has pointed out a new direction for solving the problem of food adulteration prediction. The present invention proposes a food adulteration prediction method and system based on a Bayesian network, which can efficiently save computing power by processing food data in different states through hierarchical early warning prediction, introduce incremental learning and dynamic topology optimization to dynamically update the Bayesian network, enhance the adaptability of the model to rapidly changing adulteration patterns, improve the dynamic assessment and accurate prediction of food adulteration risks, effectively overcome the deficiencies of traditional technologies, provide a new and efficient adulteration prediction tool for the food industry, help to timely discover potential adulteration risks, reduce the occurrence of food safety incidents, and improve the overall safety level of the food industry. Summary of the Invention

[0004] The purpose of the present invention is to provide a food adulteration prediction method and system based on a Bayesian network.

[0005] To achieve the above object, the present invention is implemented according to the following technical solution:

[0006] The present invention includes the following steps:

[0007] Obtain historical food data and historical production and sales data to construct a food adulteration database; the food adulteration database includes an initial food adulteration Bayesian network and a corresponding network characterization vector;

[0008] Extract the food information and production and sales logs of the food to be predicted, perform data processing to obtain food data and production and sales data, and form a set of food chain features by combining the food data features and the production and sales data features;

[0009] Determine the characteristics of the standard food chain, calculate the comprehensive similarity between the food chain characteristics and the standard food chain characteristics, and determine the adulteration warning data and adulteration risk data based on the comprehensive similarity;

[0010] Conduct adulteration warning directly based on the adulteration warning data, determine the adulteration characterization vector and selection strategy based on the adulteration risk data, and obtain the predicted food adulteration Bayesian network according to the selection strategy; the selection strategy includes reusing the initial food adulteration Bayesian network, optimizing the initial food adulteration Bayesian network, and creating a new initial food adulteration Bayesian network;

[0011] Input the adulteration risk data into the predicted food adulteration Bayesian network to obtain the adulteration category and adulteration probability, and update the food adulteration database using the predicted food adulteration Bayesian network and the corresponding adulteration characterization vector.

[0012] Furthermore, the method for constructing the food adulteration database includes:

[0013] Obtain historical food data and historical production and sales data; the historical food data includes historical food quality inspection data, historical food image features, historical food public opinion features, historical food categories, and historical food adulteration conclusions; the historical production and sales data includes historical raw material procurement data, historical processing technologies, historical storage and transportation data, and historical sales records;

[0014] Align the data according to food categories, time, and region to form a spatio-temporal data cube, and determine the observation nodes, hidden nodes, and decision nodes according to the data categories; the observation nodes include quality inspection indicators, image feature similarity, public opinion risk values, raw material price volatility, number of times of exceeding storage and transportation temperature standards, deviation of storage and transportation trajectory points, sterilization temperature compliance rate, compliance of additive use, channel type, and abnormal return rate index; the hidden node is the adulteration motivation; the decision node is the adulteration category;

[0015] Define strong causal relationships using domain experience to determine the key edges, search for the optimal structure based on the K2 algorithm and the hill climbing algorithm, and perform dynamic pruning based on time-decayed mutual information and intervention experiment causal stability to determine the Bayesian network structure. The dynamic pruning conditions are:

[0016]

[0017] where KeepEdge(X→Y) is the condition for retaining the edge (X→Y) from the observation node X to Y, I τ (X,Y) is the time-decayed mutual information, θ I is the time-decayed mutual information threshold, S causal (X,Y) is the causal stability score, θ Sis the causal stability score threshold, T is the total length of the time window, λ is the decay coefficient, I(X t ,Y t ) is the mutual information value at time t within the time window, used to measure the correlation between variables X t and Y t at time t, N int is the number of intervention experiments, is the quantification of the causal effect, P[Y|do(X=x k )] is the probability distribution of Y under the intervention operation x k of X;

[0018] Input the spatio-temporal data cube into the Bayesian network structure for parameter learning to generate a conditional probability table, and determine the initial food adulteration Bayesian network according to the Bayesian network structure and the conditional probability table; when the spatio-temporal data cube is complete, the maximum likelihood estimation is used to generate the conditional probability table; when the spatio-temporal data cube is missing, the Bayesian estimation combined with the Dirichlet prior is used to generate the conditional probability table, and the expression is:

[0019]

[0020] where P(X=x i |Pa(X)=pa j ) is the conditional probability that node X takes the value x j when the parent node Pa(X) takes the value pa i , is the number of observations where X=x i and Pa(X)=pa j , is the number of observations where Pa(X)=pa j , is the dynamic prior parameter at time t, K is the number of all possible values of X, C source is the data source confidence weight, β is the weight coefficient of the data source confidence, γ is the weight coefficient of the time decay, δ is the weight coefficient of the expert knowledge injection, ExpertPrior(x i ,pa j ) is the expert-prescribed prior probability when X=x i and Pa(X)=pa j ;

[0021] Extract the food category, production and sales time, and production and sales region from the historical food data and historical production and sales data, splice them into a network representation vector, associate the network representation vector with the corresponding initial food adulteration Bayesian network, and construct different category initial food adulteration Bayesian networks and network representation vectors according to different category historical food data and historical production and sales data to form a food adulteration database.

[0022] Furthermore, the method for determining the food chain characteristics includes:

[0023] Extract the food production and sales log, and use the Logstash parser to extract the key fields of the food production and sales log and output structured data to obtain production and sales data; the production and sales data includes raw material procurement data, processing technology, storage and transportation data, and sales records.

[0024] The food data includes food quality inspection data, food image features, food public opinion features, and food categories. The method for obtaining and processing food information to obtain food data includes:

[0025] Use a social crawler engine to obtain food public opinion reflections, and input the food public opinion reflections into the BERT model to generate food public opinion features; the food public opinion features include sentiment polarity scores and risk keyword frequencies.

[0026] Obtain food quality inspection data by docking with the regulatory agency database and the manufacturer's backup database.

[0027] Obtain food images taken at the production line and sales terminals, extract the color histogram of the food images to output the proportion of the main color, use ResNet-50 to perform deep features on the food images to output food texture features and shape features, use the ResNet classification model to process the food images to obtain raw material features, and combine the proportion of the main color, food texture features, shape features, and raw material features to form food image features.

[0028] Combine the food data features and production and sales data features to form a set of food chain features.

[0029] Furthermore, the method for determining adulteration warning data and adulteration risk data includes:

[0030] Determine standard food data features and standard production and sales data features according to the manufacturer's production specifications, merchant sales specifications, and market supervision requirements, and form a set of standard food chain features; the data categories and dimensions of the standard food chain features are the same as those of the food chain features.

[0031] Calculate the comprehensive similarity between the food chain feature vector and the standard food chain feature vector. The expression is:

[0032]

[0033] where Sim line is the comprehensive similarity between the food chain feature and the standard food chain feature, S food is the similarity between the food data feature and the standard food data feature, S supplyis the similarity between the production and sales data characteristics and the standard production and sales data characteristics, θ is the adjustment factor, η is the collaborative gain coefficient, I(·) is the indicator function, which takes 1 when the condition is met and 0 otherwise. is the similarity threshold of food data is the similarity threshold of production and sales data, n f is the quantity of food data is an element of the food data feature vector is an element of the standard food data feature vector, w i is the weight of the i-th element is the balance weight between the Gaussian kernel similarity and the improved Jaccard coefficient is the continuous feature vector of food data is the continuous feature vector of standard food data, σ is the smoothing parameter in the Gaussian kernel similarity is the set of discrete features of food data is the set of discrete features of standard food data, ε is the smoothing factor, φ2 is the balance weight between the Euclidean distance and the geographical grid similarity is the production and sales data feature vector of data type and the standard production and sales data feature vector of data type similarity is the production and sales data feature vector of coordinate type and the standard production and sales data feature vector of coordinate type similarity, ρ is the collaborative gain coefficient, N cover is the number of grids where two regions overlap, N total is the total number of grids of two regions, δ is the logistics cost weight coefficient is the production and sales logistics cost is the standard production and sales logistics cost;

[0034] Define the data corresponding to the food chain characteristics with a comprehensive similarity greater than 0.8 to the standard food chain characteristics as safe food data;

[0035] Define the data corresponding to the food chain characteristics with a comprehensive similarity less than 0.5 to the standard food chain characteristics as adulteration warning data;

[0036] Conversely, define the data corresponding to the remaining food chain characteristics as adulteration risk data.

[0037] Further, the method for directly giving adulteration warnings includes:

[0038] Determine early warning indicators according to the characteristics of the food chain. Taking the mean of the historical normal data of each early warning indicator as the center of the sphere, set a sliding window to dynamically adjust the radius, construct a multi-dimensional sphere to perform multi-sphere screening to obtain adulterated data, calculate the ratio of the adulterated data to the corresponding sphere radius to obtain the degree of adulteration, input the degree of adulteration and the corresponding early warning indicators into the food domain word bag model to obtain the adulteration category and adulteration weight, generate an adulteration feature vector according to the adulteration category and adulteration weight, calculate the cosine similarity between the adulteration feature vector and the historical adulteration feature vectors in the adulteration conclusion library to obtain the food adulteration conclusion, and conduct food adulteration early warning according to the food adulteration conclusion.

[0039] Further, the method for obtaining the Bayesian network for predicting food adulteration includes:

[0040] Extract the food category, production and sales time, and production and sales region in the adulteration risk data, splice them into an adulteration characterization vector, calculate the cosine similarity between the adulteration characterization vector and the network characterization vectors in the food adulteration database, and determine the selection strategy according to the maximum cosine similarity;

[0041] When the maximum cosine similarity is greater than 0.8, directly select the initial food adulteration Bayesian network associated with the network characterization vector corresponding to the maximum cosine similarity as the Bayesian network for predicting food adulteration;

[0042] When the maximum cosine similarity is less than 0.5, directly create a new Bayesian network for predicting food adulteration;

[0043] On the contrary, update the initial food adulteration Bayesian network associated with the network characterization vector corresponding to the maximum cosine similarity to obtain the Bayesian network for predicting food adulteration. The specific steps are as follows:

[0044] Set a dynamic sliding window for food data and production and sales data for incremental learning, and update the conditional probability table of the initial food adulteration Bayesian network. The expression is:

[0045]

[0046] where P new (X|Pa(X)) is the updated conditional probability table of node X when the parent node takes Pa(X), is the weight of the new and old data fusion ratio, Conf date and Var date are the data source confidence and data variance within the dynamic sliding window respectively, N new is the number of observations of the new data, N old is the number of observations of the old data, P new,date is the conditional probability estimate based on the new data, and P old is the old conditional probability table;

[0047] Calculate the causal KL divergence and the mutual information of the edges of the initial food adulteration Bayesian network after incremental update. The expression is as follows:

[0048]

[0049] Among them, KL causal (P∥Q) is the causal KL divergence after incremental update, which is used to measure the difference between two probability distributions P and Q. KL(P∥Q) is the traditional KL divergence, and ζ is the causal effect weight coefficient. is the quantification of the causal effect. is the expected value of Y under the intervention X = x i , and M int is the number of interventions;

[0050] Adjust the network nodes and edges according to the causal KL divergence KL causal (P∥Q) and the mutual information I(X,Y) of the edge X→Y. Use Bayesian optimization to select the optimal subnet structure to obtain the predicted food adulteration Bayesian network.

[0051] In the second aspect, a food adulteration prediction system based on a Bayesian network includes:

[0052] Database module: used to obtain historical food data and historical production and sales data to construct a food adulteration database; the food adulteration database includes the initial food adulteration Bayesian network and the corresponding network characterization vector;

[0053] Data module: used to process the food information and production and sales logs of the food to be predicted to obtain the food chain characteristics, calculate the comprehensive similarity between the food chain characteristics and the standard food chain characteristics, and determine the adulteration warning data and adulteration risk data according to the comprehensive similarity;

[0054] Network selection module: used to determine the adulteration characterization vector and selection strategy according to the adulteration risk data, and obtain the predicted food adulteration Bayesian network according to the selection strategy; the selection strategy includes reusing the initial food adulteration Bayesian network, optimizing the initial food adulteration Bayesian network, and creating a new initial food adulteration Bayesian network;

[0055] Warning and prediction module: used to directly conduct adulteration warnings according to the adulteration warning data, input the adulteration risk data into the predicted food adulteration Bayesian network to obtain the adulteration category and adulteration probability, and update the food adulteration database with the predicted food adulteration Bayesian network and the corresponding adulteration characterization vector;

[0056] Management module: used to store, manage, and view the food adulteration database, the adulteration warning data, the adulteration category, and the adulteration probability, and adjust the food production and sales process operations according to the adulteration warning results and adulteration prediction results.

[0057] The beneficial effects of the present invention are:

[0058] The present invention is a food adulteration prediction method and system based on Bayesian network. Compared with the prior art, the present invention has the following technical effects:

[0059] The present invention can improve the data preprocessing capability in food adulteration prediction by constructing a food adulteration database, extracting food chain features, calculating comprehensive similarity, hierarchical early warning prediction and dynamically updating the Bayesian network steps, and can enhance the adaptability of multi-source food data in the model, thereby improving the efficiency and accuracy of food adulteration prediction. The food adulteration prediction technology is optimized, which can greatly save resources and improve work efficiency, and provide more reliable technical support for food adulteration prediction. In the early stage of food production and sales, potential adulteration risks are discovered in time, and early warnings are issued in advance, so as to buy precious time for regulatory authorities to take measures and for enterprises to adjust production and sales strategies, and ensure food safety. The present invention can adapt to different food adulteration prediction systems and the prediction needs of food adulteration of different users, and has certain universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 The present invention is a flowchart of the steps of a method for predicting food adulteration based on a Bayesian network. DETAILED DESCRIPTION

[0061] The present invention is further described below by means of specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.

[0062] The present invention provides a method and system for predicting food adulteration based on a Bayesian network, comprising the following steps:

[0063] like Figure 1 As shown, in this embodiment, the following steps are included:

[0064] Acquire historical food data and historical production and sales data to build a food adulteration database; the food adulteration database includes an initial food adulteration Bayesian network and a corresponding network representation vector;

[0065] Extract the food information and production and marketing logs of the food to be predicted, perform data processing to obtain food data and production and marketing data, and combine the food data features and production and marketing data features into a set of food chain features;

[0066] Determining standard food chain characteristics, calculating the comprehensive similarity between the food chain characteristics and the standard food chain characteristics, and determining adulteration warning data and adulteration risk data according to the comprehensive similarity;

[0067] Perform adulteration warning directly based on the adulteration warning data, determine the adulteration characterization vector and selection strategy according to the adulteration risk data, and obtain the predicted food adulteration Bayesian network according to the selection strategy; the selection strategy includes reusing the initial food adulteration Bayesian network, optimizing the initial food adulteration Bayesian network, and creating a new initial food adulteration Bayesian network;

[0068] Input the adulteration risk data into the predicted food adulteration Bayesian network to obtain the adulteration category and adulteration probability, and update the food adulteration database by using the predicted food adulteration Bayesian network and the corresponding adulteration characterization vector.

[0069] In this embodiment, the method for constructing the food adulteration database includes:

[0070] Obtain historical food data and historical production and sales data; the historical food data includes historical food quality inspection data, historical food image features, historical food public opinion features, historical food categories, and historical food adulteration conclusions; the historical production and sales data includes historical raw material procurement data, historical processing techniques, historical storage and transportation data, and historical sales records;

[0071] Align the data according to food categories, time, and region to form a spatio-temporal data cube, and determine the observation nodes, hidden nodes, and decision nodes according to the data categories; the observation nodes include quality inspection indicators, image feature similarity, public opinion risk value, raw material price volatility, number of times of exceeding the storage and transportation temperature standard, offset of storage and transportation trajectory points, sterilization temperature compliance rate, compliance of additive use, channel type, and abnormal return rate index; the hidden node is the adulteration motive; the decision node is the adulteration category;

[0072] Use domain experience to define strong causal relationships to determine the key edges, search for the optimal structure based on the K2 algorithm and the hill-climbing algorithm, and perform dynamic pruning according to the time-decayed mutual information and intervention experiment causal stability to determine the Bayesian network structure. The dynamic pruning conditions are:

[0073]

[0074] Where KeepEdge(X→Y) is the condition for retaining the edge (X→Y) from the observation node X to Y, I τ (X,Y) is the time-decayed mutual information, θ I is the time-decayed mutual information threshold, S causal (X,Y) is the causal stability score, θ S is the causal stability score threshold, T is the total length of the time window, λ is the decay coefficient, I(X t ,Y t ) is the mutual information value at time t within the time window, used to measure the variables X t and Y tThe relevance, N int is the number of intervention experiments, is the quantification of the causal effect, P[Y|do(X=x k )] is the probability distribution of Y under the intervention operation x of X k ;

[0075] Input the spatio-temporal data cube into the Bayesian network structure for parameter learning to generate the conditional probability table, and determine the initial food adulteration Bayesian network according to the Bayesian network structure and the conditional probability table; when the spatio-temporal data cube is complete, the maximum likelihood estimation is used to generate the conditional probability table; when the spatio-temporal data cube is missing, the Bayesian estimation combined with the Dirichlet prior is used to generate the conditional probability table, and the expression is:

[0076]

[0077] where P(X=x i |Pa(X)=pa j ) is the conditional probability that the node X takes the value x when the parent node Pa(X) takes the value pa j , i is the number of observations of X=x and Pa(X)=pa i , j is the number of observations of Pa(X)=pa , j is the dynamic prior parameter at time t, K is the number of all possible values of X, C source is the data source confidence weight, β is the weight coefficient of the data source confidence, γ is the weight coefficient of the time decay, δ is the weight coefficient of the expert knowledge injection, ExpertPrior(x i ,pa j ) is the expert preset prior probability when X=x i and Pa(X)=pa j ;

[0078] Extract the food category, production and sales time, and production and sales region from the historical food data and historical production and sales data, splice them into a network representation vector, associate the network representation vector with the corresponding initial food adulteration Bayesian network, and construct different category initial food adulteration Bayesian networks and network representation vectors according to different category historical food data and historical production and sales data to form a food adulteration database;

[0079] ​In the actual evaluation, historical food data and historical production and sales data of a canned fruit production enterprise in the past three years were obtained, including: 800 pieces of historical food quality inspection data (freshness of fruits, heavy metal content, additives, spectral data) and historical food image features (shape of fruits, texture features, proportion of main colors, raw material features) were obtained from the quality inspection department, 300 pieces of food public opinion information (including "fruits with peculiar smell", "fruits not fresh", "abnormal color", "fruits too sweet", etc.) were collected by using a social crawler engine, the food categories (such as canned yellow peaches, canned strawberries, etc.) and 80 corresponding historical food adulteration conclusions (whether adulterated and specific adulteration types, such as using inferior yellow peaches, adding illegal pigments, etc.) were clarified, 400 pieces of historical raw material purchase data (including the origin of raw materials, purchase price, purchase quantity) were obtained from the enterprise's purchase records, 200 pieces of historical processing technology data (including processing temperature, time, additive addition sequence) were obtained from the factory machine logs, 300 pieces of historical storage and transportation data (covering storage and transportation temperature, humidity, duration, route) were obtained from the storage and transportation logs, and 150 pieces of historical sales records (including sales regions, sales volume, purchase volume) were obtained from the sales logs;

[0080] Define critical edges by determining strong causal relationships based on domain experience. For example, if the raw material purchase price is too low, inferior fruits may be used, so there is a critical edge between the raw material price volatility and the adulteration motivation; set the total length T of the time window to 36 months, the decay coefficient λ to 0.4, the time-decayed mutual information threshold θ I to 0.25, the causal stability score threshold θ S to 0.35, and the number of intervention experiments N int to 40. Taking the fruit freshness X and the adulteration conclusion Y as an example, calculate the time-decayed mutual information I τ (X, Y) to be 0.3 and the causal stability score S causal (X, Y) to be 0.5, then retain the edge (X → Y);

[0081] Obtain the data source confidence weight C source to be 0.7 for the official database, 0.3 for the forum social news posts respectively, the weight coefficient β of the data source confidence to be 0.4, the weight coefficient γ of the time decay to be 0.2, and the weight coefficient δ of the expert knowledge injection to be 0.4. Input the spatio-temporal data cube into the Bayesian network structure for parameter learning to generate a conditional probability table;

[0082] Taking the food category, production and sales time, and production and sales region in a set of historical food data and historical production and sales data as an example, obtain the network representation vector [canned strawberries, 20240415, Yantai, Shandong, Zhengzhou, Henan].

[0083] In this embodiment, the method for determining the food chain characteristics includes:

[0084] Extract the food production and sales log, and use the Logstash parser to extract the key fields of the food production and sales log and output structured data to obtain production and sales data; the production and sales data includes raw material procurement data, processing technology, storage and transportation data, and sales records;

[0085] The food data includes food quality inspection data, food image features, food public opinion features, and food categories. The methods for obtaining and processing food information to obtain food data include:

[0086] Use a social crawler engine to obtain food public opinion responses, and input the food public opinion responses into the BERT model to generate food public opinion features; the food public opinion features include sentiment polarity scores and the frequency of risk keywords;

[0087] Obtain food quality inspection data by docking with the regulatory agency database and the manufacturer's backup database;

[0088] Obtain food images taken by the production line and sales terminals, extract the color histogram of the food images to output the proportion of the main color, use ResNet-50 to perform deep features on the food images to output food texture features and shape features, use the ResNet classification model to process the food images to obtain raw material features, and combine the proportion of the main color, food texture features, shape features, and raw material features to form food image features;

[0089] Combine the food data features and the production and sales data features to form a set of food chain features;

[0090] In actual evaluation, obtain the food information and production and sales logs of a certain batch of strawberry cans and yellow peach cans of a certain fruit canning enterprise;

[0091] Taking strawberry cans as an example: Use the Logstash parser to process the food production and sales log to obtain production and sales data [price - 20%, purchase volume + 50%, processing temperature - 5°C, processing time + 10 min, storage and transportation temperature 4°C, humidity exceeded the standard 3 times, sales volume - 30%, small supermarkets / farmers' markets, third / fourth-tier cities]; use a social crawler engine and the BERT model to generate food public opinion features [sentiment polarity score - 0.3, strange smell / 5 times, not fresh / 3 times]; dock with the official database to obtain food quality inspection data [soluble solids 15%, vitamin C 30 mg / 100 g, unknown additive 5 mg / 100 g]; obtain food image processing to obtain food image features (proportion of the main color, texture features - graininess and uniformity / smoothness and uniformity, shape features - roundness / deformity rate / size uniformity, raw material features - strawberry feature confidence / other fruit feature confidence / unknown fruit feature confidence) [red 70%, 0.2, 0.3, 0.1, 0.4, 0.8, 0.6, 0.7, 0.9, 0.1, 0.05].

[0092] In this embodiment, the method for determining adulteration warning data and adulteration risk data includes:

[0093] Determine the standard food data characteristics and standard production and sales data characteristics according to the manufacturer's production specifications, the merchant's sales specifications, and the market supervision requirements, and form a set of standard food chain characteristics; the data categories and dimensions of the standard food chain characteristics are consistent with those of the food chain characteristics;

[0094] Calculate the comprehensive similarity between the food chain feature vector and the standard food chain feature vector, and the expression is:

[0095]

[0096] where Sim line is the comprehensive similarity between the food chain feature and the standard food chain feature, S food is the similarity between the food data characteristics and the standard food data characteristics, S supply is the similarity between the production and sales data characteristics and the standard production and sales data characteristics, θ is the adjustment factor, η is the collaborative gain coefficient, I(·) is the indicator function, taking 1 when the condition is met and 0 otherwise, is the food data similarity threshold, is the production and sales data similarity threshold, n f is the number of food data, is the element of the food data feature vector, is the element of the standard food data feature vector, w i is the weight of the i-th element, is the balance weight between the Gaussian kernel similarity and the improved Jaccard coefficient, is the continuous feature vector of the food data, is the continuous feature vector of the standard food data, σ is the smoothing parameter in the Gaussian kernel similarity, is the set of discrete features of the food data, is the set of discrete features of the standard food data, ε is the smoothing factor, φ2 is the balance weight between the Euclidean distance and the geographical grid similarity, is the feature vector of the data-type production and sales data and the similarity between the standard data-type production and sales data feature vector , is the feature vector of the coordinate-type production and sales data and the similarity between the standard coordinate-type production and sales data feature vector , ρ is the collaborative gain coefficient, N cover is the number of grids where the two regions overlap, N total is the total number of grids in the two regions, δ is the logistics cost weight coefficient, is the production and sales logistics cost, is the standard production and sales logistics cost;

[0097] Define the data corresponding to the food chain features with a comprehensive similarity greater than 0.8 to the standard food chain features as safe food data;

[0098] Define the data corresponding to the food chain features with a comprehensive similarity less than 0.5 to the standard food chain features as adulteration warning data;

[0099] Conversely, define the data corresponding to the remaining food chain features as adulteration risk data;

[0100] In the actual evaluation, take the adjustment factor θ as 1.1, the collaborative gain coefficient η as 0.7, the food data similarity threshold as 0.65, the production and sales data similarity threshold as 0.65, the balance weight of the Gaussian kernel similarity and the improved Jaccard coefficient as 0.5, the smoothing parameter σ in the Gaussian kernel similarity as 1.2, the smoothing factor ε as 0.02, the balance weight φ2 of the Euclidean distance and the geographical grid similarity as 0.6, the collaborative gain coefficient ρ as 0.4, and the logistics cost weight coefficient δ as 0.25. Calculate that the comprehensive similarities of the food chain feature vectors of a certain batch of strawberry cans and yellow peach cans of a certain fruit can production enterprise to the standard food chain feature vector are 0.66 and 0.39 respectively, then their corresponding data are divided into adulteration risk data and adulteration warning data.

[0101] In this embodiment, the method for directly giving adulteration warning includes:

[0102] Determine the warning indicators according to the food chain features. With the mean of the historical normal data of each warning indicator as the center of the sphere, set a sliding window to dynamically adjust the radius, construct a multi-dimensional sphere for multi-sphere screening to obtain adulterated data, calculate the ratio of the adulterated data to the corresponding sphere radius to obtain the degree of adulteration, input the degree of adulteration and the corresponding warning indicators into the food domain word bag model to obtain the adulteration category and adulteration weight, generate an adulteration feature vector according to the adulteration category and adulteration weight, calculate the cosine similarity between the adulteration feature vector and the historical adulteration feature vector in the adulteration conclusion library to obtain the food adulteration conclusion, and give food adulteration warning according to the food adulteration conclusion;

[0103] In the actual evaluation, for the adulteration warning data of a certain batch of yellow peach cans of a certain fruit can production enterprise, the specific steps are as follows:

[0104] According to the characteristics of the food chain, the early warning indicators are determined as the freshness of yellow peaches, the compliance of additive use, the number of times of excessive storage and transportation humidity, the shape characteristics, etc. The historical normal data mean of this brand of canned yellow peaches is collected, the sliding window is set as the past 1 month, a multi-dimensional sphere is constructed for multi-sphere screening, and the ratio of the adulterated data to the corresponding sphere radius is calculated (the adulteration degree of yellow peach freshness is 1.3, the adulteration degree of yellow peach raw materials is 1.2, and the adulteration degree of yellow peach flavor is 1.3). The adulteration degree and the corresponding early warning indicators are input into the word bag model in the food field to obtain the adulteration category and adulteration weight [inferior fruits, 0.6, other types of fruits, 0.7, saccharin, 0.7]. The maximum cosine similarity between this adulteration feature vector and the historical adulteration feature vectors in the adulteration conclusion library is 0.81 (exceeding the set conclusion threshold of 0.75), then it is determined that there is a risk of adulteration with inferior yellow peaches and additives in this batch of canned yellow peaches, and food adulteration early warning is carried out.

[0105] In this embodiment, the method for obtaining the predicted food adulteration Bayesian network includes:

[0106] Extract the food category, production and sales time, and production and sales region in the adulteration risk data, splice them into an adulteration characterization vector, calculate the cosine similarity between the adulteration characterization vector and the network characterization vector in the food adulteration database, and determine the selection strategy according to the maximum cosine similarity;

[0107] When the maximum cosine similarity is greater than 0.8, directly select the initial food adulteration Bayesian network associated with the network characterization vector corresponding to the maximum cosine similarity as the predicted food adulteration Bayesian network;

[0108] When the maximum cosine similarity is less than 0.5, directly create a predicted food adulteration Bayesian network;

[0109] Conversely, update the initial food adulteration Bayesian network associated with the network characterization vector corresponding to the maximum cosine similarity to obtain the predicted food adulteration Bayesian network. The specific steps are as follows:

[0110] Set a dynamic sliding window for food data and production and sales data for incremental learning, and update the conditional probability table of the initial food adulteration Bayesian network. The expression is:

[0111]

[0112] where P new (X|Pa(X)) is the updated conditional probability table of node X when the parent node takes Pa(X), is the weight of the new and old data fusion ratio, Conf date and Var date are the data source confidence and data variance within the dynamic sliding window respectively, N new is the number of observations of the new data, N oldis the number of observations of old data, P new,date is the conditional probability estimate based on new data, P old is the old conditional probability table;

[0113] Calculate the causal KL divergence and the mutual information of edges of the initial food adulteration Bayesian network after incremental update. The expression is:

[0114]

[0115] where KL causal (P∥Q) is the causal KL divergence after incremental update, which is used to measure the difference between two probability distributions P and Q. KL(P∥Q) is the traditional KL divergence, and ζ is the causal effect weight coefficient. is the quantification of the causal effect, is at the intervention X = x i is the expected value of Y, M int is the number of interventions;

[0116] According to the causal KL divergence KL causal (P∥Q) and the mutual information I(X,Y) of the edge X→Y, adjust the network nodes and edges, and use Bayesian optimization to select the optimal subnet structure to obtain the predicted food adulteration Bayesian network;

[0117] In the actual evaluation, according to the food category in the adulteration risk data of a batch of strawberry canned food produced by a certain fruit canning enterprise (the production and sales time and the production and sales region are spliced into an adulteration characterization vector [strawberry canned food, 20230329, Yantai, Shandong, Xi'an, Shaanxi]), calculate that the maximum cosine similarity between the adulteration characterization vector and the network characterization vector in the food adulteration database is 0.7, and select the initial food adulteration Bayesian network associated with the network characterization vector corresponding to the maximum cosine similarity for update;

[0118] Set a dynamic sliding window for food data and production and sales data for incremental learning. The number of new data observations N new is 15, and the number of old data observations N old is 65. The confidence level Conf date of the data source within the dynamic sliding window is 0.75, and the data variance Var date is 0.06. Update the conditional probability table of the initial food adulteration Bayesian network according to the formula;

[0119] Take the causal effect weight coefficient ζ as 0.5, calculate the causal KL divergence and the mutual information of edges of the initial food adulteration Bayesian network after incremental update. When the causal KL divergence KL causalWhen (P∥Q) is greater than 0.2, new network nodes are added. When the mutual information I(X,Y) of the edge X→Y is less than 0.05, the edge is deleted. Bayesian optimization is used to select the optimal subnet structure, and finally a Bayesian network for predicting food adulteration is obtained;

[0120] The adulteration risk data corresponding to canned strawberries is input into the Bayesian network for predicting food adulteration to obtain the adulteration category and adulteration probability: inclusion of inferior fruits - 50%, use of damaged fruits with damaged parts cut off - 70%, treatment of immature fruits with pigments and saccharin - 79%;

[0121] The Bayesian network for predicting food adulteration and the corresponding adulteration characterization vector are used to update the food adulteration database.

[0122] In a second aspect, a food adulteration prediction system based on a Bayesian network includes:

[0123] Database module: used to obtain historical food data and historical production and sales data to construct a food adulteration database; the food adulteration database includes an initial food adulteration Bayesian network and a corresponding network characterization vector;

[0124] Data module: used to process the food information and production and sales logs of the food to be predicted to obtain food chain characteristics, used to calculate the comprehensive similarity between the food chain characteristics and the standard food chain characteristics, and determine adulteration warning data and adulteration risk data according to the comprehensive similarity;

[0125] Network selection module: used to determine an adulteration characterization vector and a selection strategy according to the adulteration risk data, and obtain a Bayesian network for predicting food adulteration according to the selection strategy; the selection strategy includes reusing the initial food adulteration Bayesian network, optimizing the initial food adulteration Bayesian network, and creating a new initial food adulteration Bayesian network;

[0126] Warning and prediction module: used to directly perform adulteration warnings according to the adulteration warning data, used to input the adulteration risk data into the Bayesian network for predicting food adulteration to obtain the adulteration category and adulteration probability, and use the Bayesian network for predicting food adulteration and the corresponding adulteration characterization vector to update the food adulteration database;

[0127] Management module: used to store, manage, and view the food adulteration database, the adulteration warning data, the adulteration category, and the adulteration probability, and adjust the food production and sales process operations according to the adulteration warning results and adulteration prediction results

[0128] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A food adulteration prediction method based on a Bayesian network, characterized in that, It includes the following steps: S1. Obtain historical food data and historical production and sales data to construct a food adulteration database; the food adulteration database includes an initial food adulteration Bayesian network and a corresponding network characterization vector; S2. Extract the food information and production and sales logs of the food to be predicted, perform data processing to obtain food data and production and sales data, and form a set of food chain features by combining the food data features and the production and sales data features; S3. Determine the standard food chain features, calculate the comprehensive similarity between the food chain features and the standard food chain features, and determine the adulteration warning data and adulteration risk data according to the comprehensive similarity; S4. Directly conduct adulteration warnings according to the adulteration warning data, determine the adulteration characterization vector and selection strategy according to the adulteration risk data, and obtain the predicted food adulteration Bayesian network according to the selection strategy; the selection strategy includes reusing the initial food adulteration Bayesian network, optimizing the initial food adulteration Bayesian network, and creating a new initial food adulteration Bayesian network; S5. Input the adulteration risk data into the predicted food adulteration Bayesian network to obtain the adulteration category and adulteration probability, and update the food adulteration database by using the predicted food adulteration Bayesian network and the corresponding adulteration characterization vector.

2. The food adulteration prediction method based on a Bayesian network according to claim 1, wherein The method for constructing the food adulteration database includes: Obtain historical food data and historical production and sales data; the historical food data includes historical food quality inspection data, historical food image features, historical food public opinion features, historical food categories, and historical food adulteration conclusions; the historical production and sales data includes historical raw material procurement data, historical processing techniques, historical storage and transportation data, and historical sales records; Align the data according to food categories, time, and region to form a spatio-temporal data cube, and determine the observation nodes, hidden nodes, and decision nodes according to the data categories; the observation nodes include quality inspection indicators, image feature similarity, public opinion risk values, raw material price volatility, number of times of exceeding the storage and transportation temperature standard, deviation of storage and transportation trajectory points, sterilization temperature compliance rate, compliance of additive use, channel type, and abnormal return rate index; the hidden node is the adulteration motivation; the decision node is the adulteration category; Define strong causal relationships using domain experience to determine the key edges, search for the optimal structure based on the K2 algorithm and the hill-climbing algorithm, and perform dynamic pruning according to the time-decay mutual information and the causal stability of intervention experiments to determine the Bayesian network structure. The dynamic pruning conditions are: Where KeepEdge(X→Y) is the condition for retaining the edge (X→Y) from the observed node X to Y, I τ (X, Y) is the time-decaying mutual information, θ I is the time-decaying mutual information threshold, S causal (X, Y) is the causal stability score, θ S is the causal stability score threshold, T is the total length of the time window, λ is the decay coefficient, I(X t , Y t ) is the mutual information value at time t within the time window, used to measure the correlation between variables X t and Y t at time t, N int is the number of intervention experiments, is the quantification of the causal effect, P[Y|do(X = x k )] is the probability distribution of Y under the intervention operation x k of X; Input the spatio-temporal data cube into the Bayesian network structure for parameter learning to generate a conditional probability table, and determine the initial food adulteration Bayesian network according to the Bayesian network structure and the conditional probability table; when the spatio-temporal data cube is complete, the maximum likelihood estimation is used to generate the conditional probability table; when the spatio-temporal data cube is missing, the Bayesian estimation combined with the Dirichlet prior is used to generate the conditional probability table, and the expression is: Where P(X=x i |Pa(X)=pa j ) is the value of pa at the parent node Pa(X) j When node X takes value x i The conditional probability of X = x i And Pa(X)=pa j The number of observations, Pa(X)=pa j The number of observations, is the dynamic prior parameter at time t, K is the number of all possible values of X, C source is the data source confidence weight, β is the weight coefficient of data source confidence, γ is the weight coefficient of time decay, δ is the weight coefficient of expert knowledge injection, ExpertPrior(x i ,pa j ) is X=x i And Pa(X)=pa j The expert preset prior probability at the time; Extract the food category, production and sales time, and production and sales region from the historical food data and historical production and sales data, splice them into a network representation vector, associate the network representation vector with the corresponding initial food adulteration Bayesian network, and construct different categories of initial food adulteration Bayesian networks and network representation vectors according to different categories of historical food data and historical production and sales data to form a food adulteration database.

3. The food adulteration prediction method based on Bayesian network according to claim 1, wherein, The method for determining the food chain characteristics includes: Extract the food production and sales log, and use the Logstash parser to extract the key fields of the food production and sales log to output structured data to obtain production and sales data; the production and sales data includes raw material procurement data, processing technology, storage and transportation data, and sales records. The food data includes food quality inspection data, food image characteristics, food public opinion characteristics, and food category. The methods for obtaining and processing food information to obtain food data include: Use a social crawler engine to obtain food public opinion responses, and input the food public opinion responses into the BERT model to generate food public opinion characteristics; the food public opinion characteristics include sentiment polarity scores and risk keyword frequencies. Obtain food quality inspection data by docking with the regulatory agency database and the manufacturer backup database. Obtain food images taken by the production line and sales terminals, extract the color histogram of the food images to output the proportion of the main color, use ResNet-50 to perform deep features on the food images to output food texture characteristics and shape characteristics, use the ResNet classification model to process the food images to obtain raw material characteristics, and combine the proportion of the main color, food texture characteristics, shape characteristics, and raw material characteristics to form food image characteristics. Combine the food data characteristics and production and sales data characteristics into a set of food chain characteristics.

4. The food adulteration prediction method based on Bayesian network according to claim 1, wherein, The method for determining adulteration warning data and adulteration risk data includes: Determine the standard food data characteristics and standard production and sales data characteristics according to the manufacturer's production specifications, merchant sales specifications, and market supervision requirements, and form a set of standard food chain characteristics; the data categories and dimensions of the standard food chain characteristics are the same as those of the food chain characteristics. Calculate the comprehensive similarity between the characteristic vector of the food chain and the standard characteristic vector of the food chain, and the expression is as follows: Among them, Sim line is the comprehensive similarity between the food chain characteristics and the standard food chain characteristics, S food is the similarity between the food data characteristics and the standard food data characteristics, S supply is the similarity between the production and sales data characteristics and the standard production and sales data characteristics, is the adjustment factor, η is the collaborative gain coefficient, I(·) is the indicator function, which takes 1 when the condition is met and 0 otherwise, is the food data similarity threshold, is the production and sales data similarity threshold, n f is the quantity of food data, is an element of the food data feature vector, is an element of the standard food data feature vector, w i is the weight of the i-th element, is the balance weight between the Gaussian kernel similarity and the improved Jaccard coefficient, is the continuous feature vector of food data, is the continuous feature vector of standard food data, σ is the smoothing parameter in the Gaussian kernel similarity, is the set of discrete features of food data, is the set of discrete features of standard food data, ε is the smoothing factor, φ2 is the balance weight between the Euclidean distance and the geographical grid similarity, is the feature vector of data-type production and sales data and the similarity with the standard data-type production and sales data feature vector is the feature vector of coordinate-type production and sales data and the similarity with the standard coordinate-type production and sales data feature vector , ρ is the collaborative gain coefficient, N cover is the number of grids where two regions overlap, N total is the total number of grids in two regions, δ is the logistics cost weight coefficient, is the production and sales logistics cost, is the standard production and sales logistics cost;​ Define the data corresponding to the food chain characteristics with a comprehensive similarity greater than 0.8 to the standard food chain characteristics as safe food data. Define the data corresponding to the food chain characteristics with a comprehensive similarity less than 0.5 to the standard food chain characteristics as adulteration warning data. Conversely, define the data corresponding to the remaining food chain characteristics as adulteration risk data.

5. The method for predicting food adulteration based on a Bayesian network according to claim 1, wherein The method for directly conducting adulteration warnings includes: Determine warning indicators according to the food chain characteristics, use the mean of the historical normal data of each warning indicator as the center of the sphere, set a sliding window to dynamically adjust the radius, construct a multi-dimensional sphere for multi-sphere screening to obtain adulterated data, calculate the ratio of the adulterated data to the corresponding sphere radius to obtain the degree of adulteration, input the degree of adulteration and the corresponding warning indicators into the food domain word bag model to obtain the adulteration category and adulteration weight, generate an adulteration feature vector according to the adulteration category and adulteration weight, calculate the cosine similarity between the adulteration feature vector and the historical adulteration feature vector in the adulteration conclusion library to obtain the food adulteration conclusion, and conduct food adulteration warnings according to the food adulteration conclusion.

6. The food adulteration prediction method based on Bayesian network according to claim 1, wherein The method for obtaining the predicted food adulteration Bayesian network includes: Extract the food category, production and sales time, and production and sales region from the adulteration risk data, splice them into an adulteration characterization vector, calculate the cosine similarity between the adulteration characterization vector and the network characterization vector in the food adulteration database, and determine the selection strategy according to the maximum cosine similarity; When the maximum cosine similarity is greater than 0.8, directly select the initial food adulteration Bayesian network associated with the network characterization vector corresponding to the maximum cosine similarity as the predicted food adulteration Bayesian network; When the maximum cosine similarity is less than 0.5, directly create a predicted food adulteration Bayesian network; Otherwise, update the initial food adulteration Bayesian network associated with the network characterization vector corresponding to the maximum cosine similarity to obtain the predicted food adulteration Bayesian network. The specific steps are as follows: Set dynamic sliding windows for food data and production and sales data for incremental learning, and update the conditional probability table of the initial food adulteration Bayesian network. The expression is: Where P new (X|Pa(X)) is the updated conditional probability table of node X when the parent node takes Pa(X), is the weight of the new and old data fusion ratio, Conf date and Var date are the data source confidence and data variance within the dynamic sliding window respectively, N new is the number of observations of the new data, N old is the number of observations of the old data, P new,date is the conditional probability estimate based on the new data, P old is the old conditional probability table; Calculate the causal KL divergence and edge mutual information of the initial food adulteration Bayesian network after incremental update. The expression is: Among them, KL causal (P∥Q) is the causal KL divergence after incremental update, which is used to measure the difference between two probability distributions P and Q. KL(P∥Q) is the traditional KL divergence, and ζ is the causal effect weight coefficient. is the quantification of the causal effect. is the expected value of Y under the intervention X = x i and M int is the number of interventions. According to the causal KL divergence KL causal (P∥Q) and the mutual information I(X, Y) of the edge X→Y to adjust the network nodes and edges, and use Bayesian optimization to select the optimal subnet structure to obtain the Bayesian network for predicting food adulteration.

7. A food adulteration prediction system based on a Bayesian network for performing the method according to any one of claims 1-6, characterized in that, Include: Database module: used to obtain historical food data and historical production and sales data to construct a food adulteration database; the food adulteration database includes an initial food adulteration Bayesian network and the corresponding network characterization vector; Data module: used to process the food information and production and sales logs of the food to be predicted to obtain food chain characteristics, calculate the comprehensive similarity between the food chain characteristics and the standard food chain characteristics, and determine the adulteration warning data and adulteration risk data according to the comprehensive similarity; Network selection module: used to determine the adulteration characterization vector and selection strategy according to the adulteration risk data, and obtain the predicted food adulteration Bayesian network according to the selection strategy; the selection strategy includes reusing the initial food adulteration Bayesian network, optimizing the initial food adulteration Bayesian network, and creating a new initial food adulteration Bayesian network; Warning and prediction module: used to directly conduct adulteration warnings according to the adulteration warning data, input the adulteration risk data into the predicted food adulteration Bayesian network to obtain the adulteration category and adulteration probability, and update the food adulteration database with the predicted food adulteration Bayesian network and the corresponding adulteration characterization vector; Management module: used to store, manage, and view the food adulteration database, the adulteration warning data, the adulteration category, and the adulteration probability, and adjust the food production and sales process operations according to the adulteration warning results and adulteration prediction results.

Citation Information

Patent Citations

  • Fruit and vegetable food safety risk prediction method

    CN107918837A

  • Food safety risk early warning model based on AHP-EW and AE-RNN fusion and establishment method thereof

    CN115392618A

  • Food enterprise risk grading intelligent management and control system

    CN115526546A

  • Risk early warning method and system based on food operation risk information base

    CN118863552A

  • Risk early warning method based on food management risk information base

    CN119647975A