Multi-level food safety risk evaluation method based on clustering and entropy weight method
A multi-level food safety risk assessment model was constructed by Mini-Batch K-Means clustering and entropy weight method, which solved the problems of subjectivity and lack of adaptability in food safety risk assessment in existing technologies, realized scientific and systematic evaluation and dynamic monitoring of food safety risks, and improved the accuracy and efficiency of supervision.
Patent Information
- Application Number
- CN202510724010.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-02
- Publication Date
- 2025-09-12
AI Technical Summary
Existing food safety risk assessment methods are highly subjective, lack systematicity and adaptability, are difficult to adapt to large-scale multi-dimensional data processing, and cannot distinguish between multi-level risks such as provinces, cities, and districts, leading to problems such as missed judgments and misjudgments.
The Mini-Batch K-Means clustering algorithm is used for unsupervised risk level classification, and the entropy weight method is used to objectively assign weights of risk factors. A multi-level food safety risk assessment model is constructed. Combined with data mining and statistical modeling techniques, risk assessment from single-point samples to regional overall is achieved.
It improves the scientificity and systematic nature of risk identification, enhances the spatial adaptability of the model, supports dynamic monitoring and accurate early warning, and improves the effectiveness of regulatory decision-making.
Smart Images

Figure CN120634240A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of food safety, and specifically is a multi-level food safety risk assessment method based on clustering and entropy weight method. Background Art
[0002] Food safety is a major livelihood issue affecting all of society. In recent years, with the continuous improvement of the regulatory system, my country's food safety situation has generally improved. However, recent food safety incidents have demonstrated that the risk of localized outbreaks of food safety issues persists. This not only directly threatens public health but also seriously undermines consumer trust in the food industry, negatively impacting social stability and economic development.
[0003] In food regulation, food safety risk assessment is a crucial tool for identifying potential safety hazards, guiding regulatory enforcement, and optimizing resource allocation. By analyzing and assessing sampling data and conducting risk assessments, regulatory authorities can promptly identify high-risk areas and food categories, enabling targeted sampling, risk warnings, and market regulation. However, existing food safety risk assessment methods still have limitations. First, traditional risk assessment methods often rely on expert scoring and empirical judgment, which are highly subjective and difficult to adapt to the scientific processing of large-scale, multi-dimensional data. Second, such methods often lack systematic evaluation models, making it difficult to distinguish risk profiles at multiple levels, such as provinces, cities, and districts (counties). Finally, given the diverse food categories, diverse sampling items, and inconsistent regulatory levels across regions, a single indicator or empirical model cannot fully reflect the true risk level, leading to problems such as missed and misjudgment.
[0004] Therefore, there is an urgent need to establish an objective, scientific and operational food safety risk assessment method to improve the accuracy of risk identification and the effectiveness of regulatory decisions. Summary of the Invention
[0005] To address the limitations of existing technologies, this paper proposes a multi-level food safety risk assessment method based on clustering and entropy weighting. This method, based on real food inspection data, leverages data mining and statistical modeling techniques. Using the Mini-Batch K-Means clustering algorithm, it performs unsupervised risk classification of samples, effectively overcoming the shortcomings of traditional manual scoring methods in terms of accuracy and consistency. Furthermore, an entropy weighting method is introduced to objectively assign weights to various food categories and their risk factors, measuring the importance of each risk indicator from an information theory perspective, and ultimately constructing a comprehensive risk scoring model.
[0006] At the same time, this method's technical implementation covers multiple levels, including provincial, prefectural, and district (county) levels. It can assess food safety risks from single-point samples to regional scales, significantly enhancing the model's spatial adaptability and resolution. Compared with traditional methods, this approach not only improves the scientific and systematic nature of risk identification but also effectively supports multiple regulatory requirements, including dynamic monitoring, precise early warning, and targeted supervision.
[0007] The present invention provides a multi-level food safety risk assessment method based on clustering and entropy weight method, which includes the following steps:
[0008] Step 1: Obtain food data, including food type, sampling date, sampling location, inspection items and inspection results, and clean the food data.
[0009] Step 2: Preliminary risk level determination: Use Mini-Batch K-Means to cluster the samples and assign a preliminary risk level label to each record.
[0010] Step 3: Count the number of risk levels for different food types at each level.
[0011] Step 4: Use the entropy weight method to calculate the comprehensive risk score of each food type at this level.
[0012] Step 5: Use the entropy weight method to further obtain the overall food safety risk level of each province, city, district (county).
[0013] Step 6: Use K-Means clustering to obtain the final food risk level.
[0014] The step 1 comprises:
[0015] Step 101: Data Collection
[0016] Collect food sampling data, including food type, sampling time, sampling location (including three levels: provincial, prefecture-level city and district (county)) and inspection items and results (such as lead, cadmium, mercury and other content data).
[0017] Step 102: Data Cleaning
[0018] The data must be processed before use, including deleting the sampling items with missing inspection results, and eliminating fields that are not related to risk assessment (such as "sample specifications", "implementation standards", etc.), and retaining core fields such as "food categories", "sample names", "inspection conclusions", and "geographical locations".
[0019] The step 2 includes:
[0020] Step 201: normalize the collected data.
[0021] Step 202: Use the Mini-Batch K-Means method to perform a preliminary risk classification for each piece of data.
[0022] The step 3 includes:
[0023] Step 301: Count the number of risk levels of each food in each province.
[0024] Step 302: Count the number of food items at each risk level in each prefecture-level city.
[0025] Step 303: Count the number of risk levels of each food in each county (city, district).
[0026] The step 4 includes:
[0027] Step 401: Through the statistics in step 3, different locations under each level can form a frequency matrix Y = [y ij ] and normalize it.
[0028] Step 402: Calculate the comprehensive score of each food category at each level using the entropy weight method.
[0029] The step 5 comprises:
[0030] Step 501: Calculate the risk comprehensive scores of different foods at different locations at each level to form a risk comprehensive score matrix S = [s ij ].
[0031] Step 502: Calculate the total risk score of each region using the entropy weight method again.
[0032] The step 6 includes:
[0033] Step 601: Normalize the comprehensive risk score vector obtained in step 5.
[0034] Step 602: Use K-Means clustering to divide the final food risk level and obtain the food safety risk level of each province, prefecture-level city, district (county). BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flow chart of a multi-level food safety risk assessment method based on clustering and entropy weight method. DETAILED DESCRIPTION
[0036] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0037] like Figure 1As shown in the figure, a food sampling decision-making method based on graph attention network and Critic-Topsis of the present invention includes the following 6 steps.
[0038] Step 1: Obtain food data, including food type, sampling date, sampling location, inspection items and inspection results, and clean the food data.
[0039] Step 101: Data Collection. The sampling data used in this article is from China's grain and processed fruit and vegetable product inspections from 2017 to 2019, totaling over 1.2 million data points. Grain data includes rice, processed grains, rice flour, other grain flour products, other milled grain products, general-purpose wheat flour, special wheat flour, corn flour, corn flakes, corn grits, etc.; fruit and vegetable data includes various common fruits such as watermelon, apples, pears, and peaches, as well as common vegetables such as cabbage, eggplant, and bean sprouts. The sampling data includes information such as the sample's time, region, inspection items, and inspection results.
[0040] Step 102: Data Cleaning. 1) Since the ratio of detected data to total data is greater than 60%, all "not detected" and " / " are replaced with LOD (Limit of Detection); 2) Special symbols such as < and ≤ are removed; 3) Text is removed; 4) Data such as "not detected" and "qualitatively detected" are considered invalid and are deleted; 5) Due to different test items, the same sample may have multiple test data points. Data cleaning is performed to merge multiple test data points for the same sample into a single data point.
[0041] The data after partial cleaning is shown in Table 1.
[0042] Table 1 Example of food sampling data structure
[0043]
[0044] Step 2: Preliminary risk level determination: Use Mini-Batch K-Means to cluster the samples and assign a preliminary risk level label to each record.
[0045] Step 201: Data normalization processing.
[0046] The data is normalized using formula (1):
[0047]
[0048] Among them, x ij is the value of the mth test item of the i-th sample. max(x j ) and min(x j ) are the maximum and minimum values of the j-th test item respectively.
[0049] Step 202: Use the Mini-Batch K-Means method to divide the risk levels into k levels.
[0050] Randomly select k sample points from the data set as the initial cluster centers μ1,μ2,…,μ k , set the size b of the mini-batch and the maximum number of iterations, and finally perform iterative calculations. In the iteration, we randomly select a mini-batch (size b) of data, calculate the Euclidean distance between each sample point and each cluster center, and find the nearest cluster center μ c , and record the cumulative update times n of the c-th cluster center c The calculation formula of Euclidean distance is shown in (2), and the update formula of cluster center is shown in (3).
[0051]
[0052] Where d is the feature dimension, μ ji is the value of the cluster center in the i-th dimension.
[0053]
[0054] Where η is the learning rate,
[0055] Table 2 shows some of the results after using Mini-Batch K-Means to perform preliminary risk level classification on each sample.
[0056] Table 2 Preliminary risk level classification (partial results)
[0057]
[0058] Step 3: Count the number of risk levels for different food types at each level.
[0059] Step 301: Count the number of risk levels of each food in each province. Assuming that the risk level is divided into k levels, the number of risk levels of the βth food type in the αth province is n. αβ1 ,n αβ2 ,…,n αβk .
[0060] Step 302: Similarly, count the number of risk levels of each food in each prefecture-level city. Assuming that the risk level is divided into k levels, the number of risk levels of the βth food type in the αth prefecture-level city in a province is n. αβ1 ,n αβ2 ,…,n αβk .
[0061] Step 303: Similarly, count the number of risk levels of each type of food in each county (city, district). Assuming that the risk level is divided into k levels, the number of risk levels of the βth type of food in the αth prefecture-level city of a certain prefecture-level city is n. αβ1 ,n αβ2 ,…,n αβk .
[0062] Table 3 shows some of the statistical results based on Shanghai as an example.
[0063] Table 3 Statistical results of some risk levels (taking Shanghai as an example)
[0064]
[0065] Step 4: Use the entropy weight method to calculate the comprehensive risk score of each food type at this level.
[0066] Step 401: Data normalization processing.
[0067] Through the statistics of step 3, different locations at each level can form a frequency matrix Y=[y ij ],y ij represents the frequency of the i-th food category at the j-th risk level. It is also normalized using formula (1).
[0068] Step 402: Calculate the comprehensive score of each food category at each level using the entropy weight method.
[0069] Assuming there are m types of food, first use formula (4) to calculate the proportion matrix.
[0070]
[0071] Then, the information entropy is calculated using formula (5):
[0072]
[0073] Then use formula (6) to calculate the information entropy:
[0074]
[0075] Finally, the comprehensive score is calculated using formula (7):
[0076]
[0077] Among them S i That is, the comprehensive risk score for each food type at this level.
[0078] As shown in Table 4, the comprehensive scores of various foods in Shanghai are displayed.
[0079] Table 4 Comprehensive scores of various foods in Shanghai (partial results)
[0080]
[0081] Step 5: Use the entropy weight method to further obtain the overall food safety risk level of each province, city, district (county).
[0082] Step 501: Calculate the risk comprehensive scores of different foods at different locations at each level to form a risk comprehensive score matrix S = [s ij ],s ij represents the comprehensive risk score of the jth food type at the i-th location.
[0083] Step 502: Use the entropy weight method again to calculate the total risk score for each region. The calculation formula is the same as in step 402. Finally, the total risk score is obtained:
[0084]
[0085] Step 6: Use K-Means clustering to obtain the final food risk level.
[0086] Step 601: Data normalization processing.
[0087] After step 5, the comprehensive risk score vector of each region is obtained, and these scores form a two-dimensional matrix R i =[r i1 ,r i2 ,…,r im ]. Similarly, use formula (1) to normalize them.
[0088] Step 602: Use K-Means clustering to classify the final risk level.
[0089] First, initialize the cluster center and randomly select k sample points μ1, μ2,…, μ k As the initial center. Then calculate the Euclidean distance between each sample point and all centers using the same formula (2), and assign it to the cluster to which the nearest center belongs.
[0090] Then continuously iterate and update the cluster center. The update formula is as follows:
[0091]
[0092] where |C j | represents the number of samples in the jth cluster.
[0093] Finally, the center of each cluster is compared with the scores of other clusters, and the regions are assigned to corresponding risk levels according to the size of the scores, so as to finally obtain the food safety risk levels of each province, prefecture-level city, district (county).
[0094] After completing steps 5 and 6, we can obtain a comprehensive regional score and the final food safety risk level. The comprehensive scores and food safety risk levels for each province are shown in Table 5. The comprehensive scores and risk levels for prefecture-level cities, taking Jiangsu as an example, are shown in Table 6.
[0095] Table 5 Comprehensive scores and food safety risk levels of provinces
[0096]
[0097] Table 6 Comprehensive regional scores and food safety risk levels in Jiangsu Province
[0098]
[0099]
Claims
1. A multi-level food safety risk assessment method based on clustering and entropy weight method includes the following steps: Step 1: Obtain food data, including food type, sampling date, sampling location, inspection items and inspection results, and clean the food data. The step 1 comprises: Step 101: Data Collection Collect food sampling data, including food type, sampling time, sampling location (including three levels: provincial, prefecture-level city and district (county)) and inspection items and results (such as lead, cadmium, mercury and other content data). Step 102: Data Cleaning The data must be processed before use, including deleting the inspection items with missing inspection results, and eliminating fields that are not related to risk assessment (such as "sample specifications", "implementation standards", etc.), and retaining core fields such as "food category", "sample name", "inspection conclusion", and "geographic location". Step 2: Preliminary risk level determination: Use Mini-Batch K-Means to cluster the samples and assign a preliminary risk level label to each record. Step 3: Count the number of risk levels for different food types at each level. Step 4: Use the Entropy Weight Method to calculate the comprehensive risk score of each food type at this level. Step 5: Use the entropy weight method to further obtain the overall food safety risk level of each province, city, district (county). Step 6: Use K-Means clustering to obtain the final food risk level.
2. The multi-level food safety risk assessment method based on clustering and entropy weight method according to claim 1, wherein step 2 comprises: Step 201: Data normalization processing. The data is normalized using formula (1): Among them, x ij is the value of the mth test item of the i-th sample. max(x j ) and min(x j ) are the maximum and minimum values of the j-th test item respectively. Step 202: Use the Mini-Batch K-Means method to divide the risk levels into k levels. Randomly select k sample points from the data set as the initial cluster centers μ1,μ2,…,μ k , set the size b of the mini-batch and the maximum number of iterations, and finally perform iterative calculations. In the iteration, we randomly select a mini-batch (size b) of data, calculate the Euclidean distance between each sample point and each cluster center, and find the nearest cluster center μ c , and record the cumulative update times n of the c-th cluster center c The calculation formula of Euclidean distance is shown in (2), and the update formula of cluster center is shown in (3). Where d is the feature dimension, μ ji is the value of the cluster center in the i-th dimension. Where η is the learning rate, 3. The multi-level food safety risk assessment method based on clustering and entropy weight method according to claim 1, wherein step 3 comprises: Step 301: Count the number of risk levels of each food in each province. Assuming that the risk level is divided into k levels, the number of risk levels of the βth food type in the αth province is n. αβ1 ,n αβ2 ,…,n αβk . Step 302: Similarly, count the number of risk levels of each food in each prefecture-level city. Assuming that the risk level is divided into k levels, the number of risk levels of the βth food type in the αth prefecture-level city in a province is n. αβ1 ,n αβ2 ,…,n αβk . Step 303: Similarly, count the number of risk levels of each type of food in each county (city, district). Assuming that the risk level is divided into k levels, the number of risk levels of the βth type of food in the αth prefecture-level city of a certain prefecture-level city is n. αβ1 ,n αβ2 ,…,n αβk .
4. The multi-level food safety risk assessment method based on clustering and entropy weight method according to claim 1, wherein step 4 comprises: Step 401: Data normalization processing. Through the statistics of step 3, different locations at each level can form a frequency matrix Y=[y ij ],y ij represents the frequency of the i-th food category at the j-th risk level. It is also normalized using formula (1). Step 402: Calculate the comprehensive score of each food category at each level using the entropy weight method. Assuming there are m types of food, first use formula (4) to calculate the proportion matrix. Then, the information entropy is calculated using formula (5): Then use formula (6) to calculate the information entropy: Finally, the comprehensive score is calculated using formula (7): Among them S i That is, the comprehensive risk score for each food type at this level.
5. The multi-level food safety risk assessment method based on clustering and entropy weight method according to claim 1, wherein step 5 comprises: Step 501: Calculate the risk comprehensive scores of different foods at different locations at each level to form a risk comprehensive score matrix S = [s ij ],s ij represents the comprehensive risk score of the jth food type at the i-th location. Step 502: Use the entropy weight method again to calculate the total risk score for each region. The calculation formula is the same as in step 402. Finally, the total risk score is obtained:
6. The multi-level food safety risk assessment method based on clustering and entropy weight method according to claim 1, wherein step 6 comprises: Step 601: Data normalization processing. After step 5, the comprehensive risk score vector of each region is obtained, and these scores form a two-dimensional matrix R i =[r i1 ,r i2 ,…,r im ]. Similarly, use formula (1) to normalize them. Step 602: Use K-Means clustering to classify the final risk level. First, initialize the cluster center and randomly select k sample points μ1, μ2,…, μ k As the initial center. Then calculate the Euclidean distance between each sample point and all centers using the same formula (2), and assign it to the cluster to which the nearest center belongs. Then continuously iterate and update the cluster center. The update formula is as follows: where |C j | represents the number of samples in the jth cluster. Finally, the center of each cluster is compared with the scores of other clusters, and the regions are assigned to corresponding risk levels according to the size of the scores, so as to finally obtain the food safety risk levels of each province, prefecture-level city, district (county).
Citation Information
Cited By
Food-borne strain propagation risk early warning system and method based on clustering analysis
CN121659246A
Foodborne bacterial strain transmission risk early warning system and method based on cluster analysis
CN121659246B