Food safety intelligent supervision method and system based on big data

By collecting multi-source data and building a risk prediction model, and by using image recognition and reinforcement learning to optimize parameters, the problems of data dispersion and response lag in traditional food safety supervision have been solved, achieving efficient and accurate food safety supervision.

CN120975402APending Publication Date: 2025-11-18BEIJING YELLOW ELEPHANT FOOD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511335024.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional food safety supervision methods rely on limited data sources and manual processing, which makes it difficult to meet the needs of modern food safety management and faces pain points such as data dispersion, delayed response and difficulty in traceability.

Method used

By collecting multi-source data from production, distribution, and regulatory processes, image recognition algorithms are used to generate appearance anomaly scores. Combined with data cleaning and standardization, a risk prediction model is constructed, risk model parameters are dynamically optimized, and a reinforcement learning framework is adopted to improve regulatory efficiency.

Benefits of technology

It has achieved high efficiency and precision in food safety supervision, enabling the rapid identification of high-risk entities and triggering enhanced regulatory measures, thereby improving the long-term predictive accuracy and responsiveness of food safety supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975402A_ABST
    Figure CN120975402A_ABST
Patent Text Reader

Abstract

The invention discloses a food safety intelligent supervision method and system based on big data, and the method comprises the following steps: collecting production, circulation and supervision link data, and generating an appearance abnormality score through image recognition based on the data; cleaning the collected data, and outputting a data quality score through a quality evaluation formula according to the cleaned data; constructing a risk prediction model according to the cleaned data set, and calculating and outputting a comprehensive risk score and a key influence index of each food; counting the high-risk food quantity and the single-food risk value of each enterprise and region, identifying a high-risk subject through a double-threshold rule, and emphatically marking and enhancing supervision; and dynamically optimizing risk model parameters through interaction of a reinforcement learning framework, experience state perception, action selection, reward calculation and strategy network updating according to the comprehensive risk score and the actual supervision cost. The risk prediction model is trained according to the data of each link of food safety, and the efficiency of food safety supervision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of food, in particular to a food safety intelligent supervision method and system based on big data. BACKGROUND

[0002] Food safety is the core link of people's livelihood project. With the acceleration of globalization and urbanization, the food supply chain has become increasingly complex, and food safety problems have become more prominent. Traditional food safety supervision methods often rely on limited data sources and manual processing methods, which are difficult to meet the needs of modern food safety management. The traditional supervision mode faces the pain points of scattered data, delayed response, and difficult traceability. In recent years, with the development of information technology, especially the application of big data technology, new means and ideas have been provided for food safety supervision. Based on this, the present application provides a food safety intelligent supervision method and system based on big data. SUMMARY

[0003] The present application provides a food safety intelligent supervision method based on big data, characterized by comprising: S10, collecting multi-source data of production, circulation and supervision links, generating appearance abnormal score through image recognition algorithm according to food appearance in data and standard template to assist supervision sampling; S20, data cleaning and standardization processing are performed on the collected data, and data quality score is output based on effective data proportion, noise level, image definition and cleaning time consumption in the data through quality evaluation formula; S30, according to the cleaned data set, a risk prediction model is constructed by fusing historical event influence value, external environment, data update frequency and time decay factor, and the comprehensive risk score and key influence index of each food are calculated and output; S40, according to the comprehensive risk score, the number of high-risk foods and the risk value of single food of each enterprise and region are counted, and the high-risk subject is identified by setting double threshold rule, and the key is marked and the enhanced supervision measures are triggered; S50, according to the comprehensive risk score and the actual supervision cost, through the interaction process of state perception, action selection, reward calculation and strategy network update in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and supervision efficiency.

[0004] The food safety intelligent supervision method based on big data as described above, wherein multi-source data of production, circulation and supervision links are collected, appearance abnormal score is generated through image recognition algorithm according to food appearance in data and standard template to assist supervision sampling, and the specific sub-steps are as follows: Multi-source structured and unstructured data of production, circulation and supervision links are collected through government systems, Internet of Things devices and web crawlers respectively; The image recognition algorithm is used to compare the on-site food image with the standard template, and the packaging integrity, color abnormality and defect density characteristics are extracted; Based on the image analysis result, the appearance abnormality score is calculated as the high-risk food preliminary screening basis to assist the priority decision of sampling inspection.

[0005] The data cleaning and standardization processing of the collected data is performed, the data quality score is output based on the quality evaluation formula of the effective data proportion, noise level, image definition and cleaning time, and the specific sub-steps are as follows: The multi-source data is cleaned and standardized by methods such as duplicate checking, error correction, completion and unstructured information extraction; The extraction, conversion and loading tools are used to unify the data format and coding rules, and the cross-link correlation of production, circulation and supervision data is realized through unique identification; Based on the effective data proportion, noise level, image definition and cleaning time, the data quality score is calculated and output through the quality evaluation formula.

[0006] The risk prediction model is constructed based on the cleaned data set, which integrates the historical event influence value, external environment, data update frequency and time decay factor, and the comprehensive risk score and key influence index of each food are calculated and output, and the specific sub-steps are as follows: The frequency and severity of past food safety problems in the data chain determine the historical event influence value, which is used as the core risk factor; The dynamic weighted risk prediction model is constructed based on the historical event influence value, the comprehensive risk score of each food is calculated by running the model, and the key indicators affecting the score are output through feature analysis.

[0007] According to the comprehensive risk score, the number of high-risk foods and the risk value of each food of each enterprise and region are calculated, the double threshold rule is set to identify high-risk subjects, and the key markers are triggered to trigger enhanced supervision measures, and the specific sub-steps are as follows: The enterprises and regions where the foods with high risk values are located, and the enterprises and regions where the foods with high risk values are located are calculated; Set the single food risk threshold and high-risk category number threshold, and meet one of them to determine the high-risk subject, mark the identified high-risk subject, and automatically trigger the increased sampling frequency and other enhanced supervision measures.

[0008] The food safety intelligent supervision method based on big data as described above, wherein according to the comprehensive risk score and the actual supervision cost, through the interactive process of state perception, action selection, reward calculation and policy network updating in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and supervision efficiency, and the method comprises the following sub-steps: The risk model parameters are set as adjustable action space, and the agent selects the adjustment strategy according to the current prediction effect and supervision cost; The risk score improvement and the supervision cost reduction are taken as positive rewards, and an environment feedback mechanism is constructed to realize strategy evaluation; The strategy network is updated through experience replay and policy gradient algorithm, and the model parameters are continuously optimized to improve the long-term prediction performance.

[0009] The food safety intelligent supervision method based on big data as described above, wherein the risk model parameters are set as adjustable action space, and the agent selects the adjustment strategy according to the current prediction effect and supervision cost, and the method comprises the following sub-steps: The key parameters in the risk prediction model are defined as a limited action space, so that the agent can fine-tune in a small range, and the model instability caused by drastic parameter changes is avoided; An adaptive mechanism is introduced to dynamically adjust the action step according to the reward feedback, so as to accelerate the convergence while maintaining the system stability; A parameter freezing strategy is set to temporarily lock the parameters that perform stably, reduce redundant adjustment, and improve the model interpretability and operation reliability.

[0010] The application also provides a food safety intelligent supervision system based on big data, which comprises: A data acquisition module: acquiring multi-source data of production, circulation and supervision links, generating an appearance abnormality score through an image recognition algorithm according to the food appearance in the data and a standard template to assist supervision sampling; A cleaning and standardization module: performing data cleaning and standardization processing on the collected data, and outputting a data quality score based on the effective data proportion, noise level, image definition and cleaning time consumption in the data through a quality evaluation formula; A model construction module: constructing a risk prediction model based on the cleaned data set, the influence value of historical events, external environment, data update frequency and time decay factor, calculating and outputting the comprehensive risk score and key influence index of each food; A statistical marking module: according to the comprehensive risk score, the number of high-risk foods and the risk value of each food of each enterprise and region are counted, and a double threshold rule is set to identify high-risk subjects, and key marking is performed to trigger enhanced supervision measures; An optimization adjustment parameter module: according to the comprehensive risk score and the actual supervision cost, through the interaction process of state perception, action selection, reward calculation and policy network update in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and supervision efficiency.

[0011] The application realizes the beneficial effects as follows: the application collects data related to each link of food safety, trains a risk prediction model according to the data, and performs supervision according to the model output, thereby improving the efficiency of food safety supervision. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0013] Figure 1 is a flow chart of a food safety intelligent supervision method based on big data provided by the first embodiment of the present application.

[0014] Figure 2 is a schematic diagram of a food safety intelligent supervision system based on big data provided by the second embodiment of the present application. DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0016] Embodiment one As shown in the figure, the first embodiment of the present application provides a food safety intelligent supervision method based on big data, which comprises: Figure 1 S10, collecting multi-source data of production, circulation and supervision links, generating appearance abnormality score through image recognition algorithm according to food appearance in the data and standard template to assist supervision sampling. S11, collecting data of production, circulation and supervision links respectively.

[0017] Collecting production link data, including food production enterprise data, production process data and quality control data.

[0018]

[0019] ​The registration information of food production enterprises is obtained through the enterprise registration system and the supplier filing system of the government regulatory platform, including enterprise name, address, production license scope, legal person information, etc. At the same time, the information of food raw material suppliers is obtained, including raw material origin, variety, qualification certification, etc.; using Internet of Things technology, sensors are deployed in the production workshop to collect production environment parameters in real time, including temperature, humidity, air quality, etc., and operation data of production equipment, including equipment start-stop time, running time, key process parameters, etc.; quality control data is the internal quality inspection report of the enterprise, including raw material, semi-finished product and finished product detection data, covering physical and chemical indicators, microbial indicators, etc. At the same time, quality traceability information in the production process is recorded, including raw material batch, production batch, product flow, etc., to facilitate accurate traceability when problems occur.

[0020] Collect data in the circulation link, including logistics data, warehouse data, and market transaction data.

[0021] Cooperate with logistics enterprises to obtain detailed information during food transportation, including transportation route, transportation time, transportation tool, storage condition, etc. By installing intelligent sensors on transportation vehicles and warehouses, real-time monitoring of environmental parameters is realized to ensure that food is transported and stored under suitable conditions; collect warehouse inventory information, including food types, quantities, shelf life, in-out warehouse time, etc. At the same time, obtain environmental monitoring data of the warehouse, including temperature and humidity, ventilation conditions, etc., to evaluate the impact of warehouse conditions on food quality; collect sales data from various food trading markets, supermarkets, and e-commerce platforms, including sales time, sales location, sales price, sales volume, etc., as well as market promotion activities and product recall information.

[0022] Collect regulatory link data, including sampling data, administrative penalty data, and public opinion data.

[0023] Integrate sampling information from food regulatory departments at all levels, including sampling location, sampling object, sampling project, test result, unqualified reason, etc.; collect administrative penalty records of food production and operation enterprises, including illegal and irregular behavior, penalty basis, and penalty result information; use web crawler technology to collect food safety-related public opinion information from social media platforms, news websites, and food safety forums. Through natural language processing technology, analyze public opinion data to extract information such as public attention to food safety hot events, consumer complaints and feedback, etc., to timely discover potential food safety problems and public opinion risks.

[0024] S12, analyze the appearance of the food from the aspects of complete packaging, color abnormalities, and defect density through image recognition algorithms, and calculate the appearance abnormality score.

[0025] Use image recognition technology to analyze food appearance and packaging defects to assist manual sampling. The specific formula is wherein, is the food appearance anomaly score, is the packaging integrity coefficient, evaluating structural defects such as packaging deformation, label skew, is the mean luminance of image x, x being the current captured food packaging image, is the mean luminance of image y, y being the standard undamaged template image, is the variance of image x, used to measure the texture complexity of the image, is the variance of image y, being the texture reference of the standard packaging, is the covariance of image x and image y, used to detect structural deformation, is the luminance stability constant, preventing calculation overflow in low luminance areas, is the contrast stability constant, preventing calculation overflow in low contrast areas. The packaging damage degree is intact, slight deformation, obvious defect, severe damage. is the color anomaly coefficient, used to quantify the color difference between the current food image and the standard food image, identifying abnormal situations such as food discoloration, corruption, , is the image lightness, detecting overall light changes in the food, is the standard image lightness, is the image red-green chroma, is the standard image red-green chroma, is the image yellow-blue chroma, is the standard image yellow-blue chroma, is the color threshold, when linearly reacting color difference, when it is a severe discoloration, the food has a high risk of corruption. is the defect density coefficient, n being the number of defects detected in the image, , is the number of pixels of the i-th defect, is the total number of pixels of the food packaging in the image, is the defect type weight, distinguishing the risk level of different defects, is the edge enhancement base, is the decay adjustment coefficient, is the shortest pixel distance of the i-th defect to the packaging edge, is the edge-sensitive radius, defining the width range of the edge area.

[0026] S20, data cleaning and standardization processing is performed on the collected data, and based on the effective data proportion, noise level, image clarity, and cleaning time in the data, a data quality score is output through a quality evaluation formula.

[0027] The collected data of each link is cleaned to remove duplicate data, error data and incomplete data. Duplicate detection algorithms are used to identify and delete duplicate test reports; obviously abnormal data collected by sensors are corrected or deleted; records missing key information, such as product information missing production batch number, are processed by data completion or deletion; unstructured text information is extracted through natural language processing; highly blurred food pictures are rephotographed.

[0028] Different sources and formats of data are standardized by data extraction, conversion and loading tools, and unified data format and coding rules are used to facilitate data integration and analysis.

[0029] The data of production, circulation and supervision are associated by unique identification, and the product batch of production link is associated with the transportation record, sales record and supervision record of circulation link, realizing the information tracing and integration of food from production to consumption, and building a complete food safety data chain.

[0030] The quality evaluation formula is The quality of the data cleaning is evaluated, The data cleaning quality score is obtained, and if the score exceeds the set threshold, the data can be used for subsequent use. The condition probability of the cleaned effective data is The amount of cleaned effective data is The total amount of original data is The condition probability of noise data is The amount of noise data is The sharpness incentive coefficient is The image sharpness score is The maximum image sharpness is The time penalty coefficient is The cleaning time is The maximum allowed cleaning time is

[0031] S30, according to the cleaned data set, a risk prediction model is constructed by integrating historical event influence value, external environment, data update frequency and time decay factor, and the comprehensive risk score and key influence index of each food are calculated and output.

[0032] The cleaned data is used to construct a risk prediction model to realize intelligent prediction of food safety risk. The model is based on input data and real-time monitoring data to predict the safety risk of each food, and the specific formula is , , The risk prediction comprehensive score is the final output indicator, reflecting the comprehensive risk level of the food; ReLU is used as the activation function, which is responsible for converting the linear combination of inputs into a non-linear output, ensuring that the score will not appear negative, while retaining sensitivity to positive risk; Date is the cleaned food dataset. The feature vector of the jth food data chain includes production, circulation, supervision, and other multi-dimensional data, ensuring that the model can fully capture the risk characteristics of the food throughout its life cycle. The weight of the jth data chain, The historical event impact value is determined by the frequency and severity of past food safety issues in the data chain, , This includes image recognition defect impact, public opinion complaint impact, unqualified sampling impact, and supply chain risk impact. The supervision cycle, The number of defects in the time series is the number of packaging defects detected daily by image detection, The defect severity coefficient is a function of time, The historical maximum number of defects. The number of consumer complaints, The historical maximum number of complaints, The sentiment analysis score, The Beta function distribution shape parameter quantifies the uncertainty of the impact of sentiment intensity on complaints through a probability model. The stronger the negative emotions, the greater the impact of complaints. The number of historical unqualified sampling, The interval between the event occurrence time and the current time, The maximum allowed time interval, The decay coefficient combines the sampling results and time factors to reflect the long-term impact and recent urgency of historical problems. The supply chain risk impact is determined by the number of supply chain disruption events. The more disruptions, the greater the impact; The historical event maximum value is used for standardization to avoid the impact of absolute value differences on the score. External environmental factors include seasonal epidemics, policy changes, and other uncontrollable variables. The data update frequency refers to the real-time update frequency of data in each link of the food chain. The maximum allowed update frequency is used for normalization. The time interval represents the time difference between the current time and the last update time of the data chain. The maximum allowed time interval is used to limit the impact of outdated data on risk scoring. To punish the coefficient, control the dynamic influence of historical events, external factors, update frequency and time decay on risk score respectively, adapt to the needs of different regulatory scenarios.

[0033] The cycle model outputs the comprehensive risk score of each food and the key indicators that have a large proportion of influence score.

[0034] S40, according to the comprehensive risk score, the number of high-risk foods of each enterprise and region and the risk value of individual food are counted, and high-risk subjects are identified by setting double threshold rules, and key markers are triggered to trigger enhanced regulatory measures.

[0035] The enterprises and regions where the food with high risk value and the enterprises and regions where the food with high risk value are counted, and the risk status of various foods, enterprises and regions is displayed in real time through the monitoring platform. High-risk objects are marked and warned. The number of risk food categories exceeding the threshold and the risk value of individual food exceeding the threshold are high-risk. For high-risk enterprises and regions, increase the frequency and intensity of sampling inspection; for food categories with obvious continuous risk rising trend, carry out special rectification action in advance.

[0036] When a batch of food is found to have safety problems, the raw material supplier, production enterprise, circulation channel and affected downstream enterprises and consumers can be quickly traced and analyzed through the food data chain in the data set. According to the results of traceability analysis, the problem food is quickly recalled and processed to prevent the problem from spreading.

[0037] S50, according to the comprehensive risk score and the actual regulatory cost, through the interaction process of state perception, action selection, reward calculation and policy network update in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and regulatory efficiency.

[0038] Introduce reinforcement learning framework to dynamically adjust model parameters The core of reinforcement learning framework is to let the agent make decisions by perceiving the state of the environment. In the food safety supervision scenario, the state should be the following key information: current risk score , the current risk value output by the model as the feedback benchmark; sampling unqualified rate, consumer complaint volume, supply chain risk level, seasonal factor, historical event influence value, external environmental factor, data update frequency, time decay factor. These state variables together constitute the perception ability of the agent, so that it can fully understand the complexity of the current regulatory environment.

[0039] Define the adjustable parameter set in the risk prediction model as the action space of the agent, including but not limited to the weight coefficient of each risk factor in the model, the time decay factor, the data update frequency adjustment coefficient, etc. The action space is continuous or multi-discrete, ensuring that the agent has sufficient adjustment freedom to explore the optimal configuration.

[0040] To ensure the balance between flexibility and stability of parameter adjustment, a strict action constraint mechanism is set: the single adjustment range of all parameters is limited within , that is, each action can only make a small range of disturbance to the original parameter value. An adaptive step control strategy is introduced: when it is detected that consecutive adjustments all bring positive rewards, the upper limit of the adjustment range can be moderately relaxed to speed up the convergence; on the contrary, if consecutive negative rewards or a sudden drop in model prediction accuracy occur, the action range is automatically reduced to enhance the system robustness. To avoid system shock caused by frequent modification of key core parameters, a parameter freezing mechanism is set: for parameter combinations that have been running stably and have performed well for a long time, their entry into the action space is limited within a certain period, and only secondary parameters are allowed to participate in optimization, thereby improving the explainability and running stability of the overall system.

[0041] The reward function is the core driving force for the agent to learn, and needs to clearly reflect the degree of realization of the regulatory objectives. Through the reward function, the agent optimizes four parameters to make the risk score more accurate and the regulatory cost controllable. The reward calculation formula is , is the difference between the future risk score and the current score, is the regulatory cost weight, is the change in regulatory cost, including inspection frequency, recall rate, etc., is the stability weight, is the stability of parameter adjustment, and the smaller the parameter change range, the higher the reward.

[0042] The agent continuously optimizes the parameter adjustment strategy through interaction with the environment. The initialization parameters set the initial values of the parameters and record the current state. Based on the current state, the agent selects an adjustment strategy from the action space. The adjusted parameters are substituted into the prediction model to calculate the future risk score and regulatory efficiency. If the difference in risk score decreases and the regulatory cost decreases, the environment returns a positive reward; otherwise, a negative reward is returned. The state, action, reward and other information are stored in the experience replay buffer for subsequent training. Data is sampled from the experience replay, and the policy network is optimized through the policy gradient algorithm. During training, the agent will gradually learn which parameters to adjust in which states to be most effective.

[0043] Through reinforcement learning, the model can adaptively adjust the parameters and dynamically respond to changes in the regulatory environment, ultimately achieving the accuracy and robustness of risk prediction.

[0044] Embodiment Two As shown in Figure 2 , the embodiment two of the present application provides a food safety intelligent supervision system based on big data, comprising: Data collection module: Collect multi-source data from production, circulation, and supervision links. According to the food appearance in the data and the standard template, generate an appearance abnormality score through image recognition algorithm to assist supervision sampling. Including collection sub-module, image recognition sub-module.

[0045] Collection sub-module: Used for collecting data from production, circulation, and supervision links.

[0046] Collect production link data, including food production enterprise data, production process data, and quality control data.

[0047] Obtain food production enterprise registration information through the enterprise registration system of the government supervision platform and the supplier filing system, including enterprise name, address, production license scope, legal person information, etc. At the same time, obtain food raw material supplier information, including raw material origin, variety, qualification certification, etc.; use Internet of Things technology to deploy sensors in production workshops to collect production environment parameters in real time, including temperature, humidity, air quality, etc., and production equipment operation data, including equipment start-stop time, running time, key process parameters, etc.; quality control data is the enterprise internal quality detection report, including raw material, semi-finished product and finished product detection data, covering physicochemical indicators, microbial indicators, etc. At the same time, record the quality traceability information in the production process, including raw material batch, production batch, product flow direction, etc., to facilitate accurate traceability when problems occur.

[0048] Collect circulation link data, including logistics data, warehouse data, and market transaction data.

[0049] Cooperate with logistics enterprises to obtain detailed information during food transportation, including transportation route, transportation time, transportation tool, storage condition, etc. Through the installation of intelligent sensors in transportation vehicles and warehouses, real-time monitoring of environmental parameters is realized to ensure that food is transported and stored under suitable conditions; collect warehouse inventory information, including food types, quantity, shelf life, warehouse entry and exit time, etc. At the same time, obtain warehouse environmental monitoring data, including temperature and humidity, ventilation conditions, etc., to evaluate the impact of warehouse conditions on food quality; collect sales data from various food trading markets, supermarkets, and e-commerce platforms, including sales time, sales location, sales price, sales volume, etc., as well as market promotion activities and product recall information.

[0050] Collect supervision link data, including sampling data, administrative punishment data, and public opinion data.

[0051] Integrate the sampling information of food regulatory departments at all levels, including sampling location, sampling object, sampling project, test result, unqualified reason, etc.; collect the administrative punishment records of food production and operation enterprises, including illegal and irregular behavior, penalty basis, penalty result, etc. information; use web crawler technology to collect public opinion information related to food safety from social media platforms, news websites, food safety forums and other channels. Analyze the public opinion data through natural language processing technology, extract the public's attention to food safety hot events, consumer complaints and feedback, and timely identify potential food safety problems and public opinion risks.

[0052] Image recognition sub-module: analyze the appearance of food from the aspects of complete packaging, color abnormalities, and defect density through image recognition algorithms, and calculate the appearance abnormality score.

[0053] Cleaning and standardization module: clean and standardize the collected data, and based on the proportion of valid data, noise level, image clarity, and cleaning time, output the data quality score through the quality evaluation formula.

[0054] Clean the collected data at each link, remove duplicate data, incorrect data and incomplete data. Identify and delete duplicate test reports through a duplicate checking algorithm; correct or exclude obviously abnormal data collected by sensors; for records missing key information, such as product information missing production batch number, through data completion or deletion processing; extract unstructured text information through natural language processing; re-shoot highly blurred food pictures.

[0055] Standardize the data of different sources and formats using data extraction, transformation and loading tools, unify the data format and coding rules, and facilitate data integration and analysis.

[0056] Correlate the data from production, circulation, supervision and other links through unique identifiers, correlate the product batch in the production link with the transportation records, sales records in the circulation link and the sampling records in the supervision link, realize the information tracing and integration of food from production to consumption, and build a complete food safety data chain.

[0057] Model building module: according to the cleaned data set, build a risk prediction model that integrates historical event influence value, external environment, data update frequency and time decay factor, calculate and output the comprehensive risk score and key influence index of each food.

[0058] Statistical marking module: according to the comprehensive risk score, count the number of high-risk foods and the risk value of individual foods for each enterprise and region, identify high-risk subjects through double-threshold rules, mark them and trigger enhanced supervision measures.

[0059] The enterprises and regions where the food with high statistical risk value are located and the enterprises and regions where the risk food categories are more are displayed in real time through the monitoring platform, and the risk status of various food, enterprises and regions is marked and warned. The food categories exceeding the threshold number and the single food risk value exceeding the threshold are high risk, and the frequency and intensity of sampling inspection are increased for high-risk enterprises and regions; for food categories with obvious continuous risk rising trend, special rectification action is carried out in advance.

[0060] When it is found that a batch of food has safety problems, the raw material suppliers, production enterprises, circulation channels and affected downstream enterprises and consumers can be quickly traced and analyzed through the food data chain in the data set. According to the traceability analysis results, the problem food is quickly recalled and processed to prevent the problem from spreading.

[0061] Optimization adjustment parameter module: according to the comprehensive risk score and the actual supervision cost, through the interaction process of state perception, action selection, reward calculation and policy network update in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and supervision efficiency.

[0062] The reinforcement learning framework is introduced to dynamically adjust the model parameters. The core of the reinforcement learning framework is to let the agent make decisions by perceiving the state of the environment. In the food safety supervision scene, the state should be the following key information: the current risk score, the current risk value output by the model as the feedback benchmark; sampling unqualified rate, consumer complaint quantity, supply chain risk level, seasonal factor, historical event influence value, external environment factor, data update frequency, time decay factor. These state variables together constitute the perception ability of the agent, so that it can fully understand the complexity of the current supervision environment.

[0063] The adjustable parameter set in the risk prediction model is defined as the action space of the agent, including but not limited to the weight coefficient of each risk factor in the model, the time decay factor, the data update frequency adjustment coefficient, etc. The action space is continuous or multi-discrete space, which ensures that the agent has sufficient adjustment freedom to explore the optimal configuration.

[0064] In order to ensure the balance between flexibility and stability of parameter adjustment, a strict action constraint mechanism is set: the single adjustment amplitude of all parameters is limited to The range of each action is only a small perturbation of the original parameter value. An adaptive step size control strategy is introduced: when it is detected that consecutive adjustments bring positive rewards, the upper limit of the adjustment range can be moderately relaxed to speed up convergence; on the contrary, if consecutive negative rewards or a sudden drop in model prediction accuracy occur, the action range is automatically reduced to enhance system robustness. To avoid system shock caused by frequent modification of key core parameters, a parameter freezing mechanism is set: for parameter combinations that have been running stably and performing well for a long time, their entry into the action space is limited within a certain period, and only secondary parameters are allowed to participate in optimization, thereby improving the explainability and running stability of the overall system.

[0065] The reward function is the core driving force for agent learning, and needs to clearly reflect the degree of implementation of regulatory objectives. Through the reward function, the agent optimizes four parameters to make the risk score more accurate and the regulatory cost controllable.

[0066] The agent continuously optimizes the parameter adjustment strategy through interaction with the environment. The initialization parameters set the initial values for the parameters and record the current state. Based on the current state, the agent selects an adjustment strategy from the action space. The adjusted parameters are substituted into the prediction model to calculate the future risk score and regulatory efficiency. If the difference in risk score decreases and the regulatory cost decreases, the environment returns a positive reward; otherwise, a negative reward is returned. The state, action, reward, and other information are stored in the experience replay buffer for subsequent training. Data is sampled from the experience replay, and the policy network is optimized through the policy gradient algorithm. During training, the agent will gradually learn which parameters to adjust in which states to be most effective. Through reinforcement learning, the model can adaptively adjust parameters and dynamically respond to changes in the regulatory environment, ultimately achieving the accuracy and robustness of risk prediction.

[0067] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present application should be included within the scope of protection of the present application.

Claims

1. A big data-based food safety intelligent supervision method and system, characterized in that, Comprise: S10, collect multi-source data of production, circulation and supervision links, generate appearance abnormality score through image recognition algorithm according to food appearance in data and standard template to assist supervision sampling; S20, data cleaning and standardization processing is carried out on the collected data, and the data quality score is output through the quality evaluation formula based on the effective data proportion, noise level, image definition and cleaning time consumption in the data; S30, according to the cleaned data set, a risk prediction model is constructed by integrating historical event influence value, external environment, data update frequency and time decay factor, and the comprehensive risk score and key influence index of each food are calculated and output; S40, according to the comprehensive risk score, the number of high-risk foods of each enterprise and region and the risk value of single food are counted, and the high-risk subject is identified through setting double threshold rule, and the key mark is triggered to trigger the strengthened supervision measures; S50, according to the comprehensive risk score and the actual supervision cost, through the interaction process of state perception, action selection, reward calculation and strategy network update in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and supervision efficiency.

2. The big data-based food safety intelligent supervision method according to claim 1, characterized in that, Collect multi-source data of production, circulation and supervision links, generate appearance abnormality score through image recognition algorithm according to food appearance in data and standard template to assist supervision sampling, which is divided into the following sub-steps: Multi-source structured and unstructured data of production, circulation and supervision links are collected through government systems, Internet of Things devices and network crawlers; The image recognition algorithm is used to compare the food images taken on site with the standard template, and the packaging integrity, color abnormality and defect density characteristics are extracted; Based on the image analysis results, the appearance abnormality score is calculated as the basis for preliminary screening of high-risk foods, which assists in the priority decision of sampling.

3. The big data-based food safety intelligent supervision method according to claim 1, characterized in that, Data cleaning and standardization processing is carried out on the collected data, and the data quality score is output through the quality evaluation formula based on the effective data proportion, noise level, image definition and cleaning time consumption in the data, which is divided into the following sub-steps: Through duplicate checking, error correction, completion and unstructured information extraction method, multi-source data is cleaned and standardized; The extraction, conversion and loading tools are used to unify the data format and coding rules, and the cross-link association of production, circulation and supervision data is realized through unique identification; Based on the effective data proportion, noise level, image definition and cleaning time consumption, the data quality score is calculated and output through the quality evaluation formula.

4. The big data-based food safety intelligent supervision method according to claim 1, characterized in that, According to the cleaned data set, a risk prediction model is constructed by integrating historical event influence value, external environment, data update frequency and time decay factor, and the comprehensive risk score and key influence index of each food are calculated and output, which is divided into the following sub-steps: According to the frequency and severity of past food safety problems in the data chain, the historical event influence value is determined and used as the core risk factor; According to the historical event influence value, a dynamically weighted risk prediction model is constructed, the model is run to calculate the comprehensive risk score of each food, and the key indicators influencing the score are output through feature analysis.

5. The big data-based food safety intelligent supervision method according to claim 1, characterized in that, According to the comprehensive risk score, the number of high-risk food of each enterprise and region and the risk value of individual food are counted, and the high-risk subject is identified by setting double threshold rules to trigger enhanced supervision measures. The specific steps are as follows: Count the enterprises and regions where the food with high risk value are located, and the enterprises and regions with many kinds of risk food; Set single food risk threshold and high-risk category quantity threshold. If one of them is met, it is determined as a high-risk subject. The identified high-risk subject is marked and the supervision measures are triggered.

6. The big data-based food safety intelligent supervision method according to claim 1, wherein, According to the comprehensive risk score and the actual supervision cost, through the interaction process of state perception, action selection, reward calculation and policy network update in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and supervision efficiency. The specific steps are as follows: Set the risk model parameters as adjustable action space. The agent selects adjustment strategy according to the current prediction effect and supervision cost. Build an environmental feedback mechanism to evaluate the strategy by improving the risk score and reducing the supervision cost. Update the strategy network through experience replay and policy gradient algorithm to continuously optimize the model parameters and improve the long-term prediction performance.

7. The big data-based food safety intelligent supervision method according to claim 6, characterized in that, Set the risk model parameters as adjustable action space. The agent selects adjustment strategy according to the current prediction effect and supervision cost. The specific steps are as follows: Define the key parameters in the risk prediction model as a limited action space to ensure that the agent makes fine adjustments within a small range and avoid model instability caused by drastic parameter changes. Introduce an adaptive mechanism to dynamically adjust the action step according to the reward feedback, which accelerates the convergence while maintaining system stability. Set the parameter freezing strategy to temporarily lock the parameters that perform stably, reduce redundant adjustments, and improve the model interpretability and operational reliability.

8. A big data-based food safety intelligent supervision system, characterized in that, It includes: Data collection module: Collect multi-source data from production, circulation and supervision links. According to the food appearance and standard template in the data, generate appearance anomaly score through image recognition algorithm to assist supervision sampling; Cleaning and standardization module: Clean and standardize the collected data. Based on the effective data proportion, noise level, image clarity and cleaning time in the data, output data quality score through quality evaluation formula; Model construction module: According to the cleaned data set, construct a risk prediction model that integrates historical event influence value, external environment, data update frequency and time decay factor, calculate and output the comprehensive risk score and key influence indicators of each food; Statistical marking module: According to the comprehensive risk score, count the number of high-risk food of each enterprise and region and the risk value of individual food. Identify high-risk subjects by setting double threshold rules, mark them and trigger enhanced supervision measures; Optimization and adjustment parameter module: According to the comprehensive risk score and the actual supervision cost, through the interaction process of state perception, action selection, reward calculation and policy network update in the reinforcement learning framework, the risk model parameters are dynamically optimized to improve the long-term prediction accuracy and supervision efficiency.