Food safety incident intelligent early warning system based on multi-source data analysis
Through the intelligent early warning system for food safety incidents integrating multi-source data analysis and random forest models, the problem of untimely and inaccurate pollution source detection in the existing technology is solved, and timely identification and early warning of various potential pollution sources in the food production environment is achieved, and the level of food safety management is improved.
Patent Information
- Application Number
- CN202510079664.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN119990819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of food safety early warning, and in particular to an intelligent early warning system for food safety incidents based on multi-source data analysis. Background Art
[0002] As a key issue of public health and social stability, food safety has always been a core issue of great concern to the government and various industries. In the field of food safety, pollution source control in the production process is one of the most important research directions. Ensuring that the food production process is not contaminated by harmful substances or microorganisms is a basic requirement for food production safety. On a more specific level, with the development of modern food production and processing technology, complex production environments and diverse food ingredients have made the detection and control of pollution sources increasingly complex. For pollution sources in these complex environments, how to accurately and efficiently detect and analyze potential pollution risks has become one of the key issues in improving the level of food safety management.
[0003] At present, the detection of pollution sources in the food production process mostly relies on traditional single sensor data or manual inspection and monitoring. Common detection methods include the use of temperature and humidity sensors, air quality sensors, and manual sampling inspections. Although these methods can detect pollution sources to a certain extent, they have great limitations.
[0004] The above-mentioned deficiencies of traditional food safety testing methods have resulted in the failure to timely and accurately identify and warn of the risks of contamination sources in many practical application scenarios. Once an environmental factor is abnormal, it may lead to contamination of finished food products or even a large-scale food safety incident. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides an intelligent early warning system for food safety incidents based on multi-source data analysis, which solves the problems mentioned in the background technology.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an intelligent early warning system for food safety incidents based on multi-source data analysis, including a data acquisition module, a data preprocessing module, a pollution source risk assessment module, an early warning detection and alarm module, an intelligent decision support module and a feedback and optimization module;
[0007] The data acquisition module collects raw data from sensors and acquisition devices, including temperature and humidity data TH, air quality index data PC and material composition data Cm, and performs fitting to form a safety data set SPA;
[0008] The data preprocessing module is responsible for cleaning and normalizing the security data set SPA to obtain a normalized data set SPG;
[0009] The pollution source risk assessment module uses a random forest model to perform a risk assessment on the normalized data set SPG to obtain a pollution risk value Risk;
[0010] The early warning detection and alarm module monitors and detects in real time whether abnormal fluctuations occur according to the obtained pollution risk value Risk, and obtains the early warning state At;
[0011] The intelligent decision support module provides adjustment measures according to the detected warning state At;
[0012] The feedback and optimization module adjusts the random forest model according to abnormal fluctuations and adjustment measures.
[0013] Preferably, the data acquisition module includes a data acquisition unit and a data fitting unit;
[0014] The data acquisition unit collects temperature and humidity data TH, air quality index data PC and material composition data Cm from sensors and acquisition devices;
[0015] The temperature and humidity data TH includes temperature data T and humidity data H. The temperature data T is collected by a temperature sensor, and the humidity data H is collected by a humidity sensor.
[0016] Air quality index data PC includes PM2.5 concentration and CO2 concentration, and is collected by air quality sensors;
[0017] Material composition data Cm includes chemical composition Cem, bacterial content Cbt and heavy metal content Chm. The heavy metal content Chm is collected by atomic absorption spectrometer, the bacterial content Cbt is collected by microbial detector, and the number of bacteria is calculated by ATP content. The chemical composition Cem is collected by spectrum analyzer.
[0018] The data fitting unit summarizes the collected temperature and humidity data TH, air quality index data PC and material composition data Cm to form a safety data set SPA.
[0019] Preferably, the data preprocessing module includes a data cleaning unit and a data normalization unit;
[0020] The data cleaning unit performs missing value filling and outlier detection on the security data set SPA to obtain a cleaned data set SPC;
[0021] Missing value imputation involves using interpolation to fill in missing data;
[0022] Outlier detection uses the Z-Score method to detect outliers in the data and remove them;
[0023] The data normalization unit converts the cleaned data set SPC into a unified standard range, eliminates the dimensional differences between different data sources, and obtains a normalized data set SPG;
[0024] The normalized data set SPG is obtained by the following formula:
[0025]
[0026] Wherein, SPGd represents the d-th data item in the normalized data set, SPCd represents the d-th data item in the cleaned data set, SPCdmax represents the peak value of the d-th data item in the cleaned data set, and SPCdmin represents the valley value of the d-th data item in the cleaned data set.
[0027] Preferably, the pollution source risk assessment module includes a feature extraction unit and a random forest model assessment unit;
[0028] The feature extraction unit extracts features from the normalized data set SPG, including environmental monitoring data features, material composition change trend features, and cross-contamination features, and forms a contamination feature set F;
[0029] Environmental monitoring data characteristics include temperature change rate ΔT, humidity change rate ΔH, and air quality change rate ΔPC;
[0030] The temperature change rate ΔT is obtained by the ratio of the difference between the temperature data at time t and the temperature data at time t-1 to the time interval Δt;
[0031] The humidity change rate ΔH is obtained by the ratio of the difference between the humidity data at time t and the humidity data at time t-1 to the time interval Δt;
[0032] The air quality change rate ΔPC is obtained by the ratio of the difference between the air quality index data at time t and the air quality index data at time t-1 to the time interval Δt;
[0033] The characteristics of the changing trends of material composition include the mean value of chemical composition μCem, the fluctuation of heavy metal content σChm and the mean value of bacteria μCbt;
[0034] The chemical component mean μCem is obtained by extracting the average concentration of the chemical component Cem;
[0035] The bacterial mean μCbt is obtained by recording the average level of bacterial content:
[0036] The heavy metal content fluctuation σChm is obtained by the following formula:
[0037]
[0038] In the formula, Chmi represents the heavy metal content at the i-th moment, μChm represents the mean value of the heavy metal content, and N represents the total number of collected moments;
[0039] The cross-contamination feature is obtained by extracting the correlation coefficient between pollutants and using the Pearson correlation coefficient to measure the impact between pollution sources:
[0040] The cross contamination feature extraction formula is:
[0041]
[0042] Wherein, Corr(D1, D2) represents the Pearson correlation coefficient between pollutant D1 and pollutant D2. The pollutants include PM2.5, CO2, chemical composition Cem, bacterial content Cbt and heavy metal content Chm. j represents the data point, m represents the total number of data points, μD1 represents the mean of pollutant D1, μD2 represents the mean of pollutant D2, D1,j represents the jth data point of pollutant D1, and D2,j represents the jth data point of pollutant D2.
[0043] Preferably, the random forest model evaluation unit trains the pollution source risk assessment model based on the random forest algorithm according to the pollution feature set F, trains multiple decision trees, and makes a final decision by weighted averaging the prediction results of each tree;
[0044] By using the pollution feature set F to train the random forest model, multiple decision trees are constructed. The training data of each tree is F(a, tr), the output of the decision tree is y(tr), and the pollution risk value Risk is calculated;
[0045] The pollution risk value Risk is expressed by the following formula:
[0046]
[0047] In the formula, y(a, tr) represents the prediction result of the ath tree, and M represents the total number of trees;
[0048] Compare the obtained contamination risk value Risk with the risk threshold TRi to determine the cross-contamination risk status;
[0049] The cross contamination risk status is obtained by matching in the following ways:
[0050] When the contamination risk value Risk ≥ risk threshold TRi, it indicates that there is a risk of cross contamination.
[0051] Preferably, the early warning detection and alarm module includes a risk fluctuation detection unit and an alarm judgment unit;
[0052] The risk fluctuation detection unit monitors the real-time fluctuation of the pollution risk value Risk, detects whether there is an abnormal fluctuation state, analyzes the change trend of the pollution risk value Risk, calculates the fluctuation amplitude ΔRisk, and compares it with the preset fluctuation threshold ΔTR to determine the abnormal fluctuation state;
[0053] The fluctuation range ΔRisk is obtained by the following formula:
[0054] ΔRisk=|Risk(t)-Risk(t-1)|;
[0055] In the formula, Risk(t) represents the pollution risk value at time t, and Risk(t-1) represents the pollution risk value at time t-1;
[0056] The abnormal fluctuation state is obtained by matching in the following way:
[0057] When the fluctuation amplitude ΔRisk ≥ the fluctuation threshold ΔTR, it indicates that abnormal fluctuation occurs;
[0058] When the fluctuation amplitude ΔRisk is less than the fluctuation threshold ΔTR, it indicates that there is no abnormal fluctuation.
[0059] Preferably, the alarm judgment unit judges the triggering of the warning state At according to the acquired pollution risk value Risk and the fluctuation range ΔRisk;
[0060] The warning state At is obtained by the following formula:
[0061]
[0062] In the formula, yes means triggering an alert, and no means not triggering an alert.
[0063] Preferably, the intelligent decision support module includes a decision judgment unit and an adjustment measure generation unit;
[0064] The decision-making unit obtains a judgment result based on the detected warning state At and related parameters, including temperature and humidity data TH, air quality index data PC and material composition data Cm, to determine whether adjustment measures need to be taken;
[0065] When the warning state At is yse, the temperature and humidity data TH is judged and compared with the preset temperature threshold AT and the preset humidity threshold AH to judge the relationship between the temperature and humidity data TH and the pollution risk;
[0066] When the pollution risk value Risk ≥ risk threshold TRi, and the temperature data T ≥ temperature threshold AT, and the humidity data H ≥ humidity threshold AH, it means that the pollution risk is related to the temperature and humidity data TH, specifically, the temperature data T is abnormal and the humidity data H is abnormal;
[0067] When the warning state At is yse, the relationship between the air quality index data PC and the material composition data Cm is judged, and compared with the preset air quality threshold TPC to judge the relationship between the air quality, material composition and pollution risk;
[0068] When the pollution risk value Risk ≥ risk threshold TRi, and the air quality index data PC > air quality threshold TPC, and harmful components are detected in the material composition data Cm, it means that the pollution risk is related to the air quality index data PC and the material composition data Cm, specifically, the air quality index data PC is abnormal and harmful components are detected.
[0069] Preferably, the adjustment measure generating unit adjusts the measure according to the judgment result;
[0070] When the contamination risk is related to the temperature and humidity data TH, the ventilation is adjusted;
[0071] When the air quality index data PC and material composition data Cm are related to pollution risks, production will be suspended and materials will be replaced.
[0072] Preferably, the feedback and optimization module includes a model feedback unit and a model updating unit;
[0073] The model feedback unit is responsible for collecting the fluctuation amplitude ΔRisk, the pollution risk value Risk and the pollution feature set F, and inputting them into the random forest model for retraining, and calculating the pollution error Error between the pollution risk value YRisk predicted by the model and the pollution risk value Risk actually measured;
[0074] The contamination error Error is obtained by the following formula:
[0075] Error = |YRisk-Risk|;
[0076] The contamination error Error is recorded to form an error set FEr, the long-term performance of the random forest model is evaluated, and compared with the preset error threshold TErr;
[0077] When the contamination error Error>error threshold TErr, the model update is triggered to update the random forest model;
[0078] The model updating unit continuously adjusts the random forest model according to the contamination error Error as a feedback signal;
[0079] The fluctuation range ΔRisk, the pollution risk value Risk and the pollution feature set F are combined and fitted into the data training set Fnew, and then brought into the random forest model. The random forest model is retrained, and the prediction ability of the random forest model is adjusted by adjusting the decision tree, and a new prediction result y(tr, new) is obtained, and the pollution risk value Risk is recalculated;
[0080] The new prediction result y(tr, new) is obtained by the following formula:
[0081] y(tr,new)=RF(Fnew);
[0082] Where RF(Fnew) represents the random forest model trained using the data training set Fnew.
[0083] The present invention provides an intelligent early warning system for food safety incidents based on multi-source data analysis, which has the following beneficial effects:
[0084] (1) When the system is running, it can enhance the adaptability and flexibility of the model through continuous self-optimization and adjustment. As the amount of data increases, the system can become more accurate and efficient, avoiding the problems caused by long-term reliance on fixed rules and outdated models. By integrating data from multiple sensors, the system can promptly detect multiple potential sources of pollution and identify the risk of cross-contamination, avoiding risk factors that may be overlooked by traditional methods. The system's intelligent decision support function greatly reduces the burden of manual intervention, improves the timeliness and accuracy of response, and avoids delays or errors caused by human judgment.
[0085] Early warning can significantly improve response speed, reduce the probability of food safety problems, reduce safety hazards in food production, and protect the health and safety of consumers. Using the random forest model to conduct pollution source risk assessment enables the system to identify potential pollution risk points based on a large number of features and complex relationships, and provide more accurate warnings, thereby effectively preventing the occurrence of food safety incidents. High-quality data input is the basis for accurate warning. Through preprocessing, the system can reduce the misleading caused by inaccurate data and ensure the scientificity and reliability of subsequent analysis and decision-making.
[0086] (2) Through the integration of multi-source data, the system can comprehensively monitor key parameters in the food production process and detect potential risks from multiple sources of contamination in advance, rather than relying solely on a single piece of data. This diversified data collection method can effectively avoid the data blind spots in traditional monitoring methods and improve the accuracy and reliability of the system.
[0087] Each link of data preprocessing provides high-quality data support for subsequent pollution source risk assessment and intelligent decision-making. Through accurate risk assessment, the system can detect potential safety hazards in real time and issue warnings in a timely manner, providing sufficient response time for food safety problems that may arise during the production process, thereby reducing the incidence of food safety accidents. By aggregating key data from different sources, the system can evaluate the comprehensive safety status of the food production environment instead of relying solely on a single data source. This multi-dimensional aggregation method enables the system to identify potential cross-contamination risks and safety issues caused by multiple factors, improving the accuracy and comprehensiveness of warnings.
[0088] (3) The extraction of cross-contamination features makes up for the shortcomings of traditional single pollution source analysis methods and can identify the interactions and impacts between multiple pollutants. By capturing the correlation between pollution sources, the system can identify cross-contamination problems in complex environments earlier and provide more comprehensive and detailed data support for the early warning system. Especially in a multi-pollution source environment, this feature extraction method enables the system to promptly identify and respond to the complex risks brought about by cross-contamination.
[0089] The multi-dimensional feature set F provides a rich source of information for the random forest model, enabling the model to make more accurate risk assessments based on complex and ever-changing data. Through the ensemble learning method of the random forest model, the system can handle various complex relationships and accurately identify potential pollution risks, greatly improving the accuracy and robustness of risk assessments.
[0090] (4) By comparing the pollution risk value Risk with the set risk threshold TRi, determine whether there is a cross-contamination risk. When the Risk value exceeds the set threshold, the system automatically determines that there is a cross-contamination risk and issues a corresponding warning. By setting the pollution risk threshold TRi, the system can flexibly control the sensitivity of the risk warning. When the pollution risk reaches a certain threshold, the system can issue a warning in time to prevent the cross-contamination risk from causing serious impact on the production environment. This judgment mechanism based on the risk threshold enables the system to activate the warning function in time when the pollution source risk changes, thereby ensuring food safety in the production process. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 The present invention is a block diagram and flow chart of an intelligent early warning system for food safety incidents based on multi-source data analysis. DETAILED DESCRIPTION
[0092] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0093] Example 1
[0094] The present invention provides an intelligent early warning system for food safety incidents based on multi-source data analysis. Figure 1 , including data acquisition module, data preprocessing module, pollution source risk assessment module, early warning detection and alarm module, intelligent decision support module and feedback and optimization module;
[0095] The data acquisition module collects raw data from sensors and acquisition devices, including temperature and humidity data TH, air quality index data PC and material composition data Cm, and performs fitting to form a safety data set SPA;
[0096] The data preprocessing module is responsible for cleaning and normalizing the security data set SPA to obtain a normalized data set SPG;
[0097] The pollution source risk assessment module uses a random forest model to perform a risk assessment on the normalized data set SPG to obtain a pollution risk value Risk;
[0098] The early warning detection and alarm module monitors and detects in real time whether abnormal fluctuations occur according to the obtained pollution risk value Risk, and obtains the early warning state At;
[0099] The intelligent decision support module provides adjustment measures according to the detected warning state At;
[0100] The feedback and optimization module adjusts the random forest model according to abnormal fluctuations and adjustment measures.
[0101] In this embodiment, by integrating multiple sensors, environmental monitoring equipment and external data sources, the system can collect temperature and humidity data TH, air quality indicators PC and material composition data Cm in real time. Compared with the traditional monitoring method of a single data source, this multi-source data collection method can more comprehensively capture possible sources of pollution in the food production process. Diversified data input enables the system to more accurately depict the dynamic changes in the production environment and reduce errors caused by a single factor. The data preprocessing module cleans and normalizes the collected safety data set, thereby eliminating the interference of data noise and outliers, ensuring the accuracy and reliability of subsequent analysis. Cleaning and normalization make all types of data consistent, which is convenient for model analysis and risk assessment.
[0102] The pollution source risk assessment module uses the random forest model to analyze the normalized data and assess the potential pollution risk. Different from the traditional simple threshold comparison method, the random forest model improves the accuracy and stability of the prediction through the integration of multiple decision trees. The early warning detection and alarm module monitors the changes in the pollution risk value in real time and detects whether there are abnormal fluctuations. Once the risk value exceeds the preset threshold, the system immediately issues an early warning to prompt relevant personnel to intervene in time. Through real-time monitoring and automatic detection, the system can identify and deal with possible safety hazards in advance to avoid the spread and expansion of the incident.
[0103] The intelligent decision support module provides adjustment measures based on the real-time warning status and automatically generates response plans. For example, if abnormal temperature and humidity are detected and the risk of contamination is high, the system will recommend increasing ventilation; if the material composition contains harmful substances, the system will recommend suspending production and replacing materials. Through the feedback and optimization module, the system can continuously adjust and optimize the random forest model based on abnormal fluctuations monitored in real time and the adjustment measures taken. By continuously learning new data, the system continuously improves its decision-making ability and prediction accuracy, allowing the model to adapt to the changing production environment.
[0104] Example 2
[0105] This embodiment is explained in Example 1, please refer to Figure 1 ,Specifically: the data acquisition module includes a data acquisition unit and a data fitting unit;
[0106] The data acquisition unit collects temperature and humidity data TH, air quality index data PC and material composition data Cm from sensors and acquisition devices;
[0107] The temperature and humidity data TH includes temperature data T and humidity data H. The temperature data T is collected by a temperature sensor, and the humidity data H is collected by a humidity sensor.
[0108] Air quality index data PC includes PM2.5 concentration and CO2 concentration, and is collected by air quality sensors;
[0109] Material composition data Cm includes chemical composition Cem, bacterial content Cbt and heavy metal content Chm. The heavy metal content Chm is collected by atomic absorption spectrometer, the bacterial content Cbt is collected by microbial detector, and the number of bacteria is calculated by ATP content. The chemical composition Cem is collected by spectrum analyzer.
[0110] The data fitting unit summarizes the collected temperature and humidity data TH, air quality index data PC and material composition data Cm to form a safety data set SPA.
[0111] The data preprocessing module includes a data cleaning unit and a data normalization unit;
[0112] The data cleaning unit performs missing value filling and outlier detection on the security data set SPA to obtain a cleaned data set SPC;
[0113] Missing value imputation involves using interpolation to fill in missing data;
[0114] Outlier detection uses the Z-Score method to detect outliers in the data and remove them;
[0115] The data normalization unit converts the cleaned data set SPC into a unified standard range, eliminates the dimensional differences between different data sources, and obtains a normalized data set SPG;
[0116] The normalized data set SPG is obtained by the following formula:
[0117]
[0118] Wherein, SPGd represents the d-th data item in the normalized data set, SPCd represents the d-th data item in the cleaned data set, SPCdmax represents the peak value of the d-th data item in the cleaned data set, and SPCdmin represents the valley value of the d-th data item in the cleaned data set.
[0119] In this embodiment, the data acquisition module integrates a variety of sensors and environmental monitoring equipment, including temperature and humidity sensors, air quality sensors, material composition detection instruments, etc., which can comprehensively monitor changes in the production environment and material composition. Specifically, temperature and humidity data are collected through temperature and humidity sensors, air quality data are obtained through air quality sensors for PM2.5 concentration and CO2 concentration, and material composition data includes the collection of chemical composition, bacterial content, and heavy metal content.
[0120] The data cleaning unit is responsible for filling missing values and detecting outliers in the collected raw data to ensure data quality. Missing data is filled by interpolation, and outliers are detected and eliminated by the Z-Score method, thereby removing noise and interference in the data and ensuring data accuracy and consistency. The data normalization unit standardizes the cleaned data, converts data from different sources to a unified standard range, and eliminates dimensional differences between different data sources. This process ensures that the weight of each data item in subsequent analysis is consistent, thereby improving the accuracy of the model.
[0121] After data cleaning and normalization, the generated normalized data set SPG can be used as the input of the pollution source risk assessment model to provide accurate data for the risk assessment module. The random forest model can assess the risk of pollution sources based on these data and warn of potential safety hazards in advance. Through continuous optimization and adjustment, the system can accumulate more data in long-term operation and improve its ability to predict various complex situations.
[0122] Example 3
[0123] This embodiment is explained in Example 2. Please refer to Figure 1 ,Specifically: the pollution source risk assessment module includes a feature extraction unit and a random forest model evaluation unit;
[0124] The feature extraction unit extracts features from the normalized data set SPG, including environmental monitoring data features, material composition change trend features, and cross-contamination features, and forms a contamination feature set F;
[0125] Environmental monitoring data characteristics include temperature change rate ΔT, humidity change rate ΔH, and air quality change rate ΔPC;
[0126] The temperature change rate ΔT is obtained by the ratio of the difference between the temperature data at time t and the temperature data at time t-1 to the time interval Δt;
[0127] The humidity change rate ΔH is obtained by the ratio of the difference between the humidity data at time t and the humidity data at time t-1 to the time interval Δt;
[0128] The air quality change rate ΔPC is obtained by the ratio of the difference between the air quality index data at time t and the air quality index data at time t-1 to the time interval Δt;
[0129] The characteristics of the changing trends of material composition include the mean value of chemical composition μCem, the fluctuation of heavy metal content σChm and the mean value of bacteria μCbt;
[0130] The chemical component mean μCem is obtained by extracting the average concentration of the chemical component Cem;
[0131] The bacterial mean μCbt is obtained by recording the average level of bacterial content:
[0132] The heavy metal content fluctuation σChm is obtained by the following formula:
[0133]
[0134] In the formula, Chmi represents the heavy metal content at the i-th moment, μChm represents the mean value of the heavy metal content, and N represents the total number of collected moments;
[0135] The cross-contamination feature is obtained by extracting the correlation coefficient between pollutants and using the Pearson correlation coefficient to measure the impact between pollution sources:
[0136] The cross contamination feature extraction formula is:
[0137]
[0138] Wherein, Corr(D1, D2) represents the Pearson correlation coefficient between pollutant D1 and pollutant D2. Pollutants include PM2.5, CO2, chemical composition Cem, bacterial content Cbt and heavy metal content Chm, j represents the data point, m represents the total number of data points, μD1 represents the mean of pollutant D1, μD2 represents the mean of pollutant D2, D1,j represents the jth data point of pollutant D1, and D2,j represents the jth data point of pollutant D2.
[0139] In this embodiment, the feature extraction unit provides dynamic environmental change information for subsequent risk assessment of pollution sources by calculating environmental monitoring data features such as temperature change rate ΔT, humidity change rate ΔH, and air quality change rate ΔPC. By paying attention to these change trends, the system can more accurately identify abnormal fluctuations in the environment and discover possible pollution risks in advance.
[0140] The feature extraction unit also extracts the change trend characteristics of the material composition, including the chemical composition mean μCem, heavy metal content fluctuation σChm and bacterial mean μCbt. These material composition characteristics can reflect the quality fluctuation and potential pollution problems in the production process. In particular, the calculation of the heavy metal content fluctuation σChm can provide the system with more comprehensive trend change information by considering all data at the time of collection.
[0141] The cross-contamination feature extracts the correlation coefficient between pollutants and uses the Pearson correlation coefficient to measure the mutual influence between different pollution sources. This method can capture the potential correlation between different pollution sources and reveal the risks under the joint action of multiple pollution sources. By extracting the characteristics of environmental monitoring data, material composition change characteristics and cross-contamination characteristics, the feature extraction unit combines these characteristics into a pollution feature set F, which provides input data for the random forest model evaluation unit. The random forest model can efficiently process and evaluate these multidimensional characteristics, generate pollution risk values, and further accurately identify and evaluate pollution sources.
[0142] The system can provide flexible early warning and decision support based on environmental changes, material characteristics and the relationship between pollution sources, which enables it to cope with various complex situations in practical applications. Whether in the case of a single pollution source or in a cross-contamination environment, the system can make accurate risk assessments and timely response suggestions.
[0143] Compared with static monitoring, this method of extracting environmental features based on dynamic changes can more keenly capture small fluctuations in environmental changes, thereby improving the sensitivity and accuracy of pollution source risk assessment. In-depth analysis of material composition enables the system to not only monitor the concentration of existing pollutants, but also identify potential risks in changes in material composition. Fluctuations in heavy metal content can reflect the instability of material sources or potential contamination in processing. In this way, the system can identify abnormal fluctuations in material composition in advance and prevent pollutants from entering the production process through raw materials.
[0144] Example 4
[0145] This embodiment is explained in Example 3, please refer to Figure 1 Specifically: the random forest model evaluation unit trains the pollution source risk assessment model based on the random forest algorithm according to the pollution feature set F, trains multiple decision trees, and makes a final decision by weighted averaging the prediction results of each tree;
[0146] By using the pollution feature set F to train the random forest model, multiple decision trees are constructed. The training data of each tree is F(a, tr), the output of the decision tree is y(tr), and the pollution risk value Risk is calculated;
[0147] The pollution risk value Risk is expressed by the following formula:
[0148]
[0149] In the formula, y(a, tr) represents the prediction result of the ath tree, and M represents the total number of trees;
[0150] Compare the obtained contamination risk value Risk with the risk threshold TRi to determine the cross-contamination risk status;
[0151] The cross contamination risk status is obtained by matching in the following ways:
[0152] When the contamination risk value Risk ≥ risk threshold TRi, it indicates that there is a risk of cross contamination.
[0153] The early warning detection and alarm module includes a risk fluctuation detection unit and an alarm judgment unit;
[0154] The risk fluctuation detection unit monitors the real-time fluctuation of the pollution risk value Risk, detects whether there is an abnormal fluctuation state, analyzes the change trend of the pollution risk value Risk, calculates the fluctuation amplitude ΔRisk, and compares it with the preset fluctuation threshold ΔTR to determine the abnormal fluctuation state;
[0155] The fluctuation range ΔRisk is obtained by the following formula:
[0156] ΔRisk=|Risk(t)-Risk(t-1)|;
[0157] In the formula, Risk(t) represents the pollution risk value at time t, and Risk(t-1) represents the pollution risk value at time t-1;
[0158] The abnormal fluctuation state is obtained by matching in the following way:
[0159] When the fluctuation amplitude ΔRisk ≥ the fluctuation threshold ΔTR, it indicates that abnormal fluctuation occurs;
[0160] When the fluctuation amplitude ΔRisk is less than the fluctuation threshold ΔTR, it indicates that there is no abnormal fluctuation.
[0161] The alarm judgment unit judges the triggering of the warning state At according to the obtained pollution risk value Risk and the fluctuation range ΔRisk;
[0162] The warning state At is obtained by the following formula:
[0163]
[0164] In the formula, yes means triggering an alert, and no means not triggering an alert.
[0165] In this embodiment, a random forest algorithm is used to perform pollution source risk assessment by training multiple decision trees. Each decision tree is trained according to the pollution feature set F, and then the prediction results of each tree are weighted averaged to finally calculate the pollution risk value Risk. The ensemble learning method of random forest reduces the overfitting problem of a single tree through the cooperation of multiple decision trees, and improves the accuracy of the system's risk assessment of pollution sources in complex environments. Through the training and weighted averaging of multiple decision trees, the system can better capture the complex relationship between different features and pollution sources, and improve the reliability of pollution risk prediction. Compared with the traditional single decision tree model, random forests show stronger robustness and accuracy when processing multi-dimensional and nonlinear data.
[0166] By comparing the pollution risk value Risk with the set risk threshold TRi, it is determined whether there is a risk of cross contamination. When the Risk value exceeds the set threshold, the system automatically determines that there is a risk of cross contamination and issues a corresponding warning. By setting the pollution risk threshold TRi, the system can flexibly control the sensitivity of the risk warning. When the pollution risk reaches a certain threshold, the system can issue a warning in time to prevent the cross contamination risk from causing serious impact on the production environment. This judgment mechanism based on the risk threshold enables the system to activate the warning function in time when the risk of the pollution source changes, thereby ensuring food safety in the production process.
[0167] In order to improve the response speed to changes in pollution sources, this embodiment further introduces a fluctuation detection mechanism for pollution risk values. The risk fluctuation detection unit monitors the real-time fluctuation of the pollution risk value Risk, analyzes the changing trend of the pollution risk value, calculates the fluctuation amplitude ΔRisk, and compares it with the preset fluctuation threshold ΔTR to determine whether abnormal fluctuations occur. This real-time fluctuation monitoring enables the system to respond to instantaneous changes in pollution risk values and promptly identify abnormal fluctuations in pollution sources. When the risk fluctuation exceeds the normal range, the system can quickly make adjustments or issue an early warning to effectively avoid the spread or aggravation of potential pollution risks. This makes the system more adaptable in a dynamic environment, able to respond to small changes in the environment, and provide real-time and accurate risk assessments.
[0168] By comparing the fluctuation amplitude ΔRisk with the preset fluctuation threshold ΔTR, the system can determine whether there is an abnormal fluctuation in the pollution risk, and trigger the warning state At through the alarm judgment unit. When the fluctuation amplitude exceeds the preset threshold, the system will issue a warning signal to prompt the operator to take measures. The introduction of the fluctuation detection and early warning mechanism enables the system to identify small fluctuations in the pollution source and trigger an early warning in time based on this fluctuation. This can not only help users discover problems earlier, but also provide a strong data basis for subsequent decision support and optimization adjustments by analyzing the trend of pollution risk fluctuations. Compared with traditional static monitoring, the fluctuation-based early warning mechanism can respond to environmental changes more sensitively and reduce the risk of sudden food safety problems.
[0169] Example 5
[0170] This embodiment is explained in Example 4. Please refer to Figure 1 ,Specifically: the intelligent decision support module includes a decision judgment unit and an adjustment measure generating unit;
[0171] The decision-making unit obtains a judgment result based on the detected warning state At and related parameters, including temperature and humidity data TH, air quality index data PC and material composition data Cm, to determine whether adjustment measures need to be taken;
[0172] When the warning state At is yse, the temperature and humidity data TH is judged and compared with the preset temperature threshold AT and the preset humidity threshold AH to judge the relationship between the temperature and humidity data TH and the pollution risk;
[0173] When the pollution risk value Risk ≥ risk threshold TRi, and the temperature data T ≥ temperature threshold AT, and the humidity data H ≥ humidity threshold AH, it means that the pollution risk is related to the temperature and humidity data TH, specifically, the temperature data T is abnormal and the humidity data H is abnormal;
[0174] When the warning state At is yse, the relationship between the air quality index data PC and the material composition data Cm is judged, and compared with the preset air quality threshold TPC to judge the relationship between the air quality, material composition and pollution risk;
[0175] When the pollution risk value Risk ≥ risk threshold TRi, and the air quality index data PC > air quality threshold TPC, and harmful components are detected in the material composition data Cm, it means that the pollution risk is related to the air quality index data PC and the material composition data Cm, specifically, the air quality index data PC is abnormal and harmful components are detected.
[0176] The adjustment measure generating unit adjusts the measures according to the judgment result;
[0177] When the contamination risk is related to the temperature and humidity data TH, the ventilation is adjusted;
[0178] When the air quality index data PC and material composition data Cm are related to pollution risks, production will be suspended and materials will be replaced.
[0179] The feedback and optimization module includes a model feedback unit and a model updating unit;
[0180] The model feedback unit is responsible for collecting the fluctuation amplitude ΔRisk, the pollution risk value Risk and the pollution feature set F, and inputting them into the random forest model for retraining, and calculating the pollution error Error between the pollution risk value YRisk predicted by the model and the pollution risk value Risk actually measured;
[0181] The contamination error Error is obtained by the following formula:
[0182] Error = |YRisk-Risk|;
[0183] The contamination error Error is recorded to form an error set FEr, the long-term performance of the random forest model is evaluated, and compared with the preset error threshold TErr;
[0184] When the contamination error Error>error threshold TErr, the model update is triggered to update the random forest model;
[0185] The model updating unit continuously adjusts the random forest model according to the contamination error Error as a feedback signal;
[0186] The fluctuation range ΔRisk, the pollution risk value Risk and the pollution feature set F are combined and fitted into the data training set Fnew, and then brought into the random forest model. The random forest model is retrained, and the prediction ability of the random forest model is adjusted by adjusting the decision tree, and a new prediction result y(tr, new) is obtained, and the pollution risk value Risk is recalculated;
[0187] The new prediction result y(tr, new) is obtained by the following formula:
[0188] y(tr,new)=RF(Fnew);
[0189] Where RF(Fnew) represents the random forest model trained using the data training set Fnew.
[0190] In this embodiment, the intelligent decision support module of this embodiment uses the detected warning state At and various sensor data, including temperature and humidity data TH, air quality index data PC and material composition data Cm, to make a judgment to determine whether adjustment measures need to be taken. According to the relationship between different pollution source risks and environmental parameters, the intelligent decision support module can automatically judge and provide targeted adjustment measures. This module can accurately determine the root cause of the pollution source, such as the pollution risk caused by abnormal temperature and humidity, or the risk caused by abnormal air quality and material composition. By combining the actual production environment data with the preset threshold for real-time comparison, the system can provide flexible and accurate response strategies, such as adjusting ventilation, suspending production and replacing materials.
[0191] When the warning status At is "yes", the system can identify the association between pollution risk and temperature, humidity or air quality based on the pollution risk value and environmental data. If the temperature and humidity are abnormal and the pollution risk is high, the system will automatically propose adjustment measures, such as improving ventilation; if the air quality does not meet the standard and there are harmful substances in the material composition, the system will recommend suspending production and replacing materials. This mechanism ensures the dynamic adjustment of the production environment and responds to changes that may cause pollution in a timely manner. For example, high temperature and humidity may accelerate the growth of bacteria, or when the humidity and air quality indicators are poor, harmful components in the material composition may be more active, bringing cross-contamination risks. By automatically adjusting the production environment or material sources through the system, food safety problems caused by operational delays can be effectively avoided. This improves the response speed of the system and the accuracy of adjustment measures, which is particularly important in real-time monitoring and rapid response environments.
[0192] This embodiment introduces a pollution error feedback mechanism. By calculating the difference between the pollution risk value YRisk predicted by the model and the actually measured pollution risk value Risk, the system can record the error in real time and determine whether the model needs to be updated. When the pollution error exceeds the set threshold, the system will trigger a model update, adjust the decision tree, and update the algorithm and decision rules; this pollution error feedback mechanism ensures the continuous learning and optimization of the system, and can self-adapt and improve. Each error between the model prediction and the actual data will be recorded and evaluated, so that the deficiencies of the model can be discovered in time and adjusted. Through dynamic optimization, the system can gradually improve its adaptability to different production environments, pollution source types, and data noise, making pollution source risk assessment more accurate and reducing misjudgments due to data bias. Compared with static models, this feedback and optimization mechanism significantly improves the long-term performance and accuracy of the system.
[0193] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent early warning system for food safety incidents based on multi-source data analysis, characterized in that: It includes data acquisition module, data preprocessing module, pollution source risk assessment module, early warning detection and alarm module, intelligent decision support module and feedback and optimization module; The data acquisition module collects raw data from sensors and acquisition devices, including temperature and humidity data TH, air quality index data PC and material composition data Cm, and performs fitting to form a safety data set SPA; The data preprocessing module is responsible for cleaning and normalizing the security data set SPA to obtain a normalized data set SPG; The pollution source risk assessment module uses a random forest model to perform a risk assessment on the normalized data set SPG to obtain a pollution risk value Risk; The early warning detection and alarm module monitors and detects in real time whether abnormal fluctuations occur according to the obtained pollution risk value Risk, and obtains the early warning state At; The intelligent decision support module provides adjustment measures according to the detected warning state At; The feedback and optimization module adjusts the random forest model according to abnormal fluctuations and adjustment measures.
2. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 1 is characterized in that: The data acquisition module includes a data acquisition unit and a data fitting unit; The data acquisition unit collects temperature and humidity data TH, air quality index data PC and material composition data Cm from sensors and acquisition equipment, the sensors include temperature sensors, humidity sensors and air quality sensors; the acquisition equipment includes atomic absorption spectrometers and spectrum analyzers; The temperature and humidity data TH includes temperature data T and humidity data H. The temperature data T is collected by a temperature sensor, and the humidity data H is collected by a humidity sensor. Air quality index data PC includes PM2.5 concentration and CO2 concentration, and is collected by air quality sensors; Material composition data Cm includes chemical composition Cem, bacterial content Cbt and heavy metal content Chm. The heavy metal content Chm is collected by atomic absorption spectrometer, the bacterial content Cbt is collected by microbial detector, and the number of bacteria is calculated by ATP content. The chemical composition Cem is collected by spectrum analyzer. The data fitting unit summarizes the collected temperature and humidity data TH, air quality index data PC and material composition data Cm to form a safety data set SPA.
3. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 1 is characterized in that: The data preprocessing module includes a data cleaning unit and a data normalization unit; The data cleaning unit performs missing value filling and outlier detection on the security data set SPA to obtain a cleaned data set SPC; Missing value imputation involves using interpolation to fill in missing data; Outlier detection uses the Z-Score method to detect outliers in the data and remove them; The data normalization unit converts the cleaned data set SPC into a unified standard range, eliminates the dimensional differences between different data sources, and obtains a normalized data set SPG; The normalized data set SPG is obtained by the following formula: Wherein, SPGd represents the d-th data item in the normalized data set, SPCd represents the d-th data item in the cleaned data set, SPCdmax represents the peak value of the d-th data item in the cleaned data set, and SPCdmin represents the valley value of the d-th data item in the cleaned data set.
4. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 2 is characterized in that: The pollution source risk assessment module includes a feature extraction unit and a random forest model assessment unit; The feature extraction unit extracts features from the normalized data set SPG, including environmental monitoring data features, material composition change trend features, and cross-contamination features, and forms a contamination feature set F; Environmental monitoring data characteristics include temperature change rate ΔT, humidity change rate ΔH, and air quality change rate ΔPC; The temperature change rate ΔT is obtained by the ratio of the difference between the temperature data at time t and the temperature data at time t-1 to the time interval Δt; The humidity change rate ΔH is obtained by the ratio of the difference between the humidity data at time t and the humidity data at time t-1 to the time interval Δt; The air quality change rate ΔPC is obtained by the ratio of the difference between the air quality index data at time t and the air quality index data at time t-1 to the time interval Δt; The characteristics of the changing trends of material composition include the mean value of chemical composition μCem, the fluctuation of heavy metal content σChm and the mean value of bacteria μCbt; The chemical component mean μCem is obtained by extracting the average concentration of the chemical component Cem; The bacterial mean μCbt is obtained by recording the average level of bacterial content: The heavy metal content fluctuation σChm is obtained by the following formula: In the formula, Chmi represents the heavy metal content at the i-th moment, μChm represents the mean value of the heavy metal content, and N represents the total number of collected moments; The cross-contamination feature is obtained by extracting the correlation coefficient between pollutants and using the Pearson correlation coefficient to measure the impact between pollution sources: The cross contamination feature extraction formula is: Wherein, Corr(D1, D2) represents the Pearson correlation coefficient between pollutant D1 and pollutant D2, j represents the data point, m represents the total number of data points, μD1 represents the mean of pollutant D1, μD2 represents the mean of pollutant D2, D1,j represents the jth data point of pollutant D1, and D2,j represents the jth data point of pollutant D2.
5. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 4 is characterized in that: The random forest model evaluation unit trains the pollution source risk assessment model based on the random forest algorithm according to the pollution feature set F, trains multiple decision trees, and makes a final decision by weighted averaging the prediction results of each tree; By using the pollution feature set F to train the random forest model, multiple decision trees are constructed. The training data of each tree is F(a, tr), and the output of the decision tree is y(tr). The pollution risk value Risk is calculated and compared with the preset risk threshold TRi to determine the cross-contamination risk status; F(a, tr) represents the ath feature data in the feature set F, and y(tr) represents the prediction result of the random forest model; The pollution risk value Risk is expressed by the following formula: In the formula, y(a, tr) represents the prediction result of the ath tree, and M represents the total number of trees; Compare the obtained contamination risk value Risk with the risk threshold TRi to determine the cross-contamination risk status; The cross contamination risk status is obtained by matching in the following ways: When the contamination risk value Risk ≥ risk threshold TRi, it indicates that there is a risk of cross contamination.
6. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 5 is characterized in that: The early warning detection and alarm module includes a risk fluctuation detection unit and an alarm judgment unit; The risk fluctuation detection unit monitors the real-time fluctuation of the pollution risk value Risk, detects whether there is an abnormal fluctuation state, analyzes the change trend of the pollution risk value Risk, calculates the fluctuation amplitude ΔRisk, and compares it with the preset fluctuation threshold ΔTR to determine the abnormal fluctuation state; The fluctuation range ΔRisk is obtained by the following formula: ΔRisk=|Risk(t)-Risk(t-1)|; In the formula, Risk(t) represents the pollution risk value at time t, and Risk(t-1) represents the pollution risk value at time t-1; The abnormal fluctuation state is obtained by matching in the following way: When the fluctuation amplitude ΔRisk ≥ the fluctuation threshold ΔTR, it indicates that abnormal fluctuation occurs; When the fluctuation amplitude ΔRisk is less than the fluctuation threshold ΔTR, it indicates that there is no abnormal fluctuation.
7. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 6 is characterized in that: The alarm judgment unit judges the triggering of the warning state At according to the obtained pollution risk value Risk and the fluctuation range ΔRisk; The warning state At is obtained by the following formula: In the formula, yes means triggering an alert, and no means not triggering an alert.
8. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 7 is characterized in that: The intelligent decision support module includes a decision judgment unit and an adjustment measure generation unit; The decision-making unit obtains a judgment result based on the detected warning state At and related parameters, including temperature and humidity data TH, air quality index data PC and material composition data Cm, to determine whether adjustment measures need to be taken; When the warning state At is yse, the temperature and humidity data TH is judged and compared with the preset temperature threshold AT and the preset humidity threshold AH to judge the relationship between the temperature and humidity data TH and the pollution risk; When the pollution risk value Risk ≥ risk threshold TRi, and the temperature data T ≥ temperature threshold AT, and the humidity data H ≥ humidity threshold AH, it means that the pollution risk is related to the temperature and humidity data TH, specifically, the temperature data T is abnormal and the humidity data H is abnormal; When the warning state At is yse, the relationship between the air quality index data PC and the material composition data Cm is judged, and compared with the preset air quality threshold TPC to judge the relationship between the air quality, material composition and pollution risk; When the pollution risk value Risk ≥ risk threshold TRi, and the air quality index data PC > air quality threshold TPC, and harmful components are detected in the material composition data Cm, it means that the pollution risk is related to the air quality index data PC and the material composition data Cm, specifically, the air quality index data PC is abnormal and harmful components are detected.
9. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 8 is characterized in that: The adjustment measure generating unit adjusts the measures according to the judgment result; When the contamination risk is related to the temperature and humidity data TH, the ventilation is adjusted; When the air quality index data PC and material composition data Cm are related to pollution risks, production will be suspended and materials will be replaced.
10. The intelligent early warning system for food safety incidents based on multi-source data analysis according to claim 9, characterized in that: The feedback and optimization module includes a model feedback unit and a model updating unit; The model feedback unit is responsible for collecting the fluctuation amplitude ΔRisk, the pollution risk value Risk and the pollution feature set F, and inputting them into the random forest model for retraining, and calculating the pollution error Error between the pollution risk value YRisk predicted by the model and the pollution risk value Risk actually measured; The contamination error Error is obtained by the following formula: Error = |YRisk-Risk|; The contamination error Error is recorded to form an error set FEr, the long-term performance of the random forest model is evaluated, and compared with the preset error threshold TErr; When the contamination error Error>error threshold TErr, the model update is triggered to update the random forest model; The model updating unit adjusts the random forest model according to the contamination error Error as a feedback signal; The fluctuation range ΔRisk, the pollution risk value Risk and the pollution feature set F are combined and fitted into the data training set Fnew, and then brought into the random forest model. The random forest model is retrained, and the prediction ability of the random forest model is adjusted by adjusting the decision tree, and a new prediction result y(tr, new) is obtained, and the pollution risk value Risk is recalculated; The new prediction result y(tr, new) is obtained by the following formula: y(tr,new)=RF(Fnew); Where RF(Fnew) represents the random forest model trained using the data training set Fnew.
Citation Information
Cited By
Early warning method and system for microbial pollution of cheese
CN120927411A
A method and system for early warning of cheese microbial contamination
CN120927411B