A food safety risk early warning method and system
By collecting and processing multi-source supply chain data, using algorithms such as clustering, cross-verification and multi-dimensional correlation analysis, potential bias data are identified and isolated, triggering factors are extracted, and dynamic risk models are constructed, which solves the problem that the warning results caused by closed-loop bias in the food safety supply chain are inconsistent with the actual risks, and efficient and accurate food safety risk warning is achieved.
Patent Information
- Application Number
- CN202510289279.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-12
AI Technical Summary
In the prior art, the integration of multi-source data in the food safety supply chain may generate a closed-loop bias in data, resulting in the early warning results that are inconsistent with the actual risks, and the potential hidden dangers in the food safety supply chain cannot be discovered in a timely manner.
By collecting multi-source supply chain data, grouping using clustering algorithms, setting a differential comparison threshold to mark anomaly data and isolating potential bias sources. External control data were introduced and supply chain data with significant differences in control were marked by cross-validation algorithms. A multi-dimensional correlation analysis algorithm is used to extract trigger features and assign weights to each trigger feature separately. Based on the triggering elements, the data is processed in batches and time periods, and small batch sampling data is introduced as a reference. Multi-grained analysis algorithm is used to analyze the significant abnormalities of small samples and the pattern deviation of large batches of data to screen local high-risk points. Build a dynamic risk model, adjust the weight of the triggering elements through the iterative update mechanism and output a preliminary warning. Based on the number of abnormal signals in the preliminary warning results and the weight of the triggering elements, determine whether the food safety risk warning is triggered.
It effectively solves the problems of insufficient data integration and abnormal information being ignored or diluted, improves the accuracy and timeliness of food safety risk warnings, and can accurately identify high-risk data points and issue risk warnings in a timely manner.
Smart Images

Figure CN119785566B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of food supervision, and more specifically, to a food safety risk early warning method and system thereof. Background Art
[0002] As the supply chain becomes more digitalized, food companies and regulatory authorities often use the same or compatible data platforms for traceability and early warning. Sometimes, due to the large number of supply chain entities and the large-scale interconnection of data collection systems, "data closed-loop bias" may occur: that is, the internal system forms an "information island" or "data self-consistency" appearance due to systematic deviations in some collection equipment, upload mechanisms or algorithm models, while the real external risks are ignored or diluted.
[0003] In the existing technology, the integration of multi-source data in the food safety supply chain may cause data closed-loop bias, resulting in early warning results that are inconsistent with actual risks and the inability to timely discover potential hidden dangers in the food safety supply chain. Summary of the invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a food safety risk early warning method and system thereof to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A food safety risk early warning method comprises the following steps:
[0007] S1: Collect multi-source supply chain data and group them using clustering algorithms, set difference comparison thresholds to mark abnormal data and isolate potential bias sources;
[0008] S2: Introduce external control data and use a cross-validation algorithm to mark supply chain data with significant control differences;
[0009] S3: Apply multi-dimensional correlation analysis algorithm to the supply chain data with significant comparison differences to extract trigger factors, and assign weights to each trigger factor;
[0010] S4: Based on the triggering factors, the supply chain data with significant control differences are processed by batches and time periods, and small batch sampling data are introduced as a reference;
[0011] S5: Use multi-granularity analysis algorithms to link significant anomalies in small samples with pattern deviations in large batches of data to screen out local high-risk points that may be masked by the overall trend;
[0012] S6: Build a dynamic risk model based on local high-risk points and their triggering factors, adjust the weights of triggering factors through an iterative update mechanism, and output a preliminary warning;
[0013] S7: Based on the number of abnormal signals in the preliminary warning results and the weights of the triggering factors, determine whether a food safety risk warning is triggered.
[0014] In a preferred embodiment, S1 comprises the following steps:
[0015] S101: Collect raw data from all links of the supply chain, including sensor logs, production records, storage records, transportation records, and sales records of the five elements of man-machine-material-method-environment;
[0016] S102: Clean the collected raw data, remove duplicate data and unify the data format;
[0017] S103: using a clustering algorithm to group the cleaned original data, and dividing the cleaned original data into multiple cluster groups;
[0018] S104: Setting a difference comparison threshold and calculating a data difference index based on the data of each cluster group;
[0019] S105: Mark the data whose data difference index exceeds the difference comparison threshold as abnormal, and isolate the potential biased data.
[0020] In a preferred embodiment, S2 comprises the following steps:
[0021] S201: Collect external control data from historical compliance records and third-party test reports;
[0022] S202: Perform data cleaning on the collected external control data, remove duplicate records and unify the data format;
[0023] S203: using a cross-validation algorithm to compare the external control data with the supply chain data processed in step S1;
[0024] S204: Based on the cross-validation results, supply chain data that are significantly different from external control data are marked to form a supply chain data set with significant control differences.
[0025] In a preferred embodiment, S3 comprises the following steps:
[0026] S301: Use the supply chain data with significant control differences as the input data set for multi-dimensional association analysis;
[0027] S302: Process the input data set using a multi-dimensional correlation analysis algorithm, comprehensively count the numerical characteristics of each detection indicator in the supply chain data, and compare the correlation between each detection indicator;
[0028] S303: According to the processing results of the multi-dimensional correlation analysis algorithm, the triggering factors reflecting the abnormal fluctuation trend of the data are identified, and the numerical characteristics and distribution of each triggering factor are recorded;
[0029] S304: Allocate a weight to each extracted triggering factor according to a preset weight allocation rule, and the allocated weight reflects the relative importance of each triggering factor in the risk warning.
[0030] In a preferred embodiment, S4 comprises the following steps:
[0031] S401: Determine supply chain data segmentation standards according to triggering factors, including batch standards and time period standards;
[0032] S402: Segmenting the supply chain data with significant comparison differences according to the segmentation standard to form multiple data segments;
[0033] S403: Collecting small batch sampling data corresponding to each data segment, where the small batch sampling data includes local sampling results and related testing information;
[0034] S404: Introduce small batch sampling data into data segmentation as a reference to ensure that local data characteristics are compared with the overall pattern.
[0035] In a preferred embodiment, S5 comprises the following steps:
[0036] S501: extracting features from the supply chain data with significant comparison differences after segmentation, and calculating the mean, variance and standard deviation of each segmented data;
[0037] S502: Statistically analyze the collected small batch sampling data to calculate the frequency, extreme value and dispersion degree of abnormal fluctuations of local data;
[0038] S503: using a multi-granularity analysis algorithm to respectively calculate a large batch data pattern deviation index and a small batch data anomaly index;
[0039] S504: Compare the small batch data anomaly index with the large batch data pattern deviation index according to the preset linkage rules, and screen and mark local high-risk data points.
[0040] In a preferred embodiment, the small batch data anomaly index is compared with the large batch data pattern deviation index according to the preset linkage rules, and the local high-risk data points are screened and marked, specifically:
[0041] For each data segment, calculate the linkage ratio: ;in, is the linkage ratio, represents the large batch data mode deviation indicator, Represents the small batch data anomaly indicator, To prevent small positive numbers from dividing by zero;
[0042] like Exceeding the preset threshold , then the relevant data points in the data segment are marked as local high-risk data points.
[0043] In a preferred embodiment, S6 comprises the following steps:
[0044] S601: Integrate the selected local high-risk data points and corresponding triggering factors into an input data set of a dynamic risk model;
[0045] S602: construct a dynamic risk model based on the input data set of the dynamic risk model using a statistical analysis method to form an initial risk assessment framework;
[0046] S603: Using an iterative update mechanism to adjust the weights of each trigger factor in the dynamic risk model in real time, and updating the weight value according to the change of risk indicators;
[0047] S604: Generate a preliminary warning signal and record warning data based on the output results of the dynamic risk model.
[0048] In a preferred embodiment, S7 comprises the following steps:
[0049] S701: extracting the number of preliminary warning signals and weight data of each triggering factor from the preliminary warning results output by the dynamic risk model;
[0050] S702: Count the preliminary warning signals in the preliminary warning results, and calculate the cumulative number of preliminary warning signals as a risk reference indicator;
[0051] S703: using a weighted calculation method to multiply the accumulated number of preliminary warning signals by the weight of each triggering factor and then sum them up to obtain a comprehensive risk judgment index;
[0052] S704: Compare the comprehensive risk judgment index with the preset food risk threshold to determine whether the conditions for triggering food safety risk warning are met, and output a warning signal.
[0053] On the other hand, the present invention provides a food safety risk early warning system, including a data collection and grouping module, an external data verification module, a trigger factor extraction module, a data segmentation processing module, a linkage analysis and screening module, a dynamic risk modeling module, and a risk warning determination module;
[0054] Data collection and grouping module: collects multi-source supply chain data and groups them using clustering algorithms, sets difference comparison thresholds to mark abnormal data and isolate potential sources of bias;
[0055] External data verification module: introduce external control data and mark supply chain data with significant control differences through cross-validation algorithms;
[0056] Trigger factor extraction module: Apply multi-dimensional correlation analysis algorithm to the supply chain data with significant comparison differences to extract trigger factors, and assign weights to each trigger factor;
[0057] Data segmentation processing module: Based on trigger factors, the supply chain data with significant comparison differences are segmented by batches and time periods, and small batch sampling data are introduced as a reference;
[0058] Linkage analysis and screening module: Uses multi-granularity analysis algorithms to link the significant anomalies of small samples with the pattern deviations of large batches of data to screen local high-risk points that may be masked by the overall trend;
[0059] Dynamic risk modeling module: builds a dynamic risk model based on local high-risk points and their triggering factors, adjusts the weights of triggering factors through an iterative update mechanism, and outputs preliminary warnings;
[0060] Risk warning determination module: Combined with the number of abnormal signals in the preliminary warning results and the weight of the triggering factors, it determines whether a food safety risk warning is triggered.
[0061] The technical effects and advantages of a food safety risk early warning method and system of the present invention are as follows:
[0062] 1. The present invention provides a risk warning method for the data closed-loop bias problem that may exist in the food safety supply chain, which effectively solves the problem in the prior art that the warning results are inconsistent with the actual risks due to insufficient data integration and the neglect or dilution of abnormal information. By introducing clustering algorithms, difference comparison and abnormal marking mechanisms in the multi-source supply chain data collection and preprocessing stages, it is possible to quickly identify potential systematic biases in the early stages of data analysis, isolate abnormal data, and ensure the accuracy of subsequent analysis data. In addition, the introduction of external control data and the application of cross-validation algorithms enhance the multi-level comparison capabilities of data, and can mark data that is significantly different from external standards, providing a reliable data basis for subsequent association analysis and high-risk point screening. Through multi-dimensional association analysis and fine-grained comparison of small batches of data, the present invention breaks through the limitations of the existing reliance on the average effect of large batches of data, which makes local anomalies difficult to identify, and significantly improves the ability to accurately identify high-risk data.
[0063] 2. The dynamic risk model of the present invention has adaptive capabilities by adjusting the weights of triggering factors in real time, so that the model can be updated according to the dynamic changes in risks at different stages of the supply chain, thereby improving the accuracy and timeliness of early warning. Through the iterative update mechanism, the model can continuously optimize the weight distribution of each triggering factor according to new risk information, avoiding the problems of fixed weights and poor model adaptability in traditional risk warnings, and enhancing the ability to respond to new risks or systematic deviations. In addition, the comprehensive risk judgment mechanism combines the cumulative number of early warning signals with the weights of triggering factors to achieve accurate quantification of local abnormal events and overall risk levels. When the comprehensive risk exceeds the preset threshold, a risk warning is issued in a timely manner to ensure that enterprises and regulatory agencies can take timely intervention measures before potential risks are upgraded to food safety accidents, thereby ensuring the overall safety of the supply chain. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 A flowchart of a food safety risk early warning method of the present invention;
[0065] Figure 2 It is a structural schematic diagram of a food safety risk early warning system of the present invention. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] Embodiment 1: Figure 1 The present invention provides a food safety risk early warning method, comprising the following steps:
[0068] S1: Collect multi-source supply chain data and group them using clustering algorithms, set difference comparison thresholds to mark abnormal data and isolate potential sources of bias.
[0069] S2: Introduce external control data and use a cross-validation algorithm to mark supply chain data with significant control differences.
[0070] S3: Apply multi-dimensional correlation analysis algorithm to the supply chain data with significant comparison differences to extract trigger factors, and assign weights to each trigger factor respectively.
[0071] S4: Based on the triggering factors, the supply chain data with significant comparison differences are segmented by batches and time periods, and small batch sampling data are introduced as a reference.
[0072] S5: Use a multi-granularity analysis algorithm to conduct a linkage analysis of significant anomalies in small samples and pattern deviations in large batches of data to screen out local high-risk points that may be masked by the overall trend.
[0073] S6: Build a dynamic risk model based on local high-risk points and their triggering factors, adjust the weights of triggering factors through an iterative update mechanism, and output a preliminary warning.
[0074] S7: Based on the number of abnormal signals in the preliminary warning results and the weights of the triggering factors, determine whether a food safety risk warning is triggered.
[0075] S1 includes the following steps:
[0076] S101: Collect raw data from all links of the supply chain, including sensor logs, production records, storage records, transportation records, and sales records of the five elements of man-machine-material-method-environment;
[0077] Obtain unprocessed raw data from various links such as food production, storage, transportation and sales. Raw data includes but is not limited to real-time log data from temperature and humidity sensors, operation records during the production process, and vehicle positioning and environmental monitoring data during transportation. Here, "raw data" refers to the unprocessed data directly output by the data acquisition device. Its purpose is to ensure the comprehensiveness and diversity of data sources so that the integrity and reliability of the data can be processed and analyzed later.
[0078] S102: Clean the collected raw data to remove duplicate data and unify the data format.
[0079] The raw data collected by S101 is preprocessed, including data deduplication, noise data removal and format standardization. Data deduplication refers to filtering the same information collected repeatedly; unified data format requires converting data from different links and sources into a unified structured format to ensure that each data item has consistent semantics and format during subsequent processing, so as to facilitate subsequent algorithm processing. The data obtained after data cleaning here still retains the characteristics of the original information, but the structure is more standardized and convenient for calculation.
[0080] S103: using a clustering algorithm to group the cleaned original data, and dividing the cleaned original data into a plurality of cluster groups.
[0081] The data processed by S102 is automatically grouped using a clustering algorithm, and the data is divided into multiple cluster groups based on the similarity between the data items. The clustering algorithm can use a K-means algorithm, a hierarchical clustering algorithm, or other clustering methods suitable for big data analysis. Here, "cluster group" means a group of data that has similar statistical characteristics or behavior patterns, with the purpose of providing a basis for the subsequent calculation of the data differences within each group, thereby providing a grouping basis for the identification of abnormal data.
[0082] S104: Setting a difference comparison threshold, and calculating a data difference index based on the data of each cluster group.
[0083] According to the data of each cluster group obtained in S103, a predetermined difference comparison threshold is set, and statistical analysis is performed on the data in each cluster group to calculate the data difference index of each group of data. The difference comparison threshold is a value determined based on historical data statistical results or preset standards, and is used to determine whether there is abnormal fluctuation in the data in the same cluster group.
[0084] The calculation method of data difference index involves the mean, variance and other statistics of the data. For example, the calculation method of data difference index is to calculate the standard deviation of the data within the cluster group. The standard deviation is used to measure the degree of dispersion of the data item relative to the mean. The specific process includes calculating the deviation between each data item and the mean within the group, and then finding the square mean of all deviations and taking the square root to obtain the standard deviation value. The larger the standard deviation, the more discrete the data distribution. If the standard deviation exceeds the preset difference comparison threshold, the group of data is abnormal.
[0085] S105: Mark the data whose data difference index exceeds the difference comparison threshold as abnormal, and isolate the potential biased data.
[0086] According to the data difference index calculated in S104, all data in each cluster group are compared. Any data whose data difference index exceeds the difference comparison threshold is marked as abnormal data and isolated from the normal data set. Here, "abnormal data" refers to abnormal data that may be caused by collection errors, equipment deviations, or upload mechanism deviations. At the same time, these abnormal data may cause subsequent risk warning results to deviate from actual risks. By isolating abnormal data, the subsequent data analysis process can be effectively prevented from being interfered by potential biased data, thereby providing a reliable data basis for subsequent external comparison and association analysis.
[0087] S2 includes the following steps:
[0088] S201: Collect external control data from historical compliance records and third-party test reports to ensure that the collected external control data covers all relevant testing links.
[0089] The sources of external control data collected include historical compliance records and third-party test reports, specifically quality test data related to production, transportation, storage and other links in the supply chain. Historical compliance records can include quality test reports from within the company or regulatory authorities, and third-party test reports can be provided by certification bodies or external laboratories. External control data covers key indicators such as temperature, humidity, and chemical composition test results, ensuring that the data source is comprehensive and comprehensive, and can be effectively compared with data from all links in the supply chain.
[0090] S202: Perform data cleaning on the collected external control data, remove duplicate records and unify the data format to form an external control data set.
[0091] Data cleaning is performed on the collected external control data, duplicate records are filtered, and the external control data is converted according to a unified data format to form an external control data set with a unified format and standard content. During the data cleaning process, missing data is marked and processed to ensure data consistency during subsequent algorithm processing.
[0092] S203: Use a cross-validation algorithm to compare the external control data with the supply chain data processed in step S1 to ensure the consistency of the data structure and content of both parties during the comparison process.
[0093] The cross-validation algorithm is used to compare the cleaned external control data with the supply chain data processed in step S1 item by item. The core calculation of cross-validation is: ;in, is the comparison difference value; Corresponding items for supply chain data; Corresponding items for external control data; Used to normalize difference values to ensure that data of different dimensions are comparable.
[0094] During the comparison process, the comparison difference value is calculated item by item and it is determined whether it exceeds the set threshold. ,if , then the data is considered to have significant differences and proceed to the next step.
[0095] Threshold It is determined based on the statistical results of historical data, the mean and standard deviation of various detection indicators, through regression analysis and risk level classification, and with reference to actual test data to ensure accurate differentiation ability, so as to ensure high-precision and timely risk identification.
[0096] S204: Based on the cross-validation results, supply chain data that are significantly different from external control data are marked to form a supply chain data set with significant control differences.
[0097] According to the comparison difference value calculated in S203, all The data items are marked as supply chain data with significant control differences. Supply chain data with significant control differences include supply chain data that may have significant deviations from external standards in terms of test results, environmental conditions, or recording methods. These supply chain data with significant control differences are classified into the supply chain data set with significant control differences for subsequent multi-dimensional correlation analysis. During the marking process, the corresponding data sources, difference indicators, and comparison results will be retained to facilitate subsequent traceability and analysis.
[0098] S3 includes the following steps:
[0099] S301: Use the supply chain data with significant comparison differences as the input data set for multi-dimensional association analysis.
[0100] The supply chain data with significant control differences are used as the input data set, which includes various testing indicator data, such as environmental monitoring data, production parameter data and key testing data during transportation, to ensure that the input data set is comprehensive and representative, providing a reliable data basis for subsequent correlation analysis.
[0101] S302: Use a multi-dimensional correlation analysis algorithm to process the input data set, comprehensively count the numerical characteristics of each detection indicator in the supply chain data, and compare the correlation between each detection indicator.
[0102] By statistically analyzing the numerical characteristics of each detection indicator in the input data set, we can obtain the statistical parameters such as the mean and variance of each indicator data, and use the correlation analysis formula to calculate. The specific formula is as follows: ;in, Indicates The value of the first detection indicator in the data item; Indicates the arithmetic mean of the detection index; Indicates The value of the second detection indicator in the data item; Indicates the arithmetic mean of the detection index; Indicates the total number of data items; Represents the calculated correlation coefficient. This formula is used to quantitatively compare the linear correlation between various test indicators to ensure that the relationship between data is accurately reflected in the multi-dimensional data statistics process.
[0103] S303: According to the processing results of the multi-dimensional correlation analysis algorithm, the triggering factors reflecting the abnormal fluctuation trend of the data are identified, and the numerical characteristics and distribution corresponding to each triggering factor are recorded.
[0104] Based on the correlation analysis results obtained in S302, data items with significant abnormal fluctuations in a certain dimension are selected as trigger factors. For example, in the dimension of ambient temperature monitoring, when the temperature values of some data items in a batch of data exceed two standard deviations of the average temperature of the batch of data and show obvious fluctuations, these data items are selected as trigger factors to ensure that abnormalities are identified in a timely manner.
[0105] For each trigger factor, its numerical characteristics are recorded in detail, including the mean, dispersion and distribution range of each detection indicator, to ensure that the information of each trigger factor is fully disclosed, thereby providing accurate data support for subsequent risk warnings.
[0106] S304: Allocate a weight to each extracted triggering factor according to a preset weight allocation rule, and the allocated weight reflects the relative importance of each triggering factor in the risk warning.
[0107] According to the pre-set weight allocation rules, the triggering factors extracted in S303 are weighted one by one. The weight allocation rules are determined based on the statistical results of historical data and the risk assessment standards formulated by experts. The weight value of each triggering factor reflects its impact on the overall supply chain risk warning. The specific values, calculation basis and statistical parameters of the weight allocation are recorded to ensure that the allocation process is fully open.
[0108] For example, if a trigger factor has a significant correlation among various detection indicators, it will be given a higher weight; otherwise, it will be given a lower weight. This rule ensures that the assigned weights can objectively reflect the influence of each trigger factor on the overall risk warning model and provide a quantitative basis for the subsequent risk model construction.
[0109] S4 includes the following steps:
[0110] S401: Determine supply chain data segmentation standards based on triggering factors, including batch standards and time period standards.
[0111] According to the triggering factors extracted in the previous steps, by counting the collection time and batch information of the supply chain data, the data segmentation standards are established, including batch division standards and time period division standards. The batch division standards are determined based on information such as production batch numbers and transportation batches; the time period division standards are based on the specific time of data collection, such as daily, weekly or monthly divisions. The data segmentation standards ensure that the data within each segment is consistent and comparable, providing an objective basis for subsequent local anomaly detection.
[0112] S402: Segment the supply chain data with significant comparison differences according to the segmentation standard to form multiple data segments.
[0113] According to the batch and time period standards determined in S401, the supply chain data with significant control differences obtained in step S2 are segmented and divided into several data segments. Each data segment corresponds to a certain production batch or collection time interval, which can reflect the statistical characteristics of the detection indicators in the period or batch, and lay a data foundation for subsequent analysis.
[0114] S403: Collect small batch sampling data corresponding to each data segment, where the small batch sampling data includes local sampling results and related testing information.
[0115] For each data segment, the corresponding small batch sampling data is collected from the on-site sampling records. The small batch sampling data is obtained through on-site random sampling or a predetermined sampling plan. Its content includes local test results and environmental monitoring information, which can truly reflect the characteristics and fluctuations of local data in the supply chain, thereby supplementing the local anomalies that may be missed in large batches of data due to the average effect.
[0116] S404: Introduce small batch sampling data into data segmentation as a reference to ensure that local data characteristics are compared with the overall pattern.
[0117] The small batch sampling data collected in S403 is compared with each data segment processed in S402 to establish a reference relationship between local data and overall data. This reference relationship helps to find local high-risk points in the overall data that may be masked by the average effect through abnormal fluctuations in local data, thereby providing more accurate input data for subsequent risk warnings.
[0118] S5 includes the following steps:
[0119] S501: Extract features from the supply chain data with significant comparison differences after segmentation processing, and calculate the mean, variance and standard deviation of each segmented data.
[0120] The specific operations include: calculating the arithmetic mean (mean), data dispersion index (variance) and data fluctuation quantitative index (standard deviation) of each detection index in each data segment. For example, setting the variable Indicates the arithmetic mean of a detection index in a data segment, variable Represents the variance of the detection index, variable It represents the standard deviation of the detection index, where all parameters are obtained through detailed statistical analysis. The above statistical results reflect the overall trend and data fluctuation of large batches of data in each segment, providing a quantitative basis for subsequent local anomaly detection.
[0121] S502: Perform statistical analysis on the collected small batch sampling data to calculate the frequency, extreme value and dispersion degree of abnormal fluctuations in local data.
[0122] For the small batch data obtained from the on-site sampling, a separate statistical analysis is performed. The main operations include: calculating the abnormal fluctuation frequency of the detection index in the local data, determining the extreme value in the local data (i.e. the difference between the local highest value and the lowest value), and quantifying the degree of dispersion of the local data. To ensure the accuracy of the description, the variables are set Indicates the frequency of abnormal fluctuations in local data, variable Indicates local data extremes, variables Indicates the degree of dispersion of local data. Through the above statistical methods, it is possible to identify which data items in the local data show obvious abnormal fluctuations compared with the overall data, providing local characteristic parameters for subsequent multi-granularity analysis.
[0123] S503: Calculate the large batch data pattern deviation index and the small batch data anomaly index respectively using a multi-granularity analysis algorithm.
[0124] This sub-step uses a multi-granularity analysis algorithm to quantitatively calculate the statistical features of the large batch data obtained in S501 and the abnormal features of the small batch data obtained in S502. Specifically, for large batch data, its pattern deviation index is calculated to reflect the overall fluctuation trend of the data in each segment; and for small batch data, the abnormal index is calculated to quantitatively describe the degree of abnormal fluctuation of local data. To this end, set the variable Represents a large batch of data mode deviation indicators, variables Represents the abnormal index of small batch data. A unified statistical formula is used in the calculation process to ensure that the two sets of indicators are comparable in value. For example, the following formula can be used for quantitative calculation: ; ; The "+1" in the formula is used to prevent the denominator from being zero. All variables have been defined in the previous sub-steps to ensure the accuracy and consistency of the calculation results.
[0125] S504: Compare the small batch data anomaly index with the large batch data pattern deviation index according to the preset linkage rules, and screen and mark local high-risk data points.
[0126] According to the preset linkage rules, the small batch data anomaly index calculated in S503 is Deviation indicators from bulk data patterns The preset linkage rules are determined by historical statistical data and risk assessment standards. The rules require that when local abnormal indicators Exceeding the large batch mode deviation index When the linkage ratio is greater than a certain ratio (e.g. 1.5 times or other specific values), the local data is considered to be significantly abnormal and thus marked as a local high-risk data point. The specific operation is: for each data segment, calculate the linkage ratio ,in ;in, To prevent small positive numbers from dividing by zero; if Exceeding the preset threshold (e.g. 1.5), the relevant data points in the data segment are marked as local high-risk data points, and their detailed statistical information is recorded, including the segment to which they belong, key detection indicators and linkage ratios. Etc. This marking process ensures that the screening of local high-risk data points is fully disclosed and provides reliable data input for the construction of subsequent risk warning models.
[0127] The preset threshold It is determined by analyzing typical risk cases in historical data, the deviation distribution of various detection indicators and multiple measured data results, ensuring that normal fluctuations and significant anomalies can be effectively distinguished, while having high adaptability and recognition accuracy.
[0128] S6 includes the following steps:
[0129] S601: Integrate the selected local high-risk data points and corresponding triggering factors into an input data set of the dynamic risk model.
[0130] In this sub-step, all relevant information is integrated into a complete data input set based on the local high-risk data points and corresponding trigger factors obtained in the previous step. The input data set includes the detection index value, abnormal fluctuation parameters, and the corresponding trigger factor identification and preliminary weight information for each local high-risk data point. All data are accompanied by time stamps and batch information to ensure data consistency and traceability, providing a comprehensive basis for the subsequent construction of a dynamic risk model.
[0131] S602: Use statistical analysis methods to construct a dynamic risk model based on the input data set of the dynamic risk model to form an initial risk assessment framework.
[0132] The input data set obtained in S601 is processed by statistical analysis method to construct an initial dynamic risk model. The specific operation includes weighted average calculation of the risk indicators of each local data to obtain a comprehensive risk score. This process can be quantitatively described by the following formula: ;in, represents the comprehensive risk score, i.e., the overall risk score output by the dynamic risk model; Indicates The weight of each trigger factor in the dynamic risk model, whose value reflects the importance of the trigger factor to the overall risk assessment; Indicates The risk indicator value corresponding to each trigger factor reflects the risk level of the trigger factor; Represents the total number of trigger factors, that is, the number of trigger factors participating in the calculation of the dynamic risk model.
[0133] For example, in food cold chain management, a trigger factor is abnormal temperature, and its corresponding risk index value can be expressed by the deviation between the actual measured temperature and the standard temperature. Assuming the standard temperature is 4°C and the actual temperature is 7°C, the risk index value is 3°C. Similarly, in the transportation process, if the trigger factor is transportation delay, its risk index value can be calculated by the difference between the actual transportation time and the scheduled time. The larger the difference, the higher the risk.
[0134] Each parameter is determined based on the statistical results of the input data to ensure that the dynamic risk model can truly reflect the comprehensive risk level of each local high-risk point and constitute an initial risk assessment framework.
[0135] S603: Use an iterative update mechanism to adjust the weights of each trigger factor in the dynamic risk model in real time, and update the weight value according to changes in risk indicators.
[0136] The iterative update mechanism is introduced to adjust the weights of each trigger factor in the dynamic risk model in real time. The specific method is that in each update cycle, the weights of each trigger factor are corrected according to the deviation between the current risk indicator and the reference risk level. The update process can be expressed as follows: ;in, Indicates the latest weight of the first trigger element after the current update cycle; Indicates the weight of the trigger factor in the previous update cycle; is the learning rate parameter, whose value is determined by historical data statistics; Indicates the actual risk indicator value of the trigger factor in the current period; Indicates the predetermined reference risk indicator value.
[0137] Through multiple iterations, the dynamic risk model can continuously adapt to changes in supply chain risks and continuously adjust the relative importance of each trigger factor.
[0138] S604: Generate a preliminary warning signal and record warning data based on the output results of the dynamic risk model.
[0139] Based on the comprehensive risk score output by the dynamic risk model obtained in S602 and S603, the score is compared with the preset risk threshold. If the comprehensive risk score exceeds the preset risk threshold, a preliminary warning signal (i.e., an abnormal signal) is generated. The preliminary warning signal includes the comprehensive risk score, the current weights of the relevant triggering factors, and the data segmentation information. All warning data are systematically recorded to facilitate subsequent risk monitoring and verification. The recorded data includes the weight change record after each iterative update, the time series of the risk score, and the detailed description of the triggering high risk points, so as to ensure that the entire warning process is fully open and provide data support for subsequent dynamic risk management.
[0140] Among them, the preset risk threshold is set based on historical risk event data, the risk score distribution of each triggering factor and expert experience. The critical value that can effectively distinguish normal data from high-risk data is determined through statistical analysis, ensuring high risk identification accuracy and adaptability in different scenarios.
[0141] S7 includes the following steps:
[0142] S701: Extract the number of preliminary warning signals and the weight data of each triggering factor from the preliminary warning results output by the dynamic risk model.
[0143] The number of preliminary warning signals refers to the number of events that trigger warnings in different data segments, and the weight data of the triggering factors reflects the relative importance of each triggering factor in the risk assessment. To ensure data consistency, this step also retains the time stamp, batch information, and detection indicators of each data point, thereby forming a complete risk assessment input information.
[0144] S702: Count the preliminary warning signals in the preliminary warning results and calculate the cumulative number of preliminary warning signals as a risk reference indicator.
[0145] Perform statistical analysis on the number of preliminary warning signals extracted in S701, and calculate the total number of warning signals that appear in a predetermined detection cycle or each data segment. The cumulative number of preliminary warning signals is used as a risk reference indicator to quantify the frequency of local risk events in the entire supply chain. For example, if the number of warning signals detected in a certain data segment is 5, the cumulative number of warning signals in the segment is recorded as 5; if warning signals appear in multiple consecutive data segments, the number of warning signals in each segment is accumulated to form a basis for global risk assessment.
[0146] S703: A weighted calculation method is used to multiply the cumulative number of preliminary warning signals by the weight of each triggering factor and then sum them up to obtain a comprehensive risk judgment index.
[0147] According to the cumulative number of preliminary warning signals obtained by statistics in S702 and the weight data of each triggering factor recorded in S701, the weighted operation method is used for calculation. The specific operation is to multiply the cumulative number of preliminary warning signals corresponding to each triggering factor by its pre-set weight, and then accumulate all the product results to obtain a comprehensive risk judgment index reflecting the overall risk level. The comprehensive risk judgment index can quantitatively describe the contribution of local abnormal events to the overall risk status. For example, if the weight of a triggering factor is 0.4 and the cumulative number of preliminary warning signals in the segment is 10 times, the risk score contributed by the triggering factor is 4 points; after the risk scores of all triggering factors are accumulated, the comprehensive risk judgment index is obtained, which is used as the basis for the next risk judgment.
[0148] S704: Compare the comprehensive risk judgment index with the preset food risk threshold to determine whether the conditions for triggering food safety risk warning are met, and output a warning signal.
[0149] The comprehensive risk judgment index calculated in S703 is strictly compared with the preset food risk threshold. The preset food risk threshold is a critical value determined by historical data statistics and expert experience to distinguish normal data fluctuations from abnormally high risk states.
[0150] When the comprehensive risk judgment index exceeds the preset food risk threshold, it is determined that the risk event in the current supply chain has reached the conditions for triggering a food safety risk warning. At this time, the system will generate a corresponding warning signal and automatically record the detailed information of the warning signal, including the value of the comprehensive risk judgment index, the weight data of each triggering factor, and the cumulative number of warning signals. The output warning signal will serve as an important basis for supply chain risk management, prompting relevant responsible personnel to take corresponding risk control measures in a timely manner to ensure that food safety risks are effectively monitored and managed.
[0151] It is worth noting that the method of this embodiment is applicable to food categories such as liquor, aquatic products, vegetable products, and beverages, and can achieve real-time monitoring and accurate identification of food safety risks, providing data support for internal self-inspection and risk verification of enterprises. The entire embodiment achieves continuous improvement in risk warning and control through data collection, indicator extraction, verification feedback, and model updating, which can effectively narrow the scope of risk points and reduce potential food safety hazards, meeting the technical requirements for full-process management of food safety.
[0152] Example 2: The difference between Example 2 of the present invention and Example 1 is that this example introduces a food safety risk early warning system.
[0153] Figure 2A structural schematic diagram of a food safety risk early warning system of the present invention is given, which includes a data collection and grouping module, an external data verification module, a trigger factor extraction module, a data segmentation processing module, a linkage analysis and screening module, a dynamic risk modeling module and a risk warning determination module.
[0154] Data collection and grouping module: collects multi-source supply chain data and groups them using clustering algorithms, sets difference comparison thresholds to mark abnormal data and isolate potential sources of bias.
[0155] External data verification module: Introduce external control data and mark supply chain data with significant control differences through cross-validation algorithms.
[0156] Trigger factor extraction module: Apply multi-dimensional correlation analysis algorithm to the supply chain data with significant comparison differences to extract trigger factors, and assign weights to each trigger factor respectively.
[0157] Data segmentation processing module: Based on trigger factors, supply chain data with significant comparison differences are segmented by batches and time periods, and small batch sampling data are introduced as a reference.
[0158] Linkage analysis and screening module: Uses a multi-granularity analysis algorithm to conduct linkage analysis on significant anomalies of small samples and pattern deviations of large batches of data, screening out local high-risk points that may be masked by the overall trend.
[0159] Dynamic risk modeling module: Build a dynamic risk model based on local high-risk points and their triggering factors, adjust the weights of triggering factors through an iterative update mechanism and output preliminary warnings.
[0160] Risk warning determination module: Combined with the number of abnormal signals in the preliminary warning results and the weight of the triggering factors, it determines whether a food safety risk warning is triggered.
[0161] The above formulas are all dimensionless and numerical calculations. The formula is a formula that is closest to the actual situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.
[0162] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or may be transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state hard disk.
[0163] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0164] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0165] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0166] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0168] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0169] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0170] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A food safety risk early warning method, characterized in that: The steps include: S1: Collect multi-source supply chain data and group them using clustering algorithms, set difference comparison thresholds to mark abnormal data and isolate potential bias sources; S2: Introduce external control data and use a cross-validation algorithm to mark supply chain data with significant control differences; S3: Apply multi-dimensional correlation analysis algorithm to the supply chain data with significant comparison differences to extract trigger factors, and assign weights to each trigger factor; S4: Based on the triggering factors, the supply chain data with significant control differences are processed by batches and time periods, and small batch sampling data are introduced as a reference; S5: Use multi-granularity analysis algorithms to link significant anomalies in small samples with pattern deviations in large batches of data to screen out local high-risk points that may be masked by the overall trend; S6: Build a dynamic risk model based on local high-risk points and their triggering factors, adjust the weights of triggering factors through an iterative update mechanism, and output a preliminary warning; S7: Based on the number of abnormal signals in the preliminary warning results and the weights of the triggering factors, determine whether a food safety risk warning is triggered.
2. A food safety risk early warning method according to claim 1, characterized in that: S1 includes the following steps: S101: Collect raw data from all links of the supply chain, including sensor logs, production records, storage records, transportation records, and sales records of the five elements of man-machine-material-method-environment; S102: Clean the collected raw data, remove duplicate data and unify the data format; S103: using a clustering algorithm to group the cleaned original data, and dividing the cleaned original data into multiple cluster groups; S104: Setting a difference comparison threshold and calculating a data difference index based on the data of each cluster group; S105: Mark the data whose data difference index exceeds the difference comparison threshold as abnormal, and isolate the potential biased data.
3. A food safety risk early warning method according to claim 1, characterized in that: S2 includes the following steps: S201: Collect external control data from historical compliance records and third-party test reports; S202: Perform data cleaning on the collected external control data, remove duplicate records and unify the data format; S203: using a cross-validation algorithm to compare the external control data with the supply chain data processed in step S1; S204: Based on the cross-validation results, supply chain data that are significantly different from external control data are marked to form a supply chain data set with significant control differences.
4. A food safety risk early warning method according to claim 1, characterized in that: S3 includes the following steps: S301: Use the supply chain data with significant control differences as the input data set for multi-dimensional association analysis; S302: Process the input data set using a multi-dimensional correlation analysis algorithm, comprehensively count the numerical characteristics of each detection indicator in the supply chain data, and compare the correlation between each detection indicator; S303: According to the processing results of the multi-dimensional correlation analysis algorithm, the triggering factors reflecting the abnormal fluctuation trend of the data are identified, and the numerical characteristics and distribution of each triggering factor are recorded; S304: Allocate a weight to each extracted triggering factor according to a preset weight allocation rule, and the allocated weight reflects the relative importance of each triggering factor in the risk warning.
5. A food safety risk early warning method according to claim 1, characterized in that: S4 includes the following steps: S401: Determine supply chain data segmentation standards according to triggering factors, including batch standards and time period standards; S402: Segmenting the supply chain data with significant comparison differences according to the segmentation standard to form multiple data segments; S403: Collecting small batch sampling data corresponding to each data segment, where the small batch sampling data includes local sampling results and related testing information; S404: Introduce small batch sampling data into data segmentation as a reference to ensure that local data characteristics are compared with the overall pattern.
6. A food safety risk early warning method according to claim 1, characterized in that: S5 includes the following steps: S501: extracting features from the supply chain data with significant comparison differences after segmentation, and calculating the mean, variance and standard deviation of each segmented data; S502: Statistically analyze the collected small batch sampling data to calculate the frequency, extreme value and dispersion degree of abnormal fluctuations of local data; S503: using a multi-granularity analysis algorithm to respectively calculate a large batch data pattern deviation index and a small batch data anomaly index; S504: Compare the small batch data anomaly index with the large batch data pattern deviation index according to the preset linkage rules, and screen and mark local high-risk data points.
7. A food safety risk early warning method according to claim 6, characterized in that: Compare the small batch data anomaly indicators with the large batch data pattern deviation indicators according to the preset linkage rules, and filter and mark local high-risk data points, specifically: For each data segment, calculate the linkage ratio: ;in, is the linkage ratio, represents the large batch data mode deviation indicator, Represents the small batch data anomaly indicator, To prevent small positive numbers from dividing by zero; like Exceeding the preset threshold , then the relevant data points in the data segment are marked as local high-risk data points.
8. A food safety risk early warning method according to claim 1, characterized in that: S6 includes the following steps: S601: Integrate the selected local high-risk data points and corresponding triggering factors into an input data set of a dynamic risk model; S602: construct a dynamic risk model based on the input data set of the dynamic risk model using a statistical analysis method to form an initial risk assessment framework; S603: Using an iterative update mechanism to adjust the weights of each trigger factor in the dynamic risk model in real time, and updating the weight value according to the change of risk indicators; S604: Generate a preliminary warning signal and record warning data based on the output results of the dynamic risk model.
9. A food safety risk early warning method according to claim 1, characterized in that: S7 includes the following steps: S701: extracting the number of preliminary warning signals and weight data of each triggering factor from the preliminary warning results output by the dynamic risk model; S702: Count the preliminary warning signals in the preliminary warning results, and calculate the cumulative number of preliminary warning signals as a risk reference indicator; S703: using a weighted calculation method to multiply the accumulated number of preliminary warning signals by the weight of each triggering factor and then sum them up to obtain a comprehensive risk judgment index; S704: Compare the comprehensive risk judgment index with the preset food risk threshold to determine whether the conditions for triggering food safety risk warning are met, and output a warning signal.
10. A food safety risk early warning system, used to implement a food safety risk early warning method according to any one of claims 1 to 9, characterized in that: It includes data collection and grouping module, external data verification module, trigger factor extraction module, data segmentation processing module, linkage analysis and screening module, dynamic risk modeling module and risk warning determination module; Data collection and grouping module: collects multi-source supply chain data and groups them using clustering algorithms, sets difference comparison thresholds to mark abnormal data and isolate potential sources of bias; External data verification module: introduce external control data and mark supply chain data with significant control differences through cross-validation algorithms; Trigger factor extraction module: Apply multi-dimensional correlation analysis algorithm to the supply chain data with significant comparison differences to extract trigger factors, and assign weights to each trigger factor; Data segmentation processing module: Based on trigger factors, the supply chain data with significant comparison differences are segmented by batches and time periods, and small batch sampling data are introduced as a reference; Linkage analysis and screening module: Uses multi-granularity analysis algorithms to link the significant anomalies of small samples with the pattern deviations of large batches of data to screen local high-risk points that may be masked by the overall trend; Dynamic risk modeling module: builds a dynamic risk model based on local high-risk points and their triggering factors, adjusts the weights of triggering factors through an iterative update mechanism, and outputs preliminary warnings; Risk warning determination module: Combined with the number of abnormal signals in the preliminary warning results and the weight of the triggering factors, it determines whether a food safety risk warning is triggered.
Citation Information
Patent Citations
A Physical and Chemical Data Analysis System for Food Safety Risk Monitoring
AU2020103340A4
TabNet-GRA-based food safety risk prediction method and visual analysis system
CN116933928A