Food safety risk early warning system and method based on multi-source heterogeneous data
By constructing time evolution vectors and knowledge graphs, and combining multi-source data analysis, the early warning strategy is dynamically adjusted, which solves the problem of insufficient time series analysis in food safety data processing of existing systems, and realizes timely identification and accurate assessment of food safety risks.
Patent Information
- Application Number
- CN202511947193.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-24
AI Technical Summary
Existing food safety risk early warning systems lack effective time-series analysis mechanisms when dealing with complex and dynamically changing food safety data. This results in the inability to capture long-term trends and local fluctuations in the data in a timely manner, making it difficult to cope with real-time changes and affecting the accuracy and timeliness of risk assessment.
A food safety risk early warning system based on multi-source heterogeneous data is adopted. Through a time evolution vector construction module, a knowledge graph deep embedding vector generation module, and a food safety risk curve construction module, combined with multi-source data and dynamic knowledge graph, time series analysis and causal relationship mining are carried out to dynamically adjust the risk early warning strategy.
It enables timely identification of potential food safety hazards, improves the accuracy and flexibility of early warning, and can respond promptly to emergencies. By assessing food safety risks from multiple dimensions, it enhances the comprehensiveness and accuracy of predictions.
Smart Images

Figure CN121724432A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data analytics, and in particular to a food safety risk early warning system and method based on multi-source heterogeneous data. Background Technology
[0002] With the accelerating pace of globalization and urbanization, food production, processing, transportation, and consumption are continuously expanding, leading to increasingly complex supply chain structures and food safety issues characterized by high frequency, diversity, and suddenness. Traditional food safety supervision models, relying primarily on manual inspections, on-site sampling, and post-event analysis, suffer from limitations such as poor timeliness, limited coverage, and difficulty in timely detection of potential risks, making them insufficient to meet the demands of modern food supply chains for efficient and accurate risk monitoring. Against this backdrop, food safety risk early warning technology has gradually become a focus of research and application, enabling the early identification and rapid response to food safety hazards through multi-source data collection, real-time monitoring, big data analysis, and intelligent prediction.
[0003] In existing technologies, food safety risk early warning systems collect multi-source data from production, processing, logistics, environmental monitoring, and laboratory testing to achieve dynamic analysis and assessment of food safety incidents. This allows for the identification of potential pollution sources and cross-contamination risks, and intelligent decision support reduces human intervention, improving the timeliness and accuracy of early warnings. Meanwhile, knowledge graph-based risk assessment methods for food safety incidents have emerged. The core of these methods is to input a food safety knowledge graph containing entities, entity relationships, and entity characteristics into a risk assessment model. This model assesses the risk level of target entities, performs cause-and-effect reasoning, and predicts consequences. By extracting multi-level features and performing associative reasoning on entities and their relationships within the knowledge graph, it achieves dynamic assessment and outcome prediction of food safety incidents, thereby supporting food safety management decisions for major events or specific scenarios and improving the accuracy of risk assessment.
[0004] The above-mentioned technology has at least the following technical problems: Existing systems lack effective time-series analysis mechanisms when processing complex and dynamically changing food safety data. This results in the inability to capture long-term trends and local fluctuations in the data in a timely manner, which in turn affects the prediction and monitoring of potential risks. They are unable to cope with real-time changes such as temperature fluctuations and are prone to missing critical early warning opportunities. The resulting insufficient detection of data changes leads to a relatively simple causal relationship analysis of food safety data. There is a lack of in-depth modeling of the mutual influence between different entity nodes, making it difficult to capture abnormal fluctuations and sudden risks in food safety data in real time, which affects the accuracy and timeliness of risk assessment. Summary of the Invention
[0005] On the one hand, a food safety risk early warning system based on multi-source heterogeneous data is provided, which includes: The temporal evolution vector construction module is used to collect and analyze various original food safety datasets, obtain data for each time period of the original food safety data, construct temporal evolution vectors for each time period of the original food safety data, analyze the variance and cosine similarity of the food safety data, and adjust the evolution window length accordingly.
[0006] The knowledge graph deep embedding vector generation module is used to construct a food safety knowledge graph based on time evolution vectors. It inputs the entity nodes and their attributes of the knowledge graph into the time sequence graph reasoning model and analyzes the deep embedding vector of each entity node.
[0007] The food safety risk curve construction module is used to generate food safety risk curves for entity nodes based on the deep embedding vectors of each entity node, and to determine the warning level.
[0008] On the other hand, a food safety risk early warning method based on multi-source heterogeneous data is provided, which includes: We collect and analyze various original food safety datasets to obtain data for each time period. For each time period of the original food safety datasets, we construct the time evolution vectors of the original food safety datasets, analyze the variance and cosine similarity of the food safety data, and adjust the evolution window length accordingly.
[0009] A food safety knowledge graph is constructed based on time evolution vectors. The entity nodes and their attributes of the knowledge graph are input into the time sequence graph reasoning model, and the deep embedding vectors of each entity node are analyzed.
[0010] Based on the deep embedding vectors of each entity node, a food safety risk curve for the entity node is generated to determine the warning level.
[0011] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. The food safety risk early warning system based on multi-source heterogeneous data provided by this invention systematically analyzes food safety risks by combining multi-source data and dynamic knowledge graphs. It can identify potential safety hazards in a timely manner. Through in-depth analysis of time series data and accurate temporal reasoning, it can dynamically adjust risk early warning strategies to avoid the occurrence of sudden food safety problems. Based on historical and real-time data analysis, it can identify potential safety hazards in each link and identify potential food safety problems in a timely manner. Through in-depth analysis of the time evolution of food safety data, it not only considers the timeliness and volatility of the data, but also combines the verification of historical patterns to ensure the accuracy and reliability of the early warning system. Furthermore, by integrating various risk values, it can more comprehensively predict and respond to complex food safety risk scenarios.
[0012] 2. This invention can form a clear causal chain through the construction of a knowledge graph, realize the mining of causal relationships, make the relationship between various food safety data more transparent, solve the problem of unclear causal relationships between different datasets, and Granger causality test discovers potential causal associations between data. This not only increases the accuracy of early warning, but also provides structured information for subsequent risk assessment and improves the accuracy of prediction.
[0013] 3. This invention comprehensively analyzes the historical and real-time data of each entity node to dynamically assess food safety risks, solving the problem that traditional static early warning cannot cope with rapid changes. It calculates the time-series risk curve based on the data evolution window and removes short-term fluctuations through exponential smoothing to present a more stable risk curve. Based on this curve, the risk level can be accurately determined and early warning can be issued in a timely manner. By adopting dynamic time-series analysis, the early warning threshold can be automatically adjusted, making it more flexible and accurate in responding to various emergencies.
[0014] 4. This invention integrates risk and mutation risk from multiple dimensions, comprehensively considering the changing characteristics of food safety data from multiple perspectives, including long-term stability changes and sudden mutation risks. It can more comprehensively and accurately assess the overall risk level of food safety, no longer relying solely on a single risk assessment indicator, but comprehensively considering the multiple changing characteristics of the data. This makes the risk assessment results more comprehensive and detailed, and better adaptable to various types of risk changes. Whether it is a stable trend change or a sudden abnormal fluctuation, it can be identified and handled in a timely manner, thereby enhancing the accuracy of early warning. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This application provides a structural diagram of a food safety risk early warning system based on multi-source heterogeneous data, as shown in the embodiments of the present application. Figure 2 A flowchart illustrating the construction of the time evolution vector for a food safety risk early warning system based on multi-source heterogeneous data, provided in this application embodiment; Figure 3 This is a flowchart illustrating the construction of a food safety knowledge graph for a food safety risk early warning system based on multi-source heterogeneous data, as provided in an embodiment of this application. Figure 4 This application provides a flowchart of the steps for generating a deep embedding vector in a food safety risk early warning system based on multi-source heterogeneous data. Figure 5 A flowchart of a food safety risk early warning method based on multi-source heterogeneous data provided in this application embodiment. Detailed Implementation
[0017] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of the present disclosure are shown in the drawings, it should be understood that embodiments of the present disclosure may be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure.
[0018] It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure. In the description of the embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "this embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects.
[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0020] Embodiment 1 of the present invention Figure 1 The diagram shown is a structural diagram of a food safety risk early warning system based on multi-source heterogeneous data provided in this application embodiment. The system includes: a time evolution vector construction module, a knowledge graph deep embedding vector generation module, and a food safety risk curve construction module.
[0021] In this embodiment, the data requirements are clearly defined. Taking food safety in cold chain food transportation as an example, environmental data such as temperature, humidity, and gas concentration are acquired through sensors and monitoring systems. Real-time geographical location and transportation route data are obtained through GPS and vehicle monitoring systems. Waybill information also needs to be extracted from the warehousing and transportation management system to understand the transportation requirements, temperature control standards, sample numbers, and batch numbers of different foods. The product status is monitored through electronic tags, including the food's production date, shelf life, and temperature control history, to ensure the quality of food during transportation.
[0022] Preliminary checks are performed on various types of raw food safety data, including food safety data temperature sets, food safety data humidity sets, and food safety data production batch number sets, to ensure data integrity and consistency. According to the preset standard format in the database, various types of raw food safety data are extracted and mapped to standard fields. Entity node identifiers are usually extracted from the data identifiers (such as sample number and batch number) to identify each data instance. Feature variables include relevant measurement data (such as temperature, humidity, and gas concentration). These values are classified and normalized into numerical data. The original timestamps are parsed according to the time field of the data to ensure consistent time format. Data sources are extracted from the institution name, source field, etc. in the dataset to indicate the data collection source.
[0023] We collect and organize timestamp information from various types of raw food safety datasets, use a clock offset estimation model to compare and analyze the timestamps in each type of raw food safety dataset, and identify clock errors caused by clock drift or time synchronization problems by calculating the time differences between various types of raw food safety datasets.
[0024] Clock skew estimation is a mathematical model used to identify and correct time errors between different data. The input of this model is usually a timestamp sequence from various original food safety datasets, and the output is the clock skew between the data sources. The clock skew estimation model uses the least mean square error method to perform statistical analysis on all time differences and calculate the skew amount, i.e., the clock error, to help the system correct timestamps and ensure the consistency of various original food safety datasets in the time dimension. In this way, the clock skew estimation model provides a reliable foundation for subsequent data synchronization and time series analysis.
[0025] The clock error of each type of original food safety dataset is recorded as the timestamp correction amount of each type of original food safety dataset. The timestamp correction amount of each type of original food safety dataset is added to the timestamp of each type of original food safety dataset to obtain the timestamp update value of each type of original food safety dataset. That is, the timestamps of all datasets are aligned in the time dimension, ensuring that the timestamps of different datasets can be synchronized, forming various aligned food safety datasets that are synchronized in the time dimension.
[0026] Figure 2 The flowchart for constructing the time evolution vector of the food safety risk early warning system based on multi-source heterogeneous data provided in this application embodiment includes: the aligned dataset of each entity node is divided into multiple time period data according to the preset time window parameters, and the time series feature extraction module is called to calculate the data statistical indicators of each time period to obtain a multi-dimensional feature vector representing the time series features. Through the sliding window mechanism, the feature vectors of adjacent time periods are superimposed in time order to obtain the time evolution vector.
[0027] For each type of original food safety data, it is treated as an entity node, and a time evolution vector is constructed based on each aligned food safety dataset under that entity node.
[0028] For each aligned food safety dataset of an entity node, the aligned food safety dataset is divided into data for each time period according to a preset time window parameter.
[0029] The preset logic for time window parameters in the database is set based on the frequency of data change. It typically includes the evolution window length and evolution step size. The evolution window length refers to the time span of data in each time period, determining the amount of historical data included in each analysis. The evolution step size is the interval at which the window moves along the time axis, determining the time distance between each analysis. The preset logic is usually determined based on the data sampling frequency, the required analysis accuracy, and the dynamic characteristics of the time series. The minimum time unit of the parameter is determined based on the sampling frequency, and the core period of the data is identified, ensuring that the window length covers at least one complete period to capture effective patterns. Based on historical data analysis, the evolution window is set... The specific ranges for window length, evolution step size, and analytical precision requirements are determined by transforming these parameters using predefined mapping functions or rules. For example, linear mapping, piecewise functions, or nonlinear mapping can be used. The evolution window length, evolution step size, and analytical precision requirements are input into the mapping function or rule to find or calculate the corresponding evolution window length and evolution step size. Through this mapping logic, for food safety datasets, the evolution window length and evolution step size are set according to the frequency of data changes. The pre-setting of these parameters ensures the consistency of time period divisions and provides an effective data foundation for subsequent time-series feature extraction and evolutionary analysis.
[0030] The aligned food safety dataset is divided into multiple time periods based on the evolution window length, with each time period corresponding to a specific time interval.
[0031] For each time period, the time-series feature extraction module in the database is invoked. This module is a tool or algorithm used to automatically extract meaningful features and patterns from time-series data. Its main function is to process the raw time-series data and extract statistical features that reflect the data's changing patterns. Through this module, the raw data can be transformed into a more informative feature vector. Statistical indicators for that time period are calculated, including the linear trend slope, the maximum and minimum values of each raw food safety data point, and the rate of change. These indicators reflect the basic characteristics and fluctuations of the data within that time period. The mean and standard deviation of each statistical indicator are calculated. Then, based on the Pearson correlation coefficient formula, the linear relationship between two features is evaluated. The Pearson correlation coefficient is obtained by calculating the covariance and standardizing it. The four statistical indicators are arranged in descending order of feature correlation, ensuring that highly correlated features are grouped together to better reflect the dynamic changes and regularities of the data. This results in a multi-dimensional vector containing all the time-series features within that time period, denoted as the feature vector. Each dimension of the feature vector represents a temporal feature, which integrates multiple information from the data within that time period, providing a structured feature representation for subsequent construction of time evolution vectors and time series analysis.
[0032] To fit the optimal straight line for the data within the time window, each sampling time is first recorded as a time point, and its corresponding food safety data value is recorded as a data value. The average of all time points is calculated and recorded as the time mean, and the average of all data values is recorded as the data mean. For each data point, the difference between its time point and the time mean is calculated and recorded as the time deviation value, and the difference between its data value and the data mean is recorded as the data deviation value. The calculation of the linear trend slope is essentially a measure of the degree of coordinated change between these two difference sequences. That is, the sum of the products of the time deviation values and the data deviation values of all data points is divided by the sum of the squares of the differences between all time points and the time mean. The slope of the fitted optimal straight line is then calculated, which gives the rate of change of the data over time, i.e., the slope of the linear trend. The maximum and minimum values are obtained by finding the maximum and minimum values of the original food safety data within the time period, directly obtaining the extreme values of the data fluctuation, thus reflecting the range of change of the data within the time period. The rate of change is obtained by dividing the difference between the maximum and minimum values within the time period by the length of the time period. If the slope of the optimal straight line is positive, the rate of change is positive; if the slope of the optimal straight line is negative, the rate of change is negative.
[0033] For an aligned food safety dataset, an initial window is determined based on the timestamp sequence and the evolution window length, covering several feature vectors within a time period. These feature vectors are then sequentially superimposed to form a new temporal evolution vector, which represents the comprehensive features of the data within that time period. Next, the sliding window moves forward by the evolution step size to cover the next time period.
[0034] In one specific embodiment, suppose we have a time-series dataset on food safety, recording temperature data every minute. To capture the dynamic characteristics of temperature within each time period, we first define a time window (e.g., 30 minutes), and the collected raw temperature sequence is (-18.0, -18.5, -17.0, -15.5, -14.0, -13.5, -12.0). We then extract multiple statistical features from this time period to form a feature vector. Within each time period, we calculate statistical indicators such as the linear trend slope, maximum and minimum values, and rate of change, obtaining a linear trend slope of +0.20℃ / minute, a maximum temperature of -12℃, a minimum temperature of -18.5℃, and a rate of change of 0.22℃ / minute. Based on these statistical results, we construct the feature vector for that time period, i.e., [+0.20, -12, -18.5, 0.22], which contains all the key features within that time period. Then, using a sliding window mechanism, the window is moved forward in 15-minute increments to cover the next time period.
[0035] When constructing the temporal evolution vector, for a time period of data in an aligned food safety dataset, the variance of each food safety data point within that time period is first calculated, which measures the degree of fluctuation of each food safety data point within that time period. Then, the mean variance of each food safety data point is processed to obtain the variance of the food safety data within that time period, representing the overall level of data fluctuation within that time period. Next, the cosine similarity between the feature vector within the current time window and the feature vector of the long-term historical pattern is calculated. This similarity value is used to assess the similarity between the current data and the historical pattern, with a value ranging from -1 to 1; the closer the value is to 1, the more similar it is. Finally, the evolution window length is dynamically adjusted based on the quantification results of the food safety data variance and cosine similarity.
[0036] The feature vectors of long-term historical patterns in a database are usually constructed by analyzing the statistical features of long-term time periods in historical datasets. This involves extracting a preset historical time span from the database, collecting data for that historical time span, extracting multiple time-series features from this data, performing statistical analysis on these features, and obtaining the long-term average value or other aggregated indicators of each feature in the historical data. This forms the feature vectors of the long-term historical pattern. These feature vectors represent the overall change patterns and trends of the historical data and are used as a benchmark pattern for comparison with the data of the current time period.
[0037] The historical time span in the database is preset to be the average time span during which warning information appears in historical food safety data.
[0038] The preset data volatility threshold and cosine similarity threshold in the database are typically set based on a comprehensive analysis of historical data and specific considerations of system requirements. The data volatility threshold is usually set by calculating the standard deviation or variance of historical data, reflecting the data's fluctuation range under normal conditions. By calculating the mean and standard deviation of historical data, a reasonable fluctuation range is determined, usually the mean plus or minus a certain multiple of the standard deviation. The set threshold should reflect the typical fluctuation range of food safety data under normal operating conditions; fluctuations exceeding this range indicate potential anomalies or risks. The cosine similarity threshold is usually set based on the similarity analysis between long-term historical data and current data features. By calculating the cosine similarity between the feature vectors of long-term historical patterns and the feature vectors of real-time data, a similarity threshold is set, representing the degree of consistency between the current data and historical patterns. If the similarity is high, it indicates that the current data is consistent with historical patterns, and the system is functioning normally; if the similarity is low, it may mean that the current data is abnormal or deviates from the normal pattern, requiring further monitoring or intervention. Both thresholds need to be adjusted according to industry standards, the statistical characteristics of historical data, and the needs of actual application scenarios to ensure timely identification and response to potential food safety risks.
[0039] When the variance of food safety data is detected to be greater than or equal to the preset data volatility threshold in the database, it means that the fluctuation of food safety data in that time period has increased abnormally, which may indicate potential risks or unstable factors. Such data with large fluctuations usually indicate that the food safety environment has changed significantly, which may lead to an increase in food safety hazards. This is recorded as the first condition.
[0040] Having a cosine similarity greater than or equal to the preset cosine similarity threshold in the database means that the feature vector of the data in the current time period is highly similar to the feature vector of the long-term historical pattern. This indicates that the trend of the current data is consistent with or close to the historical pattern. This situation usually means that the food safety environment has remained stable in this time period without significant abnormal fluctuations or risks. This is recorded as the second condition.
[0041] If the first and second conditions exist, it means that although the fluctuation pattern conforms to the expected historical pattern, the current data is highly volatile, indicating that although the system has experienced a complete anomaly, there is a certain degree of risk or instability. In this case, further monitoring and analysis are required. Therefore, the data volatility deviation value is obtained by subtracting the data volatility threshold from the variance of the food safety data, and the cosine similarity deviation value is obtained by subtracting the cosine similarity threshold from the cosine similarity value. The evolution window shortening value is determined based on the data volatility deviation value and the cosine similarity deviation value.
[0042] Based on historical data analysis, specific ranges are set for volatility deviation and cosine similarity deviation values. The volatility deviation value, between 0 and 1, represents the difference between the current data fluctuation and historical data, while the cosine similarity deviation value, also between 0 and 1, represents the degree of similarity between the current data and historical patterns. Based on these values, a predefined mapping function or rule is used for transformation, such as linear mapping, piecewise functions, or nonlinear mapping. The volatility deviation and cosine similarity deviation values are input into the mapping function or rule to find or calculate the corresponding evolution window shortening value, which is then applied to the evolution window update. This achieves adaptive control and adjustment of the evolution window length. Through this mapping logic, the database can dynamically adjust the evolution window according to data changes, ensuring a timely and accurate response to changes in current food safety data.
[0043] If only the first condition exists, it means that the data volatility in the current time period has increased abnormally, which may indicate risks or instability. In this case, it is necessary to pay attention to the degree of data volatility and adjust subsequent operations based on this volatility. The evolution window shortening value is then determined based on the data volatility deviation value.
[0044] If only the second condition exists, it means that the data characteristics of the current time period are highly similar to the long-term historical pattern, indicating that the current data change trend is consistent with the historical data and there is no significant abnormal fluctuation. In this case, the current food safety data can be considered to be in a stable state. The evolution window length can be appropriately extended to improve the computational efficiency. The extension value of the evolution window is determined based on the cosine similarity deviation value.
[0045] The mapping logic for determining the evolution window lengthening value based on the cosine similarity deviation value is usually achieved by setting a predefined relationship and calculating the cosine similarity deviation value between the current time period data and the long-term historical pattern. This reflects the difference between the current data and the historical pattern. The specific mapping method is usually implemented through nonlinear mapping or weighting functions. For example, by setting the threshold range and the proportional relationship of the lengthening magnitude, it is ensured that the adjustment of the evolution window can accurately reflect the stability and change requirements of the data under different similarity deviation values.
[0046] If neither of the two conditions exists, it means that although the current data is not very similar to the historical pattern, the data fluctuation is not significant. It can be considered that the current data change is stable and there is no significant risk or abnormal fluctuation. Therefore, the evolution window will not be adjusted, the original evolution window length will be maintained, and data analysis and processing will continue according to the current settings. This situation usually means that the data change trend is gentle and can run stably without much adjustment or intervention.
[0047] The new evolution window value is obtained by adding the shortened or lengthened evolution window value to the evolution window length, and then applied to the data in the next time period.
[0048] like Figure 3 The flowchart for constructing a food safety knowledge graph for a food safety risk early warning system based on multi-source heterogeneous data provided in this application embodiment includes: calling each entity node and filling the calculated time evolution vector into the attributes of each entity node; filtering out entity node pairs that co-occur in the same time period; extracting their historical numerical sequences; using Granger causality test to analyze the causal relationship between entity node pairs; estimating the regression coefficient of the lag period through regression model and least squares method; calculating the significance index; using this to determine whether to create a causal relationship edge for the entity node pair in the knowledge graph; and marking it as an analyzed entity node pair to prevent duplicate analysis. The above steps are repeated to complete the analysis of all entity nodes and their causal relationships, thus constructing a complete food safety knowledge graph.
[0049] S1 calls each entity node and fills the calculated time evolution vector into the entity node attributes of each entity node to construct a food safety knowledge graph.
[0050] Each entity node's attributes include linear trend slope, maximum and minimum values, and rate of change. These attributes are multidimensional and represent various characteristics of the entity node in the time series. As time goes by and the data is continuously updated, the entity node's attributes are dynamically updated based on new time period data and time evolution vectors. Each time an update occurs, the new time evolution vector replaces the old entity node attributes, thereby reflecting the latest food safety data characteristics.
[0051] S2, under a certain entity node, filter out all entity nodes that co-occur in time, arbitrarily select two different entity nodes to generate entity node pairs. These entity node pairs refer to entity nodes that exist in the same time period. Their change trends may be related or mutually influential. For each pair of entity nodes, the system will extract the historical value sequence from the attributes of all continuous entity nodes associated with them. These entity node attributes include the original food safety data in each time period.
[0052] S3. Using the Granger causality test to analyze these two historical numerical sequences can help determine whether one sequence can predict the future changes of the other, thus revealing the causal relationship between them. The lag period is selected using the AIC (Akaike Information Criterion), and then the Granger causality test is performed based on the regression model. The least squares method is used to estimate the regression model. The least squares method finds the best fit to the data by minimizing the sum of squares of the residuals. The regression coefficient for each lag period is calculated, which represents the strength of the influence of the lagged data on the current data. The t-value is obtained by calculating the ratio of the standard error to the coefficient value of each regression coefficient. That is, a t-test is performed on each regression coefficient to test whether the influence of each lag period is significant. The obtained t-value is recorded as the significance index.
[0053] S4. If the significance index is greater than or equal to the preset significance threshold in the database, it means that there is a significant causal relationship between the two entity nodes in statistics, that is, there is a causal dependency, which can be represented and archived in the knowledge graph as the basis for further analysis and decision-making. Then, a new directional causal relationship edge is created between the two entity nodes in the knowledge graph, and a significance index is assigned to the relationship. The relationship is marked as a marked causal entity node pair and removed in the next entity node pair causal analysis.
[0054] In one specific embodiment, assuming that the causal analysis of entity nodes A and B has been completed and a significant causal relationship between A and B has been found with a significance index of 0.03, then an undirected causal relationship edge from A to B is created in the knowledge graph and the significance index of the edge is assigned to 0.03.
[0055] Mark the entity node pair A and B as a marked causal entity node pair and add them to the marked list.
[0056] In subsequent analysis, when the system detects the entity nodes A and B, it will skip them to avoid repeating the causal analysis.
[0057] If the significance index is less than the preset significance threshold in the database, it means that there is no significant causal relationship between the two entity nodes. Skip this pair of entity nodes and continue to analyze the next pair of entity nodes to ensure that causal relationship edges are created only when they are statistically significant.
[0058] S5, repeat S2 to S4, and analyze all causal relationship edges of this entity node.
[0059] S6. Repeat S1 to S5 to analyze all entity nodes and complete the construction of the food safety knowledge graph.
[0060] like Figure 4The diagram shows the steps of generating a deep embedding vector for a food safety risk early warning system based on multi-source heterogeneous data provided in this application embodiment. The steps include: After completing the construction of the food safety knowledge graph, the entity nodes, causal relationship edges, and their attributes in the graph are input into a temporal graph inference model. For each entity node, the temporal graph inference model first performs intra-temporal attention calculation on the food safety data within each evolution window, extracting the most significant local patterns and generating a condensed vector. Then, it performs cross-temporal attention calculation on the food safety data between different evolution windows to obtain a context vector reflecting historical influence and evolutionary background. Finally, based on a preset fusion ratio factor in the database and information from causal relationship edges, it fuses the condensed vector and the context vector to generate a deep embedding vector for each entity node.
[0061] After the food safety knowledge graph is constructed, the entity nodes, causal relationship edges and their attributes in the graph are input into the time series graph inference model. For each entity node, a temporal position code is added to each entity node, including the time segment to which the entity node attribute belongs. The time series graph inference model will perform intra-temporal attention calculation on the food safety data in each evolution window of the entity node. The importance of the data at each time step in the sequence is calculated using a self-attention mechanism. The weight of each time step can be calculated through multiplicative attention. Based on the calculated attention weights, the time series data in the window are weighted and summed to obtain the most significant local pattern in the window.
[0062] The temporal graph inference model is a hybrid artificial intelligence model specifically designed for analyzing and predicting risk transmission in dynamic systems. It integrates time-series data with entity relationship networks for inference. In this model, temporal refers to the continuous monitoring of various risk indicators of food safety entity nodes, forming a time evolution sequence; graph refers to the heterogeneous network constructed based on the physical, logical, or business relationships between entities, where nodes represent entities and edges represent relationships; inference refers to the model simultaneously capturing temporal patterns and network structure to simulate the propagation and evolution of risks in both time and network topology dimensions, thereby enabling the tracing, real-time assessment, and early warning of complex risk chains.
[0063] The specific steps for multiplicative attention calculation are as follows: First, for the query vector at each time step and the key vectors at all time steps, calculate the similarity between them. Usually, the dot product multiplication method is used for calculation, that is, by calculating the dot product between the query vector and each key vector, a similarity score is obtained. These scores represent the relevance or matching degree between the query vector and the key vector. The softmax function is applied to transform these similarity scores into normalized attention weights, which represent the relative importance or influence between each time step.
[0064] A cross-time-segment attention mechanism is used to assign a weight to the features of each time segment. The weight value reflects the historical impact of the data in that time segment on the current entity node state. The calculated attention weights are used to weight and sum the features of each time segment to generate a context vector, thus obtaining the context vector of each entity node, which represents the historical impact and evolutionary background of that entity node.
[0065] Based on the preset fusion ratio factor in the database, the condensed vectors and context vectors are fused together with the causal relationship edges to generate the deep embedding vectors of each entity node.
[0066] Based on the preset fusion ratio factor in the database, and combined with causal relationship edges, the specific steps for generating the deep embedding vector of each entity node by fusing each condensed vector and each context vector are as follows: First, extract the condensed vector and context vector of each entity node from the time sequence graph inference model, where the condensed vector represents the local pattern and the context vector represents the historical influence and evolutionary background. Then, according to the preset fusion ratio factor in the database, which controls the weight of the condensed vector and the context vector in the final fused vector, the condensed vector and the context vector are weighted and summed according to this ratio. Where A represents the deep embedding vector of the entity node, B represents the condensed vector of the entity node, C represents the context vector of the entity node, and λ represents the obtained fusion scaling factor.
[0067] The database's preset fusion scaling factor controls the weight ratio of the condensed vector and context vector in the final deep embedding vector. This factor is typically a value between 0 and 1, representing the relative importance of the condensed vector and context vector during fusion. In the database's preset logic, the fusion scaling factor is adjusted based on the characteristics of different types of entity nodes and data. For certain types of entity nodes, the condensed vector may be given higher weight, indicating that the local pattern of that entity node is more important; while for other entity nodes, the context vector may be given more weight to reflect its historical evolution and background influence. The value of the fusion scaling factor can be set based on historical data, the type of entity node, and the experience of domain experts, or optimized through a data-driven approach. Ultimately, the preset fusion scaling factor ensures a balance of various types of information during the fusion process, enabling the final deep embedding vector to fully capture the multidimensional features of the entity nodes.
[0068] Traverse all entity nodes to obtain the deep embedding vector of each entity node, and simultaneously perform fusion scaling factor update.
[0069] For data from a given entity node across different time periods, the data for each time period is extracted and its corresponding time evolution vector is calculated. This vector contains the values of various statistical characteristics within the time period, such as the slope of the linear trend, maximum and minimum values, and rate of change. The time evolution vector is then quantified. Specifically, the variance of each element in the vector is calculated to measure the volatility of the data within that time period. By calculating the variance of all elements and averaging them, the vector variance of the data within that time period is obtained. This value reflects the dispersion and volatility of the data within that time period.
[0070] When the calculated vector variance is greater than or equal to the preset vector variance threshold in the database, it indicates that the data in this time period is highly volatile. The time series inference model needs to pay more attention to the data in this time period. Therefore, the vector variance threshold is subtracted from the vector variance to obtain the vector variance deviation value. Based on this, the weight increase value of the condensed vector is determined. This increase value will make the condensed vector in this time period occupy a greater weight in the fusion process, so as to highlight the significant patterns in this time period.
[0071] When the vector variance is less than the vector variance threshold, it indicates that the data in that time period has low volatility. The time series inference model should reduce its focus on the data in that time period and pay more attention to the data in the context. The vector variance deviation value is obtained by subtracting the vector variance threshold from the vector variance. Based on this, the weight reduction value of the condensed vector is determined. By reducing the weight of the data in that time period in the condensed vector, the model can effectively filter out low-volatility data, avoid over-reliance on noisy data, and achieve adaptive filtering.
[0072] In the database, the logic for obtaining weight increase and decrease values is implemented through a series of rules and functions to map vector variance deviation values to weight increase and decrease values. First, the database transforms the vector variance deviation value into the corresponding weight increase and decrease values according to preset mapping rules or functions. Common mapping rules include linear mapping, nonlinear mapping, and piecewise functions. The specific mapping relationship is preset in the database by control parameters, such as proportional coefficients or piecewise thresholds. Linear mapping transforms vector variance deviation values into weight increase and decrease values through simple proportional relationships, while nonlinear mapping adjusts them through more complex functions (such as exponential or logarithmic functions) to make larger deviation values produce stronger responses, emphasizing the importance of highly volatile data. Piecewise functions adjust to different degrees according to different deviation value ranges to ensure that the time series inference model can flexibly increase or decrease weights according to changes in actual data, thereby optimizing the analysis process and adjusting the model's focus in real time. In this way, the system can automatically adjust the response of the time series inference model according to real-time data changes, ensuring that the analysis results can effectively reflect the dynamic changes of the data.
[0073] The obtained value of increasing or decreasing the weight of the condensed vector is multiplied by the fusion scaling factor to obtain the updated value of the fusion scaling factor, which is then applied to the data in the next time period.
[0074] The variance of food safety data is detected by a sliding window, and the deviation value of food safety data variance is obtained by subtracting the data volatility threshold from the variance of food safety data.
[0075] When the detected variance deviation value of food safety data exceeds the preset data volatility deviation value threshold in the database, it indicates that the volatility of the current data is abnormal. It is necessary to trace back to the historical data according to the preset traceability step size and retrieve the variance of historical food safety data.
[0076] The preset data volatility deviation thresholds in the database typically rely on statistical analysis of historical data, experimental data, or on-site monitoring data. By analyzing the volatility characteristics of food safety data under different environmental conditions, volatility features can be extracted. Based on these volatility features, multiple thresholds can be set to determine the stability and volatility of the data in practical applications. By analyzing historical data, the significant range of data fluctuations before and after a food safety issue can be identified, thereby determining the volatility range for early warning. These volatility deviation thresholds are usually stored in the database for real-time comparison and analysis of food safety data.
[0077] The cumulative risk value of an entity node is obtained by subtracting the variance of historical food safety data from the variance deviation value of food safety data.
[0078] When the detected variance deviation value of food safety data is less than or equal to the preset data volatility deviation value threshold in the database, it indicates that the window volatility is acceptable, and the next window of detection can proceed.
[0079] Traverse all entity nodes to obtain the cumulative risk value of each entity node, and record it as the total risk value sequence.
[0080] By applying a weighted average to the total risk value series with a given smoothing factor, and then exponentially smoothing the total risk value series, the smoothed total risk value series is mapped to the time axis using Matplotlib (Matplotlib plotting library) to create a time series risk curve.
[0081] The preset smoothing factor is typically determined based on the experience of domain experts or through historical data analysis. In databases, the smoothing factor value is generally between 0 and 1. A smaller smoothing factor means that more of the influence of historical data is considered, which is suitable for situations with slower changes; a larger smoothing factor means that current data has a greater impact on the prediction result, which is suitable for situations with faster changes. The smoothing factor can be preset in several ways: First, a suitable smoothing factor can be selected based on the volatility and frequency of change of historical data; second, domain experts can manually set the smoothing factor value based on their understanding of the system's dynamic behavior; finally, automated methods can also be used, such as model validation or cross-validation, to optimize the smoothing factor with the goal of minimizing prediction error. The preset smoothing factor values in the database are usually stored as control parameters and called during time series data processing to ensure that the risk assessment model can perform appropriate smoothing under different conditions.
[0082] Matplotlib is a plotting library for the Python programming language. It provides a flexible and powerful way to create static, dynamic, and interactive graphs. Matplotlib supports various graph types, including line charts, scatter plots, bar charts, pie charts, histograms, etc., making it easy to use for data visualization. Users can customize the appearance of charts, such as colors, labels, titles, and axes, through simple commands and configurations. The core module of Matplotlib is graph drawing, which provides a plotting interface that allows users to draw graphs with concise syntax. In addition, Matplotlib supports output to various file formats, such as PNG, PDF, and SVG.
[0083] Extract the first, second, and third risk warning lines preset in the database.
[0084] The preset logic of the first, second, and third risk warning lines in the database is usually set based on historical data analysis, domain expert experience, and actual application needs. When setting these risk lines, control parameters in the database, such as the volatility of historical data, risk change trends, and safety boundaries set by experts, are used as references to ensure that different levels of risk events can be effectively monitored and responded to.
[0085] If the time-series risk curve does not exceed the first risk warning line, it means that the current risk level of the food safety system is in a low range and has not reached any warning threshold. This indicates that the current monitored food safety data or related risk factors have changed little and the system is operating in a relatively safe state. In this case, no risk warning will be triggered, and the warning level is determined to be no warning.
[0086] If the time-series risk curve has a portion that exceeds the first risk warning line but does not exceed the second risk warning line, it means that the current risk level has exceeded the slight warning threshold, but has not yet reached the level of medium risk. That is, although the risk level has fluctuated or become abnormal, it does not constitute an emergency. In this case, the warning level is determined to be a slight risk warning.
[0087] If the time-series risk curve has a portion that exceeds the second risk warning line but does not exceed the third risk warning line, it means that the current risk level has exceeded the medium risk range, but has not yet reached the highest warning level. This indicates that the current risk level is high and there may be a relatively serious potential threat, which requires sufficient attention. The warning level is determined to be a medium risk warning.
[0088] If the time-series risk curve exceeds the third risk warning line, it means that the current risk level has reached the highest warning threshold, indicating that the system's risk level is very high and there is a potential major threat or crisis that may seriously affect the safety of facilities, personnel, or the environment. Immediate emergency response measures are required, and the warning level is determined to be a severe risk warning.
[0089] In Embodiment 2 of the present invention, while keeping everything else unchanged from Embodiment 1, the total risk value sequence can also be calculated by calculating the mutation risk value. The specific analysis method is as follows: Based on the deep embedding vectors of entity nodes, the gradient of the deep embedding vector within each time period is calculated, and the difference of the deep embedding vector between adjacent time periods is calculated. That is, the difference of the deep embedding vector between two consecutive time periods is calculated in each dimension to obtain the gradient value of each dimension. Then, the Euclidean norm is used to aggregate the entire difference vector into a scalar representing the overall displacement amplitude. Finally, the overall displacement amplitude is divided by the corresponding time interval to obtain the overall rate of change of the embedding vector in the feature space per unit time, which reflects the rate of change of the embedding vector with time.
[0090] The Euclidean norm, also known as the L2 norm, is the most commonly used method to measure the length or size of a vector in a vector space. It is calculated as the square root of the sum of the squares of the vector's components. For a multidimensional vector, its Euclidean norm is the geometric straight-line distance from the origin to the endpoint of the vector in the multidimensional space. In the scenario of calculating the rate of change of deeply embedded vectors, by taking the Euclidean norm of the difference between two embedded vectors in adjacent time periods, the overall straight-line distance between the two in the feature space can be accurately measured, thereby avoiding the problem of the cancellation of the change directions in different dimensions. This yields a scalar value that reflects the comprehensive displacement amplitude of the vector, providing a stable and reliable geometric basis for quantifying the intensity of state transitions.
[0091] The preset gradient change rate threshold in the database is usually set based on historical data analysis, risk assessment models, and the experience of domain experts. First, by analyzing the gradient changes in historical time series data, gradient change patterns within the normal fluctuation range are identified. These change rates within the normal fluctuation range are used as a benchmark to determine a reasonable threshold. The setting of this threshold needs to take into account the natural volatility, periodic changes, and possible abnormal fluctuations of the data. Domain experts may adjust it according to different risk levels and application scenarios.
[0092] If the gradient rate of change is greater than the gradient rate of change threshold, it means that the deep embedding vector of the entity node has changed significantly in a short period of time, exceeding the normal fluctuation range. This usually indicates that there is a large change or mutation in the system, which may be caused by external shocks, abnormal events or drastic changes in internal trends. In this case, a mutation point is determined to have occurred.
[0093] If the gradient rate of change is less than or equal to the gradient rate of change threshold, it means that the deep embedding vector of the entity node changes relatively smoothly and is within the normal fluctuation range. This indicates that the changes in the system are in line with expectations and there are no violent abnormal fluctuations or sudden events. Therefore, it is judged that no mutation point has occurred.
[0094] Traverse all entity nodes to detect all mutation points, subtract the gradient change rate threshold from the gradient change rate to obtain the initial mutation risk value, amplify the calculated initial mutation risk value according to the preset mutation enhancement factor in the database, and multiply the initial mutation risk value by the mutation enhancement factor to obtain the mutation risk value of the entity.
[0095] The mutation enhancement factor is a weighted coefficient used to emphasize the impact of a mutation point. It is usually determined through historical data analysis, domain expert experience, and the magnitude of the mutation. First, based on historical data, the impact of different types of mutations on the system is assessed. If similar mutations in the past have caused significant system changes, the enhancement factor for the current mutation will be set larger to amplify its risk impact. Second, domain experts set reasonable enhancement factors based on long-term practical experience and understanding of the specific system. The mutation enhancement factor is usually proportional to the magnitude of the mutation. When the mutation magnitude is large, the enhancement factor is increased to ensure timely response.
[0096] Iterate through all detected mutation points and calculate the risk value for each mutation point following the steps described above. This results in a set of mutation risk values, which represent the degree of impact of different mutation points on the overall system risk.
[0097] In Embodiment 3 of the present invention, based on Embodiment 1 or Embodiment 2, the total risk value sequence can be calculated by calculating the mutation risk value and the cumulative risk value. The specific analysis method is as follows: After calculating the mutation risk value and the cumulative risk value, they are arranged in chronological order based on the time period in which the entity node's deep embedding vector is located. For a time period, if there is only one risk value, its corresponding risk value is used as the total risk value sequence for that time period. If both mutation risk value and cumulative risk value exist, the average value is processed and recorded as the total risk value sequence for that time period.
[0098] like Figure 5 The flowchart of the food safety risk early warning method based on multi-source heterogeneous data provided in this application embodiment includes: acquiring various original food safety datasets; synchronizing timestamps from different data sources using a clock skew estimation model; performing detailed analysis of food safety data for each entity node using a time evolution vector and temporal feature extraction module; dynamically adjusting the length of the evolution window by quantifying statistical indicators such as variance and cosine similarity of the food safety data; deriving causal relationships through Granger causality tests to further enrich the entity nodes and their causal relationship edges in the knowledge graph; subsequently, calculating the deep embedding vector of the entity node using a temporal graph inference model; combining the data from each evolution window to generate a condensed vector and context vector based on an attention mechanism, completing the deep embedding of the entity node; and, based on the calculation of volatility and mutation risk, smoothing the total risk value sequence and determining the early warning level according to a preset risk warning line in the database, ultimately achieving real-time monitoring and dynamic early warning.
[0099] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the above functions can be divided into different functional modules to complete all or part of the functions described above.
[0100] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0101] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units, located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0103] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the solution, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0104] The above content is only a specific implementation of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the protection scope of this application.
Claims
1. A food safety risk early warning system based on multi-source heterogeneous data, characterized in that, The system includes: a time evolution vector construction module, a knowledge graph deep embedding vector generation module, and a food safety risk curve construction module; The time evolution vector construction module is used to collect and analyze various original food safety datasets, obtain data for each time period of various original food safety datasets, construct time evolution vectors for various original food safety datasets for each time period of various original food safety datasets, analyze the variance and cosine similarity of food safety data and adjust the evolution window length accordingly. The knowledge graph deep embedding vector generation module is used to construct a food safety knowledge graph based on time evolution vectors, input the entity nodes and their attributes of the knowledge graph into the time sequence graph reasoning model, and analyze the deep embedding vector of each entity node. The food safety risk curve construction module is used to generate food safety risk curves for entity nodes based on the deep embedding vectors of each entity node, and to determine the warning level.
2. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 1, characterized in that: The specific analysis methods for collecting and analyzing various original food safety datasets are as follows: Obtain various raw food safety datasets; Perform structured parsing to uniformly convert various raw food safety datasets into a standard internal format that includes entity node identifiers, feature variables, original timestamps, and data sources; The built-in clock offset estimation model is invoked to automatically identify clock errors between different original food safety datasets; Based on clock errors, the timestamps of various original food safety datasets are corrected to form various aligned food safety datasets that are synchronized in the time dimension.
3. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 1, characterized in that: The specific analysis method for constructing time evolution vectors of various types of raw food safety data is as follows: Each type of original food safety data is recorded as an entity node, and a time evolution vector is constructed for each aligned food safety dataset of each entity node. For each aligned food safety dataset of an entity node, the aligned food safety dataset is divided into data for each time period according to a preset time window parameter; For data from each time period, the time series feature extraction module is invoked to calculate several statistical indicators within the data of each time period. The statistical indicators are then arranged according to feature correlation to obtain a multi-dimensional vector representing the time series features, denoted as the feature vector. By using a sliding window mechanism, feature vectors of data from adjacent time periods are superimposed in chronological order according to the evolution window length and evolution step size to obtain the temporal evolution vectors of various types of original food safety data.
4. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 1, characterized in that: The specific analysis method for analyzing the variance and cosine similarity of food safety data and adjusting the evolution window length accordingly is as follows: For a given time period of data, the variance of each food safety data point within that time period is calculated. This variance is then processed based on the mean variance of each food safety data point and recorded as the food safety data variance. The cosine similarity is calculated by comparing the feature vector of the current time window with the feature vector of the long-term historical pattern. The evolution window length is then dynamically adjusted based on the quantification results of these two indicators. When the variance of food safety data is detected to be greater than or equal to the preset data volatility threshold in the database, it is recorded as the first condition; when the cosine similarity is greater than or equal to the preset cosine similarity threshold in the database, it is recorded as the second condition. If the first and second conditions exist, the evolution window shortening value is determined based on the data volatility deviation value obtained from the food safety data variance and data volatility threshold, as well as the cosine similarity deviation value obtained from the cosine similarity and cosine similarity threshold. If only the first condition exists, then the evolution window shortening value is determined based on the data volatility deviation value; If only the second condition exists, the evolution window lengthening value is determined based on the cosine similarity deviation value. If neither of the two conditions exists, the original evolution window length is maintained; A new evolution window value is obtained based on the shortened or lengthened evolution window value and applied to the data in the next time period.
5. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 1, characterized in that: The specific analysis method for constructing a food safety knowledge graph based on time evolution vectors is as follows: S1, calls each entity node, whose entity node attributes are filled by the time evolution vector; S2, under a certain entity node, filter out all entity nodes that co-occur in time, generate each entity node pair, and for each pair of entity nodes, extract the historical value sequence of the entity node from the attributes of all consecutive entity nodes associated with them. S3. Granger causality test was used to analyze these two historical numerical sequences to obtain significance indicators. S4. If the significance index is greater than or equal to the preset significance threshold in the database, a new directional causal relationship edge is created between the two entity nodes in the knowledge graph, and the relationship is assigned a significance index and marked as a marked causal entity node pair, which is then removed in the next entity node pair causal analysis. If the significance index is less than the preset significance threshold in the database, then proceed to the next entity node pair causal analysis; S5, repeat S2 to S4, and analyze all causal relationship edges of this entity node; S6. Repeat S1 to S5 to analyze all entity nodes and complete the construction of the food safety knowledge graph.
6. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 1, characterized in that: The specific analysis method for the deep embedding vector of each entity node is as follows: After completing the construction of the food safety knowledge graph, the entity nodes, causal relationship edges and their attributes in the graph are input into the time sequence graph reasoning model. For an entity node, the temporal graph inference model first performs temporal attention calculation on the food safety data within each evolution window length of the entity node, generating vectors that highlight each evolution window, denoted as each condensation vector; Perform cross-time segment attention computation on food safety data between different evolution windows of entity nodes to generate vectors representing historical influence and evolutionary background, denoted as context vectors; Based on the preset fusion ratio factor in the database, the condensed vectors and context vectors are fused together with the causal relationship edges to generate the deep embedding vectors of each entity node. Traverse all entity nodes to obtain the deep embedding vector of each entity node, and simultaneously perform fusion scaling factor update.
7. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 6, characterized in that: The specific analysis method for updating the fusion scaling factor is as follows: For data of a certain entity node in different time periods, the corresponding time evolution vector is obtained for a time period of data. This time evolution vector contains the values of each statistical feature within the time period of data. The volatility within the vector is quantified: the variance of each element of the vector is calculated to obtain the degree of dispersion of the features within the segment. The variance of each element of the vector is synthesized to obtain the vector variance. Extract the preset vector variance threshold from the database; When the vector variance is greater than or equal to the vector variance threshold, the value of the concentrated vector weight improvement is determined based on the vector variance and the vector variance deviation value. When the vector variance is less than the vector variance threshold, the reduction value of the condensed vector weight is determined based on the deviation between the vector variance and the vector variance threshold, thereby achieving adaptive filtering. The updated value of the fusion scaling factor is obtained by increasing or decreasing the weight of the condensed vector, and then applied to the data in the next time period.
8. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 1, characterized in that: The method for generating food safety risk curves for entity nodes based on their deep embedding vectors and determining early warning levels is as follows: Calculate the total risk value series, smooth the total risk value series, and use the exponential smoothing method to remove short-term fluctuations to obtain the time series risk curve; Extract the preset first risk warning line, second risk warning line, and third risk warning line from the database; If the time-series risk curve does not exceed the first risk warning line, no warning will be issued; If the time-series risk curve has a portion that exceeds the first risk warning line but does not exceed the second risk warning line, the warning level is determined to be a minor risk warning. If the time-series risk curve contains a portion that exceeds the second risk warning line but does not exceed the third risk warning line, the warning level is determined to be a medium-risk warning. If the time-series risk curve contains a portion that exceeds the third risk warning line, the warning level is determined to be a severe risk warning.
9. The food safety risk early warning system based on multi-source heterogeneous data as described in claim 8, characterized in that: The specific analysis method for calculating the total risk value sequence is as follows: The sequence for calculating the total risk value includes the use of standardized cumulative risk value and / or mutation risk value; The specific analysis method for the accumulated risk value is as follows: The variance of food safety data is detected by a sliding window, and the deviation value of food safety data variance is obtained based on the variance of food safety data and the data volatility threshold. When the variance deviation of food safety data is detected to be greater than the preset data volatility deviation threshold in the database, the historical food safety data variance is traced back according to the preset traceability step size in the database. The cumulative risk value of the entity node is obtained based on the variance of historical food safety data and the deviation of food safety data variance. Traverse all entity nodes to obtain the cumulative risk value of each entity node; The specific analysis method for the mutation risk value is as follows: Calculate the gradient rate of change of the deep embedding vector of the entity node. If the gradient rate of change is greater than the preset gradient rate of change threshold in the database, it is determined that a mutation point has been generated. If the gradient rate of change is less than or equal to the preset gradient rate of change threshold in the database, it is determined that no mutation point has been generated. Iterate through all entity nodes, and for each detected mutation point, amplify it according to the preset mutation enhancement factor based on the magnitude of its deviation from the preset normal fluctuation range to obtain each mutation risk value.
10. A food safety risk early warning method based on multi-source heterogeneous data, applied to the food safety risk early warning system based on multi-source heterogeneous data as described in any one of claims 1-9, characterized in that, The method includes; Collect and analyze various original food safety datasets to obtain data for each time period of each type of original food safety data. For each time period of each type of original food safety data, construct the time evolution vector of each type of original food safety data, analyze the variance and cosine similarity of food safety data, and adjust the evolution window length accordingly. A food safety knowledge graph is constructed based on time evolution vectors. The entity nodes and their attributes of the knowledge graph are input into the time sequence graph reasoning model to analyze the deep embedding vector of each entity node. Based on the deep embedding vectors of each entity node, a food safety risk curve for the entity node is generated to determine the warning level.