Data analysis system and method for tracing information management platform

By dividing the refrigerated product transportation nodes, using differentiated missing value filling and standardization processing, building a dynamic threshold pool, and combining causal graphs and Bayesian formulas, the problems of low data quality and inaccurate anomaly judgment in refrigerated product transportation traceability are solved, and accurate anomaly source positioning and management optimization are achieved.

CN120822897APending Publication Date: 2025-10-21上海市大数据中心

Patent Information

Application Number
CN202511339899.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies fail to effectively consider the environmental characteristics of different types of nodes in the data analysis of refrigerated product transportation traceability, resulting in low data quality, poor adaptability, fixed safety thresholds that are not dynamically updated, anomaly judgments that are easily affected by instantaneous fluctuations, lack of causal relationship modeling, difficulty in accurately locating the source of quality anomalies, and impact on management efficiency and targeting.

Method used

The entire transportation process is divided into nodes. A differentiated missing value filling strategy and standardized processing are adopted to construct a dynamic threshold pool. The proportion of nodes exceeding the threshold is calculated by combining the node flow time. A causal graph is constructed and the source of anomalies is located in reverse through Bayesian formula. The influence weight of nodes is quantified and key nodes are selected for optimization.

Benefits of technology

It improves data quality and the accuracy of anomaly judgment, avoids misjudgment of instantaneous fluctuations, achieves accurate tracing of the source of anomalies, and improves the efficiency and pertinence of transportation management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822897A_ABST
    Figure CN120822897A_ABST
Patent Text Reader

Abstract

The invention discloses a data analysis system and method for tracing an information management platform, and relates to the technical field of data analysis, and the method comprises the following steps: dividing full-process transportation nodes, and deploying a sensor collection environment; according to the difference of node types, different missing value filling strategies are adopted, and environmental parameters after interpolation are subjected to standardization processing; screening qualified batch data according to quality scores, constructing environmental parameter safety thresholds of each node to form a threshold pool, and iteratively updating the thresholds according to the qualified batch data; determining an effective monitoring time period by combining node circulation time, calculating a standard environmental parameter over-threshold proportion, and judging whether the node is abnormal according to a proportion threshold value; and constructing a node causal map, reversely positioning an abnormal source through a Bayesian formula, calculating a node influence weight, and screening key nodes for feedback optimization. The condition that in the prior art, effective traceability is difficult when refrigerated goods transportation is abnormal can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to a data analysis system and method for a traceability information management platform. Background Art

[0002] In the entire supply chain of refrigerated products (such as fresh produce and pharmaceutical cold chain products), product quality is influenced by multiple factors, including the transportation environment (such as temperature and humidity), circulation efficiency, and operational specifications. Therefore, information traceability and management throughout the entire transportation process are crucial. The transportation process for these products typically encompasses multiple nodes, including static storage (such as departure warehousing and transit cold storage), dynamic transportation (such as refrigerated truck trunk transport), and short-term operations (such as loading and unloading, transit handling). Sensors must be deployed to collect environmental parameters, node flow information, and product quality data to enable visual monitoring of the transportation process and provide data support for product quality assurance.

[0003] However, existing technologies have obvious deficiencies in data analysis for refrigerated product transportation traceability: First, insufficient consideration is given to the differences in environmental characteristics of different types of nodes (static storage, dynamic transportation, and short-term operations), and unified data processing strategies (such as missing value filling and standardization methods) are often adopted, resulting in low data quality and poor adaptability; second, safety thresholds are mostly fixed values, which are not dynamically updated in combination with actual qualified batch data, making it difficult to adapt to dynamic changes in the transportation environment and affecting the accuracy of abnormality judgment; third, abnormality judgment often relies on instantaneous over-threshold data, and is not combined with the over-threshold ratio of the node during the effective monitoring period, making it susceptible to misjudgment due to accidental fluctuations; fourth, the causal relationship between nodes is ignored, and there is a lack of modeling of the influence relationship between the previous node and the current node, which makes it difficult to accurately locate the source when quality anomalies occur, and there is a lack of quantitative basis for the screening of key optimization nodes, affecting the efficiency and pertinence of traceability management. Summary of the Invention

[0004] The purpose of the present invention is to provide a data analysis system and method for a traceability information management platform to solve the problems raised in the prior art.

[0005] To achieve the above object, the present invention provides the following technical solution: a data analysis method for a traceability information management platform, the method comprising the following steps: Step 1: Divide the entire process into transport nodes, including static storage, dynamic transport, and short-term operation nodes. Deploy sensors to collect environmental, node flow, and quality data, and define the transmission method. Step 2: Different missing value filling strategies are used depending on the node type. When the data is incomplete, manual filling is performed and the environmental parameters after interpolation are standardized. Step 3: Filter qualified batch data according to quality scores, build safety thresholds for environmental parameters of each node to form a threshold pool, and iteratively update the thresholds using qualified batch data; Step 4: Determine the effective monitoring period based on the node turnover time, calculate the threshold-exceeding ratio of the standardized environmental parameters, and determine whether the node is abnormal based on the ratio threshold; Step 5: Build a node causal graph, use the Bayesian formula to reversely locate the source of the anomaly, calculate the node impact weight, and screen key nodes for feedback optimization.

[0006] In step 1, the entire refrigerated product transportation process is divided into n transportation nodes according to the process characteristics, which are expressed as: [h1,h2,…,h n ]; where n represents the number of transport nodes; h1,h2,…,h n Represent the 1st, 2nd,…, nth transport nodes respectively; Transport nodes cover all scenarios, including static storage nodes (such as departure warehousing, transit cold storage, etc.), dynamic transport nodes (such as refrigerated truck transportation, cold chain logistics, etc.), and short-term operation nodes (such as loading and unloading, transit handling, etc.); For static storage nodes, dynamic transportation nodes, and short-term operation nodes, sensors are deployed to collect environmental parameter data; Set the data collection interval Δt for static storage nodes and dynamic transport nodes; set the data collection interval Δt1 for short-time operation nodes; Obtain environmental data, including: node number, sensor ID, timestamp, and environmental parameters; Get node flow data, including: previous node, current node, entry time, and exit time; Obtain quality data, including: detection time, quality score; Static storage nodes transmit data through wired or wireless networks, dynamic transport nodes and short-time operation nodes transmit data through wireless networks, local cache is enabled in weak network environments, and missing data is automatically retransmitted after the network is restored.

[0007] In step 2, based on the characteristic differences of transport node types and combined with the collection interval, a differentiated missing value filling strategy is defined: For static storage nodes, the environment is stable and the data fluctuation is small. The missing values ​​rely on the weighted fusion of the previous and next time series data: when the environmental parameter data at a certain time point t is missing, and the data at the adjacent acquisition time points (t-Δt, t+Δt) are valid, the filling formula is: V(t)=w1×V(t-Δt)+w2×V(t+Δt); Where V(t) represents the environmental parameter value at time t; V(t-Δt) and V(t+Δt) represent the environmental parameter values ​​at time (t-Δt) and (t+Δt), respectively; w1 and w2 are weight coefficients; For dynamic transport nodes, the environment of dynamic transport nodes changes with location and driving status. The data has a time series trend. The missing values ​​need to retain the trend characteristics: when the environmental parameter data at a certain time point t is missing, and the data of the adjacent collection time points (t-Δt, t+Δt) are valid, linear interpolation is used to fill in the missing values: V(t)=V(t-Δt)+(t-(t-Δt)) / ((t+Δt)-(t-Δt))×(V(t+Δt)-V(t-Δt)); Short-term operation nodes have short operation time, high data continuity requirements, and missing data need to be quickly supplemented: when the missing interval is Δt1, the environmental parameter value of the previous valid time point is used; For static storage nodes or dynamic transport nodes, when environmental parameter data at a certain point in time is missing and the data at the adjacent collection time points are not all valid, or for short-term operation nodes, when the missing interval is greater than Δt1, a transmission threshold time limit ΔT is set. If no data is received within the ΔT time limit, the missing data is fed back and manually filled in. Standardize the interpolated environmental parameter data.

[0008] In step 3, historical data is filtered based on the quality score: a quality score threshold θ is set (used to filter qualified batch data), and batch data with a quality score not lower than θ is filtered; environmental parameter safety thresholds are constructed according to node type as a benchmark for abnormality judgment: Static storage node: calculate the mean μ of the standardized environmental parameters s and standard deviation σ s , threshold L s =μ s +a1σ s ; Dynamic transport node: Calculate the mean μ of the standardized environmental parameters d and standard deviation σ d , threshold L d =μ d +a2σ d ; Short-term operation node: Calculate the mean μ of the standardized environmental parameters o and standard deviation σ o , threshold L o =μ o +a3σ o ; Among them, a1, a2, and a3 represent the threshold standard deviation coefficients of static storage nodes, dynamic transportation nodes, and short-time operation nodes, respectively; Store threshold parameters to form a threshold pool; After each batch of transportation is completed, if the quality score ≥ θ, the mean and standard deviation of the corresponding node type are updated using the standardized environmental parameters of the current qualified batch: μ new=w3×μ old +w4×μ batch ; σ new =w3×σ old +w4×σ batch ; Among them, μ old , σ old represents the historical threshold parameter, μ batch , σ batch are the parameters of the current qualified batch; w3 and w4 are the weight coefficients of the threshold iteration (historical data accounts for a higher proportion to ensure threshold stability).

[0009] In step 4, based on the normalized data and the threshold pool, the abnormality level of each node is calculated: Combine the entry and exit times in the node flow data to determine the node h j The effective monitoring period P j =[t j,i ,t j,o ]; where j∈{1,2,…,n}; h j represents the jth transport node; t j,i Represents node h j Entry time, t j,o Represents node h j Time of departure; For node h j , calculate the threshold-exceeding ratio of standardized environmental parameters: E j =Σ t∈Pj I(V n (t)>L j ) / N j ; Where I() represents the indicator function, which is 1 when the condition is met and 0 otherwise; L j Represents node h j The threshold value; N j Represents node h j The effective number of acquisitions; V n (t) represents the standardized environmental parameter data; Set the ratio threshold ΔE (critical value for determining node abnormality), when E j When ≥ΔE, define node h j abnormal.

[0010] In step 5, based on the relationship between the previous node and the current node in the node flow data, the inter-node causal graph G(V,E) is constructed: Node set V: contains all transport nodes [h1,h2,…,h n ], each node attribute is associated with its historical threshold ratio sequence E j(t) (sorted by batch time) and historical quality correlation; Among them, the historical quality correlation is the correlation frequency between node anomalies and quality scores less than θ; Edge set E: edge (h j-1 ,h j ) represents the preceding node h j-1 For the current node h j The influence relationship, edge weight w j-1,j Indicates the impact intensity, w j-1,j ∈[0,1]; Among them, the weight w j-1,j By statistical h j-1 After threshold h j Calculation of the probability of exceeding the threshold; When the quality score is less than θ, it is determined to be a quality anomaly, and the source of the anomaly is located reversely using the Bayesian formula: Calculate the initial value of prior risk P(R j ), represented by node h j The ratio of the number of batches that are abnormal and result in a quality score less than θ1 to the total number of batches, indicating that node h j risk of congenital anomalies; Combined with the quality score Q, update the abnormal risk probability of each node: P(R j |Q)=P(Q|R j )×P(R j ) / Σ k P(Q|R k )×P(R k ); Among them, P(Q|R j ) is the posterior risk, indicating that node h j In the event of an abnormality, the conditional probability that the quality score is less than θ (based on historical data statistics); k∈{1,2,…,n}; Σ k P(Q|R k )×P(R k ) represents the total contribution probability of all nodes to the current quality anomaly Q<θ; P(R j |Q) is the posterior risk; P(R j |Q) represents the probability that a node is the source of the abnormality, given that the quality of the current batch is known to be abnormal. It is used to quantify the possibility that each node causes quality abnormalities. The higher the value, the higher the credibility of the node as the source. Screening posterior risk P(R j |Q)≥θ1, combined with the edge weight w of the causal graph G(V,E) j-1,j , eliminate isolated occasional anomalies and determine the abnormal source node h j *; (When there is no valid abnormal source node, an alarm will be pushed to prompt manual verification); When node h j P(R j |Q)≥θ1, but the edge weight w of the previous node j-1,j When it is less than the set threshold θ2, it is determined to be an isolated abnormality; Comprehensive node posterior risk P(R j |Q) and the actual abnormality rate, calculate the influence weight of each node: W j =w5×P(R j |Q)+w6×node h j Actual abnormal batch quantity / total batch quantity; Among them, w5 and w6 represent the weight coefficients that affect the weight calculation; Filter W j Nodes with a value greater than or equal to the set threshold θ3 are regarded as influencing nodes and are given priority for feedback and optimization reminders.

[0011] A data analysis system for a traceability information management platform, the system includes a data acquisition module, a data preprocessing module, a threshold management module, an anomaly detection module and a traceability optimization module; The data acquisition module is used to divide the entire process of transportation nodes and deploy the sensor collection environment; the data preprocessing module is used to standardize the environmental parameters after interpolation by adopting different missing value filling strategies according to the differences in node types; the threshold management module is used to screen qualified batch data according to the quality score, build the safety threshold of the environmental parameters of each node to form a threshold pool, and iteratively update the threshold according to the qualified batch data; the anomaly detection module is used to determine the effective monitoring period based on the node flow time, calculate the proportion of standardized environmental parameters exceeding the threshold, and determine whether the node is abnormal based on the proportion threshold; the traceability optimization module is used to construct a node causal map, reversely locate the source of the anomaly through the Bayesian formula, calculate the node impact weight, and screen key nodes for feedback optimization.

[0012] The data acquisition module includes a node division unit, a sensor deployment unit and a data transmission unit; The node division unit is used to divide the entire process of transportation nodes and generate a node sequence; the sensor deployment unit is used to deploy environmental parameter sensors and set the collection interval according to the node type; the data transmission unit is used to transmit environmental parameters and node flow data, and enable cache supplementary transmission when the network is weak; The data preprocessing module includes a missing filling unit and a normalization unit; The missing value filling unit is used to fill missing values ​​according to node type differences; the standardization unit is used to standardize the interpolated environmental parameters.

[0013] The threshold management module includes a data screening unit, a threshold calculation unit and a threshold update unit; The data screening unit is used to screen qualified batch data according to quality scores; the threshold calculation unit is used to construct environmental parameter safety thresholds according to node types and store them in a threshold pool; the threshold updating unit is used to iteratively update threshold parameters according to qualified batch data; The anomaly detection module includes a time period determination unit, an over-threshold calculation unit, and an anomaly determination unit; The time period determination unit is used to determine the effective monitoring time period in combination with the node flow time; the threshold exceeding calculation unit is used to calculate the threshold exceeding ratio of the standardized environmental parameter; and the abnormality determination unit is used to determine whether the node is abnormal based on the threshold exceeding ratio.

[0014] The traceability optimization module includes a graph construction unit, a risk calculation unit, a source positioning unit, and a weight calculation unit; The graph construction unit is used to construct a causal graph based on the node flow relationship; the risk calculation unit is used to calculate the posterior risk probability of the node through the Bayesian formula; the source positioning unit is used to eliminate isolated anomalies in combination with the graph and locate the abnormal source node; the weight calculation unit is used to calculate the node impact weight and screen the key optimization nodes.

[0015] Compared with the existing technology, the beneficial effects of the present invention are: the present invention can improve data quality and provide a reliable basis for subsequent analysis by designing differentiated missing value filling and standardization strategies for static, dynamic and short-term operation nodes; the present invention can make the abnormality judgment benchmark adaptive to changes in the transportation environment and improve accuracy by constructing a dynamic threshold pool and iteratively updating the threshold in combination with qualified batch data; the present invention can avoid misjudgment of instantaneous fluctuations and improve the accuracy of abnormal node identification by calculating the over-threshold ratio in combination with the effective monitoring period of the node; the present invention can achieve accurate tracing of the source of the abnormality by constructing a node causal graph and combining it with the Bayesian formula to reversely locate the source of the abnormality, and at the same time quantify the node influence weight to screen key nodes, thereby improving the efficiency of tracing management and optimization targeting. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of the entire process of a data analysis method for a traceability information management platform according to the present invention; Figure 2 The present invention is a flow chart of a data analysis system for a traceability information management platform. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0018] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution, a data analysis method for a traceability information management platform, the method comprising the following steps: Step 1: Divide the entire process into transport nodes, including static storage, dynamic transport, and short-term operation nodes. Deploy sensors to collect environmental, node flow, and quality data, and define the transmission method. Step 2: Different missing value filling strategies are used depending on the node type. When the data is incomplete, manual filling is performed and the environmental parameters after interpolation are standardized. Step 3: Filter qualified batch data according to quality scores, build safety thresholds for environmental parameters of each node to form a threshold pool, and iteratively update the thresholds using qualified batch data; Step 4: Determine the effective monitoring period based on the node turnover time, calculate the threshold-exceeding ratio of the standardized environmental parameters, and determine whether the node is abnormal based on the ratio threshold; Step 5: Build a node causal graph, use the Bayesian formula to reversely locate the source of the anomaly, calculate the node impact weight, and screen key nodes for feedback optimization.

[0019] In step 1, the entire refrigerated product transportation process is divided into n transportation nodes according to the process characteristics, which are expressed as: [h1,h2,…,h n ]; where n represents the number of transport nodes; h1,h2,…,h n Represent the 1st, 2nd,…, nth transport nodes respectively; Transport nodes cover all scenarios, including static storage nodes (such as departure warehousing, transit cold storage, etc.), dynamic transport nodes (such as refrigerated truck transportation, cold chain logistics, etc.), and short-term operation nodes (such as loading and unloading, transit handling, etc.); For static storage nodes, dynamic transportation nodes, and short-term operation nodes, sensors are deployed to collect environmental parameter data; Set the data collection interval Δt for static storage nodes and dynamic transport nodes; set the data collection interval Δt1 for short-time operation nodes; Obtain environmental data, including: node number, sensor ID, timestamp, and environmental parameters; Get node flow data, including: previous node, current node, entry time, and exit time; Obtain quality data, including: detection time, quality score; Static storage nodes transmit data through wired or wireless networks, dynamic transport nodes and short-time operation nodes transmit data through wireless networks, local cache is enabled in weak network environments, and missing data is automatically retransmitted after the network is restored.

[0020] In step 2, based on the characteristic differences of transport node types and combined with the collection interval, a differentiated missing value filling strategy is defined: For static storage nodes, the environment is stable and the data fluctuation is small. The missing values ​​rely on the weighted fusion of the previous and next time series data: when the environmental parameter data at a certain time point t is missing, and the data at the adjacent acquisition time points (t-Δt, t+Δt) are valid, the filling formula is: V(t)=w1×V(t-Δt)+w2×V(t+Δt); Where V(t) represents the environmental parameter value at time t; V(t-Δt) and V(t+Δt) represent the environmental parameter values ​​at time (t-Δt) and (t+Δt), respectively; w1 and w2 are weight coefficients; For dynamic transport nodes, the environment of dynamic transport nodes changes with location and driving status. The data has a time series trend. The missing values ​​need to retain the trend characteristics: when the environmental parameter data at a certain time point t is missing, and the data of the adjacent collection time points (t-Δt, t+Δt) are valid, linear interpolation is used to fill in the missing values: V(t)=V(t-Δt)+(t-(t-Δt)) / ((t+Δt)-(t-Δt))×(V(t+Δt)-V(t-Δt)); Short-term operation nodes have short operation time, high data continuity requirements, and missing data need to be quickly supplemented: when the missing interval is Δt1, the environmental parameter value of the previous valid time point is used; For static storage nodes or dynamic transport nodes, when environmental parameter data at a certain point in time is missing and the data at the adjacent collection time points are not all valid, or for short-term operation nodes, when the missing interval is greater than Δt1, a transmission threshold time limit ΔT is set. If no data is received within the ΔT time limit, the missing data is fed back and manually filled in. Standardize the interpolated environmental parameter data.

[0021] In step 3, historical data is filtered based on the quality score: a quality score threshold θ is set (used to filter qualified batch data), and batch data with a quality score not lower than θ is filtered; environmental parameter safety thresholds are constructed according to node type as a benchmark for abnormality judgment: Static storage node: calculate the mean μ of the standardized environmental parameters s and standard deviation σ s , threshold L s =μ s +a1σ s ; Dynamic transport node: Calculate the mean μ of the standardized environmental parameters d and standard deviation σ d , threshold L d =μ d +a2σ d ; Short-term operation node: Calculate the mean μ of the standardized environmental parameters o and standard deviation σ o , threshold L o =μ o +a3σ o ; Among them, a1, a2, and a3 represent the threshold standard deviation coefficients of static storage nodes, dynamic transportation nodes, and short-time operation nodes, respectively; Store threshold parameters to form a threshold pool; After each batch of transportation is completed, if the quality score ≥ θ, the mean and standard deviation of the corresponding node type are updated using the standardized environmental parameters of the current qualified batch: μ new =w3×μ old +w4×μ batch ; σ new =w3×σ old +w4×σ batch ; Among them, μ old , σ old represents the historical threshold parameter, μ batch , σ batch are the parameters of the current qualified batch; w3 and w4 are the weight coefficients of the threshold iteration (historical data accounts for a higher proportion to ensure threshold stability).

[0022] In step 4, based on the normalized data and the threshold pool, the abnormality level of each node is calculated: Combine the entry and exit times in the node flow data to determine the node h j The effective monitoring period P j =[t j,i ,t j,o ]; where j∈{1,2,…,n}; h j represents the jth transport node; t j,i Represents node h j Entry time, t j,o Represents node h j Time of departure; For node h j , calculate the threshold-exceeding ratio of standardized environmental parameters: E j =Σ t∈Pj I(V n (t)>L j ) / N j ; Where I() represents the indicator function, which is 1 when the condition is met and 0 otherwise; L j Represents node h j The threshold value of N j Represents node h j The effective number of acquisitions; V n (t) represents the standardized environmental parameter data; Set the ratio threshold ΔE (critical value for determining node abnormality), when E j When ≥ΔE, define node h j abnormal.

[0023] In step 5, based on the relationship between the previous node and the current node in the node flow data, the inter-node causal graph G(V,E) is constructed: Node set V: contains all transport nodes [h1,h2,…,h n ], each node attribute is associated with its historical threshold ratio sequence E j (t) (sorted by batch time) and historical quality correlation; Among them, the historical quality correlation is the correlation frequency between node anomalies and quality scores less than θ; Edge set E: edge (h j-1 ,h j ) represents the preceding node h j-1 For the current node h j The influence relationship, edge weight w j-1,j Indicates the impact intensity, w j-1,j ∈[0,1]; Among them, the weight w j-1,j By statistical h j-1 After threshold h j Calculation of the probability of exceeding the threshold; When the quality score is less than θ, it is determined to be a quality anomaly, and the source of the anomaly is located reversely using the Bayesian formula: Calculate the initial value of prior risk P(R j ), represented by node h j The ratio of the number of batches that are abnormal and result in a quality score less than θ1 to the total number of batches, indicating that node h j risk of congenital anomalies; Combined with the quality score Q, update the abnormal risk probability of each node: P(R j |Q)=P(Q|R j )×P(R j ) / Σ k P(Q|R k )×P(R k ); Among them, P(Q|R j ) is the posterior risk, indicating that node hj In the event of an abnormality, the conditional probability that the quality score is less than θ (based on historical data statistics); k∈{1,2,…,n}; Σ k P(Q|R k )×P(R k ) represents the total contribution probability of all nodes to the current quality anomaly Q<θ; P(R j |Q) is the posterior risk; P(R j |Q) represents the probability that a node is the source of the abnormality, given that the quality of the current batch is known to be abnormal. It is used to quantify the possibility that each node causes quality abnormalities. The higher the value, the higher the credibility of the node as the source. Screening posterior risk P(R j |Q)≥θ1, combined with the edge weight w of the causal graph G(V,E) j-1,j , eliminate isolated occasional anomalies and determine the abnormal source node h j * ; (When there is no valid abnormal source node, an alarm will be pushed to prompt manual verification); When node h j P(R j |Q)≥θ1, but the edge weight w of the previous node j-1,j When it is less than the set threshold θ2, it is determined to be an isolated abnormality; Comprehensive node posterior risk P(R j |Q) and the actual abnormality rate, calculate the influence weight of each node: W j =w5×P(R j |Q)+w6×node h j Actual abnormal batch quantity / total batch quantity; Among them, w5 and w6 represent the weight coefficients that affect the weight calculation; Filter W j Nodes with a value greater than or equal to the set threshold θ3 are regarded as influencing nodes and are given priority for feedback and optimization reminders.

[0024] A data analysis system for a traceability information management platform, the system includes a data acquisition module, a data preprocessing module, a threshold management module, an anomaly detection module and a traceability optimization module; The data acquisition module is used to divide the entire process of transportation nodes and deploy the sensor collection environment; the data preprocessing module is used to standardize the environmental parameters after interpolation by adopting different missing value filling strategies according to the differences in node types; the threshold management module is used to screen qualified batch data according to the quality score, build the safety threshold of the environmental parameters of each node to form a threshold pool, and iteratively update the threshold according to the qualified batch data; the anomaly detection module is used to determine the effective monitoring period based on the node flow time, calculate the proportion of standardized environmental parameters exceeding the threshold, and determine whether the node is abnormal based on the proportion threshold; the traceability optimization module is used to construct a node causal map, reversely locate the source of the anomaly through the Bayesian formula, calculate the node impact weight, and screen key nodes for feedback optimization.

[0025] The data acquisition module includes a node division unit, a sensor deployment unit and a data transmission unit; The node division unit is used to divide the entire process of transportation nodes and generate a node sequence; the sensor deployment unit is used to deploy environmental parameter sensors and set the collection interval according to the node type; the data transmission unit is used to transmit environmental parameters and node flow data, and enable cache supplementary transmission when the network is weak; The data preprocessing module includes a missing filling unit and a normalization unit; The missing value filling unit is used to fill missing values ​​according to node type differences; the standardization unit is used to standardize the interpolated environmental parameters.

[0026] The threshold management module includes a data screening unit, a threshold calculation unit and a threshold update unit; The data screening unit is used to screen qualified batch data according to quality scores; the threshold calculation unit is used to construct environmental parameter safety thresholds according to node types and store them in a threshold pool; the threshold updating unit is used to iteratively update threshold parameters according to qualified batch data; The anomaly detection module includes a time period determination unit, an over-threshold calculation unit, and an anomaly determination unit; The time period determination unit is used to determine the effective monitoring time period in combination with the node flow time; the threshold exceeding calculation unit is used to calculate the threshold exceeding ratio of the standardized environmental parameter; and the abnormality determination unit is used to determine whether the node is abnormal based on the threshold exceeding ratio.

[0027] The traceability optimization module includes a graph construction unit, a risk calculation unit, a source positioning unit, and a weight calculation unit; The graph construction unit is used to construct a causal graph based on the node flow relationship; the risk calculation unit is used to calculate the posterior risk probability of the node through the Bayesian formula; the source positioning unit is used to eliminate isolated anomalies in combination with the graph and locate the abnormal source node; the weight calculation unit is used to calculate the node impact weight and screen the key optimization nodes.

[0028] In this embodiment, the full-process traceability information management of a certain refrigerated meat (fresh meat) is specifically implemented as follows: Step 1: Transport node division and data collection; The entire process of fresh meat from the slaughterhouse to the terminal supermarket is divided into 6 transportation nodes: [h1, h2, h3, h4, h5, h6], among which: h1 (departure warehouse), h3 (transit cold storage), h5 (terminal cold storage) are static storage nodes; h2 (long-distance refrigerated truck transportation) and h4 (short-distance refrigerated truck transshipment) are dynamic transportation nodes; h6 (terminal loading and unloading) is a short-term operation node.

[0029] Data collection settings: Environmental parameters: temperature, humidity; The collection interval of static storage nodes (h1, h3, h5) is Δt = 5 minutes; the collection interval of dynamic transportation nodes (h2, h4) is Δt = 5 minutes; the collection interval of short-time operation nodes (h6) is Δt1 = 1 minute; Node flow data: records the predecessor node of each node (for example, the predecessor node of h2 is h1, the predecessor node of h3 is h2, etc.), entry time, and exit time; Quality data: Each batch is scored after arriving at the terminal (with a maximum score of 100, including indicators such as freshness and microbial content); Transmission method: static nodes use wired + WiFi transmission, dynamic nodes and short-term nodes use 4G / 5G transmission; when the network is weak (such as remote sections), local SD card cache is enabled, and the data will be automatically uploaded within 1 hour after the network is restored.

[0030] Step 2: Missing value filling and standardization: Differentiated missing value filling strategy: Static storage nodes (h1, h3, h5): Due to the stable environment, missing values ​​are filled using weighted time series, with weights w1=0.5 and w2=0.5 (in a static environment, the impact of the previous and next time series on the missing points is similar, with no obvious trend deviation). Dynamic transport nodes (h2, h4): Because temperature fluctuates with the road section (e.g., open air / tunnel) and has a trend, missing values ​​are interpolated linearly (e.g., if time t is missing in h2, it is linearly calculated based on the temperature at t-5 minutes and t+5 minutes); Short-term operation node (h6): The operation time is usually ≤30 minutes, and the data continuity requirement is high. If the missing interval is ≤1 minute, the data of the previous valid time point will be used; When the data is incomplete: set the transmission threshold time limit ΔT = 30 minutes (if no supplementary data is received after 30 minutes), trigger manual filling (such as contacting the driver / warehouse manager to supplement the records).

[0031] Standardization processing: All interpolated temperature and humidity data are standardized using Z-score (the mean of the standardized data is 0 and the standard deviation is 1).

[0032] Step 3: Qualified batch screening and threshold pool construction; The quality score threshold θ = 80 points (qualified batch data with a score ≥ 80 points are screened); Calculation of safety thresholds for environmental parameters of each node: Static storage nodes (h1, h3, h5): threshold standard deviation coefficient a1=1.2 (static storage time is long (such as h1 is usually stored for 6-12 hours), slightly larger fluctuations are allowed, 1.2 times the standard deviation can cover more than 90% of qualified data), threshold L s =μ s +1.2σ s ; Dynamic transport nodes (h2, h4): threshold standard deviation coefficient a2 = 1.5 (temperature fluctuations in dynamic transport are slightly larger (such as opening the door of a refrigerated truck for ventilation), 1.5 times the standard deviation can cover more than 85% of qualified data), threshold L d =μ d +1.5σ d ; Short-term operation node (h6): threshold standard deviation coefficient a3 = 0.8 (the temperature is prone to sudden rise during loading and unloading, and needs to be strictly controlled. 0.8 times the standard deviation can cover more than 95% of qualified data), threshold L o =μ o +0.8σ o ; Iterative threshold update: weights w3 = 0.8, w4 = 0.2 (historical data (w3) accounts for a higher proportion to prevent single-batch outliers from interfering with threshold stability. The current qualified batch (w4) only fine-tunes the threshold to adapt to long-term trends).

[0033] Step 4: Node abnormality determination; Effective monitoring period: determined based on the node entry / exit time, such as [8:00, 14:00] for h1, [14:30, 18:30] for h2, etc. Calculation of threshold-exceeding ratio: Calculate the ratio of the number of times the standardized environmental parameters exceed the threshold within the effective period for each node (E j ); Ratio threshold ΔE = 10% (in historical qualified batches, the quality is stable when the threshold ratio is ≤ 10%, and exceeding it may affect the quality). j When the percentage is ≥10%, the node is considered abnormal.

[0034] Step 5: Abnormal source location and optimization feedback; Causal graph construction: Node set V: contains h1-h6, attribute association of each node's historical threshold ratio sequence (such as h2's E in the past three months)j sequence) and historical quality correlation (e.g., when h2 is abnormal, the frequency of quality score < 80 points is 15%); edge set E: edge weight reflects the influence of the previous node on the current node, such as w 1,2 =0.7 (abnormal h1 temperature may lead to high initial h2 temperature), w 5,6 =0.8 (the temperature of the cold storage in h5 directly affects the meat temperature during loading and unloading in h6) (the physical connection between the previous node and the current node is close, the more direct the connection, the higher the weight); Abnormal source location: Prior risk P(R j ): h2=8%, h4=9% (dynamic transportation has a complex environment and a higher risk of congenital abnormalities), h1=5%, h3=4% (static storage has high stability); The posterior risk screening threshold θ1 = 30% (only nodes with a probability of causing quality abnormality ≥ 30% are retained); Isolated anomaly determination: edge weight threshold θ2 = 0.3 (when the influence of the previous node on the current node is less than 30%, it is considered an isolated anomaly); Impact weight and optimization: Impact weight coefficients w5=0.6 and w6=0.4 (the posterior risk better reflects the abnormal correlation of the current batch, while the actual abnormal rate reflects the long-term stability, so the former has a slightly higher weight); Screening threshold θ3 = 40% (W j Nodes with a value of ≥40% are the key optimization targets), such as W of h2 in a batch j =55%, the refrigerated truck temperature control system calibration is given priority.

[0035] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A data analysis method for a traceability information management platform, characterized by: The method comprises the following steps: Step 1: Divide the entire transportation process of refrigerated agricultural products into nodes and deploy a sensor collection environment; Step 2: According to the differences in node types, different missing value filling strategies are used to standardize the interpolated environmental parameters; Step 3: Filter qualified batch data according to quality scores, build safety thresholds for environmental parameters of each node to form a threshold pool, and iteratively update the thresholds based on qualified batch data; Step 4: Determine the effective monitoring period based on the node turnover time, calculate the threshold-exceeding ratio of the standardized environmental parameters, and determine whether the node is abnormal based on the ratio threshold; Step 5: Build a node causal graph, use the Bayesian formula to reversely locate the source of the anomaly, calculate the node impact weight, and screen key nodes for feedback optimization.

2. A data analysis method for a traceability information management platform according to claim 1, characterized in that: In step 1, the entire refrigerated product transportation process is divided into n transportation nodes according to the process characteristics, which are expressed as: [h1,h2,…,h n ]; where n represents the number of transport nodes; h1,h2,…,h n Represent the 1st, 2nd,…, nth transport nodes respectively; Transport nodes cover all scenarios, including static storage nodes, dynamic transport nodes, and short-term operation nodes; For static storage nodes, dynamic transportation nodes, and short-term operation nodes, sensors are deployed to collect environmental parameter data; Set the data collection interval Δt for static storage nodes and dynamic transport nodes; set the data collection interval Δt1 for short-time operation nodes; Obtain environmental data, including: node number, sensor ID, timestamp, and environmental parameters; Get node flow data, including: previous node, current node, entry time, and exit time; Obtain quality data, including: detection time, quality score; Static storage nodes transmit data through wired or wireless networks, dynamic transport nodes and short-time operation nodes transmit data through wireless networks, local cache is enabled in weak network environments, and missing data is automatically retransmitted after the network is restored.

3. A data analysis method for a traceability information management platform according to claim 2, characterized in that: In step 2, based on the characteristic differences of transport node types and combined with the collection interval, a differentiated missing value filling strategy is defined: For static storage nodes: When the environmental parameter data at a certain time point t is missing, and the data at the adjacent collection time points (t-Δt, t+Δt) are valid, the filling formula is: V(t)=w1×V(t-Δt)+w2×V(t+Δt); Where V(t) represents the environmental parameter value at time t; V(t-Δt) and V(t+Δt) represent the environmental parameter values ​​at time (t-Δt) and (t+Δt), respectively; w1 and w2 are weight coefficients; For dynamic transportation nodes: When the environmental parameter data at a certain time point t is missing, and the data at the adjacent collection time points (t-Δt, t+Δt) are valid, linear interpolation is used to fill in the data: V(t)=V(t-Δt)+(t-(t-Δt)) / ((t+Δt)-(t-Δt))×(V(t+Δt)-V(t-Δt)); Short-term operation node: When the missing interval is Δt1, the environmental parameter value at the previous valid time point is used; For static storage nodes or dynamic transport nodes, when environmental parameter data at a certain point in time is missing and the data at the adjacent collection time points are not all valid, or for short-term operation nodes, when the missing interval is greater than Δt1, a transmission threshold time limit ΔT is set. If no data is received within the ΔT time limit, the missing data is fed back and manually filled in. Standardize the interpolated environmental parameter data.

4. A data analysis method for a traceability information management platform according to claim 3, characterized in that: In step 3, historical data is filtered based on the quality score: a quality score threshold θ is set to filter batches of data with a quality score not lower than θ; environmental parameter safety thresholds are constructed by node type as a benchmark for abnormality judgment: Static storage node: calculate the mean μ of the standardized environmental parameters s and standard deviation σ s , threshold L s =μ s +a1σ s ; Dynamic transport node: Calculate the mean μ of the standardized environmental parameters d and standard deviation σ d , threshold L d =μ d +a2σ d ; Short-term operation node: Calculate the mean μ of the standardized environmental parameters o and standard deviation σ o , threshold L o =μ o +a3σ o ; Among them, a1, a2, and a3 represent the threshold standard deviation coefficients of static storage nodes, dynamic transportation nodes, and short-time operation nodes, respectively; Store threshold parameters to form a threshold pool; After each batch of transportation is completed, if the quality score ≥ θ, the mean and standard deviation of the corresponding node type are updated using the standardized environmental parameters of the current qualified batch: μ new =w3×μ old +w4×μ batch ; σ new =w3×σ old +w4×σ batch ; Among them, μ old , σ old represents the historical threshold parameter, μ batch , σ batch is the parameter of the current qualified batch; w3 and w4 are the weight coefficients of the threshold iteration respectively.

5. The data analysis method for a traceability information management platform according to claim 4, characterized in that: In step 4, based on the normalized data and the threshold pool, the abnormality level of each node is calculated: Combine the entry and exit times in the node flow data to determine the node h j The effective monitoring period P j =[t j,i ,t j,o ]; where j∈{1,2,…,n}; h j represents the jth transport node; t j,i Represents node h j Entry time, t j,o Represents node h j Time of departure; For node h j , calculate the threshold-exceeding ratio of standardized environmental parameters: E j =Σ t∈Pj I(V n (t)>L j ) / N j ; Where I() represents the indicator function, which is 1 when the condition is met and 0 otherwise; L j Represents node h j The threshold value of N j Represents node h j The effective number of acquisitions; V n (t) represents the standardized environmental parameter data; Set the ratio threshold ΔE, when E j When ≥ΔE, define node h j abnormal.

6. The data analysis method for a traceability information management platform according to claim 5, characterized in that: In step 5, based on the relationship between the previous node and the current node in the node flow data, the inter-node causal graph G(V,E) is constructed: Node set V: contains all transport nodes [h1,h2,…,h n ], each node attribute is associated with its historical threshold ratio sequence E j (t) correlation with historical quality; Among them, the historical quality correlation is the correlation frequency between node anomalies and quality scores less than θ; Edge set E: edge (h j-1 ,h j ) represents the preceding node h j-1 For the current node h j The influence relationship, edge weight w j-1,j Indicates the impact intensity, w j-1,j ∈[0,1]; Among them, the weight w j-1,j By statistical h j-1 After threshold h j Calculation of the probability of exceeding the threshold; When the quality score is less than θ, it is determined to be a quality anomaly, and the source of the anomaly is located reversely using the Bayesian formula: Calculate the initial value of prior risk P(R j ), represented by node h j The ratio of the number of batches that are abnormal and result in a quality score less than θ1 to the total number of batches, indicating that node h j risk of congenital anomalies; Combined with the quality score Q, update the abnormal risk probability of each node: P(R j |Q)=P(Q|R j )×P(R j ) / Σ k P(Q|R k )×P(R k ); Among them, P(Q|R j ) is the posterior risk, indicating that node h j The conditional probability that the quality score is less than θ when abnormal; k∈{1,2,…,n}; Σ k P(Q|R k )×P(R k ) represents the total contribution probability of all nodes to the current quality anomaly Q<θ; P(R j |Q) is the posterior risk; Screening posterior risk P(R j |Q)≥θ1, combined with the edge weight w of the causal graph G(V,E) j-1,j , eliminate isolated occasional anomalies and determine the abnormal source node h j * ; When node h j P(R j |Q)≥θ1, but the edge weight w of the previous node j-1,j When it is less than the set threshold θ2, it is determined to be an isolated abnormality; Comprehensive node posterior risk P(R j |Q) and the actual abnormality rate, calculate the influence weight of each node: W j =w5×P(R j |Q)+w6×node h j Actual abnormal batch quantity / total batch quantity; Among them, w5 and w6 represent the weight coefficients that affect the weight calculation; Filter W j Nodes with a value greater than or equal to the set threshold θ3 are regarded as influencing nodes and are given priority for feedback and optimization reminders.

7. A data analysis system for a traceability information management platform, applied to the data analysis method for a traceability information management platform according to any one of claims 1 to 6, characterized in that: The system includes a data acquisition module, a data preprocessing module, a threshold management module, an anomaly detection module and a traceability optimization module; The data acquisition module is used to divide the entire process of transportation nodes and deploy the sensor collection environment; the data preprocessing module is used to standardize the environmental parameters after interpolation by adopting different missing value filling strategies according to the differences in node types; the threshold management module is used to screen qualified batch data according to the quality score, build the safety threshold of the environmental parameters of each node to form a threshold pool, and iteratively update the threshold according to the qualified batch data; the anomaly detection module is used to determine the effective monitoring period based on the node flow time, calculate the proportion of standardized environmental parameters exceeding the threshold, and determine whether the node is abnormal based on the proportion threshold; the traceability optimization module is used to construct a node causal map, reversely locate the source of the anomaly through the Bayesian formula, calculate the node impact weight, and screen key nodes for feedback optimization.

8. The data analysis system for a traceability information management platform according to claim 7, characterized in that: The data acquisition module includes a node division unit, a sensor deployment unit and a data transmission unit; The node division unit is used to divide the entire process of transportation nodes and generate a node sequence; the sensor deployment unit is used to deploy environmental parameter sensors and set the collection interval according to the node type; the data transmission unit is used to transmit environmental parameters and node flow data, and enable cache supplementary transmission when the network is weak; The data preprocessing module includes a missing filling unit and a normalization unit; The missing value filling unit is used to fill missing values ​​according to node type differences; the standardization unit is used to standardize the interpolated environmental parameters.

9. The data analysis system for a traceability information management platform according to claim 8, characterized in that: The threshold management module includes a data screening unit, a threshold calculation unit and a threshold update unit; The data screening unit is used to screen qualified batch data according to quality scores; the threshold calculation unit is used to construct environmental parameter safety thresholds according to node types and store them in a threshold pool; the threshold updating unit is used to iteratively update threshold parameters according to qualified batch data; The anomaly detection module includes a time period determination unit, an over-threshold calculation unit, and an anomaly determination unit; The time period determination unit is used to determine the effective monitoring time period in combination with the node flow time; the threshold exceeding calculation unit is used to calculate the threshold exceeding ratio of the standardized environmental parameter; and the abnormality determination unit is used to determine whether the node is abnormal based on the threshold exceeding ratio.

10. The data analysis system for a traceability information management platform according to claim 9, characterized in that: The traceability optimization module includes a graph construction unit, a risk calculation unit, a source positioning unit, and a weight calculation unit; The graph construction unit is used to construct a causal graph based on the node flow relationship; the risk calculation unit is used to calculate the posterior risk probability of the node using the Bayesian formula; The source location unit is used to eliminate isolated anomalies in combination with the graph and locate the abnormal source node; The weight calculation unit is used to calculate the node influence weight and select the key optimization nodes.

Citation Information

Patent Citations

  • Meat product quality safety full-period intelligent tracing method and system

    CN118982362A

  • Medical quality management method and system based on big data and artificial intelligence

    CN119557779A

  • Food supply chain quality safety tracing method and system

    CN119809664A

  • Cold-chain logistics path traceability management system

    CN120258660A

  • Cold chain product whole-process wireless traceability system based on Internet of Things

    CN120612104A

Cited By

  • Comprehensive test device and test method applied to power supply filter

    CN121763170A

  • A comprehensive testing device and testing method applied to a power filter

    CN121763170B