Agricultural product supply chain traceability data integrity verification method based on machine learning
By constructing a multi-dimensional evaluation matrix and a liquid neural network model, the problems of low accuracy and low computational efficiency in verifying the integrity of agricultural product supply chain traceability data were solved, achieving efficient and accurate data integrity verification.
Patent Information
- Application Number
- CN202511154989.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for verifying the integrity of agricultural product supply chain traceability data have low accuracy and low computational efficiency, making it difficult to meet the needs of complex and ever-changing supply chain environments.
A multi-dimensional evaluation matrix system and a liquid neural network data fusion verification model are constructed, including a stability evaluation matrix, a loss risk matrix, a jump detection matrix, and an anomaly identification matrix. Combined with a liquid neural network with embedded nonlinear activation functions, multi-source data fusion processing and verification are performed.
It enables accurate verification of agricultural product supply chain traceability data, improves verification accuracy and optimizes computational efficiency, and adapts to the real-time requirements of complex supply chain structures.
Smart Images

Figure CN120996835A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of agricultural product supply chain technology, and more specifically, relates to a method for verifying the integrity of agricultural product supply chain traceability data based on machine learning. Background Technology
[0002] Agricultural product supply chain traceability is a crucial technological means to ensure food safety and protect consumer rights. Traditional agricultural product supply chain traceability data verification primarily relies on manual inspection, simple database comparisons, and basic statistical analysis methods. This involves setting up data collection points at various supply chain nodes to collect basic information such as temperature, humidity, location, and time, and then storing and querying this information using traditional relational databases. However, traditional technologies have significant limitations when facing complex and ever-changing supply chain environments. On the one hand, simple data comparisons cannot effectively identify subtle damage and abnormal data jumps during transmission. On the other hand, traditional verification methods lack the ability to fuse and process multi-source heterogeneous data, making it difficult to construct a complete supply chain integrity assessment system. In existing technologies, due to the lack of intelligent data fusion mechanisms and adaptive verification strategies, the accuracy rate of agricultural product supply chain traceability data verification is generally low when facing complex supply chain structures. Furthermore, traditional verification algorithms have high computational complexity, making it difficult to meet real-time requirements. In other words, existing technologies suffer from low accuracy and computational inefficiency in verifying the integrity of agricultural product supply chain traceability data. Summary of the Invention
[0003] In view of this, the present invention provides a machine learning-based method for verifying the integrity of agricultural product supply chain traceability data, which can solve the technical problems of low accuracy and low computational efficiency in the existing technology for verifying the integrity of agricultural product supply chain traceability data.
[0004] This invention is implemented as follows: It provides a machine learning-based method for verifying the integrity of agricultural product supply chain traceability data. This method achieves accurate verification of the integrity of agricultural product supply chain traceability data by constructing a multi-dimensional evaluation matrix system and a liquid neural network data fusion verification model. The method includes: establishing an agricultural product supply chain data collection system; constructing a data classification and processing architecture; establishing an agricultural product damage prediction model; constructing a multi-dimensional evaluation matrix system; designing an agricultural product traceability verification game model; constructing an agricultural product data fusion verification model using a liquid neural network with embedded nonlinear activation functions; conducting small-scale verification experiments; and performing integrity verification calculations.
[0005] The steps for establishing an agricultural product supply chain data collection system specifically involve setting up multiple data collection nodes throughout the entire process of agricultural products from planting to sales, collecting environmental parameters, logistics information, and quality inspection data to form a raw dataset, and then dividing the collected data into three levels according to timeliness: real-time data, near-real-time data, and historical data.
[0006] Specifically, the step of constructing the data classification and processing architecture involves dividing the original dataset into three categories according to its complexity and importance: simplified data, rapid data, and detailed data. Simplified data contains basic identification information, rapid data contains key quality indicators, and detailed data contains complete traceability chain information.
[0007] Specifically, the step of establishing a damage prediction model for agricultural products involves analyzing historical damage data to construct a damage occurrence pattern model, identifying vulnerable links and damage patterns, and providing a benchmark for subsequent integrity verification.
[0008] The multi-dimensional evaluation matrix system includes a stability evaluation matrix, a loss risk matrix, a jump detection matrix, and an anomaly identification matrix. The stability evaluation matrix is established based on the stability performance of agricultural products in different links of the supply chain, and the loss risk matrix is constructed based on the probability and degree of loss of agricultural products in each link of the supply chain.
[0009] The jump detection matrix is established by analyzing the jump characteristics of data in the time series, the anomaly identification matrix is constructed by combining the characteristics of agricultural products and the rules of the supply chain, and the four evaluation matrices work together to form a three-dimensional integrity evaluation system.
[0010] The agricultural product traceability verification game model includes an upper-level model that aims to maximize data integrity and a lower-level model that aims to optimize computational efficiency. The optimal strategy for integrity verification is achieved through dual-level optimization, and the upper-level model and the lower-level model are associated through coupling terms.
[0011] The agricultural product data fusion verification model is constructed using a liquid neural network with embedded nonlinear activation functions to fuse multi-source data. The state parameters of the liquid neurons in the model are dynamically adjusted according to three parameters: the number of supply chain nodes, data timeliness, and data importance.
[0012] The liquid neural network comprises a liquid neuron reservoir layer, a nonlinear activation function embedding layer, a state update layer, and an output mapping layer. The liquid neuron reservoir layer contains randomly connected liquid neuron nodes, and the nonlinear activation function embedding layer embeds the hyperbolic tangent activation function into the state transition process of the liquid neurons.
[0013] Specifically, the steps of the small-scale verification experiment involve selecting agricultural product samples to collect data at different stages of the supply chain, and determining the weight coefficient range of the stability assessment matrix, the probability threshold range of the loss risk matrix, the sensitivity parameter range of the jump detection matrix, and the confidence threshold range of the anomaly identification matrix through experiments.
[0014] The state update layer dynamically updates the state parameters of the liquid neurons based on the input data and the feedback signal from the liquid neuron reservoir layer. The output mapping layer maps the state vector of the liquid neuron reservoir layer to the final data fusion result. The state parameters of the liquid neurons include membrane potential parameters, threshold parameters, and decay parameters.
[0015] The stability assessment matrix quantifies the data stability of agricultural products at each stage of the supply chain. Rows represent supply chain stages, including planting, processing, storage, transportation, and sales; columns represent data dimensions, including temperature, humidity, location, time, and quality data. The loss risk matrix assesses the risk level of data loss or damage in the agricultural product supply chain. Its structure is a two-dimensional mapping between supply chain stages and risk factors, including equipment failure, network interruption, human error, environmental change, and system anomaly. The jump detection matrix identifies anomalous jumps in agricultural product data over time. Rows correspond to data types, including temperature, humidity, location, time, and quality data; columns correspond to time windows of varying lengths. The anomaly identification matrix identifies abnormal data patterns that deviate from the normal supply chain patterns of agricultural products. This matrix employs a multi-layered structure: the first layer categorizes data dimensions (numerical, textual, temporal, and spatial); the second layer categorizes anomaly types (numerical anomalies, logical anomalies, temporal anomalies, and correlation anomalies).
[0016] Specifically, the step of performing integrity verification calculation involves inputting the fused data into the verification algorithm, calculating the comprehensive score of the stability assessment matrix, loss risk matrix, jump detection matrix, and anomaly identification matrix, and outputting the integrity verification result of the agricultural product supply chain traceability data.
[0017] This invention addresses the technical problems of low accuracy and computational efficiency in verifying the integrity of agricultural product supply chain traceability data by constructing a multi-dimensional evaluation matrix system based on machine learning and a liquid neural network data fusion model embedding nonlinear activation functions. The invention employs a comprehensive evaluation mechanism combining a stability evaluation matrix, a loss risk matrix, a jump detection matrix, and anomaly identification matrix, enabling accurate identification of data integrity issues from multiple dimensions. Simultaneously, a two-layer game optimization model is used to synergistically optimize verification accuracy and computational efficiency, effectively overcoming the low accuracy problem caused by a single evaluation standard in traditional methods. The liquid neural network designed in this invention possesses dynamic state adjustment capabilities, adaptively adjusting neuron parameters according to supply chain complexity and data quality, achieving efficient multi-source data fusion processing and significantly improving computational efficiency. In summary, this invention solves the technical problems of low accuracy and computational efficiency in verifying the integrity of agricultural product supply chain traceability data in existing technologies. Attached Figure Description
[0018] Figure 1 This is a flowchart of the method of the present invention.
[0019] Figure 2 This is a schematic diagram of the neural network structure of the agricultural product data fusion verification model involved in the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0021] like Figure 1 The diagram shown is a flowchart of a machine learning-based method for verifying the integrity of agricultural product supply chain traceability data, provided by this invention. This method includes the following steps:
[0022] S01. Establish an agricultural product supply chain data collection system, set up multiple data collection nodes throughout the entire process of agricultural products from planting to sales, collect environmental parameters, logistics information, and quality inspection data to form a raw dataset, and divide the collected data into three levels according to timeliness: real-time data, near-real-time data, and historical data.
[0023] S02. Construct a data classification and processing architecture, and divide the original dataset into three categories according to complexity and importance: simplified data, fast data, and detailed data. Simplified data contains basic identification information, fast data contains key quality indicators, and detailed data contains complete traceability chain information.
[0024] S03. Establish a damage prediction model for agricultural products. By analyzing historical damage data, construct a damage occurrence pattern model, identify vulnerable links and damage modes, and provide a benchmark for subsequent integrity verification.
[0025] S04. Construct a multi-dimensional evaluation matrix system, including a stability evaluation matrix, a loss risk matrix, a jump detection matrix, and an anomaly identification matrix. The stability evaluation matrix is established based on the stability performance of agricultural products in different links of the supply chain. The loss risk matrix is constructed based on the probability and degree of loss of agricultural products in each link of the supply chain. The jump detection matrix is established by analyzing the jump characteristics of data in the time series. The anomaly identification matrix is constructed by combining the characteristics of agricultural products and the rules of the supply chain.
[0026] S05. Design a game theory model for agricultural product traceability verification, including an upper-level model that aims to maximize data integrity and a lower-level model that aims to optimize computational efficiency. The optimal strategy for integrity verification is achieved through two-level optimization.
[0027] S06. A liquid neural network with embedded nonlinear activation function is used to construct an agricultural product data fusion verification model to fuse multi-source data. The state parameters of the liquid neurons in the model are dynamically adjusted according to three parameters: the number of supply chain nodes, data timeliness, and data importance.
[0028] S07. Conduct a small-scale verification experiment. Select 100 agricultural product samples and collect data at 3 different supply chain links. Through the experiment, determine the weight coefficient range of the stability assessment matrix as weight coefficient ∈ [0.1, 0.9), the probability threshold range of the loss risk matrix as probability threshold ∈ [0.05, 0.95], the sensitivity parameter range of the jump detection matrix as sensitivity parameter ∈ [1.2, 2.8], and the confidence threshold range of the anomaly identification matrix as confidence threshold ∈ (0.7, 0.98].
[0029] S08. Perform integrity verification calculation. Input the fused data into the verification algorithm. By calculating the comprehensive score of the stability assessment matrix, loss risk matrix, jump detection matrix and anomaly identification matrix, output the integrity verification result of agricultural product supply chain traceability data.
[0030] The stability assessment matrix is used to quantify the data stability of agricultural products at each stage of the supply chain. The rows of the matrix represent the supply chain stages, including planting, processing, storage, transportation, and sales. The columns represent data dimensions, including temperature, humidity, location, time, and quality data. The element values reflect the stability weight coefficients of the corresponding stage and dimension combinations. The weight coefficients are calculated through variance analysis and trend stability analysis of historical data. The weight coefficient values range from [0.1, 0.9], and the higher the value, the stronger the data stability of the corresponding stage and dimension combination.
[0031] The loss risk matrix is used to assess the risk level of data loss or damage in the agricultural product supply chain. The matrix structure is a two-dimensional mapping between supply chain links and risk factors. Supply chain links include planting, processing, storage, transportation, and sales. Risk factors include equipment failure factors, network interruption factors, human error factors, environmental change factors, and system anomaly factors. The element value represents the probability value of data loss caused by the corresponding risk factor in the corresponding link. The probability value is determined based on the statistics of historical loss events and the correlation analysis of risk factors. The probability value ranges from [0.05, 0.95]. The higher the value, the greater the possibility of data loss caused by the combination of the corresponding link and risk factor.
[0032] The jump detection matrix is used to identify abnormal jump behavior in agricultural product data over time. The rows of the matrix correspond to data types including temperature data, humidity data, location data, time data, and quality data. The columns correspond to time windows including 1-minute windows, 5-minute windows, 15-minute windows, 30-minute windows, and 60-minute windows. The element values are set as the jump threshold for each data type within the corresponding time window. The jump threshold values range from [1.2, 2.8]. When the actual data change exceeds the jump threshold, the jump detection mechanism is triggered. The higher the jump threshold value, the greater the allowable change range for the corresponding data type within the corresponding time window.
[0033] The anomaly identification matrix is used to identify abnormal data patterns that deviate from the normal supply chain patterns of agricultural products. The matrix adopts a multi-layer structure design. The first layer is the data dimension classification, including numerical dimension, text dimension, time dimension and spatial dimension. The second layer is the anomaly type classification, including numerical anomaly, logical anomaly, time sequence anomaly and correlation anomaly. The element value is the identification confidence threshold of the corresponding dimension and anomaly type combination. The confidence threshold value range is confidence threshold ∈ [0.7, 0.98]. The higher the value, the more stringent the requirement for the accuracy of anomaly identification of the corresponding dimension and anomaly type combination.
[0034] The objective function of the upper-level model is to maximize the accuracy of data integrity verification. The objective function includes a data coverage parameter, a verification accuracy parameter, a timeliness weight parameter, and a stability coefficient parameter. The data coverage parameter is derived from the proportion of valid data in the original dataset to the total data. The verification accuracy parameter is derived from the accuracy statistics of historical verification results. The timeliness weight parameter is derived from the importance scores of data with different timeliness. The stability coefficient parameter is derived from the average weight value of the stability evaluation matrix. The product of the logarithmic values of the first term, the data coverage parameter and the verification accuracy parameter, reflects the nonlinear promoting effect of data coverage on verification accuracy. The product of the second term, the timeliness weight parameter and the negative exponent of the stability coefficient parameter, reflects the diminishing marginal contribution of timeliness to accuracy in a stable environment. The product of the sine function of the third term, the data coverage parameter, the verification accuracy parameter, and the timeliness weight parameter, represents the periodic optimization effect produced by multi-parameter coupling. The constraints are that the data processing time does not exceed a set threshold and the verification accuracy is not less than 90%.
[0035] The objective function of the lower-level model is to minimize computational resource consumption. The objective function includes processing time parameters, memory usage parameters, network bandwidth parameters, and computational complexity parameters. The processing time parameter is derived from the actual time consumption statistics of each step of the system execution. The memory usage parameter is derived from the memory usage monitoring during program operation. The network bandwidth parameter is derived from the bandwidth usage measurement during data transmission. The computational complexity parameter is derived from the theoretical analysis of the algorithm's time and space complexity. The product of the square root of the first term of the objective function, processing time parameter and memory usage parameter, reflects the nonlinear correlation between time and memory resource consumption. The product of the cosine function of the second term, network bandwidth parameter and computational complexity parameter, reflects the periodic resource demand changes of network transmission under different computational complexities. The product of the logarithmic values of the third term, processing time parameter, memory usage parameter, and network bandwidth parameter, represents the composite effect of the three resource synergistic optimizations. The constraints are that the system response time does not exceed 3 seconds and the resource utilization rate does not exceed 80%.
[0036] The coupling term is reflected in the negative correlation between the verification accuracy parameter of the upper-level model and the processing time parameter of the lower-level model. The verification accuracy parameter changes with the square inverse of the processing time parameter. When the lower-level model is optimized to reduce the processing time parameter, the verification accuracy parameter of the upper-level model will be improved. The output result of the processing time parameter is used for the calculation of the coupling term of the objective function of the upper-level model, and the output result of the verification accuracy parameter is used for the performance evaluation of the constraint conditions of the lower-level model.
[0037] The specific structure of the agricultural product data fusion verification model is a liquid neural network with embedded nonlinear activation functions. It includes a liquid neuron reservoir layer, a nonlinear activation function embedding layer, a state update layer, and an output mapping layer. The liquid neuron reservoir layer contains 500 randomly connected liquid neuron nodes. The nonlinear activation function embedding layer embeds the hyperbolic tangent activation function into the state transition process of the liquid neurons. The state update layer dynamically updates the state parameters of the liquid neurons according to the input data and the feedback signal from the liquid neuron reservoir layer. The output mapping layer maps the state vector of the liquid neuron reservoir layer to the final data fusion result. The state parameters of the liquid neurons include membrane potential parameters, threshold parameters, and decay parameters.
[0038] The steps for establishing the training dataset for the agricultural product data fusion verification model specifically include collecting historical traceability data covering different agricultural product types, different supply chain links, and different seasonal periods as positive samples; manually constructing negative sample data containing missing data, data errors, and timestamp anomalies; standardizing and feature-engineering the collected data; and dividing the dataset into training, validation, and test sets in an 8:1:1 ratio to ensure that the distribution of agricultural product types and data quality is uniform in each subset. The total size of the training dataset is 50,000 samples, including 30,000 positive samples and 20,000 negative samples.
[0039] The specific steps of training the agricultural product data fusion verification model include: initializing model parameters by setting the learning rate to 0.001 using the AdamW optimizer; constructing a composite loss function using the cross-entropy loss function combined with the data integrity loss function; training the model using batch training and monitoring model performance on the validation set; triggering an early stopping mechanism when the validation set loss does not decrease for 10 consecutive epochs; evaluating the generalization performance of the model on the test set and saving the optimal model parameters after training is completed; setting the training batch size to 64 and the maximum number of training epochs to 200.
[0040] The liquid neuron state adjustment function is used to adjust the state parameters of the liquid neuron in the agricultural product data fusion verification model. The function calculates a comprehensive state score based on four data: the number of supply chain nodes, data timeliness, data integrity rate, and historical verification accuracy rate. The comprehensive state score is derived from the weighted average of the four input parameters. The number of supply chain nodes is derived from the total number of nodes in the current supply chain path. The data timeliness is derived from the time difference between the current data and the collection time. The data integrity rate is derived from the proportion of non-missing data in the current dataset to the total data. The historical verification accuracy rate is derived from the statistical average of the accuracy rates of historical verification results. The comprehensive state score is used for subsequent adjustment calculations of the liquid neuron state parameters.
[0041] Specifically, when the comprehensive state score S ∈ [0, 0.2), a low-sensitivity mode is used to adjust the state parameters of the liquid neuron. In this mode, the membrane potential parameter of the liquid neuron is set to a lower value to reduce the neuron's sensitivity to input signals, the threshold parameter is set to a higher value to increase the threshold for neuron activation, and the decay parameter is set to a larger value to accelerate the decay rate of state information. The low-sensitivity mode is suitable for situations with poor data quality or simple supply chain structures, reducing the activity level of the liquid neuron to avoid the amplification and propagation of erroneous information. When the comprehensive state score S ∈ [0.2, 0.4), a medium-low sensitivity mode is used to adjust... For fluid-state neurons, the membrane potential parameter is moderately increased to enhance signal reception, the threshold parameter is moderately decreased to allow more neurons to participate in computation, and the decay parameter is moderately reduced to prolong the retention time of state information. In the low-to-medium sensitivity mode, the network's information processing capacity is appropriately enhanced while maintaining stability. When the overall state score S ∈ [0.4, 0.6), the fluid-state neuron state parameters are adjusted using a medium sensitivity mode, with the membrane potential parameter set to a medium value to balance response speed and stability, the threshold parameter set to a medium value to achieve a moderate activation frequency, and the decay parameter set to a medium value to maintain a reasonable... The medium sensitivity mode is suitable for standard processing scenarios where both data quality and supply chain complexity are at a medium level. When the comprehensive state score S∈[0.6, 0.8), the medium-high sensitivity mode is used to adjust the state parameters of the liquid neurons. The membrane potential parameter is further increased to enhance the neurons' ability to capture subtle signal changes, the threshold parameter is further decreased to activate more neurons to participate in complex pattern recognition, and the decay parameter is further decreased to maintain state memory for a longer period of time. The medium-high sensitivity mode is suitable for processing needs of high-quality data and complex supply chain structures, and captures more useful information by increasing the activity of the liquid neurons. When the comprehensive state score S∈[0.8, 1.0], the high sensitivity mode is used to adjust the state parameters of the liquid neurons. The membrane potential parameter is set to the highest value to achieve the maximum response capability to the input signal, the threshold parameter is set to the lowest value to activate the most number of neurons, and the decay parameter is set to the minimum value to achieve the longest state retention time. The high sensitivity mode is suitable for the precise processing requirements of high-quality data and high-complexity supply chains, and achieves the optimal data fusion effect by maximizing the computing power of the liquid neurons. The comprehensive state score output is used to guide the allocation of weights of each evaluation matrix in subsequent verification calculation steps.
[0042] The small-scale verification experiment involved selecting 100 samples each of three major categories of agricultural products: vegetables, fruits, and grains. Data was collected from each sample at the planting, processing, and sales stages. Five types of data were collected for each sample: temperature, humidity, location, time, and quality. During the experiment, 10% missing data, 5% data errors, and 3% timestamp anomalies were artificially introduced to simulate data problems in a real supply chain. Statistical analysis determined the distribution of weight coefficients for each stage and data dimension combination in the stability assessment matrix. The average weight coefficient for the planting stage was 0.65, for the processing stage it was 0.72, for the sales stage it was 0.58, for temperature data it was 0.71, for humidity data it was 0.69, for location data it was 0.63, for time data it was 0.75, and for quality data it was 0.68.
[0043] In the small-scale verification experiment, the probability parameters of the loss risk matrix were determined by statistically analyzing the frequency of data loss caused by each risk factor in each stage. In the planting stage, the loss probability of equipment failure was 0.15, the loss probability of network interruption was 0.22, the loss probability of human operation was 0.18, the loss probability of environmental change was 0.25, and the loss probability of system anomaly was 0.12. In the processing stage, the loss probabilities of each risk factor were 0.18, 0.19, 0.23, 0.16, and 0.21, respectively. In the sales stage, the loss probabilities of each risk factor were 0.13, 0.17, 0.26, 0.14, and 0.19, respectively.
[0044] In the small-scale verification experiment, the sensitivity parameters of the jump detection matrix were determined by analyzing the statistical changes in the data types within different time windows. The jump thresholds for temperature data were 1.8 within a 1-minute window, 2.1 within a 5-minute window, 2.3 within a 15-minute window, 2.5 within a 30-minute window, and 2.7 within a 60-minute window. The jump thresholds for humidity data within each time window were 1.6, 1.9, 2.2, 2.4, and 2.6, respectively. The jump thresholds for location data within each time window were 1.4, 1.7, 2.0, 2.2, and 2.4, respectively. The jump thresholds for time data within each time window were 1.2, 1.5, 1.8, 2.0, and 2.2, respectively. The jump thresholds for quality data within each time window were 2.0, 2.3, 2.5, 2.7, and 2.8, respectively.
[0045] In the small-scale verification experiment, the confidence threshold of the anomaly identification matrix was determined by analyzing the identification accuracy of each dimension and anomaly type combination. The confidence threshold for the numerical dimension and numerical anomaly combination was 0.85, the confidence threshold for the combination with logical anomaly was 0.78, the confidence threshold for the combination with temporal anomaly was 0.82, and the confidence threshold for the combination with correlation anomaly was 0.79. The confidence thresholds for the text dimension and each anomaly type combination were 0.76, 0.88, 0.73, and 0.81, respectively. The confidence thresholds for the time dimension and each anomaly type combination were 0.72, 0.75, 0.92, and 0.77, respectively. The confidence thresholds for the spatial dimension and each anomaly type combination were 0.83, 0.71, 0.86, and 0.94, respectively. These confidence thresholds are used to guide the identification and processing of abnormal data during the integrity verification calculation process.
[0046] The specific implementation methods of the above steps are described in detail below.
[0047] The specific implementation of step S01 involves establishing a data collection system covering the entire lifecycle of agricultural products. This system utilizes IoT sensor technology and distributed data acquisition principles to achieve multi-node data acquisition. First, environmental monitoring sensors are deployed at the planting stage to collect environmental parameters such as soil temperature and humidity, light intensity, and rainfall. The collection frequency is set to once every 5 minutes to ensure the continuity and integrity of environmental data. Second, temperature, humidity, and pressure sensors are installed at the processing stage to collect process parameters. The accuracy requirement is that the temperature error does not exceed ±0.5℃ and the humidity error does not exceed ±2% relative humidity. Then, GPS locators and RFID tags are deployed at the storage and transportation stages to collect location information and logistics trajectories. The positioning accuracy requirement is at the meter level to ensure the traceability of agricultural product flow. Finally, at the sales stage, quality inspection data and sales information are collected through QR code scanning and weighing equipment. Quality inspection data includes key indicators such as weight, appearance, and quality grade. The collected raw data is classified according to timeliness requirements: real-time data refers to data transmitted immediately after collection, with a delay of no more than 1 second; near-real-time data refers to data transmitted within 5 minutes of collection; and historical data refers to stored data older than 1 hour. Timestamps are used to ensure the temporal integrity of the data.
[0048] The specific implementation of step S02 involves constructing a three-tier classification and processing architecture based on data importance and complexity. This architecture employs data preprocessing techniques and hierarchical storage principles to achieve effective data management. First, the original dataset undergoes data cleaning and format standardization, removing duplicate data, correcting erroneous data, and supplementing missing data. The accuracy of data cleaning is required to reach over 98%. Second, the data is classified according to its business importance. Simplified data includes basic identification information for agricultural products, such as product number, production date, and place of origin, accounting for approximately 30% of the total data, with a storage period of 1 year. Rapid data includes key quality indicators such as pesticide residues, heavy metal content, and microbial indicators, accounting for approximately 50% of the total data, with a storage period of 3 years. Detailed data includes complete traceability chain information such as planting records, processing technology, transportation conditions, and test reports, accounting for approximately 20% of the total data, with a storage period of 5 years. Then, a data indexing mechanism is established, using hash tables and balanced binary tree structures to achieve fast data retrieval, with query response time controlled within 100 milliseconds. Finally, a data backup and recovery mechanism was established, and distributed storage technology was used to ensure the reliability and availability of data. The data redundancy was set to 3 copies, and the recovery time was controlled within 1 hour.
[0049] The specific implementation of step S03 involves establishing a predictive model based on historical damage data. This model employs time series analysis and machine learning algorithms to identify and predict damage patterns. First, historical damage data of agricultural products at different stages of the supply chain is collected, including damage type, damage severity, occurrence time, and environmental conditions. The historical data collection period is the past three years, with a sample size of no less than 10,000. Second, feature extraction and data preprocessing are performed on the historical damage data, extracting key features such as temperature variation amplitude, humidity variation trend, transportation duration, and storage conditions. The feature dimension is controlled to within 50 to avoid overfitting due to excessive dimensionality. Then, a Long Short-Term Memory (LSTM) network algorithm is used to construct the damage prediction model. This algorithm can effectively handle long-term dependencies in time series data. The model contains three hidden layers, each with 128 neurons, and the training rounds are set to 200. Next, 80% of the historical data is used as the training set, and 20% as the test set for model training and validation. Cross-validation is used during training to avoid overfitting, and the validation accuracy is required to reach above 85%. Finally, a damage risk assessment mechanism was established, which quantifies the damage risk of different links in the supply chain based on the prediction results. The risk level is divided into three levels: low risk, medium risk, and high risk, with risk thresholds set at 0.3, 0.6, and 0.8, respectively.
[0050] The specific implementation of step S04 involves constructing a matrix system based on multi-dimensional assessment. This system employs statistical analysis methods and weight allocation algorithms to quantitatively assess the risks at each stage of the supply chain. First, a stability assessment matrix is established. This matrix uses analysis of variance and trend stability analysis to calculate the stability weight coefficients of each stage and data dimension combination. The matrix has a 5×5 dimension, with rows representing the five supply chain stages of planting, processing, storage, transportation, and sales, and columns representing the five data dimensions of temperature, humidity, location, time, and quality. The weight coefficients are determined by calculating the standard deviation and coefficient of variation of historical data, with values controlled between 0.1 and 0.9. Second, a loss risk matrix is established. This matrix uses risk factor correlation analysis and historical loss event statistics to determine the loss probability of different risk factors at each stage. The matrix also has a 5×5 dimension, with rows representing the five supply chain stages and columns representing the five risk factors of equipment failure, network interruption, human error, environmental change, and system anomaly. The probability values are calculated based on Bayesian statistical methods, with values controlled between 0.05 and 0.95. Then, a jump detection matrix is established. This matrix uses a sliding window algorithm and anomaly detection methods to identify jump behaviors in data over time series. The matrix is 5×5, with rows representing five data types and columns representing five time windows: 1 minute, 5 minutes, 15 minutes, 30 minutes, and 60 minutes. The jump threshold is determined by calculating the statistical distribution of the data change rate. Jump detection is triggered when the actual change rate exceeds 1.5 times the threshold. The sensitivity parameter is controlled between 1.2 and 2.8. Finally, an anomaly identification matrix is established. This matrix uses a multi-level classification structure and confidence assessment methods to identify abnormal data patterns that deviate from normal patterns. The matrix adopts a 4×4 two-level structure. The first level includes four data dimensions: numerical, textual, temporal, and spatial. The second level includes four anomaly types: numerical anomalies, logical anomalies, temporal anomalies, and correlation anomalies. The confidence threshold is determined through cross-validation and accuracy statistics, with a value range between 0.7 and 0.98.
[0051] The specific implementation of step S05 involves designing a game theory model based on a two-layer optimization structure. This model uses game theory principles and a multi-objective optimization algorithm to optimize the verification strategy. First, an upper-layer model is constructed. This model aims to maximize the accuracy of data integrity verification. The objective function includes four parameters: data coverage, verification accuracy, timeliness weight, and stability coefficient. Data coverage is determined by calculating the proportion of valid data to the total data. Verification accuracy is determined by statistically analyzing the accuracy of historical verification results. The timeliness weight is determined based on the importance scores of different timeliness data. The stability coefficient is determined by calculating the average weight value of the stability evaluation matrix. The objective function uses a non-linear combination form. The first term reflects the non-linear promoting effect of data coverage on verification accuracy; the second term reflects the diminishing marginal contribution effect of timeliness in a stable environment; and the third term represents the periodic optimization effect generated by the coupling of multiple parameters. Secondly, a lower-level model is constructed, aiming to minimize computational resource consumption. The objective function includes four parameters: processing time, memory usage, network bandwidth, and computational complexity. Processing time is determined through statistical analysis of the actual time consumed by each step of the system execution; memory usage is determined through monitoring memory usage during program execution; network bandwidth is determined through bandwidth usage measurement during data transmission; and computational complexity is determined through theoretical analysis of the algorithm's time and space complexity. Then, a coupling mechanism between the upper and lower-level models is established. There is a negative correlation between the verification accuracy parameter of the upper-level model and the processing time parameter of the lower-level model. Optimizing the lower-level model to reduce processing time will improve the verification accuracy of the upper-level model. This coupling relationship is described by the inverse square function, ensuring coordinated optimization of the two models. Finally, a genetic algorithm is used to solve the bi-level optimization problem. The algorithm parameters are set as follows: population size 100, crossover probability 0.8, mutation probability 0.01, number of generations 200. The constraints are: data processing time does not exceed a set threshold, verification accuracy is not less than 90%, system response time does not exceed 3 seconds, and resource utilization does not exceed 80%.
[0052] The specific implementation of step S06 involves constructing a data fusion and verification model based on a liquid neural network. This model employs the principles of a liquid state machine and a nonlinear activation function to achieve effective fusion and verification of multi-source data. First, a liquid neuron reservoir layer is designed, containing 500 randomly connected liquid neuron nodes. Each neuron has three state parameters: membrane potential, threshold, and decay. The connection weights between neurons are randomly initialized using a Gaussian distribution, and the connection density is set to 10% to ensure network sparsity and computational efficiency. Second, a nonlinear activation function embedding layer is designed. This layer embeds the hyperbolic tangent activation function into the state transition process of the liquid neurons. The slope parameter of the activation function is set to 1.0, and the bias parameter is set to 0.0. The activation function can effectively handle nonlinear mapping relationships, improving the model's expressive power. Then, a state update layer is designed. This layer dynamically updates the state parameters of the liquid neurons based on the input data and the feedback signal from the reservoir layer. The update rule adopts the form of a differential equation, with a time constant of 10 milliseconds and an update frequency of 100 Hz to ensure the smoothness and continuity of state changes. Next, an output mapping layer is designed. This layer maps the state vectors of the reserve pool layer to the final data fusion result. The mapping matrix is obtained through least squares training. The output dimension is determined according to the actual application requirements, usually set as the dimension of the number of classifications or the regression target. Finally, a liquid neuron state regulation mechanism is established. A comprehensive state score is calculated based on four parameters: the number of supply chain nodes, data timeliness, data completeness, and historical verification accuracy. The score ranges from 0 to 1, corresponding to different sensitivity modes: low sensitivity mode is suitable for scores of 0 to 0.2, medium sensitivity mode is suitable for scores of 0.4 to 0.6, and high sensitivity mode is suitable for scores of 0.8 to 1.0.
[0053] The specific implementation of step S07 involves conducting a small-scale verification experiment. This experiment uses the controlled variable method and statistical analysis to verify the effectiveness and reliability of the parameters of each evaluation matrix. First, 100 samples from each of the three major categories of agricultural products—vegetables, fruits, and grains—are selected as experimental subjects. Sample selection follows the principle of random sampling to ensure the representativeness and diversity of the samples. Each category of agricultural products includes samples from different varieties, origins, and seasons. Second, data collection points are set up at three key stages: planting, processing, and sales. For each sample, five types of data—temperature, humidity, location, time, and quality—are collected at each stage. Data collection equipment includes temperature and humidity sensors, GPS locators, RFID tags, and quality testing instruments. The collection accuracy and frequency are performed according to the requirements of step S01. Then, data problems are artificially introduced to simulate a real supply chain environment: a 10% data loss rate is introduced to simulate data loss due to equipment failure or network interruption; a 5% data error rate is introduced to simulate insufficient sensor accuracy or human error; and a 3% timestamp anomaly rate is introduced to simulate system clock asynchrony or data transmission delay. Next, statistical analysis was used to determine the parameter distribution patterns of each evaluation matrix. The weight coefficients of the stability evaluation matrix were determined by calculating the variance and coefficient of variation of each stage and data dimension combination. The average weight coefficient for the planting stage was 0.65, for the processing stage it was 0.72, and for the sales stage it was 0.58. The average weight coefficient for the temperature data dimension was 0.71, for the humidity data dimension it was 0.69, for the location data dimension it was 0.63, for the time data dimension it was 0.75, and for the quality data dimension it was 0.68. The probability parameters of the loss risk matrix were determined by statistically analyzing the frequency of data loss caused by each risk factor. In the planting stage, the loss probabilities of the five risk factors—equipment failure, network interruption, human error, environmental change, and system anomaly—were 0.15, 0.22, 0.18, 0.25, and 0.12, respectively. The sensitivity parameters of the jump detection matrix were determined by analyzing the statistical changes in the data types within different time windows. The jump thresholds for temperature data within 1-minute, 5-minute, 15-minute, 30-minute, and 60-minute windows were 1.8, 2.1, 2.3, 2.5, and 2.7, respectively. The confidence thresholds of the anomaly identification matrix were determined by analyzing the identification accuracy of combinations of various dimensions and anomaly types. The confidence thresholds for combinations of numerical dimensions with numerical anomalies, logical anomalies, temporal anomalies, and correlation anomalies were 0.85, 0.78, 0.82, and 0.79, respectively.
[0054] The specific implementation of step S08 involves performing integrity verification calculations. This calculation employs a multi-matrix comprehensive evaluation method and a weighted fusion algorithm to quantitatively assess the integrity of agricultural product supply chain traceability data. First, the data processed in step S06 is input into the verification algorithm. The algorithm uses a parallel computing architecture to simultaneously calculate the evaluation results of the stability assessment matrix, loss risk matrix, jump detection matrix, and anomaly identification matrix, with the calculation time controlled within one second to ensure real-time verification. Next, the stability assessment matrix is calculated by comparing current data with historical data to determine the stability score for each link and data dimension combination. The score calculation formula is based on weighting coefficients and the degree of data variation, with scores ranging from 0 to 1; higher values indicate better data stability. Then, the loss risk matrix is calculated by analyzing the risk factor status of the current supply chain links and calculating the probability of data loss caused by different risk factors at each link. The probability calculation is based on historical statistical data and current environmental conditions. When the probability exceeds 0.5, a risk warning mechanism is triggered. Next, the jump detection matrix is calculated, using a sliding window algorithm to detect anomalous jumps in the data over time. When the data change exceeds a jump threshold, it is marked as an anomaly. The sensitivity of jump detection is dynamically adjusted based on the data type and time window. Then, the anomaly identification matrix is calculated, using a multi-level classification algorithm to identify anomalous data patterns deviating from normal trends. When the confidence level of anomaly identification exceeds a set threshold, it is marked as anomalous data. The anomaly handling strategy includes data correction, data labeling, and data exclusion. Finally, a weighted average method is used to calculate the comprehensive score of the four matrices. The weight allocation is determined based on the experimental results of step S07: the stability assessment matrix has a weight of 0.3, the loss risk matrix has a weight of 0.25, the jump detection matrix has a weight of 0.25, and the anomaly identification matrix has a weight of 0.2. The comprehensive score ranges from 0 to 1. A score above 0.8 indicates good data integrity, a score below 0.6 indicates problems with data integrity, and a score between 0.6 and 0.8 indicates average data integrity.
[0055] Nong Ru Figure 2As shown, the detailed structure of the product data fusion verification model includes four core components: a liquid neuron reservoir layer, a nonlinear activation function embedding layer, a state update layer, and an output mapping layer. The liquid neuron reservoir layer, as the core computational unit of the model, contains 500 randomly connected liquid neuron nodes. Each neuron simulates the membrane potential change process of a biological neuron and has three state variables: membrane potential parameter, threshold parameter, and decay parameter. Neurons form a complex network topology through sparse connections. The connection weights follow a Gaussian distribution with a mean of 0 and a variance of 0.1, and the connection density is controlled at around 10% to ensure sufficient computational power while avoiding increased computational complexity due to over-connection. The nonlinear activation function embedding layer incorporates the hyperbolic tangent activation function into the state transition process of the liquid neuron. The mathematical form of the hyperbolic tangent function is that the output value equals the hyperbolic tangent of the input value. This function exhibits an S-shaped curve characteristic, capable of mapping any real number to the interval between -1 and +1. The derivative of the function reaches its maximum value near zero, which is beneficial for gradient propagation and parameter optimization. Compared to traditional step functions and linear functions, the hyperbolic tangent function can better handle nonlinear mapping relationships, improving the model's ability to express complex data patterns. The state update layer is responsible for dynamically updating the state parameters of the liquid neuron based on the input data and feedback signals from the reservoir layer. The update process follows the dynamic evolution law of differential equations. The update of membrane potential parameters considers the combined influence of input current, leakage current, and synaptic current. The time constant is set to 10 milliseconds, and the update frequency is set to 100 Hz to ensure the continuity and smoothness of state changes. The output mapping layer maps the high-dimensional state vector of the reservoir layer to the low-dimensional data fusion result. The mapping process adopts a linear transformation method. The mapping matrix is trained by a supervised learning method. The training objective is to minimize the mean square error between the predicted output and the true label. The dimension of the mapping matrix is 500 × the output dimension, where the output dimension is determined according to the specific application requirements.
[0056] The detailed steps for establishing the training dataset for the agricultural product data fusion and verification model include four stages: data collection, data construction, data preprocessing, and data partitioning. In the data collection stage, authentic historical traceability data was obtained from multiple agricultural product production enterprises and supply chain management companies. The data covers four major categories of agricultural products: vegetables, fruits, grains, and livestock products, spanning the past five years and covering major agricultural production areas nationwide. The data content includes complete supply chain information such as planting records, processing techniques, storage conditions, transportation routes, quality inspection reports, and sales information. A total of 30,000 positive samples were collected, with each sample containing an average of 50 data fields. In the data construction stage, negative samples containing various data problems were manually constructed. Construction methods included randomly deleting some data fields to simulate missing data, randomly modifying data values to simulate data errors, randomly adjusting timestamps to simulate time anomalies, and randomly inserting unreasonable data to simulate logical anomalies. A total of 20,000 negative samples were constructed, with the ratio of negative samples to positive samples controlled at 2:3 to ensure the balance of the training data. The data preprocessing stage involves standardization and feature engineering of the collected and constructed data. Standardization includes data format unification, unit conversion, missing value imputation, and outlier handling. Feature engineering includes feature extraction, feature selection, feature transformation, and feature combination, ultimately forming a standardized dataset with 100 feature dimensions, where each feature's value is standardized to between 0 and 1. In the data partitioning stage, the standardized dataset is randomly divided into training, validation, and test sets in an 8:1:1 ratio. The training set contains 40,000 samples for model parameter training, the validation set contains 5,000 samples for model performance monitoring and hyperparameter tuning, and the test set contains 5,000 samples for model generalization performance evaluation. This partitioning process ensures a uniform distribution of agricultural product types and data quality across all subsets, avoiding the impact of data bias on model performance.
[0057] The reason why the agricultural product data fusion verification model is suitable for solving the technical problem of verifying the integrity of traceability data in the agricultural product supply chain lies in its unique liquid computing principle and multi-source data fusion capability. Traditional data integrity verification methods mainly adopt static rule matching and simple statistical analysis, such as data consistency detection based on hash verification and anomaly detection based on threshold judgment. These methods have obvious limitations when dealing with the complex and ever-changing data environment of the agricultural product supply chain. As an emerging computing model, liquid neural networks have three significant characteristics: dynamic time-varying, nonlinear mapping, and short-term memory. They can effectively handle the data characteristics of agricultural product supply chains that are highly time-series, highly random, and complexly correlated. Compared with traditional feedforward neural networks, liquid neural networks do not require strict hierarchical structures and backpropagation training. Instead, they achieve information processing through the dynamic evolution of liquid neurons in a reservoir. This computing method is more suitable for processing continuously changing time-series data and non-stationary random data. Compared with rule-based expert systems, liquid neural networks can automatically learn the implicit patterns and relationships in the data without the need for manually formulating complex rule bases, and have stronger adaptability and generalization ability. The model enhances its ability to express complex data patterns by embedding nonlinear activation functions, achieves comprehensive evaluation of different types of data problems through multi-dimensional evaluation matrices, and balances verification accuracy and computational efficiency through two-level game optimization. These technical features enable the model to effectively address the technical challenges in verifying the integrity of agricultural product supply chain traceability data.
[0058] The key technical ideas of this invention mainly include dynamic data fusion technology based on liquid neural networks, comprehensive evaluation technology based on multi-dimensional matrices, and optimization strategy technology based on a two-layer game model. Dynamic data fusion technology based on liquid neural networks has significant technical advantages over traditional static data processing methods. Traditional methods typically use fixed weighting coefficients and preset fusion rules to process multi-source data, which cannot adapt to the dynamic changes in data characteristics and real-time changes in environmental conditions in the agricultural product supply chain. Liquid neural networks, however, achieve adaptive data fusion through the dynamic state evolution of liquid neurons in a reservoir, and can dynamically adjust the fusion strategy according to the timeliness, completeness, and importance of the data, significantly improving the accuracy and robustness of data fusion. Comprehensive evaluation technology based on multi-dimensional matrices has a more comprehensive evaluation capability than single-indicator evaluation methods. Traditional methods usually only focus on the single dimension of data consistency or completeness, making it difficult to fully reflect the complexity and diversity of data problems in the agricultural product supply chain. Multi-dimensional matrix evaluation technology, by constructing an evaluation matrix with four dimensions—stability, risk, abruptness, and anomaly—can comprehensively evaluate the integrity status of data from different perspectives, effectively identify various types of data problems, and improve the accuracy and reliability of the evaluation results. Compared to single-objective optimization methods, the optimization strategy based on a two-level game model offers better balance. Traditional methods typically consider only one objective—verification accuracy or computational efficiency—making it difficult to find the optimal balance between the two. The two-level game model, through a game mechanism where the upper-level model pursues maximizing verification accuracy while the lower-level model pursues optimal computational efficiency, optimizes computational resource allocation while ensuring verification quality, achieving an optimal balance between verification effectiveness and computational cost. The synergistic effect of these three key technological approaches has produced significant technological advantages over existing technologies. Liquid neural networks provide powerful data processing capabilities, multi-dimensional matrices offer a comprehensive evaluation framework, and the two-level game model provides an optimized strategy mechanism. These three elements work together to form a complete technical system that not only improves the accuracy and efficiency of verifying the integrity of agricultural product supply chain traceability data but also enhances the system's adaptability to complex environments and its ability to handle anomalies, providing more reliable and efficient technical support for agricultural product quality and safety traceability.
[0059] It should be noted that the agricultural product supply chain involves multiple stages, including planting, processing, storage, transportation, and sales. The data generated at each stage varies significantly in format, accuracy, and collection frequency, making it difficult for traditional data processing methods to effectively integrate these heterogeneous data sources. This invention establishes a data classification and processing architecture, dividing the original dataset into three categories based on complexity and importance: simplified data, rapid data, and detailed data. It employs a liquid neural network with a reservoir layer containing 500 randomly connected liquid neuron nodes, enabling adaptive processing of input data of different types and structures. Furthermore, a nonlinear activation function embedding layer integrates the hyperbolic tangent activation function into the state transition process of the liquid neurons, effectively solving the problem of feature extraction and fusion of heterogeneous data. The supply chain structures of different agricultural products vary greatly in complexity, ranging from simple direct sales by farmers to complex multi-level distribution networks. Traditional fixed-parameter validation methods cannot adapt to this dynamic change. The liquid neuron state adjustment function designed in this invention calculates a comprehensive state score based on four parameters: the number of supply chain nodes, data timeliness, data integrity rate, and historical verification accuracy. Based on the score in different ranges, it dynamically adjusts the membrane potential, threshold, and attenuation parameters of the liquid neuron using five modes: low sensitivity, low-to-medium sensitivity, medium sensitivity, medium-to-high sensitivity, and high sensitivity. This achieves adaptive adjustment of the verification strategy according to the complexity of the supply chain, ensuring that false alarms caused by oversensitivity are avoided in simple supply chains and that sufficient detection accuracy is maintained in complex supply chains.
[0060] Specifically, the principle of this invention is as follows: This invention can solve the technical problems of low accuracy and low computational efficiency in verifying the integrity of agricultural product supply chain traceability data. Its fundamental principle lies in establishing a multi-level intelligent data integrity assessment system and an adaptive computational optimization mechanism. First, the fundamental reason why traditional single verification standards cannot comprehensively assess data integrity is that agricultural product supply chain data has multi-dimensional characteristics and complex spatiotemporal correlations. This invention constructs a comprehensive assessment system with four dimensions: a stability assessment matrix, a loss risk matrix, a jump detection matrix, and an anomaly identification matrix. This system can conduct a three-dimensional assessment from multiple perspectives, including data stability, risk probability, temporal anomalies, and pattern deviations, effectively capturing integrity issues missed by traditional methods. Second, the fundamental reason for low computational efficiency is the lack of adaptive optimization mechanisms in traditional algorithms. The two-layer game optimization model designed in this invention establishes a dynamic balance mechanism between verification accuracy and computational efficiency through collaborative optimization, where the upper-layer model pursues maximizing verification accuracy and the lower-layer model pursues minimizing computational resource consumption. This achieves the minimization of computational complexity while ensuring verification accuracy. Furthermore, the liquid neural network employed in this invention possesses memory capabilities and dynamic adaptability, automatically adjusting the membrane potential, threshold, and decay parameters of the liquid neurons based on the comprehensive state score. This allows the network to adopt appropriate sensitivity modes for supply chain scenarios of varying complexity, avoiding overfitting in low-quality data environments while ensuring accurate processing capabilities in highly complex scenarios. Finally, through the design of embedding nonlinear activation functions, the liquid neural network can handle the nonlinear characteristics of agricultural product data, improving the fusion effect of multi-source heterogeneous data, thereby achieving a dual improvement in verification accuracy and computational efficiency.
[0061] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.
[0062] In this embodiment, the specific implementation of steps S01-S03 is the same as described above, and will not be repeated in detail here.
[0063] The specific implementation of step S04 is to construct a matrix system based on multi-dimensional evaluation, wherein the stability evaluation matrix M s The mathematical expression is:
[0064]
[0065] In the formula, w ij Let w be the stability weighting coefficient for the combination of the i-th supply chain link and the j-th data dimension. The value of i ranges from 1 to 5, corresponding to the planting, processing, storage, transportation, and sales links respectively. The value of j ranges from 1 to 5, corresponding to the temperature, humidity, location, time, and quality data dimensions respectively. ij The calculation formula is:
[0066]
[0067] In the formula, λ represents the standardized variance of the j-th data dimension in the i-th stage; ij Let be the absolute value of the trend slope of the j-th data dimension in the i-th stage; α is the weight of the variance term, with a value of 0.6; β is the weight of the trend term, with a value of 0.4. The calculation formula is:
[0068]
[0069] In the formula, σ ij σ represents the original variance value of the j-th data dimension in the i-th stage; max σ represents the maximum variance across all combinations of elements and dimensions. min This represents the minimum variance value among all combinations of elements and dimensions. Where the original variance value σ is... ij The calculation formula is:
[0070]
[0071] In the formula, x ijk For the i-th stage, the k-th observation in the j-th data dimension; λ represents the corresponding sample mean; n is the sample size. The absolute value of the trend slope |λ ij The formula for calculating | is:
[0072]
[0073] In the formula, t k Let k be the time sequence number of the k-th observation point, where k ranges from 1 to n.
[0074] Loss Risk Matrix M r The mathematical expression is:
[0075]
[0076] In the formula, p ij Let p be the probability value of data loss caused by the j-th risk factor in the i-th supply chain link. Risk factors include equipment failure, network interruption, human error, environmental changes, and system anomalies. ij Calculated using Bayesian statistical methods:
[0077]
[0078] In the formula, L(D) ij |H ij ) for assuming H ij Observational data D under the condition ijLikelihood function; P(H ij P(D) represents the prior probability; ij The marginal probability is L(D). ij |H ij The formula for calculating ) is:
[0079]
[0080] In the formula, d ijk f(d) represents the observation of the k-th loss event; ijk |h ij ) represents the probability density function under given assumptions; m represents the total number of observed events.
[0081] Jump detection matrix M j The mathematical expression is:
[0082]
[0083] In the formula, θ ij Let θ be the threshold for the jump of the i-th data type within the j-th time window. ij Calculated based on the statistical characteristics of the normal distribution:
[0084]
[0085] In the formula, μ ij This represents the average change of the i-th data type within the j-th time window; is the corresponding standard deviation; κ is the safety factor, ranging from 1.2 to 2.8. The mean of the variation is μ. ij The calculation formula is:
[0086]
[0087] In the formula, y ijk Let be the data value of the i-th data type at time k within the j-th time window; l is the number of data points within the time window. (Corresponding standard deviation) The calculation formula is:
[0088]
[0089] Anomaly detection matrix M a Expressed using a multi-layered structure:
[0090]
[0091] In the formula, γ ijThe confidence threshold for identifying the combination of the i-th data dimension and the j-th anomaly type is defined. Data dimensions include numerical, text, time, and spatial dimensions, while anomaly types include numerical anomalies, logical anomalies, time-series anomalies, and correlation anomalies.
[0092] The specific implementation of step S05 is to design a game model based on a two-layer optimization structure, where the objective function F of the upper-layer model is... upper Expressed as:
[0093]
[0094] In the formula, C cov For data coverage parameters; A acc To verify the accuracy parameters; W time For timeliness weighting parameters; S stab η1, η2, and η3 are stability coefficient parameters; they are weighting coefficients with values of 0.4, 0.35, and 0.25, respectively. Data coverage parameter C. cov The calculation formula is:
[0095]
[0096] In the formula, N valid N represents the number of valid data points. total Total number of data points. Verify accuracy parameter A. acc The calculation formula is:
[0097]
[0098] In the formula, N correct To verify the correct number of data points; N verified This represents the total number of validated data points. The stability coefficient parameter S... stab The calculation formula is:
[0099]
[0100] In the formula, w ij is the weight coefficient in the i-th row and j-th column of the stability evaluation matrix.
[0101] The objective function F of the lower-level model lower Expressed as:
[0102]
[0103] In the formula, T proc For processing time parameters; M mem This is a memory usage parameter; B band For network bandwidth parameters; C compξ1, ξ2, and ξ3 are the computational complexity parameters; ξ1, ξ2, and ξ3 are the weighting coefficients, with values of 0.45, 0.3, and 0.25, respectively.
[0104] Coupling term Q couple Expressed as:
[0105]
[0106] In the formula, the coupling term reflects the negative correlation between the verification accuracy parameter of the upper-level model and the processing time parameter of the lower-level model.
[0107] The specific implementation of step S06 is to construct a data fusion verification model based on a liquid neural network, and the liquid neuron state adjustment function S comp Expressed as:
[0108] S comp =ω1·N node +ω2·T fresh +ω3·R complete +ω4·H acc ;
[0109] In the formula, N node T represents the total number of nodes in the current supply chain path being processed; fresh For data timeliness; R complete For data integrity rate; H acc To verify historical accuracy, ω1, v2, ω3, and ω4 are weighting coefficients, with values of 0.25, 0.25, 0.25, and 0.25, respectively. The data completeness rate R... complete The calculation formula is:
[0110]
[0111] In the formula, N non-missing N represents the number of non-missing data points. expected Given the desired amount of complete data. Validate historical accuracy H. acc The calculation formula is:
[0112]
[0113] In the formula, acc u Let be the accuracy of the u-th historical validation; h be the number of historical validations. The data timeliness T... fresh The calculation formula is:
[0114]
[0115] In the formula, t current The current time is in seconds; t collectτ represents the data acquisition time in seconds; τ is the time decay constant, with a value of 3600 seconds.
[0116] The specific implementation method of step S07 is the same as described above, and will not be repeated in detail here.
[0117] The specific implementation of step S08 is to perform integrity verification calculation and comprehensively evaluate the score. final The calculation formula is:
[0118] Score final =δ1·Score s +δ2·Score r +δ3·Score j +δ4·Score a ;
[0119] In the formula, Score s The score is used to assess stability. r Score is the score used to assess the risk of loss. j The score is used to evaluate transition detection. a The score represents the anomaly detection evaluation score; δ1, δ2, δ3, and δ4 are weighting coefficients, with values of 0.3, 0.25, 0.25, and 0.2, respectively. The stability evaluation score is also included. s The calculation formula is:
[0120]
[0121] In the formula, f ij This is the stability factor for the j-th data dimension in the i-th stage, ranging from 0 to 1, representing the stability of the current data relative to historical data. (Score for loss risk assessment) r The calculation formula is:
[0122]
[0123] In the formula, r ij This represents the current risk status indicator value for the j-th risk factor in the i-th stage, and its value can be either 0 or 1. Jump Detection Evaluation Score j The calculation formula is:
[0124]
[0125] In the formula, N jump N represents the number of detected transition anomalies. check Total number of detections. Anomaly detection evaluation score. a The calculation formula is:
[0126]
[0127] In the formula, N anomaly N represents the number of abnormal data points identified. total_data This represents the total number of data points.
[0128] It should be explained that the specific implementation method for adjusting the state parameters of the liquid neuron is based on the comprehensive state score value S. comp Using piecewise function Φ(S) comp Adjustments are made accordingly.
[0129]
[0130] In the formula, φ low , φ mid-low , φ mid , φ mid-high , φ high These are parameter configuration vectors corresponding to low sensitivity, low-to-medium sensitivity, medium sensitivity, medium-to-high sensitivity, and high sensitivity modes, respectively. The parameter configuration vector φ includes the membrane potential parameter V. mem Threshold parameter V th Attenuation parameter α decay Three components:
[0131] φ=(V mem V th α decay );
[0132] The low-sensitivity mode is configured as φ low = (0.2, 0.8, 0.9), medium sensitivity mode configured as φ mid = (0.5, 0.5, 0.5), high sensitivity mode configured as φ high = (0.9, 0.2, 0.1).
[0133] The principles and effects of each formula are explained below. The formula for calculating the weight coefficients of the stability assessment matrix uses a combination of analysis of variance and trend analysis, through standardized variance values. The absolute value of the trend slope |λ reflects the degree of dispersion of the data. ij | Reflecting the trend of data changes, the weighted combination of the two This formula can comprehensively assess the stability characteristics of data in both time and space dimensions. Compared with traditional single-indicator assessment methods, it can more accurately identify data quality issues and improve the reliability of subsequent integrity verification.
[0134] The Bayesian probability calculation formula for the loss risk matrix is based on the principle of fusing prior knowledge and observed data, using the Bayesian formula. To achieve dynamic updates of risk probabilities, the likelihood function is used. This method reflects the degree to which observed data supports the hypothesis. Compared with traditional frequency statistics methods, it is more adaptable and can provide more accurate risk assessment results when historical data is limited, effectively improving the accuracy and timeliness of risk warning.
[0135] The threshold calculation formula for the jump detection matrix is based on the normal distribution statistical theory, and is obtained through the threshold formula. Determine the jump detection threshold, where the average change amplitude is... and standard deviation The introduction of a safety factor k, which reflects the statistical characteristics of data changes, ensures a balance between detection sensitivity and false alarm rate. Compared with the fixed threshold method, this formula can adaptively adjust the detection parameters according to the statistical characteristics of the data, significantly improving the accuracy and robustness of jump detection.
[0136] The objective function of the upper-level model adopts a nonlinear combination form. The first term η1·C cov ·ln(A acc The logarithmic function of ) reflects the non-linear promoting effect of data coverage on verification accuracy, the second term The exponential function reflects the diminishing marginal contribution of timeliness in a stable environment, and the third term η3·C cov ·A acc ·sin(W time The sine function represents the periodic optimization effect generated by multi-parameter coupling. Compared with the traditional linear combination, this function design can better capture the complex relationship between parameters and achieve the maximum optimization of verification accuracy. The dimensionless characteristics of the weight coefficients η1, η2, and η3 ensure the dimensionality consistency of the objective function.
[0137] The objective function of the lower-level model also adopts a nonlinear combination form. By combining square root, cosine, and logarithmic functions to reflect the nonlinear correlation consumption and periodic demand changes among different resources, this design can more accurately describe the consumption pattern of computing resources compared to simple weighted sums, thereby optimizing computing efficiency. The weight coefficients ξ1, ξ2, and ξ3 are normalized to ensure the uniformity of dimensions.
[0138] Coupled term function By establishing a negative correlation between upper and lower level models through the inverse square relationship, a coordinated optimization of verification accuracy and processing time is achieved. Compared with independent optimization, this coupling mechanism can avoid local optima and obtain a globally optimal verification strategy. Here, verification accuracy is a dimensionless parameter, and the inverse square of processing time has a s...-2 The dimension of the coupling term is s. -2 .
[0139] The state regulation function of the liquid neuron adopts the weighted average principle S comp =ω1·N node +ω2·T fresh +ω3·R complete +ω4·H acc The study comprehensively considers the impact of four factors on neuron state: number of nodes, timeliness, completeness rate, and historical accuracy. Timeliness is assessed using an exponential decay function. Reflecting the impact of time on data value, this function design, compared to fixed parameter configuration, can dynamically adjust the neuron state according to the actual situation, improving the adaptability and accuracy of data fusion. All parameters are normalized to ensure dimensionless characteristics.
[0140] Sensitivity adjustment mechanism Φ(S) of piecewise functions comp Different parameter configuration strategies φ=(V) are adopted according to different ranges of the comprehensive score. mem V th α decay This method achieves the optimal matching between neural network computing power and data quality. Compared with single parameter configuration, it can achieve the best processing effect under different data quality conditions, significantly improving the overall performance and robustness of the system. Among them, the membrane potential parameter, threshold parameter, and decay parameter are all dimensionless parameters, ensuring the dimensionality consistency of the configuration vector.
[0141] To better understand and implement this invention, the following is a specific application scenario of Example 2: The technical team first established an agricultural product supply chain data collection system. In the vegetable base planting stage, 128 sensor nodes were set up to collect soil temperature, humidity, pH value, and light intensity data every 10 minutes. In the pre-processing stage, 32 data collection points were configured to record washing temperature, sorting weight, and packaging time information in real time. In the cold chain storage stage, 64 temperature and humidity monitoring devices were deployed, collecting data every 5 minutes. In the transportation stage, GPS positioning systems and temperature monitoring devices were installed on 18 refrigerated trucks, uploading location and temperature data every minute. In the sales stage, a sales record system was set up in 48 stores to update inventory and sales information in real time. The entire system generates approximately 2.4 × 10⁻⁶ data points per day. 6 One original data record.
[0142] The technical team categorized the collected data into three levels based on timeliness. Real-time data includes urgent information such as abnormal temperature alarms and location deviation alerts, requiring processing within 30 seconds. Near real-time data includes important information such as environmental parameter change trends and logistics status updates, with a processing time limit of 5 minutes. Historical data includes basic information such as daily environmental records and logistics trajectory archives, which can be processed in batches within 1 hour. The data was also categorized into three types based on complexity: simplified data, rapid data, and detailed data, with simplified data accounting for 45% of the total, rapid data for 35%, and detailed data for 20%.
[0143] To establish a damage prediction model for agricultural products, the technical team analyzed the damage data of 8,640 batches of organic tomatoes over the past two years. Statistics showed that the damage rate during the planting stage was 3.2%, mainly due to pests, diseases, and extreme weather. The damage rate during pre-processing was 1.8%, primarily caused by mechanical damage. The damage rate during storage was 4.5%, with temperature fluctuations being the main factor. The damage rate during transportation was 2.1%, with packaging damage caused by bumpy transport. The damage rate during sales was 1.4%, mainly due to excessive storage time. The damage prediction model built based on this historical data achieved an accuracy rate of 87.3%, providing a reliable benchmark for subsequent integrity verification.
[0144] The technical team constructed a multi-dimensional evaluation matrix system, comprising four core components: a stability evaluation matrix, a loss and risk matrix, a jump detection matrix, and an anomaly identification matrix. The weight coefficient distribution of the stability evaluation matrix is shown in Table 1.
[0145] Table 1. Distribution of weight coefficients in the stability assessment matrix
[0146] Supply chain Temperature data Humidity data Location data Time data Quality data Planting 0.72 0.68 0.45 0.82 0.75 Preprocessing 0.65 0.63 0.58 0.79 0.71 Storage 0.85 0.78 0.52 0.73 0.69 Transportation 0.71 0.64 0.89 0.76 0.62 Sales process 0.59 0.61 0.47 0.68 0.74
[0147] The probability parameter distribution of the loss risk matrix is shown in Table 2:
[0148] Table 2 Distribution of probability parameters for the loss risk matrix
[0149] Supply chain Equipment failure Network interruption Human operation Environmental change System malfunction Planting 0.14 0.21 0.17 0.26 0.11 Preprocessing 0.19 0.18 0.24 0.15 0.22 Storage 0.16 0.23 0.19 0.28 0.13 Transportation 0.25 0.17 0.22 0.21 0.18 Sales process 0.12 0.16 0.27 0.14 0.20
[0150] The sensitivity parameters of the jump detection matrix are shown in Table 3:
[0151] Table 3 Sensitivity parameters of the transition detection matrix
[0152] Data types 1-minute window 5-minute window 15-minute window 30-minute window 60-minute window Temperature data 1.9 2.2 2.4 2.6 2.8 Humidity data 1.7 2.0 2.3 2.5 2.7 Location data 1.5 1.8 2.1 2.3 2.5 Time data 1.3 1.6 1.9 2.1 2.3 Quality data 2.1 2.4 2.6 2.8 2.8
[0153] The confidence thresholds for the anomaly detection matrix are shown in Table 4:
[0154] Table 4 Confidence Thresholds for Anomaly Detection Matrix
[0155] Data Dimensions Numerical anomalies Logical exception Timing anomalies Association anomalies Numerical Dimensions 0.86 0.79 0.83 0.80 Text Dimension 0.77 0.89 0.74 0.82 Time dimension 0.73 0.76 0.93 0.78 Spatial dimension 0.84 0.72 0.87 0.95
[0156] The technical team designed a game theory model for agricultural product traceability verification, employing a two-layer optimization strategy to achieve optimal results in integrity verification. The upper-layer model aims to maximize data integrity, comprehensively considering four key parameters: data coverage (0.92), verification accuracy (0.89), timeliness weight (0.76), and stability coefficient (0.71). The lower-layer model aims to optimize computational efficiency, considering a processing time of 2.1 seconds, memory usage of 485MB, network bandwidth of 32Mbps, and computational complexity of O(n^2). 2 Four resource parameters. Through coupling term design, when the processing time is reduced to 1.8 seconds by optimizing the lower-level model, the verification accuracy of the upper-level model is improved to 0.91, achieving synergistic optimization of efficiency and accuracy.
[0157] The core agricultural product data fusion and verification model employs a liquid neural network architecture with embedded nonlinear activation functions. This network contains 500 randomly connected liquid neurons and uses a hyperbolic tangent activation function to handle nonlinear mapping relationships. The technical team constructed a training dataset containing 48,000 samples, including 28,800 positive samples and 19,200 negative samples, covering 24 common data integrity problem patterns. Using the AdamW optimizer with a learning rate of 0.001 and a batch size of 64, the model converged after 156 training epochs.
[0158] The liquid neuron state regulation function calculates a comprehensive state score based on four parameters: the number of supply chain nodes, data timeliness, data completeness, and historical verification accuracy. In actual operation, a low-sensitivity mode is used when the comprehensive score is 0.15, with the membrane potential parameter set to -65mV, the threshold parameter set to -45mV, and the attenuation parameter set to 0.85. A medium-low sensitivity mode is used when the comprehensive score is 0.35, with the corresponding parameters adjusted to -58mV, -48mV, and 0.72. A medium sensitivity mode is used when the comprehensive score is 0.55, with the parameters set to -52mV, -52mV, and 0.65. A medium-high sensitivity mode is used when the comprehensive score is 0.72, with the parameters adjusted to -46mV, -56mV, and 0.58. A high sensitivity mode is used when the comprehensive score is 0.89, with the parameters set to -40mV, -60mV, and 0.45.
[0159] The technical team conducted a three-month comprehensive verification experiment, processing traceability data from three batches totaling 1,800 tons of organic tomatoes. The experiment simulated real-world problems such as 12% missing data, 6% data errors, and 4% timestamp anomalies. The system successfully identified 98.7% of data integrity issues, with a false alarm rate controlled below 2.3%. The average processing time was 1.9 seconds, memory usage remained stable at 456MB, and network bandwidth utilization was 75%, meeting the performance requirements for real-time processing.
[0160] It should be noted that the variables involved in this invention are explained in detail in Tables 5 and 6.
[0161] Table 5. Variable Explanation Table (Part 1)
[0162]
[0163]
[0164] Table 6. Variable Explanation Table (Part Two)
[0165]
[0166] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for verifying the integrity of agricultural product supply chain traceability data based on machine learning, characterized in that, Accurate verification of the integrity of agricultural product supply chain traceability data is achieved by constructing a multi-dimensional evaluation matrix system and a liquid neural network data fusion verification model, including establishing an agricultural product supply chain data collection system and constructing a data classification and processing architecture; Establish a damage prediction model for agricultural products; construct a multi-dimensional evaluation matrix system; design a game-theoretic model for agricultural product traceability verification; construct an agricultural product data fusion verification model using a liquid neural network with embedded nonlinear activation functions; conduct small-scale verification experiments; and perform integrity verification calculations.
2. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 1, characterized in that, The steps for establishing an agricultural product supply chain data collection system are as follows: set up multiple data collection nodes throughout the entire process of agricultural products from planting to sales, collect environmental parameters, logistics information, and quality inspection data to form a raw dataset, and divide the collected data into three levels according to timeliness: real-time data, near-real-time data, and historical data.
3. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 2, characterized in that, The steps for constructing the data classification and processing architecture specifically involve dividing the original dataset into three categories based on complexity and importance: simplified data, rapid data, and detailed data. Simplified data contains basic identification information, rapid data contains key quality indicators, and detailed data contains complete traceability chain information.
4. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 3, characterized in that, The steps for establishing a damage prediction model for agricultural products specifically involve analyzing historical damage data to construct a damage occurrence pattern model and identifying vulnerable links and damage modes.
5. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 4, characterized in that, The multi-dimensional evaluation matrix system includes a stability evaluation matrix, a loss risk matrix, a jump detection matrix, and an anomaly identification matrix. The stability evaluation matrix is established based on the stability performance of agricultural products in different links of the supply chain, and the loss risk matrix is constructed based on the probability and degree of loss of agricultural products in each link of the supply chain.
6. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 5, characterized in that, The jump detection matrix is established by analyzing the jump characteristics of data in the time series, and the anomaly identification matrix is constructed by combining the characteristics of agricultural products and the rules of the supply chain.
7. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 6, characterized in that, The agricultural product traceability verification game model includes an upper-level model that aims to maximize data integrity and a lower-level model that aims to optimize computational efficiency.
8. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 7, characterized in that, The agricultural product data fusion verification model is constructed using a liquid neural network with embedded nonlinear activation functions to fuse multi-source data. The state parameters of the liquid neurons in the model are dynamically adjusted according to three parameters: the number of supply chain nodes, data timeliness, and data importance.
9. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 8, characterized in that, The liquid neural network comprises a liquid neuron reservoir layer, a nonlinear activation function embedding layer, a state update layer, and an output mapping layer. The liquid neuron reservoir layer contains randomly connected liquid neuron nodes, and the nonlinear activation function embedding layer embeds the hyperbolic tangent activation function into the state transition process of the liquid neurons.
10. The method for verifying the integrity of agricultural product supply chain traceability data based on machine learning according to claim 9, characterized in that, The steps of the small-scale verification experiment are as follows: select agricultural product samples to collect data at different links in the supply chain, and determine the weight coefficient range of the stability assessment matrix, the probability threshold range of the loss risk matrix, the sensitivity parameter range of the jump detection matrix, and the confidence threshold range of the anomaly identification matrix through experiments.
Citation Information
Cited By
Integrity verification and quality evaluation method before agricultural Internet of Things data uplink
CN121456920A
An integrity verification and quality evaluation method for agricultural internet of things data before on-chain
CN121456920B
Agricultural product quality safety data traceability analysis system based on big data
CN121936984A
Agricultural product quality and safety data traceability analysis system based on big data
CN121936984B