A big data prediction method and system for cold chain transportation status
Patent Information
- Application Number
- CN202610873571.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-17
AI Technical Summary
[0006]为解决现有冷链运输状态预测技术中存在的由于简单采样策略无法识别并区分由外部扰动引发的波动数据与模型自身预测不确定性高的疑难数据,且在模型训练阶段丢弃了数据的优先级先验知识,从而限制模型预测性能提升的技术问题,本发明在如下的多个方面中提供方案
本发明能够从外部物理因果影响与内部算法认知短板两个高阶维度对海量时序数据的潜在价值进行深度解剖,并通过基于皮尔逊相关系数的滑动归因机制,实现了复杂冷链工况下数据价值权重的平滑自适应融合,大大降低了无效与平稳数据的算力浪费。
Smart Images

Figure CN122414960B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology. More specifically, this invention relates to a big data prediction method and system for cold chain transportation status. Background Technology
[0002] In the field of cold chain transportation, to achieve accurate prediction of the internal microenvironment of transportation equipment, existing technologies typically deploy multimodal sensors to collect multi-dimensional time-series data such as temperature, humidity, and door opening / closing in real time. Recurrent neural networks, long short-term memory networks, and attention-based Transformer models are then applied to this data to predict the future state of the cold chain microenvironment. To alleviate the computational burden on the model caused by massive amounts of stable data, data sampling has become a necessary preprocessing step before model training.
[0003] However, existing cold chain microenvironment state prediction technologies generally employ uniform sampling, random sampling, or threshold sampling methods based on a single statistical indicator in the data sampling stage. These crude sampling methods fail to identify and distinguish two core high-value data types at the underlying logic level: one type is state fluctuation data caused by external disturbance events, such as opening and closing doors, which has a clear spatiotemporal causal chain; the other type is difficult data with weak model cognitive ability and extremely high prediction uncertainty during the fitting process. Even if some high-value data can be screened out, the prior knowledge of the importance or priority of these data points is usually directly discarded in the subsequent model training stage. The self-attention mechanism of the existing Transformer network can only rely on the numerical characteristics of the data itself to perform correlation calculations, treating all input sampling points as equally important, and cannot target and focus on key influencing factors, thus limiting the improvement of prediction accuracy.
[0004] Referring to Chinese patent document CN120031471B, entitled "A Method and System for Cold Chain Transportation Demand Analysis Based on Big Data Prediction," this method acquires historical cold chain order data, extracts macro-level transportation demand characteristics such as cargo type and temperature sensitivity, transportation timeliness, and equipment energy consumption, and inputs these characteristics into a pre-trained spatiotemporal prediction model to generate multi-dimensional cold chain demand parameters. Ultimately, a dynamic scheduling scheme is generated based on these parameters. This technical solution addresses the macro-level optimization problem of transportation resource allocation and route planning.
[0005] However, the aforementioned patent documents focus on overall transportation demand and route planning, without addressing the identification of specific physical event fluctuations in high-frequency time-series data of the in-vehicle microenvironment, nor the analysis of the model's own cognitive blind spots. They fail to identify high-value data from the two dimensions of external physical event influence and internal model cognitive shortcomings, and thus cannot effectively inject this data value prior knowledge into the deep learning model to target and guide attention. Summary of the Invention
[0006] To address the technical problems in existing cold chain transportation status prediction technologies, such as the inability of simple sampling strategies to identify and distinguish between fluctuating data caused by external disturbances and difficult data with high uncertainty in the model's own predictions, and the discarding of prior knowledge of data priority during the model training phase, thereby limiting the improvement of model prediction performance, this invention provides solutions in the following aspects.
[0007] In a first aspect, the present invention provides a big data prediction method for cold chain transportation status, comprising: acquiring multidimensional time-series cold chain status data of a historical observation period and encoding it into a feature vector sequence, and inputting it into a benchmark prediction model; calculating the event disturbance priority and model uncertainty priority at each time step based on the spatial distance and time difference decay rules of door opening and closing events and the model prediction difference; extracting a prediction error sequence from the predicted and actual values of historical time steps based on the benchmark prediction model, and calculating the correlation index between the error sequence and the historical event disturbance priority sequence within a historical sliding time window to obtain an attribution coefficient sequence; weightedly fusing the event disturbance priority and the model uncertainty priority based on the attribution coefficient sequence to generate a sampling weight sequence; acquiring cold chain data to be predicted and encoding it into a feature sequence to be predicted, generating a high-value data subset by non-uniform sampling of the feature sequence to be predicted according to the sampling weight sequence, and inputting it into a weight-enhanced attention network to output an initial score matrix; generating a weight adjustment matrix based on the weights corresponding to the subset, and multiplying it element-wise with the normalized initial score matrix to obtain the target attention weight, so as to output the future state prediction result.
[0008] This invention calculates the priority of event disturbances, representing the impact of external physical events, and the priority of model uncertainty, representing cognitive blind spots, respectively. It then innovatively utilizes the cross-correlation of prediction errors for attribution-weighted fusion, accurately selecting high-value training data that possesses both causal value and information gain. Furthermore, by injecting the sampling weight sequence into the self-attention computation layer using element-wise multiplication tensor operations, it successfully transforms prior knowledge of data value into a hard constraint guiding the model's focus. This significantly reduces the computational overhead of redundant data while achieving a breakthrough improvement in the accuracy and anti-interference capability of cold chain transportation status prediction.
[0009] Preferably, both the multidimensional time-series cold chain status data of the historical observation period and the cold chain data to be predicted include location coordinates, door opening and closing events, temperature readings, and humidity readings; the encoding into a feature vector sequence and a feature sequence to be predicted includes: spatially encoding the location coordinates and processing the door opening and closing events into binary features; concatenating the encoded location coordinates, the binary features, and the temperature and humidity readings, and encoding them according to time steps to form the feature vector sequence or the feature sequence to be predicted; the baseline prediction model is a long short-term memory network.
[0010] This invention constructs a holographic physical mapping space by stitching and encoding three-dimensional space, discrete events and continuous temperature and humidity readings at the bottom layer, enabling the long short-term memory network to fully capture the multimodal long-distance temporal dependencies of cold chain microenvironment changes.
[0011] Preferably, the step of obtaining cold chain multidimensional time-series status data for historical observation periods includes: performing time-series-based linear interpolation on the cold chain multidimensional time-series status data to fill missing values; and using a normalized threshold method to remove outliers from the cold chain multidimensional time-series status data.
[0012] This invention employs a time-based linear interpolation method, which can preserve the continuous physical characteristics of environmental parameters changing over time in the multidimensional time-series cold chain status data to the greatest extent possible, and avoid introducing artificial data breaks and gradient abrupt changes.
[0013] Preferably, the calculation of the event disturbance priority at each time step includes: calculating the spatial Euclidean distance based on the timestamp of the current time step and the location coordinates, as well as the timestamps of the historical door opening and closing events and the location coordinates; substituting the spatial Euclidean distance and the time difference into a preset exponential decay function to calculate the impact intensity of a single event; and accumulating all the impact intensities of the single events to obtain the event disturbance priority.
[0014] This invention employs a nonlinear exponential decay function with a defined attenuation coefficient and performs a historical accumulation operation, which can perfectly fit the spatial damping characteristics and time lag effect of thermodynamic transfer within the cold chain compartment, and accurately assess the impact potential energy of external physical disturbances on the current microenvironment state.
[0015] Preferably, the model uncertainty priority is obtained as follows: Calculate the partial derivative matrix of the mean squared error loss of the benchmark prediction model with respect to the activation values of the last hidden layer of the benchmark prediction model; perform singular value decomposition on the partial derivative matrix to obtain a set of singular values; calculate the probability distribution based on the set of singular values; if, in an extreme case, all partial derivative matrices are zero, resulting in a zero denominator for the calculated probability distribution, then each probability distribution is set as a uniform distribution; calculate the normalized Shannon entropy of the probability distribution as the model uncertainty priority; or, calculate the output variance based on multiple random forward propagations of the benchmark prediction model, and use the output variance as the model uncertainty priority.
[0016] On the one hand, this invention can achieve high-precision theoretical detection of cognitive blind spots by decomposing the singular value of the underlying gradient space; on the other hand, it can also achieve lightweight engineering approximation evaluation by using the output variance of multiple random forward propagations, thus balancing the high-precision theoretical rigor of the central cloud with the engineering real-time performance of edge cold chain vehicle-mounted equipment with limited computing power.
[0017] Preferably, the step of calculating the correlation index between the error sequence and the historical event perturbation priority sequence to obtain the attribution coefficient sequence includes: extracting subsequences from the prediction error sequence and the event perturbation priority sequence within the historical sliding time window; calculating the Pearson correlation coefficient of the subsequence and using it as the correlation index; and truncating the negative values in the Pearson correlation coefficient to zero to obtain the attribution coefficient sequence.
[0018] This invention employs a fixed time window to calculate the local Pearson correlation coefficient and performs a strict negative correlation truncation to zero operation, effectively filtering out irrelevant random noise. It ensures that the system establishes an attribution relationship only when the model's prediction error is significant and shows a positive linear causal correlation with the external door opening / closing event, thereby significantly improving the statistical reliability of dynamic attribution analysis.
[0019] Preferably, the sampling weight sequence is generated by weighted fusion of the event perturbation priority and the model uncertainty priority based on the attribution coefficient sequence, including: mapping the attribution coefficient sequence to an event perturbation weight sequence through a sigmoid activation function; and using the event perturbation weight sequence and the corresponding complement weights, performing a weighted summation of the event perturbation priority sequence and the model uncertainty priority sequence to obtain the sampling weight sequence.
[0020] Preferably, generating a high-value data subset by non-uniformly sampling the feature sequence to be predicted according to the sampling weight sequence includes: normalizing the sampling weight sequence to obtain a sampling probability sequence for each time step; determining the total number of samples according to a preset sampling ratio and the total length of the feature sequence to be predicted; and extracting data points corresponding to the total number of samples without replacement according to the sampling probability sequence to form the high-value data subset.
[0021] Preferably, inputting the feature vector sequence into the baseline prediction model includes: using mean squared error as the loss function, and using an adaptive moment estimation optimizer to train the long short-term memory network through forward and backward propagation to obtain the baseline prediction model that has initially converged.
[0022] Secondly, the present invention provides a big data prediction system for cold chain transportation status, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned big data prediction method for cold chain transportation status is implemented.
[0023] By adopting the above technical solution, a big data prediction method for cold chain transportation status is generated into a computer program and stored in a memory for loading and execution by a processor. Terminal devices are then created based on the memory and processor for convenient use.
[0024] The beneficial effects of this invention are as follows: This invention can deeply analyze the potential value of massive time-series data from two high-order dimensions: external physical causal influences and internal algorithmic cognitive shortcomings. Through a sliding attribution mechanism based on Pearson correlation coefficient, it achieves smooth adaptive fusion of data value weights under complex cold chain conditions, greatly reducing the computational waste of invalid and stable data.
[0025] Furthermore, this invention no longer stops data value assessment at the preprocessing stage, but treats it as a prior knowledge weight matrix, and performs hard element-wise intervention on the underlying tensor dimension and the Softmax normalized score matrix of the Transformer architecture. In this way, without violating the mathematical premise of attention convergence limit, it forcibly guides the deep neural network to target and focus on high-frequency disturbances and difficult doubts, thereby improving the model's sensitivity and prediction accuracy to sudden cold chain anomalies. Attached Figure Description
[0026] Figure 1 This is a flowchart of a big data prediction method for the status of cold chain transportation; Figure 2 It is a graph showing the changes in the collected data; Figure 3 This is a bar chart comparing results from different experimental errors. Detailed Implementation
[0027] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0028] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0029] This invention discloses a big data prediction method for cold chain transportation status, referring to... Figure 1 This includes steps S1-S3: S1. Obtain the cold chain data to be predicted and encode it into a feature sequence to be predicted.
[0030] The system acquires multidimensional time-series cold chain status data from historical observation periods and encodes it into a feature vector sequence by time step. This feature vector sequence is then input into the baseline prediction model. Additionally, during the online prediction phase, the cold chain data to be predicted is acquired and encoded into a feature sequence to be predicted.
[0031] In an optional embodiment, the IoT gateway first periodically collects raw data streams from temperature sensors, humidity sensors, GPS modules, and door magnetic sensors deployed in the cold chain transportation equipment using the MQTT protocol or HTTP polling method to form cold chain multidimensional time-series status data. The cold chain multidimensional time-series status data, stored as a CSV file or database, is then loaded into a DataFrame structure. The cold chain multidimensional time-series status data includes location coordinates, door opening and closing events, temperature readings, and humidity readings.
[0032] Reference Figure 2 This paper presents a schematic diagram of 25 hours of data collection on the internal status of cold chain transportation equipment during a certain historical observation period. As shown by the solid line in the figure and the dashed line in the figure, the temperature change trajectory shows that when an external physical disturbance occurs, the original stable thermodynamic microenvironment inside the carriage is disrupted. The temperature readings show a significant lag increase and fluctuation after the event. This physical phenomenon indicates that a specific event is the core causal variable that causes drastic changes in the cold chain microenvironment. This also provides a practical basis for prioritizing the extraction of high-value data affected by the event in subsequent steps.
[0033] The acquired multidimensional time-series status data of the cold chain is subjected to time-series-based linear interpolation to fill missing values. Specifically, the interpolation method of DataFrame is used and the mode parameter is set to time to perform interpolation filling on the multidimensional time-series status data of the cold chain. This can preserve the continuous physical characteristics of temperature and humidity gradually changing over time in the multidimensional time-series status data of the cold chain to the greatest extent, and avoid introducing artificial data discontinuities and gradient abrupt changes.
[0034] Subsequently, a normalized thresholding method was used to remove outliers from the multidimensional time-series cold chain status data. Specifically, a Z-score-based normalized thresholding method was used to identify and remove outliers in the temperature and humidity readings of the multidimensional time-series cold chain status data, and Z-score normalization was performed on the temperature and humidity readings.
[0035] After preprocessing, feature engineering and time-step encoding are performed. The location coordinates are then encoded in three-dimensional space, specifically by pre-reading the physical internal dimensions of the cold chain transport equipment. ,Width ,high The three-dimensional position coordinates (x, y, z) are spatially normalized based on the physical internal size parameters, as shown in the formula. , , This process ensures that the normalized position coordinate values fall within the [0, 1] interval. Simultaneously, door opening and closing events are processed as binary features, with 1 for door open and 0 for door closed. Finally, the encoded position coordinates, binary features, and Z-score-normalized temperature and humidity readings are directly concatenated at each time step using the `hstack` function from the NumPy library, resulting in a precise, unified 6-dimensional feature vector. The position coordinates occupy 3 dimensions, the binary features 1 dimension, the temperature reading 1 dimension, and the humidity reading 1 dimension. These are concatenated at each time step to form a 3D tensor with the shape of sample number × time step size × feature dimension, ultimately encoding a feature vector sequence at each time step.
[0036] After obtaining the feature vector sequence, it is input into the baseline prediction model. The baseline prediction model is specifically a Long Short-Term Memory (LSTM) network. In actual deployment, the PyTorch framework is used to implement the baseline prediction model, specifically calling `torch.nn.LSTM` to construct a sequence prediction structure containing two LSM hidden layers and one fully connected output layer. The number of units in each hidden layer of the LSM network is set to 128. The input layer dimension of the baseline prediction model matches the 6-dimensional features of the feature vector sequence, and the output layer of the baseline prediction model outputs the predicted temperature or humidity scalar value for the next time step. In the model training phase, mean squared error is used as the loss function, and an adaptive moment estimation optimizer is used to train the LSM network through forward and backward propagation. During training, the initial learning rate is set to 0.001, the batch size is set to 64, and training is performed on the complete dataset for 50 epochs, ultimately obtaining a preliminarily converged baseline prediction model.
[0037] A two-layer LSTM structure with 128 units per layer is used, and an adaptive moment estimation optimizer is executed for 50 cycles on the full dataset to balance the ability to capture multi-dimensional time-series features of the cold chain with computing power. Without causing model overfitting, the long-distance memory dependencies of 6-dimensional multi-source time-series features are fully extracted.
[0038] S2. Generate the sampling weight sequence.
[0039] The event disturbance priority and model uncertainty priority are calculated at each time step to form the event disturbance priority sequence and the model uncertainty priority sequence. The prediction error sequence is obtained based on the predicted and actual values of the benchmark prediction model. A historical sliding time window is set, and the Pearson correlation coefficient between the prediction error sequence and the event disturbance priority sequence is calculated within the historical sliding time window to obtain the attribution coefficient sequence. The event disturbance priority sequence and the model uncertainty priority sequence are weighted and fused based on the attribution coefficient sequence to generate the sampling weight sequence.
[0040] In an optional embodiment, the event perturbation priority for each time step is first calculated. For any data point in the feature vector sequence... Based on the timestamp of the current time step and position coordinates And historical events involving the opening and closing of doors. timestamp and position coordinates Calculate the spatial Euclidean distance in three-dimensional space. The spatial Euclidean distance is compared with the time difference. Substituting the pre-defined exponential decay function, the intensity of a single event's impact is calculated. The introduced spatiotemporal impact decay model relationship is as follows: Among them, the time difference is set. The unit is h, the unit of spatial Euclidean distance is m, and the time decay coefficient is... The preferred range is [0.05, 0.5] and the dimension is h. -1 In this embodiment, the value is 0.1; spatial attenuation coefficient The preferred range is [0.1, 1] and the dimension is m. -1 In this embodiment, the value is 0.5. Then, the `sum` function from the NumPy library is used to accumulate the impact strength of all historical door opening and closing events on the single event occurring at the current time step, using the following formula: The initial event perturbation priorities for each time step are obtained. To avoid weight imbalance during subsequent feature fusion due to unlimited accumulated values, the Min-Max normalization method is applied to all initial event perturbation priorities to map their value range to the [0, 1] interval, and the normalized event perturbation priority sequence is obtained by integrating them according to the time dimension.
[0041] For example, a data point in the feature vector sequence The timestamp is 50 hours, and the location coordinates are (1, 2, 0.5) m. This is a record of historical door opening and closing events. Occurred in The location coordinates are (0, 0, 0) m. Then the time difference... Spatial Euclidean distance Substituting into the exponential decay function, the calculated impact strength of a single event is 0.26. If other historical door opening and closing events exist, the calculation continues and the values are accumulated. The event disturbance priorities of each time step are integrated according to the time dimension to form an event disturbance priority sequence.
[0042] It can accurately assess the cumulative impact potential energy of external physical disturbances on the current microenvironment, improving the fidelity of event disturbance priority in depicting physical reality.
[0043] Simultaneously, the model uncertainty priority is calculated for each time step. The current time step and previous time steps are selected. The length of the data points is The local feature vector sequence is input into the baseline prediction model for forward propagation, where the window size is... The example value is 24. The mean squared error loss between the predicted and actual values is calculated using forward propagation. ,in, This represents the predicted value output by the baseline prediction model. This represents the corresponding true value. The `requires_grad` attribute of the hidden layer state tensor is set to true. Then, backpropagation is performed on the mean squared error loss, calculating the partial derivatives of the mean squared error loss with respect to the state activation values at all time steps in the last hidden layer of the baseline prediction model. These partial derivatives are collected to form a dimension... The partial derivative matrix is then subjected to singular value decomposition to obtain the set of singular values. The number of singular values is bounded. Normalize the singular value set and calculate the probability distribution. It should be noted that, in extreme cases, if the partial derivative matrix is all zero, it will cause the denominator to... If the value is zero, then define the probability distributions. To ensure a uniform distribution, the Shannon entropy of the normalized singular value set is calculated. This normalized Shannon entropy is then used as the model uncertainty priority. Specifically, the normalized Shannon entropy is calculated by dividing the original Shannon entropy by the singular value count limit of the partial derivative matrix under a given window length. It is achieved by the logarithm of the equation, and its calculation relationship is as follows:
[0044] in, This indicates the priority of model uncertainty, and its value range is mapped to the interval [0, 1]. The first one obtained from the above calculation The probability distribution corresponding to each singular value; This is the aforementioned limit on the number of singular values; if Then it is agreed .
[0045] The model uncertainty priority of each time step is integrated according to the time dimension to form a model uncertainty priority sequence.
[0046] Furthermore, the correlation index between the prediction error sequence of the baseline prediction model and the priority sequence of event disturbances is calculated. A historical sliding time window of fixed width is set, and the width of the historical sliding time window is denoted as... In this embodiment, the value is taken as 60, which is in the case of a width Within the historical sliding time window, corresponding subsequences are extracted from the prediction error sequence and the event perturbation priority sequence, respectively. The correlation index between the two extracted subsequences is calculated; in this embodiment, the Pearson correlation coefficient is specifically used. The extracted prediction error subsequence is denoted as... The event perturbation priority subsequence is denoted as Using relational expressions Calculate the Pearson correlation coefficient between the two extracted subsequences, where Describing covariance, The standard deviation operator indicates that if the variance of the subsequence within a local window is zero, resulting in a zero standard deviation in the denominator, then the Pearson correlation coefficient within that sliding window is directly determined. Zero. The negative or no-correlation values in the calculated Pearson correlation coefficient are forcibly truncated to zero, as expressed in the formula: This leads to the attribution coefficient sequence.
[0047] Finally, a sampling weight sequence is generated by weighting and fusing the event perturbation priority sequence and the model uncertainty priority sequence based on the attribution coefficient sequence. First, the attribution coefficient sequence is nonlinearly mapped using a sigmoid activation function to generate the event perturbation weight sequence. The mapping relationship is as follows: The scaling factor controls the steepness of the weight switching. The preferred range is [5, 20], and in this embodiment, the value is 10, with an offset. The preferred range is [0.3, 0.7], and in this embodiment, the value is 0.5. Subsequently, the calculated event perturbation weight sequence is used... and the corresponding complement weights The normalized event disturbance priority sequence and the model uncertainty priority sequence are summed point-by-point with weights. The calculation formula is as follows: This generates a sampling weight sequence that combines considerations of event disturbance priority and model uncertainty priority.
[0048] For example, if the attribution coefficient corresponding to a certain data point in the feature vector sequence is 0.8, then the weight value in the calculated event perturbation weight sequence is substituted into it. The attribution coefficient is 0.95, meaning the sampling weight for this data point is primarily determined by the event disturbance priority; if the attribution coefficient is 0.1, then the calculated weight value... The value is 0.02, and the sampling weight corresponding to this data point is mainly determined by the model uncertainty priority.
[0049] It should be noted that the aforementioned model uncertainty calculation method based on singular value decomposition of partial derivative matrices can provide extremely high evaluation accuracy and is suitable for central cloud servers with strong computing power. In some cold chain edge vehicle computing terminal application scenarios with limited computing power, the model uncertainty prioritization can also be replaced by lightweight calculation methods based on the output variance of multiple forward propagations, such as the Monte Carlo dropout method, to balance prediction accuracy and engineering real-time requirements.
[0050] S3. Utilize target attention weights to output future state prediction results.
[0051] A high-value data subset is generated by non-uniformly sampling the current feature vector sequence to be predicted based on the sampling weight sequence. The high-value data subset is input into the weight-enhanced attention network to output the initial score matrix. Based on the weights of the corresponding high-value data subset in the sampling weight sequence, a weight adjustment matrix with the same dimension as the initial score matrix is generated by dimensional expansion. The weight adjustment matrix is multiplied element-wise with the normalized initial score matrix to obtain the target attention weight. The future state prediction result is output using the target attention weight.
[0052] In an optional embodiment, after acquiring the cold chain data to be predicted and encoding it into a feature sequence to be predicted, a non-uniform sampling operation is performed on the current feature vector sequence to be predicted based on a pre-generated sampling weight sequence derived from historical data analysis. The sampling weight sequence corresponding to the full time step is then normalized, and the relationship is calculated as follows: During implementation, if any time step occurs... If all values approach zero, resulting in a zero denominator, the process degenerates into using a uniform sampling strategy to obtain the sampling probability sequence for each time step. Then, based on a preset sampling ratio... and the total length of the feature vector sequence The total number of samples is determined by the following formula: Among them, the preset sampling ratio The preferred range is [0.1, 0.4], and in this embodiment, the value is 0.2. Subsequently, the corresponding total number of samples is drawn without replacement according to the sampling probability sequence. The data points constitute a high-value subset of data. For example, for a sequence containing 10,000 data points, if... If the value is 0.2, then 2000 points are precisely extracted to form a high-value data subset composed of 6-dimensional feature vectors.
[0053] Next, the high-value data subset is input into a weighted attention network. Specifically, this network interacts with the high-value data subset using a self-attention mechanism, outputting an initial score matrix. Then, the weights corresponding to the high-value data subset in the sampled weight sequence are extracted, and a weight adjustment matrix with the same dimension as the initial score matrix is generated through dimensionality expansion, such as using a tensor broadcasting mechanism. Finally, the weight adjustment matrix is multiplied element-wise with the initial score matrix after Softmax normalization, thereby transforming the prior data value into a hard constraint guiding the network's focus, resulting in the target attention weights.
[0054] Finally, during the forward inference prediction phase, the network outputs feature vectors for one or more future time steps. Temperature or humidity scalar values are extracted from these vectors, and the normalized values are converted into physically meaningful cold chain environment values using an inverse transformation method. The resulting future state predictions are then output using the target attention weights.
[0055] To further verify the effectiveness and advancement of the big data prediction method for cold chain transportation status of this invention, the following explanation of the beneficial effects is based on specific experimental data and comparative models.
[0056] In the specific experimental implementation, six consecutive months of multi-point sensor data from inside refrigerated trucks, provided by a cold chain logistics company, were used as the test benchmark, totaling 100,000 time-step data points. 80% of the data was used as the training set, and the remaining 20% as the test set. The prediction task was to predict the temperature for the next hour. The evaluation metrics used were mean absolute error (MAE) and root mean square error (RMSE). In all experiments involving sampling, the data sampling ratio was uniformly set to 20%.
[0057] To demonstrate the superiority of the feature vector sequence after non-uniform sampling and fusion of event perturbation priority and model uncertainty priority, the following five sets of comparative experimental models were set up, referring to... Figure 3 This is a diagram comparing the prediction errors of different experiments. The solid-color bars represent the mean absolute error (MAE), and the shaded bars represent the root mean square error (RMSE). After training on the full dataset, the benchmark Long Short-Term Memory network achieved a MAE of 0.451℃ and an RMSE of 0.658℃ on the test set.
[0058] Using 20% of the data for training the baseline model, the performance decreased, with MAE reaching 0.583℃ and RMSE reaching 0.824℃.
[0059] In the first ablation experiment, the weight-enhanced attention network trained using only event perturbation priority sampling achieved a MAE of 0.392℃ and an RMSE of 0.575℃.
[0060] In ablation experiment two, the model trained using only model uncertainty prioritization had a MAE of 0.365℃ and an RMSE of 0.521℃. The complete method proposed in this invention, which involves sampling with weighted fusion of two priorities and training with a weighted enhanced attention network, achieves optimal performance with a MAE of 0.284℃ and an RMSE of 0.413℃.
[0061] Compared to random sampling, all priority-based sampling models showed improved performance, demonstrating the necessity of targeted screening of high-value data. Both ablation experiment models outperformed the baseline models trained on the complete dataset, demonstrating the effectiveness of event perturbation priority and model uncertainty priority in guiding model convergence, with model uncertainty priority contributing slightly more than event perturbation priority.
[0062] Crucially, the complete solution model proposed in this invention achieves a significant performance improvement compared to the two single-priority ablation experiment models. Its MAE index is further reduced by 22.2% compared to the second-best ablation experiment. This demonstrates the superiority of the mechanism of dynamically weighting and fusing event perturbation priority and model uncertainty priority through attribution analysis. Simultaneously, the weighted attention network effectively utilizes the prior weights of high-value data, perfectly achieving coordinated targeted attention to key physical events and model cognitive difficulties. This significantly reduces training computational costs while breaking through the accuracy bottleneck of traditional cold chain prediction.
[0063] This invention also discloses a big data prediction system for cold chain transportation status, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a big data prediction method for cold chain transportation status according to the present invention is implemented.
[0064] The aforementioned big data prediction system for cold chain transportation status also includes other components well known to those skilled in the art, such as communication buses and communication interfaces. Their settings and functions are known in the art and will not be described in detail here.
[0065] In the description of this specification, "multiple" or "several" means at least two, such as two, three or more, unless otherwise expressly and specifically defined.
Claims
1. A big data prediction method for cold chain transportation status, characterized in that, This includes: acquiring multidimensional time-series cold chain status data from historical observation periods and encoding it into a feature vector sequence, then inputting it into a benchmark prediction model, wherein the benchmark prediction model is a long short-term memory network; Calculate the event disturbance priority and model uncertainty priority for each time step; extract the prediction error sequence from the predicted and actual values of the historical time steps based on the benchmark prediction model, and calculate the correlation index between the error sequence and the historical event disturbance priority sequence within the historical sliding time window to obtain the attribution coefficient sequence. Based on the weighted fusion of the event perturbation priority and the model uncertainty priority according to the attribution coefficient sequence, a sampling weight sequence is generated, including: mapping the attribution coefficient sequence to an event perturbation weight sequence through a sigmoid activation function; using the event perturbation weight sequence and the corresponding complement weight sequence, performing a weighted summation of the event perturbation priority sequence and the model uncertainty priority sequence to obtain the sampling weight sequence; acquiring the cold chain data to be predicted and encoding it into a feature sequence to be predicted; performing non-uniform sampling on the feature sequence to be predicted according to the sampling weight sequence to generate a high-value data subset, and inputting it into a weight-enhanced attention network to output an initial score matrix; generating a weight adjustment matrix based on the weights corresponding to the subset through dimensional expansion, and multiplying it element-wise with the normalized initial score matrix to obtain the target attention weight, so as to output the future state prediction result; the multidimensional time-series state data of the cold chain during the historical observation period and the cold chain data to be predicted both include location coordinates, door opening and closing events, temperature readings, and humidity. The process involves: reading the degree value; calculating the event disturbance priority at each time step, including: calculating the spatial Euclidean distance based on the timestamp of the current time step and the location coordinates, as well as the timestamps of the historical door opening and closing events and the location coordinates; substituting the spatial Euclidean distance and the time difference into a preset exponential decay function to calculate the intensity of a single event's impact, and summing all the intensity of a single event's impact to obtain the event disturbance priority; obtaining the model uncertainty priority through the following method: calculating the partial derivative matrix of the mean squared error loss of the benchmark prediction model with respect to the state activation values in the last hidden layer of the benchmark prediction model; performing singular value decomposition on the partial derivative matrix to obtain a set of singular values; calculating the probability distribution based on the set of singular values; if, in extreme cases, all partial derivative matrices are zero, resulting in a zero denominator for the calculated probability distribution, then setting each probability distribution as a uniform distribution; calculating the normalized Shannon entropy of the probability distribution as the model uncertainty priority; or, calculating the output variance based on multiple random forward propagations of the benchmark prediction model, and using the output variance as the model uncertainty priority.
2. The big data prediction method for cold chain transportation status according to claim 1, characterized in that, Encoding into a feature vector sequence or a feature sequence to be predicted includes: spatially encoding the location coordinates and processing the door opening / closing event into binary features; concatenating the encoded location coordinates, the binary features, the temperature reading, and the humidity reading, and encoding them according to time steps to form the feature vector sequence or the feature sequence to be predicted.
3. The big data prediction method for cold chain transportation status according to claim 1, characterized in that, The process of obtaining multidimensional time-series status data of the cold chain during historical observation periods includes: performing time-series-based linear interpolation on the multidimensional time-series status data of the cold chain to fill in missing values; and using a normalized threshold method to remove outliers from the multidimensional time-series status data of the cold chain.
4. The big data prediction method for cold chain transportation status according to claim 1, characterized in that, The step of calculating the correlation index between the error sequence and the historical event perturbation priority sequence to obtain the attribution coefficient sequence includes: extracting subsequences from the prediction error sequence and the event perturbation priority sequence within the historical sliding time window; calculating the Pearson correlation coefficient between the subsequences of the prediction error sequence and the subsequences of the event perturbation priority sequence, and using it as the correlation index; and truncating the negative values in the Pearson correlation coefficient to zero to obtain the attribution coefficient sequence.
5. The big data prediction method for cold chain transportation status according to claim 1, characterized in that, The high-value data subset is generated by non-uniformly sampling the feature sequence to be predicted based on the sampling weight sequence, including: normalizing the sampling weight sequence to obtain the sampling probability sequence at each time step; determining the total number of samples based on a preset sampling ratio and the total length of the feature sequence to be predicted; and extracting data points corresponding to the total number of samples without replacement according to the sampling probability sequence to form the high-value data subset.
6. The big data prediction method for cold chain transportation status according to claim 2, characterized in that, The feature vector sequence is input into the baseline prediction model, which includes: using mean squared error as the loss function, and using an adaptive moment estimation optimizer to train the long short-term memory network through forward and backward propagation to obtain the baseline prediction model that has initially converged.
7. A big data prediction system for cold chain transportation status, characterized in that, include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement a big data prediction method for cold chain transportation status according to any one of claims 1-6.
Citation Information
Patent Citations
Cold Chain Transportation Demand Analysis Method and System Based on Big Data Prediction
CN120031471B
Temperature prediction method and system for multi-mode AIGC cold-chain logistics monitoring platform
CN119477139A
Cold chain transportation temperature control monitoring method and system based on artificial intelligence
CN120746427A