Cluster type supply chain emergency early warning and linkage decision-making method based on deep learning
By combining deep learning methods with 2D/3D convolution and temporal networks, the problem of real-time risk classification and coordinated handling of multi-source heterogeneous high-dimensional data in the supply chain was solved, achieving efficient and accurate risk warning and decision support, and improving the resilience and operational efficiency of the supply chain system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-09
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to achieve real-time risk classification and coordinated response in multi-source, heterogeneous, non-stationary, nonlinear, and high-dimensional supply chain data, resulting in high false alarm rates, insufficient model adaptability, and poor decision execution.
By employing a deep learning approach that combines 2D/3D convolutions with temporal networks, and through hierarchical thresholding, hysteresis banding, and cost-sensitive optimization, a cross-chain linkage model is constructed to achieve data governance, spatiotemporal fusion, dynamic evaluation, and hierarchical alerting, forming a closed-loop early warning and decision support system.
It improves the accuracy and feasibility of supply chain risk early warning, reduces tiered fluctuations and false alarm rates, and enhances the resilience and operational efficiency of the supply chain system in complex environments.
Smart Images

Figure CN121808541A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of supply chain management and computer data processing technology, in particular to a cluster type supply chain emergency early warning and linkage decision method based on deep learning. The method runs on a computing device and a computer readable medium, comprehensively uses time-space data fusion, convolutional neural network (CNN) and continuous / variable depth time series network (such as long short-term memory network LSTM, etc.), models and infers multi-source heterogeneous data from supply, demand and logistics links, realizes real-time monitoring, dynamic evaluation and hierarchical disposal of emergency risks. The method is suitable for global manufacturing, cross-border logistics, emergency material guarantee, urban emergency supply and large-scale activity guarantee, etc. which have high requirements for timeliness and stability, and can provide continuous, interpretable and executable early warning and decision support for supply chain operation under complex external environment and resource constraints. TECHNICAL BACKGROUND
[0002] With globalization, digitization and increasing uncertainty, the supply chain presents the characteristics of diversified participants, complex network structure and high volatility of running environment. Extreme weather, public health events, geopolitics, transportation congestion, shortage of key raw materials, demand side mutation and price volatility, etc. will affect the balance of supply and demand and logistics operation in different levels and different regions in various forms. Traditional supply chain management mainly relies on rule threshold, static safety stock, linear or quasi-linear planning model and offline report analysis, which is difficult to reflect the linkage effect of cross-link and cross-region in time and accurately, and the response lags behind when facing long tail risks and rapid changes, which is difficult to meet the needs of fine and real-time risk prevention and control.
[0003] The progress of deep learning in pattern recognition and time series prediction provides a new technical path for supply chain risk early warning. But the existing applications mainly focus on single link or single data modality, using fixed threshold or static model for alarm, which is difficult to solve the following problems: (1) Data level: multi-source heterogeneous, different scales, inconsistent time delays and existence of missing, abnormal and caliber changes, which easily leads to feature drift and statistical mismatch; (2) Model level: risk transmission has spatial correlation and cross-chain cooperation, which is non-stationary, nonlinear and mutation in time series, and traditional single chain or low-dimensional model is difficult to describe the interaction structure of "time × chain number × risk dimension"; (3) Engineering level: online scene has strict constraints on end-to-end time delay, throughput and stability, which needs to realize high-frequency update and continuous learning under limited computing power; (4) Decision level: false alarm and missed alarm costs are asymmetric, threshold selection lacks cost sensitivity, and alarm and disposal strategy library does not form a closed loop, leading to "alarm-execution" disconnection and insufficient operability.
[0004] To address the aforementioned issues, existing technologies have attempted to incorporate statistical process control, time-series models such as ARIMA / VAR, tree-based machine learning algorithms, or deep networks for demand forecasting and inventory optimization. However, limitations remain in cross-chain coupling, online adaptation, and coordinated response: First, the lack of hierarchical thresholds and hysteresis mechanisms in system design leads to high levels of hierarchical jitter and false alarm rates. Second, two-dimensional convolutions struggle to effectively integrate the interaction between the time dimension and multiple link dimensions, resulting in insufficient three-dimensional / spatiotemporal modeling. Third, under conditions of concept drift and data distribution changes, models and thresholds lack joint recalibration and online incremental learning strategies. Fourth, standardized interfaces and case retrieval processes have not been established for model output and business processing, hindering knowledge accumulation and reuse. These shortcomings directly impact the overall performance of supply chain systems in real-time monitoring, risk classification, strategy execution, and effect evaluation.
[0005] Therefore, there is an urgent need for a technical solution that achieves a closed loop of "data governance—spatiotemporal fusion—dynamic assessment—tiered alerting—case-based decision-making—feedback learning" under a unified system. This solution involves building an interpretable and traceable comprehensive risk measurement system based on an indicator system covering supply, demand, and logistics; employing joint modeling of 2D / 3D convolutional and temporal networks to uncover cross-chain linkage patterns; introducing hierarchical thresholds, hysteresis bands, and cost-sensitive optimization to reduce tiered jitter and false negative / false negative costs; achieving continuous learning and threshold recalibration in online scenarios through incremental / migratory / distillation mechanisms; and integrating early warning and case-based strategy libraries to output executable handling suggestions and priority rankings, forming a closed-loop control system from monitoring to execution. This invention proposes specific implementation methods and system processes to address the above technical needs, aiming to improve the resilience and operational efficiency of the supply chain system in complex environments. Summary of the Invention
[0006] This invention aims to provide a cluster-based supply chain emergency early warning and coordinated decision-making method based on deep learning. It is designed for complex supply networks involving multiple enterprises, regions, and links, and addresses the shortcomings of existing technologies in high-dimensional heterogeneous data fusion, real-time risk classification, model adaptation to environmental changes, and low-latency coordinated response. This reduces false alarms / missed alarms and classification jitter, and improves the accuracy, interpretability, and executability of early warnings.
[0007] To address the aforementioned issues, this invention proposes the following overall solution, detailed in the invention description: Based on risk levels and an indicator system, and combining time-space multi-source data fusion and deep time-series assessment, it triggers linked alarms and case-based decision support when thresholds are reached, and achieves closed-loop iterative optimization through continuous monitoring, feedback learning, and online adaptation. Compared to existing methods, this invention reduces hierarchical jitter and cost risks through "layered thresholds + hysteresis bands" and cost-sensitive optimization; identifies cross-chain linkage patterns of "time × number of chains × dimension" through 2D / 3D convolution; adapts to distribution drift through online / incremental learning and threshold recalibration; and transforms early warnings into executable strategies using a case-based library, significantly improving supply chain resilience and decision-making efficiency.
[0008] Step 1: Establish a risk level and evaluation index system. First, systematically characterize and structurally model the three risk sources (supply, demand, logistics, and log) and their transmission paths within the enterprise and across organizations. This modeling covers order fulfillment, production, supply, sales coordination, transportation, warehousing operations, and external environmental shocks, unifying risk control signals from different sources into a multi-dimensional index space. Based on this index space, a five-level risk mechanism (normal, roughly normal, alert, abnormal, severe) is constructed, specifying transfer rules and hysteresis intervals between levels to avoid frequent jumps caused by short-term noise. To facilitate online calculation, a multi-dimensional index vector x is collected and aggregated at each discrete observation time t. t This vector is organized hierarchically by subject area sub-processes and data quality, including both raw metrics and derived ratios, as well as cross-sectional comparison indicators, to comprehensively reflect the operational status of both the supply and demand sides and the logistics side. A comprehensive risk score R is calculated for the above vector under the influence of the parameter weight vector w. t As shown in equation (1) (1).
[0009] Where t represents the discrete time index or detection period, for example, a time slot of fifteen minutes, thirty minutes, or one hour. Each component of the column vector composed of J indicators collected or summarized at time t These can be derived from third-party IoT sensor platforms or statistical data derived from business systems. The non-negative weight vector, matched with each indicator, represents the relative impact of each indicator on overall risk. To ensure interpretability, it can be normalized. , The risk mapping function maps the real-number field score to a specified interval while maintaining monotonicity. A common form is... When a function or piecewise linear function is used hour .
[0010] To determine the dividing point of the five-level mechanism, this invention is based on a comprehensive risk score R, using historical observations or observations within a sliding detection window. t Empirical distribution determines the threshold set And each level corresponds to a specific grade. Specifically, the threshold is determined using the quantile method, as shown in equation (2). (2)
[0011] in Indicates score Inverse mapping of the empirical cumulative distribution function For a pre-defined quantile sequence, such as ({0.2, 0.4, 0.6, 0.8}), the score is divided into five segments (K=5) to represent the number of levels. This method can adapt to the statistical scale of scores under different organizational industries and seasons, making the threshold objective and transferable.
[0012] To balance the costs of false positives and false negatives, a cost-sensitive optimization criterion is introduced in the threshold setting stage. The threshold is solved by annotations or weak supervision labels in the backtesting window. Its objective function is shown in equation (3). (3).
[0013] in This refers to the number of missed cases that occurred within the backtesting window, i.e., the number of samples with a high true level but a low false positive rate. The number of false positives is the number of samples with a lower true value but a higher false positive value. The unit cost of the two types of errors is given by the business department in combination with the resource constraints, SLA default risk and public opinion cost. The larger the value, the lower the tolerance. The goal of optimization is to find a set of thresholds that minimizes the overall loss under the given cost trade-off.
[0014] During online operation, a hysteresis-based level transfer mechanism is employed to suppress level jitter caused by short-term fluctuations. Specifically, an upgrade threshold is set. With the downgrade threshold And guarantee when When the upgrade is triggered Downgrading is allowed at times The range remains unchanged. Thresholds and hysteresis widths are stratified and calibrated according to industry, region, season, and other dimensions, and consistency checks are performed. Tolerances are allowed for holidays, extreme weather, and rare dates. In addition, data quality gates and small sample quantile smoothing strategies are combined to ensure that the classification is stable, interpretable, and traceable.
[0015] To ensure dimensional comparability among indicators from different sources and improve fusion quality, the original indicators were standardized after unifying their definitions. The value of the j-th indicator at time t... The Z-transform normalization is shown in equation (4). (4).
[0016] in The historical mean of indicator j over the base period The standard deviation is obtained by this transformation. Dimensionless quantities with zero mean and unit variance can be directly incorporated into multi-index fusion.
[0017] After standardization, the original fusion score is obtained by weighted summation of multiple indicators in a linearly interpretable manner, as shown in equation (5). (5).
[0018] in The non-negative weight of the j-th indicator reflects its contribution to the overall risk. The weight can be assigned by experts or learned automatically using data-driven methods.
[0019] To suppress the impact of short-term noise and random fluctuations on the rating, an exponential smoothing score S is introduced on top of the comprehensive score. t And as one of the main inputs for the final judgment level, its update relationship is as shown in equation (6). (6)
[0020] in A larger smoothing coefficient value indicates a higher weight for the current observation. The initial value for the smoothed score of the previous observation period can be the historical average or estimated using the short-window average.
[0021] To objectively determine the relative importance of each indicator, one implementation method uses the entropy method to estimate the weights. First, the probability distribution of the values of indicator j is estimated using binning or histograms. Calculate information entropy As shown in equation (7). (7)
[0022] Where m is the number of bins, a larger value indicates higher resolution and lower information entropy, meaning the index shows more significant differences between samples and contributes more information. Subsequently, the weights are obtained by inverse normalization of the entropy as shown in equation (8). (8)
[0023] The denominator is the sum of the anti-entropy of all indicators to ensure that the sum of the weights is one.
[0024] Step 2: Processing distributed high-dimensional data streams of clustered supply chains based on spatiotemporal data fusion and deep convolutional networks. In this step, the indicator tensor is modeled as a high-order array containing time, link, spatial, and channel dimensions. Local spatial structure and inter-period interaction features are extracted through joint two-dimensional and three-dimensional convolutions, supplemented by normalized residual connections and channel attention to enhance expressive power. The training process adopts a cosine annealing learning rate strategy, and its scheduling with each round e is as shown in equation (9). (9)
[0025] in Initial learning rate Let E be the lower bound of the learning rate and the total number of training epochs. This strategy maintains a large learning rate in the early stages of training to facilitate exploration, and gradually decreases it in the later stages to achieve stable convergence. The overall objective function is defined by the task loss. Based on this, an L2 regularization term is added to suppress overfitting as shown in equation (10). (10)
[0026] in Trainable parameters of the network The regularization coefficient is chosen by the validation set.
[0027] When the network undertakes self-supervised or semi-supervised tasks such as missing completion and anomaly detection, an autoencoder reconstruction loss can be introduced to improve the ability to characterize normal patterns, as shown in Equation (11). (11)
[0028] in For input samples Output for network reconstruction To balance the coefficients, by minimizing this loss, the network can learn the mainstream structure of the data, and once a distribution deviation occurs, it can be identified through reconstruction error.
[0029] In specific convolution calculations, two-dimensional convolution is used to extract the spatial local structure within a single period, and its calculation expression is shown in equation (12). (12)
[0030] Where (x) is the input tensor, y is the output tensor, (K) is the convolution kernel, (i,j) is the spatial coordinate, (u,v) is the relative displacement, and (c,c') is the channel index. 3D convolution further introduces a time window to extract the interactive features of time and space, and its calculation formula is shown in equation (13). (13)
[0031] in The relative time displacement can be captured by sliding on the time axis to capture the time dependence caused by link fluctuations and scheduling changes.
[0032] To address the continuous evolution of the business environment, this invention introduces a dynamic optimization and distribution drift adaptive mechanism during the online phase. Network parameters are continuously adjusted based on the actual results returned from the alarm and handling closed loop, employing cost-sensitive empirical risk minimization as its objective, as shown in equation (14). (14)
[0033] in and They represent the parameters respectively. The count or ratio of false negatives and false positives is determined. To accelerate the adaptation and suppress overfitting, a knowledge distillation strategy is adopted to transfer the representations of the historical stable model or the larger teacher model to the current online student model. The distillation loss is as shown in Equation (15). (15)
[0034] in and These are the log-odds vectors or latent space embeddings of the teacher and student networks on sample i, respectively. The classification output uses Softmax with temperature T to control the smoothness of the probability distribution, as expressed in equation (16). (16)
[0035] in The predicted probability of the kth level To address the issue that the logarithmic probability T is a temperature parameter, when T is greater than one, the probability is smoother and more conducive to distillation.
[0036] Step 3: Perform dynamic evaluation and probability-based ranking using deep temporal networks. Feed the fused features from Step 2 into sequence models such as LSTM or Transformer to establish intertemporal dependencies, thereby obtaining a temporally consistent probability sequence of ranks. And combine the hysteresis rule defined in step one with the exponential smoothing score S t Implement final-level decision-making. To avoid frequent alarms caused by short-term noise, a multi-scale rolling window is used to smooth and verify the consistency of probabilities and scores. For objects that reach the attention level or above, the impact domain and handling priority are further estimated, and a quantitative prediction of recovery time (TTR) is given to provide a basis for resource orchestration.
[0037] Step 4: Trigger linked alarms and case-specific decision support, and perform continuous monitoring and visualization. When an object reaches the attention level or higher, the system automatically generates a comprehensive score R that includes key indicators for the current level. t With S t and probability vector The system generates alarm reports and pushes summaries of the causes affecting affected objects and suggested handling strategies to responsible personnel via email instant messaging or a work order system. Simultaneously, it retrieves historical similar events from the case library and strategy library based on similarity, returning a set of executable strategies and predicted effects to form a closed-loop verification. To achieve multi-dimensional drill-down, a visual mapping from nodes and time to risk scores is constructed in GIS. Its definition is as shown in equation (17). Users can conduct spatial distribution analysis and observe risk hotspots and evolution trajectories by dimensions such as organizational category, region, or channel. (17)
[0038] To achieve policy adaptation, the system collects false alarm and false alarm annotations, handling effects and SLA indicators under time delay constraints during continuous operation, and performs multi-objective optimization. Its overall objective is shown in equation (18). (18)
[0039] in and These are the false alarm rate and the false alarm rate, respectively. For accuracy The weighting coefficients for these three factors are set by the business side based on risk appetite. To accommodate the differentiated needs of different users or scenarios, user preference vectors are used. The personalized increment mapped to the threshold enables the online increment threshold recalibration as shown in equation (19). (19)
[0040] in The preference coefficient vector that needs to be learned This indicates that the new level threshold, influenced by user preferences, is continuously learned to make the threshold more closely match the risk tolerance level and handling capability boundary of business personnel.
[0041] To ensure the verifiability of the data loop, this invention embeds lineage tracking and caliber versioning management into the data link, records all caliber changes and audit logs, and implements stratified smoothing and quality gate control for stratified samples. Abnormal missing values are preprocessed to remove duplicates, noise, and perform consistency checks, thereby ensuring that the data entering the grading link is verifiable, traceable, and auditable.
[0042] Through the above steps, this invention unifies heterogeneous data from supply, demand, and logistics into a computable spatiotemporal representation without altering existing business systems. An online mechanism centered on cost sensitivity and hysteresis judgment enables stable early warning of sudden event risks. Simultaneously, feedback learning and personalized thresholds provide adaptive support for different organizations and scenarios, thereby improving the accuracy and executability of alarms while ensuring interpretability and traceability. Attached Figure Description
[0043] Figure 1 A flowchart illustrating the overall solution for a cluster-based supply chain emergency early warning and coordinated decision-making method based on deep learning;
[0044] Figure 2 This is a technical roadmap for a cluster-based supply chain emergency early warning and coordinated decision-making method based on deep learning.
[0045] Figure 3 A flowchart of a multi-data fusion algorithm for emergency risk early warning decision indicators based on convolutional neural networks;
[0046] Figure 4 This is a flowchart of a sudden event risk prediction algorithm based on a deep recurrent neural network. Detailed Implementation
[0047] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.
[0048] See Figure 1 This invention provides an overall flowchart of a cluster-based supply chain emergency early warning and decision-making method based on deep learning. The process includes steps S101 to S105 sequentially along the data flow direction, and sets up parallel handling branches at the risk assessment node.
[0049] S101 (Indicator Data Collection): Collect multi-source indicator data from the supply, demand, and logistics links of the clustered supply chain to form a raw data stream with timestamps as the primary key. Data sources may include internal business systems, IoT terminals, and publicly available external data.
[0050] S102 (Data Governance and Aggregation): Perform standardization, missing data completion, anomaly detection and standardization on raw data, and aggregate multi-source heterogeneous data into a unified analysis dataset for subsequent model input.
[0051] S103 (Algorithm Analysis and Risk Level Assessment): Based on convolutional neural networks and temporal networks (such as LSTM), spatiotemporal features are extracted and fused from the data in step S102 to calculate the comprehensive risk score R. t and the probability of each level Output the current risk level (L_t). The safety range is determined by a set of hierarchical thresholds. Hysteresis rules are pre-defined and used for stability classification determination.
[0052] S104 (Judgment and Branch Handling): When the assessment results exceed the safety range (e.g., ... or L t When the level is not lower than "Attention", the process enters a parallel branch:
[0053] S104a (Warning Notification): Automatically generates warning information including level, key indicators and brief causes, and pushes it to the corresponding managers / positions;
[0054] S104b (Solution Matching): Retrieves historical events and solutions similar to the current situation from the database / policy library, and returns the matching response strategy and its key execution points. The above two branches can be executed in parallel.
[0055] S105 (Recording and Closing the Loop): Record and archive warnings, actions, and results. Use the feedback for threshold recalibration and online model updates for subsequent rolling evaluation and continuous optimization. The flowchart above visually illustrates the complete closed loop from indicator data collection—spatiotemporal feature analysis and classification—threshold determination—warning notification / solution matching—result write-back. The steps shown can be combined or split according to business scenarios during implementation, and all operate within the constraint of not exceeding the preset end-to-end latency.
[0056] See Figure 2 This embodiment presents a technical roadmap for a cluster-based supply chain emergency early warning and decision-making method based on deep learning. The roadmap, from left to right along the data flow direction, includes: a two-dimensional convolutional feature extraction unit, a three-dimensional convolutional spatiotemporal fusion unit (3D-CNN), a temporal evaluation unit (LSTM), a classification and ranking unit (fully connected + softmax), and a multi-dimensional cloud model case library linkage unit. Specifically, the supply risk dimension, demand risk dimension, and logistics risk dimension from each supply chain are respectively connected to a shared or independent two-dimensional convolutional encoder (CNN), outputting corresponding local feature maps. These features are then time-aligned and stacked into a three-dimensional tensor according to the "link × risk dimension" structure, serving as the input to the 3D-CNN.
[0057] A 3D convolutional spatiotemporal fusion unit performs multi-scale convolution and nonlinear mapping on the 3D tensor to extract the interaction features of "time × chain number × risk dimension" and obtain a spatiotemporal fusion representation. This representation is sequentially fed into a temporal evaluation unit (LSTM), which models the dynamic dependencies across windows and outputs a sequence representation. Subsequently, a classification and rating unit consists of several fully connected layers, with a classification head at the end having 5 softmax nodes, outputting a set of predicted probabilities for five risk levels (normal, mostly normal, attention, abnormal, and severe). When the judgment level reaches a preset threshold (e.g., "attention" level) or the corresponding probability exceeds a threshold. When this happens, two types of parallel actions are triggered: first, an early warning notification, which pushes alarm information containing the level, key indicators and cause summary to designated management personnel; second, a solution matching, which retrieves historical events and their handling solutions that are similar to the current situation from the multi-dimensional cloud model case library and returns the key points for execution.
[0058] To ensure model input caliber and training stability, normalization / standardization is performed on each metric during the training phase, and a sliding time window is used to convert the sequence into supervised learning samples. The window length (W) and step size (h) are configurable. The loss function can be either mean absolute error (MAE) or cross-entropy, and the optimizer preferably uses Adam. During training and validation, the loss curve and performance metrics such as accuracy / recall are recorded, and the optimal weights are saved based on early stopping and learning rate annealing strategies so that they can be reproduced in the testing and online inference phases.
[0059] Figure 2 The technical route shown forms an end-to-end process through "local representation of two-dimensional convolution - cross-chain spatiotemporal fusion of three-dimensional convolution - dynamic modeling of LSTM - five-level probability classification - case library linkage handling". It can output a stable risk level under the condition of meeting online latency constraints and provide actionable strategy suggestions for emergency response. It is suitable for real-time early warning and decision support scenarios in clustered supply chains.
[0060] See Figure 3 This invention provides a multi-data fusion method for emergency risk early warning decision indicators based on convolutional neural networks. It extracts, fuses, and discriminates features from multi-source spatiotemporal data in clustered supply chains to output risk levels and link early warning and response plans. The algorithm is described in detail below:
[0061] Step 1: Data Acquisition and Classification. Training and testing samples are obtained from historical business and monitoring data. This data includes, but is not limited to, order in transit volume, inventory holdings, transportation timeliness, node throughput, anomaly reports, and external event tags. Based on a predetermined evaluation system, the samples are classified into five risk levels: Normal, Generally Normal, Caution, Abnormal, and Severe.
[0062] Step 2: Network parameter initialization. Initialize the weights and thresholds (biases) of the convolutional neural network; preferably, set the initial values to 0 or near-zero constants to ensure that the parameters of neurons in the same layer are consistent in the initial state, which facilitates subsequent end-to-end training;
[0063] Step 3: Two-dimensional convolutional feature extraction. Data from each supply chain is input into a two-dimensional convolutional neural network for feature extraction based on three risk dimensions. To improve stability over time, the 24-hour data is divided into four subsamples every six hours, outputting two-dimensional feature maps corresponding to each risk dimension.
[0064] Step 4: Constructing the 3D Spatiotemporal Fusion Input. The four sets of data from the two supply chains obtained in Step 3, across three risk dimensions, are aggregated daily and used as input to a 3D convolutional neural network. The three axes correspond to the risk dimension, the number of supply chains, and time (four 6-hour windows). The network completes one forward computation each time it reads one day's worth of data.
[0065] Step 5: Hardline Layer Channel Construction. Five channels are extracted from the hardline layer: risk level channel, risk dimension x-gradient channel, supply chain coordinate gradient channel, x-optical flow channel, and y-optical flow channel. The first three channels are directly obtained from feature calculations, while the optical flow channel is extracted twice in adjacent time slices, resulting in a total of 18 feature matrices as input.
[0066] Step 6: First convolutional layer processing. The five channels from step 5 are convolved using 2×2×2 three-dimensional convolutional kernels to obtain feature matrices (20). Multiple three-dimensional convolutional kernels are set to achieve multi-path feature extraction. The number and spatial size of the obtained feature matrices are determined by formulas (21) and (22), respectively. (20)
[0067] In the formula, The xyz eigenvalues are the eigenvalues of the j-th element in the i-th layer of the feature matrix. Since this patent studies a clustered supply chain, using two supply chains as an example, x represents the three risk dimensions of the emergency risk warning decision assessment indicators, y represents the weekly data of the two supply chains, and z represents the number of weeks within a certain period, i.e., the time dimension. P i This represents the size of the 3D convolution kernel in the risk dimension, Q. i R represents the size of the 3D convolution kernel in the dimension of the number of supply chain items. i This represents the size of the 3D convolution kernel in the time dimension. This represents the weights of the convolutional kernel connected to the previous layer. b represents the activation function. ij This indicates the bias of the convolutional layer. Because a single 3D convolutional kernel can only extract features of one class, and the weights of the kernels are shared, this layer can use n different 3D convolutional kernels to extract multiple features, resulting in n feature matrices. Here, two 3D convolutional kernels are used. To ensure that the extracted features tend to be globalized after being extracted layer by layer, convolutional neural networks typically use multiple convolutional layers for repeated extraction. (twenty one)
[0068] In the formula, r1 is the size of the convolution kernel in the time dimension, and n is the number of different convolution kernels. (twenty two)
[0069] In the formula, a and b are the data input sizes, and p1 and q2 are the sizes of the convolution kernel risk dimension and supply chain dimension.
[0070] Step 7: First downsampling layer. Perform max pooling downsampling on the feature matrix (1); after downsampling, the number of matrices remains unchanged, but the length, width and time dimension are reduced to 1 / 2 of the original size to reduce redundancy and enhance translation invariance;
[0071] Step 8: Second convolutional layer augmentation. Perform three-dimensional convolution again on the n downsampled feature matrices; to improve the representation ability, use convolution kernels with different parameter configurations to perform parallel convolution on each group of features, and the number and size of the output feature matrices are determined by formula (23) and formula (24) respectively. (twenty three)
[0072] In the formula, r2 is the size of the convolutional kernel in the time dimension, and m is the number of different convolutional kernels in the current layer. (twenty four).
[0073] In the formula, a2 and b2 are the data input sizes, and p2 and q2 are the sizes of the convolution kernel risk dimension and supply chain dimension, respectively.
[0074] Step 9: Second downsampling layer. Repeat the max pooling downsampling process to further reduce the size of the feature matrix, suppress noise, and preserve the dominant spatiotemporal pattern;
[0075] Step 10: Terminal Convolution Aggregation. A 2×2×1 three-dimensional convolution kernel is used to convolve the output of Step 9 to aggregate local temporal dependencies along the time axis and preserve the differentiated patterns of the risk dimension and the link axis, resulting in the final high-order spatiotemporal feature matrix;
[0076] Step 11: Output layer and nonlinear mapping. The input of the output layer neuron is multiplied by the corresponding weight, summed and then thresholded. The feature response of each risk level is calculated through the activation function, and its calculation expression is shown in formula (25). Then, a nonlinear mapping to the label space is realized through a fully connected layer to output the final risk assessment result. (25)
[0077] Step Twelve: Early Warning Triggering and Plan Linkage. When the predicted risk level reaches "Attention" or above, the system automatically triggers an early warning, sends a notification to management personnel, and retrieves a matching emergency response plan from the database for coordinated execution. To reduce frequent alarms caused by boundary samples, dual thresholds and a minimum duration can be set. Constraints are an optional strategy.
[0078] See Figure 4 This embodiment presents a process for predicting the risk of sudden events based on a deep recurrent neural network (LSTM). The process includes steps S401 to S406 sequentially along the data flow direction, and sets up an optional branch S401a for feature selection and reconstruction, which is used to compress and enhance the original indicators before training.
[0079] S401 Data Preparation and Data Segmentation. Extract multi-source indicators relevant to risk assessment from historical data of the clustered supply chain to construct a chronologically ordered sample set. Segment the data chronologically to form a training set and a test set (optionally, a validation set). Divide the samples into five categories based on risk level labels: normal, largely normal, caution, abnormal, and severe. Preferably, stratified sampling is used to maintain the consistency and representativeness of the sample proportions at each level.
[0080] S401a Feature Selection and Reconstruction (Optional). Performs correlation screening, principal component / autoencoder dimensionality reduction, or business rule reconstruction on the original indicators, outputting a compact and robust feature set to improve the generalization ability and training efficiency of subsequent models.
[0081] S402 Normalization and Standardization Processing. Perform uniform caliber and scaling transformations on all indicators to ensure that the inputs are within the same numerical range (e.g., mean-variance standardization or interval scaling); save the normalization parameters (mean, variance, or scaling factor) to maintain consistency between testing and online inference phases.
[0082] S403 Supervised Learning Sample Construction. The time series is converted into a supervised learning format, using a sliding time window of length (W=5) days as input to predict the risk level label for day (W+1) (i.e., day 6); resulting in a format like... The sample pairs. Sample windows can overlap to increase the effective sample size; for class imbalance, undersampling / oversampling or class weighting can be used to balance the imbalance.
[0083] S404 LSTM Network Initialization. Establish a network structure including an input layer, LSTM hidden layers, and an output layer. Set hyperparameters such as input dimension, number of hidden units, number of layers, and dropout rate. The output layer uses a fully connected classification head with 5 softmax nodes. Parameter initialization can use the Xavier / He method; optionally, load pre-trained weights from a similar scenario and fine-tune them. Define the loss function as Mean Absolute Error (MAE) or its weighted variant, and select Adam as the optimizer. Record the normalization and network hyperparameters for future experiment reproduction.
[0084] S405 Training and Performance Evaluation. Iterative training is performed according to the set number of training epochs and batch size, tracking and saving the loss curves during the training / validation phases; model performance is evaluated on the test set, outputting metrics such as MAE, accuracy / recall, confusion matrix, and (optional) AUC-PR, and plotting the training and testing loss changes to visually represent convergence and generalization; when validation metrics fail to reach preset thresholds, hyperparameter search or early stopping strategies are executed, and a snapshot of the optimal weights is saved.
[0085] S406 Inference Level Determination and Alarm Linkage. The latest normalized input is fed into the trained LSTM to calculate the hidden state and obtain the output logits. After softmax, the predicted probabilities {p} of the five risk levels are obtained. kThe system determines the level based on preset thresholds / delay strategies. If the determination result reaches or exceeds the "attention" level, an alert is automatically triggered: a notification containing the level, key indicators, and a brief explanation of the cause is generated and pushed to the relevant responsible person; at the same time, similar events and response plans are searched in the case and strategy library, and handling suggestions are returned and the execution results are recorded for subsequent online learning and threshold recalibration.
Claims
1. A cluster-based supply chain emergency early warning and coordinated decision-making method based on deep learning, characterized in that, Running on a computer system equipped with an edge data acquisition gateway, unified clock synchronization, message queue or streaming computing engine, CPU / GPU inference module, and GIS rendering module, the method sequentially includes data aggregation and governance, risk level and threshold configuration, fusion scoring and hysteresis judgment, spatiotemporal feature extraction, time series evaluation and probability output, linkage alarm and case-based decision-making, online adaptive and threshold recalibration, and visualization and auditing closed-loop links. Specifically, it involves collecting multi-source heterogeneous data from the supply, demand, and logistics sides through industrial interfaces or APIs and performing standardization and time alignment, while using a sliding window length for missing values. Imputation is performed and robust anomaly handling is applied. The processed structured indicators are written to the in-memory indicator cache. A five-level threshold set is initialized based on the upper and lower quantiles according to historical or sliding window data. And introduces parameters for the cost of false positives and false negatives. Joint optimization yields threshold configurations that can be stratified by industry, region, and season; the governed multidimensional indicators are weighted by vector. Calculate risk score And with smoothing coefficient Exponential smoothing Set upgrade threshold With the downgrade threshold And satisfy ,when Upgrade in time, when When downgraded, when Maintain the level in time; construct higher-order tensors from multi-period, multi-link, and multi-regional indicators. ,in For time steps, For the number of links or business links, For the number of regions or nodes, The system uses the number of indicator channels and a deep network containing 2D and 3D convolutions, normalization, residual connections, and channel attention to extract spatial local structure and spatiotemporal interaction features. These spatiotemporal features are input into a deep temporal network to obtain the probability distribution of future time windows and are combined with the hysteresis judgment to form the final level decision. When an object reaches a preset level or probability threshold, an alarm is generated containing the current level, key indicators, and scoring information. Based on embedded similarity, a handling strategy and priority are retrieved from the case library and strategy library to form an executable solution. Under the condition of meeting end-to-end latency and throughput constraints, the spatiotemporal and temporal models are updated online according to the closed-loop feedback of the handling, and the threshold is recalibrated incrementally according to the user's risk preference. The risk distribution, hotspots, and evolution trajectory of nodes and time dimensions are displayed in GIS, and field lineage and caliber versions are recorded in the data link to generate audit logs to achieve an interpretable and traceable closed loop from monitoring to execution, and under the same computing power and data caliber, the end-to-end alarm latency is no greater than [missing value]. Alarm jitter standard deviation is no greater than and weighted cost The decline is no less than .
2. The method according to claim 1, characterized in that, The missing value handling in the data aggregation and governance process combines sliding window imputation with similar sample imputation, while outlier handling combines quantile truncation with robust statistical correction. Furthermore, the governed data is processed in batches of varying sizes. Write the aforementioned memory metrics cache.
3. The method according to claim 1, characterized in that, The joint optimization of the risk level and threshold configuration process is based on a sliding window dataset, and is achieved by minimizing... get And introduce tolerance terms for holidays and extreme weather scenarios. To improve stability.
4. The method according to claim 1, characterized in that, Risk scoring in the fusion scoring and hysteresis classification process ,in For the standardized index vector, exponential smoothing is used to denote the index. and to meet Hysteresis band suppression level jitter.
5. The method according to claim 1, characterized in that, In the spatiotemporal feature extraction stage, two-dimensional convolution is used for single-period spatial structure modeling, and three-dimensional convolution is used for time-space interaction modeling. The training stage adopts learning rate scheduling and weight regularization, and self-supervised loss or semi-supervised loss based on reconstruction can be introduced to enhance the characterization of normal patterns and the ability to identify distribution deviations.
6. The method according to claim 1, characterized in that, The deep temporal network in the temporal evaluation and probability output stages is a long short-term memory network or a sequence model based on a self-attention mechanism. The probability output adopts a softened distribution with a temperature parameter to improve the uncertainty calibration and knowledge distillation effect.
7. The method according to claim 1, characterized in that, The linked alarm and case decision-making process uses structured event features and embedded representations as retrieval keys, combines resource constraints and processing time limits to generate an execution list, and provides a prediction of the average recovery time to support dynamic adjustment of resource orchestration, linkage scheduling, and strategy priority.
8. The method according to claim 1, characterized in that, The online adaptive and threshold recalibration process is performed periodically or triggered by a specific mechanism. The model update method is any one of incremental learning, transfer learning, or knowledge distillation, and the user's risk preference is mapped to a personalized incremental threshold. To achieve recalibration.
9. The method according to claim 1, characterized in that, The visualization and auditing closed-loop process is configured with quality gates, which include at least missing and abnormal preprocessing, duplicate data removal, cross-system consistency verification and quantile smoothing, and record field sources, processing steps and version changes through data lineage tracing to ensure that the results are verifiable, traceable and auditable.
10. The method according to claim 1, characterized in that, When an object reaches the "attention" level or above, the system estimates the potential impact domain and handling priority, and sorts and dynamically adjusts the strategy in real time based on the level probability and historical handling effects, so as to improve the accuracy of early warning and the feasibility of execution while meeting the time delay constraint.