A warehouse cascading failure prediction method based on propagation probability graphs
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本发明提供一种基于传播概率图的仓储连锁故障预测方法,以解决现有仓储系统的故障诊断技术对设备故障进行检测时,故障传播诊断和预测的准确性不高而导致连锁故障响应滞后,造成产能损失的技术问题,旨在为运维人员提供数据驱动的预测性决策支持
[0006]与现有技术相比,本发明基于传播概率图的仓储连锁故障预测方通过构建表征物理连接关系的物理拓扑有向图与独立的故障传播概率有向图,形成双图架构,在仓储领域实现了设备间故障传播关系的量化建模,使运维决策从经验判断转变为基于传播概率的数据驱动;通过以异常事件关闭为触发条件的在线增量学习过程,利用物理拓扑有向图所约束的候选范围进行传播统计,使传播概率图能够随系统运行持续演化、自动适应设备布局和负载的动态变化,同时,通过物理拓扑可达性候选范围约束确保学习到的传播关系具有仓储物料流转的物理因果基础,有效过滤了设备之间仅因时间巧合产生的伪关联,在设备密集场景下显著提升了故障传播概率有向图学习的准确;且通过基于故障传播概率有向图对指定故障设备进行故障传播分析,能够在故障发生初期即根据传播概率预测最可能的传播路径,或输出量化的下游风险设备列表和上游原因设备列表,为运维人员提供前瞻性的决策支持,从而缩短故障处置时间、降低连锁故障造成的产能损失。
Smart Images

Figure CN122365318B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent diagnostic technology for warehousing and logistics, and in particular to a method for predicting warehousing chain failures based on propagation probability graphs. Background Technology
[0002] With the rapid development of intelligent manufacturing and e-commerce, the automation level of modern warehousing systems is constantly improving. A large number of automated equipment such as conveyor lines, stacker cranes, elevators, AGVs, and shuttle cars constitute a complex material handling network. These devices form a close upstream and downstream dependency relationship through material flow. When one piece of equipment fails, its impact often spreads along the material flow, leading to system-level cascading failures and causing large-scale capacity losses.
[0003] However, existing fault diagnosis technologies for warehousing systems typically operate on a per-device basis, employing threshold-based alarm mechanisms. Each device is independently assessed for abnormal status. When it is necessary to evaluate the potential impact of a device failure, maintenance personnel can only rely on experience, which is inefficient and fails to provide quantifiable propagation probabilities and expected delays. Alternatively, they may rely on manually defined rules or expert knowledge bases, but static rules are difficult to adapt to dynamic changes in equipment layout and load, affecting the accuracy of fault propagation diagnosis and prediction. This leads to delayed responses to cascading faults and resulting in capacity losses. Summary of the Invention
[0004] This invention provides a warehouse cascading failure prediction method based on propagation probability graphs to address the technical problem that existing fault diagnosis technologies for warehouse systems have low accuracy in detecting and predicting equipment faults, leading to delayed response to cascading failures and resulting in capacity loss. The invention aims to provide data-driven predictive decision support for operation and maintenance personnel.
[0005] To address the aforementioned technical problems, this invention provides a method for predicting warehouse cascading failures based on propagation probability graphs, comprising: Construct a physical topology directed graph that represents the physical connection relationship between storage equipment and a fault propagation probability directed graph that shares a set of nodes with the physical topology directed graph but has an independent set of edges; When an abnormal event triggers shutdown, based on the candidate range constrained by the physical topology directed graph, the fault propagation events between devices are counted, and the propagation probabilities in the fault propagation probability directed graph are updated to perform an online incremental learning process; wherein, the candidate range is a preset hop reachable set centered on the device that triggered shutdown in the physical topology directed graph, and the preset hop reachable set is a set of all nodes that can be reached by traversing the directed edges of the physical topology directed graph in both the forward and reverse directions, starting from any node; Based on the directed graph of the fault propagation probability, fault propagation analysis is performed on the specified faulty device to predict the fault propagation path, and / or, output a list of downstream risky devices and / or a list of upstream cause devices.
[0006] Compared with existing technologies, this invention's warehouse cascading failure prediction method based on propagation probability graphs constructs a physical topology directed graph representing physical connections and an independent fault propagation probability directed graph, forming a dual-graph architecture. This enables quantitative modeling of fault propagation relationships between equipment in the warehousing field, transforming operational decisions from experience-based judgments to data-driven decisions based on propagation probability. Through an online incremental learning process triggered by abnormal event shutdown, propagation statistics are performed using the candidate range constrained by the physical topology directed graph. This allows the propagation probability graph to continuously evolve with system operation, automatically adapting to dynamic changes in equipment layout and load. Furthermore, through physical topology... The reachability candidate range constraint ensures that the learned propagation relationships have a physical causal basis in the flow of warehouse materials, effectively filtering out false associations between devices that are merely coincidental in time. This significantly improves the accuracy of directed graph learning of fault propagation probability in dense equipment scenarios. Furthermore, by performing fault propagation analysis on specified faulty devices based on the directed graph of fault propagation probability, the most likely propagation path can be predicted based on the propagation probability in the early stages of a fault, or a quantified list of downstream risk devices and upstream cause devices can be output, providing forward-looking decision support for operation and maintenance personnel. This shortens the fault handling time and reduces the production capacity loss caused by cascading faults. Attached Figure Description
[0007] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a flowchart illustrating a warehouse cascading failure prediction method based on a propagation probability diagram according to the present invention. Figure 2 yes Figure 1 The diagram shows the specific process flow of S30 in the warehouse cascading failure prediction method based on the propagation probability diagram. Figure 3 yes Figure 2 The diagram shows the specific process flow of S34 in the warehouse cascading failure prediction method based on the propagation probability diagram. Figure 4 yes Figure 1 The diagram shows the specific process flow of S40 in the warehouse cascading failure prediction method based on the propagation probability diagram. Figure 5 yes Figure 1The diagram shows another specific process flow diagram of S40 in the warehouse cascading failure prediction method based on the propagation probability diagram. Detailed Implementation
[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0010] The warehousing system includes various automated equipment, including but not limited to: conveyor lines, stacker cranes, elevators, AGVs (Automated Guided Vehicles), and shuttles. Multiple devices are scheduled and controlled by a Warehouse Control System (WCS). The material flow direction between devices is defined through configuration files or routing tables in a database.
[0011] Reference Figure 1 , Figure 1 This is a flowchart illustrating the warehouse cascading failure prediction method based on propagation probability graphs according to the present invention. In the embodiment shown in the figure, the warehouse cascading failure prediction method based on propagation probability graphs includes: S10. Construct a physical topology directed graph representing the physical connection relationship between storage equipment and a fault propagation probability directed graph that shares a set of nodes with the physical topology directed graph but has an independent set of edges.
[0012] In this step, a physical topology directed graph is constructed, with each device in the warehousing system as a node and the material flow direction between devices as directed edges. Where V is the set of nodes, This is a set of directed edges. Specifically, the device routing table, which defines the material flow connection relationships between devices, can be read from the configuration file or database of the warehousing system. Each device is treated as a node V in the graph, and each node records the device identifier and device type attributes. The material flow connections between devices are treated as directed edges in the graph. That is, if material can flow directly from device a to device b, then a directed edge (a, b) is added.
[0013] Preferably, the constructed physical topology directed graph is managed in a thread-safe singleton pattern to ensure global uniqueness and support concurrent access.
[0014] At the same time, initialize a directed graph of fault propagation probability. This graph is related to the physical topology directed graph. They have the same set of nodes V (i.e., all the devices in the warehousing system), but different set of edges. It is maintained independently and initialized as an empty set. Each directed edge (a,b) in the diagram is associated with a propagation edge data structure, which stores propagation statistics, including source device identifier, target device identifier, number of propagation occurrences, total number of observations, the propagation probability that the target device also fails after the source device fails, average propagation delay, list of delayed samples, and last update timestamp.
[0015] In this invention, a dual-graph framework is constructed. A directed graph of physical topology is used to depict the static physical connections between devices as an objective physical constraint. A directed graph of fault propagation probability is used to depict the dynamic fault propagation statistical laws between devices learned from historical data. Both graphs together provide support for subsequent fault propagation analysis and prediction.
[0016] S20. Based on the physical topology directed graph, pre-calculate the preset hop count reachable set for each node.
[0017] In this embodiment, the preset hop count reachable set is a set formed by traversing all nodes reachable by the preset hop count along both the forward and reverse directions of the directed edges of the physical topology directed graph, starting from any node. The preset hop count is preferably two hops. For each device node v in the dataset, its two-hop reachable set is pre-calculated. This serves as a constraint on the candidate range of statistical device-to-device fault propagation events during subsequent online incremental learning. Understandably, the preset number of hops can also be three hops, and can be adjusted according to actual needs.
[0018] This step specifically includes: for each device node v in the directed graph of the physical topology, according to the formula... Calculate its two-hop reachable set; where, Let v be the set of all successor nodes that can be reached from node v by following the positive direction of the directed edge (i.e., the direction from the predecessor to the successor) in one hop. From The set of all nodes that can be reached by continuing forward one hop from each node in the sequence. Let v be the set of all predecessor nodes that can be reached from node v by one hop along the reverse direction of the directed edge (i.e., from the successor to the predecessor). From The set of all nodes reachable by each node in the two-hop reachable set, continuing in the reverse direction with one more hop. In this embodiment, the two-hop reachable set... for , , , The union of four sets, then the set of nodes that remove node v itself.
[0019] In this embodiment, during pre-computation, starting from node v, all successor nodes reachable in one hop can be traversed along the positive direction of the directed edge, denoted as . ; and then from For each node in the list, continue traversing forward to all nodes reachable in one hop, denoted as . This enables the computation of the forward two-hop reachable set; simultaneously, it allows traversing all one-hop reachable predecessor nodes from node v along the reverse direction of the directed edges, denoted as... ; and then from For each node in the list, continue traversing backwards to all nodes reachable in one hop, denoted as . To calculate the reverse two-hop reachable set, and then take... , , , The union of four sets, with node v removed, yields a two-hop reachable set. .
[0020] Understandably, this two-hop reachable set basically excludes physically unreachable device pairs. Subsequently, propagation statistics are only performed on device pairs within the two-hop reachable set, limiting the range of device pairs for online incremental learning to a physically reasonable range, which can effectively filter out false associations caused by temporal coincidences.
[0021] Furthermore, the pre-computed two-hop reachability set for each device node can be cached in, for example, a hash dictionary to avoid repeated graph traversal calculations. In this cache, the key is the device identifier, and the value is the set of two-hop reachable device identifiers. When the physical topology changes (e.g., adding a device or adjusting routes), all cached data can be cleared through a cache invalidation mechanism, and the cached data will be recalculated and cached on the next query.
[0022] S30. When an abnormal event is triggered to shut down, the fault propagation events between devices are counted, and the propagation probability in the directed graph of the fault propagation probability is updated, with the preset hop count reachable set of each device node as a constraint, in order to perform an online incremental learning process.
[0023] In this embodiment, the operating status data of each device in the warehousing system is collected periodically, and abnormal events are generated through state transition detection. When an abnormal event is detected that the device has recovered from an abnormal state to a normal state and is shut down, an online incremental learning process is automatically executed, using the shutdown event as the trigger condition and the preset hop reachability set of the device that triggered the shutdown as the candidate range, to continuously evolve the directed graph of fault propagation probability.
[0024] like Figure 2 As shown, step S30 specifically includes the following steps S31-S34: S31. Periodically collect the operating status data of each device. When a device abnormal event is detected to have recovered from an abnormal state to a normal state, confirm that the abnormal event has been triggered and shut down.
[0025] In this invention, operational status data of each device in the warehousing system is periodically collected. Status transition detection checks whether abnormal event closure conditions are triggered, thereby confirming whether an abnormal event has been closed. Abnormal events can include at least one of equipment failure events, task queuing abnormal events, processing performance abnormal events, and intelligent detection abnormal events. The operational status data can include device operation data, the number of queued tasks (i.e., the number of subtasks currently queued for processing on the device), the number of tasks currently executing, the average processing time (i.e., the average time to complete a subtask within the most recent time window), and currently associated task identifiers (including a list of strings representing the identifiers of the main tasks currently executing or waiting on the device). Understandably, the device operation data can be compared with preset operating thresholds to determine the current operating status of the device, which can be one of normal operation, idle, fault, or blocked.
[0026] Specifically, based on the operating status data and preset state transition detection rules, a sliding window time-series tracker can detect the state transitions of abnormal events of each device. The state transition detection rules include at least the following: when the current operating status of the device is not faulty and the previous cycle was faulty; when the current number of tasks waiting in the queue of the device does not exceed a preset threshold and the previous cycle exceeded it; when the current average processing time of the device does not exceed the product of a preset baseline processing time and a preset multiplier and the previous cycle exceeded it; and / or, when no detection result is received from the AI anomaly detection module marking the device as abnormal and it was marked in the previous cycle, the abnormal event closure condition is confirmed to be triggered.
[0027] In this step, when an abnormal event is detected that the device has been shut down due to recovery from an abnormal state to a normal state, an online incremental learning process is performed based on the information of these shut-down events. Preferably, the online incremental learning process adopts a thread-safe mechanism to ensure that concurrent read and write operations on the directed graph of fault propagation probability do not generate data races in a multi-threaded or multi-process environment, thus guaranteeing the consistency of the directed graph data.
[0028] Furthermore, an event index dictionary can be created for all abnormal events within the current time window, categorized by device identifier. The key is the device identifier, and the value is the list of abnormal events for that device within the window. This accelerates the subsequent event search process, and closed abnormal events can be moved to the closed event list. When triggering the detection of abnormal event closure conditions, the closed event list or the event list within the current time window can be checked first. If either the closed event list or the event list within the current time window is empty, it indicates that the closed abnormal events have already undergone online incremental learning and will not undergo further learning.
[0029] S32. Taking the device that is shut down due to an abnormal event as the source device, obtain each device in the preset hop reach set of the source device based on the physical topology directed graph as a candidate target device, and check whether there is a directed edge between the source device and each candidate target device in the fault propagation probability directed graph, and increment the total number of observations between the source device and each candidate target device by one.
[0030] In this step, for each closed abnormal event (denoted as the source event), its device identifier is extracted as the source device. Extract its timestamp as the source time. And based on the physical topology directed graph, obtain the two-hop reachable set of the source device. For each device in the two-hop reachable set, it is considered a candidate target device. For each candidate target device, it is checked in the directed graph of fault propagation probability whether a directed edge from the source device to the candidate target device already exists. If not, a new propagation edge data structure is created, and the total number of observations for the pair of devices is incremented by one.
[0031] In this invention, when a source device experiences an anomaly, each physically reachable candidate target device is observed to determine whether it also experiences an anomaly. Regardless of whether the candidate target device experiences an anomaly, the observation count for both the source and candidate target devices is increased.
[0032] S33. Within the maximum propagation delay window after the abnormal event of the source device occurs, if a candidate target device has an abnormal event with a timestamp later than that of the source device and meets the timing conditions, then the propagation occurrence count between the source device and the candidate target device is incremented by one, and the current propagation delay sample is recorded; wherein, the same pair of devices is counted at most once in one learning iteration.
[0033] In this step, within a preset maximum propagation delay time (e.g., 120 seconds) after the abnormal event of the source device occurs, it is checked whether each candidate target device in its two-hop reachable set has also experienced an abnormal event.
[0034] Specifically, for each candidate target device, the list of abnormal events for that device within the window is queried, and the difference between the timestamp of each candidate target event and the timestamp of the source device is calculated, i.e., the propagation delay. ,in, For the timestamp of the candidate target event. The timestamp of the source device is used. A valid propagation event is confirmed when Δt meets the following two conditions: a) Δt > 0, that is, the time of the anomaly of the target device is later than the time of the anomaly of the source device, which satisfies the causal timing requirement and excludes the self-correlation of events at the same time; b) Δt ≤ preset maximum propagation delay time, that is, the propagation delay does not exceed the preset maximum propagation delay threshold (e.g., 120 seconds). Anomalies exceeding this time are more likely to be independent events rather than causal propagation. That is, the difference between the timestamp of the anomaly event of the candidate target device and the timestamp of the anomaly event of the original device satisfies the timing condition of not exceeding the maximum propagation delay.
[0035] When an abnormal event occurring between the source device and any candidate target device meets the above conditions, the propagation count for that pair of devices is incremented by one, and the current propagation delay value is adjusted accordingly. Add it to the delayed sample list of that side. In this step, the principle of recording at most one propagation event per source device-candidate target device pair in one learning iteration is followed. That is, when a candidate target device has multiple abnormal events that meet the timing conditions within the maximum propagation delay window, only the abnormal event with the earliest timestamp is counted as one propagation between the pair of devices. This avoids the problem of multiple abnormal events generated by frequent interruptions of the candidate target device within the window, which could artificially inflate the propagation count and cause statistical bias. For example, in a certain learning iteration, source device A experiences an abnormality at 10:00:00 and recovers at 10:05:00, triggering learning. Its downstream device B experiences three abnormal interruptions at 10:01:00, 10:02:30, and 10:04:00. The system only counts the first abnormality at 10:01:00 as one valid propagation from A to B, and the subsequent two are not counted.
[0036] In this embodiment, the delay sample list employs a sliding window management strategy, recording only the most recent N (e.g., 50) propagation delay samples for each pair of devices. When a new propagation delay sample is added and the list length exceeds the preset sample capacity limit N, the oldest delay sample is automatically discarded, ensuring the list length never exceeds N. This strategy maintains statistical validity while controlling memory usage and naturally biases the calculated average propagation delay towards the most recent system operating conditions, thus providing a certain degree of timeliness.
[0037] S34. When the cumulative total number of observations between the source device and any candidate target device reaches a preset minimum observation threshold, calculate the propagation probability and average propagation delay between the pair of devices, and update the data of the corresponding edge in the directed graph of the fault propagation probability.
[0038] like Figure 3 As shown, step S34 specifically includes the following steps S341-S343: S341. Maintain a dirty edge set. In each learning iteration, add the directed edges whose propagation count or total observation count changes to the dirty edge set. S342. After the current learning iteration ends, traverse each directed edge in the dirty edge set, and for the edges whose cumulative total number of observations reaches the preset minimum number of observations threshold, recalculate the propagation probability and average propagation delay and update them in the fault propagation probability directed graph.
[0039] In this invention, after each learning iteration, only the edges in the dirty edge set are traversed for probability updates, rather than traversing all edges. The average propagation delay is calculated based on the current delay sample list for that edge, and is the arithmetic mean of all retained samples in the list.
[0040] In this step, according to the formula Calculate the propagation probability, and according to the formula Calculate the average propagation delay, where, Let be the propagation probability from device a to device b. This represents the number of propagation events between device a and device b. This represents the total number of observations between device a and device b. Let S be the average propagation delay of the device to (a, b), and S be the set of retained delay samples. For the sample size, This is the k-th delayed sample.
[0041] Furthermore, after the calculation is completed, the updated propagation probability and average propagation delay are written into the data structure of the corresponding edges of the directed graph of the fault propagation probability. Based on the updated edge data, the directed graph data structure of the directed graph of the fault propagation probability is reconstructed, and edges with a propagation probability greater than zero and a cumulative total number of observations reaching the threshold are added to the directed graph data structure. The edge weight is set to the propagation probability value, and attributes such as average delay and propagation count are recorded to ensure that the graph data structure used for subsequent fault propagation analysis (such as path search and risk query) is consistent with the latest data.
[0042] Specifically, the formula for calculating the edge weight of each edge in the directed graph of the fault propagation probability is as follows: ; in, Let be the edge weight from device a to device b. For the probability of propagation, Average propagation delay, To prevent extremely small positive numbers from being divided by zero.
[0043] S343. After the update is complete, clear the dirty edge set.
[0044] In this step, the dirty edge set is cleared to prepare for the next learning iteration.
[0045] For steps S341-S343 above, this embodiment employs a dirty edge set optimization strategy for incremental probability updates. In each learning iteration, whenever the total number of observations or propagation occurrences of a directed edge changes due to the aforementioned statistical operations, the edge's identifier (a key composed of the source device and the candidate target device) is added to the dirty edge set. After the current learning iteration ends, only the edges in the dirty edge set are traversed to recalculate probabilities and delays, rather than traversing all edges in the directed graph of fault propagation probability, thus significantly reducing computational overhead.
[0046] Furthermore, in some embodiments, the warehouse cascading failure prediction method based on the propagation probability graph can also persistently store the updated directed graph of failure propagation probability. Specifically, an atomic write strategy is adopted to serialize the data of the directed graph of failure propagation probability into a structured data format and write it to a temporary file. The temporary file is then used to replace the original persistent storage file through a file renaming operation. That is, all edge data is serialized into a structured data format, and the most recent 20 delayed samples for each edge are preferably retained to control the file size. The data is first written to a temporary file (with a temporary suffix), and after successful writing, the file is renamed to replace the original file, ensuring that the data is not corrupted in the event of a system failure during the writing process. Historical data can be loaded from the persistent file when the system starts, enabling knowledge retention across restarts.
[0047] S40. Based on the directed graph of the fault propagation probability, perform fault propagation analysis on the specified faulty equipment to predict the fault propagation path, and / or output a list of downstream risky equipment and / or a list of upstream cause equipment.
[0048] In this invention, fault propagation analysis may include fault propagation path prediction, downstream risk assessment, and / or upstream cause tracing.
[0049] like Figure 4 As shown, step S40, which involves performing fault propagation analysis on a specified faulty device based on the directed graph of the fault propagation probability to predict the fault propagation path, specifically includes the following steps S41a-S44a: S41a. Starting from the specified faulty device, a greedy search algorithm is used to expand the path in the directed graph of the fault propagation probability.
[0050] In this step, a greedy search algorithm is used to predict the fault propagation path based on the learned directed graph of fault propagation probability in order to assess the potential impact of equipment failure.
[0051] S42a. In each expansion step, obtain all successor nodes of the current node, and select the node with the highest propagation probability that has not been visited as the next hop; when there are multiple successor nodes with the same propagation probability, select the node with the shortest average propagation delay as the next hop.
[0052] In this step, all successor nodes of the current node (i.e., the specified faulty device) in the directed graph of fault propagation probability are traversed. The node with the highest propagation probability that has not been visited is selected as the next hop. When there are multiple successor nodes with the same propagation probability, the node with the shortest average propagation delay is selected as the next hop to ensure the determinism of the decision. At the same time, the propagation probability and average propagation delay of the node selected as the next hop are recorded.
[0053] S43a. Repeat the above extension step S42a until the termination condition is met. The termination condition includes: there is no unvisited successor node, the propagation probability of the next hop is lower than a preset probability threshold, or the number of extended hops reaches a preset maximum number of hops.
[0054] In this step, the search process continues until the propagation probability of the next hop is lower than a preset probability threshold (such as 0.1), or the preset maximum number of hops is reached, or there are no unvisited successor nodes.
[0055] S44a. Output the propagation path from the specified faulty device to each node, and the propagation probability and cumulative propagation delay of each segment on the path.
[0056] This step outputs the complete propagation path from the specified faulty device to each node via it, the propagation probability at each step, and the cumulative propagation delay. The cumulative propagation probability of this path can be calculated using the following formula: ; in, For source devices, For the end point of the path, Let be the propagation probability between adjacent nodes on the path. The time complexity of this greedy search algorithm is O(maximum number of hops × ...). ),in The maximum out-degree of a node in the directed graph of fault propagation probability is usually a single digit in warehouse topology, so the algorithm is extremely efficient and can be completed in milliseconds.
[0057] Furthermore, such as Figure 5 As shown, step S40, which involves performing fault propagation analysis on a specified faulty device based on the directed graph of the fault propagation probability to output a list of downstream risky devices, specifically includes the following steps S41b-S42b: S41b: Obtain all successor nodes of the specified faulty device in the directed graph of fault propagation probability, along with their corresponding propagation probabilities and average propagation delays. Sort them in descending order of propagation probability and output a list of downstream risk devices consisting of the top K highest-risk downstream devices.
[0058] In this invention, for downstream risk assessment, all successor nodes of a specified faulty device in a directed graph of propagation probability, along with their propagation probabilities and average propagation delays, are obtained and sorted in descending order of propagation probability. A list of the top K highest-risk downstream devices is then output. Each list element may contain a candidate target device identifier, propagation probability, average propagation delay, historical propagation count, and a data source marker. The data source marker can be "learned," indicating that the propagation probability and average propagation delay between the device pairs are statistically derived from historical operational data through an online incremental learning process.
[0059] S42b. When the specified faulty device does not have a valid outgoing edge in the directed graph of the propagation probability, it reverts to the directed graph of the physical topology and outputs the devices in the set of devices that can be reached by the device's preset hop count as candidate downstream risk devices.
[0060] In this step, a valid outgoing edge refers to a directed edge in the fault propagation probability directed graph that simultaneously satisfies the following two conditions, pointing from one device node to another: (1) The propagation probability of this edge is greater than zero, that is, there is indeed a statistically recorded fault propagation event between the device pairs; (2) The cumulative total number of observations on this side reaches the preset minimum number of observations threshold, that is, the propagation probability value has statistical credibility.
[0061] In this invention, when there are no valid edges originating from a specified faulty device in the directed graph of fault propagation probability (possibly because the device is still in the cold start phase or has not accumulated enough data), the system reverts to the directed graph of physical topology. Devices in the two-hop reachable set are output as candidate downstream risk devices, and their propagation probability is marked as zero. The data source can be marked as the physical topology. This backoff mechanism is used to solve the cold start problem in the early stages of online learning, ensuring that even when propagation statistics are insufficient, qualitative reference results based on physical topology can still be provided to maintenance personnel.
[0062] Preferably, in this step, upstream cause tracing can also be performed. This involves analyzing the fault propagation of a specified faulty device based on the directed graph of fault propagation probability, and outputting a list of upstream cause devices. Specifically, all predecessor nodes of the specified faulty device in the directed graph of fault propagation probability and their corresponding propagation probabilities are obtained, sorted in descending order of propagation probability, and a list of upstream cause devices consisting of the top K upstream devices is output. Each list element may include the source device identifier, propagation probability, average propagation delay, and historical propagation count.
[0063] Based on the above steps, the fault propagation probability directed graph can be used to predict the fault propagation path, assess downstream risks, and / or trace upstream causes for a specified faulty device. In this embodiment, the fault propagation probability directed graph can be used to predict the fault propagation path, assess downstream risks, and trace upstream causes for a specified faulty device. When a fault is detected in a device, maintenance personnel can immediately query the list of downstream risky devices for that device and take preventative measures (such as adjusting routes, reducing load, etc.) for high-risk downstream devices in advance, intervening before the cascading fault spreads. Furthermore, the list of upstream cause devices can be queried to assist maintenance personnel in quickly locating the possible source of the fault.
[0064] In some embodiments, the directed graph of fault propagation probability can also be exported and displayed. This involves providing the structured data of the directed graph for front-end visualization, including a list of all nodes involved in the propagation relationship, a list of valid edges sorted in descending order of propagation probability, and for each edge, its source and target devices, propagation probability, average latency, number of propagations, total number of valid edges, and total number of observations. Furthermore, a summary of the operational status of the directed graph can be displayed, including the total number of edges containing non-compliant edges, the number of valid outgoing edges, the average propagation probability, the highest propagation probability, and the number of nodes and edges in the current directed graph data structure. This allows operations personnel to monitor the progress of the online learning process and the data quality of the directed graph of fault propagation probability.
[0065] In summary, the warehouse cascading failure prediction method based on propagation probability graphs of this invention constructs a physical topology directed graph representing physical connections and an independent fault propagation probability directed graph, forming a dual-graph architecture. This enables quantitative modeling of fault propagation relationships between equipment in the warehousing field, transforming operational decisions from traditional qualitative experience-based judgments to data-driven decisions based on propagation probability. Through an online incremental learning process triggered by abnormal event shutdown, propagation statistics are performed using the candidate range constrained by the physical topology directed graph. This allows the propagation probability graph to continuously evolve with system operation, automatically adapting to dynamic changes in equipment layout and load, reducing maintenance costs by minimizing manual maintenance. Meanwhile, by constraining the candidate range of physical topology reachability, the learned propagation relationships are ensured to have a physical causal basis in the flow of warehouse materials, effectively filtering out false associations between devices that are merely coincidental in time. This significantly improves the accuracy of directed graph learning of fault propagation probability in dense device scenarios. Furthermore, by providing propagation path prediction and bidirectional risk analysis based on the directed graph of fault propagation probability, the most likely propagation path can be predicted based on the propagation probability at the early stage of a fault, and / or a quantified list of downstream risk devices and upstream cause devices can be output, providing forward-looking decision support for operation and maintenance personnel, thereby shortening fault handling time and reducing capacity loss caused by cascading faults.
[0066] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0067] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for predicting cascading failures in warehouses based on propagation probability graphs, characterized in that, include: Construct a physical topology directed graph that represents the physical connection relationship between storage equipment and a fault propagation probability directed graph that shares a set of nodes with the physical topology directed graph but has an independent set of edges; When an abnormal event triggers shutdown, based on the candidate range constrained by the physical topology directed graph, the fault propagation events between devices are counted, and the propagation probabilities in the fault propagation probability directed graph are updated to perform an online incremental learning process; wherein, the candidate range is a preset hop reachable set centered on the device that triggered shutdown in the physical topology directed graph, and the preset hop reachable set is a set of all nodes that can be reached by traversing the directed edges of the physical topology directed graph in both the forward and reverse directions, starting from any node; Based on the directed graph of the fault propagation probability, fault propagation analysis is performed on the specified faulty device to predict the fault propagation path, and / or, output a list of downstream risky devices and / or a list of upstream cause devices. Specifically, the step of, when an abnormal event triggers shutdown, statistically analyzing the fault propagation events between devices based on the candidate range constrained by the physical topology directed graph, and updating the propagation probability in the fault propagation probability directed graph, includes: Periodically collect the operating status data of each device, and when a device returns to normal after an abnormal event is detected, confirm that the abnormal event has been triggered and shut down. Taking the device that is shut down due to an abnormal event as the source device, each device in the preset hop reach set of the source device is obtained based on the physical topology directed graph and used as a candidate target device. In the fault propagation probability directed graph, it is checked whether there is a directed edge between the source device and each candidate target device, and the total number of observations between the source device and the candidate target device is incremented by one. Within the maximum propagation delay window after the abnormal event of the source device occurs, if a candidate target device has an abnormal event with a timestamp later than that of the source device and meets the timing conditions, the propagation count between the source device and the candidate target device is incremented by one, and the current propagation delay sample is recorded; wherein, the same pair of devices is counted at most once in one learning iteration; When the total number of observations between the source device and any candidate target device reaches a preset minimum observation threshold, the propagation probability and average propagation delay between the pair of devices are calculated, and the data of the corresponding edge in the directed graph of the fault propagation probability is updated.
2. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 1, characterized in that, The following is included prior to the statement that the shutdown is triggered by an abnormal event: Based on the physical topology directed graph, a preset hop count reachable set for each node is pre-calculated to serve as a candidate range constraint for statistically analyzing fault propagation events between devices during the online incremental learning process.
3. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 2, characterized in that, The preset hop count is two hops, wherein the pre-calculation of the preset hop count reachable set for each node based on the physical topology directed graph specifically includes: For each device node v in the directed graph of the physical topology, according to the formula Calculate its two-hop reachable set; where, Let v be the set of all successor nodes that can be reached from node v by one hop along the directed edge in the positive direction. From The set of all nodes that can be reached by continuing forward one hop from each node in the sequence. Let v be the set of all predecessor nodes that can be reached from node v by one hop along the directed edge in the opposite direction. From The set of all nodes that can be reached by continuing in the reverse direction with one more hop from each node in the array.
4. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 1, characterized in that, The formula for calculating the propagation probability is: in, Let be the propagation probability from device a to device b. This represents the number of propagation events between device a and device b. This represents the total number of observations between device a and device b.
5. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 1, characterized in that, When the cumulative total number of observations between the source device and any candidate target device reaches a preset minimum observation threshold, the propagation probability and average propagation delay between the pair of devices are calculated, and the data of the corresponding edges in the directed graph of the fault propagation probability are updated, specifically including: Maintain a dirty edge set. In each learning iteration, add directed edges whose propagation count or total observation count changes to the dirty edge set. After the current learning iteration is completed, each directed edge in the dirty edge set is traversed. For the edges whose cumulative total number of observations reaches the preset minimum number of observations threshold, the propagation probability and average propagation delay are recalculated and updated in the fault propagation probability directed graph. After the update is complete, clear the set of dirty edges.
6. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 1, characterized in that, The step of performing fault propagation analysis on a specified faulty device based on the directed graph of fault propagation probability to predict the fault propagation path specifically includes: Starting from a specified faulty device, a greedy search algorithm is used to expand the path in the directed graph of the fault propagation probability. In each expansion step, obtain all successor nodes of the current node, and select the node with the highest propagation probability that has not been visited as the next hop; when there are multiple successor nodes with the same propagation probability, select the node with the shortest average propagation delay as the next hop. Repeat the above extension steps until the termination conditions are met. The termination conditions include: there are no unvisited successor nodes, the propagation probability of the next hop is lower than a preset probability threshold, or the number of extended hops reaches a preset maximum number of hops. Output the propagation path from the specified faulty device to each node, as well as the propagation probability and cumulative propagation delay of each segment on the path.
7. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 1, characterized in that, The step of performing fault propagation analysis on a specified faulty device based on the directed graph of fault propagation probability to output a list of downstream risky devices specifically includes: Obtain all successor nodes of the specified faulty device in the directed graph of fault propagation probability, along with their corresponding propagation probabilities and average propagation delays. Sort them in descending order of propagation probability and output a list of downstream risk devices consisting of the top K highest-risk downstream devices. When the specified faulty device has no valid outgoing edges in the directed graph of propagation probability, it reverts to the directed graph of physical topology and outputs devices in the set of devices reachable by the device's preset hop count as candidate downstream risk devices.
8. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 1, characterized in that, The step of performing fault propagation analysis on a specified faulty device based on the directed graph of fault propagation probability to output a list of upstream cause devices specifically includes: Obtain all predecessor nodes of the specified faulty device in the directed graph of fault propagation probability and their corresponding propagation probabilities, sort them in descending order of propagation probability, and output a list of upstream cause devices consisting of the top K upstream devices.
9. The warehouse cascading failure prediction method based on propagation probability graphs as described in claim 1, characterized in that, The warehouse cascading failure prediction method based on propagation probability graphs further includes: An atomic write strategy is adopted to serialize the data of the directed graph of the fault propagation probability into a structured data format and write it to a temporary file. The original persistent storage file is replaced by the temporary file through a file renaming operation.
Citation Information
Patent Citations
Power supply equipment fault prediction method and device based on deep learning
CN121278536A
Fault diagnosis method and device, storage medium and electronic equipment
CN121585525A