Warehouse equipment system fault chain detection and root cause positioning method and device
Patent Information
- Application Number
- CN202610847888.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-12
AI Technical Summary
[0003]本发明提供一种仓储设备系统级故障链检测及根因定位方法、装置,以解决现有技术中无法实现仓储设备系统级故障链的自动检测还原,导致根因定位困难、故障响应时间显著延长,系统停机损失扩大的技术问题
[0006]与现有技术相比,本发明仓储设备系统级故障链检测及根因定位方法通过事件分组合并算法结合时序接近性与物理拓扑可达约束的双重条件进行跨设备异常事件聚类,有效过滤了时间巧合导致的伪关联,实现了系统级故障链的高精度自动发现,且通过融合时序先发性、图结构中心性和历史故障传播概率三个维度的加权评分,从因果时序、结构位置和历史规律三方面协同定位根因,提高根因定位准确率,实现精准定位根因,并结合时间戳推断故障传播路径,还可基于各故障链的路径指纹编码实现传播方向敏感的模式匹配,从而智能识别反复性故障并提前预警。
Smart Images

Figure CN122388880B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of warehousing and logistics technology, and more specifically to a method and apparatus for detecting and locating root causes of system-level fault chains in warehousing equipment. Background Technology
[0002] With the rapid development of intelligent manufacturing and e-commerce, the automation level of modern warehousing and logistics systems is constantly improving. Numerous automated equipment, such as conveyor lines, stacker cranes, elevators, AGVs (Automated Guided Vehicles), and shuttles, constitute complex material handling networks. These devices form close upstream and downstream dependencies through material flow. When one device malfunctions or experiences performance abnormalities, its impact often propagates along the material flow to related upstream and downstream devices, leading to system-wide chain reactions. However, existing fault diagnosis technologies for automated warehousing systems typically operate on a single device basis, employing threshold-based alarm mechanisms. Each device is independently assessed for abnormality. If a stacker crane malfunctions, causing congestion on its upstream conveyor lines, multiple independent alarm messages will be generated. Maintenance personnel find it difficult to reconstruct the fault propagation chain from a large number of isolated alarms, resulting in difficulties in root cause localization, significantly prolonged fault response time, and increased system downtime losses. Summary of the Invention
[0003] This invention provides a method and apparatus for detecting and locating system-level fault chains in warehousing equipment, thereby solving the technical problem that the prior art cannot achieve automatic detection and reconstruction of system-level fault chains in warehousing equipment, which leads to difficulties in locating root causes, significantly extended fault response time, and increased system downtime losses.
[0004] To address the aforementioned technical problems, in a first aspect, the present invention provides a method for system-level fault chain detection and root cause localization in warehousing equipment, comprising: Periodically collect operational status data from various devices in the warehousing system; Based on the operating status data of each device, detect whether a state transition event has occurred in each device, maintain an active abnormal event set based on the state transition events, and generate an event record containing device identifier, event type, and timestamp; Active abnormal events within a preset time window are acquired. Based on the event grouping and merging algorithm, the active abnormal events are clustered under the dual constraints of temporal proximity and physical topological reachability to construct a fault chain. A three-dimensional weighted root cause scoring algorithm that integrates temporal pre-emptivity, graph structure centrality, and fault propagation probability is used to calculate the root cause score of each device in each fault chain. The root cause device is identified based on the root cause score, and the fault propagation path is inferred from the root cause device based on the root cause score and timestamp. Based on the device sequence in the fault chain, a path fingerprint code is generated according to the fault propagation path. Fault chains with the same path fingerprint code are grouped into the same fault mode, and a recurring fault mode alarm is triggered when preset conditions are met.
[0005] Secondly, a system-level fault chain detection and root cause localization device for warehousing equipment is provided, comprising: The data acquisition module is used to periodically collect operating status data from various devices in the warehousing system. The timing tracking module is used to detect whether a state transition event has occurred in each device based on the operating status data of each device, maintain an active abnormal event set based on the state transition event, and generate an event record containing device identifier, event type and timestamp. The fault chain clustering module is used to obtain active abnormal events within a preset time window. Based on the event grouping and merging algorithm, the active abnormal events are clustered under the dual constraints of temporal proximity and physical topological reachability to construct a fault chain. The root cause localization module is used to calculate the root cause score of each device in each fault chain using a three-dimensional weighted root cause scoring algorithm that integrates temporal pre-emptivity, graph structure centrality, and fault propagation probability, and to determine the root cause device based on the root cause score, and to infer the fault propagation path based on the root cause device and the root cause score and timestamp. The pattern recognition module is used to generate path fingerprint codes based on the device sequence in the fault chain according to the fault propagation path, classify fault chains with the same path fingerprint codes into the same fault mode, and trigger a recurring fault mode alarm when a preset condition is met.
[0006] Compared with existing technologies, the system-level fault chain detection and root cause localization method for warehousing equipment of this invention uses an event grouping and merging algorithm combined with the dual conditions of temporal proximity and physical topology reachability to perform cross-device abnormal event clustering. This effectively filters out false associations caused by temporal coincidences, achieving high-precision automatic discovery of system-level fault chains. Furthermore, by integrating weighted scores from three dimensions—temporal precession, graph structure centrality, and historical fault propagation probability—the root cause is located collaboratively from three aspects: causal temporality, structural location, and historical patterns, improving the accuracy of root cause localization and achieving precise root cause location. In addition, the fault propagation path can be inferred by combining timestamps, and propagation direction-sensitive pattern matching can be achieved based on the path fingerprint encoding of each fault chain, thereby intelligently identifying recurring faults and providing early warnings. Attached Figure Description
[0007] Figure 1 This is a flowchart illustrating a system-level fault chain detection and root cause localization method for warehousing equipment according to the present invention. Figure 2 yes Figure 1The diagram shows the specific process flow of S20 in the system-level fault chain detection and root cause localization method for warehouse equipment. Figure 3 yes Figure 1 The diagram shows the specific process flow of S50 in the system-level fault chain detection and root cause localization method for warehouse equipment. Figure 4 yes Figure 1 The diagram shows the specific process flow of S60 in the system-level fault chain detection and root cause localization method for warehouse equipment. Figure 5 This is a schematic diagram of the structure of a system-level fault chain detection and root cause localization device for warehousing equipment according to the present invention. Detailed Implementation
[0008] To enable those skilled in the art to more clearly understand the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0009] Reference Figure 1 , Figure 1 This is a flowchart illustrating the system-level fault chain detection and root cause localization method for warehousing equipment according to the present invention. In the embodiment shown in the figure, the system-level fault chain detection and root cause localization method for warehousing equipment includes: S10. Periodically collect the operating status data of each device in the warehousing system.
[0010] In this invention, the warehousing system includes various automated equipment, including but not limited to: conveyor lines, stacker cranes, elevators, AGVs (Automated Guided Vehicles), and shuttles. Multiple devices are scheduled and controlled by a Warehouse Control System (WCS). The material flow direction between devices is defined through configuration files or routing tables in a database, forming a directed topology network, represented by a directed graph G=(V,E), where V is the set of nodes, E is the set of directed edges, devices are the nodes of the graph, and the connections between devices that materials can flow to are the directed edges of the graph.
[0011] In this step, a complete data collection can be performed at fixed intervals by a data collector. The collection method can include at least one of two approaches: scheduled database queries and real-time message queue subscriptions. Scheduled database queries extract real-time operating status data of each device from the warehouse control system's database using configurable SQL queries. Real-time message queue subscriptions subscribe to device status change messages via message queue middleware. When a device's status changes, the warehouse control system actively pushes the changed data, which the data collector receives, parses, and merges. For the message queue method, since the push frequency of device status change messages may be high, the data collector maintains a message aggregation buffer within the collection cycle. At the end of each collection cycle, multiple change messages for the same device in the buffer are merged into a snapshot of the device's latest status, ensuring that subsequent processing logic remains consistent with the scheduled query method. Both collection methods can be used individually or in combination, and can be flexibly switched via configuration files.
[0012] In this embodiment, the operational status data may include device operation data, the number of queued tasks (i.e., the number of subtasks currently queued for processing on the device), the number of tasks being executed, the average processing time (i.e., the average time taken to complete a subtask within the most recent time window), and currently associated task identifiers (including a list of strings representing the identifiers of the main tasks currently being executed or waiting on the device). Understandably, the device operation data can be compared with a preset operational threshold to determine the current operational status of the device. The current operational status of the device can be one of normal operation, idle, fault, or blocked. Each device in the warehousing system is assigned a device identifier, which is a globally unique string identifier for the device. Furthermore, device identifiers can be collected and associated with the collected operational status data to form a snapshot object. In some embodiments, this object can also be associated with global efficiency indicators (inbound / outbound rate, queue depth, timeout, etc.) and task path information. The task path information specifically refers to the complete sequence of device nodes traversed by each subtask from the starting device to the target device. In a warehousing system, a main task (such as an outbound order) is usually broken down into multiple subtasks (such as conveyor line handling, stacker crane picking, elevator layer changing, etc.). The task path information is for the granularity of the subtasks after decomposition, for example, "Subtask ID: Sub-12345, Path: Conveyor Line 1 → Stacker Crane 2 → Storage Location 3".
[0013] S20. Based on the operating status data of each device, detect whether a state transition event has occurred in each device, maintain an active abnormal event set based on the state transition event, and generate an event record containing device identifier, event type and timestamp.
[0014] In this embodiment, a sliding window timing tracker can be used to detect state transition events of each device, thereby maintaining a set of active abnormal events.
[0015] like Figure 2 As shown, step S20 specifically includes the following steps S21-S24: S21. Based on the operating status data of each device, detect whether the device has changed from a normal state to a preset abnormal state.
[0016] In this step, a sliding window timing tracker is used to detect whether the operating status data of each device has changed to a preset abnormal state relative to the state of the previous cycle. The preset abnormal state can be a device malfunction, task queuing abnormality, processing performance abnormality, and / or intelligent detection abnormality, etc. Therefore, the state transition event type can include at least one of the following: device malfunction event, task queuing abnormality event, processing performance abnormality event, and intelligent detection abnormality event.
[0017] Specifically, the time-series tracker internally maintains the following data structures: (a) a device status dictionary for the previous period, where the key is the device identifier and the value is the device status data structure, including running data, queue count, average time consumption, whether it has been marked as abnormal by AI, and the last update timestamp; (b) an active abnormal event index, where the key is the device identifier and the value is the currently active state transition event object for that device, with each device having at most one active abnormal event at any given time; (c) an event sliding window, a double-ended queue with a maximum capacity of 5000 entries, storing all events within the window in chronological order; and (d) a resolved event queue, a double-ended queue with a maximum capacity of 2000 entries, storing closed active abnormal events. It can be seen that the time-series tracker maintains the status of each device for the previous period. Preferably, in each acquisition period, the sliding window time-series tracker can detect various types of state transition events in priority order, that is, compare the current state with the state of the previous period and detect the transition of preset abnormal states in priority. The priority detection order is: device failure events... Task queuing exception event Handling performance exception events Intelligent detection of abnormal events.
[0018] In this embodiment, the step of detecting whether the operating status data of each device has changed to a preset abnormal state relative to the previous period using a sliding window time-series tracker specifically includes: detecting abnormal state changes of each device based on the operating status data and preset state change detection rules using a sliding window time-series tracker; the state change detection rules include at least the following: when the current operating state of a device is faulty and it was not faulty in the previous period, a device fault abnormal change occurs, and a device fault event can be generated subsequently; when the number of tasks currently queued by a device exceeds a preset threshold and it did not exceed it in the previous period, a task queuing abnormal change occurs, and a task queuing abnormal event can be generated subsequently; when the current average processing time of a device exceeds the product of a preset baseline processing time and a preset multiplier and it did not exceed it in the previous period, a processing performance abnormal change occurs, and a processing performance abnormal event can be generated; and when the detection result of the AI abnormality detection module marking the device as abnormal is received and it was not marked in the previous period, an intelligent detection abnormal change occurs, and an intelligent detection abnormal event can be generated subsequently.
[0019] S22. When the device is detected to transition from a normal state to a preset abnormal state, and there are currently no active abnormal events in the device, a state transition event corresponding to the abnormal state is generated.
[0020] In this embodiment, during detection, if an active abnormal event already exists on the device in the previous cycle, i.e., the device was in a preset abnormal state in the previous cycle, the detection of new events for that device is skipped. For devices without active abnormal events, the sliding window timing tracker detects them in the order of priority: device failure events, task queuing abnormal events, processing performance abnormal events, and intelligent detection abnormal events. When a change in the current state of the device relative to the preset abnormal state in the previous cycle is detected, a corresponding state transition event is generated. At the same time, when a device with an existing active abnormal event is detected to return to normal, the active abnormal event is closed, its duration is recorded, and it is removed from the set of active abnormal events.
[0021] Specifically, when the current operating state of the device is faulty and the operating state in the previous cycle was not faulty, a fault occurrence event is generated, with a severity fixed at 1.0 (the highest value). That is, the fault is the highest priority event. Once a fault is detected, the event is immediately registered as an active abnormal event for the device, and other types of transitions for the device are no longer detected.
[0022] While the device is not currently in a faulty state, the system continues to monitor for task queuing anomalies. These anomalies occur when the number of tasks currently queued is greater than or equal to a preset congestion threshold, and the number of tasks queued in the previous period was less than the threshold. The severity of the event is determined by a piecewise function based on a preset interval in which the number of queued tasks falls. For example, a queue surge event is generated when the number of tasks currently queued is greater than or equal to a preset congestion threshold (e.g., 5 by default), and the number of tasks queued in the previous period was less than the threshold. The severity is calculated using the following piecewise formula: ,in, The preset queuing normalization base is (in this embodiment, it can be 15.0). , , To preset the severity parameters, As a preset increment coefficient, This indicates the number of tasks waiting in the queue. Indicates the congestion threshold. This represents a severe congestion threshold (e.g., 10). Preferably, in some embodiments, , , , , The values can be 15.0, 0.7, 0.95, 0.98, and 0.03, respectively.
[0023] When the current average processing time of the device exceeds the product of the preset baseline processing time and the anomaly multiplier for that device type, and the average processing time of the previous cycle does not exceed the product, a slowdown event is generated. The severity of this event is calculated using the following formula: ,in, This indicates the current average processing time. This indicates the preset baseline time for this equipment type. For example, the preset baseline time for a conveyor line can be set to 8 seconds, for a stacker crane to 25 seconds, and for a hoist to 15 seconds. The specific time can be adjusted according to the actual situation, and the abnormality rate can be set to 1.5.
[0024] Furthermore, the sliding window time-series tracker also receives real-time inference results from an independent AI anomaly detection module in each acquisition cycle. This AI anomaly detection module is trained in an unsupervised manner, for example, using a deep learning autoencoder. During training, it only learns the multi-dimensional indicator patterns of each device under normal operating conditions (such as the number of queued tasks, processing time, and other temporal features), thereby establishing a data distribution model of normal behavior. During the inference phase, the module inputs the multi-dimensional operating indicator vectors of each device in the current acquisition cycle in real time. After compression and decoding reconstruction by the encoder, the reconstruction error between the original input and the reconstructed output is calculated as an anomaly score. When the anomaly score of a device exceeds a preset anomaly threshold (such as the 95th percentile of the reconstruction error in the training set), the device is marked as "AI-judged anomaly." After obtaining this marking result, if the device was not marked as an anomaly in the previous cycle and there are currently no other types of active anomaly events for the device, an intelligent detection anomaly event is generated. The severity is calculated based on the ratio of the anomaly score to the threshold, and this event is registered as an active anomaly event for the device. The severity calculation formula is as follows: , Indicates abnormal rating. This indicates the preset abnormal threshold.
[0025] Furthermore, when the time-series tracker detects the recovery conditions for the corresponding event in a subsequent acquisition cycle, it closes the active abnormal event and records the duration. The recovery conditions vary depending on the event type: (a) Fault occurrence event: The device's current state is no longer faulty; (b) Queue surge event: The device's current queue count drops below the congestion threshold; (c) Processing slowdown event: The device's current average processing time drops below the product of the baseline processing time and the anomaly multiplier; (d) Intelligent detection anomaly event: The AI anomaly detection module no longer marks the device as an anomaly. When the device meets the recovery conditions, its active abnormal event is closed (assuming the active state is no, the recovery timestamp and duration are recorded), a recovery event is generated, and the closed abnormal event is moved to the resolved event queue. Understandably, the generated state transition event can include not only the device identifier, event type, and timestamp, but also the severity of each event, anomaly description, triggering event data, active state flag (indicating whether the event is still ongoing), recovery timestamp, and duration.
[0026] S23. Record the state transition event as an active abnormal event of the device and add it to the active abnormal event set.
[0027] In this invention, active abnormal events are the generated state transition events.
[0028] S30. Obtain active abnormal events within a preset time window, and based on the event grouping and merging algorithm, cluster the active abnormal events with temporal proximity and physical topology reachability as dual constraints to construct a fault chain.
[0029] This step may include: obtaining the set of active abnormal events within a preset time window; when the number of events is not less than a preset minimum chain length threshold, using a disjoint-set data structure (DFS) algorithm to group two active abnormal events that simultaneously satisfy both temporal proximity constraints and physical topology reachability constraints into the same fault chain cluster to construct a fault chain; wherein, the temporal proximity constraint is that the absolute value of the difference between the timestamps of the two events does not exceed a preset temporal threshold, to avoid incorrectly associating irrelevant events with excessively large time spans; the physical topology reachability constraint is that the two devices corresponding to the two events are reachable within a preset number of hops on the directed graph of the warehouse physical topology. That is, for two events whose absolute value of the difference between their timestamps does not exceed the preset temporal threshold and whose corresponding two devices are reachable within a preset number of hops on the directed graph of the warehouse physical topology, a disjoint-set data structure (DFS) merging operation is performed, grouping them into the same group, with events within each group arranged in ascending order of timestamp. Events with a member number less than the minimum chain length threshold are filtered out. The groups are grouped, and each group retained represents a fault chain. Preferably, path compression and rank-based merging optimization strategies can be used when performing lookup and merge operations.
[0030] In this embodiment, the preset hop count is preferably two hops. Reachability within the preset hop count means that a device can reach another device by traversing no more than two hops along the forward or reverse direction of a directed edge in the physical topology directed graph G. That is, if starting from one device, reaching another device by traversing no more than two hops along a directed edge in the forward or reverse direction is considered to satisfy reachability. For example, let event... The corresponding device is a, and the event is... The corresponding device is b, and the determination is made. Whether it is valid. (Among them) The pre-computation method is as follows: Starting from node a, traverse all nodes reachable in one or two hops along the forward direction of the directed edges (i.e., from predecessor to successor), and then traverse all nodes reachable in one or two hops along the reverse direction of the directed edges (i.e., from successor to predecessor). Take the union of the two sets and remove node a itself. The pre-computation results can be stored in a cache dictionary to avoid repeated graph traversal searches and can be directly queried during clustering. The physical meaning of this constraint is: fault propagation relationships are only possible between devices that are geographically close on the material flow path. By using the physical topology, event pairs that are only coincidentally related in time but spatially unrelated are filtered out, significantly reducing the risk of spurious associations.
[0031] S40. A three-dimensional weighted root cause scoring algorithm that integrates temporal pre-emptivity, graph structure centrality, and fault propagation probability is used to calculate the root cause score of each device in each fault chain.
[0032] In this step, the three-dimensional weighted root cause scoring algorithm formula is used. Calculate the root cause score for each device in each of the aforementioned fault chains; where, For devices in the fault chain; , , These are preset weight coefficients, and satisfy... ; For equipment Timing-first scoring, based on device The earliest active anomaly timestamp is calculated relative to the position of the earliest and latest timestamps in the fault chain; For equipment Scoring the graph structure centrality in a physically topological directed graph; For equipment The probability of fault propagation, which serves as a source of fault propagation, is obtained from historical fault propagation data.
[0033] In this embodiment, the temporal precession feature is calculated based on the timestamp of the earliest abnormal event occurring in the fault chain to which the device belongs. Specifically, the precession score is determined according to the relative position of the device's earliest timestamp among the earliest and latest timestamps of all devices in the fault chain. That is, for each device d in the chain, all events in which the device participates are traversed, and the earliest timestamp is recorded as d. Let the minimum earliest timestamp of all devices in the chain be . The maximum value is Then the temporal pre-score ,in =0.001, to prevent division by zero for extremely small positive numbers. The device that first malfunctions usually gets 1.0 point (highest), and the device that malfunctions last gets 0.0 point (lowest). Intermediate devices are linearly interpolated. When there is only one device in the chain (i.e., all...) When they are the same, It is unified as 1.0.
[0034] In some embodiments, the graph structure centrality feature can be expressed as a betweenness centrality index, where betweenness centrality is the proportion of all shortest paths in the graph that pass through a given node. , where is the normalized value of the betweenness centrality of each node in the physical topological directed graph G; where The maximum value of betweenness centrality for all nodes in the graph is ε = 0.001. Furthermore, to avoid repeatedly calculating the graph centrality index during high-frequency diagnostic cycles, the calculation results of betweenness centrality are also cached. The cache validity period can be 300 seconds (5 minutes), and the cached results are returned directly within the validity period.
[0035] The fault propagation probability is calculated by statistically analyzing the co-occurrence frequency of temporally adjacent device pairs connected by a direct directed edge in the physical topology directed graph G within a fault chain, based on historical data. This can be learned online from historical fault chain data. The warehouse system maintains a directed graph of fault propagation probability parallel to the physical topology directed graph G, where each directed edge (a, b) records the conditional probability P(a→b) that device b also fails after device a fails. This probability is updated online using a frequency statistics method: whenever a fault chain is detected and recorded, the device sequences in the chain are traversed in ascending order of timestamps. For temporally adjacent device pairs (a, b) connected by a direct directed edge in the physical topology, the corresponding co-occurrence count is calculated. Increment by 1, and simultaneously increase the total trigger count with device a as the source node. Add 1; the conditional probability is calculated as follows: Only when The probability is valid only when the number of observations is not less than the preset minimum observation threshold (default is 3); otherwise, it is considered a cold start and returns a zero value. To prevent historical data from becoming outdated, the statistical calculation of fault propagation probability adopts a sliding window decay strategy: for historical records Δt days from the current time, they are weighted by an exponential decay factor exp(-λ·Δt) during counting, where the decay coefficient λ defaults to 0.1 / day. Records older than 7 days have their weight reduced to approximately 50% to ensure that the probability reflects recent fault propagation patterns. For each device d in the fault chain, the fault propagation probability is queried in the directed graph from that device as the source node to all its successor nodes. The outgoing edge probability is averaged and then multiplied by a magnification factor of 2 to enhance the distinguishability, and is limited to within 1.0. Experiments showed that in actual warehousing systems, due to the dispersed nature of fault propagation paths, the average propagation probability of a single device is typically in the range of 0.1 to 0.4. After setting the amplification factor to 2... The scores are distributed within the effective range of 0.2 to 0.8, exhibiting relatively good discriminative power. Furthermore, when in the initial operational phase, and the directed graph of fault propagation probability suffers from insufficient data (e.g., fewer than 3 observations per pair of devices), resulting in some devices receiving a score of zero, a cold start compensation strategy is implemented: traversing the events corresponding to these devices, and taking the severity value of the highest severity state transition event corresponding to the device in the current fault chain as the severity value. The alternative values allow the root cause scoring algorithm to continue to operate effectively during the data accumulation phase and gradually transition to a scoring based on the actual probability of fault propagation as data increases.
[0036] Understandably, after the scores for the three dimensions are calculated, the final root cause score for each device is obtained by fusing them according to the above three-dimensional weighted root cause scoring algorithm formula. Preferably, the temporal precession score is typically assigned the highest weight. , , The preferred values are 0.45, 0.25, and 0.30.
[0037] As described above, the temporal precession dimension can identify the earliest abnormal device in the fault chain, the graph structure centrality dimension can identify the device in the hub position in the physical topology, and the fault propagation probability dimension uses historical fault statistics to disambiguate candidate root causes, and can select the most likely fault propagation source from multiple candidate devices with similar temporal sequences based on statistical laws. The three-dimensional weighted root cause scoring algorithm achieves a complementary and synergistic effect among the three dimensions by integrating temporal precession, graph structure centrality, and propagation probability. After the integration of the three, when multiple devices become abnormal at almost the same time in the same collection period, making it impossible for the temporal precession dimension to distinguish them, the fault propagation probability dimension and the graph structure centrality dimension can provide additional judgment basis; when the graph structure centrality dimension tends to overestimate hub nodes, the temporal precession dimension and the fault propagation probability dimension can correct it; when insufficient historical data causes the fault propagation probability scoring to fail, the temporal precession and structure centrality dimensions can still ensure the effective operation of the algorithm, effectively overcoming the inherent defects of single-dimensional scoring.
[0038] S50. Based on the root cause score, determine the root cause device, and starting from the root cause device, infer the fault propagation path according to the root cause score and timestamp.
[0039] In this invention, the device with the highest root cause score is usually identified as the root cause device of the failure chain, and its score is recorded as the root cause confidence level.
[0040] like Figure 3 As shown, step S50 specifically includes the following steps S51-S53: S51. The device with the highest root cause score in each fault chain is determined as the root cause device of the fault chain, and the earliest active abnormal event timestamp of each of the other devices in the fault chain is obtained.
[0041] S52. If the earliest active abnormal event timestamps of each device are different, starting from the root cause device, the remaining devices in the fault chain are arranged in ascending order according to the timestamps of their earliest active abnormal events, forming a fault propagation path that spreads from the root cause device downstream in a time sequence.
[0042] S53. When multiple devices in a fault chain have the same earliest active anomaly timestamp, their order in the fault propagation path is determined by descending root cause scores.
[0043] Based on steps S51-S53 above, the device with the highest root cause score is taken as the starting point. The remaining devices are arranged in ascending order according to their earliest abnormal timestamps, and if timestamps are the same, they are arranged in descending order according to their root cause scores, thus forming a fault propagation path. That is, after determining the root cause device, fault propagation path inference is performed: starting with the root cause device, the remaining devices in the fault chain are arranged in ascending order according to the timestamp of their earliest abnormal event, and added to the propagation path sequentially, forming a propagation path sequence that spreads from the root cause device to downstream devices in a temporal progression. This propagation path reflects the causal transmission order of the fault among devices, providing a foundation for subsequent path fingerprint calculation and pattern matching. For example, consider a fault chain consisting of conveyor line A, stacker crane B, elevator C, and shuttle D. The earliest abnormal event timestamps of A, B, C, and D are 10:00:05, 10:00:10, 10:00:12, and 10:00:15, respectively. If A has the highest score (0.92) after three-dimensional weighted root cause scoring, then A is identified as the root cause device and serves as the starting point of the propagation path. The remaining devices are added to the path in ascending order of their earliest abnormal timestamps, resulting in the propagation path sequence A→B→C→D. If B and C have the same timestamp, the order is determined by the descending root cause score, with higher scores taking priority, thus forming a unique and reproducible propagation path.
[0044] Furthermore, after completing the propagation path inference, in some embodiments, the overall severity of the fault chain can also be calculated. The severity of the fault chain can be calculated using a logarithmic decay formula: ,in, This represents the highest single-event severity across all events in the failure chain. The logarithmic decay formula, which deduplicatizes the number of devices involved in the fault chain, significantly reduces the amplification effect when the chain length is large (e.g., Ndev > 8) compared to the linear formula. This allows fault chains of different sizes to maintain a distinguishable gradient in severity, avoiding the inability to distinguish fault events with vastly different impacts due to score saturation.
[0045] Furthermore, this invention can also classify fault chains based on the combination of event types contained in the chain: when a fault chain contains at least one fault occurrence event, it is classified as a fault cascading type, indicating a clear hardware fault that triggers a chain reaction; when all events in the fault chain are queuing surge events, it is classified as a congestion wave type, indicating that the system experiences congestion propagation due to task backlog without any substantial equipment fault; when a fault chain contains at least one processing slowdown event but does not contain fault occurrence events or queuing surge events, it is classified as a performance degradation type, indicating a system-level impact caused by a decrease in equipment processing capacity; when a fault chain contains at least one intelligent detection anomaly event but does not contain fault occurrence events, processing slowdown events, or queuing surge events, it is classified as a latent anomaly type, indicating a potential anomaly deviating from normal operating mode only detected by the AI anomaly detection model; fault chains that do not fall into any of the above categories are classified as hybrid types. This classification facilitates maintenance personnel in quickly understanding the overall nature of the fault chain and taking targeted measures.
[0046] S60. Based on the device sequence in the fault chain, generate a path fingerprint code according to the fault propagation path, classify fault chains with the same path fingerprint code into the same fault mode, and trigger a recurring fault mode alarm when a preset condition is met.
[0047] like Figure 4 As shown, step S60 specifically includes the following steps S61-S63: S61. According to the fault propagation path, each device identifier in the fault chain device sequence is concatenated into a string, and a hash value is calculated using a preset hash algorithm. The hash value of a preset length is then used as the path fingerprint code of the fault chain.
[0048] In this step, the device identifiers in the fault chain device sequence are concatenated into a string according to the order of devices in the fault propagation path, using a preset separator. A preset hash algorithm is then used to calculate the hash value of this string, and a preset length prefix of the hash value is extracted as the path fingerprint code. Preferably, the preset hash algorithm can be the MD5 algorithm, and the preset length can be the first 12 hexadecimal characters of the hash value.
[0049] Specifically, for each fault chain, the device identifiers are first concatenated into a string using the vertical bar "|" as a separator, according to the order of the devices in the fault propagation path. The hash value of this string is then calculated using the MD5 hash algorithm, and the first 12 hexadecimal characters are extracted as the path fingerprint code. This path fingerprint code preserves the directionality of fault propagation. For example, the device combination A→B→C and C→B→A will produce different fingerprint codes and be identified as different fault modes.
[0050] S62. Group fault chains with the same path fingerprint code into the same fault mode.
[0051] In this embodiment, fault chains with the same path fingerprint code are grouped into the same fault mode. The average severity and average duration of the fault mode can be updated by the moving average method. The root cause device of the fault chain is included in the root cause device frequency count of the fault mode, and the device that appears most frequently is taken as the representative root cause device of the fault mode.
[0052] The formula for calculating the average severity is as follows: ,in, The average severity level before the update. The severity of the new failure chain, The formula for calculating the cumulative occurrence count and average duration is similar. If the path fingerprint code does not exist, a new fault mode record is created.
[0053] S63. When the cumulative number of occurrences of the same fault mode reaches the preset recurrence threshold within a preset time window, the fault mode is marked as a recurring fault mode, the alarm level is upgraded, and a prompt is triggered.
[0054] In this step, the cumulative occurrence count refers to the total number of all fault chains belonging to this fault mode. Whenever a new fault chain is detected and its path fingerprint code is calculated, if the fingerprint code matches the fingerprint code of an existing mode, the cumulative occurrence count of that fault mode is incremented by 1, indicating that the fault mode has occurred once again.
[0055] Specifically, when the cumulative occurrence of a certain fault mode reaches a preset first recurrence threshold (which can be 3 times by default) for the first time, the fault mode is marked as a recurring fault mode and the alarm level is set to the attention level. When the cumulative occurrence reaches a second recurrence threshold greater than the first recurrence threshold (such as three times the first recurrence threshold), the alarm level is upgraded to the severity level, and an alarm log containing the mode identifier, cumulative occurrence, involved devices, root cause devices, and average severity can be output. The registered callback function list (such as DingTalk robot push, WeChat Work push, etc.) is traversed, and each notification method is called in turn to complete the alarm push and provide a notification.
[0056] As can be seen, the system-level fault chain detection and root cause localization method for warehousing equipment of this invention realizes a complete closed loop from equipment operating status data collection, temporal state transition detection, disjoint-set clustering to construct fault chains, three-dimensional weighted scoring for root cause localization, to path fingerprint pattern recognition and alarm escalation. Furthermore, in the above closed-loop process, cross-device abnormal event clustering is performed by combining an event grouping and merging algorithm with the dual conditions of temporal proximity and physical topology reachability constraints, effectively filtering out false associations caused by temporal coincidences, achieving high-precision automatic discovery of system-level fault chains. Moreover, by integrating weighted scoring based on three dimensions—temporal precession, graph structure centrality, and historical propagation probability—root causes are located collaboratively from three aspects: causal temporality, structural location, and historical patterns, improving the accuracy of root cause localization and achieving precise root cause localization. In addition, fault propagation paths are inferred by combining timestamps, and propagation direction-sensitive pattern matching can be achieved based on the path fingerprint encoding of each fault chain, thereby intelligently identifying recurring faults and providing early warnings. Furthermore, the entire process is embedded into the diagnostic orchestration process of each data acquisition cycle (default 10 seconds) and executed automatically, realizing real-time detection and root cause analysis of fault chains at the warehouse automation system level without manual intervention. The overall diagnostic delay is controlled within a single acquisition cycle, meeting the real-time requirements of industrial scenarios.
[0057] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0058] Reference Figure 5 , Figure 5 This is a schematic structural block diagram of the system-level fault chain detection and root cause localization device for warehousing equipment according to the present invention. In the embodiment shown in the figure, the system-level fault chain detection and root cause localization device for warehousing equipment includes: The data acquisition module 110 is used to periodically collect the operating status data of each device in the warehousing system; The timing tracking module 120 is used to detect whether a state transition event has occurred in each device based on the operating status data of each device, maintain an active abnormal event set based on the state transition event, and generate an event record containing device identifier, event type and timestamp. The fault chain clustering module 130 is used to obtain active abnormal events within a preset time window, and based on the event grouping and merging algorithm, cluster the active abnormal events with temporal proximity and physical topology reachability as dual constraints to construct a fault chain. The root cause localization module 140 is used to calculate the root cause score of each device in each fault chain using a three-dimensional weighted root cause scoring algorithm that integrates temporal pre-emptivity, graph structure centrality and fault propagation probability, and to determine the root cause device based on the root cause score, and to infer the fault propagation path based on the root cause device and the root cause score and timestamp. The pattern recognition module 150 is used to generate path fingerprint codes based on the device sequence in the fault chain according to the fault propagation path, classify fault chains with the same path fingerprint codes into the same fault mode, and trigger a recurring fault mode alarm when a preset condition is met.
[0059] In some embodiments, the timing tracking module 120 is specifically used for: Based on the operating status data of each device, detect whether the device has changed from a normal state to a preset abnormal state; When the device is detected to transition from a normal state to a preset abnormal state, and there are currently no active abnormal events on the device, a state transition event corresponding to the abnormal state is generated. The state transition events are recorded as active abnormal events of the device and added to the active abnormal event set.
[0060] In some embodiments, the timing tracking module 120 is further configured to: The sliding window timing tracker detects whether the operating status data of each device has changed to a preset abnormal state relative to the state of the previous cycle.
[0061] In some embodiments, the state transition event includes at least one of the following: device failure event, task queuing anomaly event, processing performance anomaly event, and intelligent detection anomaly event.
[0062] In some embodiments, the fault chain clustering module 130 is specifically used for: Using the disjoint-set data structure algorithm, two active abnormal events that simultaneously satisfy both temporal proximity constraints and physical topology reachability constraints are grouped into the same fault chain cluster to construct a fault chain. The temporal proximity constraint is that the absolute value of the difference between the timestamps of the two events does not exceed a preset temporal threshold, and the physical topology reachability constraint is that the two devices corresponding to the two events are reachable within a preset number of hops on the directed graph of the warehouse physical topology.
[0063] In some embodiments, the root cause localization module 140 is specifically used for: Using the three-dimensional weighted root cause scoring algorithm formula Calculate the root cause score for each device in each of the aforementioned fault chains; where, For devices in the fault chain; , , These are preset weight coefficients, and satisfy... ; For equipment Timing-first scoring, based on device The earliest active anomaly timestamp is calculated relative to the position of the earliest and latest timestamps in the fault chain; For equipment Scoring the graph structure centrality in a physically topological directed graph; For equipment The probability of fault propagation, which serves as a source of fault propagation, is obtained from historical fault propagation data.
[0064] In some embodiments, the root cause localization module 140 is further specifically used for: The device with the highest root cause score in each fault chain is identified as the root cause device of the fault chain, and the earliest active abnormal event timestamp of each of the other devices in the fault chain is obtained. If the earliest active abnormal event timestamps of each device are different, starting from the root cause device, the remaining devices in the fault chain are arranged in ascending order according to the timestamps of their earliest active abnormal events, forming a fault propagation path that spreads from the root cause device downstream in a time sequence. When multiple devices in a fault chain have the same earliest active anomaly timestamp, their order in the fault propagation path is determined by descending root cause scores.
[0065] In some embodiments, the pattern recognition module 150 is specifically used for: According to the fault propagation path, the identifier of each device in the fault chain device sequence is concatenated into a string, and a hash value is calculated using a preset hash algorithm. The hash value of a preset length is then used as the path fingerprint code of the fault chain.
[0066] In some embodiments, the pattern recognition module 150 is further configured to: When the cumulative number of occurrences of the same fault mode reaches a preset recurrence threshold within a preset time window, the fault mode is marked as a recurring fault mode, the alarm level is upgraded, and a prompt is triggered.
[0067] The warehousing equipment system-level fault chain detection and root cause localization device provided by this invention periodically collects the operating status data of each device in the warehousing system, thereby obtaining the active abnormal events existing in each device. It adopts an event grouping and merging algorithm combined with the dual conditions of temporal proximity and physical topology reachability constraints to perform cross-device abnormal event clustering, effectively filtering out false associations caused by temporal coincidences, and realizing high-precision automatic discovery of system-level fault chains. Furthermore, by integrating the weighted scoring of three dimensions—temporal precession, graph structure centrality, and historical propagation probability—it collaboratively locates the root cause from three aspects: causal temporality, structural position, and historical patterns, improving the accuracy of root cause localization and achieving precise root cause location. It also infers the fault propagation path based on the timestamp, and can achieve propagation direction-sensitive pattern matching based on the path fingerprint encoding of each fault chain, thereby intelligently identifying recurring faults and providing early warnings. It realizes a complete closed loop from equipment operating status data collection, temporal state change detection, disjoint-set clustering to construct fault chains, three-dimensional weighted scoring root cause localization, to path fingerprint pattern recognition and alarm escalation. Furthermore, the entire process is embedded into the diagnostic orchestration process of each data acquisition cycle (default 10 seconds) and executed automatically, realizing real-time detection and root cause analysis of fault chains at the warehouse automation system level without manual intervention. The overall diagnostic delay is controlled within a single acquisition cycle, meeting the real-time requirements of industrial scenarios.
[0068] Specific limitations regarding the system-level fault chain detection and root cause localization device for warehousing equipment can be found in the limitations of the system-level fault chain detection and root cause localization method for warehousing equipment mentioned above, and will not be repeated here. Each module in the aforementioned system-level fault chain detection and root cause localization device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0069] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0070] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Those skilled in the art can make various equivalent changes and improvements based on the above embodiments, and all equivalent variations or modifications made within the scope of the claims should fall within the protection scope of the present invention.
Claims
1. A method for system-level fault chain detection and root cause localization in warehousing equipment, characterized in that, include: Periodically collect operational status data from various devices in the warehousing system; Based on the operating status data of each device, detect whether a state transition event has occurred in each device, maintain an active abnormal event set based on the state transition events, and generate an event record containing device identifier, event type, and timestamp; Active abnormal events within a preset time window are acquired. Based on the event grouping and merging algorithm, the active abnormal events are clustered under the dual constraints of temporal proximity and physical topological reachability to construct a fault chain. A three-dimensional weighted root cause scoring algorithm that integrates temporal pre-emptivity, graph structure centrality, and fault propagation probability is used to calculate the root cause score of each device in each fault chain. The root cause device is identified based on the root cause score, and the fault propagation path is inferred from the root cause device based on the root cause score and timestamp. Based on the device sequence in the fault chain, a path fingerprint code is generated according to the fault propagation path. Fault chains with the same path fingerprint code are grouped into the same fault mode, and a recurring fault mode alarm is triggered when a preset condition is met. The event grouping and merging algorithm, using temporal proximity and physical topological reachability as dual constraints, clusters the active anomalous events to construct a fault chain, specifically including: Using the disjoint-set data structure algorithm, two active abnormal events that simultaneously satisfy both temporal proximity constraints and physical topology reachability constraints are grouped into the same fault chain cluster to construct a fault chain. The temporal proximity constraint is that the absolute value of the difference between the timestamps of the two events does not exceed a preset temporal threshold, and the physical topology reachability constraint is that the two devices corresponding to the two events are reachable within a preset number of hops on the directed graph of the warehouse physical topology. The step of identifying the root cause device based on the root cause score, and inferring the fault propagation path starting from the root cause device based on the root cause score and timestamp, specifically includes: The device with the highest root cause score in each fault chain is identified as the root cause device of the fault chain, and the earliest active abnormal event timestamp of each of the other devices in the fault chain is obtained. If the earliest active abnormal event timestamps of each device are different, starting from the root cause device, the remaining devices in the fault chain are arranged in ascending order according to the timestamps of their earliest active abnormal events, forming a fault propagation path that spreads from the root cause device downstream in a time sequence. When multiple devices in a fault chain have the same earliest active anomaly timestamp, their order in the fault propagation path is determined by descending root cause scores.
2. The method for system-level fault chain detection and root cause localization of warehousing equipment as described in claim 1, characterized in that, Based on the operating status data of each device, the system detects whether a state transition event has occurred in each device, and maintains a set of active abnormal events based on the state transition events. Specifically, this includes: Based on the operating status data of each device, detect whether the device has changed from a normal state to a preset abnormal state; When the device is detected to transition from a normal state to a preset abnormal state, and there are currently no active abnormal events on the device, a state transition event corresponding to the abnormal state is generated. Record the state transition event as an active anomaly event of the device, and add it to the active anomaly event. gather.
3. The method for system-level fault chain detection and root cause localization of warehousing equipment as described in claim 2, characterized in that, The step of detecting whether a device has transitioned from a normal state to a preset abnormal state based on the operating status data of each device specifically includes: The sliding window timing tracker detects whether the operating status data of each device has changed to a preset abnormal state relative to the state of the previous cycle.
4. The method for system-level fault chain detection and root cause localization of warehousing equipment as described in claim 1, characterized in that, The state transition events include at least one of the following: equipment failure events, task queuing anomaly events, processing performance anomaly events, and intelligent detection anomaly events.
5. The method for system-level fault chain detection and root cause localization of warehousing equipment as described in claim 1, characterized in that, The method employs a three-dimensional weighted root cause scoring algorithm that integrates temporal pre-emptivity, graph structure centrality, and fault propagation probability to calculate the root cause score of each device in each fault chain. Specifically, this includes: Using the three-dimensional weighted root cause scoring algorithm formula Calculate the root cause score for each device in each of the aforementioned fault chains; where, For devices in the fault chain; , , These are preset weight coefficients, and satisfy... ; For equipment Timing-first scoring, based on device The earliest active anomaly timestamp is calculated relative to the position of the earliest and latest timestamps in the fault chain; For equipment Scoring the graph structure centrality in a physically topological directed graph; For equipment The fault propagation probability score, which serves as the source of the fault, is determined by averaging the outgoing edge probabilities from the source device to each of its successor devices. The outgoing edge probabilities are calculated by statistically analyzing the co-occurrence frequency of temporally adjacent device pairs in the fault chain that are directly connected by a directed edge on the physical topology directed graph in historical data.
6. The method for system-level fault chain detection and root cause localization of warehousing equipment as described in claim 1, characterized in that, The step of generating a path fingerprint code based on the device sequence in the fault chain according to the fault propagation path specifically includes: According to the fault propagation path, the identifier of each device in the fault chain device sequence is concatenated into a string, and a hash value is calculated using a preset hash algorithm. The hash value of a preset length is then used as the path fingerprint code of the fault chain.
7. The method for system-level fault chain detection and root cause localization of warehousing equipment as described in claim 1 or 6, characterized in that, The triggering of a recurring fault mode alarm when preset conditions are met specifically includes: When the cumulative number of occurrences of the same fault mode reaches a preset recurrence threshold within a preset time window, the fault mode is marked as a recurring fault mode, the alarm level is upgraded, and a prompt is triggered.
8. A system-level fault chain detection and root cause localization device for warehousing equipment, characterized in that, include: The data acquisition module is used to periodically collect operating status data from various devices in the warehousing system. The timing tracking module is used to detect whether a state transition event has occurred in each device based on the operating status data of each device, maintain an active abnormal event set based on the state transition event, and generate an event record containing device identifier, event type and timestamp. The fault chain clustering module is used to obtain active abnormal events within a preset time window. Based on the event grouping and merging algorithm, the active abnormal events are clustered under the dual constraints of temporal proximity and physical topological reachability to construct a fault chain. Specifically, it is used to utilize the disjoint-set data structure algorithm to group two active abnormal events that simultaneously satisfy both temporal proximity constraints and physical topology reachability constraints into the same fault chain cluster to construct a fault chain; wherein, the temporal proximity constraint is that the absolute value of the difference between the timestamps of the two events does not exceed a preset temporal threshold, and the physical topology reachability constraint is that the two devices corresponding to the two events are reachable within a preset number of hops on the directed graph of the warehouse physical topology; The root cause localization module is used to calculate the root cause score of each device in each fault chain using a three-dimensional weighted root cause scoring algorithm that integrates temporal pre-emptivity, graph structure centrality, and fault propagation probability. Based on the root cause score, the module identifies the root cause device and, starting from the root cause device, infers the fault propagation path based on the root cause score and timestamp. Specifically, it identifies the device with the highest root cause score in each fault chain as the root cause device of that fault chain and obtains the earliest active anomaly event timestamp for each of the remaining devices in the fault chain. If the earliest active anomaly event timestamps of the devices are different, starting from the root cause device, the remaining devices in the fault chain are arranged in ascending order according to their earliest active anomaly event timestamps, forming a fault propagation path that spreads temporally from the root cause device downstream. When multiple devices in a fault chain have the same earliest active anomaly timestamp, their order in the fault propagation path is determined by descending root cause score. The pattern recognition module is used to generate path fingerprint codes based on the device sequence in the fault chain according to the fault propagation path, classify fault chains with the same path fingerprint codes into the same fault mode, and trigger a recurring fault mode alarm when a preset condition is met.
Citation Information
Patent Citations
Fault event positioning method based on topological model tracking analysis
CN119847116A
System fault mining method for log data
CN121579259A