Data center inspection robot monitoring analysis method and system based on machine vision

Through the combination of deep learning and graph inference algorithms, comprehensive monitoring of data center equipment status and intelligent inspection path planning are achieved, the accuracy and efficiency of equipment abnormality analysis in the existing technology are solved, and the reliability and efficiency of data center operation and maintenance are improved.

CN120279500AActive Publication Date: 2025-07-08BEIJING AMPLI INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510766712.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing data center inspection technology has shortcomings in the fusion of multi-source heterogeneous data, and it is difficult to establish the correlation between the visual state of the equipment and the operating parameters, resulting in the root cause of the abnormality due to low analysis accuracy and lack of intelligence in patrol path planning, and it is impossible to respond to the abnormal state of the equipment efficiently.

Method used

The deep learning object detection model is used to identify the device status, a multimodal data fusion analysis model is built for spatial and temporal registration, and a coordinated analysis is used to generate root cause analysis results, and the patrol path is optimized through the policy network to realize the adaptive coordinated scheduling of the patrol robot.

Benefits of technology

It improves the accuracy and efficiency of equipment abnormality detection, reduces fault diagnosis time, and improves the reliability and efficiency of data center operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279500A_ABST
    Figure CN120279500A_ABST
Patent Text Reader

Abstract

The invention provides a data center inspection robot monitoring analysis method and system based on machine vision, and relates to the technical field of inspection robots, and the method comprises the steps: obtaining a video stream and environment parameter data collected by an inspection robot, and carrying out the detection and recognition of an abnormal state through a deep learning target; and constructing a multi-modal data fusion analysis model to form a knowledge graph to generate a root cause analysis result, and planning an inspection path based on priority scores. According to the invention, intelligentization and precision of data center monitoring are realized, and inspection efficiency and fault diagnosis accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of inspection robots, and particularly to a method and system for monitoring and analyzing data center inspection robots based on machine vision. Background Art

[0002] With the continuous expansion of the scale of data centers and the improvement of the complexity of infrastructure, the traditional manual inspection method is difficult to meet the requirements of efficient and accurate monitoring of modern data centers. As a key infrastructure in the information age, the stable operation of data centers is directly related to the reliability of various online services. At present, robot inspection technology has been gradually applied to the daily operation and maintenance of data centers, and through the installation of various sensors and camera devices, it realizes the automatic monitoring of the operating status of cabinets and equipment. The data center inspection system based on machine vision can identify and analyze key information such as the status of equipment indicator lights, cable connections, and cabinet temperatures, providing real-time monitoring support for data center operation and maintenance.

[0003] There are obvious deficiencies in the existing inspection technology in the aspect of multi-source heterogeneous data fusion. Usually, visual information and environmental parameter data are analyzed as independent monitoring dimensions, and it is difficult to establish the correlation between the visual status of equipment and operating parameters, resulting in low accuracy of abnormal root cause analysis and inability to effectively identify the internal connections of equipment failures.

[0004] The traditional inspection system lacks an intelligent inspection strategy optimization mechanism. The inspection path planning is often based on a preset fixed route, and it cannot dynamically adjust the inspection priority according to the real-time status of data center equipment, making it difficult to achieve an efficient response when equipment is in an abnormal state, reducing the inspection efficiency and the timeliness of fault handling.

[0005] The existing technology mainly relies on single-modal feature extraction in the identification of equipment abnormal states. Especially in complex environments, the recognition accuracy of subtle changes such as the status of equipment indicator lights and cable connection abnormalities is not high, and it is difficult to adapt to diverse scenarios with different lighting conditions and equipment types, restricting the application scope and reliability of the inspection system. Summary of the Invention

[0006] Embodiments of the present invention provide a method and system for monitoring and analyzing data center inspection robots based on machine vision, which can solve the problems in the existing technology.

[0007] In the first aspect of the embodiments of the present invention, a method for monitoring and analyzing a data center inspection robot based on machine vision is provided, including: Obtaining the monitored video stream data collected by the data center inspection robot, extracting the real-time image of the data center cabinet area from the monitored video stream data, and synchronously collecting the temperature distribution data and equipment load data in the cabinet area as environmental parameter data; Analyze the real-time image using a deep learning object detection model to identify the display status of device indicators, the cable connection status, and the cabinet temperature display value in the data center cabinet area, and generate visual abnormal state features; Construct a multi-modal data fusion analysis model, perform spatio-temporal registration on the visual abnormal state features and the environmental parameter data, and use an improved graph reasoning algorithm to construct an equipment abnormal knowledge graph. The improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and generates a root cause analysis result; Input the root cause analysis result into a policy network to generate a patrol task priority score. Based on the patrol task priority score and the real-time layout structure of the data center, plan a patrol path, and continuously optimize the patrol strategy through a double-layer interaction learning mechanism to achieve the adaptive collaborative scheduling of patrol robots.

[0008] Analyze the real-time image using a deep learning object detection model to identify the display status of device indicators, the cable connection status, and the cabinet temperature display value in the data center cabinet area. The generated visual abnormal state features include: Use a deep learning object detection model to extract the display status of device indicators, the cable connection status, and the cabinet temperature display value, and construct a device state causal graph; Based on the device state causal graph, use the attention mechanism to calculate the feature weights for the indicator area, cable area, and temperature display area in the real-time image respectively, and perform weighted fusion on the feature weights and the visual features extracted by the deep learning object detection model to construct a multi-scale visual feature representation; For the abnormal states detected in the multi-scale visual feature representation, construct a visual feature counterfactual verification function. The visual feature counterfactual verification function performs abnormal verification by calculating the difference between the state recognition probability after image enhancement processing and the original state recognition probability, and calculates the visual similarity between the candidate abnormal area and the historical abnormal cases based on the results of the abnormal verification to generate a visual abnormal attribution confidence level; Construct an abnormal evolution path based on the visual abnormal attribution confidence level, and design a warning score calculation function based on the abnormal evolution path. The warning score calculation function comprehensively evaluates the abnormal degree of the device state and the visual feature deviation degree to generate visual abnormal state features.

[0009] Construct an abnormal evolution path based on the visual abnormal attribution confidence level. Designing a warning score calculation function based on the abnormal evolution path includes: Obtain the device abnormal state sequence, which includes abnormal visual feature vectors, device state feature vectors, and state transition probabilities. Calculate the device abnormal state sequence according to the visual abnormal attribution confidence level to generate state transition weights; Construct a multi-dimensional feature fusion function based on the abnormal state sequence of the device, and perform weighted summation on the state transition weight and the multi-dimensional feature fusion function to obtain a fused feature representation; Use the fused feature representation to construct a hierarchical early warning score calculation model, calculate the feature distance between the abnormal visual feature vector and the normal state benchmark, multiply the feature distance by the device state feature vector to obtain an abnormality degree score, calculate the change rate of the state transition probability and generate a trend score in combination with a time decay factor, and perform weighted fusion on the abnormality degree score, the trend score, and the confidence score in the comprehensive evaluation layer to generate an early warning score calculation function.

[0010] Construct a multi-modal data fusion analysis model, perform spatio-temporal registration on the visual abnormal state feature and the environmental parameter data, and use an improved graph reasoning algorithm to construct a device abnormal knowledge graph, including: Construct a multi-modal data fusion analysis model, input the visual abnormal state feature and the environmental parameter data into the multi-modal data fusion analysis model for spatio-temporal registration, and extract the timestamp identifier and spatial position identifier of the visual abnormal state feature and the environmental parameter data; Construct a temporal correlation matrix based on the timestamp identifier, construct a spatial adjacency matrix based on the spatial position identifier, and perform weighted fusion on the temporal correlation matrix and the spatial adjacency matrix to generate a spatio-temporal registration matrix; Use an improved graph reasoning algorithm to construct a device abnormal knowledge graph. The device abnormal knowledge graph includes an entity node layer, an attribute node layer, and a relationship edge set. Construct entity nodes representing different devices in the entity node layer, construct attribute nodes representing device abnormal states in the attribute node layer, construct entity relationship edges of the entity nodes based on the spatio-temporal registration matrix, construct attribute relationship edges of the attribute nodes based on device state transition rules, and use the improved graph reasoning algorithm to iteratively update the entity relationship edges and the attribute relationship edges to generate a device abnormal knowledge graph with a hierarchical structure.

[0011] The improved graph reasoning algorithm realizes collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and generates a root cause analysis result, including: Construct the knowledge graph structure of the improved graph reasoning algorithm, introduce a dynamic weight adaptive mechanism in the knowledge graph structure. The dynamic weight adaptive mechanism calculates the association strength between nodes to obtain the importance weight of the nodes, constructs an edge weight matrix based on the importance weight of the nodes, and the edge weight matrix characterizes the propagation relationship strength between nodes in the knowledge graph; Extract corresponding feature vectors from the visual anomaly data and device status data based on the edge weight matrix, combine the feature vectors with the importance weights of the nodes to calculate the correlation matrix between features, generate a unified feature representation based on the correlation matrix, and update the unified feature representation as the node attributes to the corresponding nodes in the knowledge graph structure; Construct a root cause inference path according to the updated node attributes and the edge weight matrix. Starting from the detected abnormal node, use the edge weight matrix to calculate the multi-hop propagation probability from the abnormal node to each candidate root cause node, and combine the multi-hop propagation probability with the attribute similarity between nodes to calculate the confidence score of each root cause inference path. Select the node corresponding to the root cause inference path with the highest confidence score as the root cause analysis result.

[0012] Input the root cause analysis result into the policy network to generate a patrol task priority score. Plan the patrol path according to the patrol task priority score in combination with the real-time layout structure of the data center. Continuously optimize the patrol strategy through a double-layer interactive learning mechanism. The realization of the adaptive cooperative scheduling of the patrol robot includes: Input the root cause analysis result into the policy network, calculate the influence range and time urgency of the equipment failure according to the root cause analysis result, and generate a patrol task priority score based on the influence range and the time urgency. The patrol task priority score represents the execution priority order of different patrol tasks; Plan the patrol path according to the patrol task priority score, obtain the equipment distribution information and channel constraint information in the real-time layout structure of the data center, construct an information transfer probability matrix based on the equipment distribution information and the channel constraint information, and combine the information transfer probability matrix with the patrol task priority score to generate the optimal patrol path; Construct a double-layer interactive learning mechanism to continuously optimize the patrol strategy. Collect the actual patrol results in the task evaluation layer to update the weight matrix of the policy network, and optimize the information transfer probability matrix based on the patrol execution efficiency in the path planning layer; Realize the cooperative scheduling of multiple robots based on the optimized patrol strategy. Calculate the priority ranking of multiple patrol tasks according to the updated weight matrix of the policy network, allocate patrol paths for different patrol robots based on the optimized information transfer probability matrix, and realize the adaptive cooperative scheduling of the patrol robot through the joint decision of the priority ranking result and the allocated patrol path result.

[0013] Construct an information transfer probability matrix based on the equipment distribution information and the channel constraint information, and combine the information transfer probability matrix with the patrol task priority score to generate the optimal patrol path, including: Generate an adjacency matrix for node connection based on the device distribution information and the channel constraint information. Input the query matrix, key matrix, and value matrix into the temporal attention calculation unit to obtain the temporal correlation strength between nodes, and calculate the causal relationship strength between nodes according to the adjacency matrix and the temporal correlation strength; Construct a multi-head attention scoring network based on the causal relationship strength. Input the task status information into multiple parallel attention calculation units to obtain attention weights, calculate the initial score of the inspection task priority using the attention weights, calculate the propagation influence of the task priority by combining the status information of adjacent nodes, and generate an inspection task priority score by weighted combination of the propagation influence and the initial priority score; Construct an initial information transfer probability matrix according to the device distribution information and the channel constraint information, dynamically optimize the initial information transfer probability matrix based on the inspection task priority score to obtain the information transfer probability considering the task priority; use the information transfer probability and the inspection task priority score to solve the optimal inspection path through the dynamic programming method.

[0014] In the second aspect of the embodiments of the present invention, there is provided a monitoring and analysis system for a data center inspection robot based on machine vision, including: A first unit for acquiring the monitoring video stream data collected by the data center inspection robot, extracting the real-time image of the data center cabinet area from the monitoring video stream data, and synchronously collecting the temperature distribution data and device load data in the cabinet area as environmental parameter data; A second unit for analyzing the real-time image using a deep learning object detection model to identify the display status of device indicators, the cable connection status, and the cabinet temperature display value in the data center cabinet area, and generating visual abnormal state features; A third unit for constructing a multi-modal data fusion analysis model, performing spatio-temporal registration on the visual abnormal state features and the environmental parameter data, and constructing a device abnormal knowledge graph using an improved graph reasoning algorithm. The improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and generates a root cause analysis result; A fourth unit for inputting the root cause analysis result into a policy network to generate an inspection task priority score, planning an inspection path according to the inspection task priority score in combination with the real-time layout structure of the data center, and continuously optimizing the inspection strategy through a double-layer interactive learning mechanism to achieve the adaptive collaborative scheduling of the inspection robot.

[0015] In the third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0016] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0017] The beneficial effects of this application are as follows: The method for monitoring and analyzing a data center patrol robot based on machine vision provided by the present invention uses a deep learning object detection model to identify the display status of device indicators, the cable connection status, and the cabinet temperature display value, and combines environmental parameter data to achieve a comprehensive monitoring of the status of data center devices, improving the accuracy and integrity of anomaly detection.

[0018] By constructing a multi-modal data fusion analysis model and introducing an improved graph reasoning algorithm, the spatio-temporal registration and collaborative analysis of visual anomaly state features and environmental parameter data are realized, improving the efficiency and accuracy of device anomaly root cause analysis, and effectively reducing the fault diagnosis time and the need for manual intervention.

[0019] Based on the root cause analysis results, a patrol task priority scoring mechanism and a double-layer interactive learning mechanism are used to realize the intelligent planning of the patrol path and the continuous optimization of the patrol strategy, enabling the patrol robot to perform adaptive collaborative scheduling according to the real-time conditions of the data center, and greatly improving the operation and maintenance efficiency and reliability of the data center. Description of the Drawings

[0020] Figure 1 It is a schematic flow chart of the method for monitoring and analyzing a data center patrol robot based on machine vision according to the embodiments of the present invention; Figure 2 It is a bar chart comparing the performance of the data center device status detection model according to the embodiments of the present invention; Figure 3 It is a flow chart of the method for constructing an equipment anomaly knowledge graph based on multi-modal data fusion according to the embodiments of the present invention; Figure 4 It is a logic block diagram of the improved graph reasoning algorithm according to the embodiments of the present invention; Figure 5 It is a bar chart comparing the performance of the patrol path optimization method based on temporal attention and dynamic programming according to the embodiments of the present invention. Detailed Embodiments

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] The following uses specific embodiments to elaborate on the technical solutions of the present invention in detail. These several specific embodiments can be combined with each other, and for the same or similar concepts or processes, they may not be repeated in some embodiments.

[0023] Figure 1 The following is a schematic flowchart of the method for monitoring and analyzing a data center inspection robot based on machine vision in an embodiment of the present invention. As Figure 1 shown, the method includes: Obtain the monitoring video stream data collected by the data center inspection robot, extract the real-time images of the data center cabinet area from the monitoring video stream data, and synchronously collect the temperature distribution data and equipment load data in the cabinet area as environmental parameter data; Use a deep learning object detection model to analyze the real-time images, identify the display status of device indicators, cable connection status, and cabinet temperature display values in the data center cabinet area, and generate visual abnormal state features; Construct a multi-modal data fusion analysis model, perform spatio-temporal registration on the visual abnormal state features and the environmental parameter data, and use an improved graph reasoning algorithm to construct an equipment abnormal knowledge graph. The improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and generates a root cause analysis result; Input the root cause analysis result into a policy network to generate a priority score for the inspection task. According to the priority score of the inspection task and the real-time layout structure of the data center, plan the inspection path, and continuously optimize the inspection strategy through a double-layer interactive learning mechanism to achieve the adaptive collaborative scheduling of the inspection robot.

[0024] In an alternative embodiment, using a deep learning object detection model to analyze the real-time images, identifying the display status of device indicators, cable connection status, and cabinet temperature display values in the data center cabinet area, and generating visual abnormal state features includes: Use a deep learning object detection model to extract the display status of device indicators, cable connection status, and cabinet temperature display values, and construct a causal graph of device status; Based on the device state causal graph, the attention mechanism is used to calculate the feature weights for the indicator light area, cable area, and temperature display area in the real-time image respectively, and the feature weights are weighted and fused with the visual features extracted by the deep learning object detection model to construct a multi-scale visual feature representation; For the abnormal states detected in the multi-scale visual feature representation, a visual feature counterfactual verification function is constructed. The visual feature counterfactual verification function performs abnormal verification by calculating the difference between the state recognition probability after image enhancement processing and the original state recognition probability, and calculates the visual similarity between the candidate abnormal area and the historical abnormal cases based on the results of the abnormal verification to generate the visual abnormal attribution confidence; An abnormal evolution path is constructed according to the visual abnormal attribution confidence, and a warning score calculation function is designed based on the abnormal evolution path. The warning score calculation function comprehensively evaluates the abnormal degree of the device state and the visual feature deviation degree to generate the visual abnormal state features.

[0025] Construct a device state causal graph. The system uses the YOLOv5 object detection model to analyze the real-time images of the data center cabinets collected. The YOLOv5 model is trained with 5000 labeled images, including three types of targets: indicator light status (normal green light, warning yellow light, fault red light, off state), cable connection status (loose, normal, unconnected), and temperature display value (normal range 18 - 27 °C, warning range 28 - 35 °C, dangerous range > 35 °C). The model reaches an average precision of 93.7% (mAP@0.5) on the test set. The detection results include the target category, position coordinates, confidence, and status attributes. For example, for the indicator light detection, the system not only identifies the position of the indicator light but also analyzes its color status (green: normal operation; yellow: warning state; red: fault state; gray: power off). Based on the detection results, the system constructs a device state causal graph to record the dependency relationships between device components, such as the causal chain of "host power indicator light off → network connection indicator light off".

[0026] Based on the device state causal graph, the attention mechanism is used to calculate the feature weights. For each detected area, the system calculates the spatial attention weight and the channel attention weight. The spatial attention mechanism generates an attention map through convolution operations to highlight key areas such as indicator lights, cable interfaces, and temperature displays. For example, the system assigns a weight of 0.3 to the area of a normally working green indicator light, a weight of 0.7 to the yellow indicator light in the warning state, and a weight of 0.9 to the red indicator light in the fault state.

[0027] The channel attention mechanism calculates weights for the RGB channels respectively. For the red indicator light, the weight of the R channel is 0.8, the G channel is 0.1, and the B channel is 0.1. Finally, the system performs weighted fusion of the attention weights with the visual features extracted by YOLOv5 (a feature map with a size of 256×256×256), and constructs multi-scale visual feature representations through FPN (Feature Pyramid Network), including feature maps of three scales: 256×256, 128×128, and 64×64, to adapt to the detection requirements of device components of different sizes.

[0028] For the detected abnormal states, a visual feature counterfactual verification function is constructed. When a suspected abnormal state is detected, the system conducts verification through image enhancement processing. Taking the detection of cable looseness as an example, the system performs contrast enhancement (increasing by 20%), brightness adjustment (increasing by 15%), and sharpening processing (sharpening coefficient 1.5) on the original image, and then conducts object detection again. If the recognition probability of the cable looseness state after enhancement processing increases from the original 0.75 to 0.92, with a difference of 0.17, exceeding the preset threshold of 0.15, it is confirmed as a real abnormality.

[0029] The visual similarity between the current abnormal area and the samples in the historical abnormal case library is also calculated. The similarity calculation uses the cosine distance of feature vectors. For example, the similarity of 0.85 is calculated between the currently detected red fault indicator light and the typical fault indicator light in the historical cases. Based on the abnormal verification results and historical similarities, the system generates a visual abnormal attribution confidence level. For example, for the detected indicator light abnormality, the system gives an attribution confidence level of 0.92.

[0030] An abnormal evolution path is constructed according to the visual abnormal attribution confidence level. By recording the changes in abnormal states in 10 consecutive frames of images, such as the temperature display rising from 26°C to 32°C, and at the same time the indicator light of the cooling fan changing from green to yellow, the system identifies the potential evolution path of the cooling system fault. Based on the abnormal evolution path, the system designs an early warning score calculation function to comprehensively evaluate the degree of device state abnormality and the deviation degree of visual features.

[0031] The early warning score function considers the following factors: abnormal duration (weight 0.3), severity of abnormal state (weight 0.4), similarity to historical cases (weight 0.2), and diffusion trend of abnormal state (weight 0.1). For example, for a red indicator light fault that lasts for 5 minutes, the early warning score is calculated as 0.3×(5 / 10)+0.4×0.9+0.2×0.85+0.1×0.7 = 0.77. Exceeding the preset threshold of 0.7, the system generates a high-priority visual abnormal state feature alarm. The generated visual abnormal state features include abnormal type, location information, abnormal confidence level, early warning score, and recommended handling measures, providing decision-making support for the operation and maintenance personnel of the data center.

[0032] Figure 2 Bar chart for performance comparison of the device status detection model in the embodiments of the present invention: This figure shows the comparison results of the recognition accuracies of three model methods under four different inspection task scenarios. Specifically, it includes four task scenarios: indicator light status recognition, cable connection status, temperature display value recognition, and comprehensive anomaly detection. In each scenario, the performance of the basic model (white column), the attention mechanism enhanced model (diagonal column), and the counterfactual verification enhanced model (grid column) is respectively shown. From the data, the counterfactual verification enhanced model has achieved the best results in all scenarios, with the accuracies reaching 92.5%, 91.5%, 90.0%, and 90.5% respectively. The attention mechanism enhanced model ranks second, with the accuracies in the four scenarios being 89.0%, 87.5%, 85.0%, and 86.0% respectively. The performance of the basic model is the worst, with the accuracies being 82.5%, 80.1%, 77.5%, and 79.0% respectively. The data shows that by introducing the attention mechanism and the counterfactual verification method, the recognition accuracy of the model has been significantly improved, with an average improvement of more than 10 percentage points, verifying the effectiveness of these two enhancement methods. Especially in the complex comprehensive anomaly detection task, the enhanced model has a more obvious advantage compared to the basic model.

[0033] In an alternative embodiment, constructing an anomaly evolution path according to the visual anomaly attribution confidence, and designing a warning score calculation function based on the anomaly evolution path includes: Obtaining a device anomaly state sequence, where the device anomaly state sequence includes an abnormal visual feature vector, a device state feature vector, and a state transition probability, calculating the device anomaly state sequence according to the visual anomaly attribution confidence, and generating a state transition weight; Constructing a multi-dimensional feature fusion function based on the device anomaly state sequence, and performing weighted summation of the state transition weight and the multi-dimensional feature fusion function to obtain a fusion feature representation; Using the fusion feature representation to construct a hierarchical warning score calculation model, calculating the feature distance between the abnormal visual feature vector and the normal state benchmark, multiplying the feature distance by the device state feature vector to obtain an anomaly degree score, calculating the change rate of the state transition probability and combining it with a time decay factor to generate a trend score, and performing weighted fusion of the anomaly degree score, the trend score, and the confidence score in the comprehensive evaluation layer to generate a warning score calculation function.

[0034] In an industrial production environment, a large amount of visual data will be generated during the operation of equipment. By analyzing this data, potential anomalies can be detected early. This embodiment provides a method for constructing an anomaly evolution path based on the visual anomaly attribution confidence and designing a warning score calculation function.

[0035] Obtaining the device abnormal state sequence is the basis for constructing the abnormal evolution path. This sequence consists of three key components: the abnormal visual feature vector, the device state feature vector, and the state transition probability. The abnormal visual feature vector is extracted from the device monitoring images through a deep learning model and can be represented as a 128-dimensional vector; the device state feature vector includes sensor data such as temperature, vibration, and sound, usually a 32-dimensional vector; the state transition probability describes the probability of the device transitioning from one state to another.

[0036] When calculating the device abnormal state sequence according to the visual abnormal attribution confidence, it is necessary to use a pre-trained visual attribution model. This model is based on the attention mechanism and can automatically locate the abnormal areas in the image and give the attribution confidence. For example, when the bearing on the production line shows wear, the model will identify the worn area and give a confidence value of 0.85. Subsequently, this confidence value is associated with each state in the current device state sequence to generate the state transition weight. Specifically, for an abnormality with a confidence higher than 0.7, a higher state transition weight (such as 0.8 - 1.0) is assigned; for an abnormality with a confidence between 0.3 - 0.7, a medium weight (0.4 - 0.7) is assigned; and for an abnormality with a confidence lower than 0.3, a lower weight (0.1 - 0.3) is assigned.

[0037] Constructing a multi-dimensional feature fusion function based on the device abnormal state sequence is a key step in achieving effective early warning. This fusion function adopts an adaptive weight mechanism to dynamically adjust the weights of different features during the fusion process. Specifically, first, the abnormal visual feature vector and the device state feature vector are transformed into the same-dimensional space through a non-linear mapping, such as a 64-dimensional space; then, an attention network is used to calculate the importance scores of each feature; finally, the features are weighted and summed according to these scores. Taking a certain bearing abnormal detection as an example, the visual feature obtains a weight of 0.75, the vibration feature obtains a weight of 0.15, and the temperature feature obtains a weight of 0.10. The state transition weight (such as 0.82) is weighted and summed with the result of the multi-dimensional feature fusion function to obtain the fusion feature representation, which more comprehensively reflects the current abnormal state of the device and its evolution trend.

[0038] Construct a hierarchical early warning score calculation model using fused feature representation. This model consists of three evaluation levels: the anomaly degree evaluation level, the trend evaluation level, and the comprehensive evaluation level. At the anomaly degree evaluation level, calculate the feature distance between the abnormal visual feature vector and the normal state benchmark. The specific method is to use cosine similarity to calculate the difference between the current feature and the normal benchmark feature. The normal benchmark feature is obtained through statistical analysis of a large amount of normal operation data. For example, if the cosine similarity between the current feature of a certain bearing and the normal benchmark is 0.65, it indicates a certain difference, and the feature distance is 0.35. Multiply this feature distance by the device state feature vector to obtain the anomaly degree score. If the temperature anomaly value in the device state feature vector is 0.8, then the anomaly degree score is 0.28.

[0039] At the trend evaluation level, calculate the change rate of the state transition probability and combine it with the time decay factor to generate a trend score. The change rate of the state transition probability is calculated by comparing the state transition probabilities at multiple consecutive time points. For example, the state transition probabilities of a device in the last five detection cycles are 0.1, 0.15, 0.25, 0.4, 0.55 respectively, showing an obvious upward trend, and the change rate is 0.45. The time decay factor uses an exponential decay function, making the most recent state change have a greater impact. If the decay rate is set to 0.9, then the weight of the most recent state change is 1, the previous one is 0.9, the one before that is 0.81, and so on. Combining the change rate and the decay factor, the calculated trend score is 0.38.

[0040] The confidence score directly uses the visual anomaly attribution confidence, such as 0.85 in the previous example. At the comprehensive evaluation level, weight and fuse the anomaly degree score, the trend score, and the confidence score to generate an early warning score. According to the requirements of different production scenarios, the weights of the three scores can be dynamically adjusted. In the field of high-precision manufacturing, more attention is paid to the anomaly degree, giving it a weight of 0.5, and the trend score and the confidence score each account for 0.25; while in a continuous production process, more attention is paid to the trend change, giving the trend score a weight of 0.5, and the other two each account for 0.25. In the previous bearing case, if an equal weight of 0.33 is used, then the final early warning score is 0.33×0.28 + 0.33×0.38 + 0.33×0.85 ≈ 0.50, indicating that the device is in a medium-risk state and needs to be monitored closely.

[0041] The early warning score calculation function constructed through the above steps can comprehensively consider the current abnormal state of the device, the historical evolution trend, and the credibility of the anomaly judgment, providing a scientific basis for equipment maintenance decisions and effectively preventing the occurrence of equipment failures.

[0042] In an alternative embodiment, a multi-modal data fusion analysis model is constructed to perform spatio-temporal registration on the visual anomaly state features and the environmental parameter data. Constructing an equipment anomaly knowledge graph using an improved graph reasoning algorithm includes: Construct a multi-modal data fusion analysis model, input the visual anomaly state features and the environmental parameter data into the multi-modal data fusion analysis model for spatio-temporal registration, and extract the timestamp identifiers and spatial location identifiers of the visual anomaly state features and the environmental parameter data; Construct a temporal correlation matrix based on the timestamp identifiers, construct a spatial adjacency matrix based on the spatial location identifiers, and perform weighted fusion on the temporal correlation matrix and the spatial adjacency matrix to generate a spatio-temporal registration matrix; Construct an equipment anomaly knowledge graph using an improved graph reasoning algorithm. The equipment anomaly knowledge graph includes an entity node layer, an attribute node layer, and a set of relationship edges. Construct entity nodes representing different devices in the entity node layer, construct attribute nodes representing the equipment anomaly state in the attribute node layer, construct entity relationship edges of the entity nodes based on the spatio-temporal registration matrix, construct attribute relationship edges of the attribute nodes based on the equipment state transition rules, and use the improved graph reasoning algorithm to iteratively update the entity relationship edges and the attribute relationship edges to generate an equipment anomaly knowledge graph with a hierarchical structure.

[0043] As Figure 3 shown, the method includes: Constructing a multi-modal data fusion analysis model is to perform spatio-temporal registration on visual anomaly state features and environmental parameter data, and construct an equipment anomaly knowledge graph using an improved graph reasoning algorithm. This method realizes the comprehensive perception and analysis of the equipment anomaly state through multiple key steps.

[0044] When constructing a multi-modal data fusion analysis model, it is first necessary to collect visual anomaly state features and environmental parameter data. Visual anomaly state features can be obtained by deploying high-definition cameras at industrial sites to collect equipment appearance images, such as visual anomalies such as paint peeling and joint loosening on the surface of transformers in substations. Environmental parameter data is obtained through various sensors, including parameters such as temperature, humidity, and vibration. The collected data is input into the multi-modal data fusion analysis model for spatio-temporal registration processing.

[0045] The first step in spatio-temporal registration is to extract timestamp identifiers and spatial location identifiers. For visual data, each frame of image contains the acquisition timestamp and the camera position coordinates; for environmental parameter data, each record contains the measurement time and the sensor installation location. For example, the image data of a certain transformer contains a timestamp of "2023-04-15 14:30:25" and spatial coordinates of "X=105.36, Y=78.92, Z=2.35", while the temperature sensor data in the same area contains a timestamp of "2023-04-15 14:30:20" and spatial coordinates of "X=105.40, Y=78.95, Z=2.38".

[0046] Construct a temporal correlation matrix. This matrix describes the degree of association between different data points in the time dimension. Temporal correlation can be determined by calculating the difference in timestamps. The smaller the difference, the stronger the correlation. For example, for two data points with a time difference of 5 seconds, a correlation weight of 0.95 can be assigned; while for data points with a time difference of 60 seconds, the correlation weight drops to 0.6. In practical applications, the decay function of time correlation can be adjusted according to the characteristics of different devices.

[0047] Construct a spatial adjacency matrix, which represents the degree of association between different data points in the spatial dimension. The adjacency relationship can be calculated by Euclidean distance or other spatial distance measurement methods. For example, for two data points within a spatial distance of 0.5 meters, a spatial association weight of 0.9 can be assigned; while for data points with a distance of 5 meters, the spatial association weight drops to 0.3. In an actual industrial environment, spatial association also needs to consider the physical layout of the devices and the situation of obstacles.

[0048] Weightedly fuse the temporal correlation matrix and the spatial adjacency matrix to generate a spatio-temporal registration matrix. The weighting method can adopt linear combination. For example, set the weight of the time dimension to 0.6 and the weight of the spatial dimension to 0.4 to calculate the comprehensive correlation score. In practical applications, these weights can be optimized and adjusted according to specific scenarios. For example, for devices with fast temperature changes, the weight of the time dimension needs to be increased.

[0049] Constructing an equipment anomaly knowledge graph using an improved graph reasoning algorithm is a key step. The knowledge graph consists of three main components: an entity node layer, an attribute node layer, and a set of relationship edges. In the entity node layer, entity nodes representing different devices are constructed. Each device corresponds to an entity node, which records basic information such as device ID, type, installation location, etc. For example, Transformer A can be represented as an entity node, containing attributes such as "ID=TR001", "Type=Oil-immersed transformer", "Location=Distribution Room Area B".

[0050] At the attribute node layer, construct attribute nodes representing the abnormal states of the devices. The attribute nodes contain information such as the type of abnormality, severity, duration, etc. For example, "abnormal oil temperature" can be represented as an attribute node with attributes such as "type of abnormality = temperature exceeding the standard", "severity = medium", "duration = 30 minutes".

[0051] Construct entity relationship edges for entity nodes based on the spatio-temporal registration matrix. The entity relationship edges describe relationships such as physical connections and functional dependencies between different devices. For example, a "electrical connection" relationship with a weight of 0.85 can be established between transformer A and circuit breaker B, indicating that they are closely electrically connected.

[0052] Construct attribute relationship edges for attribute nodes based on the device state transition rules. The attribute relationship edges describe causal and transformation relationships between different abnormal states. For example, a "causes" relationship with a probability of 0.75 can be established between "abnormal oil temperature" and "pressure increase", indicating that the abnormal oil temperature causes the pressure to increase.

[0053] Use an improved graph reasoning algorithm to iteratively update the entity relationship edges and attribute relationship edges. This algorithm combines the attention mechanism and propagation rules, and iteratively calculates node features and edge weights. Specifically, in each iteration, the algorithm first calculates the weighted features of neighbor nodes, then updates the representation of the central node, and at the same time adjusts the edge weights according to the changes in node states. For example, when it is found that the abnormal oil temperature of transformer A is highly correlated with the tripping event of circuit breaker B, the association weight between these two nodes will be enhanced.

[0054] After multiple rounds of iteration, a device abnormal knowledge graph with a hierarchical structure is generated. This graph not only shows the physical associations between devices, but also presents the propagation paths and causal chains of abnormal states. By analyzing features such as node centrality and community structure in the knowledge graph, key devices and core fault sources can be identified. For example, if multiple device abnormalities are related to a specific node, that node is the root cause of the fault and should be inspected and repaired first.

[0055] This method combining multi-modal data fusion and knowledge graph can effectively capture the spatio-temporal features and association rules of device abnormal states, providing strong technical support for predictive maintenance and fault diagnosis of industrial devices.

[0056] In an alternative embodiment, the improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and the root cause analysis results include: Construct the knowledge graph structure of the improved graph reasoning algorithm, and introduce a dynamic weight adaptation mechanism into the knowledge graph structure. The dynamic weight adaptation mechanism calculates the association strength between nodes to obtain the importance weights of the nodes, and constructs an edge weight matrix based on the importance weights of the nodes. The edge weight matrix represents the strength of the propagation relationship between nodes in the knowledge graph; Extract corresponding feature vectors from the visual anomaly data and device status data based on the edge weight matrix, combine the feature vectors with the importance weights of the nodes to calculate the correlation matrix between features, generate a unified feature representation based on the correlation matrix, and update the unified feature representation as the node attributes to the corresponding nodes in the knowledge graph structure; Construct a root cause inference path according to the updated node attributes and the edge weight matrix. Starting from the detected abnormal node, use the edge weight matrix to calculate the multi-hop propagation probability from the abnormal node to each candidate root cause node, combine the multi-hop propagation probability with the attribute similarity between nodes to calculate the confidence score of each root cause inference path, and select the node corresponding to the root cause inference path with the highest confidence score as the root cause analysis result.

[0057] As Figure 4 shown, the method includes: Construct a knowledge graph structure, which includes device nodes, status nodes, fault nodes and relationship edges. For example, in an industrial production line, each device such as a pump, valve, sensor, etc. can be used as a device node, temperature anomalies, pressure fluctuations, etc. as status nodes, and pipeline leaks, motor overheating, etc. as fault nodes. Nodes are connected by relationship edges such as "causes", "associated with", "belongs to", etc. to form an initial knowledge graph.

[0058] Introduce a dynamic weight adaptation mechanism into the constructed knowledge graph structure. This mechanism obtains the importance weights of nodes by calculating the association strength between nodes. Specifically, for a certain node A, calculate its connection frequency with adjacent nodes. For example, if node A and node B appear together 50 times in historical data, and with node C appear together 10 times, then the association strength of node B with respect to node A is 0.83 (50 / (50 + 10)). Repeat this process for all nodes in the whole graph to obtain the importance weight values of each node. For example, in an actual industrial scenario, the pump node has an importance weight of 0.75, while a certain auxiliary sensor node only has an importance weight of 0.3.

[0059] Construct an edge weight matrix based on the importance weights of nodes. This matrix characterizes the strength of the propagation relationship between nodes in the knowledge graph. For an edge connecting node i and node j, its weight calculation depends on the importance weights of the two end nodes and their historical co-occurrence frequencies. Taking a certain production equipment as an example, if the historical co-occurrence frequency of the pressure sensor node (importance weight 0.8) and the valve failure node (importance weight 0.7) is 0.6, then the edge weight between them can be calculated as 0.8 × 0.7 × 0.6 = 0.336. Fill the weight values of all edges into the matrix to obtain a complete edge weight matrix.

[0060] Extract corresponding feature vectors from visual anomaly data and equipment status data based on the edge weight matrix. For visual anomaly data, image processing techniques are used to extract features such as edge features, texture features, color distribution features, etc. For example, for the image data of the equipment surface, the extracted features may include the equipment surface temperature distribution feature (the proportion of the hot spot area is 25%, and the temperature gradient value is 0.4), the surface crack feature (the crack length is 5 mm, and the width is 0.8 mm), etc. For equipment status data, features such as equipment vibration frequency, current fluctuation, and temperature change rate are extracted. For example, for a certain pump equipment, the extracted status features may include vibration frequency (23 Hz), current fluctuation (±2 A), bearing temperature change rate (1.5 °C / hour), etc.

[0061] Combine the extracted feature vectors with the importance weights of nodes to calculate the correlation matrix between features. For each pair of features, calculate their degree of correlation and multiply it by the importance weights of the corresponding nodes. For example, the original correlation coefficient between the equipment vibration frequency feature and the bearing temperature change rate feature is 0.75, and the importance weights of the corresponding nodes are 0.8 and 0.7 respectively. Then the correlation value of this pair of features in the correlation matrix is 0.75 × 0.8 × 0.7 = 0.42.

[0062] Generate a unified feature representation based on the correlation matrix. Weight and fuse all features according to their correlations to obtain a feature representation with a unified dimension. For example, for a certain equipment node, its unified feature representation can be a 128-dimensional vector containing equipment status parameters and visual features, where the first 64 dimensions represent equipment status parameters such as temperature, pressure, vibration, etc., and the last 64 dimensions represent visual features such as color, texture, shape, etc. Update the generated unified feature representation as the node attribute to the corresponding node in the knowledge graph structure.

[0063] Construct a root cause inference path based on the updated node attributes and edge weight matrix. Taking the detected abnormal node as the starting point, such as the temperature abnormal node of a certain device, calculate the multi-hop propagation probability from the abnormal node to each candidate root cause node using the edge weight matrix. The multi-hop propagation probability is calculated by accumulating the edge weights. For example, the propagation probability of the three-hop path from the temperature abnormal node to the bearing wear node is 0.336×0.285×0.412 = 0.0395.

[0064] Combine the multi-hop propagation probability with the attribute similarity between nodes to calculate the confidence score of each root cause inference path. The attribute similarity between nodes is calculated by comparing the unified feature representations of the nodes. For example, the attribute similarity between the temperature abnormal node and the bearing wear node is 0.78, then the final confidence score of this path is 0.0395×0.78 = 0.03081.

[0065] Select the node corresponding to the root cause inference path with the highest confidence score as the root cause analysis result. In an actual case, the system finds that the root cause of the temperature anomaly is bearing wear (confidence 0.03081), rather than motor overheating (confidence 0.02755) or cooling system failure (confidence 0.02103). The system finally outputs bearing wear as the root cause of the temperature anomaly and gives the corresponding inference path and confidence score.

[0066] In an alternative embodiment, input the root cause analysis result into a policy network to generate a patrol task priority score, plan a patrol path according to the patrol task priority score combined with the real-time layout structure of the data center, and continuously optimize the patrol strategy through a double-layer interactive learning mechanism to achieve the adaptive cooperative scheduling of the patrol robot, including: Input the root cause analysis result into the policy network, calculate the influence range and time urgency of the equipment failure according to the root cause analysis result, and generate a patrol task priority score based on the influence range and the time urgency. The patrol task priority score represents the execution priority order of different patrol tasks; Plan a patrol path according to the patrol task priority score, obtain the equipment distribution information and channel constraint information in the real-time layout structure of the data center, construct an information transfer probability matrix based on the equipment distribution information and the channel constraint information, and combine the information transfer probability matrix with the patrol task priority score to generate an optimal patrol path; Construct a double-layer interactive learning mechanism to continuously optimize the patrol strategy. In the task evaluation layer, collect the actual patrol results to update the weight matrix of the policy network, and in the path planning layer, optimize the information transfer probability matrix based on the patrol execution efficiency; Based on the optimized inspection strategy, multi-robot collaborative scheduling is realized. The priority ranking of multiple inspection tasks is calculated according to the weight matrix of the updated policy network. The inspection paths are allocated to different inspection robots based on the optimized information transfer probability matrix. The adaptive collaborative scheduling of the inspection robots is achieved through the joint decision of the results of the priority ranking and the results of the allocated inspection paths.

[0067] The root cause analysis result is input into the policy network to generate the priority score of the inspection task. According to the priority score of the inspection task and the real-time layout structure of the data center, the inspection path is planned. The inspection strategy is continuously optimized through a two-layer interactive learning mechanism. The root cause analysis result is input into the policy network to generate the priority score of the inspection task. Specifically, a deep neural network is constructed as the policy network, which includes an input layer, multiple hidden layers, and an output layer. The input layer receives the multi-dimensional feature vector of the root cause analysis, including the fault type encoding (such as 0 represents temperature anomaly, 1 represents humidity anomaly, 2 represents air flow blockage, 3 represents power supply anomaly, etc.), the fault level (such as a numerical value from 0 to 5, the larger the value, the higher the fault level), the historical fault frequency (such as the number of faults of a certain device within 30 days), and other features. The hidden layer uses the ReLU activation function for non-linear mapping, and the output layer generates a priority score value between 0 and 100.

[0068] When evaluating the influence range of equipment failures, the system detects the number of other devices that have a dependency relationship with the faulty device. For example, a core switch failure affects 50 servers, and its influence range value is 50; while a common server failure only affects itself, and the influence range value is 1. The time urgency is calculated according to the fault development trend and the predicted time window for the fault to expand. For example, the urgency value of a server with rapidly rising temperature is 90, and the urgency value of a non-critical alarm in a stable state is 30. Through these input features, the policy network outputs a comprehensively considered priority score. For example, the overheating fault of a core network device is scored 95, and the fan alarm of a non-critical path server is scored 45.

[0069] The inspection path is planned according to the generated priority score of the inspection task. The system first obtains the real-time layout structure information of the data center, including the three-dimensional coordinate positions of the devices, the cabinet arrangement, the channel direction, etc. For the device distribution information, the system records the position coordinates (x, y, z) of each device to be inspected. For example, device A is located at (10, 20, 15), and device B is located at (10, 25, 15). For the channel constraint information, the system marks the passable areas and restricted areas. For example, the cold channel is marked as passable with 1, the hot channel is marked as restricted with 0, and the special area with direction restrictions is marked as 2. Based on this information, an information transfer probability matrix is constructed, which represents the transfer possibility from one device position to another device position.

[0070] The transfer probability between adjacent cabinet devices is 0.9, the transfer probability between devices across multiple channels is 0.3, and the transfer probability in the restricted area is 0. Combine the transfer probability matrix with the inspection task priority score, and use the improved ant colony algorithm to generate the optimal inspection path. For example, for three devices with priority scores of 95, 75, and 60 respectively, the system calculates the total utility value of various access orders. For example, the utility value of the path [Device 1 → Device 3 → Device 2] is 950, and the utility value of the path [Device 1 → Device 2 → Device 3] is 920. Finally, the access order with the highest utility value is selected as the optimal inspection path.

[0071] Construct a two-layer interactive learning mechanism to continuously optimize the inspection strategy. In the task evaluation layer, the system records the actual results of each inspection, including indicators such as the fault confirmation rate (e.g., 85% of the devices predicted to have high-risk faults actually have serious problems), the false alarm rate (e.g., 15% of the alarms are false alarms), and the average processing time (e.g., the average inspection time for each fault point is 3.5 minutes).

[0072] Based on these actual data, use the gradient descent method to update the weight matrix of the policy network. For example, when it is found that the priority of a certain type of fault is overestimated, reduce the weight value of the corresponding feature from 0.75 to 0.65. In the path planning layer, the system statistics the execution efficiency of the inspection path, including the path completion time (e.g., the expected time is 90 minutes and the actual time is 105 minutes), the path following rate (e.g., 80% of the planned path is accurately executed), etc. According to these statistical results, the system updates the information transfer probability matrix. For example, if it is found that the time taken to cross a certain channel area far exceeds the expectation, reduce the transfer probability of this area from 0.7 to 0.5. Through this two-layer learning mechanism, the system can continuously optimize the inspection strategy from the actual operation experience.

[0073] Based on the optimized inspection strategy, realize multi-robot collaborative scheduling. The system recalculates the priority ranking of all current inspection tasks to be performed according to the updated weight matrix of the policy network. For example, there are 20 inspection points in the data center, and the calculated priorities are sorted from high to low as [Device 7, Device 15, Device 3,...]. Based on the optimized information transfer probability matrix, the system uses the clustering algorithm to group the inspection points with similar geographical locations, and then assigns these groups to different inspection robots.

[0074] If there are 3 inspection robots in the data center, the system divides 20 inspection points into 3 areas. Robot A is responsible for 7 points in the northeast area, Robot B is responsible for 6 points in the southwest area, and Robot C is responsible for 7 points in the central area. During the task allocation process, the system takes into account both the priority ranking and the spatial distribution to ensure that high-priority tasks can be processed in a timely manner, while avoiding path crossing and conflict between robots. Finally, the system realizes the adaptive cooperative scheduling of inspection robots through the joint decision of the priority ranking result and the inspection path allocation result, improving the overall inspection efficiency.

[0075] In an alternative embodiment, constructing an information transfer probability matrix based on the device distribution information and the channel constraint information, and combining the information transfer probability matrix with the inspection task priority score to generate an optimal inspection path includes: Generating an adjacency matrix of node connections based on the device distribution information and the channel constraint information, inputting the query matrix, key matrix, and value matrix into the temporal attention calculation unit to obtain the temporal correlation strength between nodes, and calculating the causal relationship strength between nodes according to the adjacency matrix and the temporal correlation strength; Constructing a multi-head attention scoring network based on the causal relationship strength, inputting the task status information into multiple parallel attention calculation units to obtain attention weights, calculating the initial priority score of the inspection task using the attention weights, calculating the propagation influence of the task priority in combination with the status information of adjacent nodes, and generating an inspection task priority score by weighted combination of the propagation influence and the initial priority score; Constructing an initial information transfer probability matrix according to the device distribution information and the channel constraint information, dynamically optimizing the initial information transfer probability matrix based on the inspection task priority score to obtain an information transfer probability considering task priority; using the information transfer probability and the inspection task priority score to solve the optimal inspection path through dynamic programming method.

[0076] When generating an adjacency matrix of node connections based on the device distribution information and the channel constraint information, the system first obtains the position coordinates of all devices inside the factory and the limitation conditions of the connecting channels. For example, in a chemical plant with 10 devices, the system records the two-dimensional coordinate positions of each device, such as device 1 is located at (10, 15), device 2 is located at (25, 30), etc. At the same time, the system obtains the channel constraint information, including which devices are directly connected by channels and which channels are temporarily closed due to maintenance or safety reasons. The system constructs a 10×10 adjacency matrix according to this information. The element aij in the matrix represents whether there is a direct channel connection between device i and device j. If there is a connection, it is assigned a value of 1, otherwise 0.

[0077] When calculating the temporal attention, first extract the query matrix, key matrix, and value matrix of the device from the historical inspection data. The query matrix represents the features of the current node, and the key matrix and value matrix represent the features of the target node. Specifically, the system extracts features such as the inspection frequency, failure rate, and last inspection time of each device from the historical inspection records, and constructs a feature matrix with a dimension of 10×6 (assuming 6 features are extracted for each device).

[0078] Generate the query matrix, key matrix, and value matrix respectively through linear transformation of this feature matrix. The dimensions of the three matrices are all 10×6. The system inputs these three matrices into the temporal attention calculation unit to calculate the temporal correlation strength between nodes. The temporal correlation strength is represented as a 10×10 matrix, where the element tij represents the temporal correlation degree between device i and device j, and the numerical range is from 0 to 1.

[0079] Calculate the causal relationship strength between nodes according to the adjacency matrix and the temporal correlation strength, that is, perform element-wise multiplication of the adjacency matrix and the temporal correlation strength matrix to obtain a new 10×10 matrix, which represents the comprehensive causal relationship strength considering physical connection and temporal correlation. For example, if there is a direct channel between device 1 and device 2 (the adjacency matrix value is 1) and the temporal correlation strength is 0.8, then the causal relationship strength is 0.8. If there is no direct channel connection, the causal relationship strength is 0.

[0080] When constructing the multi-head attention scoring network based on the causal relationship strength, the system creates 4 parallel attention calculation units (taking 4-head attention as an example). The system takes the task status information of each device as input, including features such as device running time, failure risk level, and maintenance urgency, to form a 10×8 task status matrix (assuming each device has 8 status features).

[0081] Input this task status matrix into the 4 attention calculation units respectively, and each unit independently calculates the attention weights. In each attention head, the system maps the task status matrix into the query matrix, key matrix, and value matrix respectively, with dimensions of 10×2 (8-dimensional features are divided into 4 heads, 2 dimensions for each head). The system uses the causal relationship strength to constrain the attention calculation process to ensure that effective attention weights are generated only between nodes with causal relationships.

[0082] Calculate the initial priority score of the inspection task using the attention weights. Specifically, the system splices the outputs obtained from the 4 attention heads and performs a linear mapping to obtain the initial priority scores of 10 devices. For example, the score of device 1 is 85, and the score of device 2 is 63, etc. The system combines the status information of adjacent nodes to calculate the propagation influence of the task priority, that is, consider the influence of the status of neighboring devices of each device on the priority of this device.

[0083] For example, if there are high-risk devices among the adjacent devices of device 1, the priority of device 1 will be increased. The system calculates the propagation influence value through the graph propagation mechanism. For example, the propagation influence value of device 1 is +5, and the propagation influence value of device 2 is -2. The system combines the propagation influence and the initial priority score with weights to generate the final priority score of the inspection task. The weight ratio is 0.8:0.2. For example, the final priority score of device 1 is 85×0.8 + 5×0.2 = 69, and the final priority score of device 2 is 63×0.8 + (-2)×0.2 = 50.

[0084] Construct an initial information transfer probability matrix based on the device distribution information and channel constraint information. This matrix represents the probability of transferring from one device to another. The system calculates the transfer probability based on the distance between devices and channel constraints. For example, the transfer probability between devices that are relatively close and have a channel connection is relatively high. The system generates a 10×10 initial transfer probability matrix, where the element pij represents the probability of transferring from device i to device j, and the sum of all probabilities is 1.

[0085] Dynamically optimize the initial information transfer probability matrix based on the inspection task priority score. Specifically, the system normalizes the task priority score to obtain the priority weight. The system uses the priority weight to adjust the initial transfer probability, increasing the transfer probability of high-priority devices and decreasing the transfer probability of low-priority devices.

[0086] For example, if the priority score of device j is 50 and the normalized weight is 0.06, the adjusted transfer probability from device i to device j is pij×(1 + 0.06). After the system completes the adjustment of the transfer probability between all devices, each row is normalized to ensure that the sum of the probabilities of transferring from any device to other devices is 1.

[0087] Use the information transfer probability and the inspection task priority score to solve the optimal inspection path through the dynamic programming method. The system defines the state as the currently visited device and the set of already visited devices, and defines the transfer as moving from the current device to the next unvisited device. The system uses the optimized transfer probability matrix as the state transfer probability, with the objective function of maximizing the total benefit of visiting high-priority devices. The system solves the optimal path through the dynamic programming algorithm and obtains a device access sequence of length 10, such as [1, 3, 5, 7, 9, 10, 8, 6, 4, 2], indicating that the inspection personnel should visit each device in this order to optimize the inspection efficiency.

[0088] Figure 5 This is a bar chart showing the performance comparison of the inspection path optimization method based on temporal attention and dynamic programming in the embodiments of the present invention: The figure shows the performance comparison results of three different inspection path planning methods in four evaluation dimensions. The evaluation dimensions include path length optimization, time efficiency improvement, task priority handling, and comprehensive performance scoring. In the figure, the performance of the traditional path planning method is represented by white bars, the time series attention enhancement method is represented by diagonal bars, and the dynamic programming optimization method is represented by grid bars. From the data analysis, the dynamic programming optimization method has achieved the best performance in all evaluation dimensions, with a path length optimization of 92.5%, a time efficiency improvement of 91.5%, a task priority handling of 94.0%, and a comprehensive performance score of 95.0%. The performance of the time series attention enhancement method is the second best, with the performance in the four dimensions being 87.5%, 86.5%, 86.0%, and 88.5% respectively. The traditional path planning method performs the worst, with the index values in each dimension being 79.0%, 81.0%, 77.0%, and 80.1% respectively. The data shows that by introducing the time series attention mechanism and the dynamic programming optimization strategy, the overall performance of the inspection path planning has been significantly improved, with an average improvement of about 15 percentage points. Especially in terms of the comprehensive performance score, the optimized method has a more obvious advantage compared with the traditional method, fully demonstrating the effectiveness of the proposed optimization method.

[0089] In the second aspect of the embodiments of the present invention, a machine vision-based monitoring and analysis system for data center inspection robots is provided, including: A first unit for obtaining the monitoring video stream data collected by the data center inspection robot, extracting the real-time image of the data center cabinet area from the monitoring video stream data, and synchronously collecting the temperature distribution data and equipment load data in the cabinet area as environmental parameter data; A second unit for analyzing the real-time image by using a deep learning object detection model, identifying the display status of the device indicator lights, the cable connection status, and the cabinet temperature display value in the data center cabinet area, and generating visual abnormal state features; A third unit for constructing a multi-modal data fusion analysis model, performing spatio-temporal registration on the visual abnormal state features and the environmental parameter data, and constructing an equipment abnormal knowledge graph by using an improved graph reasoning algorithm. The improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and generates a root cause analysis result; A fourth unit for inputting the root cause analysis result into a policy network to generate an inspection task priority score, planning an inspection path according to the inspection task priority score in combination with the real-time layout structure of the data center, and continuously optimizing the inspection strategy through a double-layer interaction learning mechanism to realize the adaptive collaborative scheduling of the inspection robot.

[0090] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0091] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0092] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0093] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for monitoring and analyzing a data center inspection robot based on machine vision, characterized in that, Including: Obtain the monitoring video stream data collected by the inspection robot in the data center, extract the real-time image of the cabinet area in the data center from the monitoring video stream data, and synchronously collect the temperature distribution data and equipment load data in the cabinet area as environmental parameter data; Use a deep learning object detection model to analyze the real-time image, identify the display status of device indicators, cable connection status, and cabinet temperature display value in the data center cabinet area, and generate visual abnormal state features; Construct a multi-modal data fusion analysis model, perform spatio-temporal registration on the visual abnormal state features and the environmental parameter data, and use an improved graph reasoning algorithm to construct an equipment abnormal knowledge graph. The improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data through the introduction of a dynamic weight adaptive mechanism, and generates a root cause analysis result; Input the root cause analysis result into the policy network to generate an inspection task priority score, plan an inspection path according to the inspection task priority score in combination with the real-time layout structure of the data center, and continuously optimize the inspection strategy through a double-layer interactive learning mechanism to achieve the adaptive collaborative scheduling of the inspection robot.

2. The method according to claim 1, wherein Using a deep learning object detection model to analyze the real-time image, identifying the display status of device indicators, cable connection status, and cabinet temperature display value in the data center cabinet area, and generating visual abnormal state features includes: Use a deep learning object detection model to extract the display status of device indicators, cable connection status, and cabinet temperature display value, and construct a device state causal graph; Based on the device state causal graph, use the attention mechanism to calculate the feature weights for the indicator area, cable area, and temperature display area in the real-time image respectively, and perform weighted fusion on the feature weights and the visual features extracted by the deep learning object detection model to construct a multi-scale visual feature representation; For the abnormal states detected in the multi-scale visual feature representation, construct a visual feature counterfactual verification function. The visual feature counterfactual verification function performs abnormal verification by calculating the difference between the state recognition probability after image enhancement processing and the original state recognition probability, and calculates the visual similarity between the candidate abnormal area and the historical abnormal cases based on the result of the abnormal verification to generate a visual abnormal attribution confidence; Construct an abnormal evolution path according to the visual abnormal attribution confidence, and design a warning score calculation function based on the abnormal evolution path. The warning score calculation function comprehensively evaluates the abnormal degree of the device state and the visual feature deviation degree to generate visual abnormal state features.

3. The method according to claim 2, wherein Construct an abnormal evolution path according to the visual abnormal attribution confidence, and the design of a warning score calculation function based on the abnormal evolution path includes: Obtain the device abnormal state sequence, which includes abnormal visual feature vectors, device state feature vectors, and state transition probabilities, calculate the device abnormal state sequence according to the visual abnormal attribution confidence, and generate state transition weights; Construct a multi-dimensional feature fusion function based on the device abnormal state sequence, and perform weighted summation on the state transition weights and the multi-dimensional feature fusion function to obtain a fusion feature representation; Construct a hierarchical early warning score calculation model using the fused feature representation, calculate the feature distance between the abnormal visual feature vector and the normal state benchmark, multiply the feature distance by the device state feature vector to obtain the abnormal degree score, calculate the change rate of the state transition probability and generate a trend score in combination with a time decay factor, and generate an early warning score calculation function by weighted fusion of the abnormal degree score, the trend score, and the confidence score in the comprehensive evaluation layer.

4. The method according to claim 1, characterized in that Construct a multi-modal data fusion analysis model, perform spatio-temporal registration on the visual abnormal state features and the environmental parameter data, and construct a device abnormal knowledge graph using an improved graph reasoning algorithm, including: Construct a multi-modal data fusion analysis model, input the visual abnormal state features and the environmental parameter data into the multi-modal data fusion analysis model for spatio-temporal registration, and extract the timestamp identifiers and spatial location identifiers of the visual abnormal state features and the environmental parameter data; Construct a temporal correlation matrix based on the timestamp identifiers, construct a spatial adjacency matrix based on the spatial location identifiers, and perform weighted fusion on the temporal correlation matrix and the spatial adjacency matrix to generate a spatio-temporal registration matrix; Construct a device abnormal knowledge graph using an improved graph reasoning algorithm. The device abnormal knowledge graph includes an entity node layer, an attribute node layer, and a relationship edge set. Construct entity nodes representing different devices in the entity node layer, construct attribute nodes representing device abnormal states in the attribute node layer, construct entity relationship edges of the entity nodes based on the spatio-temporal registration matrix, construct attribute relationship edges of the attribute nodes based on device state transition rules, and use the improved graph reasoning algorithm to iteratively update the entity relationship edges and the attribute relationship edges to generate a device abnormal knowledge graph with a hierarchical structure.

5. The method according to claim 1, wherein The improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and generates a root cause analysis result, including: Construct the knowledge graph structure of the improved graph reasoning algorithm, introduce a dynamic weight adaptive mechanism in the knowledge graph structure. The dynamic weight adaptive mechanism calculates the association strength between nodes to obtain the importance weights of the nodes, constructs an edge weight matrix based on the importance weights of the nodes, and the edge weight matrix characterizes the propagation relationship strength between nodes in the knowledge graph; Extract corresponding feature vectors from the visual abnormal data and the device state data based on the edge weight matrix, combine the feature vectors with the importance weights of the nodes to calculate the correlation matrix between features, generate a unified feature representation based on the correlation matrix, and update the unified feature representation as the node attribute to the corresponding nodes in the knowledge graph structure; Construct a root cause inference path based on the updated node attributes and the edge weight matrix. Starting from the detected abnormal node, use the edge weight matrix to calculate the multi-hop propagation probability from the abnormal node to each candidate root cause node. Combine the multi-hop propagation probability with the attribute similarity between nodes to calculate the confidence score of each root cause inference path, and select the node corresponding to the root cause inference path with the highest confidence score as the root cause analysis result.

6. The method according to claim 1, characterized in that, Input the root cause analysis result into the policy network to generate a patrol task priority score. Plan the patrol path according to the patrol task priority score in combination with the real-time layout structure of the data center. Continuously optimize the patrol strategy through a double-layer interactive learning mechanism to achieve the adaptive cooperative scheduling of patrol robots, including: Input the root cause analysis result into the policy network, calculate the influence range and time urgency of the equipment failure based on the root cause analysis result, and generate a patrol task priority score based on the influence range and the time urgency. The patrol task priority score represents the execution priority order of different patrol tasks; Plan the patrol path according to the patrol task priority score, obtain the equipment distribution information and channel constraint information in the real-time layout structure of the data center, construct an information transfer probability matrix based on the equipment distribution information and the channel constraint information, and combine the information transfer probability matrix with the patrol task priority score to generate an optimal patrol path; Construct a double-layer interactive learning mechanism to continuously optimize the patrol strategy. In the task evaluation layer, collect the actual patrol results to update the weight matrix of the policy network, and in the path planning layer, optimize the information transfer probability matrix based on the patrol execution efficiency; Implement multi-robot cooperative scheduling based on the optimized patrol strategy. Calculate the priority ranking of multiple patrol tasks according to the weight matrix of the updated policy network, allocate patrol paths for different patrol robots based on the optimized information transfer probability matrix, and achieve the adaptive cooperative scheduling of patrol robots through the joint decision of the priority ranking result and the allocated patrol path result.

7. The method according to claim 6, wherein Construct an information transfer probability matrix based on the equipment distribution information and the channel constraint information, and combine the information transfer probability matrix with the patrol task priority score to generate an optimal patrol path, including: Generate an adjacency matrix of node connections based on the equipment distribution information and the channel constraint information. Input the query matrix, key matrix, and value matrix into the temporal attention calculation unit to obtain the temporal correlation strength between nodes. Calculate the causal relationship strength between nodes according to the adjacency matrix and the temporal correlation strength; Construct a multi-head attention scoring network based on the causal relationship strength. Input the task status information into multiple parallel attention calculation units to obtain attention weights, use the attention weights to calculate the initial priority score of the patrol task, combine the status information of adjacent nodes to calculate the propagation influence of the task priority, and generate the patrol task priority score by weighted combination of the propagation influence and the initial priority score; Construct an initial information transfer probability matrix based on the device distribution information and the channel constraint information, and dynamically optimize the initial information transfer probability matrix based on the inspection task priority score to obtain the information transfer probability considering the task priority; use the information transfer probability and the inspection task priority score to solve the optimal inspection path through the dynamic programming method.

8. A machine vision-based monitoring and analysis system for data center inspection robots, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that Including: The first unit is used to obtain the monitoring video stream data collected by the inspection robot in the data center, extract the real-time image of the cabinet area in the data center from the monitoring video stream data, and synchronously collect the temperature distribution data and equipment load data in the cabinet area as environmental parameter data; The second unit is used to analyze the real-time image by using the deep learning object detection model, identify the display status of the device indicator lights, the cable connection status, and the cabinet temperature display value in the cabinet area of the data center, and generate visual abnormal state features; The third unit is used to construct a multi-modal data fusion analysis model, perform spatio-temporal registration on the visual abnormal state features and the environmental parameter data, and use an improved graph reasoning algorithm to construct an equipment abnormal knowledge graph. The improved graph reasoning algorithm realizes the collaborative analysis of multi-source heterogeneous data by introducing a dynamic weight adaptive mechanism, and generates a root cause analysis result; The fourth unit is used to input the root cause analysis result into the policy network to generate an inspection task priority score, plan an inspection path according to the inspection task priority score in combination with the real-time layout structure of the data center, and continuously optimize the inspection strategy through a double-layer interactive learning mechanism to realize the adaptive collaborative scheduling of the inspection robot.

9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Power transmission line intelligent inspection method based on knowledge graph

    CN117521800A

  • Intelligent power plant joint inspection system and method based on video AI and robot

    CN119519156A

  • Health state assessment method for equipment based on knowledge graph attention network

    US20240403599A1

Cited By

  • Large model service chained call tracking system and root cause analysis method

    CN120598062A

  • Intelligent operation and maintenance method and system for photovoltaic inspection robot and electronic equipment

    CN120655276A

  • Automatic fault analysis method and device for computer terminal

    CN120743610A

  • Production line adjustment control system and method based on artificial intelligence

    CN120848427A

  • Power equipment robot inspection system based on AI vision

    CN120871897A