Power equipment multi-mode monitoring and intelligent early warning method based on inspection robot

By monitoring the vibration, sound, and temperature parameters of power equipment in parallel on an inspection robot, dynamically reconstructing the sensor group and sampling strategy, acquiring multimodal high-frequency data streams, and constructing a multimodal fusion network for fault diagnosis, the problem of insufficient fault mode adaptability in existing inspection technologies is solved, and efficient fault detection and early warning are achieved.

CN121769719APending Publication Date: 2026-03-31STATE GRID LIAONING ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing inspection technologies lack dynamic perception mechanisms triggered by events, making it impossible to flexibly adjust monitoring strategies. Furthermore, they are not adaptable enough to newly emerging fault modes, leading to missed or false alarms in single-modal data acquisition, and failing to fully reflect complex fault symptoms.

Method used

Based on the inspection robot, the vibration, sound and temperature parameters of power equipment are monitored in parallel. A lightweight real-time discrimination model is used to identify early anomalies, dynamically reconstruct the optimal sensor set and sampling parameter strategy, obtain time-stricken aligned multimodal high-frequency data streams, construct a multimodal fusion network for feature extraction and fusion analysis, and combine knowledge graphs and reinforcement learning for fault diagnosis.

Benefits of technology

It enables proactive sensing and scheduling of power equipment, generates targeted early warnings at the initial stage of anomalies, identifies complex anomalies and outputs interpretable fault reports. The system has adaptive optimization capabilities, improving the accuracy and efficiency of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121769719A_ABST
    Figure CN121769719A_ABST
Patent Text Reader

Abstract

The invention discloses a power equipment multi-mode monitoring and intelligent early warning method based on an inspection robot, and relates to the technical field of power equipment operation and maintenance and intelligent inspection. According to the method, vibration, sound and temperature parameters are collected in the inspection process, and an abnormal event trigger signal is generated when any parameter exceeds a self-adaptive threshold value or an early abnormal symptom is recognized by a lightweight model; the logic network is triggered to dynamically determine an optimal sensor group and sampling parameters based on the knowledge graph and reinforcement learning; after the state verification is completed, the inspection robot synchronously obtains time-aligned multi-modal high-frequency data streams through hardware, and identifies composite anomalies through a hierarchical attention fusion network; and the diagnosis engine performs fault initial judgment and causal chain reasoning in combination with the fault knowledge graph, and outputs an early warning report containing root causes and disposal suggestions. According to the method, the timeliness of anomaly identification, the reliability of multi-modal analysis and the interpretability of a diagnosis result are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power equipment operation and maintenance and intelligent inspection technology, and in particular to a method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots. Background Technology

[0002] Inspection robots are widely used in power substations and other similar locations for daily equipment inspections. They monitor parameters such as temperature, vibration, sound, and images through sensors to promptly identify potential faults. Traditional inspection methods often rely on data from fixed periods and single modes, such as periodically reading equipment temperature or taking visible light images at fixed points. This approach has two drawbacks: first, it collects a large amount of redundant data when no anomalies occur, potentially missing crucial details when brief anomalies do occur; second, single sensor information cannot comprehensively reflect the signs of complex faults, as different types of faults exhibit different characteristics across different modes, and relying on a single indicator can easily lead to missed or false alarms.

[0003] With the development of artificial intelligence and multi-sensor fusion technology, researchers have begun to explore intelligent inspection systems to improve the accuracy of anomaly detection through data fusion from multiple sensors. However, most existing methods lack event-triggered dynamic perception mechanisms, cannot flexibly adjust monitoring strategies according to real-time conditions, and are insufficiently adaptable to newly emerging fault modes. Therefore, it is necessary to provide an innovative method that can automatically activate multi-modal sensors for high-density data acquisition when equipment exhibits abnormal signs, and combine this with intelligent models for comprehensive diagnosis and early warning, thereby overcoming the shortcomings of existing technologies. Summary of the Invention

[0004] To address the above problems, this invention proposes a method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots. When abnormal signs appear in the operation of power equipment, multiple sensors on the inspection robot can be automatically triggered to collect multimodal data in a coordinated manner. The collected data is then fused and analyzed through a built-in intelligent model to identify complex abnormal patterns, locate fault locations, and issue early warnings in real time.

[0005] The present invention achieves the above objectives through the following technical solutions:

[0006] A method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots, the method comprising:

[0007] During routine inspections, the inspection robot monitors the vibration, sound, and temperature operating parameters of the power equipment in parallel. When any operating parameter exceeds the adaptive threshold and / or an early abnormal sign across modes is identified by a lightweight real-time discrimination model, an abnormal event trigger signal is generated for the target equipment.

[0008] The abnormal event trigger signal and the corresponding abnormality type information are input into a trigger logic network driven by a hybrid knowledge graph and reinforcement learning. The trigger logic network then dynamically reconstructs the optimal sensor set and sampling parameter strategy for this abnormality monitoring task.

[0009] After the optimal sensor set is ready, the inspection robot issues a collaborative acquisition command based on hardware synchronization to acquire time-aligned multimodal high-frequency data streams and constructs a multimodal fusion network based on a hierarchical attention mechanism to perform feature extraction and fusion analysis on the multimodal high-frequency data streams. When the fusion analysis result shows that there is a collaborative correlation between the abnormal patterns of multiple modes, it is determined to be a composite anomaly and fault symptom information is output.

[0010] The fault symptom information is input into a diagnostic engine that integrates data-driven and knowledge-driven approaches. The diagnostic engine performs a preliminary judgment on the fault type, traverses the fault knowledge graph, generates a diagnostic result, and outputs an interpretable early warning report containing the root cause and handling suggestions.

[0011] The end-to-end data and diagnostic results from this anomaly monitoring are used for adaptive updates, including updating the decision strategy of the triggering logic network, optimizing the adaptive threshold and lightweight real-time discrimination model, and expanding the fault knowledge graph.

[0012] As a preferred embodiment of the present invention, the process of generating an abnormal event trigger signal for the target device includes:

[0013] During the routine inspection of the inspection robot, robust statistics are calculated for vibration, sound and temperature operating parameters respectively. Based on the robust statistics, a corresponding adaptive threshold is constructed. The adaptive threshold sensitivity of the other operating parameter is dynamically adjusted by utilizing the short-term change trend of at least two of the vibration, sound and temperature operating parameters.

[0014] Based on a preset acquisition period, cross-modal joint features are extracted from the time-series data of vibration, sound, and temperature operating parameters. The cross-modal joint features are then input into a lightweight real-time discrimination model composed of convolutional units and gating units to generate anomaly scores. The anomaly scores include cross-modal inconsistency gain based on the peak time difference and short-time energy ratio of vibration, sound, and temperature signals. When the anomaly scores meet preset scoring conditions, it is considered that early abnormal signs have been identified.

[0015] When any of the operating parameters among vibration, sound, and temperature meets the adaptive threshold condition, or when early abnormal signs are identified, the inspection robot generates an abnormal event trigger signal, and marks the trigger level code and the corresponding abnormal type information according to the trigger path.

[0016] As a preferred embodiment of the present invention, the triggering logic network dynamically reconstructs the optimal sensor set for this anomaly monitoring task, including:

[0017] In the trigger logic network, a rule base and machine learning model are pre-set. When the inspection robot receives an abnormal event trigger signal and abnormality type information, it retrieves the rule mapping corresponding to the abnormality type information from the rule base.

[0018] If there is a mapping entry in the rule base that matches the anomaly type information, then the set of sensor groups associated with the mapping entry output by the rule base will be used as the optimal set of sensor groups.

[0019] If there is no mapping entry in the rule base that matches the anomaly type information, the inspection robot constructs a contextual knowledge graph consisting of target device nodes, operating parameter nodes, and historical event nodes based on the anomaly event trigger signal, anomaly type information, and contextual data generated during routine inspections. The robot extracts graph embedding representations from the contextual knowledge graph and inputs the graph embedding representations and anomaly type information into a machine learning model, which then generates a set of candidate sensor groups.

[0020] Constructing an action space based on the candidate sensor set The actions Each sensor group in the candidate sensor set corresponds one-to-one with the action space. Input the reinforcement learning decision model into the state vector describing this anomaly detection task. Under constraints, the set of actions is determined by solving the following equation:

[0021] ;

[0022] In the formula, For action In state The value of the action below;

[0023] Based on the node relationships in the context knowledge graph, actions consistent with the anomaly type information are selected from the action set to form an optimal sensor group set for solving the sampling parameters.

[0024] As a preferred embodiment of the present invention, after determining the optimal sensor set for this anomaly monitoring task, the process of solving the sampling parameter strategy for the optimal sensor set includes:

[0025] Based on the sampling capabilities and adjustable parameter ranges of each sensor in the optimal sensor set, a candidate set of sampling parameters, including sampling frequency, sampling duration, and trigger delay, is constructed. This candidate set of sampling parameters is then used to construct the sampling parameter action space. ,in The sampling parameters are mapped one-to-one with the combinations of sampling frequency, sampling duration, and trigger delay in the candidate set of sampling parameters, thus defining the parameter action space. Input reinforcement learning decision model, and define the state vector constraints that describe this anomaly detection task. The sampling parameters are obtained by solving the following formula:

[0026] ;

[0027] In the formula, Indicates sampling parameter action In state The value of the action below; This represents the information gain calculated based on the short-time data distribution corresponding to the sampling parameter actions; These are weighting coefficients;

[0028] Energy consumption, data bandwidth, and inference calculation metrics are calculated for the sampling parameter set corresponding to the obtained sampling parameter actions. Sampling parameter actions that do not meet resource constraints are eliminated. Based on the node relationships in the knowledge graph, a sampling parameter set that is consistent with the anomaly type information is selected from the sampling parameter actions that meet resource constraints. The sampling parameter set is then used as the sampling parameter strategy for this anomaly monitoring task.

[0029] As a preferred embodiment of the present invention, the optimal sensor set is ready, specifically including: while the inspection robot moves to the optimal observation pose of the target device, performing parallel state verification and initialization on the optimal sensor set, adopting a resource preemption mechanism to ensure that computing, communication and sensing resources are given priority to this anomaly monitoring task, and starting an online self-calibration process for sensors that are not ready, until all sensors generate ready marks.

[0030] As a preferred embodiment of the present invention, the process of performing parallel state verification and initialization on the optimal sensor set during the period when the inspection robot moves to the optimal observation pose of the target device includes:

[0031] During the movement, a state description vector containing working status identifiers, synchronization clock domain identifiers, and communication link identifiers is generated for the vibration sensor, sound sensor, and temperature sensor of the optimal sensor set. This vector is then combined with the robot's body posture, installation posture matrix, and observation vectors to construct a posture consistency matrix. To form a multimodal consistent state vector ;

[0032] Based on the multimodal consensus state vector Parallel state verification is performed on the optimal sensor set, including initializing the operating mode, detecting power supply status, communication link status, synchronization relationship, and attitude consistency matrix. The observation pose constraints are defined; the initialization sequence is divided according to the synchronous clock domain, an initialization channel is allocated to each sensor, and an initialization control command is issued;

[0033] Construct a multimodal link conflict detection graph composed of bandwidth occupancy, link occupancy, and interference weights. In the inspection robot's computing unit, a task-level resource preemption scheduler is activated to allocate computing, communication, and sensing resources to the resource queues corresponding to the anomaly monitoring tasks, and based on the multimodal link conflict detection map... Issue a resource preemption command to the optimal sensor group set;

[0034] For sensors that fail the condition verification, an online self-calibration process is initiated, including offset calibration, noise baseline reestimation, and synchronization alignment based on the synchronization pulse sequence, and the hardware clock offset residual is calculated. The sensor generates a ready flag by updating the local timing reference through an iterative correction loop to complete self-calibration.

[0035] Before all sensors generate a readiness flag, the inspection robot injects a pre-sampled pulse sequence into the optimal sensor set. The consistency of the trigger delay and response time of each sensor is verified. After all sensors have completed the pre-sampling pulse sequence verification, the inspection robot issues a collaborative acquisition command based on the optimal sensor set.

[0036] As a preferred embodiment of the present invention, the inspection robot issues a collaborative acquisition command based on hardware synchronization to acquire a time-strictly aligned multimodal high-frequency data stream, specifically including:

[0037] A unified hardware synchronization pulse sequence is generated in the inspection robot control module. The hardware synchronization pulse sequence includes a synchronization clock domain identifier, a trigger time identifier, and a sampling window identifier, and is accompanied by a high-frequency jitter suppression parameter, which is used to establish trigger edge consistency among the vibration sensor, sound sensor, and temperature sensor in the optimal sensor group set.

[0038] Based on the sampling parameter strategy determined by the trigger logic network, the inspection robot generates local sampling configurations for vibration sensors, sound sensors, and temperature sensors in the optimal sensor set. The local sampling configurations include the sampling frequency, sampling window length, and sampling start offset matrix based on static offset of the corresponding sampling parameter strategy, and enter high-frequency sampling mode after the collaborative acquisition command is activated.

[0039] In high-frequency sampling mode, the optimal sensor set samples vibration, sound, and temperature operating parameters to form high-frequency vibration data, high-frequency sound data, and high-frequency temperature data. A unified cross-modal timestamp is added to the sampled data based on a unified hardware synchronization pulse sequence, and cross-sensor time alignment is completed based on the synchronization clock domain. The inspection robot computing unit constructs a time alignment verification stack to verify the alignment residuals of the high-frequency vibration data, high-frequency sound data, and high-frequency temperature data frame by frame. After the verification is passed, the data is transmitted to the inspection robot computing unit in a unified data channel encoding format.

[0040] As a preferred embodiment of the present invention, the construction of a multimodal fusion network based on a hierarchical attention mechanism for feature extraction and fusion analysis of the multimodal high-frequency data stream includes:

[0041] In the computational unit of the inspection robot, a multimodal fusion network with a hierarchical attention mechanism consisting of a local attention layer and a global attention layer is constructed. The local attention layer performs feature extraction on high-frequency vibration data, high-frequency sound data and high-frequency temperature data respectively, and introduces time consistency weights generated by the time alignment verification stack to form local vibration features, local sound features and local temperature features.

[0042] The vibration local features, sound local features, and temperature local features are input into the global attention layer. A cross-modal attention weight matrix is ​​constructed in the global attention layer, and a modal weight bias vector composed of anomaly type information is added to the cross-modal attention weight matrix to form a fused feature representation.

[0043] The inspection robot's computing unit performs abnormal pattern recognition based on the fused feature representation. When there is a synergistic correlation between the vibration abnormal pattern, sound abnormal pattern and temperature abnormal pattern in the fused feature representation, fault symptom information is generated.

[0044] As a preferred embodiment of the present invention, the fault symptom information is input into a diagnostic engine that integrates data-driven and knowledge-driven approaches. The diagnostic engine performs a preliminary judgment on the fault type and traverses the causal relationship chain in the fault knowledge graph, specifically including:

[0045] Perform feature transformation on the fault symptom information to generate a fault feature vector, and generate a preliminary fault type judgment result based on the fault feature vector;

[0046] Based on the initial fault type judgment result, locate the starting node associated with the initial fault type judgment result in the fault knowledge graph, extract the node sequence containing the target device node, symptom node, event node and handling node from the starting node, and construct a causal relationship chain representation using the node sequence;

[0047] The fault feature vector and the causal relationship chain representation are used to perform relationship matching and path traversal. During the relationship matching process, a relationship function containing node embedding vectors is used. , for node pairs The relation type is encoded, where the relation function... The expression is: ;

[0048] Calculate path weights while traversing the path set. : ;

[0049] In the formula, For relational encoding matrix; , These are nodes in the causal relationship chain representation. With nodes Node embedding vectors; For activation functions; For any path in the causal relationship chain representation. A set of paths; This is the compatibility difference component between the fault feature vector and the node embedding vector in the path; These are the weighting coefficients for the compatibility difference components;

[0050] From the path set The path with the largest path weight is selected, and a diagnostic result vector is generated from the path. The diagnostic result vector includes a fault type identifier, a root cause identifier, and a treatment suggestion identifier.

[0051] An interpretable early warning report is generated based on the diagnostic result vector and output to the inspection robot control module.

[0052] As a preferred embodiment of the present invention, the process of using the full-link data and diagnostic results of this anomaly monitoring for adaptive updates includes:

[0053] A full-link feature set is constructed from the high-frequency vibration data, high-frequency sound data, and high-frequency temperature data collected during this anomaly monitoring, along with time alignment information. This full-link feature set, along with the diagnostic result vector, is input into the policy update module of the trigger logic network. In the policy update module, the cost function is applied... Calculate the cost of actions in the action space:

[0054] ;

[0055] In the formula, , , Respectively represent actions The corresponding calculation output, bandwidth output, and sensor output; , , These are weighting coefficients;

[0056] The action selection strategy of the triggering logic network is updated based on the principle of minimizing the cost function;

[0057] Based on the full-link feature set, the statistics of vibration, sound and temperature operating parameters are calculated, a threshold correction vector is constructed, the adaptive threshold is recalibrated through the threshold correction vector, and the recalibrated adaptive threshold is written into the threshold management table.

[0058] An incremental training set was constructed from the data samples input to the lightweight real-time discriminant model during this anomaly monitoring process. Based on the incremental training set, a parameter update step was performed to update the convolutional unit parameters and gating unit parameters of the lightweight real-time discriminant model.

[0059] Based on the fault type identifier, root cause identifier, and treatment suggestion identifier in the diagnostic result vector, an incremental graph structure is constructed. New nodes and new relationship types are added to the fault knowledge graph. Local recalculation of the node embedding vector is performed based on the incremental graph structure, and the recalculation results are written into the node embedding matrix of the fault knowledge graph.

[0060] The beneficial effects of this invention are as follows: By collecting vibration, sound, and temperature operating parameters in parallel during routine inspections and combining them with a lightweight real-time discriminant model to identify early cross-modal anomaly features, targeted anomaly event trigger signals can be generated at the initial stage of anomalies; by constructing a trigger logic network that integrates knowledge graphs and reinforcement learning, the optimal sensor set and sampling parameter strategy can be dynamically reconstructed according to different anomaly types, achieving proactive perception scheduling that matches equipment fault characteristics; after the optimal sensor set is ready, a unified hardware synchronization pulse sequence is used to perform collaborative acquisition to ensure that multimodal high-frequency data are strictly aligned under a unified time reference, providing a basis for subsequent fusion analysis. The system has a solid data foundation. By constructing a multimodal fusion network based on a hierarchical attention mechanism, it can identify composite anomalies with collaborative pattern characteristics under the combined effect of local features and cross-modal association weights. Furthermore, fault symptom information is input into a diagnostic engine that integrates data-driven and knowledge-driven approaches. First, a preliminary judgment of the fault type is made, and then an interpretable early warning report containing the root cause and handling suggestions is output based on the causal chain structure of the fault knowledge graph. At the same time, the trigger logic network, adaptive threshold, lightweight real-time discrimination model, and fault knowledge graph are adaptively updated using the full-link data and diagnostic results generated from this anomaly monitoring, enabling the system to have the ability to continuously optimize in subsequent inspection tasks. Attached Figure Description

[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0062] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.

[0064] like Figure 1 As shown, this is an embodiment of the present invention, which provides a method for multimodal monitoring and intelligent early warning of power equipment based on an inspection robot, including the following steps:

[0065] S1: During routine inspections, the inspection robot monitors the vibration, sound, and temperature operating parameters of the power equipment in parallel. When any operating parameter exceeds the adaptive threshold and / or an early abnormal sign across modes is identified through a lightweight real-time discrimination model, an abnormal event trigger signal is generated for the target equipment.

[0066] In this embodiment, vibration is collected using a triaxial accelerometer, sound is collected using an industrial-grade MEMS microphone, and temperature is obtained using infrared thermal imaging or a point-measurement temperature sensor. The robot monitors these three types of parameters simultaneously and compares the real-time sampled values ​​with an adaptive threshold constructed based on statistical measures. To avoid false triggering caused by transient noise and external interference, this embodiment uses robust statistical measures to construct the adaptive threshold. Robust statistical measures include the median, median absolute deviation (MAD), and distribution offset. For example, the adaptive threshold can be generated in the following form: threshold = median + coefficient × MAD, which is dynamically updated with real-time sampled data. To improve cross-modal correlation capabilities, this embodiment introduces short-term change trends between the three types of operating parameters as supplementary information. For example, if vibration shows an upward trend, the system increases the temperature threshold sensitivity; if the sound signal shows continuous spikes, the system reduces the tolerance of the vibration threshold; when the temperature changes rapidly, the vibration threshold is dynamically adjusted in conjunction with the sound trend. This multimodal linkage method ensures that the threshold values ​​of each parameter can reflect the overall operating status of the equipment.

[0067] Within a preset acquisition period, the system constructs cross-modal joint features from vibration, sound, and temperature time-series data. These features include, but are not limited to, peak time difference, energy ratio, temporal spectral envelope, rate of rapid change, and correlation coefficient within a time window. The constructed joint features serve as input to a lightweight real-time discrimination model. The model's convolutional structure extracts local patterns, and its gating structure judges short-term dynamic changes, thereby generating anomaly scores. Anomaly scores include cross-modal inconsistency gains, such as the difference between vibration peak and temperature peak times, and the short-term ratio of sound energy to vibration energy. The greater these inconsistencies, the higher the anomaly score. When the anomaly score reaches a preset threshold, the system considers early abnormal signs to have been identified. When any operating parameter meets the adaptive threshold condition, or the anomaly score meets the model's discrimination condition, the system generates an anomaly event trigger signal.

[0068] Simultaneously, the system labels trigger level codes based on trigger path information to distinguish different trigger sources, such as: threshold-only triggering, model-only triggering, and both threshold and model triggering. The trigger level codes, along with anomaly type information, are input into the trigger logic network for subsequent generation of the optimal sensor group and sampling strategy.

[0069] S2: Input the abnormal event trigger signal and the corresponding abnormality type information into the trigger logic network driven by a hybrid knowledge graph and reinforcement learning. The trigger logic network dynamically reconstructs the optimal sensor set and sampling parameter strategy for this abnormality monitoring task.

[0070] This embodiment, combined with the real-time monitoring process of power equipment by the inspection robot, illustrates how the triggering logic network, based on knowledge graphs and reinforcement learning mechanisms, determines the optimal sensor set for this anomaly monitoring task.

[0071] Before the inspection robot was put into use, the system developers pre-built a rule base and machine learning model based on the operating patterns of typical power equipment. The rule base was used to store the correspondence between anomaly types and sensor combinations, for example:

[0072] If the anomaly type is "mechanical structure anomaly", the rule base can be set to recommend the use of vibration sensors and sound sensors;

[0073] If the anomaly type is "overheating anomaly", the rule base can be configured to recommend the use of temperature sensors and sound sensors.

[0074] Machine learning models are primarily used for inference when the rule base cannot cover all anomaly types. This model, trained on historical monitoring data, can determine the applicability of different sensor combinations based on node embedding vectors.

[0075] After receiving an anomaly trigger signal, the inspection robot retrieves the anomaly type information contained in the signal. If the rule base contains a matching mapping entry for that anomaly type, the preset sensor combination in the rule base can be directly used as the optimal sensor set for this anomaly monitoring task. When the rule base does not provide a matching mapping, the system enters the knowledge graph and machine learning inference stage. Using the data accumulated by the inspection robot in routine inspection tasks, a knowledge graph with multiple nodes is constructed. The nodes of the graph include:

[0076] Target device node: Represents the attributes of power equipment, such as equipment type, equipment structure, etc.

[0077] Operating parameter node: Represents the statistical characteristics of monitored values ​​such as vibration, sound, and temperature;

[0078] Historical event nodes: Records abnormal events that have occurred in the equipment and the corresponding handling measures.

[0079] The relationships between these nodes are predefined based on the co-occurrence of monitoring data, temporal order, and engineering logic.

[0080] This embodiment encodes the knowledge graph using a graph neural network. Each node is mapped to a numerical vector representing the contextual relationship between that node and its associated nodes; these vectors are called "node embedding vectors." The system inputs the anomaly type information and the node embedding vectors into a machine learning model to generate a set of candidate sensor groups. After obtaining the candidate sensor group set, this embodiment uses a reinforcement learning decision-making strategy to select the optimal sensor group, treating each combination in the candidate sensor group set as an action, thus forming an action space. The actions Each sensor corresponds one-to-one with any sensor group in the candidate sensor set. The reinforcement learning policy is based on the action-value function. Choose the current state vector The most matching action, in the formula, For action In state The value of the action under the given conditions. State vector. It contains the following information: abnormal event trigger signal, abnormal type information, and node embedding vector obtained from the knowledge graph.

[0081] The machine learning model outputs the optimal action based on the action value function, and the sensor combination corresponding to this action is the optimal sensor set.

[0082] After determining the optimal sensor set, parameter ranges are predefined based on the specific hardware capabilities of each sensor in the optimal sensor set, for example:

[0083] Selectable sampling frequency range: hundreds of Hz to tens of kHz;

[0084] Sampling duration selectable range: tens of milliseconds to several seconds;

[0085] Trigger delay compensation range: several milliseconds.

[0086] By combining these parameters, a candidate set of sampling parameters is constructed. Each combination in the candidate set is regarded as a sampling parameter action, thus forming the sampling parameter action space. .

[0087] To better suit the sampling parameters for the current anomaly detection task, this embodiment employs a reinforcement learning strategy, selecting the sampling parameter action using the following formula: In the formula, Indicates sampling parameter action In state The value of the action below; This represents the information gain calculated based on the short-time data distribution corresponding to the sampling parameter actions; The weighting coefficients are used to adjust the influence of the information gain; this formula can guide the system to select the optimal combination of sampling frequency, sampling duration and trigger delay in the current state.

[0088] The selected sampling parameter actions are validated, including sampling energy consumption calculation, bandwidth usage analysis during the sampling process, and inference computing resource requirement analysis. If a sampling parameter action cannot meet resource constraints, the system excludes it from the set of valid actions.

[0089] By leveraging the node relationships within the knowledge graph, the system can determine the sampling feature requirements associated with the current anomaly type, for example:

[0090] Some anomaly types require higher sampling frequencies;

[0091] Some anomaly types require longer sampling times to capture slow-changing phenomena.

[0092] Finally, from the sampling parameter actions that meet the resource constraints, the action that is consistent with the anomaly type information is selected as the sampling parameter strategy for this anomaly monitoring task.

[0093] S3: After the optimal sensor set is ready, the inspection robot issues a collaborative acquisition command based on hardware synchronization to acquire multimodal high-frequency data streams with strict time alignment, and constructs a multimodal fusion network based on a hierarchical attention mechanism to perform feature extraction and fusion analysis on the multimodal high-frequency data streams. When the result of the fusion analysis shows that there is a collaborative correlation between the abnormal patterns of multiple modes, it is determined to be a composite anomaly and fault symptom information is output.

[0094] Furthermore, after triggering the logic network to output the optimal sensor set, the inspection robot marks this set as "pending initialization". As the inspection robot moves towards the target device, it triggers a readiness detection program according to the task scheduling system, causing the optimal sensor set to enter the initialization process.

[0095] As the inspection robot approaches the target equipment, it simultaneously performs status checks on all types of sensors (vibration sensors, sound sensors, and temperature sensors) in the optimal sensor set. The checks include:

[0096] Working mode status detection: The system checks whether each sensor is in a controllable working mode, including whether it has been powered on, whether it responds to the query command issued by the robot, and whether it can be switched to monitoring mode;

[0097] Communication link status detection: The inspection robot establishes a two-way communication handshake with each sensor to confirm whether the data link is available, whether the communication rate meets the subsequent high-frequency sampling requirements, and whether there is any data loss in the link;

[0098] Communication bandwidth and resource usage detection: Detect the current data bandwidth requirement of each sensor and record the proportion of that bandwidth used in the robot's internal communication network for subsequent resource preemption mechanism scheduling;

[0099] Energy state detection: The robot checks the power supply stability of each sensor, including whether the power supply meets the prerequisite for high-frequency sampling.

[0100] All of the above tests are performed in parallel to avoid delays in responding to abnormal events.

[0101] In this embodiment, the optimal sensor set can include any combination of the above three types of sensors. The inspection robot moves to the optimal observation pose of the target power equipment according to the planned path. During the movement, the robot simultaneously generates a state description vector for each sensor in the optimal sensor set, containing the following information:

[0102] The operating status indicator is used to indicate whether the sensor is in a sampleable state;

[0103] Communication link identifier, used to indicate whether the sensor can maintain stable communication with the robot body;

[0104] Communication bandwidth identifier, used for subsequent resource allocation;

[0105] Energy status indicator, used to indicate the power supply status of the sensor.

[0106] During robot movement, its body posture and gimbal orientation will change. The robot simultaneously collects its own posture parameters, such as three-axis posture angles, gimbal position, and mounting matrix, and establishes a posture consistency matrix. This is used to map the robot's current posture to the observation vectors of each sensor.

[0107] By integrating the above state description vectors with the posture consistency matrix, the robot constructs a multimodal consistent state vector. It is used to describe the overall situation of sensor status, communication status and pose relationship.

[0108] To ensure that the optimal sensor set obtains sufficient resources in subsequent sampling, a multimodal link collision detection graph is constructed. This graph records the data transmission link occupancy, available or limited bandwidth, and potential conflict relationships for each sensor. The robot's internal task scheduling system uses this detection graph to perform resource preemption, prioritizing the allocation of computing, communication, and sensing resources to the optimal sensor group set, thereby ensuring the smooth execution of subsequent sampling tasks.

[0109] For sensors that fail the condition verification, an online self-calibration process is initiated, including offset calibration (correcting sensor zero-point offset), noise baseline reestimation (re-estimating the noise baseline based on a static window), and synchronization alignment based on a synchronization pulse sequence, and calculation of hardware clock offset residuals. The sensor generates a ready flag by updating the local timing reference through an iterative correction loop to complete self-calibration.

[0110] Before all sensors generate a readiness flag, the inspection robot injects a pre-sampled pulse sequence into the optimal sensor set. The consistency of trigger delay and response time of each sensor is checked (whether the trigger delay is consistent, whether the reported sampling timestamp is consistent, whether the pre-sampling response meets the synchronization requirements, etc.). After all sensors have completed the pre-sampling pulse sequence verification, the inspection robot issues a collaborative acquisition command based on the optimal sensor group set.

[0111] The inspection robot issues collaborative data acquisition commands based on hardware synchronization to acquire multimodal, high-frequency data streams with strictly time-aligned data, specifically including:

[0112] A unified hardware synchronization pulse sequence is generated in the inspection robot control module, including a synchronization clock domain identifier, a trigger time identifier, and a sampling window identifier, and is accompanied by high-frequency jitter suppression parameters, which are used to establish trigger edge consistency among the vibration sensor, sound sensor, and temperature sensor in the optimal sensor group set.

[0113] Based on the sampling parameter strategy determined by the trigger logic network, the inspection robot generates local sampling configurations for vibration sensors, sound sensors and temperature sensors in the optimal sensor set, including the sampling frequency, sampling window length and sampling start offset matrix based on the corresponding sampling parameter strategy, and enters high-frequency sampling mode after the collaborative acquisition command is activated.

[0114] In high-frequency sampling mode, the optimal sensor group samples vibration, sound, and temperature operating parameters to form high-frequency vibration data, high-frequency sound data, and high-frequency temperature data. The sampled data is then combined with a unified hardware synchronization pulse sequence to add a cross-modal unified timestamp and complete cross-sensor time alignment based on the synchronization clock domain. The inspection robot computing unit constructs a time alignment verification stack to verify the alignment residuals of the high-frequency vibration data, high-frequency sound data, and high-frequency temperature data frame by frame. After the verification is passed, the data is transmitted to the inspection robot computing unit in a unified data channel encoding format.

[0115] In one embodiment, a multimodal fusion network based on a hierarchical attention mechanism is constructed to perform feature extraction and fusion analysis on multimodal high-frequency data streams, including:

[0116] In the computational unit of the inspection robot, a multimodal fusion network with a hierarchical attention mechanism consisting of a local attention layer and a global attention layer is constructed. The local attention layer performs feature extraction on high-frequency vibration data, high-frequency sound data and high-frequency temperature data respectively, and introduces time consistency weights generated by the time alignment verification stack to form local vibration features, local sound features and local temperature features.

[0117] The vibration local features, sound local features, and temperature local features are input into the global attention layer. A cross-modal attention weight matrix is ​​constructed in the global attention layer, and a modal weight bias vector composed of anomaly type information is added to the cross-modal attention weight matrix to form a fused feature representation.

[0118] The inspection robot performs anomaly pattern recognition based on fused feature representations. When anomalies are observed in the vibration, sound, and temperature modes, the system generates fault symptom information and submits it to the subsequent diagnostic engine.

[0119] S4: Input the fault symptom information into the diagnostic engine that integrates data-driven and knowledge-driven approaches. The diagnostic engine will make an initial judgment on the fault type, traverse the fault knowledge graph, generate diagnostic results, and output an interpretable early warning report containing the root cause and handling suggestions.

[0120] After obtaining fault symptom information, the inspection robot's computing unit first performs feature transformation and quantization processing on the fault symptom information. In specific implementation, the system maps the fault symptom information of different modalities (such as changes in vibration peak value, energy distribution of sound feature subbands, and temperature texture changes) into a multi-dimensional numerical vector according to a preset mapping rule, forming a fault feature vector. This fault feature vector is then fed as input into a data-driven fault type preliminary judgment module. This module can be a lightweight classifier, an attention-based pattern recognition algorithm, or a structured decision tree model, used to output preliminary fault type judgment results. Examples include: "Preliminary judgment of mechanical loosening faults," "Preliminary judgment of early bearing wear," and "Preliminary judgment of electrical faults due to poor contact." The generated preliminary fault type judgment results are used to subsequently locate the starting node in the knowledge-driven reasoning graph.

[0121] The fault knowledge graph consists of target equipment nodes, symptom nodes, event nodes, and processing nodes. The nodes are associated with each other based on the operating mechanism of the power equipment. For example, there is a directional relationship between fault type nodes and symptom nodes, a causal relationship between symptom nodes and event nodes, and a recommendation relationship between event nodes and processing nodes.

[0122] Based on the initial judgment of the fault type, the diagnostic engine locates the corresponding starting node in the knowledge graph. Then, according to the topology defined by the graph, it extracts the target device node, symptom node, event node, and treatment node related to the starting node to form one or more node sequences.

[0123] The system constructs a causal relationship chain representation based on these node sequences, which serves as input for performing relationship matching and path adaptation.

[0124] For any pair of nodes in the causal chain representation Call the relational function Pair of nodes The relation type is encoded, where the relation function... The expression is: ;

[0125] Calculate path weights while traversing the path set. : ;

[0126] In the formula, For relational encoding matrix; , These are nodes in the causal relationship chain representation. With nodes Node embedding vectors; For activation functions; For any path in the causal relationship chain representation. A set of paths; This is the compatibility difference component between the fault feature vector and the node embedding vector in the path; These are the weighting coefficients for the compatibility difference components;

[0127] From the path set The path with the highest path weight is selected, and a diagnostic result vector is generated from the path. The diagnostic result vector includes a fault type identifier, a root cause identifier, and a treatment suggestion identifier.

[0128] An interpretable early warning report is generated based on the diagnostic result vector and output to the inspection robot control module.

[0129] S5: Use the full-link data and diagnostic results from this anomaly monitoring for adaptive updates, including updating the decision-making strategy of the trigger logic network, optimizing the adaptive threshold and lightweight real-time discrimination model, and expanding the fault knowledge graph.

[0130] The inspection robot unifies the various types of data generated during this anomaly monitoring process into a full-link feature set, including high-frequency vibration data, high-frequency sound data, high-frequency temperature data, time alignment information of the three types of data, fault symptom information after multimodal fusion, and diagnostic result vectors obtained from causal relationship chain reasoning. This information is then used to update the strategy module in the triggering logic network.

[0131] The system inputs the full-link feature set and diagnostic result vector into the policy update module to update the action selection policy. The policy update module first updates the action selection policy for each action in the action space based on the action space of the triggering logic network. Calculate the cost of the action. Cost function. The definition is as follows:

[0132] ;

[0133] In the formula, , , Respectively represent actions The corresponding calculation output, bandwidth output, and sensor output; , , These are weighting coefficients;

[0134] Based on the cost function calculation, the diagnostic engine identifies actions with lower costs in the action space and uses these as the basis for improving the decision-making strategy of the triggering logic network. In specific implementations, the system can update the action value table, priority table, or action selection strategy parameters based on the cost ranking results.

[0135] Statistical values ​​of vibration, sound, and temperature operating parameters are extracted from the end-to-end feature set to construct a threshold correction vector. This vector may include corrections for peak change, duration, fluctuation amplitude, and contextual information. This vector is used to recalibrate the adaptive thresholds. The system updates the adaptive thresholds for vibration, sound, and temperature according to the recalibration rules and writes the updated thresholds to the threshold management table for use in the next anomaly monitoring task.

[0136] To enable the lightweight real-time discrimination model to adapt to the updated parameter distribution, the system collects the data samples input to the discrimination model during this anomaly monitoring process into an incremental training set. The incremental training set includes time-aligned multimodal data, fused fault features, and the current diagnostic result vector as label information. Based on the incremental training set, a model parameter update step is performed, incrementally training the convolutional unit parameters and gating unit parameters in the lightweight real-time discrimination model to adapt the model to the new feature distribution.

[0137] Based on the fault type identifier, root cause identifier, and treatment suggestion identifier in the diagnostic result vector, an incremental graph structure is constructed, including new nodes and new relationship types. Subsequently, the following are added to the fault knowledge graph based on the incremental graph structure:

[0138] Add new nodes: such as new symptom nodes, target device nodes, or event nodes;

[0139] New relationship types: such as new causal relationships, referential relationships, or recommendation relationships.

[0140] To maintain the uniformity and consistency of graph reasoning, the system performs local node embedding recalculation on newly added nodes and their associated existing nodes, and writes the calculated node embedding vectors into the graph node embedding matrix, so that the knowledge graph can continue to support causal relationship chain reasoning under the new structure.

[0141] In summary, the method of this invention, by combining event-driven approaches, multi-sensor collaboration, and intelligent diagnostics, can significantly improve the intelligence level of power equipment inspection systems. While ensuring real-time performance, it enhances the accuracy of anomaly detection and fault early warning, which is of great significance for reducing power equipment downtime and ensuring the safe operation of the power grid.

[0142] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots, characterized in that, The method includes: During routine inspections, the inspection robot monitors the vibration, sound, and temperature operating parameters of the power equipment in parallel. When any operating parameter exceeds the adaptive threshold and / or an early abnormal sign across modes is identified by a lightweight real-time discrimination model, an abnormal event trigger signal is generated for the target equipment. The abnormal event trigger signal and the corresponding abnormality type information are input into a trigger logic network driven by a hybrid knowledge graph and reinforcement learning. The trigger logic network then dynamically reconstructs the optimal sensor set and sampling parameter strategy for this abnormality monitoring task. After the optimal sensor set is ready, the inspection robot issues a collaborative acquisition command based on hardware synchronization to acquire time-aligned multimodal high-frequency data streams and constructs a multimodal fusion network based on a hierarchical attention mechanism to perform feature extraction and fusion analysis on the multimodal high-frequency data streams. When the fusion analysis result shows that there is a collaborative correlation between the abnormal patterns of multiple modes, it is determined to be a composite anomaly and fault symptom information is output. The fault symptom information is input into a diagnostic engine that integrates data-driven and knowledge-driven approaches. The diagnostic engine performs a preliminary judgment on the fault type, traverses the fault knowledge graph, generates a diagnostic result, and outputs an interpretable early warning report containing the root cause and handling suggestions. The end-to-end data and diagnostic results from this anomaly monitoring are used for adaptive updates, including updating the decision strategy of the triggering logic network, optimizing the adaptive threshold and lightweight real-time discrimination model, and expanding the fault knowledge graph.

2. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 1, characterized in that, The process of generating an abnormal event trigger signal for the target device includes: During the routine inspection of the inspection robot, robust statistics are calculated for vibration, sound and temperature operating parameters respectively. Based on the robust statistics, a corresponding adaptive threshold is constructed. The adaptive threshold sensitivity of the other operating parameter is dynamically adjusted by utilizing the short-term change trend of at least two of the vibration, sound and temperature operating parameters. Based on a preset acquisition period, cross-modal joint features are extracted from the time-series data of vibration, sound, and temperature operating parameters. The cross-modal joint features are then input into a lightweight real-time discrimination model composed of convolutional units and gating units to generate anomaly scores. The anomaly scores include cross-modal inconsistency gain based on the peak time difference and short-time energy ratio of vibration, sound, and temperature signals. When the anomaly scores meet preset scoring conditions, it is considered that early abnormal signs have been identified. When any of the operating parameters among vibration, sound, and temperature meets the adaptive threshold condition, or when early abnormal signs are identified, the inspection robot generates an abnormal event trigger signal, and marks the trigger level code and the corresponding abnormal type information according to the trigger path.

3. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 2, characterized in that, The triggering logic network dynamically reconstructs the optimal sensor set for this anomaly monitoring task, including: In the trigger logic network, a rule base and machine learning model are pre-set. When the inspection robot receives an abnormal event trigger signal and abnormality type information, it retrieves the rule mapping corresponding to the abnormality type information from the rule base. If there is a mapping entry in the rule base that matches the anomaly type information, then the set of sensor groups associated with the mapping entry output by the rule base will be used as the optimal set of sensor groups. If there is no mapping entry in the rule base that matches the anomaly type information, the inspection robot constructs a contextual knowledge graph consisting of target device nodes, operating parameter nodes, and historical event nodes based on the anomaly event trigger signal, anomaly type information, and contextual data generated during routine inspections. The robot extracts graph embedding representations from the contextual knowledge graph and inputs the graph embedding representations and anomaly type information into a machine learning model, which then generates a set of candidate sensor groups. Constructing an action space based on the candidate sensor set The actions Each sensor group in the candidate sensor set corresponds one-to-one with the action space. Input the reinforcement learning decision model into the state vector describing this anomaly detection task. Under constraints, the set of actions is determined by solving the following equation: ; In the formula, For action In state The value of the action below; Based on the node relationships in the context knowledge graph, actions consistent with the anomaly type information are selected from the action set to form an optimal sensor group set for solving the sampling parameters.

4. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 3, characterized in that, After determining the optimal set of sensors for this anomaly monitoring task, the process of solving the sampling parameter strategy for the optimal set of sensors includes: Based on the sampling capabilities and adjustable parameter ranges of each sensor in the optimal sensor set, a candidate set of sampling parameters, including sampling frequency, sampling duration, and trigger delay, is constructed. This candidate set of sampling parameters is then used to construct the sampling parameter action space. ,in The sampling parameters are mapped one-to-one with the combinations of sampling frequency, sampling duration, and trigger delay in the candidate set of sampling parameters, thus defining the parameter action space. Input reinforcement learning decision model, and define the state vector constraints that describe this anomaly detection task. The sampling parameters are obtained by solving the following formula: ; In the formula, Indicates sampling parameter action In state The value of the action below; This represents the information gain calculated based on the short-time data distribution corresponding to the sampling parameter actions; These are weighting coefficients; Energy consumption, data bandwidth, and inference calculation metrics are calculated for the sampling parameter set corresponding to the obtained sampling parameter actions. Sampling parameter actions that do not meet resource constraints are eliminated. Based on the node relationships in the knowledge graph, a sampling parameter set that is consistent with the anomaly type information is selected from the sampling parameter actions that meet resource constraints. The sampling parameter set is then used as the sampling parameter strategy for this anomaly monitoring task.

5. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 4, characterized in that, The process of making the optimal sensor set ready includes: while the inspection robot moves to the optimal observation pose of the target device, parallel state verification and initialization are performed on the optimal sensor set; a resource preemption mechanism is adopted to ensure that computing, communication and sensing resources are given priority to this anomaly monitoring task; and an online self-calibration process is started for the sensors that are not ready until all sensors generate ready signs.

6. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 5, characterized in that, During the process of the inspection robot moving to the optimal observation pose of the target device, the parallel state verification and initialization of the optimal sensor set includes: During the movement, a state description vector containing working status identifiers, synchronization clock domain identifiers, and communication link identifiers is generated for the vibration sensor, sound sensor, and temperature sensor of the optimal sensor set. This vector is then combined with the robot's body posture, installation posture matrix, and observation vectors to construct a posture consistency matrix. To form a multimodal consistent state vector ; Based on the multimodal consensus state vector Parallel state verification is performed on the optimal sensor set, including initializing the operating mode, detecting power supply status, communication link status, synchronization relationship, and attitude consistency matrix. The observation pose constraints; the initialization sequence is divided according to the synchronous clock domain, an initialization channel is allocated to each sensor and an initialization control command is issued; Construct a multimodal link conflict detection graph composed of bandwidth occupancy, link occupancy, and interference weights. In the inspection robot's computing unit, a task-level resource preemption scheduler is activated to allocate computing, communication, and sensing resources to the resource queues corresponding to the anomaly monitoring tasks, and based on the multimodal link conflict detection map... Issue a resource preemption command to the optimal sensor group set; For sensors that fail the condition verification, an online self-calibration process is initiated, including offset calibration, noise baseline reestimation, and synchronization alignment based on the synchronization pulse sequence, and the hardware clock offset residual is calculated. The sensor generates a ready flag by updating the local timing reference through an iterative correction loop to complete self-calibration. Before all sensors generate a readiness flag, the inspection robot injects a pre-sampled pulse sequence into the optimal sensor set. The consistency of the trigger delay and response time of each sensor is verified. After all sensors have completed the pre-sampling pulse sequence verification, the inspection robot issues a collaborative acquisition command based on the optimal sensor set.

7. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 6, characterized in that, The process by which the inspection robot issues a collaborative acquisition command based on hardware synchronization to acquire time-stricken, multimodal, high-frequency data streams includes: A unified hardware synchronization pulse sequence is generated in the inspection robot control module. The hardware synchronization pulse sequence includes a synchronization clock domain identifier, a trigger time identifier, and a sampling window identifier, and is accompanied by a high-frequency jitter suppression parameter, which is used to establish trigger edge consistency among the vibration sensor, sound sensor, and temperature sensor in the optimal sensor group set. Based on the sampling parameter strategy determined by the trigger logic network, the inspection robot generates local sampling configurations for vibration sensors, sound sensors, and temperature sensors in the optimal sensor set. The local sampling configurations include the sampling frequency, sampling window length, and sampling start offset matrix based on static offset of the corresponding sampling parameter strategy, and enter high-frequency sampling mode after the collaborative acquisition command is activated. In high-frequency sampling mode, the optimal sensor set samples vibration, sound, and temperature operating parameters to form high-frequency vibration data, high-frequency sound data, and high-frequency temperature data. A cross-modal unified timestamp is added to the sampled data according to a unified hardware synchronization pulse sequence, and cross-sensor time alignment is completed based on the synchronization clock domain. The inspection robot computing unit constructs a time alignment verification stack to verify the alignment residuals of the high-frequency vibration data, high-frequency sound data, and high-frequency temperature data frame by frame. Time consistency weights are generated based on the alignment residuals, and after the verification is passed, the data is transmitted to the inspection robot computing unit in a unified data channel encoding format.

8. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 7, characterized in that, The process of constructing a multimodal fusion network based on a hierarchical attention mechanism to extract features and perform fusion analysis on the multimodal high-frequency data stream includes: In the computational unit of the inspection robot, a multimodal fusion network with a hierarchical attention mechanism consisting of a local attention layer and a global attention layer is constructed. The local attention layer performs feature extraction on high-frequency vibration data, high-frequency sound data and high-frequency temperature data respectively, and introduces time consistency weights generated by the time alignment verification stack to form local vibration features, local sound features and local temperature features. The vibration local features, sound local features, and temperature local features are input into the global attention layer. A cross-modal attention weight matrix is ​​constructed in the global attention layer, and a modal weight bias vector composed of anomaly type information is added to the cross-modal attention weight matrix to form a fused feature representation. The inspection robot's computing unit performs abnormal pattern recognition based on the fused feature representation. When there is a synergistic correlation between the vibration abnormal pattern, sound abnormal pattern and temperature abnormal pattern in the fused feature representation, fault symptom information is generated.

9. The method for multimodal monitoring and intelligent early warning of power equipment based on inspection robots according to claim 8, characterized in that, The process of inputting the fault symptom information into a diagnostic engine that integrates data-driven and knowledge-driven approaches, performing a preliminary fault type determination through the diagnostic engine, and traversing the causal relationship chains in the fault knowledge graph includes: Perform feature transformation on the fault symptom information to generate a fault feature vector, and generate a preliminary fault type judgment result based on the fault feature vector; Based on the initial fault type judgment result, locate the starting node associated with the initial fault type judgment result in the fault knowledge graph, extract the node sequence containing the target device node, symptom node, event node and handling node from the starting node, and construct a causal relationship chain representation using the node sequence; The fault feature vector and the causal relationship chain representation are used to perform relationship matching and path traversal. During the relationship matching process, a relationship function containing node embedding vectors is used. , for node pairs The relation type is encoded, where the relation function... The expression is: ; Calculate path weights while traversing the path set. : ; In the formula, For relational encoding matrix; , These are nodes in the causal relationship chain representation. With nodes Node embedding vectors; For activation functions; For any path in the causal relationship chain representation. A set of paths; This is the compatibility difference component between the fault feature vector and the node embedding vector in the path; These are the weighting coefficients for the compatibility difference components; From path set The path with the largest path weight is selected, and a diagnostic result vector is generated from the path. The diagnostic result vector includes a fault type identifier, a root cause identifier, and a treatment suggestion identifier. An interpretable early warning report is generated based on the diagnostic result vector and output to the inspection robot control module.

10. The method for multimodal monitoring and intelligent early warning of power equipment based on an inspection robot according to claim 9, characterized in that, The process of using the full-link data and diagnostic results from this anomaly monitoring for adaptive updates includes: A full-link feature set is constructed from the high-frequency vibration data, high-frequency sound data, and high-frequency temperature data collected during this anomaly monitoring, along with time alignment information. This full-link feature set, along with the diagnostic result vector, is input into the policy update module of the trigger logic network. In the policy update module, the cost function is applied... Calculate the cost of actions in the action space: ; In the formula, , , Respectively represent actions The corresponding calculation output, bandwidth output, and sensor output; , , To calculate the weighting coefficients for the output volume, bandwidth output volume, and sensor overhead; The action selection strategy of the triggering logic network is updated based on the principle of minimizing the cost function; Based on the full-link feature set, the statistics of vibration, sound and temperature operating parameters are calculated, a threshold correction vector is constructed, the adaptive threshold is recalibrated through the threshold correction vector, and the recalibrated adaptive threshold is written into the threshold management table. An incremental training set was constructed from the data samples input to the lightweight real-time discriminant model during this anomaly monitoring process. Based on the incremental training set, a parameter update step was performed to update the convolutional unit parameters and gating unit parameters of the lightweight real-time discriminant model. Based on the fault type identifier, root cause identifier, and treatment suggestion identifier in the diagnostic result vector, an incremental graph structure is constructed. New nodes and new relationship types are added to the fault knowledge graph. Local recalculation of the node embedding vector is performed based on the incremental graph structure, and the recalculation results are written into the node embedding matrix of the fault knowledge graph.