A multimodal collaborative perception power station high-risk operation inspection method and system

By using drones and quadruped robots to work together and integrating visual, auditory, and olfactory modal data, potential hazards in power plants can be identified, solving the problem of low inspection efficiency under a single perception modality and achieving more efficient and safer inspections.

CN120236249BActive Publication Date: 2025-10-21BEIJING HUADIAN TIANREN ELECTRIC POWER CONTROL TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510726691.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-21
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Due to the limitations of a single perception modality in existing technologies, it is difficult to fully identify potential hazards in a dynamic environment, resulting in low efficiency in power plant inspections and the easy omission of potential hazards.

Method used

A multimodal collaborative perception method for high-risk power plant operations inspection is adopted. By working together with drones and quadruped robots, global and local perspectives are integrated to output a video sequence with fused perspectives. Combined with auditory and olfactory modal data, the level of dangerous behavior is identified and a linkage anomaly identification system is established.

Benefits of technology

It improves the ability to identify potential hazards, ensures timely detection and handling of potential risks, and enhances the efficiency and safety of power plant inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236249B_ABST
    Figure CN120236249B_ABST
Patent Text Reader

Abstract

The application provides a power station high-risk operation inspection method and system based on multi-modal collaborative perception, and relates to the technical field of video recognition. The method comprises the following steps: after starting the operation task of the power station, activating a UAV and a quadruped robot; starting a video acquisition unit, establishing a synchronous video stream; and fusing and aligning with a global reference coordinate system through an external synchronization signal; inputting the video sequence of the fused visual angle into a multi-view action behavior recognition network to establish a dangerous behavior level score; obtaining the hearing data and olfactory data of the quadruped robot to establish linkage abnormalities; and reporting an inspection anomaly according to the linkage abnormalities and the dangerous behavior level score. The application solves the technical problem that due to the limitations of a single perception mode, it is difficult to comprehensively identify potential dangers in a dynamic environment, resulting in low inspection efficiency. By fusing multi-modal data, potential risks can be discovered and handled in a timely manner, the accuracy of video and audio recognition is improved, and the inspection efficiency of the power station is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video recognition technology, and in particular to a multi-modal collaborative perception high-risk operation inspection method and system for power plants. Background Art

[0002] High-risk operations in power plants often involve extreme environments such as high voltage, high temperature, and toxic media. Minor operational errors or equipment anomalies can lead to serious accidents such as electric shock, explosion, and falls. In recent years, with the development of intelligent technology, power plant inspections have gradually transformed from manual inspections to automated and intelligent inspections. Currently, common inspection methods often rely on a single perception method, such as vision or sensor networks, which can detect high-risk illegal operations and potential equipment failures to a certain extent. However, these single perception methods have significant limitations, especially in dynamic, complex, and high-risk environments such as power plants. It is often difficult to fully and accurately identify all potential danger signals, which not only limits the improvement of power plant inspection efficiency but also reduces the safety of the inspection process.

[0003] In summary, the existing technology has technical problems such as difficulty in fully identifying potential dangers in a dynamic environment due to the limitations of a single perception modality, easy omission of potential dangers, and low inspection efficiency. Summary of the Invention

[0004] The purpose of this application is to provide a multi-modal collaborative perception method and system for high-risk power plant operation inspections, so as to solve the technical problems in the prior art that, due to the limitations of a single perception modality, it is difficult to fully identify potential hazards in a dynamic environment, potential hazards are easily missed, and the inspection efficiency is low.

[0005] In the first aspect, the present application provides a multimodal collaborative perception inspection method for high-risk operations in power plants, which is implemented by a multimodal collaborative perception inspection system for high-risk operations in power plants, wherein the multimodal collaborative perception inspection method for high-risk operations in power plants includes: after the power plant operation task is started, activating the drone and the quadruped robot to perform collaborative path tracking of the power plant operation task; starting the video acquisition units of the drone and the quadruped robot, performing video data acquisition during the collaborative path tracking process, and establishing a synchronous video stream; after the synchronous video stream is timestamp anchored, using an external synchronization trigger signal to time-align the time-anchored synchronous video stream, and using a global reference coordinate system to perform multi-perspective alignment of the synchronous video stream, and outputting a video sequence of a fused perspective; inputting the video sequence of the fused perspective into a multi-perspective action behavior recognition network to establish a dangerous behavior level score; obtaining the auditory modal data and olfactory modal data collected by the quadruped robot, and using the auditory modal data and olfactory modal data to establish a linkage anomaly; reporting an inspection anomaly based on the linkage anomaly and the dangerous behavior level score.

[0006] Optionally, a temporal spatial displacement path is established according to the power plant operation task; the temporal spatial displacement path is used to call the associated scene of the power plant scene to establish a spatial scene and a ground scene; and the temporal spatial displacement path is followed by path optimization in the spatial scene and the ground scene to establish a collaborative path.

[0007] Optionally, obtain device data of the video acquisition units of the drone and the quadruped robot, and configure tracking distance influence constraints based on the device data; under the tracking distance influence constraints, perform drone obstacle avoidance tracking optimization using the spatial scene to establish a first optimization path; under the tracking distance influence constraints, perform quadruped robot motion stability balance optimization using the ground scene to establish a second optimization path; and establish a collaborative path using the first optimization path and the second optimization path.

[0008] Optionally, a unified spatial coordinate system is constructed, and image features of video sequences from different perspectives are extracted and then projected into the unified spatial coordinate system; the scene skeleton extraction sub-channel of the multi-perspective action behavior recognition network is called to perform feature confidence evaluation of the video sequence projected into the unified spatial coordinate system, an anchor point set is constructed, and a sparse point cloud is generated with the anchor point set; the dense depth estimation layer of the multi-perspective action behavior recognition network is called to perform feature depth data of the video sequence projected into the unified spatial coordinate system; the sparse point cloud is projected onto the feature depth data to perform point-depth fusion, and complete local scene reconstruction and annotation; and a dangerous behavior level score is established based on the local scene reconstruction and annotation.

[0009] Optionally, the scene and character perception layer of the multi-view action behavior recognition network is called to perform power station scene and character perception segmentation within the local scene, and establish a perception segmentation result; the perception segmentation result is used to perceive character behavior actions in the video sequence and configure scene perception; the character behavior action perception and scene perception are used to score the level of dangerous behavior under scene interaction.

[0010] Optionally, a micro-motion recognition supplement mechanism is established; the micro-motion recognition supplement mechanism is used to recognize the micro-motions of character shaking and abnormal pauses in video sequences; and the micro-motion recognition results are added to the character behavior and action perception.

[0011] Optionally, after performing auditory modality and olfactory modality modeling, abnormal feature extraction is performed on the auditory modality data and olfactory modality data to establish abnormal feature extraction results; a linkage trigger rule base is established, and linkage trigger recognition of the linkage trigger rule base is performed based on the abnormal feature extraction results to establish linkage trigger recognition results; and linkage abnormalities are established using the linkage trigger recognition results.

[0012] Optionally, the linkage anomaly and the dangerous behavior level score are used to perform a trigger upgrade evaluation of the behavior in the abnormal scenario to generate a trigger upgrade evaluation result; and the inspection anomaly is reported according to the trigger upgrade evaluation result.

[0013] Optionally, the linkage anomaly is used to activate the drone for perspective transfer positioning, and a perspective transfer positioning result is established; video source tracing and identification is performed based on the perspective transfer positioning result, and a source tracing and identification result is generated; and the linkage anomaly is updated according to the source tracing and identification result.

[0014] Optionally, video preprocessing of the video data acquisition result is performed, wherein the video preprocessing includes motion compensation, image sharpening, edge enhancement, and overexposure marking; and a synchronous video stream is established based on the video preprocessing.

[0015] Optionally, the acquisition posture is recorded at each acquisition node, and the acquisition posture is used to generate a time-series shooting perspective; the time-series shooting perspective is input into a perspective correction channel, and the synchronous video stream is established according to the perspective correction channel and the video preprocessing.

[0016] In the second aspect, the present application also provides a multimodal collaborative perception power plant high-risk operation inspection system for executing a multimodal collaborative perception power plant high-risk operation inspection method as described in the first aspect, wherein the multimodal collaborative perception power plant high-risk operation inspection system includes: a collaborative tracking module for activating the UAV and the quadruped robot to perform collaborative path tracking of the power plant operation task after the power plant operation task is started; a video acquisition module for starting the video acquisition unit of the UAV and the quadruped robot, performing video data acquisition in the process of collaborative path tracking, and establishing a synchronous video stream; a synchronous alignment module for activating the UAV and the quadruped robot to perform collaborative path tracking of the power plant operation task after the power plant operation task is started; a video acquisition module for activating the video acquisition unit of the UAV and the quadruped robot, performing video data acquisition in the process of collaborative path tracking, and establishing a synchronous video stream; a synchronous alignment module for activating the video acquisition unit of the quadruped robot in the synchronous video After the stream is timestamp anchored, the time-anchored synchronous video stream is time-aligned using an external synchronization trigger signal, and the multi-perspective alignment of the synchronous video stream is performed using a global reference coordinate system to output a video sequence of a fused perspective; a hazard level assessment module is used to input the video sequence of the fused perspective into a multi-perspective action behavior recognition network to establish a hazard behavior level score; a linkage anomaly establishment module is used to obtain the auditory modal data and olfactory modal data collected by the quadruped robot, and use the auditory modal data and olfactory modal data to establish a linkage anomaly; an inspection anomaly assessment module is used to report an inspection anomaly based on the linkage anomaly and the hazard behavior level score.

[0017] One or more technical solutions provided in the present application have at least the following beneficial effects: after the power plant operation task is started, the UAV and the quadruped robot are activated to perform collaborative path tracking of the power plant operation task; the video acquisition units of the UAV and the quadruped robot are started, and video data acquisition is performed during the collaborative path tracking process to establish a synchronous video stream; after the synchronous video stream is timestamp anchored, the time-anchored synchronous video stream is time-aligned using an external synchronization trigger signal, and the multi-perspective alignment of the synchronous video stream is performed using a global reference coordinate system to output a video sequence of a fused perspective; the video sequence of the fused perspective is input into a multi-perspective action behavior recognition network to establish a dangerous behavior level score; the auditory modal data and olfactory modal data collected by the quadruped robot are obtained, and the auditory modal data and olfactory modal data are used to establish a linkage anomaly; and inspection anomalies are reported based on the linkage anomaly and the dangerous behavior level score. In other words, through the collaborative work of drones and quadruped robots, the global perspective of the drone and the local perspective of the quadruped robot are integrated, the parallax caused by the difference in perspective is eliminated, the video sequence of the fused perspective is output, the multi-perspective recognition network is called to identify human behavior and actions, and the dangerous behavior level scoring is performed in combination with scene perception. The auditory and olfactory modal data of the quadruped robot are obtained, and the linked anomaly recognition is established, which provides more dimensions for detecting potential dangers, improves the ability to identify potential dangers, ensures that potential risks are discovered and dealt with in a timely manner, and effectively improves the inspection efficiency and safety of power stations. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flowchart of a multi-modal collaborative sensing inspection method for high-risk power plant operations in this application;

[0019] Figure 2 This is a structural diagram of a multimodal collaborative perception high-risk operation inspection system for power plants applied in this application.

[0020] Explanation of the accompanying drawings: collaborative tracking module 11, video acquisition module 12, synchronization alignment module 13, danger level assessment module 14, linkage anomaly establishment module 15, inspection anomaly assessment module 16. DETAILED DESCRIPTION

[0021] This application solves the technical problem in the prior art of low inspection efficiency due to the limitations of a single perception modality, which makes it difficult to fully identify potential hazards in a dynamic environment and easily miss potential hazards, by providing a multi-modal collaborative perception inspection method and system for high-risk operations in power plants. Through the collaborative work of drones and quadruped robots, the global perspective of the drone and the local perspective of the quadruped robot are integrated to eliminate the parallax caused by the difference in perspective, output a video sequence of the fused perspective, call a multi-perspective recognition network to identify human behavior, combine scene perception to perform dangerous behavior level scoring, obtain the quadruped robot's auditory and olfactory modal data, and establish linked anomaly recognition, which provides more dimensions for detecting potential hazards, improves the ability to identify potential hazards, ensures the timely discovery and handling of potential risks, and effectively improves the inspection efficiency and safety of power plants.

[0022] For example 1, please refer to the attached Figure 1 This application provides a multi-modal collaborative sensing inspection method for high-risk power plant operations, which specifically includes the following steps:

[0023] S100: After the power plant operation mission is started, the UAV and the quadruped robot are activated to perform collaborative path tracking of the power plant operation mission.

[0024] Furthermore, S100 of the present application includes: establishing a temporal spatial displacement path according to the power plant operation task; using the temporal spatial displacement path to call the associated scene of the power plant scene to establish a spatial scene and a ground scene; executing the path optimization of the temporal spatial displacement path in the spatial scene and the ground scene to establish a collaborative path.

[0025] Furthermore, the present application also includes the following steps: obtaining device data of the video acquisition units of the drone and the quadruped robot, and configuring tracking distance influence constraints based on the device data; under the tracking distance influence constraints, performing drone obstacle avoidance tracking optimization with the spatial scene to establish a first optimization path; under the tracking distance influence constraints, performing quadruped robot motion stability balance optimization with the ground scene to establish a second optimization path; establishing a collaborative path with the first optimization path and the second optimization path.

[0026] Specifically, a temporal spatial displacement path is designed based on the specific operational tasks of the power plant, including task requirements, environmental characteristics, and work objectives. Path planning is performed based on the specific requirements of the inspection task to ensure comprehensive coverage of the inspection area. A temporal spatial displacement path refers to the motion trajectory of the drone and quadruped robot from the starting point to the end point in the power plant operation scenario over a specific time sequence. It considers the relationship between the spatial dimension (such as the specific location of the drone or quadruped robot) and the temporal dimension (the position at each moment along the path). For example, suppose a drone needs to perform an inspection task on the top of the power plant, while a quadruped robot is responsible for ground inspection. The drone's flight path from the starting point to the end point is calculated within a given time constraint, while the quadruped robot is also calculated to avoid obstacles and track the path. The drone and quadruped robot need to coordinate and complete tasks in different areas within the same time period.

[0027] Using the established temporal spatial displacement path, the power plant scene is associated with the call of the corresponding spatial scene and ground scene data. The spatial scene refers to the virtual environment or physical space in the three-dimensional space of the power plant, involving different heights, positions, equipment status, etc. For example, the spatial environment when planning the flight path of a drone takes into account factors such as the specific location of the equipment in the power plant and aerial obstacles. The ground scene mainly refers to the two-dimensional environment of the ground or robot movement within the power plant, and usually involves the planning of factors such as ground obstacles and channel width. During execution, the drone and robot will call the flight and ground scenes respectively according to the mission requirements, and automatically switch and optimize the path according to the equipment status and mission requirements.

[0028] Obtain device data for video acquisition units (such as cameras, lidar, and infrared sensors) on drones and quadruped robots, including device performance parameters such as camera resolution, field of view, and maximum detection range. Configure tracking distance constraints based on this device data. These constraints define the minimum safe distance between devices (such as drones and quadruped robots) during mission execution, or specific distance limits during mission execution. This ensures that devices do not interfere with each other during mission execution, preventing collisions or interference caused by close proximity, which could impact the smooth progress of the mission. If the drone is equipped with a camera with a narrow field of view, the tracking distance should be set relatively close to ensure that the target object is fully captured in the video. Conversely, if the camera has a wide-angle lens, the tracking distance can be set relatively far. Tracking distance constraints set a minimum safe distance between devices during path planning to ensure that the path planning results take into account the distance between devices (such as drones and quadruped robots). This prevents collisions or close proximity during mission execution, facilitates collaboration between devices, and ensures their safe operation during the joint mission.

[0029] Within the constraints of tracking distance, the drone's path is optimized based on the power plant's spatial landscape. This optimization process requires avoiding obstacles along the flight path while ensuring compliance with mission requirements. First, the safe distance between the drone and the quadruped robot, and between the drone and other equipment within the power plant, must meet the equipment requirements—that is, the minimum safe distance required by the tracking distance constraint. Within the three-dimensional virtual power plant scene, multiple flight paths are dynamically calculated for the drone from its starting point to its destination. These paths must avoid collisions with obstacles (such as towers and equipment).

[0030] The RRT algorithm (Rapidly Exploring Random Trees) is a randomized tree search algorithm for path planning, particularly well-suited for path planning problems in high-dimensional spaces. Starting from a starting point, it generates a tree through random sampling and continuously expands its branches until the tree connects to the target point. Traditional RRT algorithms do not specifically consider safe distances between devices during path search. To ensure safety between devices, the RRT algorithm incorporates tracking distance constraints. Each time a tree node is expanded, the distance between the newly generated path segment and existing segments is checked. If a new path segment is too close to another segment (such as the path of another device), it is rejected. Furthermore, the distance between the newly generated path segment and the path of the ground quadruped robot is checked to ensure that a safe distance is maintained between devices, thus avoiding collisions. If a path violates the distance constraint, the algorithm replans the path and adjusts the position of the path segments to ensure that the minimum safe distance is maintained between all segments.

[0031] During path planning, the drone must avoid obstacles in the spatial scene, including buildings, towers, pipelines, and equipment, while ensuring the path complies with tracking distance constraints. Devices such as lidar, cameras, and infrared sensors are used to acquire real-time obstacle location information and map it to the spatial scene. Based on the RRT algorithm, an optimization algorithm is used to find a feasible path and further optimize it, ensuring that it not only avoids obstacles but also remains smooth and efficient. Considering the potential for dynamic obstacles in the power plant environment (such as personnel and other robots), sensors are used to update obstacle positions in real time and dynamically adjust the path. Path planning generates multiple feasible flight paths that meet the following requirements: avoid all obstacles in the spatial scene, meet tracking distance constraints (maintaining a safe distance between the drone and other equipment), and avoid complex turns or difficult-to-execute flight maneuvers. The RRT algorithm generates multiple paths through random sampling. These paths may have different flight paths, but all meet obstacle avoidance and distance constraints.

[0032] Among multiple feasible paths, the optimal path is selected as the first optimization path. The first optimization path is usually the path with the shortest path length to reduce flight distance and save time; the path with the shortest flight time for the same path length (the straightest, smoothest flight path); and the path with the lowest flight energy efficiency to avoid high-energy flight behaviors (such as sharp turns and large altitude changes). By embedding tracking distance impact constraints in the RRT algorithm and combining them with obstacle avoidance path planning for spatial scenarios, the system can generate multiple feasible paths that meet the constraints and select the optimal path from them, ensuring that the minimum safe distance between devices is met, avoiding collisions and path interference, and selecting the path with the shortest path, minimum flight time, and lowest energy consumption, thereby optimizing the flight efficiency of the drone.

[0033] Similarly, under the influence of tracking distance, the kinematic stability balance optimization of a quadruped robot in a ground scenario requires not only avoiding ground obstacles but also ensuring its kinematic stability during movement. A quadruped robot needs to maintain balance on complex terrain and avoid tilting or tipping over, especially in the presence of irregular ground. Path optimization is used to ensure the stability of the quadruped robot while avoiding all ground obstacles. Kinematic stability balance optimization is a path optimization strategy designed for quadruped robots. It aims to ensure the robot maintains stability and balance during movement, including avoiding unstable states such as tilting and tipping during movement, especially on complex terrain or irregular ground, ensuring that the robot can smoothly complete its tasks.

[0034] Unlike drones, path planning for quadruped robots requires not only obstacle avoidance but also gait optimization and motion strategy adjustment to ensure the robot maintains stability and performs its mission in complex terrain. Ground sensors (such as lidar and ground cameras) monitor ground conditions in real time, capturing information such as surface unevenness and slope, and mapping this information to the ground scene. Based on the robot's motion model, gait planning algorithms (such as control-based gait generation algorithms) optimize the robot's gaits to ensure stable navigation through complex terrain. Specifically, the robot dynamically adjusts its gait frequency, stride length, and posture based on ground conditions to prevent falls or imbalance. The optimal motion trajectories of the robot's joints and legs ensure that the robot maintains balance with each step. For example, in areas with steep slopes, the robot may need to reduce its gait frequency and adjust its center of gravity to reduce the risk of falls.

[0035] The quadruped robot utilizes a wheel-leg hybrid locomotion structure, combining terrain adaptability and high mobility. It can switch between bipedal and quadrupedal modes, enabling it to better adapt to confined spaces and increase operating heights. By simulating the execution of different paths, the robot's kinematic stability is evaluated. This involves analyzing the robot's historical motion data to determine the robot's tilt angle, gait variations, and joint loads as it moves along the path. Selecting the optimal path from multiple feasible paths as the second optimal path not only requires meeting obstacle avoidance requirements and tracking distance constraints but also ensuring the robot maintains stability and efficiency during movement. Typically, the path that best suits the ground conditions and provides the highest stability is chosen as the final path. For example, two feasible paths are planned for the quadruped robot: one is shorter but has irregular terrain, which could cause the robot to lose balance; the other is longer but has flat terrain, allowing the robot to traverse with a more stable gait. After comprehensive evaluation, the flat path is selected as the final second optimal path.

[0036] Based on the optimized paths of the drone and quadruped robot (i.e., the first optimized path and the second optimized path), a collaborative path is established to ensure that the two devices do not interfere with each other when performing tasks and can efficiently complete the inspection tasks. A collaborative path refers to the travel path designed when multiple devices (such as drones and quadruped robots) are performing tasks simultaneously, allowing these devices to coordinate operations, avoid conflicts, and maximize task completion efficiency within a limited time and space. For example, if the drone's path and the quadruped's path intersect at a certain point, collaborative path planning will ensure that they avoid each other before or after this point to avoid collision.

[0037] By tracking distance impact constraints and obstacle avoidance tracking optimization, it is ensured that drones and quadruped robots can effectively avoid collisions and obstacle interference when performing tasks. The quadruped robot can maintain balance in complex terrain through motion stability balance optimization, reducing the risk of tipping and instability. The drone and quadruped robot successfully inspected multiple devices without interfering with each other. The entire inspection process is 30% more efficient than traditional methods, and no equipment collisions or path conflicts occurred.

[0038] After the power plant operation mission is initiated, the collaborative path tracking center, which activates the drones and quadrupeds to perform the task, is activated. A mission start signal is sent to the control modules of the drones and quadrupeds, causing them to enter an active state and activate their respective power supplies and sensors. Based on the mission requirements and real-time environmental data, the central control system assigns the planned first optimal path (the drone's obstacle avoidance and tracking path) and second optimal path (the quadruped's motion stability and balance path) to the corresponding drones and quadrupeds. Path information includes waypoints, speed, and obstacle avoidance strategies. After receiving their respective path information, the drones and quadrupeds initiate path tracking mode. The drones use onboard sensors (such as cameras, lidar, and ultrasonic sensors) to perceive their surroundings in real time and perform obstacle avoidance maneuvers while tracking the planned path. The quadrupeds utilize their own sensors and terrain adaptability to maintain motion stability in complex terrain while tracking the planned path. Once the drones and quadrupeds have completed all planned path points and collected the required data, a mission completion signal is sent back to the central control center. The central control center then issues a recovery command, and the drones and quadrupeds return to their starting point or designated recovery area to await their next mission or maintenance.

[0039] S200: Start the video acquisition units of the UAV and the quadruped robot, perform video data acquisition during the collaborative path tracking process, and establish a synchronous video stream.

[0040] Specifically, when the drone and quadruped robot begin collaborative path tracking, the video acquisition units of the drone and quadruped robot are activated. The video acquisition unit refers to the sensor equipment installed on the drone and quadruped robot, which is used to capture and record video data in real time. These sensors include high-definition cameras, infrared cameras, laser scanners, and other devices. They acquire visual information of the surrounding environment and convert it into digital signals for subsequent processing. The cameras installed on the drone and quadruped robot are configured according to mission requirements. For example, drones may be equipped with high-definition RGB cameras and infrared cameras, while quadruped robots may be equipped with close-range high-definition cameras or 360-degree panoramic cameras. For example, during power plant inspections, quadruped robots may carry high-definition RGB cameras to detect surface damage on equipment, while drones may be equipped with infrared cameras to detect heat sources or overheating equipment. Once activated, these devices begin collecting video data.

[0041] During collaborative path tracking, drones and quadruped robots need to simultaneously perform video data collection tasks, and their equipped video acquisition units continuously collect video data. For example, the drone flies over the power plant according to a predetermined flight path while collecting video data, while the quadruped robot moves along the corresponding path on the ground, simultaneously collecting video data and establishing a synchronized video stream. The video acquisition unit of each device (drone and quadruped robot) adds a timestamp to each frame of video data. Through the collaborative operation of different devices, environmental data from different angles and perspectives is acquired, and the status of power plant equipment is comprehensively identified. Abnormal conditions on power plant equipment (such as overheating and cracks) can be quickly identified and repaired in a timely manner. This not only improves inspection efficiency but also ensures the safe operation of the power plant.

[0042] Furthermore, S200 of the present application includes: performing video preprocessing of the video data acquisition results, the video preprocessing including motion compensation, image sharpening, edge enhancement, and overexposure marking; and establishing a synchronous video stream based on the video preprocessing.

[0043] The acquisition posture is recorded at each acquisition node, and the acquisition posture is used to generate a time-series shooting perspective; the time-series shooting perspective is input into a perspective correction channel, and the synchronous video stream is established according to the perspective correction channel and the video preprocessing.

[0044] Specifically, after video data is collected, it needs to go through a series of preprocessing steps, including motion compensation, image sharpening, edge enhancement, and overexposure marking, to ensure the stability of video quality and data availability. Since quadruped robots and drones may move quickly in dynamic environments, the captured images or videos may appear motion blurred or jittery, so motion compensation is required to estimate and correct the camera's motion offset based on the motion trajectory of the object in the image to reduce blur. The motion trajectory of the device is estimated by tracking feature points (such as corners, edges, etc.) in the video frame. Based on the estimated motion information, the video frames are aligned and corrected through geometric transformations (such as affine transformation, perspective transformation, etc.) to remove blur caused by device motion. The motion of each frame is corrected through motion compensation algorithms, such as feature matching-based stabilization methods.

[0045] To enhance details in video images, especially small objects or distant equipment components, image sharpening is used to increase contrast and edge definition, making key features more distinct. Image sharpening increases contrast in localized image regions, making objects and structures more prominent. It is particularly useful for identifying small targets and analyzing details in tasks such as power plant inspections. Filters (such as the Laplacian operator) extract edge information from the image and enhance the strength of these edges, making object outlines more distinct. By increasing the brightness contrast of localized image regions, the target area becomes more prominent. Sharpened images better support subsequent target recognition and classification tasks, reducing misclassifications caused by image blur or insufficient detail.

[0046] Edge enhancement enhances the strength of image contours and edges, making the structure and details of objects in the video more prominent. In complex power plant environments, edge enhancement helps clearly identify important elements such as equipment components and human movements. Detected edges are enhanced, typically by increasing their brightness and contrast to make them more distinct. The Canny edge detection algorithm is used to extract and enhance edges in the image.

[0047] When shooting in areas of strong light or high temperatures, videos may appear overexposed. Overexposure marking technology is used to mark and process these overexposed areas to avoid image information loss and ensure that image quality in critical areas is not affected. Areas in the image where the brightness exceeds a threshold are detected and marked as overexposed. These areas typically appear white or very bright and lack detail. Overexposed areas can be repaired by adjusting brightness and contrast to reduce their impact, or image restoration techniques can be used to fill in the details in these areas. For example, the pixel values ​​of an image can be converted into an image histogram. By analyzing the image's brightness histogram, it can be determined which areas have excessive brightness and, therefore, whether overexposure exists. Brightness adjustment is performed on overexposed areas to reduce the brightness value in the image and restore detail.

[0048] The capture attitude is recorded at each acquisition node. Specifically, the position, angle, and sensor attitude of each video acquisition node are recorded. This acquisition node attitude record typically includes location (e.g., GPS coordinates) and attitude angles (e.g., pitch, roll, and yaw). Drones and quadruped robots constantly change their attitude during movement, and each captured image frame is tagged with the corresponding attitude information. This captured attitude information is used to generate a time-series viewpoint for each frame, ensuring spatial and temporal alignment of the images. For example, when simultaneously capturing footage from the air and on the ground, the difference in viewpoints is recorded to facilitate subsequent temporal and spatial synchronization of the two video streams. For example, when a quadruped robot captures a video of a device failure during an inspection, the robot's position, attitude, and the angle of the video frame are recorded. When a drone captures footage from different viewpoints, the relevant drone data is recorded according to the same standards, and the corresponding time-series viewpoints are generated.

[0049] The time-series camera angle is determined by the spatial position and orientation of the device at each acquisition node. The camera angle corresponding to each acquisition node is inferred based on the device's motion trajectory, direction, and tilt angle. In other words, the camera angle and field of view at each moment in the video are calculated based on the device's posture data and time series, resulting in the time-series camera angle.

[0050] The view correction pipeline is an image processing process used to unify video frames captured at different times or from different device angles to a standard or reference viewpoint. By applying geometric transformations (such as rotation, scaling, and perspective transformation) to the video frames, it eliminates viewpoint deviations caused by changes in device posture. Because images captured by different devices at different times may have different viewpoints and angles, view correction is necessary to align video frames from different viewpoints. The goal of the view correction pipeline is to ensure that all captured video frames are aligned within a unified coordinate system or viewpoint.

[0051] The time-series camera perspective serves as input, providing the capture position and orientation of each video frame at each moment. The perspective correction pipeline adjusts the image based on this information. Specifically, based on the time-series camera perspective information for each video frame, a coordinate system transformation is performed, transferring the video data from the coordinate system at the time of capture to a unified spatial coordinate system. Through a perspective transformation (perspective matrix), video frames from different perspectives are adjusted to a standard perspective, aligning images captured by different devices or cameras. Video frames from different perspectives are reconstructed and aligned based on the same standard perspective to form a synchronized video stream.

[0052] By performing perspective correction on pre-processed video, image distortion caused by factors such as device angle differences and shooting distances is corrected, ensuring consistent visual quality in the resulting synchronized video streams. Through video pre-processing and perspective correction, video data from different sources can be precisely integrated and synchronized in a multi-device, multi-view environment, improving overall analysis accuracy and ensuring consistent presentation of the same device component or scene from different perspectives.

[0053] S300: After the synchronized video stream is timestamp anchored, an external synchronization trigger signal is used to time-align the time-anchored synchronized video stream, and a global reference coordinate system is used to align multiple perspectives of the synchronized video stream, and a video sequence of a fused perspective is output.

[0054] Specifically, the video capture unit of each device (drones and quadruped robots) adds a timestamp to each frame of video data, recording the exact time each frame was captured. This timestamp can be generated by the device's internal clock or calibrated using an external synchronization signal (such as a GPS clock or NTP protocol signal). Timestamp anchoring refers to assigning an accurate time stamp to each frame in the captured video stream. An external synchronization trigger signal is a signal from an external source, typically used to coordinate time synchronization between multiple devices (such as drones and quadruped robots). Such signals include GPS clock signals, NTP (Network Time Protocol) signals, or other precise synchronization signals to ensure that data collected by multiple devices at the same time accurately corresponds.

[0055] Using an external trigger device (such as a GPS clock module or NTP time server), each device is provided with a unified time signal. The acquisition modules of all devices begin collecting data based on this unified time signal. A central control center sends a synchronization signal, triggering the drone and quadruped robot to simultaneously begin collecting video data. After receiving the synchronization signal, the video acquisition unit aligns the captured video streams based on the timestamps, ensuring that video data from different viewpoints remains consistent. Time alignment involves synchronizing video streams from different devices (such as drones and quadruped robots) based on timestamps, ensuring that data collected at the same time matches, thus avoiding data errors caused by device clock asynchrony.

[0056] The global reference coordinate system is a unified coordinate system used to describe the spatial positions of the entire environment. In a multi-robot system, each device (such as a drone or quadruped) uses this coordinate system for positioning and navigation to ensure consistent spatial positioning across all devices. This global reference coordinate system can typically be a GPS coordinate system or a coordinate system established within a power plant. Using this global reference coordinate system, video data from different perspectives is integrated into a unified space. Each device (such as a drone or quadruped) determines its position and orientation using the global reference coordinate system and adjusts its video stream based on this coordinate system, enabling seamless integration of videos from different perspectives. For example, a quadruped captures images of the power plant ground level, while a drone captures images of the tower and overhead equipment. Using spatial alignment technology, these two video streams can be seamlessly combined within the same coordinate system.

[0057] After temporal and spatial alignment, video streams from different devices can be fused to form a unified, multi-perspective video sequence, displaying data from different viewpoints simultaneously. This fused video allows task personnel to gain a comprehensive understanding of the status of equipment within the power plant. A fused-perspective video sequence combines multiple video streams from different viewpoints into a single video sequence after temporal and spatial alignment. This allows for simultaneous observation of the environment from multiple angles, helping task personnel fully understand the on-site situation.

[0058] By anchoring timestamps in synchronized video streams, time alignment using external synchronization trigger signals, and multi-view alignment using a global reference coordinate system, the system ultimately outputs a video sequence with fused perspectives. This ensures that video data from different devices (drones and quadruped robots) are highly consistent in time and space, avoiding data errors caused by time asynchrony or spatial misalignment. This provides a comprehensive and accurate picture of the on-site situation, showing the status of power plant equipment from different perspectives. This allows for timely detection of equipment failures in the power plant (such as overheating and electrical failures) and immediate repair measures.

[0059] S400: Inputting the video sequence of the fused perspective into a multi-perspective action behavior recognition network to establish a dangerous behavior level score.

[0060] Furthermore, S400 of the present application includes: constructing a unified spatial coordinate system, extracting image features from video sequences of different perspectives, and projecting them into the unified spatial coordinate system; calling the scene skeleton extraction sub-channel of the multi-perspective action behavior recognition network, executing feature confidence evaluation of the video sequence projected into the unified spatial coordinate system, constructing an anchor point set, and generating a sparse point cloud with the anchor point set; calling the dense depth estimation layer of the multi-perspective action behavior recognition network, executing feature depth data of the video sequence projected into the unified spatial coordinate system; projecting the sparse point cloud into the feature depth data for point-depth fusion, and completing local scene reconstruction and annotation; establishing a dangerous behavior level score based on the local scene reconstruction and annotation.

[0061] Furthermore, the present application also includes the following steps: calling the scene and character perception layer of the multi-view action behavior recognition network, performing power station scene and character perception segmentation within the local scene, and establishing a perception segmentation result; using the perception segmentation result to perform character behavior action perception under the video sequence, and configure scene perception; using the character behavior action perception and scene perception to perform dangerous behavior level scoring under scene interaction.

[0062] Specifically, based on the actual geographic location and equipment layout of the power plant, a unified spatial coordinate system is established to map spatial data from different sources or perspectives (i.e., image data collected by drones and quadruped robots) into the same coordinate system, enabling comprehensive analysis of data from different perspectives. Image features are extracted from videos collected from different perspectives by various devices (such as drones and quadruped robots), and representative feature points are identified. These points can be used to describe the image content and establish correspondences between different perspectives. Image feature extraction involves identifying and extracting key visual information from images, such as edges, textures, and corners. Image feature extraction algorithms include SIFT (Scale-Invariant Feature Transform), SURF (Speeded Robust Features), and ORB (Rotation-Invariant Binary Descriptor).

[0063] After image feature extraction is complete, image feature points extracted from different perspectives (such as videos captured by drones and quadruped robots) are mapped to a unified spatial coordinate system, converting image features from image space (pixel coordinates) to actual 3D spatial coordinates. Images are calibrated and corrected using the camera's intrinsic and extrinsic parameters (such as focal length, lens distortion, camera position, and orientation). A stereo matching algorithm calculates the position of each feature point in 3D space based on images from multiple perspectives, thereby achieving spatial projection of the image. For example, after image feature extraction from drones and quadruped robots, the coordinates of these feature points in 3D space are calculated using a stereo matching algorithm and mapped to the unified coordinate system of the power plant digital twin. This means that although the feature points were originally captured from different angles, their positions in the unified coordinate system accurately correspond.

[0064] SIFT is an algorithm for extracting local features from images. It can identify key points (such as corners and edges) within an image. These key points are invariant to scale, rotation, and illumination. SURF is an improved version of the SIFT algorithm, primarily accelerating the feature point extraction process by increasing the algorithm's computational speed. It finds key points by calculating local features of the image, using techniques such as Gaussian blur and scale space to ensure feature stability, and extracting information such as the coordinates, orientation, and scale of the key points. After extracting feature points from video frames with different perspectives, a feature matching algorithm is used to match the feature points from different perspectives. Because video frames often contain noise, false matches, and outliers, the RANSAC algorithm can effectively eliminate false matches that do not conform to the model and improve matching accuracy. Two pairs of feature points are randomly selected, and the transformation relationship (such as rotation and translation) between them is calculated. This transformation relationship is then used to match other feature points, and the error between the transformed feature points and the feature points in the original image is calculated. The matching point with the smallest error is retained as the inlier point that conforms to the transformation model. The random point selection and model fitting process are repeated until the optimal transformation model is found. After completing feature matching, these matched feature points are mapped to a unified spatial coordinate system. Through camera calibration, the camera's intrinsic parameters (such as focal length, principal point, etc.) and extrinsic parameters (such as the camera's spatial position and orientation) are obtained, and the image coordinates are converted into actual spatial coordinates.

[0065] The scene skeleton extraction sub-channel of the multi-view action recognition network projects the skeleton features and scene information in each video frame into a unified spatial coordinate system and performs a confidence assessment on these features. Feature confidence assessment primarily determines the reliability of each feature point based on image quality, viewpoint matching stability, and background interference. The scene skeleton extraction sub-channel extracts a skeleton feature map (keypoint locations, joint connection information, etc.) for each video frame. Each video camera (such as those on quadruped robots and drones) has a set of intrinsic and extrinsic parameters. Intrinsic parameters include information such as focal length and optical center, while extrinsic parameters include the camera's position and pose relative to the global coordinate system. Skeleton feature points (positions in the image coordinate system) are converted to three-dimensional coordinates in the unified spatial coordinate system using the camera's intrinsic and extrinsic parameter matrices.

[0066] Feature confidence evaluation quantifies the reliability of the skeleton features extracted from each video frame. This is done based on factors such as image quality, feature stability, matching stability, and the depth information of the feature points. Feature points in video frames may be unstable in low or bright light conditions, resulting in lower confidence scores. Image noise or blur can lead to inaccurate feature extraction, so the confidence score needs to be adjusted based on the image quality score. Images captured during rapid motion may cause blur, affecting feature extraction. The stability of the skeleton features in each frame is evaluated to determine their reliability. Large changes in viewpoint within an image may cause inconsistent feature point positions for the same object across different frames. The confidence score of the feature points is adjusted based on the magnitude of the viewpoint change. If the matching relationship between the same feature points in multiple viewpoints is unstable or deviates significantly, the confidence score will be reduced. If depth information is available (e.g., using stereo vision or a depth camera), the depth map helps determine the position and stability of feature points in 3D space.

[0067] Based on the above factors, the confidence level of each feature point is calculated, and those with high confidence levels are selected as the anchor point set. These are key, stable features in the scene that can provide accurate spatial information. Specifically, based on a set confidence threshold, feature points with confidence levels above the threshold are selected to form the anchor point set. Spatial consistency testing (for example, checking whether feature points from multiple perspectives are located on the same object) further identifies stable anchor points.

[0068] Using the anchor point set and its matching relationships across different video viewpoints, a sparse point cloud is generated, representing key locations or feature points in the scene. These points have corresponding coordinates in 3D space, representing the locations of key features on the surface of an object or in the environment. While sparse point clouds contain less data than dense point clouds (which contain a large number of data points), they still provide sufficient information for subsequent 3D reconstruction and analysis.

[0069] Calling the dense depth estimation layer of the multi-view action recognition network to perform feature depth data projection of the video sequence into a unified spatial coordinate system means using the dense depth estimation layer to estimate the depth information of each pixel in the video sequence. Dense depth estimation refers to directly calculating the depth value of each pixel from the image. The video sequence projected into the unified spatial coordinate system is input into the dense depth estimation layer to obtain a depth map for each video frame, which represents the spatial position of equipment and personnel in the power plant. The depth values ​​in the depth map are usually represented using grayscale levels, with darker areas indicating farther distances and brighter areas indicating closer distances.

[0070] This depth data is projected from the local coordinate system of each camera's viewpoint into a unified spatial coordinate system, ensuring that video data from different viewpoints can be aligned in a shared coordinate system, enabling accurate 3D spatial reconstruction. By combining the depth value of each pixel with its corresponding viewpoint information (such as the camera's orientation and position), the depth information can be converted from the local coordinate system to a global coordinate system. For example, imagine a power plant where a drone captures an area containing power equipment from the air, while a quadruped robot simultaneously captures the same area from the ground. The depth data from each viewpoint is generated by a depth estimation network and projected into a global coordinate system. This allows both the drone's aerial view and the quadruped robot's ground view to be combined in a unified coordinate system, further facilitating 3D modeling and analysis of power plant equipment.

[0071] The sparse point cloud is fused with the depth estimation map, and each point in the sparse point cloud is combined with the corresponding depth information to generate a more accurate three-dimensional scene. The sparse point cloud data is aligned with the depth map, and the spatial coordinates of the point cloud are converted into actual three-dimensional coordinates by calculating the spatial position of each point and the depth value of the corresponding pixel. In the same coordinate system, the sparse point cloud is combined with the depth information in the depth map, and a dense point cloud or depth map is generated through weighted fusion to complete the reconstruction and annotation of the local scene. The various parts of the scene are annotated according to the position of the identified equipment or personnel. Point-deep fusion refers to combining the feature points in the sparse point cloud with the feature depth data, and generating a more accurate three-dimensional scene reconstruction by weighted fusion of the depth value of each feature point, improving the spatial representation of the sparse point cloud, and making it more accurate in local scene reconstruction.

[0072] The scene and person perception layer of the multi-view action recognition network segments the power plant scene and people within the local scene. Key scene elements (such as power plant equipment and building structures) and dynamic elements (such as workers) are identified from the video. Using an object detection algorithm, people are detected in each video frame to accurately locate and identify the workers. The detection algorithm typically outputs a bounding box that marks the position of each worker in the image. In addition to person detection, the scene and person perception layer also segments static elements (such as power plant equipment and buildings) within the power plant scene, dividing the image into multiple regions. Each region is assigned a label. The perceptual segmentation result generates a segmentation map, where each pixel's label indicates the category to which it belongs. The perceptual segmentation result represents various information segmented within the entire power plant scene using a deep learning algorithm, including the location and status of power plant equipment and workers.

[0073] Based on the perception and segmentation results, the system analyzes human movements in the video sequence, including both task-specific movements (such as standing, walking, squatting, and climbing) and subtle movements (such as shaking and pausing). Each behavior type is assigned a corresponding hazard level, particularly for high-risk operations (such as high-voltage contact and equipment disassembly). Simultaneously, scene perception is implemented. By training on different scenarios (such as power plant operation areas, equipment areas, and hazardous areas) and operator behaviors, the scene perception model can automatically identify different areas within the power plant and their risk levels. It identifies elements in each scene in real time and determines which areas are high-risk. For example, the high-voltage equipment area is marked as a high-risk area, while the normal operation area is marked as a safe area.

[0074] By setting risk level rules (such as contact with high-voltage equipment and entering restricted areas), the system automatically assesses the risk level based on the worker's behavior and the surrounding environment. The system then evaluates the riskiness of the behavior and generates a risk level score, such as low, medium, or high. For example, a worker entering a live area is considered a level 1 risk; a worker not wearing a hardhat is considered a level 2 risk; and a worker entering or touching high-voltage equipment is considered a level 5 risk. Based on an interactive analysis of the worker's behavior and the environment, each risk behavior is assigned a level (e.g., 1-5, with 1 being low risk and 5 being extremely high risk) according to a series of scoring rules. These scoring rules include analyzing the worker's specific behavior, such as whether they performed dangerous operations (such as touching electrical equipment or entering a dangerous area); and analyzing the environment in which they were located, such as whether they entered a high-voltage area or came into contact with hazardous materials. Based on the type of behavior and the risk level of the environment, each interaction is scored and an overall risk level is assigned, helping workers quickly identify dangerous situations and take necessary preventative measures.

[0075] By calling on the multi-view recognition network to identify human behavior and actions, combining scene perception to perform dangerous behavior level scoring, supplementing the micro-motion recognition mechanism to enhance accuracy, assigning a danger level to each behavior, improving inspection efficiency, reducing unnecessary movements, and ensuring coverage of all key areas.

[0076] Furthermore, the present application further comprises the following steps:

[0077] A micro-motion recognition supplement mechanism is established; the micro-motion recognition supplement mechanism is used to recognize micro-motions such as shaking and abnormal pauses of characters in video sequences; and the micro-motion recognition results are added to the character behavior and action perception.

[0078] Specifically, a supplementary mechanism for micro-movement recognition should be established to identify subtle changes in a worker's body movements, which typically occur over short periods of time and manifest as low-amplitude movements. This includes jitter (small vibrations caused by nervousness, cold, mechanical vibration, etc.) and abnormal pauses (brief pauses in movement caused by equipment failure, loss of balance, or unexpected problems). To capture micro-movements, the video acquisition system needs to have a high frame rate (e.g., 30 frames per second or higher) to meticulously record the worker's subtle movements, thereby improving the accuracy of micro-movement recognition. Micro-movements are typically subtle changes over a short period of time, so real-time analysis of several consecutive frames (e.g., 5-10 frames) is required.

[0079] Using continuous frames from high-frame-rate video sequences, the micro-motion recognition supplementary mechanism performs subtle motion analysis. By calculating minute displacements and posture changes in the operator's joints and limbs, the system determines whether there is involuntary body tremor—rapid, small movements of a body part (such as the hand, head, or shoulder) over a short period of time. An abnormal pause typically occurs when an operator suddenly pauses in a certain position, exceeding the normal operating time. For example, while operating equipment, an operator may experience slight hand tremors due to nervousness or cold. The micro-motion recognition supplementary mechanism captures this subtle tremor in the video sequence and flags it as a potential signal of "physical discomfort."

[0080] Abnormal pauses can be caused by equipment failure, sudden loss of balance, fatigue, or other emergencies. Time series analysis is used to determine changes in the operator's speed and position. If the operator's movement speed approaches zero within a short period of time and no normal movement transitions (such as switching operations or adjusting equipment) occur, an abnormal pause is considered to have occurred.

[0081] Once micro-movements (such as jitters or unusual pauses) are identified and incorporated into the person's behavioral perception data, the micro-movement recognition and behavioral perception results together form the operator's current behavioral model. Micro-movements (such as jitters and pauses) may reflect changes in the operator's health or equipment status. Therefore, the presence of micro-movements in the behavioral perception model may influence the operator's behavioral assessment and further adjust their risk level score.

[0082] By introducing a supplementary mechanism for micro-motion recognition, subtle movements of workers (such as shaking and abnormal pauses) are captured and integrated with human behavior perception results, further enhancing the ability to detect potential risks. By establishing a supplementary mechanism for micro-motion recognition, subtle movements such as shaking and abnormal pauses of workers are identified in real time and integrated into human behavior perception, thereby improving the ability to detect potential risks.

[0083] S500: Acquire auditory modal data and olfactory modal data collected by the quadruped robot, and establish linkage anomaly using the auditory modal data and olfactory modal data.

[0084] Furthermore, S500 of the present application includes: after executing auditory modality and olfactory modality modeling, performing abnormal feature extraction on the auditory modality data and olfactory modality data, and establishing abnormal feature extraction results; establishing a linkage trigger rule base, performing linkage trigger identification of the linkage trigger rule base based on the abnormal feature extraction results, and establishing a linkage trigger identification result; and using the linkage trigger identification result to establish a linkage abnormality.

[0085] Specifically, auditory and olfactory modal modeling is performed. Auditory modal modeling includes abnormal characteristics of arc sound / discharge sound, such as high-frequency sharpness / sudden shortness, indicating possible risks of insulation breakdown and poor contact; abnormal characteristics of vibration noise changes, such as high-frequency changes or offsets, indicating possible risks of loose equipment and dropped tools; abnormal characteristics of unstructured speech, such as cries for help and abnormal shouting, indicating possible risks of worker errors or emergencies. Olfactory modal modeling includes abnormal signal characteristics of ozone, such as instantaneous concentration increases, indicating possible hidden dangers of arc discharge and byproducts; abnormal signal characteristics of alkanes / odor residues, such as continuous exceeding of standards or abnormal fluctuations, indicating possible hidden dangers of flammable gas leakage / oil evaporation; abnormal signal characteristics of burnt odor residual compounds, such as long-tail continuous or repeated changes, indicating possible hidden dangers of equipment short circuit / high-temperature aging.

[0086] Auditory modal data refers to sound data collected by the quadruped robot's sensors or microphones. In complex environments like power plants, sound signals can reveal a wealth of critical information, such as the normal operation of equipment, abnormal equipment noise, personnel shouting, and alarm sounds. Olfactory modal data refers to gas information collected by the robot's gas sensors (such as gas sensor arrays). In high-risk environments like power plants, certain gases (such as toxic and flammable gases, or chemical odors from electrical equipment) can be early warning signs of equipment failure, leakage, or fire. The quadruped robot is equipped with gas sensors to collect this information, monitoring the gas composition of the surrounding environment in real time.

[0087] The sound signal is converted into a spectrogram through Fourier transform, from which information such as frequency, amplitude, and waveform is extracted. Based on the time domain waveform of the audio signal, the sound characteristics such as amplitude, amplitude, and frequency are analyzed. Abnormal noise is identified by comparing it with known normal sound characteristics. For example, sudden high-frequency noise indicates a device malfunction. The gas concentration value output by the gas sensor is obtained, and a gas concentration distribution model within the normal range is established based on historical data. When the sensor detects a concentration outside the normal range, it is judged as abnormal. For example, if the gas sensor detects a sudden increase in carbon monoxide concentration (such as reaching 50ppm), it means that a certain device is leaking or burning abnormally, and further confirmation is required. Features that can reflect abnormal conditions are extracted from the sound data and gas data and annotated.

[0088] Based on multiple sensor information and conditional triggers, a linkage trigger rule library is established. Each rule usually consists of multiple sensor input data and certain specific conditions. When these conditions are met, the corresponding linkage operation is triggered. For example, some examples of the linkage trigger rule library are shown in Table 1:

[0089] Table 1 Linkage trigger rule library

[0090]

[0091] Combining abnormal characteristics of auditory and olfactory modalities, for example, abnormal noise (auditory abnormality) generated by a device and a gas leak (olfactory abnormality) may occur simultaneously. These conditions require automatic linkage based on pre-defined rules, triggering an alarm or executing other emergency actions. For example, if a quadruped robot detects high-frequency noise emanating from electrical equipment accompanied by rising ammonia concentrations, the rule base can trigger linkage recognition based on the combination of high-frequency noise and exceeding ammonia concentrations, generating a warning of a dangerous event.

[0092] Based on an established linkage trigger rule library, the system extracts abnormal features from auditory and olfactory modal data to identify which trigger condition is met and outputs the linkage trigger identification result. For example, in a power plant environment, if the robot detects abnormal noise and gas leak signals, the linkage rule is triggered, issuing a warning of a possible equipment failure with an associated leakage risk and initiating appropriate safety measures.

[0093] Based on the linkage trigger recognition results, linkage anomalies are established. This means that when anomaly information captured by multiple sensors (such as vision, hearing, and smell) corroborates each other, potential multimodal anomaly signals can be promptly identified. This avoids missed reports due to the inability of a single perception modality to fully identify risks, thereby improving safety and response efficiency during inspections. By acquiring auditory and olfactory modal data collected by the quadruped robot, the auditory and olfactory perception processes are simulated to extract useful information from the raw data, identify and extract features that represent anomalies or abnormal conditions, determine when a linkage reaction should be triggered, determine whether a linkage anomaly exists, and take appropriate measures.

[0094] S600: Reporting inspection anomalies based on the linkage anomaly and the dangerous behavior level score.

[0095] Furthermore, S600 of the present application includes: performing a trigger upgrade evaluation of the behavior in the abnormal scenario based on the linkage abnormality and the dangerous behavior level score, and generating a trigger upgrade evaluation result; and reporting the inspection abnormality according to the trigger upgrade evaluation result.

[0096] Specifically, based on the determined linkage anomaly and dangerous behavior level scores, a determination is made as to whether the abnormal behavior requires escalation. For example, if a robot detects linkage anomaly signals of abnormal noise and toxic gas leaks during an inspection, and also identifies an employee performing high-risk operations (such as not wearing safety gear near high-voltage equipment), the dangerous behavior score indicates that the behavior requires an escalation assessment, resulting in a higher-risk safety alert or response. Anomalies are identified in multimodal data, such as abnormal noise or gas leaks, or other environmental changes. Dangerous behavior scores are assigned to the human behaviors in the scenario to assess their risk level. Based on the combination of linkage anomaly and dangerous behavior level scores, if a behavior has a higher risk level (such as incorrect operation or high-risk operation), it triggers a higher-level alert or response. Escalation is performed according to the rules, raising the alert's priority or initiating emergency response measures.

[0097] The trigger escalation evaluation result is a safety warning derived from a comprehensive assessment of linked anomalies and dangerous behavior ratings. For example, if the assessed behavior is potentially high-risk, the alert level is upgraded to a higher level. Based on the dangerous behavior and anomaly type, specific response recommendations are generated, such as immediately ceasing operations, activating the emergency exhaust system, or evacuating personnel. Based on the trigger escalation evaluation results, an inspection anomaly report is generated, summarizing the anomalies in the current scenario and detailing the anomaly type, triggering cause, hazard level, and recommended emergency response measures. The report is automatically sent to the relevant operator or control center to ensure that safety issues are addressed promptly.

[0098] Inspection anomalies also include personnel behavior warnings and scenario-based leak warnings. Human behavior warnings primarily monitor and analyze personnel behavior during power plant inspections, promptly identifying potentially dangerous behaviors and issuing alerts. Human behavior is perceived through video streams and sensor data (such as auditory and olfactory modal data). Each identified behavior is assigned a risk level score, which is used to assess the threat to safety based on various criteria. When an identified behavior is assessed as medium or high risk, a behavior warning is triggered, providing immediate feedback to power plant managers or inspectors via the system interface, audible alarms, or notifications. The warning results include a detailed description of the dangerous behavior and recommended actions.

[0099] Leakage warning uses environmental monitoring and sensor data to provide real-time monitoring and early warning of potential hazardous gas leaks or other dangerous scenarios within a power plant. This includes leak source detection, leak assessment, and emergency response. A quadruped robot and other sensor devices (such as gas sensors and temperature sensors) continuously monitor various indicators within the power plant environment, particularly concentration changes of toxic and hazardous gases (such as ammonia, carbon monoxide, and hydrogen sulfide), to assess the plant's safety environment in real time. Once an anomaly is detected in the collected sensor data, comprehensive analysis is performed in conjunction with scene image information to determine whether a leak is dangerous. For example, if a gas sensor detects a gas concentration exceeding a safe value, combined with image analysis from a visual sensor (e.g., sparks, steam leaks, etc.), it determines whether it is an actual leak.

[0100] When a potential leak is confirmed, a leak alarm is triggered if the concentration of hazardous gas exceeds a preset safety threshold. Image recognition technology is used to pinpoint the leak source, such as a pipe crack or equipment failure. In some cases, dangerous human behavior and a leak can occur simultaneously. For example, a mishandled operation while handling equipment could cause a gas leak. If both warnings are triggered simultaneously, a multi-level response is implemented, such as activating both the gas exhaust system and the evacuation alarm.

[0101] Furthermore, the present application further comprises the following steps:

[0102] The linkage anomaly is used to activate the drone to perform perspective transfer positioning and establish a perspective transfer positioning result; video source tracing and identification are performed based on the perspective transfer positioning result to generate a source tracing and identification result; and the linkage anomaly is updated according to the source tracing and identification result.

[0103] Specifically, a coordinated anomaly refers to a combined anomaly condition identified through the mutual verification of abnormal signals collected by various sensors (such as olfaction, hearing, and vision). Once a coordinated anomaly is detected (e.g., high concentrations of toxic gas and equipment failure occurring simultaneously), a drone is automatically activated for further investigation. When a coordinated anomaly occurs, the drone automatically departs from a predetermined location or the nearest launch point and flies to the anomaly area. During flight, the drone locates the anomaly area based on its geographic location and real-time data, adjusting its flight path and camera angle to achieve the optimal viewing angle, ensuring detailed monitoring and recording of the anomaly scene.

[0104] After the drone performs viewpoint shift positioning, it performs video source identification on the captured video stream data. Using known viewpoints and time points, it can trace the cause and development of the abnormal event from the video data. Based on the video timestamps and the viewpoint shift positioning results, video streams from different viewpoints are temporally and spatially aligned, allowing simultaneous processing of different perspectives of the same event. Based on abnormal behavior or environmental changes (such as human misconduct or equipment failure) captured in the video, combined with known historical data and abnormal signals, the abnormal development trajectory in the video data is analyzed to generate source identification results. Video source identification can track the occurrence of abnormal events and clearly identify their source and development. For example, suppose a fire at a power plant was caused by equipment failure. Using the drone's viewpoint shift positioning, the fire's specific location was confirmed. Video source identification can trace the equipment's condition prior to the fire (such as overheating or a short circuit) and analyze the specific cause of the fire.

[0105] Based on the video source identification results, combined with data from multimodal sensors, more accurate anomaly features are added to update the linkage anomaly, adjusting the response priority or triggering mechanism. For example, if the source identification shows that the fire caused by the equipment failure is more serious than originally judged, a higher level of response will be adjusted. By activating the drone for perspective shift positioning based on the linkage anomaly and using the perspective shift positioning results for video source identification, the occurrence process of abnormal events in high-risk environments such as power plants can be more accurately traced, helping to identify the source of the anomaly, predict potential risks, and improve the efficiency of emergency response.

[0106] In summary, the multimodal collaborative sensing method for high-risk power plant operations provided in this application has the following beneficial effects:

[0107] After the power plant operation task is started, the UAV and the quadruped robot are activated to perform collaborative path tracking of the power plant operation task; the video acquisition units of the UAV and the quadruped robot are started, and video data acquisition is performed during the collaborative path tracking process to establish a synchronous video stream; after the synchronous video stream is timestamp anchored, the time-anchored synchronous video stream is time-aligned using an external synchronization trigger signal, and the multi-perspective alignment of the synchronous video stream is performed using a global reference coordinate system to output a video sequence with a fused perspective; the video sequence with a fused perspective is input into a multi-perspective action behavior recognition network to establish a dangerous behavior level score; the auditory modal data and olfactory modal data collected by the quadruped robot are obtained, and the auditory modal data and olfactory modal data are used to establish a linkage anomaly; and the inspection anomaly is reported according to the linkage anomaly and the dangerous behavior level score. In other words, through the collaborative work of drones and quadruped robots, the global perspective of the drone and the local perspective of the quadruped robot are integrated, the parallax caused by the difference in perspective is eliminated, the video sequence of the fused perspective is output, the multi-perspective recognition network is called to identify human behavior and actions, and the dangerous behavior level scoring is performed in combination with scene perception. The auditory and olfactory modal data of the quadruped robot are obtained, and the linked anomaly recognition is established, which provides more dimensions for detecting potential dangers, improves the ability to identify potential dangers, ensures that potential risks are discovered and dealt with in a timely manner, and effectively improves the inspection efficiency and safety of power stations.

[0108] Example 2: Based on the same inventive concept as the above-mentioned Example 1, this application also provides a multi-modal collaborative sensing high-risk operation inspection system for power stations. Figure 2 ,include:

[0109] The collaborative tracking module 11 is used to activate the UAV and the quadruped robot to perform collaborative path tracking of the power plant operation task after the power plant operation task is started; the video acquisition module 12 is used to start the video acquisition units of the UAV and the quadruped robot, perform video data acquisition during the collaborative path tracking process, and establish a synchronous video stream; the synchronous alignment module 13 is used to use an external synchronous trigger signal to time-align the time-anchored synchronous video stream after the synchronous video stream is timestamp anchored, and use the global reference coordinate system to perform multi-perspective alignment of the synchronous video stream, and output a video sequence of a fused perspective; the danger level assessment module 14 is used to input the video sequence of the fused perspective into the multi-perspective action behavior recognition network to establish a danger behavior level score; the linkage anomaly establishment module 15 is used to obtain the auditory modal data and olfactory modal data collected by the quadruped robot, and use the auditory modal data and olfactory modal data to establish a linkage anomaly; the inspection anomaly assessment module 16 is used to report an inspection anomaly based on the linkage anomaly and the danger behavior level score.

[0110] Furthermore, the collaborative tracking module 11 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: establish a temporal spatial displacement path based on the power plant operation task; use the temporal spatial displacement path to perform associated scene calls on the power plant scene to establish a spatial scene and a ground scene; perform path optimization of the temporal spatial displacement path in the spatial scene and the ground scene to establish a collaborative path.

[0111] Furthermore, the collaborative tracking module 11 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: obtain equipment data of the video acquisition units of the drone and the quadruped robot, and configure the tracking distance influence constraint according to the equipment data; under the tracking distance influence constraint, perform drone obstacle avoidance tracking optimization with the spatial scene to establish a first optimization path; under the tracking distance influence constraint, perform quadruped robot motion stability balance optimization with the ground scene to establish a second optimization path; and establish a collaborative path with the first optimization path and the second optimization path.

[0112] Furthermore, the video acquisition module 12 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: perform video preprocessing of video data acquisition results, the video preprocessing including motion compensation, image sharpening, edge enhancement, and overexposure marking; and establish a synchronous video stream based on the video preprocessing.

[0113] Furthermore, the video acquisition module 12 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: record the acquisition posture at each acquisition node, and use the acquisition posture to generate a time-series shooting perspective; input the time-series shooting perspective into the perspective correction channel, and establish the synchronous video stream based on the perspective correction channel and the video preprocessing.

[0114] Furthermore, the hazard level assessment module 14 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: construct a unified spatial coordinate system, extract image features from video sequences of different perspectives, and project them into the unified spatial coordinate system; call the scene skeleton extraction sub-channel of the multi-perspective action behavior recognition network to perform feature confidence evaluation of the video sequence projected to the unified spatial coordinate system, construct an anchor point set, and generate a sparse point cloud with the anchor point set; call the dense depth estimation layer of the multi-perspective action behavior recognition network to execute feature depth data of the video sequence projected to the unified spatial coordinate system; project the sparse point cloud to the feature depth data to perform point-depth fusion, and complete local scene reconstruction and annotation; establish a hazard behavior level score based on the local scene reconstruction and annotation.

[0115] Furthermore, the hazard level assessment module 14 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: call the scene and character perception layer of the multi-view action behavior recognition network, perform power plant scene and character perception segmentation within the local scene, and establish a perception segmentation result; use the perception segmentation result to perform character behavior action perception in the video sequence, and configure scene perception; use the character behavior action perception and scene perception to perform hazard behavior level scoring under scene interaction.

[0116] Furthermore, the hazard level assessment module 14 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: establish a micro-motion recognition supplement mechanism; use the micro-motion recognition supplement mechanism to identify human jitter and abnormal pause micro-motions in video sequences; and add micro-motion recognition results to human behavior and action perception.

[0117] Furthermore, the linkage abnormality establishment module 15 in the multi-modal collaborative perception power plant high-risk operation inspection system is also used to: after executing auditory modality and olfactory modality modeling, perform abnormal feature extraction on the auditory modality data and olfactory modality data to establish abnormal feature extraction results; establish a linkage trigger rule base, perform linkage trigger identification of the linkage trigger rule base based on the abnormal feature extraction results, and establish a linkage trigger identification result; and use the linkage trigger identification result to establish a linkage abnormality.

[0118] Furthermore, the inspection anomaly assessment module 16 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: perform trigger upgrade evaluation of behavior in abnormal scenarios based on the linkage anomaly and the dangerous behavior level score, and generate a trigger upgrade evaluation result; and report the inspection anomaly based on the trigger upgrade evaluation result.

[0119] Furthermore, the inspection anomaly assessment module 16 in the multimodal collaborative perception power plant high-risk operation inspection system is also used to: use the linkage anomaly to activate the drone to perform perspective transfer positioning, and establish a perspective transfer positioning result; perform video source tracing and identification based on the perspective transfer positioning result, and generate a source tracing identification result; and update the linkage anomaly according to the source tracing and identification result.

[0120] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. Figure 1The multimodal collaborative perception power plant high-risk operation inspection method and specific examples in Example 1 are also applicable to the multimodal collaborative perception power plant high-risk operation inspection system of this embodiment. Through the above detailed description of the multimodal collaborative perception power plant high-risk operation inspection method, those skilled in the art can clearly understand the multimodal collaborative perception power plant high-risk operation inspection system of this embodiment, so for the sake of brevity of the specification, it will not be described in detail here.

[0121] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

[0122] Obviously, for those skilled in the art, several improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A multi-modal collaborative sensing inspection method for high-risk power plant operations, characterized by: include: After the power plant operation mission is initiated, the UAV and quadruped robot are activated to perform collaborative path tracking for the power plant operation mission; Start the video acquisition units of the UAV and quadruped robot to perform video data acquisition during the collaborative path tracking process and establish synchronized video streams; After the synchronized video stream is timestamp anchored, the time-anchored synchronized video stream is time-aligned using an external synchronization trigger signal, and the multi-view alignment of the synchronized video stream is performed using a global reference coordinate system, and a video sequence of a fused view is output; The video sequence of the fused perspective is input into the multi-perspective action behavior recognition network to establish a dangerous behavior level score; Acquiring auditory modal data and olfactory modal data collected by the quadruped robot, and establishing linkage anomalies using the auditory modal data and olfactory modal data; Report inspection anomalies based on the linkage anomaly and the dangerous behavior level score; The method of inputting the fused perspective video sequence into a multi-perspective action recognition network to establish a dangerous behavior level score includes: Constructing a unified spatial coordinate system, extracting image features from video sequences of different perspectives, and projecting the extracted images into the unified spatial coordinate system; Invoking the scene skeleton extraction subchannel of the multi-view action recognition network, performing feature confidence evaluation on the video sequence projected into a unified spatial coordinate system, constructing an anchor point set, and generating a sparse point cloud using the anchor point set; Call the dense depth estimation layer of the multi-view action recognition network to perform feature depth data projection of the video sequence into a unified spatial coordinate system; Projecting the sparse point cloud onto feature depth data for point-depth fusion to complete local scene reconstruction and annotation; Establish a dangerous behavior level score based on local scene reconstruction and annotation.

2. A multimodal collaborative sensing high-risk power plant operation inspection method according to claim 1, characterized in that: The step of establishing a dangerous behavior level score based on local scene reconstruction and annotation includes: Call the scene and person perception layer of the multi-view action recognition network to perform perception segmentation of the power station scene and people in the local scene and establish the perception segmentation result; Using the perception and segmentation results to perceive human behavior in the video sequence and configure scene perception; The character behavior and action perception and scene perception are used to score the level of dangerous behavior under scene interaction.

3. A multimodal collaborative sensing high-risk power plant operation inspection method according to claim 2, characterized in that: The method of using the perception segmentation result to perceive the human behavior in the video sequence includes: Establish a supplementary mechanism for micro-movement recognition; The micro-motion recognition supplement mechanism is used to recognize the shaking and abnormal pause micro-motions of the characters in the video sequence; Add micro-motion recognition results to human behavior perception.

4. The multimodal collaborative sensing high-risk operation inspection method for power plants according to claim 1, characterized in that: The collaborative path tracking of activating the UAV and the quadruped robot to perform power station operation tasks includes: Establishing a temporal spatial displacement path according to the power station operation task; Using the temporal spatial displacement path, the power station scene is associated with the scene call to establish a space scene and a ground scene; Optimizing the temporal spatial displacement path is performed in the spatial scene and the ground scene to establish a collaborative path.

5. A multi-modal collaborative sensing high-risk operation inspection method for power plants according to claim 4, characterized in that: The performing of path optimization of the temporal spatial displacement path in the space scene and the ground scene to establish a collaborative path includes: Obtain device data of the video acquisition units of the drone and quadruped robot, and configure tracking distance impact constraints based on the device data; Under the influence constraint of tracking distance, the obstacle avoidance tracking optimization of the UAV is performed in the spatial scene to establish a first optimization path; Under the influence constraint of tracking distance, the motion stability balance of the quadruped robot is optimized in the ground scene to establish a second optimization path; A collaborative path is established using the first optimizing path and the second optimizing path.

6. The multimodal collaborative sensing high-risk operation inspection method for power plants according to claim 1, characterized in that: The establishing of linkage abnormality by using the auditory modality data and the olfactory modality data includes: After performing auditory modality and olfactory modality modeling, abnormal feature extraction is performed on the auditory modality data and the olfactory modality data to establish abnormal feature extraction results; Establishing a linkage trigger rule library, performing linkage trigger identification of the linkage trigger rule library based on the abnormal feature extraction result, and establishing a linkage trigger identification result; The linkage trigger identification result is used to establish a linkage anomaly.

7. The multimodal collaborative sensing high-risk power plant operation inspection method according to claim 6, characterized in that: The reporting of inspection anomalies according to the linkage anomaly and the dangerous behavior level score includes: Performing a trigger upgrade evaluation of the behavior in the abnormal scenario based on the linkage anomaly and the dangerous behavior level score, and generating a trigger upgrade evaluation result; Report inspection anomalies based on the trigger upgrade evaluation results.

8. The multi-modal collaborative sensing high-risk operation inspection method for power plants according to claim 7, characterized in that: Before reporting the inspection anomaly based on the linkage anomaly and the dangerous behavior level score, the method includes: Utilizing the linkage anomaly to activate the drone for perspective shift positioning, and establishing a perspective shift positioning result; Perform video source identification based on the perspective shift positioning result to generate a source identification result; The linkage anomaly is updated according to the traceability identification result.

9. The multimodal collaborative sensing high-risk operation inspection method for power plants according to claim 1, characterized in that: The step of performing video data acquisition and establishing a synchronized video stream during the collaborative path tracing process includes: Perform video preprocessing of the video data acquisition results, wherein the video preprocessing includes motion compensation, image sharpening, edge enhancement, and overexposure marking; A synchronized video stream is established based on the video preprocessing.

10. The multi-modal collaborative sensing high-risk operation inspection method for power plants according to claim 9, characterized in that: The establishing of a synchronous video stream based on the video preprocessing comprises: Recording the acquisition posture at each acquisition node, and using the acquisition posture to generate a time-series shooting perspective; The time-series shooting perspective is input into a perspective correction channel, and the synchronous video stream is established according to the perspective correction channel and the video preprocessing.

11. A multi-modal collaborative sensing high-risk operation inspection system for power plants, characterized by: The method for inspecting high-risk operations of a power plant using multimodal collaborative sensing is implemented according to any one of claims 1 to 10. The method comprises: The collaborative tracking module is used to activate the collaborative path tracking of the UAV and quadruped robot to perform the power station operation task after the power station operation task is started; The video acquisition module is used to start the video acquisition units of the UAV and quadruped robot, perform video data acquisition during the collaborative path tracking process, and establish synchronized video streams; A synchronization alignment module is used to, after the synchronized video stream is timestamp-anchored, perform time alignment of the time-anchored synchronized video stream using an external synchronization trigger signal, and perform multi-view alignment of the synchronized video stream using a global reference coordinate system, and output a video sequence of a fused view; The danger level assessment module is used to input the video sequence of the fused perspective into the multi-perspective action recognition network to establish a dangerous behavior level score; A linkage anomaly establishment module is used to obtain auditory modal data and olfactory modal data collected by the quadruped robot, and establish a linkage anomaly using the auditory modal data and olfactory modal data; The inspection anomaly assessment module is used to report inspection anomalies based on the linkage anomaly and the dangerous behavior level score.

Citation Information

Patent Citations

  • Air-ground cooperative intelligent inspection robot and inspection method

    CN111300372A

  • Cloud-side collaborative 1+6 + N intelligent inspection system and method for intelligent thermal power plant

    CN118644906A

  • Power distribution network equipment hidden danger detection method based on large-view-field spliced video data and related device

    CN120013921A