Multi-mode collaborative awareness power station high-risk operation inspection method and system

Through the collaborative work of drones and quadruped robots, visual, auditory and olfactory data are integrated to identify potential hazards in high-risk operations of the power station, solving the problem of inefficiency under a single perception mode, and achieving more efficient and safe patrols.

CN120236249AActive Publication Date: 2025-07-01BEIJING HUADIAN TIANREN ELECTRIC POWER CONTROL TECH

Patent Information

Application Number
CN202510726691.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Due to the limitations of a single perception mode in the prior art, it is difficult to fully identify potential hazards in high-risk operations of power plants in dynamic environments, resulting in inefficient inspections.

Method used

The multimodal collaborative perception method is adopted to work collaboratively by drone and quadruped robot, integrating video data from global perspectives and local perspectives, combining auditory and olfactory modal data to establish dangerous behavior level scores and linkage abnormality recognition.

Benefits of technology

It improves the ability to identify potential hazards, ensures timely detection and handling of potential risks, and improves the efficiency and safety of power station inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236249A_ABST
    Figure CN120236249A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode cooperative sensing power station high-risk operation inspection method and system, and relates to the technical field of video recognition, and the method comprises the steps: activating an unmanned plane and a quadruped robot after a power station operation task is started; starting a video acquisition unit, and establishing a synchronous video stream; carrying out fusion alignment with the global reference coordinate system through an external synchronization signal; inputting the video sequence of the fused view angle into a multi-view angle action behavior recognition network, and establishing a dangerous behavior grade score; auditory data and olfactory data of the quadruped robot are obtained, and linkage abnormity is established; and polling abnormity is reported according to linkage abnormity and dangerous behavior grade scores. Through the method and the device, the technical problem of low inspection efficiency caused by difficulty in comprehensively identifying potential risks in a dynamic environment due to limitation of a single sensing mode is solved, and the inspection efficiency of a power station is improved by fusing multi-mode data, timely finding and processing the potential risks and improving the accuracy of video and audio identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video recognition technology, and in particular, to a multi-modal collaborative perception method and system for high-risk operation inspection in power plants. Background Art

[0002] High-risk operations in power plants usually involve extreme environments such as high voltage, high temperature, and toxic media. Subtle operation errors or equipment abnormalities may trigger serious accidents such as electric shock, explosion, and fall. In recent years, with the development of intelligent technologies, power plant inspections have gradually transformed from manual inspections to automated and intelligent inspections. Currently, common inspection methods usually rely on a single perception method, such as vision or sensor networks, etc., which can detect high-risk illegal operations and potential equipment failures to a certain extent. However, these single perception means have significant limitations. Especially in a dynamic and complex high-risk environment like a power plant, it is often difficult to comprehensively and accurately identify all potential danger signals, which not only limits the improvement of power plant inspection efficiency but also reduces the safety of the inspection process.

[0003] In summary, there is a technical problem in the prior art that due to the limitations of a single perception modality, it is difficult to comprehensively identify potential dangers in a dynamic environment, potential dangers are easily missed, resulting in low inspection efficiency. Summary of the Invention

[0004] The purpose of this application is to provide a multi-modal collaborative perception method and system for high-risk operation inspection in power plants to solve the technical problem in the prior art that due to the limitations of a single perception modality, it is difficult to comprehensively identify potential dangers in a dynamic environment, potential dangers are easily missed, resulting in low inspection efficiency.

[0005] In a first aspect, the present application provides a method for inspecting high-risk operations in a power station with multimodal collaborative perception. The method for inspecting high-risk operations in a power station with multimodal collaborative perception is implemented through a system for inspecting high-risk operations in a power station with multimodal collaborative perception. Among them, the method for inspecting high-risk operations in a power station with multimodal collaborative perception includes: after the power station operation task is started, activating the drone and the quadruped robot to perform collaborative path tracking of the power station operation task; starting the video acquisition units of the drone and the quadruped robot, performing video data acquisition during the collaborative path tracking, and establishing a synchronous video stream; after timestamp anchoring of the synchronous video stream, using an external synchronous trigger signal to perform time alignment of the synchronized video stream with time anchoring, and using a global reference coordinate system to perform multi-view alignment of the synchronous video stream, and outputting a video sequence with a fused view; inputting the video sequence with the fused view into a multi-view action behavior recognition network to establish a risk behavior level score; obtaining the auditory modality data and olfactory modality data collected by the quadruped robot, and establishing a linkage anomaly using the auditory modality data and olfactory modality data; reporting an inspection anomaly according to the linkage anomaly and the risk behavior level score.

[0006] Optionally, establish a temporal-spatial displacement path according to the power station operation task; use the temporal-spatial displacement path to call associated scenes of the power station scene, and establish a spatial scene and a ground scene; perform follow-up path optimization of the temporal-spatial displacement path under the spatial scene and the ground scene, and establish a collaborative path.

[0007] Optionally, obtain the device data of the video acquisition units of the drone and the quadruped robot, and configure the tracking distance influence constraint according to the device data; under the tracking distance influence constraint, perform obstacle avoidance tracking optimization of the drone with the spatial scene to establish a first optimization path; under the tracking distance influence constraint, perform motion stability balance optimization of the quadruped robot with the ground scene to establish a second optimization path; establish a collaborative path with the first optimization path and the second optimization path.

[0008] Optionally, construct a unified space coordinate system, extract image features from video sequences with different views, and project them into the unified space coordinate system; call the scene skeleton extraction sub-channel of the multi-view action behavior recognition network to perform feature confidence evaluation of the video sequence projected into the unified space coordinate system, construct an anchor point set, and generate a sparse point cloud with the anchor point set; call the dense depth estimation layer of the multi-view action behavior recognition network to perform feature depth data of the video sequence projected into the unified space coordinate system; project the sparse point cloud into the feature depth data for point-depth fusion to complete local scene reconstruction annotation; establish a risk behavior level score according to the local scene reconstruction annotation.

[0009] Optionally, call the scene and character perception layer of the multi-view action behavior recognition network to perform power station scene and character perception segmentation within a local scene, and establish perception segmentation results; use the perception segmentation results to perform character behavior action perception under a video sequence, and configure scene perception; use the character behavior action perception and scene perception to perform risk behavior level scoring under scene interaction.

[0010] Optionally, establish a micro-action recognition supplement mechanism; use the micro-action recognition supplement mechanism to recognize character jitter and abnormal pause micro-actions under a video sequence; add the micro-action recognition results to the character behavior action perception.

[0011] Optionally, after performing auditory modality and olfactory modality modeling, extract abnormal features from the auditory modality data and olfactory modality data, and establish abnormal feature extraction results; establish a linkage trigger rule library, and perform linkage trigger recognition of the linkage trigger rule library based on the abnormal feature extraction results to establish linkage trigger recognition results; use the linkage trigger recognition results to establish linkage anomalies.

[0012] Optionally, perform a trigger upgrade evaluation of the behavior in an abnormal scene based on the linkage anomaly and the risk behavior level score to generate a trigger upgrade evaluation result; report a patrol inspection anomaly according to the trigger upgrade evaluation result.

[0013] Optionally, use the linkage anomaly to activate a drone for perspective transfer positioning, and establish perspective transfer positioning results; perform video traceability recognition based on the perspective transfer positioning results to generate traceability recognition results; update the linkage anomaly according to the traceability recognition results.

[0014] Optionally, perform video preprocessing on the video data acquisition results, and the video preprocessing includes motion compensation, image sharpening, edge enhancement, and overexposure marking; establish a synchronous video stream based on the video preprocessing.

[0015] Optionally, record the acquisition pose at each acquisition node, and use the acquisition pose to generate a time-series shooting perspective; input the time-series shooting perspective into a perspective correction channel, and establish the synchronous video stream according to the perspective correction channel and the video preprocessing.

[0016] Second aspect, the present application also provides a multi-modal collaborative perception power plant high-risk operation inspection system for performing a multi-modal collaborative perception power plant high-risk operation inspection method as described in the first aspect. Wherein, the multi-modal collaborative perception power plant high-risk operation inspection system includes: a collaborative tracking module, configured to activate a drone and a quadruped robot to perform collaborative path tracking of the power plant operation task after the power plant operation task is started; a video acquisition module, configured to start the video acquisition units of the drone and the quadruped robot, perform video data acquisition during the collaborative path tracking, and establish a synchronized video stream; a synchronization alignment module, configured to perform time alignment of the synchronized video stream with time anchoring using an external synchronization trigger signal after time stamping the synchronized video stream, and perform multi-view alignment of the synchronized video stream using a global reference coordinate system, and output a video sequence with a fused view; a danger level assessment module, configured to input the video sequence with the fused view into a multi-view action behavior recognition network to establish a danger behavior level score; a linkage anomaly establishment module, configured to obtain the auditory modality data and olfactory modality data collected by the quadruped robot, and establish a linkage anomaly using the auditory modality data and the olfactory modality data; an inspection anomaly assessment module, configured to report an inspection anomaly according to the linkage anomaly and the danger behavior level score.

[0017] One or more technical solutions provided in the present application have at least the following beneficial effects: After the power plant operation task is started, activate a drone and a quadruped robot to perform collaborative path tracking of the power plant operation task; start the video acquisition units of the drone and the quadruped robot, perform video data acquisition during the collaborative path tracking, and establish a synchronized video stream; after time stamping the synchronized video stream, perform time alignment of the synchronized video stream with time anchoring using an external synchronization trigger signal, and perform multi-view alignment of the synchronized video stream using a global reference coordinate system, and output a video sequence with a fused view; input the video sequence with the fused view into a multi-view action behavior recognition network to establish a danger behavior level score; obtain the auditory modality data and olfactory modality data collected by the quadruped robot, and establish a linkage anomaly using the auditory modality data and the olfactory modality data; report an inspection anomaly according to the linkage anomaly and the danger behavior level score. That is to say, through the collaborative work of the drone and the quadruped robot, fuse the global view of the drone and the local view of the quadruped robot, eliminate the parallax caused by the perspective difference, output a video sequence with a fused view, call the multi-view recognition network to identify the human behavior actions, perform a danger behavior level score in combination with scene perception, obtain the auditory and olfactory modality data of the quadruped robot, establish a linkage anomaly recognition, provide more dimensions for detecting potential dangers, improve the ability to identify potential dangers, ensure timely discovery and handling of potential risks, and effectively improve the inspection efficiency and safety of the power plant. Description of the Drawings

[0018] Figure 1 This is a schematic flow chart of a multi-modal collaborative perception method for high-risk operation inspection in a power station according to the present application; Figure 2 This is a schematic structural diagram of a multi-modal collaborative perception system for high-risk operation inspection in a power station according to the present application.

[0019] Explanation of reference numerals: collaborative tracking module 11, video acquisition module 12, synchronization alignment module 13, danger level assessment module 14, linkage anomaly establishment module 15, inspection anomaly assessment module 16. Detailed implementation manners

[0020] By providing a multi-modal collaborative perception method and system for high-risk operation inspection in a power station, the present application solves the technical problem in the prior art that due to the limitations of a single perception modality, it is difficult to comprehensively identify potential dangers in a dynamic environment, potential dangers are easily missed, resulting in low inspection efficiency. Through the collaborative work of an unmanned aerial vehicle (UAV) and a quadruped robot, the global perspective of the UAV and the local perspective of the quadruped robot are fused to eliminate the parallax caused by the perspective difference, and a video sequence of the fused perspective is output. A multi-perspective recognition network is called to identify human behavior actions, and the danger behavior level scoring is performed in combination with scene perception. The auditory and olfactory modality data of the quadruped robot are obtained to establish linkage anomaly recognition, providing more dimensions for detecting potential dangers, improving the ability to identify potential dangers, ensuring the timely discovery and handling of potential risks, and effectively improving the inspection efficiency and safety of the power station.

[0021] Example 1. Please refer to the attached Figure 1 , the present application provides a multi-modal collaborative perception method for high-risk operation inspection in a power station, which specifically includes the following steps: S100: After the power station operation task is started, activate the UAV and the quadruped robot to perform collaborative path tracking of the power station operation task.

[0022] Further, S100 of the present application includes: establishing a time-sequence space displacement path according to the power station operation task; using the time-sequence space displacement path to perform associated scene calls on the power station scene to establish a space scene and a ground scene; performing follow-up path optimization of the time-sequence space displacement path under the space scene and the ground scene to establish a collaborative path.

[0023] Further, the present application further includes the following steps: obtaining the device data of the video acquisition units of the UAV and the quadruped robot, and configuring the tracking distance influence constraint according to the device data; under the tracking distance influence constraint, performing obstacle avoidance tracking optimization of the UAV with the space scene to establish a first optimization path; under the tracking distance influence constraint, performing motion stability balance optimization of the quadruped robot with the ground scene to establish a second optimization path; establishing a collaborative path with the first optimization path and the second optimization path.

[0024] Specifically, according to the specific operation tasks of the power station, including task requirements, environmental characteristics, and work objectives, a temporal-spatial displacement path is designed. Path planning is carried out according to the specific requirements of the inspection task to ensure full coverage of the inspection area. The temporal-spatial displacement path refers to the movement trajectory of the unmanned aerial vehicle (UAV) and the quadruped robot from the starting point to the ending point in the power station operation scenario in a certain time sequence, considering the relationship between the spatial dimension (such as the specific position where the UAV or quadruped robot is located) and the time dimension (the position at each moment on the path). For example, assume that the UAV needs to perform inspection tasks at the top of the power station, and the quadruped robot is responsible for ground inspection. Calculate the flight path of the UAV from the starting point to the ending point within the given time limit, and at the same time calculate how the quadruped robot bypasses obstacles for path tracking. The UAV and the quadruped robot need to coordinate and complete tasks in different areas within the same time period.

[0025] Using the established temporal-spatial displacement path, relevant scene calls are made for the power station scene, and the corresponding spatial scene and ground scene data are called. The spatial scene refers to the virtual environment or physical space in the three-dimensional space of the power station, involving different heights, positions, equipment states, etc. For example, the spatial environment during the UAV flight path planning, considering factors such as the specific positions of the equipment in the power station and aerial obstacles. The ground scene mainly refers to the two-dimensional environment on the ground in the power station or where the robot travels, usually involving the planning of factors such as ground obstacles and passage widths. During the execution process, the UAV and the robot will respectively call the flight and ground scenes according to the task requirements, and automatically switch and optimize the path according to the equipment state and task needs.

[0026] Obtain the device data of the video acquisition units (such as cameras, lidar, infrared sensors, etc.) on the UAV and the quadruped robot, including the performance parameters of the devices, such as camera resolution, field of view angle, maximum detection distance, etc. Configure the tracking distance impact constraint according to the device data, that is, the minimum safe distance between devices (such as the UAV or quadruped robot) during the task execution or the specific distance limit during the task execution, to ensure that the devices do not interfere with each other during the task execution, and avoid collisions or interference due to being too close, thus affecting the smooth progress of the task. If the UAV is equipped with a camera with a narrow field of view angle, then the tracking distance needs to be set relatively close to ensure that the target object can be completely captured in the video. On the contrary, if the camera has a wide-angle lens, the tracking distance can be set relatively far. The tracking distance impact constraint refers to setting the minimum safe distance between devices during the path planning process to ensure that the path planning result takes into account the distance between devices (such as the UAV and the quadruped robot), avoiding collisions or getting too close during the task execution, helping the cooperation between devices, and ensuring their safe operation in the common task.

[0027] Under the influence of the tracking distance constraint, according to the spatial scenario of the power station, path optimization is carried out for the UAV. During this optimization process, it is necessary to avoid obstacles on the flight path while ensuring compliance with the mission requirements. First of all, the safety distances between the UAV and the quadruped robot, as well as between the UAV and other equipment in the power station, must meet the equipment requirements, that is, the minimum safety distance that must meet the tracking distance influence constraint. In the 3D virtual scenario of the power station, multiple flight paths of the UAV from the starting point to the ending point are dynamically calculated, and these flight paths need to avoid colliding with obstacles (such as towers, equipment, etc.).

[0028] The RRT algorithm (Rapidly-exploring Random Tree) is a random tree search algorithm for path planning, especially suitable for path planning problems in high-dimensional spaces. Starting from the starting point, a tree is generated by random sampling, and the branches of the tree are continuously expanded until the tree is connected to the target point. In the traditional RRT algorithm, the safety distance between devices is not particularly considered during path search. To ensure the safety between devices, the tracking distance influence constraint is added to the RRT algorithm. When expanding each tree node, check the distance between the newly generated path segment and the existing path segments. If the new path segment is too close to other path segments (such as the path of another device), it will not be adopted. At the same time, it is necessary to detect whether the distance between the newly generated path segment and the path of the quadruped robot on the ground meets the tracking distance influence constraint to ensure that a certain safety distance is always maintained between devices, thus avoiding collisions. When it is found that the path violates the distance constraint, the algorithm will re-plan the path and adjust the position of the path segment to ensure that the minimum safety distance is maintained between all path segments.

[0029] During the path planning process, the UAV needs to avoid obstacles in the spatial scenario, including buildings, towers, pipelines, equipment, etc., while ensuring that the path meets the tracking distance influence constraint. Devices such as lidar, cameras, and infrared sensors are used to obtain the position information of obstacles in real time and map it to the spatial scenario. Based on the RRT algorithm, an optimization algorithm is used. While finding a feasible path, through further optimization of the path, it is ensured that the path not only avoids obstacles but is also smoother and more efficient. Considering the possible dynamic obstacles (such as staff, other robots, etc.) in the power station environment, it is necessary to update the position of the obstacles in real time through sensors and dynamically adjust the path. Through path planning, multiple feasible flight paths are generated, and these paths meet the following conditions: avoiding all obstacles in the spatial scenario, meeting the tracking distance influence constraint (maintaining a safe distance between the UAV and other devices), and the path avoiding overly complex turns or difficult-to-execute flight behaviors, etc. The RRT algorithm generates multiple paths through random sampling, and these paths may have different flight routes, but all meet the obstacle avoidance and distance constraints.

[0030] Among multiple feasible paths, the optimal path is selected as the first optimized path. The first optimized path is usually the path with the shortest length to reduce the flight distance and save time; the path with the shortest flight time (the straightest and smoothest flight path) under the same path length; in addition, it is also the path with the lowest flight energy efficiency to avoid high-energy-consuming flight behaviors (such as sharp turns, large altitude changes, etc.). By embedding the tracking distance influence constraint in the RRT algorithm and combining the obstacle avoidance path planning of the spatial scene, the system can generate multiple feasible paths that meet the constraint conditions and select the optimal path to ensure that the minimum safety distance between devices is satisfied, avoiding collisions and path interference, and selecting the shortest path, the minimum flight time, and the path with the lowest energy consumption, thus optimizing the flight efficiency of the unmanned aerial vehicle.

[0031] Similarly, under the tracking distance influence constraint, the motion stability balance optimization of the quadruped robot is carried out in the ground scene. In addition to avoiding ground obstacles, it is also necessary to ensure its motion stability during movement. The quadruped robot needs to maintain balance on complex terrains and avoid tilting or tipping over, especially in the presence of irregular ground. Through path optimization, the stability of the quadruped robot is ensured while avoiding all ground obstacles. The motion stability balance optimization is a path optimization strategy designed for the quadruped robot, aiming to ensure that the robot can maintain stability and balance during movement, including avoiding unstable states such as tilting and tipping over during the robot's movement, especially on complex terrains or irregular ground, and ensuring that the robot can complete tasks smoothly.

[0032] Different from the unmanned aerial vehicle, the path planning of the quadruped robot not only needs to consider obstacle avoidance, but also needs to optimize the gait and adjust the motion strategy to ensure that the robot can maintain stability and perform tasks in complex terrains. Ground sensors (such as lidar, ground cameras, etc.) are used to detect the ground conditions in real time, obtain information such as ground unevenness and slopes, and map them into the ground scene. Based on the robot's motion model, gait planning algorithms (such as control-based gait generation algorithms) are used to optimize the robot's steps to ensure that the robot can pass through complex terrains stably. Specifically, the robot dynamically adjusts the frequency, stride, and posture of the steps according to the ground conditions to prevent falling or losing balance. According to the optimal motion trajectories of the robot's various joints and legs, it is ensured that each step of the robot can maintain balance. For example, in areas with a large slope, the robot may need to reduce the step frequency, adjust the center of gravity position, and reduce the risk of falling.

[0033] The quadruped robot adopts a wheel-leg composite mobile structure, which combines terrain adaptability and moving speed. It can switch between bipedal and quadrupedal modes, better adapt to narrow spaces, and increase the working height. By simulating the execution process of different paths, the motion stability of each path is evaluated, that is, through the analysis of the historical motion data of the quadruped robot, the tilt angle, gait changes, joint loads, etc. of the robot when walking on the path are determined. Select the optimal path from multiple feasible paths as the second optimized path, which not only needs to meet the obstacle avoidance requirements and tracking distance constraints, but also needs to ensure that the robot can maintain stability and efficiency during the movement. Usually, the path that is most suitable for the ground conditions and can best ensure stability is selected as the final path. For example, two feasible paths are planned for the quadruped robot. One path is shorter, but there is irregular ground, and the robot may lose balance due to this; the other path is longer, but the ground is flat, and the robot can pass with a more stable gait. Through comprehensive evaluation, the flat path is selected as the final second optimized path.

[0034] According to the optimized paths of the drone and the quadruped robot (i.e., the first optimized path and the second optimized path), a collaborative path is established to ensure that the two devices do not interfere with each other during the task execution and can efficiently complete the inspection task. The collaborative path refers to the travel path designed when multiple devices (such as drones and quadruped robots) perform tasks simultaneously, enabling these devices to coordinate their operations, avoid conflicts, and maximize the task completion efficiency within limited time and space. For example, if the paths of the drone and the quadruped robot intersect at a certain point, then the collaborative path planning will ensure that they avoid each other before or after this point to avoid collisions.

[0035] By influencing the tracking distance constraints and obstacle avoidance tracking optimization, it is ensured that the drone and the quadruped robot can effectively avoid collisions and interference from obstacles during the task execution. The quadruped robot can maintain balance through motion stability balance optimization, reduce the risk of tipping over and instability in complex terrains. Without interfering with each other, the drone and the quadruped robot successfully inspected multiple devices. The entire inspection process improved the efficiency by 30% compared with the traditional method, and there were no device collisions or path conflicts.

[0036] After the power plant operation task is started, the collaborative path tracking center that activates the drone and the quadruped robot to execute the power plant operation task is started. The task start signal is sent to the control modules of the drone and the quadruped robot, entering the activation state, and their respective power supplies and sensors are started. The central control system, according to the task requirements and real-time environmental data, assigns the first optimized path (drone obstacle avoidance tracking path) and the second optimized path (quadruped robot motion stability balance path) respectively planned for the drone and the quadruped robot to the corresponding drone and quadruped robot. The path information includes waypoints, speeds, obstacle avoidance strategies, etc. After receiving their respective path information, the drone and the quadruped robot start the path tracking mode. The drone uses on-board sensors (such as cameras, lidars, ultrasonic sensors, etc.) to sense the surrounding environment in real time, perform obstacle avoidance operations, and at the same time track the predetermined path; the quadruped robot uses its own sensors and terrain adaptation capabilities to maintain motion stability in complex ground environments and at the same time track the predetermined path. When the drone and the quadruped robot complete all the predetermined path points and collect the required data, the task completion signal is sent back to the central control center, and the central control center sends a recovery instruction. The drone and the quadruped robot return to the starting point or the designated recovery area, waiting for the next task or for maintenance.

[0037] S200: Start the video acquisition units of the drone and the quadruped robot, perform video data acquisition during the collaborative path tracking, and establish a synchronized video stream.

[0038] Specifically, when the drone and the quadruped robot start to execute collaborative path tracking, the video acquisition units of the drone and the quadruped robot are started. The video acquisition unit refers to the sensor devices installed on the drone and the quadruped robot, which are used to capture and record video data in real time, including high-definition cameras, infrared cameras, laser scanners, etc., obtain visual information of the surrounding environment, and convert it into digital signals for subsequent processing. The cameras installed on the drone and the quadruped robot are configured according to the task requirements. For example, the drone is equipped with a high-definition RGB camera and an infrared camera, while the quadruped robot is equipped with a close-range high-definition camera or a 360-degree panoramic camera. For example, in the power plant inspection task, the quadruped robot carries a high-definition RGB camera to detect damage on the surface of the equipment, while the drone is equipped with an infrared camera to detect heat sources or overheated equipment. After being started, these devices begin to acquire video data.

[0039] During the collaborative path tracking process, the drone and the quadruped robot need to simultaneously perform video data collection tasks, and the video capture units they are equipped with will continuously collect video data. For example, the drone flies over the power station along a predetermined flight path while collecting video data, and the quadruped robot moves along the corresponding path on the ground while synchronously collecting video data to establish a synchronized video stream. The video capture unit of each device (drone and quadruped robot) will attach a timestamp to each frame of video data. Through the collaborative work of different devices, environmental data from different angles and perspectives can be obtained, the equipment status of the power station can be comprehensively identified, abnormal statuses (such as equipment overheating, cracks, etc.) on the power station equipment can be quickly identified, and repairs can be carried out in a timely manner, which not only improves the inspection efficiency but also ensures the safe operation of the power station.

[0040] Furthermore, S200 of this application includes: performing video preprocessing on the video data collection results, where the video preprocessing includes motion compensation, image sharpening, edge enhancement, overexposure marking; and establishing a synchronized video stream based on the video preprocessing.

[0041] Record the acquisition pose at each acquisition node, generate a sequential shooting perspective using the acquisition pose; input the sequential shooting perspective into the perspective correction channel, and establish the synchronized video stream based on the perspective correction channel and the video preprocessing.

[0042] Specifically, after the video data is collected, a series of preprocessing steps are required, including motion compensation, image sharpening, edge enhancement, overexposure marking, to ensure the stability of the video quality and the usability of the data. Since the quadruped robot and the drone may move quickly in a dynamic environment, the captured images or videos may exhibit motion blur or jitter. Therefore, motion compensation is needed. According to the motion trajectory of the objects in the image, the motion offset of the camera is estimated and corrected to reduce the blur phenomenon. By tracking feature points (such as corner points, edges, etc.) in the video frames, the motion trajectory of the equipment is estimated. According to the estimated motion information, through geometric transformations (such as affine transformation, perspective transformation, etc.), the video frames are aligned and corrected to remove the blur phenomenon caused by the equipment movement. Through motion compensation algorithms, such as the stabilization method based on feature matching, the motion of each frame of the image is corrected.

[0043] To enhance the details in video images, especially small objects or equipment components at a distance, image sharpening is used to increase the contrast and edge sharpness of the image, making key features more prominent. Image sharpening makes objects and structures in the image more prominent by increasing the contrast in local regions of the image, and is particularly suitable for small target recognition and detail analysis in tasks such as power plant inspection. Edge information of the image is extracted through filters (such as the Laplace operator, etc.) and the intensity of these edges is enhanced, making the object contours more obvious. By increasing the brightness contrast in local regions of the image, the target area becomes more prominent. The sharpened image can better support subsequent target recognition and classification tasks, reducing misjudgments caused by image blurring or insufficient details.

[0044] Edge enhancement makes the structure and details of objects in the video more prominent by enhancing the intensity of contour lines and edges in the image. In a complex power plant environment, enhancing edges helps to clearly identify important elements such as equipment components and human actions. The detected edges are enhanced, usually by increasing the brightness contrast of the edges to make the object edges more obvious. The edges in the image are extracted through the Canny edge detection algorithm and the edges are strengthened.

[0045] When shooting in strong light or high-temperature areas, overexposure may occur in the video. Through overexposure marking technology, these overexposed areas are marked and processed to avoid loss of image information and ensure that the image quality is not affected in key areas. Regions in the image with brightness exceeding the threshold are detected and marked as overexposed areas. Usually, these areas will appear white or very bright and lack details. By adjusting the brightness and contrast, the overexposed areas are repaired to reduce the impact of overexposure, or image inpainting technology is used to fill in the details of these areas. For example, the pixel values of the image are converted into an image histogram. By analyzing the brightness histogram of the image, it is judged which areas have too high brightness, and then whether overexposure exists is judged. The brightness of the overexposed areas is adjusted to reduce the brightness value in the image and restore the details.

[0046] The acquisition pose is recorded at each acquisition node, that is, the position, angle of each video acquisition node, and the pose of the sensor are all recorded. The pose record of the acquisition node usually includes the position (such as GPS coordinates), pose angles (such as pitch angle, roll angle, and yaw angle), etc. UAVs and quadruped robots will continuously change their poses during movement, and each frame of the acquired image will be marked with the corresponding pose information. Through the acquired pose information, the sequential shooting perspective corresponding to each frame of the image is generated to ensure the alignment of the image in time and space. For example, when shooting synchronously in the air and on the ground, record the respective perspective differences to facilitate subsequent time synchronization and spatial alignment of the two video streams. For example, when a quadruped robot captures a video of a device failure during inspection, record the position and pose of the robot and the shooting angle of this video frame. When a UAV shoots from different perspectives, record the relevant data of the UAV according to the same standard and generate the corresponding sequential shooting perspective.

[0047] The sequential shooting perspective refers to the perspective jointly determined by the spatial position and orientation of the device at each acquisition node. According to the movement trajectory, direction, and tilt angle of the device, the camera perspective corresponding to each acquisition node is deduced. That is to say, based on the pose data and time series of the device, the shooting angle and field of view of the video at each moment are deduced to obtain the sequential shooting perspective.

[0048] The perspective correction channel is an image processing process used to unify video frames acquired at different time points or different device angles to a standard perspective or reference perspective. By performing geometric transformations (such as rotation, scaling, perspective transformation, etc.) on the video frames, the perspective deviation caused by the change of the device pose is eliminated. Since the perspectives and angles of the images acquired by different devices at different times may be different, perspective correction is required to align the video frames with different perspectives. The purpose of the perspective correction channel is to ensure that all acquired video frames are aligned in a unified coordinate system or perspective.

[0049] The sequential shooting perspective is used as input information to provide the acquisition position and orientation of the video frame at each moment, and the perspective correction channel adjusts the image based on this information. Specifically, according to the sequential shooting perspective information of each video frame, coordinate system conversion is performed to transfer the video data from the coordinate system at the time of acquisition to a unified spatial coordinate system. Through perspective transformation (perspective matrix), the video frames with different perspectives are adjusted to the standard perspective, so that the images captured by different devices or cameras are aligned. The video frames with different perspectives are reconstructed and aligned according to the same standard perspective to form a synchronized video stream.

[0050] By performing perspective correction on the preprocessed video, the image distortion caused by factors such as device angle differences and different shooting distances is corrected to ensure that the visual effects of the finally generated synchronized video stream are consistent. Through video preprocessing and perspective correction, in a multi-device and multi-perspective environment, the video data from different sources are accurately fused and synchronized, improving the overall accuracy of the analysis and ensuring that the same device components or on-site scenes captured from different perspectives can be presented consistently.

[0051] S300: After timestamp anchoring is performed on the synchronized video stream, the time alignment of the synchronized video stream with time anchoring is carried out using an external synchronization trigger signal, and the multi-perspective alignment of the synchronized video stream is carried out using a global reference coordinate system, and a video sequence with a fused perspective is output.

[0052] Specifically, the video acquisition unit of each device (drone and quadruped robot) attaches a timestamp to each frame of video data, recording the exact time of each frame acquisition, which can be generated by the internal clock of the device or calibrated by an external synchronization signal (such as a GPS clock or an NTP protocol signal). Timestamp anchoring refers to assigning an accurate time mark to each frame of image in the acquired video stream. The external synchronization trigger signal is a signal from an external source, usually used to coordinate the time synchronization between multiple devices (such as drones and quadruped robots), such as a GPS clock signal, an NTP (Network Time Protocol) signal, or other precise synchronization signals, to ensure that the data collected by multiple devices at the same moment can correspond precisely.

[0053] An external triggering device (such as a GPS clock module, an NTP time server, etc.) is used to provide a unified time signal for each device, and the acquisition modules of all devices will start acquiring data according to this unified time signal. The central control center will send a synchronization signal to trigger the drone and the quadruped robot to start acquiring video data simultaneously. After receiving the synchronization signal, the video acquisition unit performs time alignment on the acquired video stream according to the timestamp to ensure that the video data from different perspectives are consistent in time. Time alignment refers to synchronizing the video streams from different devices (such as drones and quadruped robots) according to the timestamp, so that the data collected at the same time point can match, avoiding data errors caused by the asynchronous device clocks.

[0054] The global reference coordinate system is a unified coordinate system used to describe the spatial positions of the entire environment. In a multi-robot system, each device (such as a drone, quadruped robot, etc.) will perform positioning and navigation based on this coordinate system to ensure the unity of the spatial positions of different devices. The global reference coordinate system can usually be a GPS coordinate system or a coordinate system established within the power station. By using the global reference coordinate system, video data from different perspectives is integrated into a unified space. Each device (such as a drone and a quadruped robot) determines its own position and orientation through the global reference coordinate system and adjusts the video stream according to this coordinate system, enabling seamless connection of videos from different perspectives. For example, the quadruped robot captures images on the ground of the power station, while the drone captures images of the tower and the equipment above. Through spatial alignment technology, the video streams from these two perspectives can be seamlessly combined in the same coordinate system.

[0055] After time and space alignment, the video streams from different devices can be fused to form a unified, multi-perspective video sequence, simultaneously presenting data from different perspectives. Task executors can comprehensively understand the status of the equipment in the power station by viewing this fused video. The fused perspective video sequence refers to combining multiple video streams from different perspectives after time and space alignment into a single video sequence, providing the ability to observe the environment from multiple angles simultaneously and helping task executors fully understand the on-site situation.

[0056] By performing timestamp anchoring in the synchronized video stream, time alignment using an external synchronization trigger signal, and multi-perspective alignment using the global reference coordinate system, a fused perspective video sequence is finally output, ensuring a high degree of consistency in time and space for the video data of different devices (drones and quadruped robots), avoiding data errors caused by time asynchrony or spatial misalignment, providing a comprehensive and accurate on-site situation, showing the status of the power station equipment observed from different perspectives, being able to promptly detect equipment failures in the power station (such as overheating, electrical failures, etc.), and immediately taking measures for repair.

[0057] S400: Input the fused perspective video sequence into a multi-perspective action behavior recognition network to establish a dangerous behavior level score.

[0058] Further, the S400 of the present application includes: constructing a unified spatial coordinate system, extracting image features from video sequences with different perspectives and projecting them into the unified spatial coordinate system; calling the scene skeleton extraction sub-channel of the multi-perspective action behavior recognition network to perform feature confidence evaluation on the video sequence projected into the unified spatial coordinate system, constructing an anchor point set, and generating a sparse point cloud with the anchor point set; calling the dense depth estimation layer of the multi-perspective action behavior recognition network to execute the feature depth data of the video sequence projected into the unified spatial coordinate system; projecting the sparse point cloud onto the feature depth data for point-depth fusion to complete the local scene reconstruction annotation; and establishing a dangerous behavior level score based on the local scene reconstruction annotation.

[0059] Further, the present application further includes the following steps: calling the scene person perception layer of the multi-perspective action behavior recognition network to perform power station scene and person perception segmentation within the local scene and establish the perception segmentation result; using the perception segmentation result to perform person behavior action perception under the video sequence and configure scene perception; and using the person behavior action perception and scene perception to perform dangerous behavior level scoring under scene interaction.

[0060] Specifically, according to the actual geographical location and equipment layout of the power station, a unified spatial coordinate system is determined, and spatial data from different sources or perspectives (i.e., image data collected by drones and quadruped robots) are mapped into the same coordinate system for comprehensive analysis of data from different perspectives. Image features are extracted from different perspective videos collected by each device (such as drones and quadruped robots), and representative feature points are identified. These points can be used to describe the image content and establish corresponding relationships between different perspectives. Image feature extraction is to identify and extract key visual information from images, such as edges, textures, corner points, etc. Image feature extraction algorithms include SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), etc.

[0061] After the image feature extraction is completed, the image feature points extracted from different perspectives (such as videos collected by drones and quadruped robots) will be mapped to a unified spatial coordinate system, and the image features will be transformed from the image space (pixel coordinates) to the actual three-dimensional spatial coordinates. Through the internal and external parameters of the camera (such as focal length, lens distortion, camera position and orientation, etc.), the image is calibrated and corrected. Through the stereo matching algorithm, the position of each feature point in the three-dimensional space is calculated based on the images from multiple perspectives, thus realizing the spatial projection of the image. For example, images from different angles are obtained from drones and quadruped robots. After image feature extraction, the coordinates of these feature points in the three-dimensional space are calculated through the stereo matching algorithm and mapped to a unified power plant digital twin coordinate system. That is to say, although the feature points were initially taken from different angles, their positions in the unified coordinate system are accurately corresponding.

[0062] SIFT is an algorithm for extracting local features from images, which can identify key points in the image (such as corner points, edges, etc.). These key points have scale invariance, rotation invariance, and illumination invariance; SURF is an improved version of the SIFT algorithm, mainly accelerating the feature point extraction process by improving the calculation speed of the algorithm; key points are found by calculating the local features of the image, and techniques such as Gaussian blur and scale space are used to ensure the stability of the features, and information such as the coordinates, directions, and scales of the key points are extracted. After feature points are extracted from video frames from different perspectives, the feature points from different perspectives are corresponded through the feature matching algorithm. Since there are often noises, mismatches, and outliers in video frames, the RANSAC algorithm can effectively remove the incorrect matches that do not conform to the model and improve the accuracy of the matching. Randomly select two pairs of feature points, calculate the transformation relationship between them (such as rotation, translation), use the transformation relationship to match other feature points, and calculate the error between the transformed feature points and the feature points in the original image; retain the matching points with the smallest error as the inliers that conform to the transformation model; repeat the random point selection and model fitting until the best transformation model is found. After the feature matching is completed, these matched feature points are mapped to a unified spatial coordinate system, and the internal parameters (such as focal length, principal point, etc.) and external parameters (such as the spatial position and orientation of the camera) of the camera are obtained through camera calibration, and the image coordinates are converted into actual spatial coordinates.

[0063] Call the scene skeleton extraction sub-channel of the multi-view action behavior recognition network, project the skeleton features and scene information in each video frame into a unified spatial coordinate system, and evaluate the confidence of these features. The feature confidence evaluation mainly determines the reliability of each feature point based on image quality, the stability of view matching, and the degree of background interference. Through the scene skeleton extraction sub-channel, extract the skeleton feature map (keypoint positions, joint connection information, etc.) of each video frame. Each video camera (such as the cameras on a quadruped robot and a drone) has a set of internal and external parameters. The internal parameters include information such as focal length and optical center, and the external parameters include the position and orientation of the camera relative to the global coordinate system. Through the internal and external camera parameter matrices, convert the skeleton feature points (positions in the image coordinate system) into three-dimensional coordinates in the unified spatial coordinate system.

[0064] The feature confidence evaluation quantifies the reliability of the skeleton features extracted from each video frame by quantifying according to image quality, feature stability, matching stability, depth information of feature points, etc. In low-light or high-light conditions, the feature points in the video frame may be unstable, so in this case, the feature points will be assigned a lower confidence. Image noise or blurring will cause inaccurate feature extraction, so it is necessary to adjust the confidence according to the image quality score. Images taken under fast movement may cause blurring and affect the extraction of features. Evaluate the stability of the skeleton features in each frame of the image to determine whether it is reliable. If the view changes too much in the image, it may cause the positions of feature points of the same object to be inconsistent in different frames, and adjust the confidence of the feature points according to the amplitude of the view change. If in multiple views, the matching relationship of the same feature point is unstable or has a large deviation, the confidence will be reduced. If there is depth information (such as using stereo vision or a depth camera), use the depth map to help determine the position and stability of the feature points in three-dimensional space.

[0065] According to the above elements, calculate the confidence of each feature point, and screen out those feature points with higher confidence as the anchor point set, that is, the key and stable features in the scene, which can provide accurate spatial information. Specifically, according to the set confidence threshold, select those feature points with confidence higher than the threshold to form the anchor point set. Through spatial consistency detection (such as detecting whether the feature points in multiple views are located on the same object), further screen out stable anchor points.

[0066] Using the anchor point set and its matching relationship in videos of different views, generate a sparse point cloud, which represents the key positions or feature points in the scene. These points have corresponding coordinates in three-dimensional space and represent the key feature positions on the object surface or in the environment. Compared with a dense point cloud (containing a large number of data points), the sparse point cloud has less data volume, but it can still provide enough information for subsequent 3D reconstruction and analysis.

[0067] Call the dense depth estimation layer of the multi-view action behavior recognition network to execute the feature depth data of the video sequence projected onto a unified spatial coordinate system, which means using the dense depth estimation layer to estimate the depth information of each pixel in the video sequence. Dense depth estimation refers to directly calculating the depth value of each pixel point in the image. Input the video sequence projected onto a unified spatial coordinate system into the density depth estimation layer to obtain the depth map of each video frame, which represents the spatial positions of the equipment and personnel in the power station. The depth values in the depth map are usually represented by gray levels, where darker areas indicate farther distances and brighter areas indicate closer distances.

[0068] Project these depth data from the local coordinate system of each camera view to a unified spatial coordinate system to ensure that the video data from different views can be aligned in a shared coordinate system, thus enabling accurate three-dimensional space reconstruction. By combining the depth value of each pixel with its corresponding view information (such as the orientation and position of the camera), the depth information can be transformed from the local coordinate system to the global coordinate system. For example, assume that in a power station, a drone takes pictures of the area of power equipment from the air, while a quadruped robot also takes pictures of the same area from the ground. The depth data of each view are generated by the depth estimation network and projected into the global coordinate system. In this way, the depth data from both the aerial view of the drone and the ground view of the quadruped robot can be fused under the same unified coordinate system, further helping with the three-dimensional modeling and analysis of the power station equipment.

[0069] Fuse the sparse point cloud with the depth estimation map by combining each point in the sparse point cloud with the corresponding depth information, thereby generating a more accurate three-dimensional scene. Align the data of the sparse point cloud with the depth map, and convert the spatial coordinates of the point cloud to actual three-dimensional coordinates by calculating the spatial position of each point and the depth value of the corresponding pixel. Under the same coordinate system, combine the depth information in the sparse point cloud with that in the depth map, and through weighted fusion, generate a dense point cloud or depth map to complete the local scene reconstruction annotation, and label each part of the scene according to the positions of the recognized equipment or personnel. Point-depth fusion refers to combining the feature points in the sparse point cloud with the feature depth data, and through weighted fusion of the depth values of each feature point, generating a more accurate three-dimensional scene reconstruction, improving the spatial representation of the sparse point cloud and making it more accurate in local scene reconstruction.

[0070] Call the scene and person perception layer of the multi-view action behavior recognition network to segment the power station scene and people within the local scene, identify key scene elements (such as power station equipment, building structures, etc.) and dynamic elements (such as operating personnel) from the video, perform person detection on each video frame through the object detection algorithm, and locate and identify the accurate positions of the operating personnel. The detection algorithm usually outputs a bounding box to mark the position of each operating personnel in the image. In addition to person detection, the scene and person perception layer also needs to segment the static elements (such as power station equipment, buildings, etc.) in the power station scene, divide the image into multiple regions, each region is assigned a label, and the perception segmentation result will generate a segmentation map, where the label of each pixel represents which category the pixel belongs to. The perception segmentation result refers to various types of information segmented through deep learning algorithms within the entire power station scene, including the positions and states of power station equipment and operating personnel.

[0071] Through the perception segmentation result, analyze the actions of the people in the video sequence, including obvious actions (such as standing, walking, squatting, climbing, etc.) and micro-actions (such as jittering, pausing, etc.) of the task. Each behavior type will have a corresponding hazard level assessment, especially when it comes to high-risk operations (such as high-voltage contact, equipment disassembly, etc.). At the same time, configure scene perception. Through the training of different scenes (such as the operation area, equipment area, dangerous area, etc.) of the power station and the behavior of operating personnel, the scene perception model can automatically identify different regions inside the power station and their risk levels. Identify the elements in each scene in real time and determine which regions are high-risk regions. For example, the high-voltage equipment area is marked as a high-risk region, and the normal operation area is marked as a safe region.

[0072] By setting rules for the hazard level (such as contacting high-voltage equipment, entering restricted areas, etc.), automatically evaluate the risk level according to the behavior of the operating personnel and the environmental area where they are located, evaluate the danger of the behavior, and generate a danger behavior level score, such as low risk, medium risk, high risk, etc. For example, an operating personnel invading the live area is a level 1 risk; an operating personnel not wearing a safety helmet is a level 2 risk; an operating personnel entering the high-voltage equipment area, contacting high-voltage equipment, etc. is a level 5 risk, etc. Based on the interactive analysis of behavior and environment, assign a level (such as 1 - 5 points, 1 being low risk, 5 being extremely high risk) to each dangerous behavior according to a series of scoring rules, specifically including analyzing the specific behavior of the operating personnel, such as whether they perform dangerous operations (such as touching electrical equipment, entering dangerous areas, etc.); analyzing the environment where the person is located, such as whether they enter the high-voltage area, contact dangerous substances, etc. Score each interactive behavior according to the type of behavior and the danger level of the environment, and give the overall danger level to help the staff quickly identify dangerous situations and take necessary preventive measures.

[0073] Identify the behavior and actions of personnel by calling a multi - perspective recognition network, combine scene perception to perform a risk behavior level assessment, supplement the micro - motion recognition mechanism to enhance accuracy, assign a risk level to each behavior, improve the inspection efficiency, reduce unnecessary movements, and at the same time ensure coverage of all key areas.

[0074] Furthermore, the present application further includes the following steps: Establish a micro - motion recognition supplement mechanism; use the micro - motion recognition supplement mechanism to recognize micro - motions such as personnel jitter and abnormal pauses in a video sequence; add the micro - motion recognition results to the perception of personnel behavior and actions.

[0075] Specifically, establish a micro - motion recognition supplement mechanism, that is, recognize the subtle movement changes of the operator's body, which usually occur in a short time and are manifested as low - amplitude movements, including personnel jitter (tiny vibrations caused by reasons such as nervousness, cold, mechanical vibration, etc.) and abnormal pauses (the operator no longer moves for a short time at certain moments due to equipment failure, loss of balance, or encountering sudden problems). In order to capture micro - motions, the frame rate of the video acquisition system needs to be relatively high (such as 30 frames per second or higher) to carefully record the tiny actions of the operator, thereby improving the accuracy of micro - motion recognition. Micro - motions are usually subtle changes in a short time, so real - time analysis needs to be carried out on a continuous number of frames (such as 5 - 10 consecutive frames).

[0076] Use consecutive frames in a high - frame - rate video sequence, apply the micro - motion recognition supplement mechanism for subtle motion analysis, and judge whether there is involuntary jitter of the body by calculating the tiny displacements and posture changes of the operator's joints and limbs, that is, the rapid and small - amplitude movements of a certain part of the body (such as hands, head, shoulders) in a short time. Abnormal pauses usually refer to the operator suddenly staying at a certain position for longer than the normal operation time. For example, when the operator is operating equipment, due to nervousness or cold, there is a slight jitter in the hand. Through the micro - motion recognition supplement mechanism, this tiny jitter is captured in the video sequence and marked as a potential "physical discomfort" signal.

[0077] Abnormal pauses may be due to equipment failure, the operator suddenly losing balance, fatigue, or encountering other emergencies. Use time - series analysis methods to judge the speed and position changes of the operator. If the movement speed of the operator is close to zero in a short time and no normal action transitions (such as operation switching, equipment adjustment, etc.) occur, it is considered that there is an abnormal pause.

[0078] Once micro - movements (such as jitters or abnormal pauses) are recognized, they are added to the behavioral action perception data of the person. The recognition results of micro - movements and the behavioral perception results will jointly constitute the current behavioral model of the operator. Micro - movements (such as jitters, pauses, etc.) may reflect the health status of the operator or changes in the equipment status. Therefore, in the behavioral perception model, the existence of micro - movements may affect the behavioral assessment of the operator and further adjust its risk level score.

[0079] By introducing a micro - movement recognition supplement mechanism, subtle movements of the operator (such as jitters, abnormal pauses) are captured, and this information is integrated with the results of human behavioral perception, further enhancing the ability to detect potential risks. By establishing a micro - movement recognition supplement mechanism, minute movements of the operator such as jitters and abnormal pauses are recognized in real - time, and this information is integrated into human behavioral perception, thereby improving the ability to detect potential risks.

[0080] S500: Obtain the auditory modality data and olfactory modality data collected by the quadruped robot, and establish a linkage anomaly using the auditory modality data and the olfactory modality data.

[0081] Furthermore, S500 of this application includes: After performing auditory modality and olfactory modality modeling, extract abnormal features from the auditory modality data and the olfactory modality data to establish an abnormal feature extraction result; establish a linkage trigger rule library, and perform linkage trigger recognition of the linkage trigger rule library based on the abnormal feature extraction result to establish a linkage trigger recognition result; establish a linkage anomaly using the linkage trigger recognition result.

[0082] Specifically, perform auditory modality and olfactory modality modeling. Auditory modality modeling includes that the abnormal features of arc sound / discharge sound are high - frequency sharp / sudden short, and the possible risks are insulation breakdown and poor contact; the abnormal features of vibration noise change are high - frequency change or deviation, and the possible risks are instrument loosening and tool dropping; the abnormal features of non - structural speech are calls for help and abnormal shouts, and the possible risks are worker errors or emergencies. Olfactory modality modeling includes that when the detected gas type is ozone, the abnormal signal feature is an instantaneous increase in concentration, and the possible hidden danger is arc discharge and by - products; when the detected gas type is alkanes / odor residue, the abnormal signal feature is continuous exceeding the standard or fluctuating abnormally, and the possible hidden danger is flammable gas leakage / oil evaporation; when the detected gas type is charred odor residual compounds, the abnormal signal feature is long - tail persistence or repeated change, and the possible hidden danger is equipment short - circuit / high - temperature aging.

[0083] Auditory modality data refers to the sound data collected by the sensors or microphones of the quadruped robot. In a complex environment such as a power station, the sound signal can reflect many key information, such as whether the equipment is operating normally, abnormal equipment noise, shouts of personnel, alarm sounds, etc. Olfactory modality data is the gas information collected by the robot's gas sensors (such as gas sensor arrays). In a high-risk environment like a power station, certain specific gases (such as toxic gases, combustible gases, or chemical odors of electrical equipment) may be early omens of equipment failure, leakage, or fire. The quadruped robot collects this information by installing gas sensors and monitors the gas composition of the surrounding environment in real time.

[0084] The sound signal is converted into a spectrogram through Fourier transform, and information such as frequency, amplitude, and waveform is extracted from it. Based on the time-domain waveform of the audio signal, features such as the amplitude, amplitude, and frequency of the sound are analyzed. By comparing with the known normal sound features, abnormal noises are identified. For example, a sudden high-frequency noise indicates that the equipment has failed. The gas concentration value output by the gas sensor is obtained, and according to the historical data, a gas concentration distribution model within the normal range is established. When the sensor detects a concentration beyond the normal range, it is judged as abnormal. For example, if the gas sensor detects a sudden increase in carbon monoxide concentration (such as reaching 50 ppm), it means that there is a leakage or abnormal combustion of a certain equipment, and further confirmation is required. Features that can reflect the abnormal state are extracted from the sound data and gas data and labeled.

[0085] According to multiple sensor information and condition triggers, a linkage trigger rule library is established. Each rule usually consists of multiple sensor input data and certain specific conditions. When these conditions are met, the corresponding linkage operation is triggered. For example, some examples of the linkage trigger rule library are shown in Table 1: Table 1 Linkage Trigger Rule Library

[0086] Combining the abnormal features of the auditory modality and the olfactory modality, for example, the abnormal noise generated by a certain equipment (auditory abnormality) and the leakage of a certain gas (olfactory abnormality) may occur simultaneously. It is necessary to automatically link according to the set rules and trigger an alarm or perform other emergency operations. For example, when the quadruped robot detects a high-frequency noise near a certain electrical equipment and is accompanied by an increase in the ammonia concentration, the rule library can trigger the linkage trigger recognition according to the condition of high-frequency noise + ammonia concentration exceeding the standard and output a warning of a dangerous event.

[0087] According to the established linkage trigger rule library, identify which trigger condition is met through the abnormal feature extraction results obtained from the auditory modality data and olfactory modality data, and output the linkage trigger recognition result. For example, in a power station environment, after the robot detects abnormal noises and gas leakage signals, the linkage rule is triggered, a warning of "there may be equipment failure and accompanied by leakage risk" is issued, and corresponding safety measures are initiated.

[0088] Based on the linkage trigger recognition result, establish a linkage anomaly, that is, when the abnormal information captured by multiple sensors (such as vision, hearing, and smell) corroborates each other, timely identify potential multi-modal abnormal signals, avoid false negatives caused by the inability of a single perception modality to comprehensively identify risks, and improve the safety and response efficiency during the inspection process. By obtaining the auditory modality data and olfactory modality data collected by the quadruped robot, simulate the auditory and olfactory perception processes to extract useful information from the original data, identify and extract features representing anomalies or abnormal situations, determine when the linkage reaction should be triggered, determine whether there is a linkage anomaly, and take appropriate measures.

[0089] S600: Report the inspection anomaly according to the linkage anomaly and the risk behavior level score.

[0090] Furthermore, S600 of the present application includes: performing a trigger upgrade evaluation on the behavior in the abnormal scenario based on the linkage anomaly and the risk behavior level score to generate a trigger upgrade evaluation result; reporting the inspection anomaly according to the trigger upgrade evaluation result.

[0091] Specifically, according to the determined linkage anomaly and risk behavior level score, determine whether the level of the abnormal behavior needs to be upgraded for processing. For example, if the robot discovers a linkage abnormal signal of abnormal noise and toxic gas leakage during the inspection process, and at the same time identifies that an employee is performing a high-risk operation (such as not wearing safety equipment beside high-voltage equipment), then at this time, based on the risk behavior score, it is judged that the behavior needs to be upgraded for evaluation, that is, a higher-risk safety alarm or response is carried out. Identify anomalies in multi-modal data, such as abnormal noises or gas leakage, or other environmental changes. Perform a risk behavior score on the behavior of the personnel in the scene to evaluate the risk level of these behaviors. Based on the combination of the linkage anomaly and the risk behavior level score, if the risk level of a certain behavior is relatively high (such as misoperation, high-risk operation, etc.), then this behavior triggers a higher-level alarm or response. Upgrade according to the rules to increase the priority of the alarm or take emergency response measures.

[0092] The triggered upgrade evaluation result is a safety warning result obtained by comprehensively evaluating linkage anomalies and the risk level scores of dangerous behaviors. For example, if the evaluated behavior belongs to a potentially high-risk behavior, it is upgraded to a higher alarm level. Specific response suggestions are generated based on the dangerous behaviors and anomaly types, such as immediately stopping the operation, activating the emergency exhaust system, or evacuating personnel. According to the triggered upgrade evaluation result, an inspection anomaly report is generated, which is a summary of the current scene anomalies, listing in detail information such as the anomaly type, trigger reason, risk level, and recommended emergency response measures. The report will be automatically sent to the relevant operators or the control center to ensure that safety issues are dealt with in a timely manner.

[0093] Inspection anomalies also include behavioral warnings for personnel and leakage warnings for the scene. Behavioral warnings for personnel are mainly through the monitoring and analysis of personnel behaviors during the power station inspection process to promptly detect potential dangerous behaviors and issue alarms. The behaviors of personnel are perceived through video streams and sensor data (such as auditory and olfactory modal data). Each identified behavior will be assigned a risk level score, and its threat to safety is evaluated according to different criteria. When the identified behavior is rated as medium risk or high risk, a behavioral warning is triggered and immediately fed back to the power station management or inspection personnel through the system interface, sound alarm, or notification. The warning result will include a specific description of the dangerous behavior and handling suggestions.

[0094] Leakage warning is the real-time monitoring and warning of possible harmful gas leakage or other dangerous scenarios in the power station through environmental monitoring and sensor data, including the detection of the leakage source, the assessment of the leakage, and the emergency response to the scene. Quadruped robots and other sensor devices (such as gas sensors, temperature sensors, etc.) continuously monitor various indicators in the power station environment, especially the concentration changes of toxic and harmful gases (such as ammonia, carbon monoxide, hydrogen sulfide, etc.), and evaluate the safety environment of the power station in real time. Once the abnormal sensor data is collected, it is comprehensively analyzed in combination with the scene image information to confirm whether there is a dangerous situation of leakage. For example, when the gas sensor detects that the gas concentration exceeds the safety value, it is combined with the image analysis of the visual sensor (such as sparks, steam leakage, etc.) to determine whether it is an actual leakage event.

[0095] When it is confirmed that there is a possibility of leakage, once the concentration of harmful gases exceeds the preset safety threshold, a leakage alarm is triggered. The specific location of the leakage source is determined through image recognition technology, such as a pipeline crack or equipment failure. In some cases, dangerous behaviors of personnel and leakage events may occur simultaneously. For example, personnel's misoperation during equipment handling leads to gas leakage. If warnings in both directions are triggered simultaneously, multi-level responses are carried out, such as simultaneously activating the gas discharge system and the personnel evacuation alarm.

[0096] Furthermore, the present application also includes the following steps: Use the linked anomaly to activate the drone for perspective shift positioning, and establish the perspective shift positioning result; perform video traceability identification based on the perspective shift positioning result to generate a traceability identification result; update the linked anomaly according to the traceability identification result.

[0097] Specifically, a linked anomaly refers to a comprehensive abnormal situation determined by mutual verification of abnormal signals collected by different sensors (such as smell, hearing, vision, etc.). Once a linked anomaly is detected (such as the simultaneous occurrence of a high concentration of toxic gas and equipment failure), the drone is automatically activated for further investigation. When a linked anomaly occurs, the drone automatically departs from a predetermined position or the nearest starting point and flies towards the abnormal area. During the flight of the drone, it will perform positioning based on the geographical location and real-time data of the abnormal area, adjust the flight trajectory and shooting angle to obtain the best perspective, ensuring that the abnormal site can be monitored and recorded in detail.

[0098] After the drone performs perspective shift positioning, video traceability identification is performed on the acquired video stream data. Through the known perspective and time point, the cause and development process of the abnormal event are traced from the video data. According to the time stamp of the video and the result of the perspective shift positioning, the video streams from different perspectives are aligned in time and space, so that different perspectives of the same event can be processed synchronously. According to the abnormal behavior or environmental changes (such as personnel's illegal operations, equipment failures, etc.) in the video, combined with the known historical data and abnormal signals, the abnormal development trajectory in the video data is analyzed to generate a traceability identification result. Through video traceability identification, the occurrence process of the abnormal event can be traced, and the source and development of the anomaly can be clarified. For example, assume that a fire caused by equipment failure occurred in a power station. The specific location of the fire was confirmed through the perspective shift positioning of the drone. Through video traceability identification, the equipment state before the fire (such as overheating or short circuit of the equipment) was traced, and the specific cause of the fire was analyzed.

[0099] Based on the video traceability identification result, combined with the data from multi-modal sensors, more accurate abnormal features are supplemented to update the linked anomaly, and the response priority or trigger mechanism is adjusted. For example, if the traceability identification shows that the fire caused by equipment failure is more serious than originally judged, it is adjusted to a higher-level response. By activating the drone for perspective shift positioning through the linked anomaly and using the perspective shift positioning result for video traceability identification, the occurrence process of abnormal events in high-risk environments such as power stations can be traced more accurately, helping to identify the anomaly source, predict potential risks, and improve the efficiency of emergency response.

[0100] In summary, the multi-modal collaborative perception power station high-risk operation inspection method provided by this application has the following beneficial effects: After the power station operation task is started, activate the collaborative path tracking of the drone and the quadruped robot to execute the power station operation task; start the video acquisition units of the drone and the quadruped robot, perform video data acquisition during the collaborative path tracking, and establish a synchronized video stream; after timestamp anchoring is performed on the synchronized video stream, use an external synchronization trigger signal to perform time alignment of the synchronized video stream with time anchoring, and use a global reference coordinate system to perform multi-view alignment of the synchronized video stream, and output a video sequence with a fused view; input the video sequence with the fused view into a multi-view action behavior recognition network to establish a dangerous behavior level score; obtain the auditory modality data and olfactory modality data collected by the quadruped robot, and use the auditory modality data and olfactory modality data to establish a linkage anomaly; report a patrol anomaly according to the linkage anomaly and the dangerous behavior level score. That is to say, through the collaborative work of the drone and the quadruped robot, the global view of the drone and the local view of the quadruped robot are fused, the parallax caused by the perspective difference is eliminated, a video sequence with a fused view is output, a multi-view recognition network is called to identify the human behavior actions, the dangerous behavior level score is performed in combination with scene perception, the auditory and olfactory modality data of the quadruped robot are obtained, and the linkage anomaly recognition is established, providing more dimensions for detecting potential dangers, improving the ability to identify potential dangers, ensuring timely discovery and handling of potential risks, and effectively improving the patrol efficiency and safety of the power station.

[0101] Embodiment 2. Based on the same inventive concept as in the foregoing Embodiment 1, the present application also provides a power station high-risk operation patrol inspection system with multi-modal collaborative perception. Please refer to the attached Figure 2 , including: A collaborative tracking module 11, configured to activate the collaborative path tracking of the drone and the quadruped robot to execute the power station operation task after the power station operation task is started; a video acquisition module 12, configured to start the video acquisition units of the drone and the quadruped robot, perform video data acquisition during the collaborative path tracking, and establish a synchronized video stream; a synchronization alignment module 13, configured to perform time alignment of the synchronized video stream with time anchoring by using an external synchronization trigger signal after timestamp anchoring is performed on the synchronized video stream, and perform multi-view alignment of the synchronized video stream by using a global reference coordinate system, and output a video sequence with a fused view; a danger level assessment module 14, configured to input the video sequence with the fused view into a multi-view action behavior recognition network to establish a dangerous behavior level score; a linkage anomaly establishment module 15, configured to obtain the auditory modality data and olfactory modality data collected by the quadruped robot, and establish a linkage anomaly by using the auditory modality data and the olfactory modality data; a patrol anomaly assessment module 16, configured to report a patrol anomaly according to the linkage anomaly and the dangerous behavior level score.

[0102] Furthermore, the collaborative tracking module 11 in the power plant high-risk operation inspection system with multimodal collaborative perception is further configured to: establish a temporal-spatial displacement path according to the power plant operation tasks; use the temporal-spatial displacement path to perform associated scene calls on the power plant scene, and establish a spatial scene and a ground scene; perform optimization of the follow-up path of the temporal-spatial displacement path under the spatial scene and the ground scene, and establish a collaborative path.

[0103] Furthermore, the collaborative tracking module 11 in the power plant high-risk operation inspection system with multimodal collaborative perception is further configured to: obtain the device data of the video acquisition units of the unmanned aerial vehicle and the quadruped robot, and configure the tracking distance influence constraint according to the device data; under the tracking distance influence constraint, perform obstacle avoidance tracking optimization of the unmanned aerial vehicle with the spatial scene, and establish a first optimization path; under the tracking distance influence constraint, perform motion stability balance optimization of the quadruped robot with the ground scene, and establish a second optimization path; establish a collaborative path with the first optimization path and the second optimization path.

[0104] Furthermore, the video acquisition module 12 in the power plant high-risk operation inspection system with multimodal collaborative perception is further configured to: perform video preprocessing on the video data acquisition results, and the video preprocessing includes motion compensation, image sharpening, edge enhancement, and overexposure marking; establish a synchronized video stream based on the video preprocessing.

[0105] Furthermore, the video acquisition module 12 in the power plant high-risk operation inspection system with multimodal collaborative perception is further configured to: record the acquisition posture at each acquisition node, and generate a temporal shooting perspective using the acquisition posture; input the temporal shooting perspective into the perspective correction channel, and establish the synchronized video stream according to the perspective correction channel and the video preprocessing.

[0106] Furthermore, the risk level assessment module 14 in the power plant high-risk operation inspection system with multimodal collaborative perception is further configured to: construct a unified spatial coordinate system, project the video sequences from different perspectives after image feature extraction into the unified spatial coordinate system; call the scene skeleton extraction sub-channel of the multi-perspective action behavior recognition network, perform feature confidence evaluation on the video sequences projected into the unified spatial coordinate system, construct an anchor point set, and generate a sparse point cloud with the anchor point set; call the dense depth estimation layer of the multi-perspective action behavior recognition network to perform feature depth data of the video sequences projected into the unified spatial coordinate system; project the sparse point cloud into the feature depth data for point-depth fusion to complete local scene reconstruction annotation; establish a risk behavior level score according to the local scene reconstruction annotation.

[0107] Further, the risk level assessment module 14 in the multi-modal collaborative perception power plant high-risk operation inspection system is further configured to: call the scene and person perception layer of the multi-view action behavior recognition network, perform power plant scene and person perception segmentation within a local scene, and establish a perception segmentation result; use the perception segmentation result to perform person behavior action perception under a video sequence, and configure scene perception; use the person behavior action perception and scene perception to perform risk behavior level scoring under scene interaction.

[0108] Further, the risk level assessment module 14 in the multi-modal collaborative perception power plant high-risk operation inspection system is further configured to: establish a micro-action recognition supplement mechanism; use the micro-action recognition supplement mechanism to recognize person jitter and abnormal pause micro-actions under a video sequence; add the micro-action recognition result to the person behavior action perception.

[0109] Further, the linkage anomaly establishment module 15 in the multi-modal collaborative perception power plant high-risk operation inspection system is further configured to: after performing auditory modality and olfactory modality modeling, extract abnormal features from the auditory modality data and olfactory modality data, and establish an abnormal feature extraction result; establish a linkage trigger rule library, perform linkage trigger recognition of the linkage trigger rule library based on the abnormal feature extraction result, and establish a linkage trigger recognition result; use the linkage trigger recognition result to establish a linkage anomaly.

[0110] Further, the inspection anomaly assessment module 16 in the multi-modal collaborative perception power plant high-risk operation inspection system is further configured to: perform a trigger upgrade evaluation on the behavior in an abnormal scene based on the linkage anomaly and the risk behavior level score, and generate a trigger upgrade evaluation result; report an inspection anomaly according to the trigger upgrade evaluation result.

[0111] Further, the inspection anomaly assessment module 16 in the multi-modal collaborative perception power plant high-risk operation inspection system is further configured to: use the linkage anomaly to activate a drone for perspective transfer positioning, and establish a perspective transfer positioning result; perform video traceability recognition based on the perspective transfer positioning result, and generate a traceability recognition result; update the linkage anomaly according to the traceability recognition result.

[0112] The various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The foregoing Figure 1The power plant high-risk operation inspection method and specific example in Embodiment 1 are equally applicable to the power plant high-risk operation inspection system with multimodal collaborative perception in this embodiment. Through the foregoing detailed description of the power plant high-risk operation inspection method with multimodal collaborative perception, those skilled in the art can clearly know the power plant high-risk operation inspection system with multimodal collaborative perception in this embodiment. Therefore, for the sake of brevity of the specification, it will not be elaborated herein.

[0113] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

[0114] Obviously, those skilled in the art can make several improvements and modifications to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A patrol inspection method for high-risk operations in power plants with multi-modal collaborative perception, characterized in that, Including: After the power station operation task is started, activate the drone and quadruped robot to perform collaborative path tracking for the power station operation task; Start the video acquisition units of the drone and quadruped robot, perform video data acquisition during the collaborative path tracking, and establish a synchronized video stream; After timestamp anchoring the synchronized video stream, perform time alignment of the synchronized video stream with time anchoring using an external synchronization trigger signal, and perform multi-view alignment of the synchronized video stream using a global reference coordinate system, and output a video sequence with a fused view; Input the video sequence with the fused view into a multi-view action behavior recognition network to establish a dangerous behavior level score; Obtain the auditory modality data and olfactory modality data collected by the quadruped robot, and establish a linkage anomaly using the auditory modality data and olfactory modality data; Report an inspection anomaly according to the linkage anomaly and the dangerous behavior level score.

2. The method for inspecting and patroling high-risk operations in a power station with multi-modal collaborative perception according to claim 1, wherein The inputting the video sequence with the fused view into a multi-view action behavior recognition network to establish a dangerous behavior level score includes: Construct a unified space coordinate system, extract image features from video sequences with different views, and project them into the unified space coordinate system; Call the scene skeleton extraction sub-channel of the multi-view action behavior recognition network, perform feature confidence evaluation on the video sequence projected into the unified space coordinate system, construct an anchor point set, and generate a sparse point cloud with the anchor point set; Call the dense depth estimation layer of the multi-view action behavior recognition network to perform feature depth data of the video sequence projected into the unified space coordinate system; Project the sparse point cloud into the feature depth data for point-depth fusion to complete local scene reconstruction annotation; Establish a dangerous behavior level score according to the local scene reconstruction annotation.

3. The multi-modal collaborative perception-based high-risk operation inspection method for power plants according to claim 2, wherein, The establishing a dangerous behavior level score according to the local scene reconstruction annotation includes: Call the scene person perception layer of the multi-view action behavior recognition network, perform power station scene and person perception segmentation in the local scene, and establish a perception segmentation result; Perform person behavior action perception under the video sequence using the perception segmentation result, and configure scene perception; Perform a dangerous behavior level score under scene interaction using the person behavior action perception and scene perception.

4. The method for inspecting high-risk operations in a power station with multi-modal collaborative perception according to claim 3, wherein, The performing person behavior action perception under the video sequence using the perception segmentation result includes: Establish a micro-action recognition supplementary mechanism; Perform micro-action recognition of person jitter and abnormal pause under the video sequence using the micro-action recognition supplementary mechanism; Add the micro-action recognition result to the person behavior action perception.

5. The multi-modal collaborative perception-based inspection method for high-risk operations in power plants according to claim 1, wherein, The activating the drone and quadruped robot to perform collaborative path tracking for the power station operation task includes: Establish a time-sequential space displacement path according to the power station operation task; Use the time-sequential space displacement path to call associated scenes of the power station scene, and establish a space scene and a ground scene; Perform pursuit path optimization of the time-sequential space displacement path under the space scene and the ground scene to establish a collaborative path.

6. The multi-modal collaborative perception-based inspection method for high-risk operations in power plants according to claim 5, wherein, The performing pursuit path optimization of the time-sequential space displacement path under the space scene and the ground scene to establish a collaborative path includes: Obtain the device data of the video acquisition units of the drone and the quadruped robot, and configure the tracking distance impact constraint according to the device data; Under the tracking distance impact constraint, perform obstacle avoidance tracking optimization of the drone with the spatial scene, and establish a first optimization path; Under the tracking distance impact constraint, perform motion stability balance optimization of the quadruped robot with the ground scene, and establish a second optimization path; Establish a collaborative path with the first optimization path and the second optimization path.

7. The multi-modal collaborative perception-based inspection method for high-risk operations in power plants according to claim 1, wherein The establishment of the linkage anomaly using the auditory modality data and the olfactory modality data includes: After performing auditory modality and olfactory modality modeling, extract abnormal features from the auditory modality data and the olfactory modality data, and establish an abnormal feature extraction result; Establish a linkage trigger rule library, perform linkage trigger recognition of the linkage trigger rule library based on the abnormal feature extraction result, and establish a linkage trigger recognition result; Establish a linkage anomaly using the linkage trigger recognition result.

8. The method for inspecting high-risk operations in a power station with multi-modal collaborative perception according to claim 7, characterized in that, The reporting of the inspection anomaly according to the linkage anomaly and the dangerous behavior level score includes: Perform a trigger upgrade evaluation of the behavior in the abnormal scenario with the linkage anomaly and the dangerous behavior level score, and generate a trigger upgrade evaluation result; Report the inspection anomaly according to the trigger upgrade evaluation result.

9. The multi-modal collaborative perception-based inspection method for high-risk operations in power plants according to claim 8, characterized in that, Before reporting the inspection anomaly according to the linkage anomaly and the dangerous behavior level score, it includes: Use the linkage anomaly to activate the drone for perspective shift positioning, and establish a perspective shift positioning result; Perform video traceability recognition based on the perspective shift positioning result, and generate a traceability recognition result; Update the linkage anomaly according to the traceability recognition result.

10. The method for inspecting high-risk operations in a power station with multimodal collaborative perception according to claim 1, characterized in that, The execution of video data collection during the collaborative path tracking to establish a synchronized video stream includes: Perform video preprocessing on the video data collection result, and the video preprocessing includes motion compensation, image sharpening, edge enhancement, and overexposure marking; Establish a synchronized video stream based on the video preprocessing.

11. The multi-modal collaborative perception-based high-risk operation inspection method for power plants according to claim 10, wherein, The establishment of the synchronized video stream based on the video preprocessing includes: Record the acquisition pose at each acquisition node, and generate a sequential shooting perspective using the acquisition pose; Input the sequential shooting perspective into the perspective correction channel, and establish the synchronized video stream according to the perspective correction channel and the video preprocessing.

12. A power plant high-risk operation inspection system with multimodal collaborative perception, characterized in that, Steps for implementing the method for inspecting high-risk operations in a power station with multi-modal collaborative perception according to any one of claims 1 to 11, the system for inspecting high-risk operations in a power station with multi-modal collaborative perception includes: A collaborative tracking module, configured to activate the drone and the quadruped robot to perform collaborative path tracking of the power station operation task after the power station operation task is started; A video acquisition module, configured to start the video acquisition units of the drone and the quadruped robot, perform video data collection during the collaborative path tracking, and establish a synchronized video stream; A synchronization alignment module, configured to perform time alignment of the synchronized video stream with an external synchronization trigger signal after timestamp anchoring of the synchronized video stream, and perform multi-view alignment of the synchronized video stream using a global reference coordinate system, and output a video sequence of the fused perspective; A danger level assessment module, which is used to input a video sequence from a fused perspective into a multi-perspective action and behavior recognition network to establish a danger behavior level score; An abnormal linkage establishment module, which is used to obtain auditory modality data and olfactory modality data collected by a quadruped robot, and establish an abnormal linkage by using the auditory modality data and the olfactory modality data; An inspection abnormal assessment module, which is used to report an inspection abnormality according to the abnormal linkage and the danger behavior level score.

Citation Information

Patent Citations

  • Air-ground cooperative intelligent inspection robot and inspection method

    CN111300372A

  • Cloud-side collaborative 1+6 + N intelligent inspection system and method for intelligent thermal power plant

    CN118644906A

  • Intelligent robot for power inspection and inspection method thereof

    CN119773891A

  • Power distribution network equipment hidden danger detection method based on large-view-field spliced video data and related device

    CN120013921A

  • Rapid and high-accuracy reconstruction method for power lines based on multi-source data fusion

    GB202319411D0

Cited By

  • Substation autonomous inspection and foreign matter cleaning cooperative control system and method

    CN120638653A

  • Mobile robot control device giving consideration to field navigation patrol and transportation load

    CN120821235A

  • Dual-unmanned aerial vehicle inspection method and program product

    CN121143401A

  • Security check method and system based on voice and smell fusion data recognition

    CN121598267A

  • Laser point cloud and video fusion method and device based on 3D deploy and control ball

    CN121639482A