Power distribution unmanned aerial vehicle multi-modal collaborative autonomous shooting method, system, device and medium
By using multimodal data perception and hierarchical decision optimization, the positioning and image acquisition problems of power distribution drones in complex scenarios have been solved, enabling autonomous image acquisition and efficient operation and maintenance, and improving the accuracy and efficiency of power distribution network inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
- Filing Date
- 2025-12-09
- Publication Date
- 2026-08-04
AI Technical Summary
Existing power distribution drone search and capture technologies suffer from insufficient positioning accuracy, passive decision-making, lack of closed-loop image quality, and lack of embodied interactive capabilities. This results in low success rates for search and capture in complex scenarios, making it impossible to achieve accurate image acquisition and efficient operation and maintenance.
By constructing a multimodal collaborative autonomous search and capture method, semantic feature extraction and environmental interaction perception are performed using lidar point clouds, visible light images, and inertial navigation data. An embodied perception map is constructed, and combined with hierarchical emergency-driven decision-making and real-time image quality assessment, the UAV can autonomously select search and capture behaviors and optimize its behavior.
It enables autonomous and efficient inspection of drones in complex scenarios, improves positioning accuracy and image quality, reduces operation and maintenance costs, and enhances the ability to accurately locate and photograph defective areas of power distribution equipment. It is suitable for mountainous areas and densely populated urban lines.
Smart Images

Figure CN121746965B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power distribution network inspection technology, specifically relating to a multimodal collaborative autonomous search and shooting method, system, equipment and medium for power distribution drones. Background Technology
[0002] During power distribution network inspections, drones need to accurately capture images (i.e., "search and capture") of defects such as insulator damage, broken conductor strands, bird nests, and foreign objects. This is the foundation for subsequent defect identification and maintenance decisions. Existing power distribution drone search and capture technologies largely rely on manual remote control or simple path planning, lacking the core capabilities of "perceiving the environment, adjusting behavior, and adapting to the scenario," and suffer from four major problems: 1. Insufficient positioning accuracy, resulting in semantically unsound map construction. Traditional Simultaneous Localization and Mapping (SLAM) technology is susceptible to electromagnetic interference and dense equipment (such as multiple intersecting wires) in power distribution scenarios, resulting in a positioning drift error of ≥0.5m. Furthermore, the constructed maps only contain geometric information and lack semantic labels for equipment such as "insulator-wire-tower". This causes drones to be unable to perceive the relationship between equipment attributes and the environment (such as "the spatial distance between insulators and tree branches"), making it impossible for them to actively avoid obstacles when searching for and shooting, resulting in a low equipment alignment rate.
[0003] 2. Passive decision-making in finding and capturing targets, poor adaptability to emergency scenarios. Existing search and capture decisions are mostly based on "point-to-point shooting" along a preset path, which cannot dynamically adjust behavior according to the real-time scene (e.g., shooting along the original path when blocked by tree branches, resulting in missed shots); and no environment-behavior mapping relationship has been established (e.g., "backlight scene → adjust shooting angle" "strong wind scene → reduce flight altitude"), resulting in delayed behavior response in emergency scenarios (e.g., insulator self-explosion).
[0004] 3. The image quality lacks closed-loop resolution, affecting subsequent identification. The image quality was not evaluated in real time during the search and capture process, resulting in 15% to 20% of the captured images being of low quality. Furthermore, the quality assessment was not linked to the search and capture behavior, making it impossible to achieve a closed-loop control of "quality not meeting the standard → adjusting behavior and re-execution". This led to a high rate of re-capture and increased maintenance costs.
[0005] 4. Lacks embodied interaction capabilities and has poor adaptability to complex scenarios. The existing system has not built a "perception-decision-execution-feedback" link, so the drone cannot continuously optimize its behavior according to environmental changes (such as "the first search and shooting quality is not up to standard due to obstruction → the second time actively detouring to an unobstructed location"). In complex scenarios such as mountainous areas with many obstructions and dense urban routes, the success rate of search and shooting is low, far below the actual operation and maintenance requirements. Summary of the Invention
[0006] The purpose of this invention is to address the problems in the prior art by providing a multimodal collaborative autonomous search and capture method, system, device, and medium for power distribution drones. This invention builds the perception capabilities of power distribution drones, optimizes search and capture behavior decisions, improves the feedback loop to adjust search and capture behavior in a targeted manner, enhances adaptability to complex scenarios, and enables autonomous and efficient inspection of power distribution drones.
[0007] To achieve the above objectives, the present invention provides the following technical solution: Firstly, a multimodal cooperative autonomous search and capture method for power distribution drones is provided, including: Acquire multimodal data collected by power distribution drones, extract semantic features and perform environmental interaction perception on the multimodal data, and construct an embodied perception map; Based on embodied perception maps, hierarchical emergency-driven decision-making is carried out, and the power distribution drones are enabled to autonomously select search and capture behaviors according to real-time environmental parameters based on the decision results. The system acquires images captured by the power distribution drone under its current search and capture behavior, performs real-time image quality assessment on these images, and obtains a comprehensive quality score based on the fused environmental parameters. Based on the comprehensive quality score of the fused environmental parameters, the system determines the cause of quality defects and adjusts the search and capture behavior of the power distribution drone accordingly, thereby achieving multimodal collaborative autonomous search and capture by the power distribution drone.
[0008] As a preferred embodiment, in the step of acquiring multimodal data collected by the power distribution drone, the multimodal data includes lidar point clouds for sensing the spatial position of the device and obstacles, visible light images for sensing the semantics of the device and the lighting environment, inertial navigation data for sensing flight attitude and wind speed interference, and environmental perception data as input for embodied decision-making environment parameters. The environmental sensing data includes light intensity, wind speed, occlusion rate, and device distance.
[0009] As a preferred embodiment, the method also includes a preprocessing step for the acquired multimodal data, the preprocessing including: VoxelGrid filtering is used to downsample the LiDAR point cloud, compressing the number of points to preserve the geometric features of the device and obstacles. The Retinex algorithm is used to eliminate the backlighting effect of the visible light image, and adaptive histogram equalization is used to improve the texture contrast of the device while extracting the image illumination intensity. The distance and occlusion rate between the device and obstacles are calculated based on the LiDAR point cloud, and the wind speed is retrieved by combining inertial navigation data to calibrate the environmental perception data. Finally, the LiDAR point cloud, visible light image, inertial navigation data, and environmental perception data are synchronized based on GPS timestamps.
[0010] As a preferred approach, the steps of extracting semantic features from multimodal data and perceiving environmental interactions to construct an embodied perception map include: designing a semantic feature extraction network and establishing the association between devices and the environment, including visual semantic features and point cloud semantic features associated with the environment; wherein, visual semantic features: a lightweight CNN is used to extract device semantic features from images, outputting a visual semantic probability map, adding "obstacle categories" corresponding to 6 semantic categories: "insulator, conductor, tower, transformer, obstacle, and background", and using a specific loss function to ensure the accuracy of semantic classification of devices and obstacles, realizing the perception and distinction between devices and obstacles; point cloud features associated with the environment: the visual semantic probability map is mapped to the point cloud coordinate system to obtain point cloud semantic labels, and then point cloud semantic features are extracted using PointNet, while calculating the spatial distance between devices and obstacles to construct a "device-obstacle distance matrix", and using feature matching loss to ensure the accuracy of cross-frame matching of semantic features, while avoiding drones flying within the dangerous distance between devices and obstacles, thus enhancing safety perception; By integrating semantic features, inertial navigation data, and environmental perception data into simultaneous localization and mapping (SLAM) localization optimization, a tightly coupled error function is constructed to output an embodied perception map containing environmental interaction information.
[0011] In the formula, Geometric error is used to measure the geometric spatial deviation during the positioning process. It is achieved by calculating the point cloud registration error between the current frame and the key frame. The iterative nearest point (ICP) algorithm is adopted to provide basic spatial accuracy for UAV positioning in complex power distribution scenarios. To mitigate semantic errors, semantic features are incorporated to understand the semantic information of objects in the environment. By analyzing the correlation between semantic features and positioning results, the semantic consistency of positioning is optimized, improving its adaptability to real-world application scenarios. Weighting coefficients... This reflects the importance of semantic features in the overall error function. The value was determined through cross-validation in a power distribution scenario; The inertial navigation error is calculated based on the inertial navigation pre-integration model to determine the difference between the predicted and observed poses. In power distribution environments, electromagnetic interference and wind speed are present; the inertial navigation error is used to suppress the resulting positioning drift and ensure the stability of the UAV's positioning in dynamic environments. Weighting coefficients are also included. This reflects the weighting percentage of inertial navigation data optimization. Environmental error, a embodied perception constraint, is used to calculate the deviation between environmental parameters and optimal image acquisition conditions based on the correlation between environmental perception data and device attributes. This environmental error ensures that the perception map prioritizes areas meeting the optimal image acquisition conditions, providing a basis for subsequent environmental optimization decisions. Weighting coefficients are also included. This reflects the strength of the role of environmental perception data in overall optimization; Each weighting coefficient is determined through cross-validation in a power distribution scenario to balance the contribution of different error terms to positioning optimization, thereby achieving optimal overall positioning performance; outputting a perception map. It includes device semantic tags, environmental parameters, and device-obstacle distance matrices, enabling UAVs to have a embodied perception of device attributes, environmental conditions, and spatial relationships, providing data support for decision-making.
[0012] As a preferred approach, in the step of performing hierarchical emergency-driven decision-making based on the embodied perception map and enabling the power distribution drone to autonomously select its search and capture behavior according to real-time environmental parameters based on the decision results, a four-layer decision-making model of "equipment priority - scenario emergency level - path cost - behavior adaptation" is constructed. Among them, the equipment priority is weighted based on the degree of impact of power distribution defects and dynamically adjusted in combination with the equipment status in the perception map. The expression for optimizing the emergency response level of a scenario is: ; In the formula, This is the normalized value of wind speed. To mask the normalized value, This is the normalized value for the equipment-obstacle safe distance; The path cost optimization expression is: ; In the formula, The energy cost of behavioral adjustment This represents the Euclidean distance between the drone's current position and the target position. This represents the number of obstacles in the path. , , These are the weighting coefficients for each part; The hierarchical emergency-driven decision-making process calculates a comprehensive adaptation score based on "device-scenario-path-behavior". The highest-rated combination of behaviors is selected as the target image acquisition scheme, and the calculation expression is as follows: ; In the first-level device priority filtering, retain The first layer of equipment prioritizes devices suspected of defects marked on the perception map; the second layer of scene urgency filtering will... Filter out the scene, for The first layer pre-adjusts the scene-triggered behavior; the third layer of behavior adaptation optimization, for the remaining device-scene combinations, traverses all possible search and capture behaviors, and calculates the behavior for each combination. The fourth layer of dynamic adjustment triggering mechanism stipulates that if a sudden change in environmental parameters occurs during flight, it will be updated in real time. and Recalculate and adjust behavior. Represents behavioral fit. It is the path cost; When a sudden defect in the real-time identification of the perception map is detected, a three-level embodied emergency response is triggered: Level 1 response involves suspending the current mission of the drone and reconfirming the defect; if the defect is confirmed on the second confirmation, the response proceeds to Level 2, temporarily suspending the defective equipment. Upgrade, update It will also complete emergency path replanning; if a single drone has insufficient battery life or a complex scenario, it will initiate a level three response, which will schedule surrounding drones through edge computing nodes and complete the response through collaborative behavior.
[0013] As a preferred embodiment, in the steps of acquiring the search and capture images under the current search and capture behavior of the power distribution drone, performing real-time image quality assessment on the search and capture images, and obtaining a comprehensive quality score fused with environmental parameters, the comprehensive quality score fused with environmental parameters is calculated according to the following formula:
[0014] In the formula, This represents the image sharpness score, reflecting the focus quality and detail rendering of the captured image; This represents a light quality score, used to measure whether the ambient light during shooting is sufficient and even. The coverage integrity score reflects the impact of drone shake on image stability during shooting. Represents attitude stability score; This represents the overall environmental score, which takes into account the impact of other environmental factors besides those mentioned above on the shooting quality.
[0015] As a preferred embodiment, the attitude stability score The calculation expression is as follows:
[0016] In the formula, , These represent the changes in roll and pitch angles at time t, respectively, where T is the shooting duration.
[0017] As a preferred embodiment, in the step of determining the cause of quality defects based on the comprehensive quality score of the fused environmental parameters, and adjusting the search and shooting behavior of the power distribution drone based on the cause of the quality defects to achieve multimodal collaborative autonomous search and shooting of the power distribution drone: If the quality defect is caused by "too close distance", then the action adjustment is to "reverse the drone by 1m and refocus"; if the quality defect is caused by "backlight", then the action adjustment is to "fly the drone to the side to avoid the light source and adjust the shooting angle"; if the quality defect is caused by "strong wind shaking", then the action adjustment is to "reduce the drone altitude, turn on wind-resistant mode and shorten the shooting time". After implementing the adjusted behavior and re-shooting, if the overall quality score after integrating environmental parameters still fails to meet the standard after consecutive reshoots, it is marked as "requiring manual reshooting," and the environmental parameters are recorded. The "environmental parameters - behavior adjustment - quality result" data is stored in the local knowledge base, and the behavior strategy is updated through reinforcement learning. This enables continuous optimization of embodied decision-making.
[0018] Secondly, a multimodal collaborative autonomous search and capture system for power distribution drones is provided, including: The embodied perception map construction module is used to acquire multimodal data collected by power distribution drones, extract semantic features from the multimodal data and perceive environmental interactions to construct an embodied perception map. The hierarchical emergency-driven decision-making module is used to make hierarchical emergency-driven decisions based on the embodied perception map, and enable the power distribution drone to autonomously select its search and shooting behavior according to real-time environmental parameters based on the decision results. The real-time image quality assessment and search behavior adjustment module is used to acquire search images under the current search behavior of the power distribution drone, perform real-time image quality assessment on the search images, and obtain a comprehensive quality score based on the fused environmental parameters; determine the cause of quality defects based on the comprehensive quality score based on the fused environmental parameters, and adjust the search behavior of the power distribution drone based on the cause of quality defects, so as to realize multimodal collaborative autonomous search of the power distribution drone.
[0019] As a preferred embodiment, when the embodied perception map construction module acquires multimodal data collected by the power distribution drone, the multimodal data includes lidar point clouds for sensing the spatial location of the device and obstacles, visible light images for sensing the semantics of the device and the lighting environment, inertial navigation data for sensing flight attitude and wind speed interference, and environmental perception data as input for embodied decision-making environment parameters. The environmental sensing data includes light intensity, wind speed, occlusion rate, and device distance.
[0020] As a preferred embodiment, the embodied perception map construction module preprocesses the collected multimodal data, including: VoxelGrid filtering is used to downsample the LiDAR point cloud, compressing the number of points to preserve the geometric features of the device and obstacles. The Retinex algorithm is used to eliminate the backlighting effect of the visible light image, and adaptive histogram equalization is used to improve the texture contrast of the device while extracting the image illumination intensity. The distance and occlusion rate between the device and obstacles are calculated based on the LiDAR point cloud, and the wind speed is retrieved by combining inertial navigation data to calibrate the environmental perception data. Finally, the LiDAR point cloud, visible light image, inertial navigation data, and environmental perception data are synchronized based on GPS timestamps.
[0021] As a preferred solution, the embodied perception map construction module designs a semantic feature extraction network when performing semantic feature extraction and environmental interaction perception on multimodal data, and establishes the association between devices and the environment, including visual semantic features and point cloud semantic features and their association with the environment. Specifically, for visual semantic features: a lightweight CNN is used to extract device semantic features from images, outputting a visual semantic probability map, adding "obstacle categories" corresponding to six semantic categories: "insulator, conductor, tower, transformer, obstacle, and background." A specific loss function ensures the accuracy of device and obstacle semantic classification, achieving device-obstacle perception differentiation. For point cloud features and their association with the environment: the visual semantic probability map is mapped to a point cloud coordinate system to obtain point cloud semantic labels, and then point cloud semantic features are extracted using PointNet. Simultaneously, the spatial distance between devices and obstacles is calculated to construct a "device-obstacle distance matrix." Feature matching loss ensures the accuracy of cross-frame semantic feature matching, while preventing drones from flying within dangerous distances between devices and obstacles, thus enhancing safety perception. By integrating semantic features, inertial navigation data, and environmental perception data into simultaneous localization and mapping (SLAM) localization optimization, a tightly coupled error function is constructed to output an embodied perception map containing environmental interaction information.
[0022] In the formula, Geometric error is used to measure the geometric spatial deviation during the positioning process. It is achieved by calculating the point cloud registration error between the current frame and the key frame. The iterative nearest point (ICP) algorithm is adopted to provide basic spatial accuracy for UAV positioning in complex power distribution scenarios. To mitigate semantic errors, semantic features are incorporated to understand the semantic information of objects in the environment. By analyzing the correlation between semantic features and positioning results, the semantic consistency of positioning is optimized, improving its adaptability to real-world application scenarios. Weighting coefficients... This reflects the importance of semantic features in the overall error function. The value was determined through cross-validation in a power distribution scenario; The inertial navigation error is calculated based on the inertial navigation pre-integration model to determine the difference between the predicted and observed poses. In power distribution environments, electromagnetic interference and wind speed are present; the inertial navigation error is used to suppress the resulting positioning drift and ensure the stability of the UAV's positioning in dynamic environments. Weighting coefficients are also included. This reflects the weighting percentage of inertial navigation data optimization. Environmental error, a embodied perception constraint, is used to calculate the deviation between environmental parameters and optimal image acquisition conditions based on the correlation between environmental perception data and device attributes. This environmental error ensures that the perception map prioritizes areas meeting the optimal image acquisition conditions, providing a basis for subsequent environmental optimization decisions. Weighting coefficients are also included. This reflects the strength of the role of environmental perception data in overall optimization; Each weighting coefficient is determined through cross-validation in a power distribution scenario to balance the contribution of different error terms to positioning optimization, thereby achieving optimal overall positioning performance; outputting a perception map. It includes device semantic tags, environmental parameters, and device-obstacle distance matrices, enabling UAVs to have a embodied perception of device attributes, environmental conditions, and spatial relationships, providing data support for decision-making.
[0023] As a preferred embodiment, the hierarchical emergency-driven decision-making module constructs a four-layer decision-making model of "equipment priority - scenario emergency level - path cost - behavior adaptation". Among them, the equipment priority is weighted based on the degree of impact of power distribution defects and dynamically adjusted in combination with the equipment status in the perception map. The expression for optimizing the emergency response level of a scenario is: ; In the formula, This is the normalized value of wind speed. To mask the normalized value, This is the normalized value for the equipment-obstacle safe distance; The path cost optimization expression is: ; In the formula, The energy cost of behavioral adjustment This represents the Euclidean distance between the drone's current position and the target position. This represents the number of obstacles in the path. , , These are the weighting coefficients for each part; The hierarchical emergency-driven decision-making process calculates a comprehensive adaptation score based on "device-scenario-path-behavior". The highest-rated combination of behaviors is selected as the target image acquisition scheme, and the calculation expression is as follows: ; In the first-level device priority filtering, retain The first layer of equipment prioritizes devices suspected of defects marked on the perception map; the second layer of scene urgency filtering will... Filter out the scene, for The first layer pre-adjusts the scene-triggered behavior; the third layer of behavior adaptation optimization, for the remaining device-scene combinations, traverses all possible search and capture behaviors, and calculates the behavior for each combination. The fourth layer of dynamic adjustment triggering mechanism stipulates that if a sudden change in environmental parameters occurs during flight, it will be updated in real time. and Recalculate and adjust behavior. Represents behavioral fit. It is the path cost; When a sudden defect in the real-time identification of the perception map is detected, a three-level embodied emergency response is triggered: Level 1 response involves suspending the current mission of the drone and reconfirming the defect; if the defect is confirmed on the second confirmation, the response proceeds to Level 2, temporarily suspending the defective equipment. Upgrade, update It will also complete emergency path replanning; if a single drone has insufficient battery life or a complex scenario, it will initiate a level three response, which will schedule surrounding drones through edge computing nodes and complete the response through collaborative behavior.
[0024] As a preferred embodiment, the real-time image quality assessment and target acquisition behavior adjustment module calculates the comprehensive quality score of the fused environmental parameters using the following formula:
[0025] In the formula, This represents the image sharpness score, reflecting the focus quality and detail rendering of the captured image; This represents a light quality score, used to measure whether the ambient light during shooting is sufficient and even. The coverage integrity score reflects the impact of drone shake on image stability during shooting. Represents attitude stability score; The overall environmental score takes into account the impact of other environmental factors besides those mentioned above on the shooting quality. The attitude stability score The calculation expression is as follows:
[0026] In the formula, , These represent the changes in roll and pitch angles at time t, respectively, where T is the shooting duration.
[0027] As a preferred embodiment, if the real-time image quality assessment and shooting behavior adjustment module determines that the quality defect is caused by "too close distance", it will perform the behavior adjustment of "drone retreating 1m + refocusing"; if it determines that the quality defect is caused by "backlight", it will perform the behavior adjustment of "drone flying sideways to avoid the light source + adjusting the shooting angle"; if it determines that the quality defect is caused by "strong wind shaking", it will take the behavior adjustment of "drone reducing altitude + turning on wind-resistant mode + shortening shooting time". After implementing the adjusted behavior and re-shooting, if the overall quality score after integrating environmental parameters still fails to meet the standard after consecutive reshoots, it is marked as "requiring manual reshooting," and the environmental parameters are recorded. The "environmental parameters - behavior adjustment - quality result" data is stored in the local knowledge base, and the behavior strategy is updated through reinforcement learning. This enables continuous optimization of embodied decision-making.
[0028] Thirdly, an electronic device is provided, including a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the aforementioned multimodal cooperative autonomous shooting method for power distribution drones.
[0029] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one instruction, which, when executed by a processor, implements the aforementioned multimodal cooperative autonomous search and capture method for power distribution UAVs.
[0030] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects: By integrating the environmental perception capabilities of Enhanced Semantic Simultaneous Localization and Mapping (ES-SLAM), the dynamic behavior planning capabilities of Hierarchical Emergency Driven Decision Making (HEDD), and the execution feedback capabilities of Real-Time Image Quality Assessment (RIQA) closed-loop, an embodied intelligent interaction framework is constructed. This framework addresses the pain points of "poor environmental adaptability, rigid behavior, and delayed feedback" in power distribution drone inspections. It enables precise location and imaging of defective areas in equipment such as insulators, conductors, and towers of power distribution lines. It is suitable for autonomous drone inspections in complex power distribution scenarios (such as mountainous areas with many obstructions and dense urban lines), and also supports extended applications such as collaborative image acquisition by live-line working robots in power distribution networks and AI analysis of inspection data. This invention acquires multimodal data collected by a power distribution drone, extracts semantic features and performs environmental interaction perception on the multimodal data, constructs an embodied perception map, and uses ES-SLAM to build the embodied perception map to achieve real-time perception of "device attributes-environmental status-spatial association", overcoming the shortcomings of traditional SLAM. Combined with positioning optimization, it helps the drone avoid dangerous areas. Based on the embodied perception map, hierarchical emergency-driven decision-making is carried out. According to the decision results, the power distribution drone autonomously selects its search and capture behavior based on real-time environmental parameters. The hierarchical emergency-driven decision-making HEDD achieves dynamic adaptation of "environment-behavior", and the emergency response mechanism reduces the risk of failure and maintenance costs. The method acquires the search and capture images under the current search and capture behavior of the power distribution drone, performs real-time image quality assessment on the search and capture images, obtains a comprehensive quality score fused with environmental parameters, determines the cause of quality defects based on the comprehensive quality score fused with environmental parameters, and adjusts the search and capture behavior of the power distribution drone based on the cause of quality defects. This realizes multimodal collaborative autonomous search and capture of the power distribution drone. The real-time image quality assessment RIQA feedback loop improves image quality, and self-evolution is achieved through knowledge base and reinforcement learning. The method of this invention enables power distribution drones to conduct autonomous and efficient inspections, shortening fault repair response time with high-quality images and comprehensively improving the value of power distribution operation and maintenance.
[0031] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 Flowchart of the multimodal cooperative autonomous search and capture method for power distribution drones according to an embodiment of the present invention; Figure 2 Principle architecture diagram of the multimodal cooperative autonomous search and capture method for power distribution UAVs according to an embodiment of the present invention; Figure 3 A block diagram of the multimodal collaborative autonomous search and capture system for power distribution drones according to an embodiment of the present invention. Detailed Implementation
[0034] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail. Flowcharts are used in the embodiments of this application to illustrate the operations performed by the apparatus according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more steps may be removed from these processes.
[0035] Please see Figure 1 and Figure 2 To overcome the shortcomings of existing power distribution drone (PDD) search and capture technologies, such as low positioning accuracy, passive decision-making, lack of closed-loop quality control, and absence of embodied interaction capabilities, this invention proposes a multimodal collaborative autonomous search and capture method for PDDs. This method constructs a three-in-one autonomous search and capture method for PDDs, integrating "Sensation (ES-SLAM) - Decision-Making (HEDD) - Feedback (RIQA)". The core of this method is to achieve the intelligent characteristics of the PDD—"perceiving the environment - adjusting behavior - adapting to the scene"—through multimodal environmental perception, dynamic behavioral decision-making, and real-time execution feedback. Specifically, the method of this invention includes the following steps: S1. Acquire multimodal data collected by the power distribution drone, extract semantic features and perform environmental interaction perception on the multimodal data, and construct an embodied perception map; S2. Based on the embodied perception map, hierarchical emergency-driven decision-making is carried out, and the power distribution drone is enabled to autonomously select search and shooting behavior according to real-time environmental parameters based on the decision results. S3. Acquire the search and capture images under the current search and capture behavior of the power distribution drone, perform real-time image quality assessment on the search and capture images, and obtain a comprehensive quality score based on the fused environmental parameters; determine the cause of quality defects based on the comprehensive quality score based on the fused environmental parameters, and adjust the search and capture behavior of the power distribution drone based on the cause of quality defects, so as to realize the multimodal collaborative autonomous search and capture of the power distribution drone.
[0036] In one possible implementation, step S1 employs the ES-SLAM embodied perception map construction method. Addressing the characteristics of strong electromagnetic interference and dense equipment in power distribution scenarios, it integrates multimodal data to design a perception map containing "equipment semantics + environmental parameters + safety distance," achieving coordinated optimization of localization and perception. Traditional SLAM only outputs a geometric map; this invention adds an environmental perception dimension, enabling the map to "perceive environmental risks and support behavioral decisions," solving the problem of incomplete perception in complex scenarios. This embodiment of the invention is based on the traditional SLAM framework, incorporating semantic features of power distribution equipment and environmental interaction information to achieve three-dimensional optimization of "geometric localization + semantic recognition + environmental perception." The core is to construct a perception map containing equipment attributes and obstacle information through semantic feature extraction and environmental interaction perception, providing accurate environmental data for decision-making. Step S1 specifically includes: The first step is multimodal data input and preprocessing. The input consists of multimodal data collected by the power distribution drone, including lidar point clouds used for sensing devices and the spatial location of obstacles. ( (Voxel size 0.02m, ranging accuracy ±2cm); visible light images used to sense device semantics and lighting environment. (H=1080, W=1920, 30fps); Inertial navigation data used to sense flight attitude and wind speed disturbances. (T is the time step, acceleration accuracy ±0.01m / ) (Angular velocity accuracy ±0.1° / s); and environmental perception data including parameters such as light intensity, wind speed, occlusion rate, and device distance, serving as input parameters for embodied decision-making environment parameters. In the preprocessing stage, VoxelGrid filtering is used to downsample the LiDAR point cloud, compressing the number of points to [a smaller value]. To preserve the geometric features of the equipment and obstacles, the Retinex algorithm is used to eliminate the backlighting effect of visible light images, and adaptive histogram equalization is used to improve the texture contrast of the equipment while extracting the image illumination intensity. Based on the distance and occlusion rate of the equipment and obstacles in the LiDAR point cloud, wind speed is inverted by combining IMU data to calibrate the environmental perception data to ensure that the error is ≤5%. Finally, the point cloud, image, IMU and environmental data are synchronized based on GPS timestamps to ensure that the time synchronization error is ≤1ms.
[0037] Secondly, semantic feature extraction and matching are performed. A semantic feature extraction network is designed for the specific features of power distribution equipment (such as insulator skirt texture and conductor linear structure). Simultaneously, the association between equipment and the environment is established, mainly including visual semantic features and point cloud semantic features linked to the environment. Visual semantic features: A lightweight CNN (MobileViT-XXS) is used to extract the semantic features of equipment in images, outputting a semantic probability map. A new "obstacle category" is added, corresponding to six semantic categories: "insulator, conductor, tower, transformer, obstacle, and background." A specific loss function ensures that the semantic classification accuracy of equipment and obstacles is ≥98%, achieving perceptual differentiation between equipment and obstacles. Point cloud features linked to the environment: The visual semantic probability map is mapped to a point cloud coordinate system to obtain point cloud semantic labels. PointNet is then used to extract point cloud semantic features, while simultaneously calculating the spatial distance between equipment and obstacles to construct an "equipment-obstacle distance matrix." Feature matching loss ensures that the cross-frame matching accuracy of semantic features is ≥95%, while preventing drones from flying within dangerous distances between equipment and obstacles, thus enhancing safety perception.
[0038] Finally, tightly coupled localization and perception map construction are performed. Semantic features, IMU data, and environmental perception data are integrated into SLAM localization optimization to construct a tightly coupled error function, outputting an embodied perception map containing environmental interaction information. The calculation expression is as follows: .in, Geometric error is used to measure the geometric spatial deviation during the positioning process. It is achieved by calculating the point cloud registration error between the current frame and the key frame. The Iterative Closest Point (ICP) algorithm is adopted to ensure that the geometric positioning accuracy is controlled within ≤0.05m in complex power distribution scenarios, providing a basic spatial accuracy guarantee for UAV positioning. To mitigate semantic errors, semantic features are incorporated, aiming to enable the system to understand the semantic information of objects in the environment. By analyzing the correlation between semantic features and localization results, the semantic consistency of localization is optimized, such as identifying key targets like power distribution equipment, thereby improving the adaptability of localization to real-world application scenarios. Weighting coefficients This reflects the importance of semantic features in the overall error function, and the value was determined through cross-validation in a power distribution scenario. This represents the IMU error, calculated based on the IMU pre-integration model, to determine the difference between the predicted and observed poses. In power distribution environments, electromagnetic interference and wind speed can affect positioning; this error term effectively suppresses the resulting positioning drift, ensuring the stability of the UAV's positioning in dynamic environments. Weighting coefficients. This reflects the weighting of IMU data optimization. To account for environmental errors, a new embodied perception constraint term is added. Based on the correlation between environmental perception data (illuminance, wind speed, occlusion rate) and device attributes, it calculates the deviation between environmental parameters and optimal image capture conditions. This error term ensures that the perception map prioritizes marking areas that meet the optimal image capture conditions, providing a basis for subsequent environmental optimization decisions. Weighting coefficients. This demonstrates the strength of the role of environmental perception data in overall optimization. Each coefficient was determined through cross-validation in a power distribution scenario, balancing the contributions of different error terms to positioning optimization to achieve optimal overall positioning performance. The final output is a perception map. It includes device semantic tags, environmental parameters (lighting, wind speed, occlusion rate), and device-obstacle distance matrix, enabling UAVs to have a comprehensive and embodied perception of "device attributes, environmental status, and spatial relationships", providing accurate data support for decision-making.
[0039] In one possible implementation, step S2 employs HEDD hierarchical emergency-driven decision-making to enhance decision-making capabilities. Traditional decision-making is based on preset paths and cannot be dynamically adjusted. This invention achieves adaptive behavior through a decision model, adapting to the uncertainties of power distribution scenarios. Based on the perception map output by ES-SLAM, a four-layer decision model of "equipment priority - scenario urgency - path cost - behavior adaptation" is constructed. The core is to enable the UAV to autonomously select the optimal search and capture behavior based on real-time environmental parameters through an environment-behavior mapping function and a dynamic behavior adjustment mechanism, reflecting the intelligent characteristics of "perceiving the environment → adjusting behavior".
[0040] First, the hierarchical decision-making indicators are defined. Based on the existing indicators, a mapping relationship between the environment and behavior is established by adding decision-specific parameters: Firstly, equipment priority is weighted based on the degree of impact of distribution defects (insulators 1.5, conductors 1.2, towers 1.0, transformers 0.8), and dynamically adjusted in conjunction with the equipment status in the perception map. For example, when "suspected insulator damage" is detected, the priority is temporarily adjusted... Upgraded to version 2.0. Secondly, the expression for optimizing scenario emergency response is: ,in, Normalized wind speed ( ), To occlude normalized values ( ), Normalized value of equipment-obstacle safe distance ( This ensures that the scene assessment covers four specific environmental factors: occlusion, illumination, wind speed, and safe distance. Furthermore, the path cost optimization expression is: ,in, The energy cost of behavioral adjustment This represents the Euclidean distance between the drone's current position and the target position. This represents the number of obstacles in the path.
[0041] Secondly, there is the decision function. HEDD decision-making calculates a comprehensive adaptation score based on "device-scene-path-behavior". ,Right now The highest-scoring behavior combination is selected as the target image acquisition scheme. The specific process is as follows: First, in the first-level device priority filtering, [the following is retained]... The system prioritizes devices with suspected defects, such as those marked on the perception map as "candidate areas for insulator damage." Then, a second layer of scene emergency level filtering filters out... The scene, for The first layer of behavior pre-adjustment involves adjusting the scene-triggered behavior, such as lowering the altitude before reassessing when the wind speed is too high. Then, the third layer of behavior adaptation optimization optimizes the remaining device-scene combinations, traversing all possible shooting behaviors such as "straight-line flight + hovering shooting" and "obstacle avoidance + side shooting," calculating the optimal behavior for each combination. The system selects the behavior combination corresponding to the maximum value; finally, the fourth-layer dynamic adjustment trigger mechanism stipulates that if environmental parameters such as sudden occlusion occur during flight, the system will update in real time. and The behavior should be recalculated and adjusted, and the response time should be controlled within ≤0.5s. Represents behavioral fit. Path cost Finally, there is the emergency response mechanism. When a sudden defect (such as "insulator self-explosion") identified in real-time by the perception map is detected, the system will trigger a three-level embodied emergency response: The first-level response involves the UAV suspending its current mission and reconfirming the defect within ≤1 second using visible light + infrared multimodal data; if the defect is confirmed in the second confirmation, the system will proceed to the second-level response, temporarily shutting down the defective equipment. Upgraded to 2.5, updated. (For example, "priority detour") The system can complete emergency path replanning within ≤0.4s; if it encounters complex scenarios such as insufficient battery life of a single drone or multiple obstacles in mountainous areas, it will initiate a level 3 response, and dispatch surrounding drones through edge computing nodes to complete the response within ≤3s through collaborative actions such as "main drone shooting + slave drone supplementary lighting".
[0042] In one possible implementation, step S3 employs a RIQA real-time image quality assessment closed loop. Based on the correlation between the quality of the captured image and environmental parameters, a closed loop of "real-time assessment - behavioral feedback - dynamic optimization" is constructed. The core is to achieve feedback of "quality not meeting standards → reverse behavior adjustment" through a quality-environment-behavior mapping model, rather than simply reshooting. Traditional quality assessment only judges "meets standards / does not meet standards" without feedback adjustment capabilities. This invention achieves behavioral optimization through feedback, avoiding ineffective reshoots. First, there's the extraction of multi-dimensional quality features. Building upon the existing quality features, new features closely related to the environment were added. The first is sharpness, calculated using the Laplacian operator. While maintaining the original calculation method, an environment-related rule was added, based on the sharpness value... And distance parameter When the uniformity of illumination is... And light intensity parameters At that time, it was determined to be "uneven lighting caused by backlighting". In addition, a new quality feature, attitude stability, was added, which calculates attitude fluctuations such as roll and pitch angles during shooting based on IMU data. The calculation expression is: ,in, , Let t represent the changes in roll and pitch angles at time t, where T is the shooting duration.
[0043] Secondly, a quality assessment function is constructed. A comprehensive quality score integrating environmental parameters is then developed. Its calculation expression is: . This represents the image sharpness score, reflecting the focus quality and detail rendering of the captured image; This represents a light quality score, used to measure whether the ambient light in the shooting environment is sufficient and even. The coverage integrity score reflects the impact of drone shake on image stability during shooting. This represents the overall environmental score, which takes into account the impact of other environmental factors besides those mentioned above on the shooting quality.
[0044] Finally, there's the feedback closed-loop control logic. First, after the drone completes its search and capture operation, it calculates the overall quality score within 10ms. The causes of quality defects are then determined. Based on these causes, targeted behavioral adjustment plans are generated. If the defect is attributed to "too close distance," the adjustment is to "move back 1m and refocus." If it's attributed to "backlighting," the adjustment is to "fly 3m to the side to avoid the light source and adjust the shooting angle by 30°." If it's attributed to "strong wind shaking," the adjustment is to "reduce altitude by 2m, activate wind-resistant mode, and shorten shooting time." Next, the adjusted behavior is used for reshooting. If the overall quality score still doesn't meet the standard after two consecutive reshoots, it's marked as "requiring manual reshooting," and environmental parameters are recorded for subsequent model optimization. Finally, the "environmental parameters - behavioral adjustment - quality result" are stored in a local knowledge base, and the behavioral strategy is updated through reinforcement learning. For example, in backlit scenes, the "3m side-flying" behavioral strategy The value increased from 0.9 to 0.95, achieving continuous optimization of embodied decision-making.
[0045] This invention presents a multimodal collaborative autonomous image-seeking method for power distribution drones, realizing a full-link intelligent collaborative system. Centered on an "intelligent hub," it connects three modules: ES-SLAM perception, HEDD decision-making, and RIQA feedback, achieving an end-to-end intelligent link and supporting multi-drone collaborative image-seeking and self-evolutionary learning. Existing systems have independent modules and lack collaborative logic. This invention, through the intelligent hub, achieves data exchange and logical linkage, solving the problems of "disconnect between perception and decision-making, and separation between decision-making and feedback."
[0046] The implementation of the multimodal cooperative autonomous search and capture method for power distribution UAVs of this invention is divided into three stages: "offline calibration - online embodied search and capture - performance optimization". The relevant calibration and verification steps are as follows: The specific implementation process of the method of this invention is divided into two stages: "offline preparation" and "online inspection," involving five major steps: data collection, sample construction, model training, distillation optimization, and real-time identification. These are detailed in the two stages below: (1) Offline preparation stage (steps 1 to 4) Step 1: Hardware Configuration and Multimodal Data Acquisition. A DJI M350RTK drone was selected, equipped with LiDAR, a visible light camera, an IMU, and environmental sensors. Multimodal data, including environmental parameters, behavioral data, and the correspondence between quality results, was collected in typical complex scenarios such as mountainous areas with multiple obstructions, dense urban roads, and backlighting / strong winds. This data was used to construct an embodied decision dataset. Step 2: ES-SLAM Perception Model Training. Using multimodal data labeled with device semantic tags and environmental parameters, a perception model was trained to accurately identify and locate devices and the environment, constructing a perception map. Step 3: HEDD Embodied Decision Model Training. Based on "environmental parameters-behavior-quality" samples, a deep reinforcement learning method was used to train the decision model, enabling it to generate appropriate shooting behavior plans based on environmental conditions. Step 4: RIQA Feedback Model Training. Using shooting images labeled with quality features and defect attributions, a feedback model was trained to learn the mapping relationship between quality features, environmental parameters, and defect attributions, enabling the assessment of shooting quality and problem attribution.
[0047] (2) Offline preparation stage (steps 5-6) Step 5: After takeoff, ES-SLAM rapidly processes multimodal data to construct a perception map containing information on device location, environmental parameters, and safe distance, which is then visualized at the ground station. Subsequently, based on the environmental information acquired through perception, a decision model generates corresponding shooting plans for different scenarios. Step 6: After shooting, the feedback model evaluates the shooting quality. If it is substandard, the reasons are analyzed and the behavior is adjusted for reshooting. For example, in mountainous and heavily obstructed scenarios, obstacle avoidance and side shooting methods are selected, and the flight radius is adjusted based on quality feedback; in backlit and windy scenarios, side flight, reduced altitude, and wind-resistant modes are adopted, and the shooting time is shortened; in emergency situations such as insulator self-explosion, a rapid response is initiated, the path is replanned, and high-quality shooting is completed to meet the defect identification requirements.
[0048] Please see Figure 3 Another embodiment of the present invention also proposes a multimodal cooperative autonomous search and capture system for power distribution drones, comprising: Embodied perception map construction module 301 is used to acquire multimodal data collected by power distribution drones, extract semantic features and perform environmental interaction perception on the multimodal data, and construct an embodied perception map; The hierarchical emergency-driven decision-making module 302 is used to make hierarchical emergency-driven decisions based on the embodied perception map, and enable the power distribution drone to autonomously select the search and shooting behavior according to the real-time environmental parameters based on the decision results. The real-time image quality assessment and hunting behavior adjustment module 303 is used to acquire hunting images under the current hunting behavior of the power distribution drone, perform real-time image quality assessment on the hunting images, and obtain a comprehensive quality score based on the fused environmental parameters; determine the cause of quality defects based on the comprehensive quality score based on the fused environmental parameters, and adjust the hunting behavior of the power distribution drone based on the cause of quality defects, so as to realize multimodal collaborative autonomous hunting of the power distribution drone.
[0049] In one possible implementation, when the embodied perception map construction module 301 of this embodiment acquires multimodal data collected by the power distribution drone, the multimodal data includes lidar point clouds for sensing the spatial position of the device and obstacles, visible light images for sensing the semantics of the device and the lighting environment, inertial navigation data for sensing flight attitude and wind speed interference, and environmental perception data as inputs for embodied decision-making environment parameters; the environmental perception data includes light intensity, wind speed, occlusion rate, and device distance.
[0050] In one possible implementation, the embodied perception map construction module 301 of this embodiment of the invention preprocesses the collected multimodal data, the preprocessing including: VoxelGrid filtering is used to downsample the LiDAR point cloud, compressing the number of points to preserve the geometric features of the device and obstacles. The Retinex algorithm is used to eliminate the backlighting effect of the visible light image, and adaptive histogram equalization is used to improve the texture contrast of the device while extracting the image illumination intensity. The distance and occlusion rate between the device and obstacles are calculated based on the LiDAR point cloud, and the wind speed is retrieved by combining inertial navigation data to calibrate the environmental perception data. Finally, the LiDAR point cloud, visible light image, inertial navigation data, and environmental perception data are synchronized based on GPS timestamps.
[0051] In one possible implementation, when the embodied perception map construction module 301 of this embodiment of the invention performs semantic feature extraction and environmental interaction perception on multimodal data, it designs a semantic feature extraction network and establishes the association between devices and the environment, including visual semantic features and point cloud semantic features and their association with the environment. Specifically, for visual semantic features: a lightweight CNN is used to extract device semantic features from images, outputting a visual semantic probability map, adding "obstacle categories" corresponding to six semantic categories: "insulator, conductor, tower, transformer, obstacle, and background." A specific loss function ensures the accuracy of semantic classification between devices and obstacles, achieving device-obstacle perception differentiation. For point cloud features and their association with the environment: the visual semantic probability map is mapped to a point cloud coordinate system to obtain point cloud semantic labels, and then point cloud semantic features are extracted using PointNet. Simultaneously, the spatial distance between devices and obstacles is calculated to construct a "device-obstacle distance matrix." Feature matching loss ensures the accuracy of cross-frame semantic feature matching, while preventing drones from flying within dangerous distances between devices and obstacles, thus enhancing safety perception. By integrating semantic features, inertial navigation data, and environmental perception data into simultaneous localization and mapping (SLAM) localization optimization, a tightly coupled error function is constructed to output an embodied perception map containing environmental interaction information.
[0052] In the formula, Geometric error is used to measure the geometric spatial deviation during the positioning process. It is achieved by calculating the point cloud registration error between the current frame and the key frame. The iterative nearest point (ICP) algorithm is adopted to provide basic spatial accuracy for UAV positioning in complex power distribution scenarios. To mitigate semantic errors, semantic features are incorporated to understand the semantic information of objects in the environment. By analyzing the correlation between semantic features and positioning results, the semantic consistency of positioning is optimized, improving its adaptability to real-world application scenarios. Weighting coefficients... This reflects the importance of semantic features in the overall error function. The value was determined through cross-validation in a power distribution scenario; The inertial navigation error is calculated based on the inertial navigation pre-integration model to determine the difference between the predicted and observed poses. In power distribution environments, electromagnetic interference and wind speed are present; the inertial navigation error is used to suppress the resulting positioning drift and ensure the stability of the UAV's positioning in dynamic environments. Weighting coefficients are also included. This reflects the weighting percentage of inertial navigation data optimization. Environmental error, a embodied perception constraint, is used to calculate the deviation between environmental parameters and optimal image acquisition conditions based on the correlation between environmental perception data and device attributes. This environmental error ensures that the perception map prioritizes areas meeting the optimal image acquisition conditions, providing a basis for subsequent environmental optimization decisions. Weighting coefficients are also included. This reflects the strength of the role of environmental perception data in overall optimization; Each weighting coefficient is determined through cross-validation in a power distribution scenario to balance the contribution of different error terms to positioning optimization, thereby achieving optimal overall positioning performance; outputting a perception map. It includes device semantic tags, environmental parameters, and device-obstacle distance matrices, enabling UAVs to have a embodied perception of device attributes, environmental conditions, and spatial relationships, providing data support for decision-making.
[0053] In one possible implementation, the hierarchical emergency-driven decision-making module 302 of this embodiment of the invention constructs a four-layer decision-making model of "equipment priority - scenario emergency level - path cost - behavior adaptation", wherein the equipment priority is weighted based on the degree of impact of power distribution defects and dynamically adjusted in combination with the equipment status in the perception map; The expression for optimizing the emergency response level of a scenario is: ; In the formula, This is the normalized value of wind speed. To mask the normalized value, This is the normalized value for the equipment-obstacle safe distance; The path cost optimization expression is: ; In the formula, The energy cost of behavioral adjustment This represents the Euclidean distance between the drone's current position and the target position. This represents the number of obstacles in the path. , , These are the weighting coefficients for each part; The hierarchical emergency-driven decision-making process calculates a comprehensive adaptation score based on "device-scenario-path-behavior". The highest-rated combination of behaviors is selected as the target image acquisition scheme, and the calculation expression is as follows: ; In the first-level device priority filtering, retain The first layer of equipment prioritizes devices suspected of defects marked on the perception map; the second layer of scene urgency filtering will... Filter out the scene, for The first layer pre-adjusts the scene-triggered behavior; the third layer of behavior adaptation optimization, for the remaining device-scene combinations, traverses all possible search and capture behaviors, and calculates the behavior for each combination. The fourth layer of dynamic adjustment triggering mechanism stipulates that if a sudden change in environmental parameters occurs during flight, it will be updated in real time. and Recalculate and adjust behavior. Represents behavioral fit. It is the path cost; When a sudden defect in the real-time identification of the perception map is detected, a three-level embodied emergency response is triggered: Level 1 response involves suspending the current mission of the drone and reconfirming the defect; if the defect is confirmed on the second confirmation, the response proceeds to Level 2, temporarily suspending the defective equipment. Upgrade, update It will also complete emergency path replanning; if a single drone has insufficient battery life or a complex scenario, it will initiate a level three response, which will schedule surrounding drones through edge computing nodes and complete the response through collaborative behavior.
[0054] In one possible implementation, the real-time image quality assessment and target acquisition behavior adjustment module 303 of this embodiment calculates the comprehensive quality score of the fused environmental parameters according to the following formula:
[0055] In the formula, This represents the image sharpness score, reflecting the focus quality and detail rendering of the captured image; This represents a light quality score, used to measure whether the ambient light during shooting is sufficient and even. The coverage integrity score reflects the impact of drone shake on image stability during shooting. Represents attitude stability score; The overall environmental score takes into account the impact of other environmental factors besides those mentioned above on the shooting quality. The attitude stability score The calculation expression is as follows:
[0056] In the formula, , These represent the changes in roll and pitch angles at time t, respectively, where T is the shooting duration.
[0057] In one possible implementation, if the real-time image quality assessment and target acquisition behavior adjustment module 303 of this embodiment determines that the quality defect is caused by "too close distance", it performs the behavior adjustment of "drone retreating 1m + refocusing"; if the quality defect is caused by "backlight", it performs the behavior adjustment of "drone flying sideways to avoid the light source + adjusting the shooting angle"; if the quality defect is caused by "strong wind shaking", it performs the behavior adjustment of "drone reducing altitude + turning on wind-resistant mode + shortening shooting time". After performing the adjusted behavior, the target acquisition is re-shot. If the comprehensive quality score after fusing environmental parameters is still not up to standard after continuous re-shooting, it is marked as "requiring manual re-shooting", and the environmental parameters are recorded. The "environmental parameters-behavior adjustment-quality result" data is stored in the local knowledge base, and the behavior strategy is updated through reinforcement learning. This enables continuous optimization of embodied decision-making.
[0058] The multimodal cooperative autonomous search and capture method and system for power distribution UAVs in this invention are optimized in the following aspects: 1. Building Perception Capabilities: By integrating the semantic features and geometric information of power distribution equipment through ES-SLAM technology, the positioning drift error is ≤0.1m and the accuracy of equipment semantic map construction is ≥98%. This enables UAVs to perceive equipment attributes, environmental obstacles, and spatial relationships (such as "insulator position, distance to surrounding obstacles, and light intensity"), providing environmental data support for decision-making.
[0059] 2. Achieve decision optimization: Construct an "environmental parameter-behavioral strategy" mapping model through the HEDD decision algorithm. Dynamically adjust the image acquisition behavior (path, angle, height) based on equipment priority (e.g., insulator > conductor > tower) and real-time scene (e.g., shading, lighting, wind speed). The image acquisition response time for key defects (e.g., insulator self-explosion) is ≤5s, and the image acquisition success rate in complex scenes is ≥95%.
[0060] 3. Improve the feedback loop: Real-time evaluation of image quality through RIQA closed-loop technology, with a quality compliance rate of ≥95%, and the establishment of a "quality defect - behavior adjustment" feedback mechanism (such as "low clarity → move closer to the target device" and "uneven lighting → adjust the shooting angle"), to achieve a embodied closed loop of "evaluation - feedback - behavior optimization" and avoid secondary reshoots.
[0061] 4. Enhance adaptability to complex scenarios: By collaboratively building an intelligent link of "perception-decision-execution-feedback", the drone can continuously optimize its search and shooting behavior according to environmental changes. In complex scenarios such as mountainous areas with many obstructions and dense urban routes, the success rate of autonomous search and shooting is ≥90%, meeting actual operation and maintenance needs.
[0062] Another embodiment of the present invention also proposes an electronic device, including a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the aforementioned multimodal cooperative autonomous shooting method for power distribution drones.
[0063] Another embodiment of the present invention also proposes a computer-readable storage medium storing at least one instruction, which, when executed by a processor, implements the aforementioned multimodal cooperative autonomous search and capture method for power distribution UAVs.
[0064] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. For ease of explanation, the above content only shows the parts related to the embodiments of the present invention; for specific technical details not disclosed, please refer to the method section of the embodiments of the present invention. This computer-readable storage medium is non-transitory and can be stored in storage devices formed by various electronic devices, enabling the execution process described in the method of the embodiments of the present invention.
[0065] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0066] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0067] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0068] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A multimodal cooperative autonomous search and capture method for power distribution drones, characterized in that, include: Acquire multimodal data collected by power distribution drones, extract semantic features and perform environmental interaction perception on the multimodal data, and construct an embodied perception map; Based on embodied perception maps, hierarchical emergency-driven decision-making is carried out, and the power distribution drones are enabled to autonomously select search and capture behaviors according to real-time environmental parameters based on the decision results. The system acquires images captured by the power distribution drone under its current search and capture behavior, performs real-time image quality assessment on the images, and obtains a comprehensive quality score by fusing environmental parameters. Based on the comprehensive quality score of the fused environmental parameters, the system determines the cause of quality defects and adjusts the search and capture behavior of the power distribution drone based on the cause of the quality defects, thereby realizing multimodal collaborative autonomous search and capture by the power distribution drone. In the step of making hierarchical emergency-driven decisions based on the embodied perception map and enabling the power distribution drone to autonomously select its search and capture behavior according to real-time environmental parameters based on the decision results, a four-layer decision model is constructed, which includes equipment priority, scene emergency level, path cost, and behavior adaptation. Among them, the equipment priority is weighted based on the degree of impact of power distribution defects and dynamically adjusted in combination with the equipment status in the perception map. The expression for optimizing the emergency response level of a scenario is: ; In the formula, This is the normalized value of wind speed. To mask the normalized value, This is the normalized value for the equipment-obstacle safe distance; The path cost optimization expression is: ; In the formula, The energy cost of behavioral adjustment This represents the Euclidean distance between the drone's current position and the target position. This represents the number of obstacles in the path. , , These are the weighting coefficients for each part; The hierarchical emergency-driven decision-making process utilizes a comprehensive adaptation score based on computing devices, scenarios, paths, and behaviors. The highest-scoring combination of behaviors is selected as the target image acquisition strategy, and the calculation expression is as follows: In the formula, Defective equipment; In the first-level device priority filtering, retain The first layer of equipment prioritizes devices suspected of defects marked on the perception map; the second layer of scene urgency filtering will... Filter out the scene, for The first layer pre-adjusts the scene-triggered behavior; the third layer of behavior adaptation optimization, for the remaining device-scene combinations, traverses all possible search and capture behaviors, and calculates the behavior for each combination. The fourth layer of dynamic adjustment triggering mechanism stipulates that if a sudden change in environmental parameters occurs during flight, it will be updated in real time. and Recalculate and adjust behavior. Represents behavioral fit. It is the path cost; When a sudden defect in the real-time identification of the perception map is detected, a three-level embodied emergency response is triggered: Level 1 response involves suspending the current mission of the drone and reconfirming the defect; if the defect is confirmed on the second confirmation, the response proceeds to Level 2, temporarily suspending the defective equipment. Upgrade, update It will also complete emergency path replanning; if a single drone has insufficient battery life or a complex scenario, it will initiate a level three response, which will schedule surrounding drones through edge computing nodes and complete the response through collaborative behavior.
2. The multimodal cooperative autonomous search and capture method for power distribution UAVs according to claim 1, characterized in that, In the step of acquiring multimodal data collected by the power distribution drone, the multimodal data includes lidar point clouds for sensing the spatial position of the device and obstacles, visible light images for sensing the semantics of the device and the lighting environment, inertial navigation data for sensing flight attitude and wind speed interference, and environmental perception data as input for embodied decision-making environment parameters. The environmental sensing data includes light intensity, wind speed, occlusion rate, and device distance.
3. The multimodal cooperative autonomous acquisition and shooting method for power distribution UAVs according to claim 2, characterized in that, It also includes a preprocessing step for the acquired multimodal data, which includes: VoxelGrid filtering is used to downsample the LiDAR point cloud, compressing the number of points to preserve the geometric features of the device and obstacles. The Retinex algorithm is used to eliminate the backlighting effect of the visible light image, and adaptive histogram equalization is used to improve the texture contrast of the device while extracting the image illumination intensity. The distance and occlusion rate between the device and obstacles are calculated based on the LiDAR point cloud, and the wind speed is retrieved by combining inertial navigation data to calibrate the environmental perception data. Finally, the LiDAR point cloud, visible light image, inertial navigation data, and environmental perception data are synchronized based on GPS timestamps.
4. The multimodal cooperative autonomous search and capture method for power distribution UAVs according to claim 1, characterized in that, The steps for extracting semantic features from multimodal data and perceiving environmental interactions to construct an embodied perception map include: designing a semantic feature extraction network and establishing the association between devices and the environment, including visual semantic features and point cloud semantic features associated with the environment; wherein, visual semantic features: a lightweight CNN is used to extract device semantic features from images, outputting a visual semantic probability map, adding obstacle categories, corresponding to 6 semantic categories: insulator, conductor, tower, transformer, obstacle, and background, and using a specific loss function to ensure the accuracy of device and obstacle semantic classification, realizing the perception and distinction between devices and obstacles; point cloud features associated with the environment: the visual semantic probability map is mapped to the point cloud coordinate system to obtain point cloud semantic labels, and then point cloud semantic features are extracted using PointNet, while calculating the spatial distance between devices and obstacles, constructing a device-obstacle distance matrix, and using feature matching loss to ensure the accuracy of semantic feature cross-frame matching, while avoiding drones flying within the dangerous distance between devices and obstacles, thus enhancing safety perception; By integrating semantic features, inertial navigation data, and environmental perception data into simultaneous localization, map construction and SLAM localization optimization are completed. A tightly coupled error function is constructed, and an embodied perception map containing environmental interaction information is output. In the formula, Geometric error is used to measure the geometric spatial deviation during the positioning process. It is achieved by calculating the point cloud registration error between the current frame and the key frame. The iterative nearest point (ICP) algorithm is adopted to provide basic spatial accuracy for UAV positioning in complex power distribution scenarios. To mitigate semantic errors, semantic features are incorporated to understand the semantic information of objects in the environment. By analyzing the correlation between semantic features and positioning results, the semantic consistency of positioning is optimized, improving its adaptability to real-world application scenarios. Weighting coefficients... This reflects the importance of semantic features in the overall error function. The value was determined through cross-validation in a power distribution scenario; The inertial navigation error is calculated based on the inertial navigation pre-integration model to determine the difference between the predicted and observed poses. In power distribution environments, electromagnetic interference and wind speed are present; the inertial navigation error is used to suppress the resulting positioning drift and ensure the stability of the UAV's positioning in dynamic environments. Weighting coefficients are also included. This reflects the weighting percentage of inertial navigation data optimization. Environmental error, a embodied perception constraint, is used to calculate the deviation between environmental parameters and optimal image acquisition conditions based on the correlation between environmental perception data and device attributes. This environmental error ensures that the perception map prioritizes areas meeting the optimal image acquisition conditions, providing a basis for subsequent environmental optimization decisions. Weighting coefficients are also included. This reflects the strength of the role of environmental perception data in overall optimization; Each weighting coefficient is determined through cross-validation in a power distribution scenario to balance the contribution of different error terms to positioning optimization, thereby achieving optimal overall positioning performance; outputting a perception map. It includes device semantic tags, environmental parameters, and device-obstacle distance matrices, enabling UAVs to have a embodied perception of device attributes, environmental conditions, and spatial relationships, providing data support for decision-making.
5. The multimodal cooperative autonomous hunting and shooting method for power distribution UAVs according to claim 1, characterized in that, In the step of acquiring the search and capture images under the current search and capture behavior of the power distribution drone, performing real-time image quality assessment on the search and capture images, and obtaining a comprehensive quality score by fusing environmental parameters, the comprehensive quality score by fusing environmental parameters is calculated according to the following formula: In the formula, This represents the image sharpness score, reflecting the focus quality and detail rendering of the captured image; This represents a light quality score, used to measure whether the ambient light during shooting is sufficient and even. The coverage integrity score reflects the impact of drone shake on image stability during shooting. Represents attitude stability score; This represents the overall environmental score, which takes into account the impact of other environmental factors besides those mentioned above on the shooting quality.
6. The multimodal cooperative autonomous search and capture method for power distribution UAVs according to claim 5, characterized in that, The attitude stability score The calculation expression is as follows: In the formula, , These represent the changes in roll and pitch angles at time t, respectively, where T is the shooting duration.
7. The multimodal cooperative autonomous search and capture method for power distribution UAVs according to claim 1, characterized in that, In the step of determining the cause of quality defects based on the comprehensive quality score of the fused environmental parameters, and adjusting the search and shooting behavior of the power distribution drone based on the cause of the quality defects to achieve multimodal collaborative autonomous search and shooting of the power distribution drone: If the quality defect is caused by the distance being too close, the drone will be moved back 1 meter and refocused; if the quality defect is caused by backlighting, the drone will be moved to the side to avoid the light source and the shooting angle will be adjusted; if the quality defect is caused by strong wind shaking, the drone will be moved to a lower altitude, the wind-resistant mode will be activated, and the shooting time will be shortened. After implementing the adjusted behavior, the system re-captures the image. If the overall quality score after merging environmental parameters still fails to meet the standard after continuous re-captures, it is marked as requiring manual re-capture. The environmental parameters, behavior adjustment, and quality result data are then stored in the local knowledge base. Through reinforcement learning, the behavior strategy is updated to achieve continuous optimization of embodied decision-making.
8. A multimodal cooperative autonomous search and capture system for power distribution unmanned aerial vehicles (UAVs), characterized in that, include: The embodied perception map construction module is used to acquire multimodal data collected by power distribution drones, extract semantic features from the multimodal data and perceive environmental interactions to construct an embodied perception map. The hierarchical emergency-driven decision-making module is used to make hierarchical emergency-driven decisions based on the embodied perception map, and enable the power distribution drone to autonomously select its search and shooting behavior according to real-time environmental parameters based on the decision results. The real-time image quality assessment and hunting behavior adjustment module is used to acquire hunting images under the current hunting behavior of the power distribution drone, perform real-time image quality assessment on the hunting images, and obtain a comprehensive quality score based on the fused environmental parameters; determine the cause of quality defects based on the comprehensive quality score based on the fused environmental parameters, and adjust the hunting behavior of the power distribution drone based on the cause of quality defects, so as to realize multimodal collaborative autonomous hunting of the power distribution drone. The hierarchical emergency-driven decision-making module constructs a four-layer decision-making model: equipment priority, scenario urgency, path cost, and behavior adaptation. Among them, the equipment priority is weighted based on the degree of impact of power distribution defects and dynamically adjusted in combination with the equipment status in the perception map. The expression for optimizing the emergency response level of a scenario is: ; In the formula, This is the normalized value of wind speed. To mask the normalized value, This is the normalized value for the equipment-obstacle safe distance; The path cost optimization expression is: ; In the formula, The energy cost of behavioral adjustment This represents the Euclidean distance between the drone's current position and the target position. This represents the number of obstacles in the path. , , These are the weighting coefficients for each part; The hierarchical emergency-driven decision-making process utilizes a comprehensive adaptation score based on computing devices, scenarios, paths, and behaviors. The highest-scoring combination of behaviors is selected as the target image acquisition strategy, and the calculation expression is as follows: In the formula, Defective equipment; In the first-level device priority filtering, retain The first layer of equipment prioritizes devices suspected of defects marked on the perception map; the second layer of scene urgency filtering will... Filter out the scene, for The first layer pre-adjusts the scene-triggered behavior; the third layer of behavior adaptation optimization, for the remaining device-scene combinations, traverses all possible search and capture behaviors, and calculates the behavior for each combination. The fourth layer of dynamic adjustment triggering mechanism stipulates that if a sudden change in environmental parameters occurs during flight, it will be updated in real time. and Recalculate and adjust behavior. Represents behavioral fit. It is the path cost; When a sudden defect in the real-time identification of the perception map is detected, a three-level embodied emergency response is triggered: Level 1 response involves suspending the current mission of the drone and reconfirming the defect; if the defect is confirmed on the second confirmation, the response proceeds to Level 2, temporarily suspending the defective equipment. Upgrade, update It will also complete emergency path replanning; if a single drone has insufficient battery life or a complex scenario, it will initiate a level three response, which will schedule surrounding drones through edge computing nodes and complete the response through collaborative behavior.
9. The power distribution UAV multimodal cooperative autonomous search and capture system according to claim 8, characterized in that, When the embodied perception map construction module acquires multimodal data collected by the power distribution drone, the multimodal data includes lidar point clouds for sensing the spatial location of the device and obstacles, visible light images for sensing the semantics of the device and the lighting environment, inertial navigation data for sensing flight attitude and wind speed interference, and environmental perception data as input for embodied decision-making environment parameters. The environmental sensing data includes light intensity, wind speed, occlusion rate, and device distance.
10. The power distribution UAV multimodal cooperative autonomous seeking and shooting system according to claim 9, characterized in that, The embodied perception map construction module preprocesses the collected multimodal data, including: VoxelGrid filtering is used to downsample the LiDAR point cloud, compressing the number of points to preserve the geometric features of the device and obstacles. The Retinex algorithm is used to eliminate the backlighting effect of the visible light image, and adaptive histogram equalization is used to improve the texture contrast of the device while extracting the image illumination intensity. The distance and occlusion rate between the device and obstacles are calculated based on the LiDAR point cloud, and the wind speed is retrieved by combining inertial navigation data to calibrate the environmental perception data. Finally, the LiDAR point cloud, visible light image, inertial navigation data, and environmental perception data are synchronized based on GPS timestamps.
11. The power distribution UAV multimodal cooperative autonomous seeking and shooting system according to claim 8, characterized in that, When the embodied perception map construction module extracts semantic features from multimodal data and perceives environmental interactions, it designs a semantic feature extraction network and establishes the association between devices and the environment, including visual semantic features and point cloud semantic features. Specifically, for visual semantic features, a lightweight CNN is used to extract device semantic features from images, outputting a visual semantic probability map. An obstacle category is added, corresponding to six semantic categories: insulator, conductor, tower, transformer, obstacle, and background. A specific loss function ensures the accuracy of device and obstacle semantic classification, achieving device-obstacle perception differentiation. For point cloud features and environment association, the visual semantic probability map is mapped to a point cloud coordinate system to obtain point cloud semantic labels. PointNet is then used to extract point cloud semantic features, while simultaneously calculating the spatial distance between devices and obstacles to construct a device-obstacle distance matrix. Feature matching loss ensures the accuracy of cross-frame semantic feature matching and prevents drones from flying within dangerous distances between devices and obstacles, enhancing safety perception. By integrating semantic features, inertial navigation data, and environmental perception data into simultaneous localization, map construction and SLAM localization optimization are completed. A tightly coupled error function is constructed, and an embodied perception map containing environmental interaction information is output. In the formula, Geometric error is used to measure the geometric spatial deviation during the positioning process. It is achieved by calculating the point cloud registration error between the current frame and the key frame. The iterative nearest point (ICP) algorithm is adopted to provide basic spatial accuracy for UAV positioning in complex power distribution scenarios. To mitigate semantic errors, semantic features are incorporated to understand the semantic information of objects in the environment. By analyzing the correlation between semantic features and positioning results, the semantic consistency of positioning is optimized, improving its adaptability to real-world application scenarios. Weighting coefficients... This reflects the importance of semantic features in the overall error function. The value was determined through cross-validation in a power distribution scenario; The inertial navigation error is calculated based on the inertial navigation pre-integration model to determine the difference between the predicted and observed poses. In power distribution environments, electromagnetic interference and wind speed are present; the inertial navigation error is used to suppress the resulting positioning drift and ensure the stability of the UAV's positioning in dynamic environments. Weighting coefficients are also included. This reflects the weighting percentage of inertial navigation data optimization. Environmental error, a embodied perception constraint, is used to calculate the deviation between environmental parameters and optimal image acquisition conditions based on the correlation between environmental perception data and device attributes. This environmental error ensures that the perception map prioritizes areas meeting the optimal image acquisition conditions, providing a basis for subsequent environmental optimization decisions. Weighting coefficients are also included. This reflects the strength of the role of environmental perception data in overall optimization; Each weighting coefficient is determined through cross-validation in a power distribution scenario to balance the contribution of different error terms to positioning optimization, thereby achieving optimal overall positioning performance; outputting a perception map. It includes device semantic tags, environmental parameters, and device-obstacle distance matrices, enabling UAVs to have a embodied perception of device attributes, environmental conditions, and spatial relationships, providing data support for decision-making.
12. The power distribution UAV multimodal cooperative autonomous seeking and shooting system according to claim 8, characterized in that, The real-time image quality assessment and target acquisition behavior adjustment module calculates the comprehensive quality score of the fused environmental parameters using the following formula: In the formula, This represents the image sharpness score, reflecting the focus quality and detail rendering of the captured image; This represents a light quality score, used to measure whether the ambient light during shooting is sufficient and even. The coverage integrity score reflects the impact of drone shake on image stability during shooting. Represents attitude stability score; The overall environmental score takes into account the impact of other environmental factors besides those mentioned above on the shooting quality. The attitude stability score The calculation expression is as follows: In the formula, , These represent the changes in roll and pitch angles at time t, respectively, where T is the shooting duration.
13. The power distribution UAV multimodal cooperative autonomous search and capture system according to claim 8, characterized in that, If the real-time image quality assessment and shooting behavior adjustment module determines that the quality defect is caused by the distance being too close, it will adjust the behavior by moving the drone back 1m and refocusing; if it determines that the quality defect is caused by backlighting, it will adjust the behavior by moving the drone to the side to avoid the light source and adjusting the shooting angle; if it determines that the quality defect is caused by strong wind shaking, it will adjust the behavior by reducing the drone's altitude, turning on the wind-resistant mode, and shortening the shooting time. After implementing the adjusted behavior, the system re-captures the image. If the overall quality score after merging environmental parameters still fails to meet the standard after continuous re-captures, it is marked as requiring manual re-capture. The environmental parameters, behavior adjustment, and quality result data are then stored in the local knowledge base. Through reinforcement learning, the behavior strategy is updated to achieve continuous optimization of embodied decision-making.
14. An electronic device, characterized in that, It includes a processor and a memory, the processor being used to execute a computer program stored in the memory to implement the multimodal cooperative autonomous hunting and shooting method for a power distribution UAV as described in any one of claims 1 to 7.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which, when executed by a processor, implements the multimodal cooperative autonomous hunting and shooting method for power distribution drones as described in any one of claims 1 to 7.