Unmanned aerial vehicle electric power inspection obstacle avoidance method and system
By optimizing the path through weighted fusion of multimodal sensor data and dynamic window algorithm, combined with pre-trained models to identify obstacles, the problems of insufficient robustness of sensor fusion and inflexible path planning in UAV power inspections are solved, achieving efficient obstacle avoidance and safe flight in complex environments.
Patent Information
- Application Number
- CN202510739626.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-10-17
AI Technical Summary
Existing drone power inspection technology suffers from insufficient sensor fusion robustness, low obstacle detection accuracy, inflexible path planning, and poor real-time edge computing in complex environments, resulting in low reliability of obstacle avoidance decisions and affecting the safety and efficiency of inspection tasks.
The system uses weighted fusion of multimodal sensor data to generate a global path and combines it with a dynamic window algorithm to generate a local obstacle avoidance path. It uses a pre-trained target detection model to identify unlabeled obstacles, and improves the robustness and stability of path planning through multimodal data.
It improves the reliability of obstacle avoidance decisions and inspection efficiency of drones in complex environments, can effectively identify unmarked obstacles, ensure the stability and safety of the flight path, and meet millisecond-level obstacle avoidance requirements.
Smart Images

Figure CN120803016A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to, but is not limited to, the technical field of image processing, and particularly relates to a UAV power inspection obstacle avoidance method and system. BACKGROUND
[0002] With the development of smart grid, UAV power inspection has become a key means of high-voltage line monitoring. However, the complex flight environment brings many challenges to the real-time perception and obstacle avoidance of the UAV. The power equipment is densely spaced, and there are dynamic obstacles such as flying birds, floating objects and mobile devices. These obstacles have the characteristics of randomness, diversity and non-structure, which makes the UAV prone to collision accidents during flight, seriously affecting the smooth progress of the inspection task.
[0003] The prior art has defects in many aspects. In terms of sensor fusion, the traditional method relies on linear Kalman filtering, which is difficult to effectively handle dynamic noise in complex three-dimensional space. When the sensor fails or is disturbed, the system stability is poor, the fusion weight fluctuates greatly, and the obstacle avoidance decision reliability is reduced. For example, the performance of laser radar decreases in foggy weather, visual detection is obviously affected by light, and the complementarity of lightweight adaptation and multi-modal data is not fully explored. In the obstacle detection link, the traditional model relies on fixed vocabulary and cannot identify unannotated temporary obstacles. The detection accuracy of small or low-contrast obstacles is insufficient, and the positioning robustness is not improved combined with multi-modal data. In terms of path planning, the global algorithm has redundant path, lacks flexibility, the local algorithm lacks global optimization, and relies on a single sensor, which is easy to cause positioning drift or obstacle misjudgment. In addition, the real-time detection of open vocabulary by edge device deployment has the problem of computing power, the deep learning model is difficult to deploy on embedded platform, the traditional multi-modal fusion method has poor real-time performance, and cannot meet the millisecond-level obstacle avoidance requirement. These deficiencies restrict the safety and efficiency of the UAV in complex ultra-high voltage power line inspection. SUMMARY
[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0005] The embodiment of the present application provides a UAV power inspection obstacle avoidance method and system, which improves robustness through multi-modal data weighted fusion, improves trajectory stability by combining global and local path optimization, and uses a pre-trained model to identify unannotated obstacles, thereby enhancing the safety and efficiency of UAV power inspection in complex environments.
[0006] In a first aspect, the embodiments of the present application provide a method for power inspection obstacle avoidance of a UAV, characterized in that: multi-modal sensor data of a UAV is acquired, and the multi-modal sensor data is weighted and fused to obtain multi-modal fusion data; a global path is generated based on the multi-modal fusion data; a local obstacle avoidance path is generated based on the global path by using a dynamic window algorithm; the global path and the local obstacle avoidance path are coordinated to obtain an initial flight path; real-time environment images collected when the UAV flies along the initial flight path are acquired, a pre-trained target detection model is called to detect the real-time environment images to obtain path obstacle avoidance information; and the initial flight path is adjusted according to the path obstacle avoidance information to obtain a target flight path.
[0007] In combination with the first aspect, in an embodiment of the present application, the multi-modal sensor data includes global positioning data, three-dimensional point cloud data and motion attitude data; the multi-modal sensor data is weighted and fused to obtain multi-modal fusion data, including: the confidence of the global positioning data, the three-dimensional point cloud data and the motion attitude data is calculated respectively to obtain corresponding first confidence, second confidence and third confidence; the first confidence, the second confidence and the third confidence are weighted and fused to obtain multi-modal fusion data.
[0008] In combination with the first aspect, in an embodiment of the present application, the global path is generated based on the multi-modal fusion data, including: the three-dimensional point cloud data in the multi-modal fusion data is converted into a three-dimensional map, and a two-dimensional grid map at a preset flight height is extracted based on the three-dimensional map; in the passable area of the two-dimensional grid map, a target bias search strategy is used to generate a sampling point, a target guiding direction of the sampling point is calculated by an attractive force function to obtain a sampling point sequence with target guidance; based on the sampling point sequence with target guidance, a global path is generated on the two-dimensional grid map in combination with a collision detection algorithm.
[0009] In combination with the first aspect, in an embodiment of the present application, the sampling point sequence includes a starting sampling point and an ending sampling point; the global path is generated on the two-dimensional grid map in combination with a collision detection algorithm based on the sampling point sequence with target guidance, including: starting from the starting sampling point, the sampling points with target guidance are sequentially connected to generate a preliminary global path; a breadth-first search strategy is used to sequentially perform collision detection on the sampling points on the preliminary global path starting from the starting sampling point to obtain a detection result; if the detection result represents that the current sampling point has no collision, the next sampling point is continuously detected; if the detection result represents that the current sampling point has collision, the previous sampling point of the current sampling point is added to a path table; until the ending sampling point is traversed, a global path is obtained according to the path table.
[0010] In an embodiment of the first aspect, based on the global path, the local obstacle-avoiding path is generated by using a dynamic window algorithm, which includes: decomposing the global path into a series of sub-target points; introducing a target point evaluation function into the dynamic window algorithm, combining the constraints of the sub-target points to perform forward trajectory prediction on the velocity and angular velocity combination in the sampling space, to generate a plurality of predicted motion trajectories; constructing a trajectory evaluation function, and calculating the evaluation score of each of the predicted motion trajectories by using the trajectory evaluation function; and selecting a local obstacle-avoiding path from the predicted motion trajectories according to the evaluation score.
[0011] In an embodiment of the first aspect, after the global path is obtained, the method further includes: extracting a control point sequence from the global path; performing smooth interpolation processing on the control point sequence by using a curve algorithm to obtain a smooth curve; and adjusting the parameters of the curve algorithm to obtain an optimized global path based on the smooth curve.
[0012] In an embodiment of the first aspect, the pre-trained target detection model includes a cross-stage local network layer, an image pooling attention module, a text encoding module, a semantic comparison module, and a feature fusion module connected with each other, the cross-stage local network layer is used to calculate multi-modal features of the real-time environment image features and text features; the image pooling attention module is used to perform multi-scale pooling processing on the real-time environment image to obtain multi-scale feature representation; the feature fusion module is used to fuse the multi-modal features and the multi-scale feature representation to obtain multi-modal semantic feature representation; the text encoding module is used to generate a semantic feature library containing multiple types of targets; and the semantic comparison module is used to compare the multi-modal semantic feature representation with the semantic feature library to obtain the path obstacle-avoiding information.
[0013] In an embodiment of the first aspect, the initial flight path is adjusted according to the path obstacle-avoiding information to obtain a target flight path, which includes: taking a path node in the initial flight path as a control point, and constructing a continuous and smooth initial trajectory by using a curve algorithm; combining the kinematic constraints of the unmanned aerial vehicle, constructing a dynamic obstacle-avoiding constraint condition based on the positions and obstacle-avoiding weights of obstacles in the path obstacle-avoiding information, and integrating the dynamic obstacle-avoiding constraint condition into the initial trajectory to obtain a motion trajectory with constraint conditions; iteratively optimizing the control point of the curve algorithm to obtain an optimized smooth trajectory; and when local obstacle avoidance is triggered, updating the control point of the curve algorithm and refitting the smooth trajectory to obtain a target flight path.
[0014] In a second aspect, the embodiments of the present application provide a UAV power inspection obstacle avoidance system, applied to the UAV power inspection obstacle avoidance method described above. The obstacle avoidance system comprises a multi-modal perception module, a data processing module, and a path planning module. The multi-modal perception module is configured to obtain multi-modal sensor data of the UAV, and to perform weighted fusion on the multi-modal sensor data to obtain multi-modal fusion data. The data processing module is configured to generate a global path based on the multi-modal fusion data, and to generate a local obstacle avoidance path based on the global path using a dynamic window algorithm. The path planning module is configured to perform path coordination on the global path and the local obstacle avoidance path to obtain an initial flight path, to obtain real-time environment images collected by the UAV when flying along the initial flight path, to call a pre-trained target detection model to detect the real-time environment images to obtain path obstacle avoidance information, and to adjust the initial flight path based on the path obstacle avoidance information to obtain a target flight path.
[0015] In combination with the second aspect, in an embodiment of the present application, the multi-modal perception module comprises a laser radar sensor, an inertial measurement unit, a positioning sensor, and a binocular vision camera. The laser radar sensor is configured to construct a three-dimensional point cloud map and extract obstacle contour information. The inertial measurement unit is configured to monitor the attitude and acceleration of the UAV in real time. The positioning sensor is configured to provide position information of the UAV in a global coordinate system. The binocular vision camera is configured to assist in dynamic target recognition and distance estimation.
[0016] In the embodiments of the present application, first, multi-modal sensor data of the UAV is obtained, and the multi-modal sensor data is weighted and fused to obtain multi-modal fusion data. Then, a global path is generated based on the multi-modal fusion data. Next, a local obstacle avoidance path is generated based on the global path using a dynamic window algorithm. Subsequently, path coordination is performed on the global path and the local obstacle avoidance path to obtain an initial flight path. Then, real-time environment images collected by the UAV when flying along the initial flight path are obtained, a pre-trained target detection model is called to detect the real-time environment images to obtain path obstacle avoidance information. Thereafter, the initial flight path is adjusted based on the path obstacle avoidance information to obtain a target flight path. The embodiments of the present application improve robustness through weighted fusion of multi-modal sensor data, optimize trajectory stability by generating a global path based on the fusion data and generating a local obstacle avoidance path using a dynamic window algorithm, and enhance the safety and efficiency of UAV power inspection in complex environments by identifying obstacles using a pre-trained target detection model. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a flowchart of the UAV power inspection obstacle avoidance method provided by an embodiment of the present application; Figure 2 is a schematic diagram of the UAV power inspection obstacle avoidance system provided by an embodiment of the present application Figure 1The specific flow chart of step 110; Figure 3 The specific flow chart of step 120; Figure 1 The specific flow chart of step 110; Figure 4 The specific flow chart of step 120; Figure 5 The specific flow chart of step 110; Figure 3 The specific flow chart of step 330; Figure 6 The specific flow chart of step 130; Figure 1 The specific flow chart of step 130; Figure 7 The specific flow chart of step 130; Figure 8 The specific flow chart of step 130; Figure 9 The specific flow chart of step 130; DETAILED DESCRIPTION
[0018] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0019] It should be noted that although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that in the flow chart. The terms "first", "second", etc. in the specification and claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the structure, proportion, size, etc. shown in the drawings of the specification are only used to cooperate with the content disclosed in the specification, so as to be understood and read by those skilled in the art, and do not have technical significance, any modification of the structure, change of the proportion relationship or adjustment of the size, without affecting the effect and purpose of the present application, should still fall within the scope of the technical content disclosed in the present application. At the same time, the terms such as "up", "down", "left", "right", "middle" and "one" in the specification are only used to make the description clear, and not to limit the scope of the present application, the change or adjustment of the relative relationship, without substantially changing the technical content, is also regarded as the scope of the present application.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to be limiting of this application.
[0021] With the rapid development of smart grids, unmanned aerial vehicle (UAV) power inspection has become a key means of high-voltage line monitoring. However, the complex flight environment brings many challenges to the real-time perception and obstacle avoidance of the UAV. The power equipment is densely spaced, and there are dynamic obstacles such as flying birds, floating objects, and mobile devices. These obstacles have random, diverse, and unstructured characteristics, making it easy for the UAV to collide during flight, seriously affecting the smooth progress of the inspection task.
[0022] The prior art has the problem of insufficient robustness in sensor fusion. Traditional multi-sensor fusion methods rely on linear Kalman filtering, which is difficult to effectively model dynamic noise in complex three-dimensional space. When the sensor fails or is disturbed by noise, the stability of the system is poor, and the fusion weight will fluctuate sharply, thereby reducing the reliability of the obstacle avoidance decision. For example, in foggy weather, the performance of the laser radar will be greatly reduced due to the scattering and absorption of laser beams by particles in the air; and the visual detection is obviously affected by light changes, and the detection accuracy and accuracy will be severely restricted in strong light, backlight, or low light conditions. Moreover, the prior art has not been fully explored in terms of lightweight adaptation in dynamic obstacle scenarios, and the complementarity of multi-modal data such as infrared thermal imaging and visible light in UAV power inspection has not been fully considered, and the multi-source information has not been effectively utilized to improve the perception performance and reliability of the system.
[0023] In the obstacle detection link, the existing method has obvious defects. Traditional target detection models rely on fixed vocabulary and can only identify predefined categories of obstacles, and cannot cope with unannotated temporary obstacles such as flying birds and suspended foreign objects in power inspection. At the same time, for small or low-contrast obstacles, the detection accuracy is insufficient, and it is difficult to meet the actual inspection requirements. In addition, the existing detection method does not combine multi-modal data such as laser radar point cloud to improve the positioning robustness. In complex environments, single sensor data is easily disturbed, leading to inaccurate obstacle positioning and affecting subsequent obstacle avoidance decisions.
[0024] In terms of path planning, existing algorithms also have certain limitations. Although global planning algorithms can generate feasible paths, they have problems of node redundancy and path zigzagging, and cannot adapt to changes in dynamic obstacles, lacking flexibility. Local planning algorithms have real-time obstacle avoidance capability, but have deficiencies in global path optimization, which can easily lead to path fluctuations, making the UAV flight trajectory unstable. Moreover, traditional path planning methods rely on a single sensor, and in complex scenarios, signal interference or environmental obstruction can lead to positioning drift or obstacle misjudgment, seriously affecting the safe flight of the UAV.
[0025] In addition, the existing technology has a problem of computing power in deploying an open vocabulary real-time detection of edge devices. The existing deep learning model has a large number of parameters and high computational complexity, which is difficult to deploy on a UAV embedded platform. Some existing technologies use adaptive weighted fusion and lightweight Kalman filtering to reduce computational load, but they cannot realize dynamic obstacle category expansion. At the same time, traditional multi-modal fusion methods (such as early / late fusion) do not perform model pruning and quantization compression for edge devices, resulting in insufficient real-time performance and difficulty in meeting the millisecond-level obstacle avoidance response requirement in power inspection.
[0026] In summary, the existing technology has significant deficiencies in multi-modal fusion robustness, open vocabulary detection real-time performance, dynamic optimization of path planning, and adaptability of edge computing, which restricts the safety and efficiency of UAVs in complex extra-high voltage transmission line power inspection scenarios.
[0027] In view of this, the embodiments of the present application provide a UAV power inspection obstacle avoidance method and an obstacle avoidance system. In terms of multi-modal data processing, the method obtains and weightedly fuses multi-modal sensor data, fully utilizes the complementarity of multi-source information, and effectively improves the robustness of sensor fusion compared to traditional methods. It can better cope with the performance degradation of laser radar in bad weather and visual detection in complex lighting conditions, and reduce the fluctuation of fusion weights when the sensor fails or is disturbed by noise. In terms of path planning, a global path is first generated based on multi-modal fusion data, solving the problems of node redundancy and path zigzagging of traditional global planning algorithms. Then, a local obstacle avoidance path is generated using the dynamic window algorithm, and the global path and the local obstacle avoidance path are fused and curve-fitted, taking into account global path optimization and real-time obstacle avoidance capability, avoiding path fluctuations, making the UAV flight trajectory more stable, and being able to flexibly adapt to changes in dynamic obstacles. In the obstacle detection link, a pre-trained target detection model is called to detect real-time environmental images, which can break through the limitations of traditional models relying on fixed vocabulary and effectively identify unannotated temporary obstacles (such as birds) in power inspection. Combined with multi-modal fusion data and path planning, the positioning robustness is further improved, the path obstacle avoidance information is accurately obtained, providing accurate basis for adjusting the flight path, reducing the risk of collision, and improving the safety and efficiency of the inspection task.
[0028] The embodiments of the present application are further described below with reference to the drawings.
[0029] Referring to Figure 1 , Figure 1 is a flowchart of a UAV power inspection obstacle avoidance method provided by an embodiment of the present application. The flowchart can specifically include but is not limited to steps 110 to 160.
[0030] Step 110: Obtain multi-modal sensor data of the UAV, and perform weighted fusion on the multi-modal sensor data to obtain multi-modal fusion data; Step 120: Generate a global path based on the multi-modal fusion data; Step 130: Generate a local obstacle avoidance path based on the global path by using a dynamic window algorithm; Step 140: Perform path coordination on the global path and the local obstacle avoidance path to obtain an initial flight path; Step 150: Obtain real-time environment images collected when the UAV flies along the initial flight path, call a pre-trained target detection model to detect the real-time environment images, and obtain path obstacle avoidance information; Step 160: Adjust the initial flight path according to the path obstacle avoidance information to obtain a target flight path.
[0031] The steps 110 to 160 are described in detail below.
[0032] In a feasible embodiment, the UAV is equipped with multi-modal sensors such as a laser radar sensor, an inertial measurement unit (IMU), a GPS sensor, and a binocular vision camera. Among them, the laser radar sensor scans the environment by emitting a laser beam, and obtains high-precision three-dimensional point cloud data in real time, which provides a basis for obstacle modeling and space perception; the IMU continuously monitors the motion attitude data of the UAV, including acceleration, angular velocity and other information, to ensure flight stability and attitude control; the GPS sensor is responsible for obtaining global positioning data to determine the geographical position of the UAV and assist in path planning; the binocular vision camera simulates the principle of human eye vision, and in complex power inspection scenes, it assists in the identification and distance estimation of dynamic targets (such as flying birds and mobile devices), and improves the perception ability of unstructured obstacles. The multi-sensor works cooperatively to provide multi-dimensional data support for UAV power inspection.
[0033] It can be understood that in the complex and changeable power inspection flight environment, the performance of a single sensor is easily restricted by external factors. For example, in the area with high-rise buildings, the multi-path effect caused by multiple reflections of signals leads to the drift of positioning data; in the rain and fog, the laser beam is scattered and absorbed by particles in the air, resulting in a large amount of point cloud noise, which significantly reduces the environmental perception accuracy; the strong electromagnetic field around the ultra-high voltage transmission line will interfere with the normal operation of the unmanned aerial vehicle magnetic compass, affecting the heading determination. The traditional sensing scheme based on a single sensor is difficult to guarantee the robustness of the unmanned aerial vehicle sensing system in the scene with dynamic obstacles and complex electromagnetic environment. Therefore, by weighting and fusing the multi-modal sensor data such as laser radar, GPS, and inertial measurement, the advantages of each sensor can be complemented, which can effectively improve the sensing reliability of the unmanned aerial vehicle in complex environment, and meet the demand of high-precision and high-stability sensing for power inspection.
[0034] In a feasible embodiment, the multi-modal sensor data includes global positioning data, three-dimensional point cloud data, and motion attitude data. As shown in Figure 2 The execution process of step 110 of weighting and fusing the multi-modal sensor data to obtain the multi-modal fusion data can include but is not limited to steps 210 to 220.
[0035] Step 210: respectively calculate the confidence of the global positioning data, the three-dimensional point cloud data, and the motion attitude data to obtain the corresponding first confidence, second confidence, and third confidence; Step 220: weight and fuse the first confidence, the second confidence, and the third confidence to obtain the multi-modal fusion data.
[0036] In a feasible embodiment, step 210 aims to evaluate the reliability of the global positioning data, the three-dimensional point cloud data, and the motion attitude data respectively. Specifically, the confidence of the global positioning data (such as GPS) (i.e. the first confidence) is calculated by formula (1), while the confidence of the three-dimensional point cloud data (such as laser radar) and the motion attitude data (such as IMU) (i.e. the second and third confidence) is evaluated according to formula (2). Wherein, the confidence calculation of the laser radar introduces a dynamic weight mechanism: the stability can be evaluated based on the position increment of its continuous 5 frames, once the historical increment fluctuation exceeds the preset threshold, the weight of the data in the fusion calculation will be automatically reduced; similarly, the confidence calculation of the IMU also applies formula (2), and can be combined with the error characteristics and dynamic response performance of the sensor to dynamically adjust the weight.
[0037]
[0038] In formula (1), is the current number of satellites, is the position precision dilution, is the weight coefficient. The more satellites there are and the higher the positioning accuracy, the greater the GPS confidence.
[0039]
[0040] In formula (2), Indicates the position increment of the kth frame (i.e., the displacement difference between the current frame and the previous frame), Represents the sum of the position increments of the past 5 frames (k-5 to k-1). The meaning of W includes: Consistent with the incremental trend of the past 5 frames (i.e. close to the average increment of the past 5 frames), then W≈1 (high confidence); if If the increment deviates significantly from the increment of the past 5 frames (such as suddenly increasing or decreasing), then W < 1 (low confidence).
[0041] In one feasible embodiment, after obtaining the first confidence level corresponding to the global positioning data, the second confidence level for the 3D point cloud data, and the third confidence level for the motion posture data, these confidence levels can be normalized. This normalization process maps each confidence level to a uniform range (typically [0, 1]), eliminating differences in data dimensions and ensuring comparability of confidence levels across modal data. This provides standardized input for subsequent weight assignment and fusion calculations.
[0042] In one feasible embodiment, in step 220, weights are dynamically assigned based on the confidence level of each modal data (higher confidence equals greater weight). For example, a higher weight is assigned when the GPS signal is stable, while a higher weight is assigned to the LiDAR signal in rainy or foggy weather. By integrating the three types of data through weighted summation (or algorithms such as Kalman filtering), multimodal fusion data containing position, environment, and attitude information can be generated.
[0043] In a feasible embodiment, the weighting formula is shown as formula (3):
[0044] in, is the confidence level of GPS, is the confidence of the lidar, is the confidence of IMU, is the weight of GPS, is the weight of the lidar, It should be noted that the confidence of the IMU can be dynamically adjusted based on the covariance matrix of angular velocity and acceleration to ensure that the inertial measurement unit dominates positioning in short-term high-dynamic scenarios (such as sharp turns).
[0045] It can be understood that, due to the inevitable existence of noise in the multi-modal fusion data, such as integral drift of inertial measurement unit, measurement jitter of laser radar, etc., it is necessary to utilize a filtering algorithm to smooth the data, thereby improving the positioning accuracy. The embodiment selects a state estimation filter improved based on linear Kalman filter, and optimizes the state vector through two stages of prediction and update: in the prediction stage, the next time state is estimated based on the system dynamic model; in the update stage, the prediction result is corrected combined with the latest observation data, effectively suppressing noise interference, and realizing more accurate state estimation.
[0046] Specifically, in the prediction stage:
[0047] wherein, is a state transition matrix, is a process noise covariance matrix, is a UAV control input (such as thrust instruction).
[0048] In the update stage:
[0049] wherein, is an observation matrix, is an observation noise covariance matrix. By dynamically adjusting and , different environmental noise characteristics such as GPS signal fluctuation or laser radar point cloud noise are adapted. In the requirement of real-time, the time complexity of the filtering algorithm is O(n), which meets the millisecond-level state update requirement of the UAV.
[0050] In a feasible embodiment, as shown in Figure 3 , the execution process of step 120 of generating a global path based on multi-modal fusion data can include but is not limited to steps 310 to 330.
[0051] Step 310: converting three-dimensional point cloud data in the multi-modal fusion data into a three-dimensional map, and extracting a two-dimensional grid map at a preset flight height based on the three-dimensional map; Step 320: generating a sampling point in a passable area of the two-dimensional grid map by using a target bias search strategy, calculating a target guiding direction of the sampling point by using an attractive force function, and obtaining a sampling point sequence with target guidance; Step 330: generating a global path on the two-dimensional grid map based on the sampling point sequence with target guidance, in combination with a collision detection algorithm.
[0052] In a feasible embodiment, in step 310, the 3D point cloud data in the multimodal fusion data is converted into an octree structure to construct a 3D map with spatial layering characteristics, and further processed into an offline map supporting multiple resolutions. The 3D grid map construction process is as follows: Figure 4 As shown in the figure, this process generates scene coordinate maps at six different scales based on the rasterization principle. For power inspection scenarios, when projecting a 3D map onto a 2D plane, a slice corresponding to a fixed flight altitude (30m) can be selected to generate a 2D raster map. This high-dimensional dimensionality reduction effectively reduces the amount of data required for global path planning, lowering the computational complexity of the planning algorithm and improving path planning efficiency.
[0053] In a feasible embodiment, in step 320, within the traversable area of the two-dimensional grid map, sampling points are selected with a bias toward the target direction (target bias search strategy), and then the direction of each sampling point toward the target is calculated using an attraction function. Ultimately, an ordered sequence of sampling points with a target orientation can be formed, providing a more directional node basis for path planning.
[0054] In a feasible embodiment, when executing the target bias search strategy, during the random sampling process, the P random tree can be guided to grow toward the target point with a specified probability. Specifically, when generating the random sampling point When , according to the following probability distribution:
[0055] in, represents a random sample uniformly distributed in the environment, Indicates the target point Gaussian distribution centered on is the bootstrap probability, which usually takes a value between 0 and 1.
[0056] The attraction function is used to guide the growth direction of the random tree. The design of the attraction function makes the random tree grow more towards the target point, reducing blind exploration in complex environments and improving search efficiency. The attraction function can be defined as:
[0057] in, is the current node's position, It is the direction of attraction.
[0058] In a feasible embodiment, in step 330, based on the existing target-oriented sampling point sequence (including ordered nodes pointing to the target), a collision detection algorithm can be used to verify point by point whether the path between adjacent nodes is blocked by obstacles, retain collision-free connecting edges, and finally generate a feasible global path from the starting point to the target point in the two-dimensional grid map.
[0059] In a feasible embodiment, the sampling point sequence includes a starting sampling point and an ending sampling point. Figure 5 As shown, the execution process of step 330 may include but is not limited to steps 510 to 540.
[0060] Step 510: Starting from the starting sampling point, sequentially connect the sampling points with target orientation to generate a preliminary global path; Step 520: Using a breadth-first search strategy, perform collision detection on the sampling points on the preliminary global path starting from the starting sampling point to obtain a detection result; Step 530: If the detection result indicates that there is no collision at the current sampling point, continue detecting the next sampling point; if the detection result indicates that there is a collision at the current sampling point, add the previous sampling point of the current sampling point to the path table; Step 540: until the last sampling point is traversed, the global path is obtained according to the path table.
[0061] In a feasible embodiment, in step 510, starting from the starting sampling point, each node is connected in a target-oriented sampling point sequence to form a preliminary global path. By establishing a basic framework for path planning, a traversal sequence can be provided for subsequent collision detection.
[0062] In a feasible embodiment, in step 520, a breadth-first search (BFS) is used to perform collision detection on each sampling point on the preliminary global path in a layer-by-layer expansion manner to determine whether the sampling point falls into the obstacle area of the two-dimensional grid map, and the collision status of each sampling point (collision or no collision) can be obtained.
[0063] In a feasible embodiment, in step 530, if there is no collision, that is, the current sampling point passes the detection, and the next sampling point in the detection sequence continues; if a collision occurs, the current collision point can be skipped, and the previous valid sampling point of the collision point is added to the path table, where the path table records all verified valid path nodes.
[0064] In a feasible embodiment, in step 540, when traversing to the last sampling point (target point), a collision-free global path with both goal-orientedness and obstacle avoidance capabilities can be generated based on the valid node sequence recorded in the path table, ensuring that the path extends along the preset target and bypasses all known obstacles.
[0065] Specifically, in step 520 and step 530, from the starting point Start by performing collision detection on the nodes on the path. If the current node If there is no collision, continue to detect the next node; if a collision is detected, move the previous node The path table is added. The optimized path node set { } satisfies:
[0066] wherein, is the maximum allowed distance between path nodes.
[0067] In a feasible embodiment, as shown in Figure 6 , the execution process of generating a local obstacle avoidance path based on the global path in step 130 can include but is not limited to steps 610 to 640.
[0068] Step 610: decompose the global path into a series of sub-target points; Step 620: introduce a target point evaluation function in the dynamic window algorithm, combine the constraints of the sub-target points to perform forward trajectory prediction on the velocity and angular velocity combinations in the sampling space, and generate a plurality of predicted motion trajectories; Step 630: construct a trajectory evaluation function, and calculate the evaluation scores of each predicted motion trajectory using the trajectory evaluation function; Step 640: select a local obstacle avoidance path from the predicted motion trajectories according to the evaluation scores.
[0069] In a feasible embodiment, in steps 610 to 640, first, the global path is discretized into an ordered series of sub-target points at a fixed distance (such as 5 meters) to provide clear guidance for local path planning; then, a target point evaluation function (such as the Euclidean distance from the sub-target point) is introduced in the velocity-angular velocity sampling space of the dynamic window algorithm (DWA), and a plurality of predicted motion trajectories are generated in combination with the real-time point cloud obstacle constraints of the laser radar; then, a multi-objective optimization evaluation function is constructed to comprehensively consider the distance between the trajectory endpoint and the sub-target point (target guidance), the minimum distance between the trajectory and the obstacle (safety), the trajectory curvature change rate (smoothness), and whether the velocity / acceleration meets the physical limit (dynamic feasibility) and calculate the comprehensive score. Subsequently, the trajectory with the highest score is selected as the current local obstacle avoidance path output to the unmanned aerial vehicle control system, realizing the balance between real-time obstacle avoidance and global target.
[0070] Specifically, a target point evaluation function is introduced in the traditional dynamic window algorithm evaluation function to optimize the velocity selection mechanism of the dynamic window algorithm. The target point evaluation function is shown in equation (9):
[0071] wherein, is the endpoint position of the trajectory under the velocity and angular velocity .
[0072] Trajectory prediction and evaluation: For each pair of velocity and angular velocity in the sampling space Forward trajectory prediction is performed to generate a predicted motion trajectory. The trajectory evaluation function G(v, ω) combines multiple sub-evaluation functions:
[0073] wherein, are weight coefficients corresponding to the weights of orientation, distance, velocity, and target point evaluation, respectively.
[0074] In a feasible embodiment, after obtaining the global path, further optimization can be performed by the following steps: first, screening key nodes (such as path inflection points, obstacle avoidance detour points) from the global path to construct a control point sequence; then, using a curve algorithm (such as a B-spline curve algorithm) to perform smooth interpolation on the control point sequence to generate an initial smooth curve. This algorithm can ensure high-order continuity (such as C2 continuity) of the curve through linear combination of basis functions; then, dynamically adjusting the B-spline parameters (such as order, node vector) to minimize the path length and the rate of curvature change under the premise of satisfying the kinematic constraints of the UAV (such as the maximum turning radius), and finally outputting the optimized global path.
[0075] Specifically, when using the B-spline curve algorithm to perform smooth interpolation on the control point sequence to generate an initial smooth curve, the mathematical expression of the fifth-order B-spline curve used is shown in equation (11):
[0076] wherein, is the fifth-order B-spline basis function, which is calculated by the De Boor recurrence formula. This method ensures that the trajectory is continuously derivable to the second order in the global range, i.e., the curvature is continuous, effectively improving the stability of the UAV flight. In terms of obstacle avoidance optimization, by introducing a control point weight adjustment mechanism, a potential function based on the obstacle distance field is constructed, and the influence of the control point on the trajectory shape is dynamically adjusted to make the trajectory away from the obstacle area, optimize the path length, and optimize the obstacle avoidance distance. In addition, dynamic constraint conditions are embedded in the B-spline optimization process, which can introduce the maximum speed of the UAV, the maximum acceleration , and increase the kinematic constraint penalty term in the B-spline optimization, as shown in equation (12):
[0077] wherein, is the obstacle coordinate, is the obstacle avoidance weight coefficient.
[0078] Further, to cope with dynamic environmental changes, the embodiment also designs a real-time updating mechanism: when DWA (Dynamic Window Approach) triggers local obstacle avoidance, the key nodes of the current obstacle avoidance path are immediately extracted, and the B-spline curve is refitted after fusion with the original control point sequence to ensure the global path continuity. Through this mechanism, the update delay is greatly reduced.
[0079] It can be understood that during flight, the global path can be dynamically adjusted according to real-time environmental information and obstacle avoidance requirements to ensure that the UAV safely and efficiently reaches the target point.
[0080] In a feasible embodiment, when the initial flight path is generated by path coordination fusion of the global path and the local obstacle avoidance path, the adjusted global path node set } needs to meet the constraint condition as shown in equation (13) to ensure the spatiotemporal continuity and obstacle avoidance feasibility of the path.
[0081]
[0082] wherein, is the maximum allowed offset of path adjustment.
[0083] In a feasible embodiment, during flight, the UAV collects environmental data in real time through sensors (such as laser radar, visual camera) and feeds back to the path planning system. The system dynamically updates the global path and local obstacle avoidance strategy based on real-time environmental information, and through the introduction of an adaptive weight adjustment mechanism or an incremental optimization algorithm, the updated path planning result meets the dynamic constraint condition of equation (14), which can effectively improve the robustness and environmental adaptability of path planning in complex environments.
[0084]
[0085] wherein, is the maximum allowed offset of path update, is the minimum threshold of trajectory evaluation.
[0086] The above process realizes dynamic optimization of the path through a closed-loop feedback mechanism, ensuring that the UAV can quickly generate a feasible path that takes into account global goals and local obstacle avoidance when facing sudden obstacles or environmental changes.
[0087] Referring to Figure 7 , Figure 7is an improved YOLO detection framework diagram provided by an embodiment of the present application. The detection framework process is divided into three stages of input, feature fusion and detection output. In the input stage, after inputting the image, on the one hand, multi-scale image features are extracted by using the YOLO backbone network, covering the details and semantic information of different levels of the image; on the other hand, the nouns in the text "a man, a woman and a dog are skiing" are extracted, and the text encoder is used to convert the nouns into word embeddings, while supporting the user to provide offline vocabulary integration. In the feature fusion stage, the image perception embedding and the word embedding are combined, and the multi-scale image features and the text related information are further fused by means of the visual language PAN (a specific feature fusion module), so as to realize the interactive fusion of visual and text features. In the detection output stage, the fused features are input into the text contrast head and the frame head, the former is responsible for region-text matching, and the degree of fit between the image region and the text description is calculated, and the latter performs target detection frame prediction, so as to finally determine the positions of the man, the woman, the dog and other objects in the image, and complete the detection.
[0088] Reference Figure 8 , Figure 8 is a general framework diagram of the target detection model provided by an embodiment of the present application. The target detection model is constructed based on the improved YOLO detection framework as shown in Figure 7 , and integrates a visual-semantic interaction network to realize the interaction of visual and language features, and includes a cross-stage local network layer (T-CSP layer), an image pooling attention module, a text encoding module, a semantic comparison module and a feature fusion module. The T-CSP layer (Transposed Cross Stage Partial Network) is an improved variant of the CSP layer (Cross Stage Partial Network), both of which are network structures used for optimizing feature extraction and fusion in the field of computer vision. The T-CSP layer can be used to calculate the multi-modal features of real-time environmental images and texts. The image pooling attention module is used for multi-scale pooling processing of the image to obtain multi-scale feature representation. The feature fusion module is used to fuse the multi-modal features and the multi-scale features to obtain multi-modal semantic feature representation. The text encoding module is used to generate a semantic feature library containing multiple targets. The semantic comparison module is used to compare the multi-modal semantic features with the semantic feature library to obtain path obstacle avoidance information.
[0089] It should be noted that the target detection model fuses the CSP layer guided by text and the image pooling attention module to realize the deep interaction of text and image features. Specifically, the model uses a CLIP text encoder to build a power equipment feature library covering 12 types of core equipment (such as insulators, lightning arresters) and 9 types of dynamic obstacles (broken, suspended foreign matter, etc.) standard terms and open vocabulary descriptions. Through the CSP layer, attention calculation is performed on the input image features and text embedding, and the fused features are output; at the same time, the image pooling attention module divides the multi-scale image features into 27 image block embeddings, and updates the text embedding by means of the multi-head attention mechanism. Finally, the fused image region features and the feature library are matched by cosine similarity to realize fine-grained semantic-level recognition, such as accurately locating "insulator breakage" rather than generalizing to "obstacle", effectively improving the target detection accuracy and semantic understanding ability in complex power inspection scenarios.
[0090] In the model architecture, the CSP layer performs deep fusion of the input image features and the text embedding through the attention mechanism to generate a multi-modal feature representation, and the calculation process is shown in equation (15); the image pooling attention module performs spatial pooling operation on the multi-scale image features to generate 27 image block embeddings, and updates the text embedding dynamically through the multi-head attention mechanism, and the specific calculation process is shown in equation (16).
[0091]
[0092] wherein, is a Sigmoid function, is an offline encoded text feature.
[0093]
[0094] In a feasible embodiment, the target detection model adopts a two-stage semantic fusion strategy: in the pre-training stage, the CLIP text encoder is used to map semantic labels such as "live wire" and "insulator breakage" into high-dimensional vector representations to build an offline feature library (Embedding Bank), and the text semantic information is injected into the visual language pyramid attention network (PAN) through the vocabulary embedding technology for region-level text matching; in the inference stage, an improved visual-semantic interaction network is used to realize online cross-modal alignment of image features and text features, and the feature similarity is optimized through a region-text contrast loss function, and the specific calculation process is shown in equation (17). This design enables the model to maintain real-time detection performance while having fine-grained semantic understanding ability, and can accurately identify various equipment and their abnormal states in power scenarios.
[0095]
[0096] wherein, Temperature coefficient for adjusting distribution sharpness. Thus, the generalization ability is improved, and the misclassification rate of unlabeled classes under zero-shot conditions is greatly reduced.
[0097] In a feasible embodiment, the target detection model adopts a dynamic obstacle probability distribution modeling and threat evaluation mechanism to realize real-time response to complex environments. The specific process is as follows: the confidence of the detection frame is modeled by equation (18), and the top K high-confidence obstacles (default K=10) are selected for sorting; a threat level model is constructed according to the obstacle type, speed vector, and relative distance from the UAV, and weights are assigned to different obstacles (such as birds 0.8, temporary obstacles 0.6, and suspended foreign objects 0.3); a safety threshold is set, and when the distance between the obstacle and the UAV is <5m and the threat weight is >0.5, the system automatically triggers the obstacle avoidance response. This mechanism effectively balances the safety of obstacle avoidance and the efficiency of task execution through probability modeling and multi-dimensional threat evaluation, effectively improving the autonomous flight capability of the UAV in complex power environments.
[0098]
[0099] In a feasible embodiment, the target detection model deeply integrates the improved DWA path adjustment mechanism and the IDWA (Improved Dynamic Window Approach) algorithm, embeds the anomaly score result into the speed evaluation function (as shown in equation (19)), and forms a semantic-aware local path planning framework. This optimization strategy significantly reduces the end-to-end delay from obstacle detection to obstacle avoidance instruction generation, ensuring the real-time response capability of the system in complex scenes such as forests and mountains. Through the synergistic effect of multi-sensor dynamic fusion (laser radar + vision, etc.), high-order B-spline path optimization, and real-time target detection technology, a complete logic of "environment perception-semantic understanding-path decision" is constructed, realizing safe and efficient autonomous flight in complex power inspection environments.
[0100]
[0101] In a feasible embodiment, the execution process of acquiring real-time environment images collected by the UAV when flying along the initial flight path, calling the pre-trained target detection model to detect the real-time environment images, and obtaining the path obstacle avoidance information includes: when the UAV flies along the initial flight path, the environment images are collected in real time by a visual sensor (such as a binocular vision camera) and preprocessed, and input into the pre-trained target detection model; the model extracts multi-scale image features through the T-CSP layer, cross-modal fusion with semantic features generated by the CLIP text encoder, and then processed through the image pooling attention module and the feature fusion module, and compared with the offline semantic feature library to output path obstacle avoidance information containing target categories (such as damaged insulators, flying birds), detection frame coordinates, and abnormal scores. The system accordingly models the probability distribution of the first K high-confidence obstacles, calculates the threat value in combination with the type, speed, and relative distance, and when the distance is <5 meters and the threat weight is >0.5, embeds the abnormal score into the speed evaluation function of the IDWA algorithm to generate a local obstacle avoidance path, and updates the global path through feedback to realize real-time obstacle avoidance and path optimization.
[0102] In a feasible embodiment, in the process of adjusting the initial flight path to generate the target flight path according to the path obstacle avoidance information, first, the path nodes in the initial flight path are taken as control points, and a curve algorithm is used to construct a continuous and smooth initial trajectory; then, combined with the kinematic constraints of the UAV, dynamic obstacle avoidance constraints are constructed based on the positions and obstacle avoidance weights of the obstacles in the path obstacle avoidance information, and the dynamic obstacle avoidance constraints are integrated into the initial trajectory to obtain a motion trajectory with constraints; next, the control points of the curve algorithm are iteratively optimized to obtain an optimized smooth trajectory. When local obstacle avoidance is triggered, the control points of the curve algorithm are updated, and the smooth trajectory is refitted to obtain the target flight path. It can be understood that the path optimization process takes the initial flight path nodes as control points, uses a curve algorithm to construct a continuous and smooth initial trajectory, then combines the kinematic constraints of the UAV and the positions and obstacle avoidance weights of the obstacles in the path obstacle avoidance information to construct dynamic obstacle avoidance constraints and integrate them into the initial trajectory, iteratively optimizes the control points of the curve to obtain an optimized smooth trajectory that meets the constraints, and when local obstacle avoidance is triggered, the control points are updated in real time and the trajectory is refitted to generate a target flight path that avoids obstacles and remains coherent.
[0103] Referring to Figure 9 , Figure 9is a structure diagram of an unmanned aerial vehicle power inspection obstacle avoidance system provided by an embodiment of the present application. The obstacle avoidance system can be applied to the unmanned aerial vehicle power inspection obstacle avoidance method described above, and includes a multi-modal perception module 910, a data processing module 920, and a path planning module 930. The multi-modal perception module 910 is configured to obtain multi-modal sensor data of the unmanned aerial vehicle and to obtain multi-modal fusion data by weighted fusion. The data processing module 920 is configured to generate a global path based on the multi-modal fusion data, and to generate a local obstacle avoidance path based on the global path by using a dynamic window algorithm. The path planning module 930 is configured to obtain an initial flight path by path coordination of the global path and the local obstacle avoidance path, to obtain real-time environment images collected when the unmanned aerial vehicle flies along the initial flight path, and to detect the real-time environment images by using a pre-trained target detection model to obtain path obstacle avoidance information. The initial flight path is adjusted according to the path obstacle avoidance information to obtain a target flight path, so that the unmanned aerial vehicle can fly along the target flight path to realize obstacle avoidance. Through the cooperation of multi-modal data fusion, semantic perception detection, and dynamic path optimization, the unmanned aerial vehicle can realize real-time response and autonomous obstacle avoidance in a complex environment in a power inspection scene.
[0104] In a feasible embodiment, the multi-modal perception module 910 includes a laser radar sensor, an inertial measurement unit (IMU), a positioning sensor (such as a GPS), and a binocular vision camera. The laser radar sensor can construct an environment geometric model by three-dimensional point cloud scanning, and extract an accurate contour of an obstacle. The IMU can calculate the attitude angle and linear acceleration of the unmanned aerial vehicle in real time, and provide dynamic feedback for flight control. The positioning sensor can realize position anchoring in a global coordinate system. The binocular vision camera can assist in identifying and depth estimation of a dynamic target (such as a flying bird or a floating object) by parallax calculation. After Kalman filtering fusion of multi-sensor data, a spatiotemporal consistent environment representation is output.
[0105] It should be noted that the obstacle avoidance system is applicable to the unmanned aerial vehicle power inspection obstacle avoidance method described above, and the function implementation logic and principle of each module are completely consistent with the technical solution of the method. The specific functions of the modules can be referred to the detailed description of the processes of multi-modal data fusion, semantic detection, path coordination optimization, etc. in the foregoing, and will not be repeated here.
[0106] The above description of disclosed embodiments enables one of ordinary skill in the art to make or use the application. Various modifications to these embodiments will be apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for avoiding obstacles during UAV power inspection, characterized in that: include: Acquiring multimodal sensor data from the drone, and weightedly fusing the multimodal sensor data to obtain multimodal fusion data; generating a global path based on the multimodal fusion data; Based on the global path, a local obstacle avoidance path is generated using a dynamic window algorithm; Performing path coordination on the global path and the local obstacle avoidance path to obtain an initial flight path; Acquire a real-time environment image captured when the UAV flies along the initial flight path, call a pre-trained target detection model to detect the real-time environment image, and obtain path obstacle avoidance information; The initial flight path is adjusted according to the path obstacle avoidance information to obtain a target flight path.
2. The method for avoiding obstacles during UAV power inspection according to claim 1, characterized in that: The multimodal sensor data includes global positioning data, three-dimensional point cloud data and motion posture data; The weighted fusion of the multimodal sensor data to obtain multimodal fusion data includes: Calculating confidence levels of the global positioning data, the three-dimensional point cloud data, and the motion posture data respectively to obtain corresponding first confidence levels, second confidence levels, and third confidence levels; The first confidence level, the second confidence level, and the third confidence level are weighted and fused to obtain multimodal fusion data.
3. The method for avoiding obstacles during UAV power inspection according to claim 1, characterized in that: Generating a global path based on the multimodal fusion data includes: Converting the three-dimensional point cloud data in the multimodal fusion data into a three-dimensional map, and extracting a two-dimensional grid map at a preset flight altitude based on the three-dimensional map; In the traversable area of the two-dimensional grid map, a target bias search strategy is used to generate sampling points, and a target-guided direction of the sampling points is calculated using an attraction function to obtain a target-guided sampling point sequence; Based on the target-oriented sampling point sequence and in combination with a collision detection algorithm, a global path is generated on the two-dimensional grid map.
4. The method for avoiding obstacles during UAV power inspection according to claim 3, characterized in that: The sampling point sequence includes a starting sampling point and an ending sampling point; The generating of a global path on the two-dimensional grid map based on the target-oriented sampling point sequence and in combination with a collision detection algorithm includes: Starting from the starting sampling point, sequentially connecting the sampling points with target guidance to generate a preliminary global path; Using a breadth-first search strategy, starting from the starting sampling point, sequentially perform collision detection on the sampling points on the preliminary global path to obtain a detection result; If the detection result indicates that the current sampling point has no collision, continue to detect the next sampling point; if the detection result indicates that the current sampling point has a collision, add the previous sampling point of the current sampling point to the path table; Until the last sampling point is traversed, a global path is obtained according to the path table.
5. The method for avoiding obstacles during UAV power inspection according to claim 1, characterized in that: The generating of a local obstacle avoidance path based on the global path by using a dynamic window algorithm includes: Decomposing the global path into a series of sub-goal points; A target point evaluation function is introduced into the dynamic window algorithm, and the velocity and angular velocity combination in the sampling space is predicted forward based on the constraints of the sub-target points to generate multiple predicted motion trajectories. Constructing a trajectory evaluation function, and calculating an evaluation score of each of the predicted motion trajectories using the trajectory evaluation function; A local obstacle avoidance path is selected from the predicted motion trajectory according to the evaluation score.
6. The method for avoiding obstacles in UAV power inspection according to claim 1, characterized in that: After obtaining the global path, the method further includes: extracting a sequence of control points from the global path; Using a curve algorithm to perform smooth interpolation processing on the control point sequence to obtain a smooth curve; The parameters of the curve algorithm are adjusted to obtain an optimized global path based on the smooth curve.
7. The method for avoiding obstacles during UAV power inspection according to claim 1, characterized in that: The pre-trained object detection model includes an interconnected cross-stage local network layer, an image pooling attention module, a text encoding module, a semantic comparison module and a feature fusion module, wherein the cross-stage local network layer is used to calculate the multimodal features of the real-time environment image features and text features; The image pooling attention module is used to perform multi-scale pooling processing on the real-time environment image to obtain a multi-scale feature representation; The feature fusion module is used to fuse the multimodal features and the multi-scale feature representation to obtain a multimodal semantic feature representation; the text encoding module is used to generate a semantic feature library containing multiple types of targets; and the semantic comparison module is used to compare the multimodal semantic feature representation with the semantic feature library to obtain the path obstacle avoidance information.
8. The method for avoiding obstacles during UAV power inspection according to claim 1, characterized in that: The adjusting the initial flight path according to the path obstacle avoidance information to obtain a target flight path includes: Using the path nodes in the initial flight path as control points, a curve algorithm is used to construct a continuous and smooth initial trajectory; In combination with the kinematic constraints of the UAV, dynamic obstacle avoidance constraints are constructed based on the positions and obstacle avoidance weights of the obstacles in the path obstacle avoidance information, and the dynamic obstacle avoidance constraints are integrated into the initial trajectory to obtain a motion trajectory with constraints; Iteratively optimizing the control points of the curve algorithm to obtain an optimized smooth trajectory; When local obstacle avoidance is triggered, the control points of the curve algorithm are updated, and the smooth trajectory is refitted to obtain the target flight path.
9. A UAV power inspection and obstacle avoidance system, characterized in that: The obstacle avoidance method for UAV power inspection according to any one of claims 1 to 8 is applied, wherein the obstacle avoidance system includes a multimodal perception module, a data processing module, and a path planning module; The multimodal sensing module is used to obtain multimodal sensor data of the UAV and perform weighted fusion of the multimodal sensor data to obtain multimodal fusion data; The data processing module is used to generate a global path based on the multimodal fusion data; based on the global path, a local obstacle avoidance path is generated using a dynamic window algorithm; The path planning module is used to coordinate the global path and the local obstacle avoidance path to obtain an initial flight path; obtain real-time environmental images collected when the UAV flies along the initial flight path, call a pre-trained target detection model to detect the real-time environmental images, and obtain path obstacle avoidance information; and adjust the initial flight path according to the path obstacle avoidance information to obtain a target flight path.
10. The UAV power inspection and obstacle avoidance system according to claim 9, characterized in that: The multimodal perception module includes a lidar sensor, an inertial measurement unit, a positioning sensor and a binocular vision camera. The lidar sensor is used to construct a three-dimensional point cloud map and extract obstacle contour information; the inertial measurement unit is used to monitor the drone's attitude and acceleration in real time; the positioning sensor is used to provide the drone's position information in a global coordinate system; and the binocular vision camera is used to assist in dynamic target recognition and distance estimation.
Citation Information
Patent Citations
Multi-scale-attention guiding target detection method suitable for unmanned aerial vehicle
CN117372910A
Unmanned aerial vehicle local obstacle avoidance method and system based on dynamic window algorithm
CN118838397A
Automatic ship berthing and leaving method based on video monitoring and image recognition
CN119200588A
Self-adaptive obstacle avoidance method for cooperative mechanical arm in dynamic scene
CN119238495A
Automobile automatic obstacle avoidance system based on real-time 3D target detection
CN119323777A
Cited By
Unmanned aerial vehicle flight control method and device, unmanned aerial vehicle and storage medium
CN121764155A
Unmanned aerial vehicle control method based on multi-modal fusion and related equipment
CN121765632A
Unmanned aerial vehicle control method based on multi-modal fusion and related device
CN121765632B
Unmanned aerial vehicle inspection method and system based on semantic guidance and asset value cost map
CN122086055B
Low-altitude aircraft power line obstacle avoidance detection system based on deep learning
CN122223691A