An unmanned aerial vehicle autonomous power facility inspection method based on three-dimensional semantic driving

The method of UAV autonomous power facility inspection driven by 3D semantics solves the problem of relying on manual preset routes in UAV inspection, realizes fine-grained perception and adaptive inspection of power facilities, improves inspection efficiency and coverage uniformity, and reduces operation and maintenance costs.

CN121430653BActive Publication Date: 2026-04-14湖南马栏山视频先进技术研究院有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing drone-based power facility inspections rely excessively on manually preset routes and simple geometric constraints, making it difficult to accurately reflect operational semantic requirements. They lack fine-grained perception and adaptive adjustment capabilities for key components, resulting in inspection paths that are disconnected from actual points of interest and are insensitive to environmental changes.

Method used

A method for autonomous inspection of power facilities using UAVs based on 3D semantics is adopted. By establishing a world coordinate system and combining an airborne camera and an open vocabulary visual perception network, a 3D semantic point cloud is generated. The method autonomously plans local paths and performs component-level semantic inspection, achieving fine-grained perception and adaptive adjustment of key parts.

Benefits of technology

It achieves precise adherence to operational semantics, ensures fine-grained perception of critical parts, aligns inspection paths with points of interest, possesses environmental adaptability, improves inspection efficiency and coverage uniformity, and reduces operational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121430653B_ABST
    Figure CN121430653B_ABST
Patent Text Reader

Abstract

The application relates to an unmanned aerial vehicle autonomous power facility inspection method based on three-dimensional semantic driving, and relates to the technical field of unmanned aerial vehicle cruising control.The application constructs a three-dimensional semantic representation containing key components under a unified world coordinate system, generates a three-dimensional semantic point cloud with component category labels in real time in combination with airborne visual perception, analyzes inspection requirements in an operation and maintenance regulation into an ordered component-level task sequence, and on this basis, the unmanned aerial vehicle autonomously plans an observation pose and a local flight path according to the three-dimensional semantic distribution of target components, and realizes adaptive approach and coverage of key components.The application can accurately follow operation and maintenance semantics, realize fine-grained perception of key parts, make the inspection path fit the points of attention, have environmental adaptability, effectively improve the inspection efficiency, reduce the cost, and solve the problems that current power facility unmanned aerial vehicle inspection depends on a preset mode, is difficult to meet operation and maintenance semantics, lacks key part perception and adaptability, and leads to poor inspection effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) cruise control technology, and in particular to an autonomous UAV power facility inspection method based on three-dimensional semantic driving. Background Technology

[0002] In current large-scale drone inspections of power facilities, there is a widespread over-reliance on manually preset routes and simple geometric constraints. This traditional approach has many drawbacks, failing to accurately reflect the crucial semantic requirements of maintenance procedures, such as clearly identifying key components for inspection and specifying methods for intensive inspections in abnormal situations. Furthermore, during inspections, there is a lack of granular perception capabilities for critical components such as towers, insulator strings, fittings, and blades, as well as the ability to adaptively adjust inspection strategies based on actual conditions. This directly leads to a disconnect between the inspection path and the actual points of concern, uneven inspection coverage, and a lack of sensitivity to environmental changes.

[0003] Therefore, there is an urgent need for an autonomous UAV inspection method for power facilities based on 3D semantics to solve the above problems. Summary of the Invention

[0004] To address the problems of current large-scale power facility drone inspections relying excessively on manually preset routes and simple geometric constraints, which result in difficulties in accurately reflecting operational semantic requirements, lack of fine-grained perception and adaptive adjustment capabilities for key parts, leading to disjointed inspection paths, uneven coverage, and insensitivity to environmental changes, this invention proposes a drone-based autonomous power facility inspection method driven by three-dimensional semantics.

[0005] This invention provides an autonomous UAV inspection method for power facilities based on three-dimensional semantic driving, comprising the following steps:

[0006] S1. Establish a world coordinate system, map the power facilities and drones onto the world coordinate system, and obtain the initial global 3D semantic point cloud of the facilities. and the initial pose of the drone The initial global facility three-dimensional semantic point cloud Including the initial global facility 3D point cloud and the initial global semantic label vector The drone is equipped with an onboard camera.

[0007] S2, Collection of inspection task texts After preprocessing, relevant phrases about power components are extracted and matched with the power component category set. Alignment yields component category sequence space Then according to the category sequence space Generate component-level semantic inspection sequence ;

[0008] S3. Based on the color RGB image acquired by the airborne camera at time t, obtain the pixel-level segmentation result of the target component through an open vocabulary visual perception network, and based on the depth map acquired by the airborne camera at time t... Camera intrinsic parameter matrix and the current pose of the drone After back-projecting the pixel-level segmentation results of the target component into three dimensions, the results are compared with the three-dimensional semantic point cloud of the global facility at the previous time step. The fusion process yields the current-time global facility 3D semantic point cloud. ;

[0009] S4, the first Category of electrical components awaiting inspection As the current target component category, then from the global facility 3D semantic point cloud Extract the 3D point set of the current target component category to generate representative waypoints for the current target component category. According to the current location of the drone and representative waypoints By using obstacle sets The constructed local path planning function yields a sequence of local control instructions. ;

[0010] S5. Based on the current pose of the drone Local control command sequence The control steps are executed step by step to obtain the new position and attitude of the UAV at the next moment. And update the global facility 3D semantic point cloud. Then in the new moment Filter and generate the 3D point set of the current target component And based on the three-dimensional point set of the UAV and the current target component. The closest distance between Construct a semantic inspection sequence at the component level to determine the coverage state indicator. Have all target components been inspected and completed?

[0011] Specifically, step S1 includes the following steps:

[0012] S11. Establish a world coordinate system and map the geometric point cloud data of the power facilities onto the world coordinate system to obtain a global 3D point cloud of the facilities. and a global facility 3D point cloud Initialize the corresponding initial global semantic label vector The initial three-dimensional semantic point cloud of the facility was obtained; Indicates the number of 3D points of the initial facility; Represents the set of electrical component categories; t represents time.

[0013] S12, Based on the initial position of the UAV With the initial rotation matrix Constructing the initial pose of the UAV in the world coordinate system .

[0014] Specifically, step S2 includes the following steps:

[0015] S21, Collection of inspection task texts Preprocessing yields a standardized collection of inspection task texts. ;

[0016] S22. Extract the inspection task text set Phrases related to electrical components and their association with categories of electrical components Alignment yields component category sequence space ;

[0017] S23, Component category sequence space All component categories are merged and deduplicated, and the order of different components is determined according to the explicit order words and implicit logical ordering of the operation and maintenance procedures to obtain the component-level semantic inspection sequence. .

[0018] Specifically, step S3 includes the following steps:

[0019] S31. Based on a frame of RGB image captured by the airborne camera at time t Inputting an open-vocabulary visual perception network yields pixel-level segmentation results of target components within the current field of view. ,in This represents the number of target components detected in the current frame. The component category label for the i-th target component; This is the semantic mask for the i-th target component;

[0020] S32, Traverse the pixel-level segmentation results All mask pixels For satisfying Mask pixels, based on the depth map acquired by the airborne camera at time t. Camera intrinsic parameter matrix and the current pose of the drone By backprojecting it in three dimensions, we obtain three-dimensional points in the world coordinate system, thus obtaining a set of three-dimensional points belonging to the i-th target component. ;

[0021] S33, All three-dimensional points of each target component at time t The points are stacked column-wise to form the incremental 3D point cloud matrix at the current time. And construct its corresponding incremental semantic label vector. The two are used as incremental 3D semantic point clouds ;

[0022] S34. Incremental 3D semantic point cloud Compared with the previous time step, the global facility 3D semantic point cloud The current moment's global facility 3D semantic point cloud is obtained by stitching and updating the data. .

[0023] Specifically, step S4 includes the following steps:

[0024] S41, the first Category of electrical components awaiting inspection As the current target component category, then from the global facility 3D semantic point cloud Filter out all points whose component category labels are equal to the current target component category to generate the current target component's 3D point set. l = 1, 2, ..., L; L is the total number of target components;

[0025] S42. Select a non-empty 3D point set of the current target component. The geometric center of the point cloud serves as the representative waypoint of the current target component. ;

[0026] S43. Based on the current location of the drone and the representative waypoint of the current target component Based on the global facility 3D semantic point cloud obstacle collection The constructed local path planning function yields a sequence of local control instructions. .

[0027] Specifically, step S5 includes the following steps:

[0028] S51, based on the current position of the drone Local control command sequence The control step h is based on the flight control update function to execute flight control step by step to obtain the new pose of the UAV at the new moment. And update the global facility 3D semantic point cloud through step S3. ; The new time step; h represents the control step;

[0029] S52, in every new moment Based on the global facility 3D semantic point cloud Filter out all component category labels that are equal to the current target component category. The points generate the current component's 3D point set. And based on the drone's current location Calculate the three-dimensional point set of the UAV and the current target component. The closest distance between ;

[0030] S53, Based on the nearest distance Construct coverage status indicator For the current target component category Perform semantic coverage determination of components, if Then the current target component category is considered to be... The semantic coverage requirement has been met, and the current target component category will be... Mark as completed inspection, then select the current target component category. Switch to component-level semantic inspection sequence The next target component category Proceed to step S4 until the component-level semantic inspection sequence is completed. All elements in the process are marked as having been inspected; if This indicates the current target component category. The semantic coverage has not yet been satisfied; the goal is to maintain [the target]. Proceed to step S3 to continue execution.

[0031] Specifically, the open vocabulary visual perception network described in step S31 can employ an existing open vocabulary detection and segmentation framework, the core structure of which is an image encoder plus a text encoder plus a cross-modal matching localization head, with the input being the current frame image. and the set of component categories The generated text prompts will be displayed on the network. The encoding is used to embed text and perform similarity matching with image features to locate the target element in the image. The corresponding region outputs the number of target components. and its masks and categories As the current frame image Pixel-level segmentation results of target components within the wide and high field of view.

[0032] Specifically, the local path planning function in step S43 first constructs a three-dimensional occupation grid map or Euclidean distance field map in the world coordinate system, and then uses the current UAV position... For the starting point and the destination waypoint The endpoint is based on the set of obstacles. On the map, a collision-free discrete path is obtained using any feasible planning algorithm among A*, D Lite*, and RRT. Finally, the discrete path is smoothed and subjected to dynamic feasibility processing to obtain the trajectory, which is then discretized into a sequence of control commands according to the flight control interface. Output.

[0033] Specifically, in step S52, based on the current location of the drone... Calculate the three-dimensional point set of the UAV and the current target component. The closest distance between As shown in the following formula:

[0034] ,

[0035] in It is a Euclidean norm function; Indicates the current time With the current target component category The number of relevant points; Represents the current target component's three-dimensional point set The point in the middle.

[0036] Specifically, the coverage status indication quantity mentioned in step S53 The specific values ​​include: if there exists a time... Make and Then the coverage status indicator Set to 1; otherwise, the coverage status indicator is set to 1. Take 0; where, Distance threshold for determining whether semantic coverage meets the standard.

[0037] This invention provides an autonomous UAV inspection method for power facilities based on 3D semantics. It constructs a 3D semantic representation of key components within a unified world coordinate system, and generates a 3D semantic point cloud with component category labels in real time using onboard visual perception. The inspection requirements in the operation and maintenance procedures are parsed into an ordered sequence of component-level tasks. Based on this, the UAV autonomously plans its observation pose and local flight path according to the 3D semantic distribution of the target components, achieving adaptive approach and coverage of key components. This invention accurately follows operation and maintenance semantics, achieves fine-grained perception of key parts, ensures the inspection path aligns with the points of interest, and possesses environmental adaptability. Therefore, it enables autonomous UAV inspection of large-scale power facilities driven by 3D semantic information, effectively improving inspection efficiency and reducing costs. It solves the problems of current power facility UAV inspections relying on preset modes, failing to meet operation and maintenance semantics, lacking key part perception and adaptability, resulting in poor inspection results.

[0038] Furthermore, this invention proposes a three-dimensional semantic point cloud modeling method for power facilities: components such as towers, insulator strings, fittings, and blades are uniformly mapped to the world coordinate system through open vocabulary visual perception + RGB-D back projection, forming a three-dimensional semantic point cloud with geometry and component category labels, rather than the traditional model with only geometric shapes. This enables high-precision fine-grained perception of key parts such as towers, insulator strings, fittings, and blades, and can keenly capture subtle defects and changes in these parts, such as cracks in insulator strings, corrosion of fittings, and surface damage of blades, so as to discover potential safety hazards in time, nip potential faults in the bud, and effectively ensure the safe and stable operation of power facilities.

[0039] Furthermore, this invention transforms the textual description of the operation and maintenance procedures into an executable component-level inspection sequence: by using a semantic parsing function, natural language requirements such as what to inspect first and what to inspect next are mapped into an ordered sequence of component categories. The semantic task chain directly drives subsequent perception, planning, and execution, rather than manually configuring fixed routes. This allows for a deep understanding and precise execution of the semantic instructions in the operation and maintenance procedures, clearly identifying key components to inspect, and automatically adjusting the inspection method based on different abnormal situations, thereby improving the professionalism and effectiveness of the inspection work.

[0040] Furthermore, this invention also has an autonomous pose planning and coverage determination mechanism based on three-dimensional semantic point clouds: around the current target component, the corresponding point set is extracted from the semantic point cloud and a three-dimensional representative waypoint is calculated. Combined with obstacle constraints, a local control command sequence is generated. The minimum distance from the UAV position to the component point set plus the coverage threshold is used to automatically determine whether the component has been inspected. This realizes an adaptive inspection path and coverage control directly driven by the semantic target. It can autonomously plan an inspection path that is highly consistent with the actual points of interest, avoiding the problem of path and points of interest being disconnected in the traditional mode. This ensures that the inspection coverage is comprehensive and uniform, without missing any key areas, and improves the integrity and reliability of the inspection work.

[0041] Furthermore, because this invention reduces the workload of manual intervention and preset routes, and optimizes the inspection path through autonomous planning and intelligent decision-making, it significantly shortens the inspection time and improves the inspection efficiency. Moreover, precise inspections can detect and handle problems in advance, avoiding greater losses caused by the expansion of faults. In the long run, this effectively reduces the operation and maintenance costs of power facilities. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a schematic diagram of an autonomous power facility inspection method based on three-dimensional semantic driving provided in an embodiment of the present invention. Detailed Implementation

[0044] The invention will be explained in detail through the following embodiments. The purpose of this invention is to protect all technical improvements within its scope. In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "multiple" means two or more, unless otherwise explicitly specified.

[0045] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0046] Example 1

[0047] refer to Figure 1 This embodiment provides a method for autonomous inspection of power facilities by unmanned aerial vehicles (UAVs) based on three-dimensional semantic driving, including the following steps:

[0048] S1. Establish a world coordinate system, map the power facilities and drones onto the world coordinate system, and obtain the initial global 3D semantic point cloud of the facilities. and the initial pose of the drone The initial global facility three-dimensional semantic point cloud Including the initial global facility 3D point cloud and the initial global semantic label vector The drone is equipped with an onboard camera.

[0049] S11. Establish a world coordinate system and map the geometric point cloud data of the power facilities onto the world coordinate system to obtain a global 3D point cloud of the facilities. and a global facility 3D point cloud Initialize the corresponding initial global semantic label vector The initial three-dimensional semantic point cloud of the facility was obtained; Indicates the number of 3D points of the initial facility; Represents the set of electrical component categories; t represents time.

[0050] First, select a unified global coordinate system for the area where the power facilities are located. For example, take the center of the transmission tower base or the bottom of the wind turbine tower as the origin, and vertically upwards as... The axis, the horizontal direction is defined as shaft and The axes are then used to map the collected geometric point cloud data of various power facilities (regardless of the source of the data from different surveying equipment or different model formats) to the same world coordinate system through on-site calibration and coordinate transformation. Let t=0 be the initial time, and the initial global 3D point cloud of the facilities is obtained. abbreviated as ;in This represents the number of 3D points of the initial facility, abbreviated as ;

[0051] At time t, Each column This represents the coordinates of a three-dimensional point on a power facility (such as a transmission tower, wind turbine tower, or blade) in the world coordinate system. These coordinates are a three-dimensional vector, meaning they consist of three components, and can typically be represented as: ,in Let x, y, and z represent the coordinates of the x-axis, y-axis, and z-axis in the world coordinate system at time t, respectively; n = 1, ..., N. t ; Indicates the number of three-dimensional points of the facility;

[0052] That is, at the initial time t=0, each column This represents the coordinates of a three-dimensional point on a power facility (such as a transmission tower, wind turbine tower, blade, etc.) in the world coordinate system.

[0053] Specifically, geometric point cloud data of various power facilities can be obtained by scanning point clouds with lidar or sampling point clouds from BIM / CAD models;

[0054] To subsequently bind geometric points to electrical component categories, a global 3D point cloud of the facility is used. Initialize an initial global semantic label vector abbreviated as ;in Indicates the relationship with the first Geometric points The corresponding component category label is taken from the set of electrical component categories. .

[0055] The operation and maintenance side predefines a set of electrical component categories that need to be monitored. :

[0056]

[0057] in For a component category label (such as tower body insulator string hardware blade root, etc.). This represents the total number of categories. In this step, we can... Initialize all spaces to be unlabeled or to a placeholder of a general category;

[0058] S12, Based on the initial position of the UAV With the initial rotation matrix Constructing the initial pose of the UAV in the world coordinate system :

[0059] At time t, the current pose of the UAV in the world coordinate system. Represented by the following homogeneous transformation matrix:

[0060] ,

[0061] in, Let be the current rotation matrix, and let represent the rotation matrix of the UAV at time t. The current position of the UAV represents its three-dimensional position in the world coordinate system at time t. It is a zero row vector;

[0062] Current location of the drone With the current rotation matrix It is achieved through the drone's navigation / positioning module (such as SLAM) at any given time. Directly output pose components; unify the current pose into a homogeneous matrix. ,therefore It can also be regarded as being caused by The extracted translation term. Meanwhile, yes Time from the initial The initial pose of the constructed position, and subsequent... Since the positioning module updates over time, the two are instances of the same pose at different times, and there is a clear temporal evolution relationship.

[0063] Therefore, at time t=0, based on the initial position of the UAV With the initial rotation matrix Constructing the initial pose of the UAV in the world coordinate system :

[0064] ,

[0065] This matrix is ​​used to describe the rigid transformation from the camera coordinate system / body coordinate system to the world coordinate system.

[0066] The initial position and attitude of the UAV in the world coordinate system are denoted as the position vector. With rotation matrix ;in Provide the three-dimensional coordinates of the UAV's takeoff or inspection starting point. This describes the orientation transformation of the world coordinate system relative to the UAV body coordinate system; both are used to construct the initial homogeneous pose matrix of the UAV. This variable is initialized in step S1. And it is updated over time during the inspection process;

[0067] The drone is equipped with an onboard camera, and the internal parameters of the onboard camera are provided by... intrinsic parameter matrix This is referred to as the camera intrinsic parameter matrix, which contains parameters such as focal length and principal point coordinates.

[0068] The airborne camera is an RGB-D camera;

[0069] An RGB-D camera is a sensor that can simultaneously acquire color images (RGB) and depth information. By combining a traditional color camera and a depth sensor, it provides richer 3D scene information than a single RGB camera.

[0070] Camera intrinsic parameter matrix It does not perform numerical calculations itself, but rather clarifies its relationship with the world coordinate system and the UAV's pose: in subsequent steps, it will use Normalize the pixel coordinates, map them to the camera coordinate system, and then use the current pose matrix. Transform to the world coordinate system to obtain the 3D coordinates in the 3D semantic point cloud. This is achieved by uniformly fixing the coordinates in this step. This ensures that the projection calculations at different time steps throughout the inspection process follow the same camera model, facilitating the continuous accumulation of semantic point clouds.

[0071] This step unifies large-scale power facilities, drones, and cameras into the same three-dimensional world coordinate system and initializes the core variables that will be used in subsequent semantic point cloud construction and path planning.

[0072] S2, Collection of inspection task texts After preprocessing, relevant phrases about power components are extracted and matched with the power component category set. Alignment yields component category sequence space Then according to the category sequence space Generate component-level semantic inspection sequence ;

[0073] S21, Collection of inspection task texts Preprocessing yields a standardized collection of inspection task texts. ;

[0074] The maintenance unit provides a set of inspection task descriptions based on procedures or on-site requirements, which are collectively recorded as an inspection task text set. :

[0075] in, For the number of text entries, each A text entry described in natural language (such as checking insulator strings first, then crossarm connecting bolts, and finally lightning rods on the tower top) can be considered as defined in a string space. The sequence on These texts collectively describe the semantic information of what needs to be checked and the general order of priority in this inspection mission.

[0076] First, analyze the collection of inspection task texts. Basic preprocessing is performed, including standardizing the encoding format, removing irrelevant symbols, and segmenting sentences and words, to obtain text data with a clearer structure, more standardized content, and easier subsequent in-depth analysis and mining. After preprocessing, the standardized text data is updated to the inspection task text set. middle;

[0077] S22. Extract the inspection task text set Phrases related to electrical components and their association with categories of electrical components Alignment yields component category sequence space ;

[0078] Based on this, domain rules and large-scale model reasoning are used to identify phrases related to power components in the text and associate them with a set of power component categories. Alignment, this process can be formalized as a semantic extraction function. :

[0079]

[0080] This function is used to process each piece of natural language text. Phrases such as "insulator string crossarm tower top lightning rod blade leading edge" appearing in the text are identified, and then matched with matching rules or semantic similarity to map them to specific component category labels. ; Indicates length is The component category sequence space, This represents the total number of target components obtained after analysis.

[0081] S23, Component category sequence space All component categories are merged and deduplicated, and the order of different components is determined according to the explicit order words and implicit logical ordering of the operation and maintenance procedures to obtain the component-level semantic inspection sequence. ;

[0082] After extracting key component phrases and matching their categories, the component category sequence space is then processed. The candidate component categories obtained from all text entries are merged, deduplicated, and sorted. Specifically, on the one hand, the order of different components is determined based on explicit order words (such as first, second, and last) and implicit logic (such as checking high-risk parts first, then checking regular parts) given in the operation and maintenance procedures; on the other hand, cases where the same component category appears repeatedly in multiple texts are merged, ultimately resulting in an ordered component-level semantic inspection sequence. :

[0083]

[0084] in, Indicates the first The categories of electrical components to be inspected. This is the length of the component-level semantic inspection sequence. This sequence can be considered as... The output not only retains the semantic requirements of the operation and maintenance procedures regarding what to inspect, but also provides the execution path in what order to inspect, thus providing a basis for subsequent adjustments based on the current component category. It provides a direct basis for filtering semantic point clouds and planning observation poses.

[0085] This step converts the natural language task description of the operation and maintenance procedure into an ordered list of power component-level inspections. Subsequent steps will then proceed according to this list to perceive and plan for each component.

[0086] S3. Based on the color RGB image acquired by the airborne camera at time t, obtain the pixel-level segmentation result of the target component through an open vocabulary visual perception network, and based on the depth map acquired by the airborne camera at time t... Camera intrinsic parameter matrix and the current pose of the drone After back-projecting the pixel-level segmentation results of the target component into three dimensions, the results are compared with the three-dimensional semantic point cloud of the global facility at the previous time step. The fusion process yields the current-time global facility 3D semantic point cloud. ;

[0087] The drone captures a frame of color RGB image using its onboard camera at time t. and corresponding depth map ,in These are the image height and width, respectively. Represents pixel coordinates The RGB value at that location (channel dimension is 3). This indicates the depth distance from the pixel to the camera's optical center.

[0088] It is understandable that the global facility 3D semantic point cloud includes Both are constantly updated based on time t; if time t is taken as the current time, then the global facility 3D semantic point cloud at the current time is The semantic point cloud that was constructed or updated at the previous time step (time step t-1) is represented as the three-dimensional semantic point cloud of the global facility at the previous time step, as follows:

[0089] in, This indicates the number of three-dimensional points of the facility at the previous moment. Each column represents the coordinates of a three-dimensional point in the world coordinate system. The component category label corresponding to this 3D point is taken from the set of power component categories. .exist At the current moment, the global facility 3D semantic point cloud is the initial global facility 3D point cloud. and the initial global semantic label vector At this point, there is no global facility 3D semantic point cloud from the previous moment;

[0090] S31. Based on a frame of RGB image captured by the airborne camera at time t Inputting an open-vocabulary visual perception network yields pixel-level segmentation results of target components within the current field of view. This process can be formalized as follows:

[0091]

[0092] in For open vocabulary visual perception networks; This represents the number of target components detected in the current frame. The component category label for the i-th target component indicates which type of electrical component the target component belongs to (such as insulator string, crossarm, tower body, etc.). Let be the semantic mask for the i-th target component, which is a binary matrix. Represents mask pixels Belongs to the target component, Represents mask pixels Not belonging to the target component; H is the image height; W is the image width;

[0093] This open-vocabulary detection and segmentation method allows for the direct extraction of power component category sets from images without pre-defining a fixed target set. The relevant area.

[0094] The open vocabulary visual perception network can employ existing open vocabulary detection and segmentation frameworks (e.g., a combination of an open vocabulary detector and a general segmenter, or a standalone open vocabulary segmentation network). Its core structure can be summarized as an image encoder + a text encoder + a cross-modal matching localization head (which can be connected to a segmentation head), with the input being the current frame image. and the set of component categories The generated text prompts (such as category names / synonyms) will be used by the network. The encoding is used to embed text and perform similarity matching with image features to locate the target element in the image. The corresponding region outputs the number of target components. and its masks and categories As the current frame image Pixel-level segmentation results of target components within the wide and high field of view.

[0095] Compared to conventional closed-vocabulary networks, this type of open-vocabulary method does not require retraining the classifier head for each new component category; it only needs to update / expand the classifier. The text prompts can be used to detect the problem, making it more suitable for the needs of power inspection that require scalable categories, long tails, and cross-scenario generalization.

[0096] S32, Traverse the pixel-level segmentation results All mask pixels For satisfying Mask pixels, based on the depth map acquired by the airborne camera at time t. Camera intrinsic parameter matrix and the current pose of the drone By backprojecting it in three dimensions, we obtain three-dimensional points in the world coordinate system, thus obtaining a set of three-dimensional points belonging to the i-th target component. As shown in the following formula:

[0097]

[0098]

[0099] in, For the homogeneous coordinates of the mask pixels, The inverse of the camera intrinsic parameter matrix. For mask pixels The corresponding homogeneous world coordinate vector, This indicates that the first three components of the vector are used as the three-dimensional coordinates.

[0100] S33, All three-dimensional points of each target component at time t The points are stacked column-wise to form the incremental 3D point cloud matrix at the current time. And construct its corresponding incremental semantic label vector. The two are used as incremental 3D semantic point clouds ;

[0101] After back-projecting the mask pixels to world coordinates, stack all 3D points of each target component (the i-th target component) at time t column-wise to form the incremental 3D point cloud matrix at the current time. ,in The number of new 3D points added to this frame.

[0102] Incremental 3D point cloud matrix The corresponding incremental semantic label vector is For all points belonging to the i-th target component, their label elements are all taken as... ,because ,so ;

[0103] In another possible implementation, the depth value, normal vector, or confidence level can be combined to perform simple filtering or downsampling of duplicate or noisy points without changing the above representation.

[0104] S34. Incremental 3D semantic point cloud Compared with the previous time step, the global facility 3D semantic point cloud The current moment's global facility 3D semantic point cloud is obtained by stitching and updating the data. .

[0105] The incremental 3D point cloud matrix at the current moment Compared with the previous moment, the global facility 3D point cloud Perform stitching and updating to obtain the current global facility 3D point cloud. Incremental semantic label vector Compared with the global semantic label vector at the previous time step The global semantic label vector at the current time is obtained by concatenating and updating the vector. :

[0106]

[0107]

[0108] Among them, symbols This indicates a column-wise concatenation operation; This represents the total number of points in the facility point cloud at the current moment after the update, i.e., the number of 3D points of the facility at the current moment. This represents the number of three-dimensional points of the facility at the previous moment.

[0109] By accumulating data frame by frame, a continuously enriched 3D semantic map is formed in the world coordinate system. This map preserves the geometric shape and records the component category corresponding to each point, providing a basis for subsequent analysis based on the current target component. It provides a data foundation for selecting corresponding regions and generating inspection poses in point clouds.

[0110] This step utilizes an onboard RGB-D camera from a drone to perform open-vocabulary visual perception of power facilities, and back-projects pixels belonging to various component categories onto the world coordinate system to construct and update a 3D point cloud with semantic labels. This provides a foundation for subsequent semantic-driven pose planning and coverage determination. In the following steps, a component-level semantic inspection sequence will be used. With the goal of using a global facility 3D semantic point cloud Filtering and current component category The corresponding point set is used to generate a 3D semantically driven observation pose and local flight path; then, the semantic coverage is calculated using this point cloud information and the completion of component inspection is determined.

[0111] S4, the first Category of electrical components awaiting inspection As the current target component category, then from the global facility 3D semantic point cloud Extract the 3D point set of the current target component category to generate representative waypoints for the current target component category. According to the current location of the drone and representative waypoints By using obstacle sets The constructed local path planning function yields a sequence of local control instructions. ;

[0112] Step S2 generates a component-level semantic inspection sequence. ,in Indicates the first The categories of electrical components to be inspected (e.g., insulator strings, tower top lightning rods, etc.), l=1, 2, ..., L; This represents the total number of target components. This step uses a fixed... To achieve this, we select the corresponding semantic regions from the point cloud and generate inspection poses and local paths pointing to those semantic regions.

[0113] Current moment global facility 3D point cloud Each column This represents the coordinates of a three-dimensional point on a power facility (such as a transmission tower, wind turbine tower, or blade) in the world coordinate system. These coordinates are a three-dimensional vector, meaning they consist of three components, and can typically be represented as: ,in Let x, y, and z represent the coordinates of the x-axis, y-axis, and z-axis in the world coordinate system at time t, respectively; n = 1, ..., N. t ; Indicates the number of three-dimensional points of the facility;

[0114] Based on the existing 3D semantic point cloud, focusing on the current target component category It automatically selects its three-dimensional representative area / point and generates a local control command sequence for the UAV to fly from its current position to the vicinity of the area under obstacle avoidance constraints, providing trajectory input for the subsequent autonomous execution and coverage determination of S5.

[0115] S41, the first Category of electrical components awaiting inspection As the current target component category, then from the global facility 3D semantic point cloud Filter out all points whose component category labels are equal to the current target component category to generate the current target component's 3D point set. l = 1, 2, ..., L; L is the total number of target components;

[0116]

[0117] in, For the current target component category In the current global facility 3D semantic point cloud The number of points in; if This indicates that there is currently no effective observation of the component from the current perspective, and further observation can be achieved in subsequent steps by continuing to move and supplement the perception. Gradually enriched; if Then it can be assumed that the component already has a certain geometric and semantic description at the current moment, and the current target component's three-dimensional point set It is a three-dimensional point set in the world coordinate system, which represents the current component and is suitable for generating inspection poses based on it.

[0118] S42. Select a non-empty 3D point set of the current target component. The geometric center of the point cloud serves as the representative waypoint of the current target component. ;

[0119] when At that time, that is, the current three-dimensional point set of the component Not empty; to simplify the process at the planning level, [the value will be]... The geometric information of all 3D points in the cloud is compressed into a single representative waypoint, using the geometric center of the point cloud:

[0120]

[0121] Among them, the representative waypoint of the current component It represents the three-dimensional representative position (waypoint) of the current component at the current time t, taking into account the overall distribution characteristics of the point cloud.

[0122] In another possible implementation, a safety offset or view constraint can be added near the geometric center of the point, but it should be used uniformly at the symbol level. This indicates the location of the component to be reached, facilitating subsequent correlation with the drone's location. Input them together into the path planning function.

[0123] S43. Based on the current location of the drone and the representative waypoint of the current target component Based on the global facility 3D semantic point cloud obstacle collection The constructed local path planning function yields a sequence of local control instructions. As shown in the following formula:

[0124] ,

[0125] in, The sequence of local control commands obtained from the planning; Let h be the control vector at step h, and the control vector includes the desired velocity, attitude, or position increment, etc., with dimension h. It depends on the specific flight control interface; h represents the number of planned steps; h represents the control steps, h = 1, 2, ..., H; and the current position of the UAV. As the starting point of the current path planning;

[0126] Based on the global facility 3D semantic point cloud The point distribution and labels can be used to pre-classify areas that are not allowed to be traversed, such as the tower body, the ground, and crossarms, into several obstacle objects, which are then collectively denoted as an obstacle set. :

[0127]

[0128] in Let be the number of obstacles at the current time t, and each This represents an obstacle region in the world coordinate system (which can be represented by a point cloud envelope or simplified geometry); this set is used in the planning process as an obstacle avoidance constraint.

[0129] This is a local path planning function. The following requirements must be met simultaneously during the planning process: First, the trajectory starts from the starting point. Approaching continuously and eventually reaching the target waypoint Nearby; secondly, avoid The obstacles, such as the tower and the ground, are shown in the figure. The safety distance and flight altitude constraints are met. Third, under the premise of meeting the safety constraints, the path length, number of turns, or energy consumption are optimized as much as possible.

[0130] Local path planning function The implementation can be completed according to the process of mapping—searching—trajectory generation—control output. Specifically, it starts with the global semantic point cloud. Cluster the impassable regions and expand the safety distance to obtain the obstacle set. And construct a 3D occupation grid map, Euclidean distance field map (ESDF), or navigation map in the world coordinate system; then, based on the current drone position For the starting point and the destination waypoint With the destination as the endpoint, a collision-free discrete path is obtained on the map using any feasible planning algorithm such as A*, D Lite*, or RRT. Finally, the discrete path is smoothed and subjected to dynamic feasibility processing (such as B-spline / minimum jerk) to obtain the trajectory, which is then discretized into a sequence of control commands via the flight control interface. The output, implemented as described above, belongs to A typical implementation can be achieved by selecting an equivalent planner based on computing power and platform.

[0131] The planning results are in the form of a sequence of control commands. The output will be in the form of [output format], and in subsequent steps, the drone will follow [the instructions / methods]. Control commands are executed sequentially, and the global semantic point cloud is continuously updated using the S3 mechanism during execution. Further calculate the component coverage distance and its relationship with the preset threshold The relationship between these parameters determines whether the current component inspection is complete; therefore, this step defines... , , and The symbols will be used directly in S5 and will not be redefined.

[0132] S5. Based on the current pose of the drone Local control command sequence The control steps are executed step by step to obtain the new position and attitude of the UAV at the next moment. And update the global facility 3D semantic point cloud. Then in the new moment Filter and generate the 3D point set of the current target component And based on the three-dimensional point set of the UAV and the current target component. The closest distance between Construct a semantic inspection sequence at the component level to determine the coverage state indicator. Have all target components been inspected and completed?

[0133] S51, based on the current position of the drone Local control command sequence The control step h is based on the flight control update function to execute flight control step by step to obtain the new pose of the UAV at the new moment. And update the global facility 3D semantic point cloud through step S3. ; The new time step; h represents the control step;

[0134] First, the drone follows the sequence of local control commands. Flight control is executed incrementally. To characterize the relationship between control and pose, a flight control update function can be defined:

[0135] ,

[0136] in Indicates the current pose of the drone Apply control vector The pose obtained at the next moment. For each control step Get the corresponding new moment ( The new time of the UAV is obtained by recursion through the flight control update function. Corresponding new pose of the drone at the new time. And extract the new time position of the drone from it. .

[0137] In every new moment The UAV utilizes the open-vocabulary visual perception and 3D back-projection mechanism in step S3 to re-perceive and project the current field of view, thus creating a 3D semantic point cloud of the global facilities. Replace t with This will update the semantic point cloud. and ,in The updated points allow the map to continuously accumulate and be corrected during flight.

[0138] S52, in every new moment Based on the global facility 3D semantic point cloud Filter out all component category labels that are equal to the current target component category. The points generate the current component's 3D point set. And based on the drone's current location Calculate the three-dimensional point set of the UAV and the current target component. The closest distance between ;

[0139] Referring to step S41, at each new moment Above, based on the updated global facility 3D semantic point cloud Extract the current target component category from it. The three-dimensional point set is used as the current target component's three-dimensional point set. :

[0140]

[0141] in Indicates the current time With the current target component category The number of relevant points;

[0142] Subsequently, based on the current drone position For reference, calculate the nearest distance between the UAV and the current set of 3D points of the component. :

[0143] ,

[0144] in, It is a Euclidean norm function; , used to measure the length of a three-dimensional vector; Represents the current target component's three-dimensional point set The point in the middle. This distance can be used to characterize the spatial proximity of the drone to the nearest visible point of the target component.

[0145] S53, Based on the nearest distance Construct coverage status indicator For the current target component category Perform semantic coverage determination of components, if Then the current target component category is considered to be... The semantic coverage requirement has been met, and the current target component category will be... Mark as completed inspection, then select the current target component category. Switch to component-level semantic inspection sequence The next target component category Proceed to step S4 until the component-level semantic inspection sequence is completed. All elements in the process are marked as having been inspected; if This indicates the current target component category. The semantic coverage has not yet been satisfied, so the system will not switch at this time. Instead, keep the goal as Continue executing the closed loop of perception update (S3) – planning (S4) – execution (S5);

[0146] If the current control instruction sequence has been executed There are still If so, a new path planning is triggered, and the process jumps directly to step S4 to generate a new sequence of control instructions, so as to continue to try to approach the component under the new starting point and updated point cloud conditions;

[0147] If a time exists Make and Then the coverage status indicator Set to 1; otherwise, the coverage status indicator is set to 1. Take 0; as shown in the following formula:

[0148]

[0149] in, A distance threshold for determining whether semantic coverage meets the standard; Let be a scalar real number representing the maximum permissible distance between the UAV and the target component area (e.g., if the distance between the UAV and the target component surface is less than or equal to 10 ... If the component is not effectively covered, then it is considered to have been effectively covered.

[0150] Complete the current control instruction sequence There are still The emphasis is on a checkpoint at the end of a planning and execution cycle. If the targets are still not met, a new starting point is established (i.e., after the drone completes its mission). The current position after / pose Recall S4 to generate a new one. And in the updated point cloud conditions (i.e., those accumulated through S3 during execution) , The system continues to attempt to approach and cover the component; the termination condition is the component-level semantic inspection sequence. The corresponding coverage state vector; All components are 1, meaning that each type of target component has met the coverage requirement.

[0151] Finally, the output includes a component-level semantic inspection sequence. Each category of electrical components to be inspected The state coverage vector of the coverage state indicator ;

[0152] in, Indicates the first Components The semantic coverage requirement has been met. This indicates that the project is incomplete or requires replanning.

[0153] The final three-dimensional semantic point cloud accumulated during the entire inspection process:

[0154] ,

[0155] in, This marks the end of the inspection. This represents the number of 3D points in the final point cloud, which can be used for defect identification and visualization in post-processing.

[0156] The set of pose trajectories of the drone throughout the inspection process:

[0157] ,

[0158] in These are a series of discrete location points during the inspection process. Corresponding to all components All were judged At that moment.

[0159] pose trajectory set Its function is to record the flight path of the drone during the inspection process, for purposes such as mission review, coverage visualization, mileage / energy consumption statistics, and safety auditing; where discrete location points refer to points within discrete time steps. Upsampled UAV position vector This refers to the three-dimensional position of the drone in the world coordinate system.

[0160] This embodiment provides a method for autonomous UAV inspection of power facilities based on 3D semantics. It constructs a 3D semantic representation of key components within a unified world coordinate system, and generates a 3D semantic point cloud with component category labels in real time using onboard visual perception. The inspection requirements in the operation and maintenance procedures are parsed into an ordered sequence of component-level tasks. Based on this, the UAV autonomously plans its observation pose and local flight path according to the 3D semantic distribution of the target components, achieving adaptive approach and coverage of key components. This embodiment accurately follows operation and maintenance semantics, achieves fine-grained perception of key parts, ensures the inspection path aligns with the points of interest, and possesses environmental adaptability. Therefore, it enables autonomous UAV inspection of large-scale power facilities driven by 3D semantic information, effectively improving inspection efficiency and reducing costs. It solves the problems of current power facility UAV inspections relying on preset modes, failing to meet operation and maintenance semantics, lacking key part perception and adaptability, resulting in poor inspection results.

[0161] Furthermore, this embodiment proposes a three-dimensional semantic point cloud modeling method for power facilities: components such as towers, insulator strings, fittings, and blades are uniformly mapped to the world coordinate system through open vocabulary visual perception + RGB-D back projection, forming a three-dimensional semantic point cloud with geometry and component category labels, rather than the traditional model with only geometric shapes. This enables high-precision fine-grained perception of key parts such as towers, insulator strings, fittings, and blades, and can keenly capture subtle defects and changes in these parts, such as cracks in insulator strings, corrosion of fittings, and surface damage of blades, so as to discover potential safety hazards in time, nip potential faults in the bud, and effectively ensure the safe and stable operation of power facilities.

[0162] Furthermore, this embodiment transforms the textual description of the operation and maintenance procedures into an executable component-level inspection sequence: by using a semantic parsing function, natural language requirements such as what to inspect first and what to inspect next are mapped into an ordered sequence of component categories. The semantic task chain directly drives subsequent perception, planning, and execution, rather than manually configuring fixed routes. This allows for a deep understanding and precise execution of the semantic instructions in the operation and maintenance procedures, clearly identifying key components to be inspected, and automatically adjusting the inspection method based on different abnormal situations, thereby improving the professionalism and effectiveness of the inspection work.

[0163] Furthermore, this embodiment also uses an autonomous pose planning and coverage determination mechanism based on 3D semantic point cloud: around the current target component, the corresponding point set is extracted from the semantic point cloud and a 3D representative waypoint is calculated. Combined with obstacle constraints, a local control command sequence is generated. The component is automatically determined to be inspected based on the minimum distance from the UAV position to the component point set plus the coverage threshold. This realizes an adaptive inspection path and coverage control directly driven by the semantic target. It can autonomously plan an inspection path that is highly consistent with the actual points of interest, avoiding the problem of path and points of interest being disconnected in the traditional mode. This ensures that the inspection coverage is comprehensive and uniform, without missing any key areas, and improves the integrity and reliability of the inspection work.

[0164] Furthermore, by reducing the workload of manual intervention and pre-set routes, and by optimizing inspection paths through autonomous planning and intelligent decision-making, inspection time is significantly shortened and inspection efficiency is improved. Moreover, precise inspections can detect and address problems in advance, preventing greater losses caused by the expansion of faults. In the long run, this effectively reduces the operation and maintenance costs of power facilities.

[0165] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 A process, multiple processes, and / or boxes Figure 1 Devices that specify the functions in one or more boxes.

[0166] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction device, which is implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0168] The parts of this invention not described in detail are prior art. It will be apparent to those skilled in the art that this invention is not limited to the details of the above exemplary embodiments, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects, and are intended to encompass all changes falling within the meaning and scope of equivalents within this invention.

Claims

1. A method for autonomous inspection of power facilities by unmanned aerial vehicles (UAVs) based on three-dimensional semantic driving, characterized in that, Includes the following steps: S1. Establish a world coordinate system, map the power facilities and drones onto the world coordinate system, and obtain the initial global 3D semantic point cloud of the facilities. and the initial pose of the drone The initial global facility three-dimensional semantic point cloud Including the initial global facility 3D point cloud and the initial global semantic label vector The drone is equipped with an onboard camera. S2, Collection of inspection task texts After preprocessing, relevant phrases about power components are extracted and matched with the power component category set. Alignment yields component category sequence space Then according to the category sequence space Generate component-level semantic inspection sequence ; S3. Based on the image acquired by the airborne camera at time t, obtain the pixel-level segmentation result of the target component through an open vocabulary visual perception network, and based on the depth map acquired by the airborne camera at time t... Camera intrinsic parameter matrix and the current pose of the drone After back-projecting the pixel-level segmentation results of the target component into three dimensions, the results are compared with the three-dimensional semantic point cloud of the global facility at the previous time step. The fusion process yields the current-time global facility 3D semantic point cloud. ; Step S3 specifically includes the following steps: S31. Based on a frame of RGB image captured by the airborne camera at time t Inputting an open-vocabulary visual perception network yields pixel-level segmentation results of target components within the current field of view. ,in This represents the number of target components detected in the current frame. The component category label for the i-th target component; This is the semantic mask for the i-th target component; S32, Traverse the pixel-level segmentation results All mask pixels For satisfying Mask pixels, based on the depth map acquired by the airborne camera at time t. Camera intrinsic parameter matrix and the current pose of the drone By backprojecting it in three dimensions, we obtain three-dimensional points in the world coordinate system, thus obtaining a set of three-dimensional points belonging to the i-th target component. ; S33, All three-dimensional points of each target component at time t The points are stacked column-wise to form the incremental 3D point cloud matrix at the current time. And construct its corresponding incremental semantic label vector. The two are used as incremental 3D semantic point clouds ; S34. Incremental 3D semantic point cloud Compared with the previous time step, the global facility 3D semantic point cloud The current moment's global facility 3D semantic point cloud is obtained by stitching and updating the data. ; S4, the first Category of electrical components awaiting inspection As the current target component category, then from the global facility 3D semantic point cloud Extract the 3D point set of the current target component category to generate representative waypoints for the current target component category. According to the current location of the drone and representative waypoints By using obstacle sets The constructed local path planning function yields a sequence of local control instructions. ; S5. Based on the current pose of the drone Local control command sequence The control steps are executed step by step to obtain the new position and attitude of the UAV at the next moment. And update the global facility 3D semantic point cloud. Then in the new moment Filter and generate the 3D point set of the current target component And based on the three-dimensional point set of the UAV and the current target component. The closest distance between Construct a semantic inspection sequence at the component level to determine the coverage state indicator. Have all target components been inspected and completed? 2. The method according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11. Establish a world coordinate system and map the geometric point cloud data of the power facilities onto the world coordinate system to obtain a global 3D point cloud of the facilities. and a global facility 3D point cloud Initialize the corresponding initial global semantic label vector The initial three-dimensional semantic point cloud of the facility was obtained; Indicates the number of 3D points of the initial facility; Represents the set of electrical component categories; t represents time. S12, Based on the initial position of the UAV With the initial rotation matrix Constructing the initial pose of the UAV in the world coordinate system .

3. The method according to claim 1, characterized in that, Step S2 specifically includes the following steps: S21, Collection of inspection task texts Preprocessing yields a standardized collection of inspection task texts. ; S22. Extract the inspection task text set Phrases related to electrical components and their association with categories of electrical components Alignment yields component category sequence space ; S23, Component category sequence space All component categories are merged and deduplicated, and the order of different components is determined according to the explicit order words and implicit logical ordering of the operation and maintenance procedures to obtain the component-level semantic inspection sequence. .

4. The method according to claim 1, characterized in that, Step S4 specifically includes the following steps: S41, the first Category of electrical components awaiting inspection As the current target component category, then from the global facility 3D semantic point cloud Filter out all points whose component category labels are equal to the current target component category to generate the current target component's 3D point set. l = 1, 2, ..., L; L is the total number of target components; S42. Select a non-empty 3D point set of the current target component. The geometric center of the point cloud serves as the representative waypoint of the current target component. ; S43. Based on the current location of the drone and the representative waypoint of the current target component Based on the global facility 3D semantic point cloud obstacle collection The constructed local path planning function yields a sequence of local control instructions. .

5. The method according to claim 1, characterized in that, Step S5 specifically includes the following steps: S51, based on the current position of the drone Local control command sequence The control step h is based on the flight control update function to execute flight control step by step to obtain the new pose of the UAV at the new moment. And update the global facility 3D semantic point cloud through step S3. ; The new time step; h represents the control step; S52, in every new moment Based on the global facility 3D semantic point cloud Filter out all component category labels that are equal to the current target component category. The points generate the current component's 3D point set. And based on the drone's current location Calculate the three-dimensional point set of the UAV and the current target component. The closest distance between ; S53, Based on the nearest distance Construct coverage status indicator For the current target component category Perform semantic coverage determination of components, if Then the current target component category is considered to be... The semantic coverage requirement has been met, and the current target component category will be... Mark as completed inspection, then select the current target component category. Switch to component-level semantic inspection sequence The next target component category Proceed to step S4 until the component-level semantic inspection sequence is completed. All elements in the process are marked as having been inspected; if This indicates the current target component category. The semantic coverage has not yet been satisfied; the goal is to maintain [the target]. Proceed to step S3 to continue execution.

6. The method according to claim 1, characterized in that, The open vocabulary visual perception network described in step S31 adopts an open vocabulary detection and segmentation framework. Its core structure consists of an image encoder, a text encoder, and a cross-modal matching localization head. The input is the current frame image. and the set of component categories The generated text prompts will be displayed on the network. The encoding is used to embed text and perform similarity matching with image features to locate the target element in the image. The corresponding region outputs the number of target components. and its masks and categories As the current frame image Pixel-level segmentation results of target components within the wide and high field of view.

7. The method according to claim 4, characterized in that, The local path planning function described in step S43 first constructs a three-dimensional occupation grid map or Euclidean distance field map in the world coordinate system, and then uses the current UAV position... For the starting point and the destination waypoint The endpoint is based on the set of obstacles. On the map, a collision-free discrete path is obtained using any feasible planning algorithm among A*, D Lite*, and RRT. Finally, the discrete path is smoothed and subjected to dynamic feasibility processing to obtain the trajectory, which is then discretized into a sequence of control commands according to the flight control interface. Output.

8. The method according to claim 5, characterized in that, In step S52, based on the current location of the drone Calculate the three-dimensional point set of the UAV and the current target component. The closest distance between As shown in the following formula: , in It is a Euclidean norm function; Indicates the current time With the current target component category The number of related points; Represents the current target component's three-dimensional point set The point in the middle.

9. The method according to claim 8, characterized in that, The coverage status indication quantity in step S53 The specific values ​​include: if there exists a time... Make and Then the coverage status indicator Set to 1; otherwise, the coverage status indicator is set to 1. Take 0; where, The distance threshold used to determine whether semantic coverage meets the standard.

Citation Information

Patent Citations

  • Unmanned aerial vehicle tour-inspection track acquisition method based on laser-point cloud data

    CN109141434A

  • Unmanned aerial vehicle inspection method based on high-precision positioning and visual tracking

    CN112269397A