A target ranging method and device based on three-dimensional point cloud and visual fusion

CN116500631BActive Publication Date: 2026-09-25ZHENGZHOU XINDA ADVANCED TECH RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211680944.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2026-09-25
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

然而,三维点云分割算法获取目标的准确度比图像目标检测算法的准确度低,且每次巡检都需要专业的无人机飞手控制无人机巡飞一次,成本高,且不能全时间段监测

Benefits of technology

[0038]本发明相对现有技术具有突出的实质性特点和显著的进步,具体的说,本发明预先控制无人机巡航以获取三维点云样本数据和二维视频样本影像,或者直接利用历史无人机巡航数据获取的旧三维点云样本数据和二维视频样本影像,来训练匹配坐标提取网络模型,解决了三维点云和二维图像数据的一对一匹配问题,使得后期进行隐患目标检测时只需获得二维视频影像,即可直接得到每个像素点在观察坐标系下的观测距离,提高图像目标测距的精度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116500631B_ABST
    Figure CN116500631B_ABST
Patent Text Reader

Abstract

The application provides a target ranging method and device based on three-dimensional point cloud and visual fusion, the target ranging method comprises the following steps: acquiring a two-dimensional video image on a power transmission line in real time, identifying a target area image according to the two-dimensional video image, and extracting a matching coordinate of each pixel point in the target area image by using a trained matching coordinate extraction network; performing weighted average on the matching coordinate in three dimensions of a physical coordinate system respectively to obtain a coordinate of each pixel point in a radar coordinate system, and obtaining an observation distance of each pixel point in an observation coordinate system based on the conversion of the radar coordinate system-observation coordinate system; wherein the matching coordinate extraction network is trained by using three-dimensional point cloud sample data and two-dimensional video sample images collected by a UAV. The application can solve the problem of matching between three-dimensional point cloud and two-dimensional image data, and improve the accuracy of image target ranging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional measurement technology, and more specifically, to a target ranging method and device based on the fusion of three-dimensional point clouds and vision. Background Technology

[0002] Transmission lines are a crucial component of the power grid system and a vital channel for power transmission; ensuring their stable operation is paramount. However, transmission lines are widely distributed, long, and operate under constant exposure to the natural environment. They are subjected not only to normal mechanical and electrical loads but also to interference from various external factors. Especially with urban development and the increasing number of construction sites, large construction machinery, due to insufficient safe distances from nearby power lines, frequently causes conductor discharge, leading to power outages and safety accidents. If these potential hazards are not detected and eliminated in a timely manner, they will seriously threaten the stable operation of the power grid system. Therefore, high-precision target detection and ranging methods are needed.

[0003] Traditional power transmission line inspections often rely on extensive management methods such as foot patrols, vehicle patrols, and on-site manpower monitoring, resulting in significant annual investment of manpower and resources with minimal benefits. Strengthening oversight through video image monitoring can effectively save on manpower and material costs. Currently, some line sections use AI-based image recognition to improve inspection efficiency to some extent, but this lacks target distance prediction, leads to numerous meaningless alarms, and results in a heavy workload for alarm processing.

[0004] Power grid users themselves create 3D laser point cloud data models of the lines under their jurisdiction for line data archiving or drone patrols. For example, CN115240093A presents an automatic inspection method for transmission channels based on the fusion of visible light and lidar point clouds, directly obtaining the coordinates of potential hazards through 3D point cloud segmentation. However, the accuracy of 3D point cloud segmentation algorithms in acquiring targets is lower than that of image target detection algorithms, and each inspection requires a professional drone pilot to control the drone for one patrol, resulting in high costs and the inability to monitor continuously. Therefore, maximizing the use of existing point cloud data for application innovation and management improvement is an important direction for power grid users' business exploration.

[0005] In order to solve the above problems, people have been seeking an ideal technological solution. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies by providing a target ranging method and apparatus based on the fusion of three-dimensional point clouds and vision.

[0007] To achieve the above objectives, the technical solution adopted by this invention is: a target ranging method based on the fusion of three-dimensional point clouds and vision, comprising the following steps:

[0008] Two-dimensional video images of the transmission line are acquired in real time. The target area image is identified based on the two-dimensional video image. The matching coordinate extraction network is used to extract the matching coordinates of each pixel in the target area image.

[0009] The weighted average of the matching coordinates in the three dimensions of the physical coordinate system is calculated to obtain the coordinates of each pixel in the radar coordinate system. Based on the transformation between the radar coordinate system and the observation coordinate system, the observation distance of each pixel in the observation coordinate system is obtained.

[0010] The training steps for the matching coordinate extraction network are as follows:

[0011] Using drones to collect 3D point cloud sample data and 2D video sample images, the labeled coordinates are obtained based on the mapping relationship between the 3D point cloud sample data and the 2D video sample images.

[0012] Two-dimensional image fusion data is generated based on two-dimensional video sample images and labeled coordinates;

[0013] The GAN network architecture is trained based on the fusion data of labeled coordinates and two-dimensional images to obtain the matching coordinate extraction network model.

[0014] Based on the above, the specific steps for obtaining the labeled coordinates based on the mapping relationship between 3D point cloud sample data and 2D video sample images are as follows:

[0015] Obtain 3D point cloud sample data [P] i ,X L ,Y L Z L After obtaining the two-dimensional video sample images, the two-dimensional video sample images are transformed to obtain two-dimensional image sample data [P]. i [,U,V,R,G,B];

[0016] Among them, P i ,i∈[1,n] represents a certain pose of the UAV, n is the total number of UAV flight poses in the collected data, X L ,Y L Z L Indicates pose P i The coordinates of the points are based on the radar coordinate system; U and V represent the coordinates of the points based on the pixel coordinate system, and R, G, B are the pixel values ​​corresponding to the U and V coordinates.

[0017] Using the radar coordinate system as the world coordinate system, the X coordinate in the 3D point cloud sample data is... L ,Y L Z L Transformed to points U′, V′ in pixel coordinate system, the transformation relationship is:

[0018]

[0019] In the formula, f is the camera's physical focal length, dx and dy represent the actual size of each pixel in the x and y directions, u0 and v0 represent the position of the image's center of symmetry in the pixel coordinate system, and t d The distance between the lidar and the visible light camera in the horizontal direction;

[0020] Let U′=U, V′=V, then we obtain U,V of the two-dimensional image sample data and X of the three-dimensional point cloud sample data. L ,Y L Z L The mapping relationship is established, and the pose P of the UAV is obtained based on the mapping relationship. i Below, the labeled coordinates of 3D point cloud sample data and 2D video sample images [P] i ,U,V,X L ,Y L Z L ].

[0021] Based on the above, the fusion formula used when generating 2D image fusion data from 2D video sample images and labeled coordinates is as follows:

[0022] Among them, I fuse For two-dimensional image fusion data, I Pi-1 For pose P i-1 Two-dimensional video sample images collected below, For pose P i Two-dimensional video sample images collected below, For pose P i+1 Two-dimensional video sample images acquired below; Mk i-1 for The corresponding labeled coordinates, Mk i For I Pi The corresponding labeled coordinates, Mk i+1 for The corresponding labeled coordinates.

[0023] Based on the above, the GAN network architecture includes a generator and a discriminator. The generator's network structure includes a shallow feature extraction module, a deep feature extraction module, a feature fusion module, and a result output module. The shallow feature extraction module includes two convolutional layers, and the deep feature extraction module includes 13 cascaded residual modules, each residual module including two convolutional layers and a ReLU activation function. The feature fusion module is an element-wise sum operator module. The result output module includes two CBL modules, two upsampling modules, and a 1×1 convolutional layer.

[0024] The discriminator consists of 8 convolutional layers, 2 fully connected (FC) layers, and a sigmoid activation function.

[0025] Furthermore, the present invention also provides a target ranging device based on the fusion of three-dimensional point cloud and vision, the target ranging device comprising:

[0026] The target region recognition module is used to recognize target region images based on the input two-dimensional video images;

[0027] The matching coordinate extraction network training module is used to acquire 3D point cloud sample data and 2D video sample images, obtain labeled coordinates based on the mapping relationship between the 3D point cloud sample data and the 2D video sample images, generate 2D image fusion data based on the 2D video sample images and labeled coordinates, and train the GAN network architecture based on the labeled coordinates and 2D image fusion data to obtain the matching coordinate extraction network model.

[0028] The matching coordinate extraction module is used to extract the matching coordinates of each pixel in the target region image using a trained matching coordinate extraction network.

[0029] The coordinate recovery module is used to perform a weighted average of the matched coordinates in the three dimensions of the physical coordinate system to obtain the coordinates of each pixel in the radar coordinate system.

[0030] The observation distance acquisition module is used to obtain the observation distance of each pixel in the observation coordinate system based on the transformation between the radar coordinate system and the observation coordinate system.

[0031] Furthermore, the present invention also provides a power transmission line inspection system, including a drone, a lidar and a visible light camera mounted on the drone, multiple network cameras located on the power transmission line, and the aforementioned target ranging device.

[0032] The drone flies along a built-in trajectory and adjusts its attitude;

[0033] The lidar collects three-dimensional point cloud data under different poses;

[0034] The visible light camera acquires two-dimensional video images in different poses;

[0035] The target ranging device is connected to the lidar and the visible light camera respectively, and a matching coordinate extraction network model is trained based on three-dimensional point cloud data and two-dimensional video images.

[0036] The network camera is used to acquire two-dimensional video images on the transmission line in real time;

[0037] The target ranging device is also connected to a network camera to identify target area images based on two-dimensional video images. It uses a trained matching coordinate extraction network to extract the matching coordinates of each pixel in the target area image. The matching coordinates are weighted and averaged in the three dimensions of the physical coordinate system to obtain the coordinates of each pixel in the radar coordinate system. Based on the transformation between the radar coordinate system and the observation coordinate system, the observation distance of each pixel in the observation coordinate system is obtained.

[0038] This invention has outstanding substantive features and significant progress compared to the prior art. Specifically, this invention pre-controls the drone to cruise and obtain three-dimensional point cloud sample data and two-dimensional video sample images, or directly uses old three-dimensional point cloud sample data and two-dimensional video sample images obtained from historical drone cruise data to train a matching coordinate extraction network model. This solves the one-to-one matching problem between three-dimensional point cloud and two-dimensional image data, so that when detecting hidden danger targets later, only two-dimensional video images are needed to directly obtain the observation distance of each pixel in the observation coordinate system, thereby improving the accuracy of image target ranging.

[0039] Meanwhile, since network cameras can monitor power transmission lines at all times, and the target ranging device performs real-time calculations on the two-dimensional video images obtained by the network cameras at all times, it can detect safety hazards at the first time. This solves the problem that existing power transmission line inspections all use drone patrols to monitor fixed locations and fixed time periods, which cannot detect safety hazards in a timely manner. The operation process is convenient and suitable for improving existing drone power transmission line inspections. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating the target ranging method in Embodiment 1 of the present invention.

[0041] Figure 2 This is a diagram of the GAN training network framework in Embodiment 1 of the present invention.

[0042] Figure 3 This is a diagram of the LRB residual module in Embodiment 1 of the present invention.

[0043] Figure 4 This is a diagram showing the output layer of the generator in Embodiment 1 of the present invention.

[0044] Figure 5 This is a schematic flowchart of the target ranging method in Embodiment 1 of the present invention. Detailed Implementation

[0045] The technical solution of the present invention will be further described in detail below through specific embodiments.

[0046] Example 1

[0047] like Figure 1 As shown, this embodiment provides a target ranging method based on the fusion of 3D point cloud and vision, including the following steps:

[0048] Two-dimensional video images of the transmission line are acquired in real time. The target area image is identified based on the two-dimensional video image. The matching coordinate extraction network is used to extract the matching coordinates of each pixel in the target area image.

[0049] The training steps for the matching coordinate extraction network are as follows:

[0050] Using drones to collect 3D point cloud sample data and 2D video sample images, the labeled coordinates are obtained based on the mapping relationship between the 3D point cloud sample data and the 2D video sample images.

[0051] Two-dimensional image fusion data is generated based on two-dimensional video sample images and labeled coordinates;

[0052] The GAN network architecture was trained based on the fusion data of labeled coordinates and two-dimensional images to obtain the matching coordinate extraction network model;

[0053] The matching coordinates are weighted and averaged in the three dimensions of the physical coordinate system to obtain the coordinates of each pixel in the radar coordinate system. Based on the transformation between the radar coordinate system and the observation coordinate system, the observation distance of each pixel in the observation coordinate system is obtained.

[0054] In practical implementation, the specific steps for identifying target region images based on 2D video image data are as follows: The 2D video image is input into a trained target detection network for target detection to obtain the target region image. The target detection network can be the YOLOv5 target detection network, or other target detection networks; in actual use, different detection networks can be selected according to different targets of interest.

[0055] Understandably, the targets to be monitored here could be large construction machinery, or other equipment, animals, and people that could threaten the stable operation of the power grid system.

[0056] Furthermore, in specific implementation, the steps for obtaining the labeled coordinates based on the mapping relationship between 3D point cloud sample data and 2D video sample images are as follows:

[0057] Obtain 3D point cloud sample data [P] i ,X L ,Y L Z L After obtaining the two-dimensional video sample images, the two-dimensional video sample images are transformed to obtain two-dimensional image sample data [P]. i [,U,V,R,G,B];

[0058] Among them, P i ,i∈[1,n] represents a certain pose of the UAV, n is the total number of UAV flight poses in the collected data, X L ,Y L Z L Indicates pose P i The coordinates of the points are based on the radar coordinate system; U and V represent the coordinates of the points based on the pixel coordinate system, and R, G, B are the pixel values ​​corresponding to the U and V coordinates.

[0059] Using the radar coordinate system as the world coordinate system, the X coordinate in the 3D point cloud sample data is... L ,YL,Z L (Unit: meters) Transformed to points U′, V′ in pixel coordinate system (unit: pixels), the transformation relationship is as follows:

[0060]

[0061] In the formula, f is the camera's physical focal length, dx and dy represent the actual size of each pixel in the x and y directions, u0 and v0 represent the position of the image's center of symmetry in the pixel coordinate system, and t d The distance between the lidar and the visible light camera in the horizontal direction;

[0062] This is because the hardware structure of the drone's lidar and visible light camera that collect data determines that the transformation from the radar coordinate system to the pixel coordinate system corresponding to the camera image only involves translation, not rotation.

[0063] Let U′=U, V′=V, then we obtain U,V of the two-dimensional image sample data and X of the three-dimensional point cloud sample data. L ,Y L Z L The mapping relationship is established, and the pose P of the UAV is obtained based on the mapping relationship. i Below, the labeled coordinates of 3D point cloud sample data and 2D video sample images [P] i ,U,V,X L ,Y L Z L ].

[0064] It is understandable that the labeled coordinates of 3D point cloud sample data and 2D video sample images have a many-to-one relationship. Therefore, it is necessary to further utilize the GAN network architecture to train the data in order to find matching 2D and 3D coordinates that are consistent with the distribution of labeled coordinates.

[0065] It is important to note that before training using the GAN network architecture, it is necessary to first generate 2D image fusion data based on 2D video sample images and labeled coordinates. Specifically, the fusion formula is as follows:

[0066] Among them, I fuse For fusion data of two-dimensional images, For pose P i-1 Two-dimensional video sample images collected below, For pose P i Two-dimensional video sample images collected below, For pose P i+1 Two-dimensional video sample images acquired below; Mk i-1 for The corresponding labeled coordinates, Mk i for The corresponding labeled coordinates, Mk i+1 for The corresponding labeled coordinates.

[0067] Subsequently, the fused data of the labeled coordinates and the two-dimensional image is input into the GAN network architecture for training to obtain the matching coordinate extraction network model, so that the labeled coordinates of the three-dimensional point cloud sample data and the two-dimensional video sample image are in a one-to-one relationship.

[0068] In specific implementation, such as Figure 2 As shown, the GAN network architecture includes a generator G and a discriminator D. The network structure of the generator G includes a shallow feature extraction module, a deep feature extraction module, a feature fusion module, and a result output module. The shallow feature extraction module includes two convolutional layers for extracting shallow features; the deep feature extraction module includes 13 cascaded residual modules and one convolutional layer for extracting deep features; as shown... Figure 4 As shown, each residual module includes two convolutional layers and a ReLU activation function; the feature fusion module is an element-wise sum operator module, used to fuse shallow and deep features by element-wise addition; the result output module includes two CBL modules, two upsampling modules and a 1×1 convolutional layer, wherein the CBL module includes conv+BN+relu.

[0069] The discriminator D contains 8 convolutional layers, 2 fully connected (FC) layers, and a sigmoid activation function. The 8 convolutional layers are used to extract features, the 2 FC layers are used to aggregate features, and the sigmoid activation function is used to output the discrimination probability.

[0070] In use, the labeled coordinates and fused 2D image data are input into the generator G of the GAN network architecture to generate 2D and 3D matching coordinates that are consistent with the distribution of the labeled coordinates.

[0071] The labeled coordinates and matching coordinates are input into the discriminator D of the GAN network architecture to determine the correctness of the input matching coordinates or labeled coordinates and output the discrimination result.

[0072] The generators G and D of the GAN network architecture are optimized and trained using the identification results. The final generator obtained is the matching coordinate extraction network model.

[0073] Furthermore, the specific steps for generator G to generate matching two-dimensional and three-dimensional coordinates that are consistent with the labeled coordinate distribution are as follows:

[0074]

[0075] in, It is a generator function, θ G This represents the network parameters of G, whose input is two-dimensional image fusion data. The output is the matching coordinates This indicates a concat operation, where h and w represent the height and width of the two-dimensional image.

[0076] Specifically, generator functions It consists of the extraction functions of the shallow feature extraction module, the extraction functions of the deep feature extraction module, and the output function of the result output module.

[0077] In this embodiment, the extraction function of the shallow feature extraction module is:

[0078] F shall =f1(I fuse =Relu1(W1*I) fuse )+Relu2(W2*Relu1(W1*I fuse ))

[0079] Among them, I fuse It is the input two-dimensional image fusion data, F shall The shallow features are represented by f1(·), the shallow feature extraction function is represented by f1(·), the Relu1(·) and Relu1(·) represent the Relu activation functions of each convolutional layer, and W1 and W2 represent the weight parameters of each convolutional layer in the shallow feature extraction module.

[0080] Furthermore, the deep feature extraction module extracts the depth features of the two-dimensional image fusion data from 13 cascaded LRB blocks; therefore, its extraction function is:

[0081] F deep =W d *(f 13 (f 12 (…f1(Ishall )…)))

[0082] Among them, F deep W represents deep features. d This represents the weight parameters of the last convolutional layer in this module;

[0083] Specifically, the structure diagram of each LRB block is as follows: Figure 3 As shown, the extraction formula is:

[0084]

[0085] Among them, input and Let f represent the input and output of the i-th (i = 1, 2, ..., 13) LRP block, respectively. i (·) represents the feature extraction function for the i-th LRB block; This represents the output of the (i-1)th LRP block.

[0086] After obtaining the shallow and deep features, further fusion of shallow and deep features is required to obtain fused shallow and deep features. Specifically, the process expression for fusion of shallow and deep features is as follows:

[0087] F = F shall +F deep

[0088] Where F represents the fusion feature between deep and shallow layers.

[0089] After obtaining the deep and shallow layer fusion features, the deep and shallow layer fusion features need to be sent to the result output layer to obtain the final matching coordinates.

[0090] Specifically, the structure diagram of the output layer is as follows: Figure 5 As shown, the output function of the result output layer is:

[0091] Mt = W 1×1 (up(CBL(up(CBL(F))\))

[0092] Where CBL(·) represents the conv+BN+relu operation layer, up(·) represents the upsampling convolutional layer, and W 1×1 (·) represents the weights of a 1×1 convolutional layer.

[0093] It is understandable that the purpose of generator G is to generate matching coordinates that are as close as possible to the labeled coordinates, and the purpose of discriminator D is to identify the correctness of the matching coordinates. Therefore, discriminator D uses the identification results to optimize the network parameters, and generator G uses the identification results, labeled coordinates and matching coordinates to optimize the network parameters.

[0094] Specifically, the optimization formula for generator G is:

[0095]

[0096]

[0097] G's goal is Approaching 1; It is a generator function, θ G This represents the network parameters of G.

[0098] It can be understood that the output of the discriminator D is a probability value, D(·)∈[0,1], where D(·)=1 indicates that the matching coordinate input is correct matching information; D(·)=0 indicates that the matching coordinate input is incorrect matching information; and D(·)=0.5 indicates that the matching coordinate input cannot be judged as correct.

[0099] Specifically, the optimization formula for discriminator D is as follows:

[0100]

[0101] The goal of discriminator D is to hope Approaching 1, Approaching 0, where, It is the discriminator function, θ D This represents the network parameters of D.

[0102] Since the trained matching coordinate extraction network extracts the matching coordinates of each pixel, in order to help the observer understand the observation distance, it is also necessary to obtain the observation distance of each pixel in the observation coordinate system based on the transformation between the radar coordinate system and the observation coordinate system; where the observation coordinate system is a coordinate system with the location of the person as the origin.

[0103] In practical implementation, the transformation formula between the radar coordinate system and the observation coordinate system is as follows:

[0104]

[0105]

[0106] Where, x o ,y o ,z o d represents the coordinates in the observation coordinate system. x ,d y ,d z Indicates the amount of translation change; ω,k represent the rotation angles, and the rotation matrix R is calculated with the origin of the radar coordinate system as the rotation point, and the rotations in the x, y, and z dimensions respectively. Obtained by rotating ω,k.

[0107] In summary, this embodiment utilizes existing 3D point cloud sample data and 2D video sample images to train a matching coordinate extraction network model. This allows for subsequent acquisition of only 2D video images, with the matching coordinate extraction network model extracting the matching coordinates of each pixel in the target area image within the 2D video images. Through coordinate restoration and transformation operations, the observation distance of each pixel in the observation coordinate system is directly obtained. This solves the problem of matching 3D point cloud and 2D image data, improves the accuracy of image target ranging, and is convenient to operate. It is suitable for improving existing UAV power transmission line inspection.

[0108] Furthermore, this invention eliminates the need for drones to conduct scheduled inspections at fixed times and locations. Instead, it utilizes visible light cameras for continuous monitoring and real-time calculations, enabling the immediate detection of safety hazards. This addresses the problem that existing power transmission line inspections rely on drone patrols for monitoring at fixed locations and time periods, which fails to detect safety hazards in a timely manner.

[0109] Example 2

[0110] Since the relationship between 2D coordinates and 3D coordinates is one-to-many, Example 1 uses a matching coordinate extraction network model to find one-to-one labeled coordinates. However, the found labeled coordinates may not be a correct 2D-to-3D correspondence, so it is necessary to identify the incorrect labeled coordinates. Specifically, if the image coordinates of the target region are adjacent, the corresponding 3D coordinates should also be adjacent. If a point is far away from these adjacent points, it is an outlier.

[0111] Furthermore, in order to eliminate these outliers, such as Figure 5 As shown, in this embodiment, after extracting the matching coordinates of each pixel in the target region image using the trained matching coordinate extraction network, an outlier filtering method based on the data center is used to remove outliers of the matching coordinates to obtain valid matching coordinates. Then, a weighted average is performed on the valid matching coordinates in the three dimensions of the physical coordinate system.

[0112] In practical implementation, the specific steps for using the data center-based outlier filtering method to remove outliers from the matching coordinates are as follows:

[0113] From the matching coordinates Mt=[U,V,X L ,Y L Z L Extract Mt = [X] from ] L ,Y L Z L Matrix data, data size Mt U,V =h×w;

[0114] The K-means clustering method is used to obtain the cluster center vector C = [C1, C2, C3], where C1, C2, and C3 are the three coordinate values ​​of the cluster center;

[0115] Let each row of matrix Mt subtract its corresponding cluster center vector to obtain matrix H = [X]. L -C1,Y L -C2,Z L -C3];

[0116] Performing the q-norm operation on each row of matrix H yields... q|H i | q <∞;

[0117] For vector H i Constrain the attribute values ​​to make

[0118] Choose a threshold α, and then divide the vector according to the threshold α. Elements greater than the threshold α are removed.

[0119] It is understandable that by filtering and removing outliers, the interference of outliers on target ranging can be reduced, thereby improving the accuracy of target ranging.

[0120] Example 3

[0121] This embodiment provides a target ranging device based on the fusion of 3D point cloud and vision, including:

[0122] The target region recognition module is used to recognize target region images based on the input two-dimensional video images;

[0123] The matching coordinate extraction network training module is used to acquire 3D point cloud sample data and 2D video sample images, obtain labeled coordinates based on the mapping relationship between the 3D point cloud sample data and the 2D video sample images, generate 2D image fusion data based on the 2D video sample images and labeled coordinates, and train the GAN network architecture based on the labeled coordinates and 2D image fusion data to obtain the matching coordinate extraction network model.

[0124] The matching coordinate extraction module is used to extract the matching coordinates of each pixel in the target region image using a trained matching coordinate extraction network.

[0125] The coordinate recovery module is used to perform a weighted average of the matched coordinates in the three dimensions of the physical coordinate system to obtain the coordinates of each pixel in the radar coordinate system.

[0126] The observation distance acquisition module is used to obtain the observation distance of each pixel in the observation coordinate system based on the transformation between the radar coordinate system and the observation coordinate system.

[0127] Specifically, the steps to obtain labeled coordinates based on the mapping relationship between 3D point cloud sample data and 2D video sample images are as follows:

[0128] Obtain 3D point cloud sample data [P] i ,X L ,Y L Z L After obtaining the two-dimensional video sample images, the two-dimensional video sample images are transformed to obtain two-dimensional image sample data [P]. i [,U,V,R,G,B];

[0129] Among them, P i ,i∈[1,n] represents a certain pose of the UAV, n is the total number of UAV flight poses in the collected data, X L ,Y L Z L Indicates pose P i The coordinates of the points are based on the radar coordinate system; U and V represent the coordinates of the points based on the pixel coordinate system, and R, G, B are the pixel values ​​corresponding to the U and V coordinates.

[0130] Using the radar coordinate system as the world coordinate system, the X coordinate in the 3D point cloud sample data is... L ,Y L Z L Transformed to points U′, V′ in pixel coordinate system, the transformation relationship is:

[0131]

[0132] In the formula, f is the camera's physical focal length, dx and dy represent the actual size of each pixel in the x and y directions, u0 and v0 represent the position of the image's center of symmetry in the pixel coordinate system, and t d The distance between the lidar and the visible light camera in the horizontal direction;

[0133] Let U′=U, V′=V, then we obtain U,V of the two-dimensional image sample data and X of the three-dimensional point cloud sample data. L ,Y L Z L The mapping relationship is established, and the pose P of the UAV is obtained based on the mapping relationship. i Below, the labeled coordinates of 3D point cloud sample data and 2D video sample images [P] i ,U,V,X L ,Y L Z L ].

[0134] In practical implementation, the fusion formula used when generating two-dimensional image fusion data based on two-dimensional video sample images and labeled coordinates is as follows:

[0135] Among them, I fuse For fusion data of two-dimensional images, For pose P i-1 Two-dimensional video sample images collected below, For pose P i Two-dimensional video sample images collected below, For pose P i+1 Two-dimensional video sample images acquired below; Mk i-1 for The corresponding labeled coordinates, Mk i for The corresponding labeled coordinates, Mk i+1 for The corresponding labeled coordinates.

[0136] In practical implementation, the GAN network architecture includes a generator and a discriminator. The generator's network structure includes a shallow feature extraction module, a deep feature extraction module, a feature fusion module, and a result output module. The shallow feature extraction module includes two convolutional layers, and the deep feature extraction module includes 13 cascaded residual modules, each residual module including two convolutional layers and a ReLU activation function. The feature fusion module is an element-wise sum operator module. The result output module includes two CBL modules, two upsampling modules, and a 1×1 convolutional layer.

[0137] The discriminator consists of 8 convolutional layers, 2 fully connected (FC) layers, and a sigmoid activation function.

[0138] Example 4

[0139] This embodiment provides a power transmission line inspection system, including a drone, a lidar and a visible light camera mounted on the drone, multiple network cameras located on the power transmission line, and the target ranging device based on three-dimensional point cloud and visual fusion as described in Embodiment 3.

[0140] The drone flies along a built-in trajectory and adjusts its attitude;

[0141] The lidar collects three-dimensional point cloud data under different poses;

[0142] The visible light camera acquires two-dimensional video images in different poses;

[0143] The target ranging device is connected to the lidar and the visible light camera respectively, and a matching coordinate extraction network model is trained based on three-dimensional point cloud data and two-dimensional video images.

[0144] The network camera is used to collect two-dimensional video images on the transmission line at all times;

[0145] The target ranging device is also connected to a network camera to identify target area images based on two-dimensional video images. It uses a trained matching coordinate extraction network to extract the matching coordinates of each pixel in the target area image. The matching coordinates are weighted and averaged in the three dimensions of the physical coordinate system to obtain the coordinates of each pixel in the radar coordinate system. Based on the transformation between the radar coordinate system and the observation coordinate system, the observation distance of each pixel in the observation coordinate system is obtained.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.

Claims

1. A target ranging method based on fusion of 3D point cloud and vision, characterized in that, Includes the following steps: Two-dimensional video images of the transmission line are acquired in real time. The target area image is identified based on the two-dimensional video image. The matching coordinate extraction network is used to extract the matching coordinates of each pixel in the target area image. The weighted average of the matching coordinates in the three dimensions of the physical coordinate system is calculated to obtain the coordinates of each pixel in the radar coordinate system. Based on the transformation between the radar coordinate system and the observation coordinate system, the observation distance of each pixel in the observation coordinate system is obtained. The training steps for the matching coordinate extraction network are as follows: Using drones to collect 3D point cloud sample data and 2D video sample images, the labeled coordinates are obtained based on the mapping relationship between the 3D point cloud sample data and the 2D video sample images. Two-dimensional image fusion data is generated based on two-dimensional video sample images and labeled coordinates; The GAN network architecture was trained based on the fusion data of labeled coordinates and two-dimensional images to obtain the matching coordinate extraction network model; The GAN network architecture includes a generator and a discriminator. The generator's network structure includes a shallow feature extraction module, a deep feature extraction module, a feature fusion module, and a result output module. The shallow feature extraction module includes two convolutional layers, and the deep feature extraction module includes 13 cascaded residual modules, each residual module including two convolutional layers and a ReLU activation function. The feature fusion module is an element-wise sum operator module. The result output module includes two CBL modules, two upsampling modules, and a 1×1 convolutional layer. The discriminator comprises 8 convolutional layers, 2 fully connected layers, and a sigmoid activation function; Among them, 8 convolutional layers are used to extract features, 2 fully connected layers are used to aggregate features, and the sigmoid activation function is used to output the discrimination probability; In use, the labeled coordinates and fused 2D image data are input into the generator G of the GAN network architecture to generate 2D and 3D matching coordinates that are consistent with the distribution of the labeled coordinates. The labeled coordinates and matching coordinates are input into the discriminator D of the GAN network architecture to determine the correctness of the input matching coordinates or labeled coordinates and output the discrimination result. The generators G and D of the GAN network architecture are optimized and trained using the identification results. The final generator obtained is the matching coordinate extraction network model.

2. The target ranging method according to claim 1, characterized in that, The specific steps to obtain the labeled coordinates based on the mapping relationship between 3D point cloud sample data and 2D video sample images are as follows: Obtain 3D point cloud sample data [ P i , X L , Y L , Z L After obtaining the two-dimensional video sample images, the two-dimensional video sample images are transformed to obtain two-dimensional image sample data. P i , U , V , R,G,B ]; in, P i , i ∈[1, n [This indicates a specific pose of the drone.] n This represents the total number of drone flight poses in the collected data. X L , Y L , Z L Express posture P i The coordinates of the points are based on the radar coordinate system. U , V This represents the coordinates of a point in a pixel coordinate system. R,G,B for U , V The pixel value corresponding to the coordinates; Using the radar coordinate system as the world coordinate system, the three-dimensional point cloud sample data is... X L , Y L , Z L Transformed into points in pixel coordinate system U ’ , V ’ The transformation relationship is as follows: ; In the formula, f It is the camera's physical focal length. dx and dy Indicates each pixel in x and y The actual size of the direction, u 0 and v 0 indicates the position of the image's center of symmetry in the pixel coordinate system. t d The distance between the lidar and the visible light camera in the horizontal direction; make U ’ = U , V ’ =V Two-dimensional image sample data are obtained. U , V With 3D point cloud sample data X L , Y L , Z L The mapping relationship is established, and the pose of the UAV is obtained based on the mapping relationship. P i Below, the labeled coordinates of 3D point cloud sample data and 2D video sample images [ P i , U , V , R,G,B ].

3. The target ranging method according to claim 2, characterized in that, The fusion formula used when generating 2D image fusion data based on 2D video sample images and labeled coordinates is: ; in, I fuse For fusion data of two-dimensional images, for position P i-1 Two-dimensional video sample images collected below, for position P i Two-dimensional video sample images collected below, for position P i+1 Two-dimensional video sample images collected below; Mk i-1 for The corresponding labeled coordinates Mk i for The corresponding labeled coordinates Mk i+1 for The corresponding labeled coordinates.

4. The target ranging method according to claim 1, characterized in that, The specific steps for identifying target region images from 2D video images are as follows: Input the 2D video image into a trained target detection network to perform target detection and obtain the target region image.

5. The target ranging method according to claim 2, characterized in that: After extracting the matching coordinates of each pixel in the target region image using a trained matching coordinate extraction network, an outlier point filtering method based on a data center is used to remove outliers in the matching coordinates to obtain valid matching coordinates. Then, a weighted average is calculated on the valid matching coordinates in the three dimensions of the physical coordinate system.

6. The target ranging method according to claim 5, characterized in that, The specific steps for removing outliers from matching coordinates using a data center-based outlier filtering method are as follows: From matching coordinates Mt =[ U , V , X L , Y L , Z L Take out from ] Mt =[ X L , Y L , Z L Matrix data, data size is Mt U,V = h × w ; Use the K-means clustering method to obtain cluster center vectors. C =[ C 1, C 2, C 3], of which C 1. C 2 and C 3 represents the three coordinate values ​​of the cluster center; Let matrix Mt The data in each row is subtracted from its corresponding cluster center vector to obtain a matrix. H =[ X L - C 1, Y L - C 2, Z L - C 3]; For matrix H Do each line q The order norm operation yields ; For vectors H i Constrain the attribute values ​​to make ; Select threshold α According to the threshold α vector The middle element is greater than the threshold α The removal.

7. The target ranging method according to claim 2, characterized in that, The transformation formula between the radar coordinate system and the observation coordinate system is: in, x o , y o , z o Indicates the coordinates in the observed coordinate system. d x , d y , d z Indicates the amount of translation change; φ,ω,k The rotation angle is represented by the rotation matrix R, which is based on the origin of the radar coordinate system as the rotation point. x,y,z According to the three dimensions respectively φ,ω,k Obtained by rotation.

8. A target ranging device based on fusion of 3D point cloud and vision, characterized in that, include: The target region recognition module is used to recognize target region images based on the input two-dimensional video images; The matching coordinate extraction network training module is used to acquire 3D point cloud sample data and 2D video sample images, obtain labeled coordinates based on the mapping relationship between the 3D point cloud sample data and the 2D video sample images, generate 2D image fusion data based on the 2D video sample images and labeled coordinates, and train the GAN network architecture based on the labeled coordinates and 2D image fusion data to obtain the matching coordinate extraction network model. The GAN network architecture includes a generator and a discriminator. The generator's network structure includes a shallow feature extraction module, a deep feature extraction module, a feature fusion module, and a result output module. The shallow feature extraction module includes two convolutional layers, and the deep feature extraction module includes 13 cascaded residual modules, each residual module including two convolutional layers and a ReLU activation function. The feature fusion module is an element-wise sum operator module. The result output module includes two CBL modules, two upsampling modules, and a 1×1 convolutional layer. The discriminator comprises 8 convolutional layers, 2 fully connected layers, and a sigmoid activation function; Among them, 8 convolutional layers are used to extract features, 2 fully connected layers are used to aggregate features, and the sigmoid activation function is used to output the discrimination probability; In use, the labeled coordinates and fused 2D image data are input into the generator G of the GAN network architecture to generate 2D and 3D matching coordinates that are consistent with the distribution of the labeled coordinates. The labeled coordinates and matching coordinates are input into the discriminator D of the GAN network architecture to determine the correctness of the input matching coordinates or labeled coordinates and output the discrimination result. The generators G and D of the GAN network architecture are optimized and trained using the identification results. The final generator obtained is the matching coordinate extraction network model. The matching coordinate extraction module is used to extract the matching coordinates of each pixel in the target region image using a trained matching coordinate extraction network. The coordinate recovery module is used to perform a weighted average of the matched coordinates in the three dimensions of the physical coordinate system to obtain the coordinates of each pixel in the radar coordinate system. The observation distance acquisition module is used to obtain the observation distance of each pixel in the observation coordinate system based on the transformation between the radar coordinate system and the observation coordinate system.

9. A transmission line inspection system, characterized in that: It includes a drone, a lidar and a visible light camera mounted on the drone, multiple network cameras located on the power transmission line, and the target ranging device as described in claim 8; The drone flies along a built-in trajectory and adjusts its attitude; The lidar collects three-dimensional point cloud data under different poses; The visible light camera acquires two-dimensional video images in different poses; The target ranging device is connected to the lidar and the visible light camera respectively, and a matching coordinate extraction network model is trained based on three-dimensional point cloud data and two-dimensional video images. The network camera is used to acquire two-dimensional video images on the transmission line in real time; The target ranging device is also connected to a network camera to identify target area images based on two-dimensional video images. It uses a trained matching coordinate extraction network to extract the matching coordinates of each pixel in the target area image. The matching coordinates are weighted and averaged in the three dimensions of the physical coordinate system to obtain the coordinates of each pixel in the radar coordinate system. Based on the transformation between the radar coordinate system and the observation coordinate system, the observation distance of each pixel in the observation coordinate system is obtained.

Citation Information

Patent Citations

  • Automatic power transmission channel inspection method based on fusion of visible light and laser radar point cloud

    CN115240093A

  • System and method for measuring power image distance of transmission line by unmanned aerial vehicle

    CN112525162A

  • Image processing method and system, and computer storage medium

    WO2022047625A1