Feature fusion method, system and device based on visual radar, medium and product

By adopting the feature fusion method of visual radar in the SLAM system, combining superpoint and PCPnet network to extract features, and fusion through the least squares method, the problem of insufficient robustness and accuracy of existing SLAM systems in complex environments is solved, and stable pose estimation and high-precision map construction are achieved.

CN120088606APending Publication Date: 2025-06-03INST OF ELECTRICAL ENG CHINESE ACAD OF SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510152320.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing SLAM systems have problems with insufficient robustness and accuracy in complex environments, especially in lighting changes, weak textures and dynamic environments, making it difficult to achieve stable pose estimation and high-precision map construction.

Method used

The feature fusion method based on visual radar is adopted, visual features are extracted through the superpoint network and point cloud features are extracted through the PCPnet network, and feature fusion is combined with the least squares method to generate fusion feature points to reduce the interference of environmental factors on feature extraction.

Benefits of technology

Realizing stable position estimation and high-precision map construction under variable external conditions improves the accuracy, robustness and real-time nature of the system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088606A_ABST
    Figure CN120088606A_ABST
Patent Text Reader

Abstract

The invention discloses a feature fusion method, system and device based on a visual radar, a medium and a product, and relates to the field of feature point extraction, and the method comprises the steps: obtaining original data, extracting the visual features of a to-be-processed image through employing a superpoint network, extracting the point cloud features of laser radar point cloud data through employing a PCPnet network, obtaining the point cloud curvature, and carrying out the feature fusion based on the point cloud curvature. Fusing the visual features and the laser radar point cloud data by using a feature fusion device to obtain fused feature points, aggregating the fused feature points by using a least square method to obtain line features and surface features, and obtaining a fused image according to the line features and the surface features; the invention provides a feature point fusion processing mode, interference of environmental factors on feature point extraction is reduced, stable position estimation and high-precision map construction are realized under variable external conditions, and the accuracy, robustness and real-time performance of a system in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of feature point extraction, and particularly to a feature fusion method, system, device, medium and product based on a vision radar. Background Art

[0002] With the rapid development of robots, drones and autonomous driving, the level of intelligence has been continuously improved. At the same time, the Simultaneous Localization And Mapping (SLAM) technology, as a key technology, is widely used in industrial manufacturing, smart home, disaster rescue and other scenarios. However, the complex external environment (such as rapid changes in light, occlusion, high-speed motion environment, texture loss and scenes with structural degradation) and the high requirements of micro-robot platforms for the real-time performance of SLAM algorithms have brought great challenges to the robustness and universality of existing SLAM systems. A single-sensor SLAM system is difficult to effectively cope with these challenges, often resulting in pose tracking failure and reduced mapping accuracy.

[0003] There are significant deficiencies in both radar and visual feature extraction methods in traditional SLAM systems. Radar feature extraction methods only rely on geometric features, which are not only vulnerable to interference in weak-texture scenes, but also have a large amount of computation, resulting in poor real-time performance of the system. In addition, radar feature extraction is relatively sensitive to changes in the initial point cloud. Once the initial feature positioning is inaccurate, the subsequent matching and positioning effects will be significantly affected.

[0004] In terms of visual SLAM, traditional methods (such as ORB-SLAM, VINS and DSO) usually rely on specific image features or optical flow tracking. However, these methods often exhibit the following disadvantages in weak-texture, light-changing or dynamic environments: (1) Limitations of weak texture and light changes: Visual feature extraction is difficult to reliably match feature points in cases of scarce texture or unstable lighting, resulting in an increase in pose estimation error. Low pixel utilization: Most feature extraction algorithms only use a small part of the pixels in the image (such as key points or sparse features), without making full use of the rich image information. (2) Dependence on camera stability: Traditional visual feature extraction has high requirements for the stability of the camera. It is easy to lose feature points in cases of camera jitter or rapid movement, affecting the robustness of the system. Both traditional radar and visual SLAM feature extraction methods have limitations in terms of real-time performance, accuracy and environmental adaptability.

[0005] In addition to traditional SLAM methods, the rapid development of deep neural networks in recent years has also promoted the innovation of lidar and visual SLAM technologies, and various algorithms combining SLAM and neural networks have emerged. However, these methods still have obvious deficiencies in practical applications. For example, the DeepCo algorithm regresses the pose transformation between point clouds through an end-to-end neural network structure. Although it simplifies the process of feature extraction and matching, it has low accuracy in complex environments and has high requirements for the scale of the network model. The DeepLO algorithm is an unsupervised lidar odometry method that optimizes pose estimation by using the NICP loss function (based on point cloud geometric information) through unsupervised training of front and back frame point clouds. However, this unsupervised training method has limited robustness in dynamic environments or high-noise situations.

[0006] In the field of visual SLAM, the PoseNet algorithm proposed by Kendall et al. was an early attempt to apply deep learning to indoor positioning, constructing an end-to-end positioning system. PoseNet avoids the complex registration of three-dimensional point clouds in the positioning stage and improves the running efficiency. However, due to the lack of constraints on geometric information, the network still lags behind traditional SLAM methods in terms of accuracy. In addition, the geometric loss problem of PoseNet also limits its application in precise positioning.

[0007] In summary, although the algorithms combining deep learning and SLAM show high running efficiency in some application scenarios and simplify some processing procedures, they still have significant deficiencies in terms of accuracy, robustness, and adaptability to dynamic environments. For example, visual SLAM feature points are easily affected by changes in lighting, while laser SLAM feature points show poor stability in degraded environments. These problems limit the application effect of existing SLAM systems in complex scenarios.

[0008] Therefore, there is an urgent need to provide a new method or system for tightly coupling visual radar feature extraction to reduce the interference of environmental factors on feature point extraction, and to achieve stable position estimation and high-precision map construction under changing external conditions, improving the accuracy, robustness, and real-time performance of the system in complex environments. Summary of the Invention

[0009] The purpose of this application is to provide a feature fusion method, system, device, medium, and product based on visual radar, which can reduce the interference of environmental factors on feature point extraction, and achieve stable position estimation and high-precision map construction under changing external conditions, improving the accuracy, robustness, and real-time performance of the system in complex environments.

[0010] To achieve the above purpose, this application provides the following solutions:

[0011] In a first aspect, the present application provides a feature fusion method based on a vision radar, including:

[0012] Obtain raw data; the raw data includes: an image to be processed and lidar point cloud data;

[0013] Use the superpoint network to extract the visual features of the image to be processed;

[0014] Use the PCPnet network to extract the point cloud features of the lidar point cloud data and obtain the point cloud curvature;

[0015] Based on the point cloud curvature, use a feature fuser to fuse the visual features with the lidar point cloud data to obtain fused feature points; the fused feature points include: line feature points and surface feature points; the feature fuser takes the feature points corresponding to each visual feature as the center, constructs a spherical neighborhood in the lidar point cloud data with a set radius, and performs feature point fitting within the spherical neighborhood to obtain fused feature points;

[0016] Use the least squares method to aggregate the fused feature points to obtain line features and surface features;

[0017] Obtain a fused image according to the line features and surface features.

[0018] Optionally, the obtaining of the raw data specifically includes:

[0019] Use an industrial camera to obtain the image to be processed;

[0020] Use a lidar to obtain the lidar point cloud data.

[0021] Optionally, the using the PCPnet network to extract the point cloud features of the lidar point cloud data and obtain the point cloud curvature specifically includes:

[0022] Use a spatial transformation network to normalize the lidar point cloud data to obtain normalized point clouds; based on the normalized point clouds, perform processing using a symmetric function to obtain symmetric point clouds;

[0023] Use the PCPnet network to extract the point cloud features of the symmetric point clouds and obtain the point cloud curvature.

[0024] Optionally, after using the PCPnet network to extract the point cloud features of the lidar point cloud data and obtain the point cloud curvature, it further includes:

[0025] Use the formula Convert the lidar point cloud data from a discretized state to a continuous state;

[0026] where Pimg are the three-dimensional coordinates of the visual feature; is the transformation matrix from the point cloud feature to the visual feature; P Lidar is the homogeneous representation of the point cloud feature in the lidar coordinate system.

[0027] Optionally, aggregating the fused feature points by using the least squares method to obtain line features and surface features, specifically including:

[0028] Using the least squares method to obtain the extreme points of the lidar point cloud data within the spherical neighborhood; the extreme points include: maximum points and minimum points;

[0029] Judging the feature classification of the fused feature points according to the distance between the extreme points and the fused feature points; the feature classification includes: line feature points, surface feature points and outlier points;

[0030] Aggregating the line feature points to obtain line features;

[0031] Aggregating the surface features to obtain surface features;

[0032] Removing the outlier points.

[0033] Optionally, judging the feature classification of the fused feature points according to the distance between the extreme points and the fused feature points, specifically including:

[0034] Using the formula ε 1 = ρ to obtain the first threshold; using the formula ε 2 = 2ρ to obtain the second threshold;

[0035] If the distance from the maximum point to the fused feature point is less than the first threshold, the current fused feature point is taken as the line feature point;

[0036] If the distance from the minimum point to the fused feature point is greater than the second threshold, the current fused feature point is taken as the surface feature point;

[0037] Otherwise, the current fused feature point is taken as the outlier point;

[0038] where ε 1 is the first threshold, ε 2 is the second threshold, and ρ is the local density of the fused feature point.

[0039] In a second aspect, the present application provides a feature fusion system based on vision radar, including:

[0040] A data acquisition module for acquiring original data; the original data includes: an image to be processed and lidar point cloud data;

[0041] A visual feature extraction module, configured to extract visual features of the to-be-processed image by using a superpoint network;

[0042] A point cloud curvature acquisition module, configured to extract point cloud features of the lidar point cloud data by using a PCPnet network and obtain point cloud curvature;

[0043] A feature fusion module, configured to fuse the visual features and the lidar point cloud data by using a feature fuser based on the point cloud curvature to obtain fused feature points; the fused feature points include: line feature points and surface feature points; the feature fuser takes the feature points corresponding to each visual feature as the center, constructs a spherical neighborhood in the lidar point cloud data with a set radius, and performs feature point fitting within the spherical neighborhood to obtain fused feature points;

[0044] A feature point aggregation module, configured to aggregate the fused feature points by using the least squares method to obtain line features and surface features;

[0045] A fused image acquisition module, configured to obtain a fused image according to the line features and surface features.

[0046] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the feature fusion method of the base visual lidar described in any one of the above.

[0047] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the feature fusion method of the base visual lidar described in any one of the above are implemented.

[0048] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the feature fusion method of the base visual lidar described in any one of the above are implemented.

[0049] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0050] The present application provides a feature fusion method, system, device, medium and product based on vision radar. By using multiple sensors, different types of raw data are obtained. On the basis of multi-sensor fusion, the superpoint network and the PCPnet network are introduced, and the extracted features are fused. In the present application, visual features are extracted through the unsupervised self-learning method of the superpoint network, the point cloud curvature is extracted through the PCPnet, and the line feature and the plane feature are obtained by using the least square fitting method. By effectively combining different types of raw data obtained by using multiple sensors with a deep learning model, and fusing visual features and point cloud features through the least square method, the interference of environmental factors on feature point extraction is reduced, and stable pose estimation and high-precision map construction are achieved under variable external conditions, overcoming the limitations of traditional SLAM methods, effectively solving the problem of unstable feature extraction caused by light changes and environmental degradation, and improving the accuracy, robustness and real-time performance of the system in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 Schematic flowchart of a feature fusion method based on vision radar in an embodiment of the present application;

[0053] Figure 2 Schematic diagram of an experimental vehicle and a sensor platform in an embodiment of the present application;

[0054] Figure 3 Schematic diagram of the SuperPoint structure in an embodiment of the present application;

[0055] Figure 4 Flowchart of the PCPnet network in an embodiment of the present application;

[0056] Figure 5 Schematic diagram of the preliminary fusion of radar laser point cloud data and visual features in an embodiment of the present application;

[0057] Figure 6 Schematic diagram of a feature fusion device in an embodiment of the present application;

[0058] Figure 7 Flowchart of the fitting of fused feature points in an embodiment of the present application;

[0059] Figure 8Effect comparison diagram of extraction in an embodiment of the present application; wherein, Figure 8 Part (a) is a schematic diagram of the accuracy rate of the matching algorithm, Figure 8 Part (b) is a schematic diagram of the recall rate of the algorithm matching;

[0060] Figure 9 Effect comparison diagram of the matching effects of the SuperPoint and ORB algorithms under different illuminations in an embodiment of the present application; wherein, Figure 9 Part (a) is a schematic diagram of the matching effect of SuperPoint under weak illumination, Figure 9 Part (b) is a schematic diagram of the matching effect of ORB under weak illumination, Figure 9 Part (c) is a schematic diagram of the matching effect of SuperPoint under strong illumination, Figure 9 Part (d) is a schematic diagram of the matching effect of ORB under strong illumination, Figure 9 Part (e) is a schematic diagram of the matching effect of SuperPoint under normal illumination, Figure 9 Part (f) is a schematic diagram of the matching effect of ORB under normal illumination;

[0061] Figure 10 Scene selection diagram in an embodiment of the present application; wherein, Figure 10 Part (a) is a schematic diagram of the outdoor environment of Scene 1, Figure 10 Part (b) is a schematic diagram of the outdoor environment of Scene 2, Figure 10 Part (c) is a schematic diagram of the outdoor environment of Scene 3, Figure 10 Part (d) is a schematic diagram of the indoor corridor environment with normal illumination in Scene 4, Figure 10 Part (e) is a schematic diagram of the indoor corridor environment under weak illumination in Scene 5, Figure 10 Part (f) is a schematic diagram of the indoor environment with obvious structure in Scene 6;

[0062] Figure 11 Line extraction effect diagrams of the traditional algorithm and the fusion algorithm under different illuminations in an embodiment of the present application; wherein, Figure 11 Part (a) is a schematic diagram of the line extraction effect of the traditional algorithm under weak illumination, Figure 11 Part (b) is a schematic diagram of the line extraction effect of the fusion algorithm under weak illumination, Figure 11 Part (c) is a schematic diagram of the line extraction effect of the traditional algorithm under strong illumination, Figure 11 Part (d) is a schematic diagram of the extraction effect of the fusion algorithm under strong illumination, Figure 11 Part (e) is a schematic diagram of the line extraction effect of the traditional algorithm under normal illumination, Figure 11 Part (f) is a schematic diagram of the line extraction effect of the fusion algorithm under normal illumination;

[0063] Figure 12 Schematic diagram of PcPNet normal vector visualization in an embodiment of the present application;

[0064] Figure 13 Feature extraction diagram of camera-lidar fusion in an embodiment of the present application;

[0065] Figure 14 Schematic diagram of accuracy and recall of surface feature extraction in an embodiment of the present application;

[0066] Figure 15 Schematic diagram of the structure of a computer device in an embodiment of the present application. Detailed implementation manners

[0067] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0068] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0069] In an exemplary embodiment, as Figure 1 shown, the present application provides a feature fusion method based on vision lidar, including the following S1-S6. Among them:

[0070] S1: Obtain original data; the original data includes: the to-be-processed image obtained by using a camera and the lidar point cloud data obtained by using a lidar.

[0071] S1 specifically includes:

[0072] In an exemplary embodiment, as Figure 2 shown, the original data can be collected by using an experimental vehicle platform and a computer device; the computer device is a desktop computer with a CPU of R5-5600 and a GPU of NVIDIA GTX 3060; the sensors carried by the experimental vehicle platform are a set of ADIS-16470 inertial navigation system, an industrial camera FLIR BFSPGE-31S4C, and a 16-line lidar Velodyne-16. In addition, an industrial control computer, a vehicle-mounted power supply, etc. are also carried. Among them, the industrial control computer is used for real-time processing of sensor data and real-time operation of the SLAM system; the industrial camera realizes direct image data transmission with the industrial control computer through a gigabit network cable; the inertial navigation system is connected through a CAN interface and transmitted to the industrial control computer; the lidar also uses a gigabit network cable for data transmission.

[0073] The Kalibr toolbox is used for joint calibration between the IMU and the camera. TheAutoware calibration toolbox is used for joint calibration between the camera and the lidar. Finally, the spatial transformation between the lidar and the IMU can be further calculated and transformed using the previously calculated extrinsic parameters.

[0074] S2: Use the superpoint network to extract the visual features of the image to be processed.

[0075] S2 specifically includes:

[0076] In a real environment, due to the complex external environment, the point cloud scanned by the lidar often ignores some corner points or misjudges a plane as multiple planes. Therefore, a method of using vision for pre-compensation is adopted for line and plane feature extraction. In this application, the feature extraction network of SuperPoint is used to extract feature points, and the extracted feature points are matched to the lidar point cloud for auxiliary line feature extraction.

[0077] The SuperPoint network is a self-supervised learning framework based on a fully convolutional network. The output of this network is feature points and descriptors, and its homography and estimation performance are relatively high. Moreover, the extraction of feature points in an environment with obvious illumination changes is more stable than traditional methods. The overall framework of the SuperPoint network is divided into three parts, namely, a shared encoder, a feature point extraction decoder, and a descriptor decoding network. Its overall network structure is as Figure 3 shown.

[0078] (1) Encoder:

[0079] The encoder of the SuperPoint network is similar to the VGG architecture, including 8 convolutional layers. The number of convolutional kernels in the first 4 convolutional layers is 64, and the number of convolutional kernels in the last 4 convolutional layers is 128. After each convolutional layer, a ReLU non-linear activation function is connected, and spatial downsampling pools are added after the 2nd, 4th, and 6th activation functions. In this way, the input dimension of the image is finally reduced.

[0080] (2) Key point decoder:

[0081] The encoded image with a length of H and a width of W is input into the key point decoder. First, 256 3x3 convolutional kernels are used to change the size of the image to H / 8 x W / 8 x 256. Finally, 65 1x1 convolutional kernels are used to reduce the dimension of the image to H / 8 x W / 8 x 65. After passing through the Softmax function, the channels without key point information can be removed, changing H / 8 x W / 8 x 65 to H / 8 x W / 8 x 64, and finally the image size is restored.

[0082] (3)Descriptor decoder:

[0083] The encoder output is also input into the descriptor decoder. After passing through the convolutional layer, the output is H / 8 x W / 8 x 256, and finally, the descriptor is obtained through interpolation and normalization processing.

[0084] (4)In this network, the loss function is set as:

[0085]

[0086] Among them, describes the meanings of the three components of the loss function used in the training process, namely the feature point detection loss, the descriptor loss, and the matching loss. X is the feature map output after the image passes through the SuperPoin network, X' is the resulting feature map after performing a homography transformation on the feature map, D is the feature map of the descriptor, usually a high-dimensional feature vector generated by the network in the feature extraction stage. D' is the descriptor feature map of X'. Y is the label value of the image feature points, and each pixel represents the probability of a feature point. Y' is the feature point detection result of the image X'. S is the matching relationship between the feature points, representing the geometric relationship between the corresponding feature points in X and X', and λ is a hyperparameter.

[0087]

[0088] Among them, is the loss function for the feature point extraction part, H C , W C are the total height and total width of the feature point map respectively, h and w represent the pixel height and width of the current feature point in the feature point map, and l p (x hw ; y hw ) is the loss function for a single pixel, used to measure the error between the true value y hw and the predicted value x hw .

[0089]

[0090] Among them, is the loss function for the descriptor calculation part, d hw is a bit vector that describes the feature information of the descriptor vector of the feature map D at (h, w). s hwh'w' is the matching relationship between the pixels (h, w) and (h', w') in D and D', usually 0 for non-matching and 1 for matching.

[0091] S3: Use the PCPnet network to extract the point cloud features of the lidar point cloud data and obtain the point cloud curvature.

[0092] S3 specifically includes:

[0093] The PCPnet network is used to process the lidar point cloud data obtained by the lidar, extract the preliminary features of the point cloud, and the PCPNet can well estimate the curvature of the point cloud. The PCPNet network process is as follows Figure 4 shown

[0094] Before inputting the lidar point cloud data into the PCPNet network, a spatial transformer is used to transform the lidar point cloud data into a standard pose, and then a set of functions with shared parameters are applied to each point of the lidar point cloud data. Then, a symmetric function is applied to the points to solve the problem of the unchanged order of the point cloud. Subsequently, a feature vector g is obtained through a feed-forward neural network i , and the spatial transformer is used again to act on the feature vector g i to obtain a 64x64 transformation matrix, and finally a feed-forward neural network is used to obtain h l . The calculation process is as follows

[0095] h l (p j ) = (FNN 2 °STN 2 )(g 1 (p j ), …, g 64 (p j ))

[0096] Among them, h l (p j ) is the result after feature extraction and transformation of the coordinates and initial features of point p j . STN 2 is the spatial transformer network, and FNN 2 is the feed-forward neural network. g 1 (p j ), …, g 64 (p j ) is the 64-dimensional feature obtained by applying multiple feature extraction functions g to p j .

[0097]

[0098] Among them, is the local area around the p point cloud, and H l is the sum of the transformation results of the points in all local areas

[0099]

[0100] Among them, H j is H l to H kAfter summation, a joint representation is used, and a network with three fully connected layers is employed to calculate the principal curvature direction and the central point curvature direction of the input laser point cloud.

[0101] S4: Based on the point cloud curvature, a feature fuser is used to fuse the visual features and the lidar point cloud data to obtain fused feature points; the fused feature points include: line feature points and surface feature points; the feature fuser takes the feature points corresponding to each visual feature as the center, constructs a spherical neighborhood in the lidar point cloud data with a set radius, and performs feature point fitting within the spherical neighborhood to obtain fused feature points.

[0102] S4 specifically includes:

[0103] Since the directly obtained lidar point cloud data is in a discretized state, in order to facilitate processing together with the image, the lidar point cloud data is transformed into a continuous state using a formula.

[0104]

[0105] where P Lidar is the homogeneous representation form of the three-dimensional coordinates of a point in the lidar coordinate system, is the transformation matrix from the lidar to the camera, and P img is the three-dimensional coordinates of the camera point.

[0106] As Figure 5 shown, the lidar point cloud data and the visual features are preliminarily fused, and then the point cloud is mapped to obtain the y - z plane, resulting in a matrix similar to a depth image. If the obtained point P img is not within the image size range, the point cloud that cannot be matched with the image is discarded.

[0107] To fuse the visual features and the line and surface features of the lidar point cloud data, as Figure 6 shown, a feature fuser is used to fuse the visual features and the lidar point cloud data to obtain fused feature points. This fuser can extract the line feature points and surface feature points of the point cloud based on the visual features extracted by the SuperPoint network and part of the lidar point cloud data, and output a matrix of WxHx3. The three channels of this matrix respectively display the pixel gray value, the feature index to which the pixel point belongs ( - 1 if it does not belong to the feature range), and the feature type ( - 1 if it does not belong to the feature range). Then, for a certain point cloud, it can be represented as (a, n, m), where a is the pixel value, and n and m are the corresponding indices. If it is not a feature point, it is (a, - 1, - 1).

[0108] Based on the point cloud curvature obtained from the PCPnet network, for each visual feature, the feature fuser establishes a spherical neighborhood with the visual feature as the origin and a radius of R. A KD-Tree is constructed within the neighborhood, and all lidar point cloud data within the neighborhood is searched. The visual feature and the lidar point cloud data are fused by fitting to obtain fused feature points.

[0109] S5: Aggregate the fused feature points using the least squares method to obtain line features and surface features.

[0110] S5 specifically includes:

[0111] In an exemplary embodiment, as Figure 7 shown, extract visual features and point cloud features to obtain point cloud curvature, and after fusing through the feature fuser to obtain fused feature points, aggregate the fused feature points using the least squares method. The least squares formula is as follows:

[0112]

[0113] where x i is the independent variable, c 0 is the first coefficient, c 1 is the second coefficient, c 2 is the third coefficient, k i is the true value of the sample.

[0114] Use the least squares method to obtain the maximum point and minimum point of the lidar point cloud data within the current domain, and judge the spatial distance between the two point clouds and the fused feature points. If the distance from the maximum point to the fused point is less than the first threshold ε 1 = ρ, then take the current fused feature point as a line feature point. If the distance from the minimum point to the fused feature point is greater than the second threshold ε 2 = 2ρ, where ρ is the local density of the fused feature point, related to the search radius R, and by default ρ = R / 2.

[0115] Then take the current fused feature point as a surface feature point, otherwise take the current fused feature point as an outlier.

[0116] Through this method, the curve of the plane where the visual feature points are located can be obtained. By comparing with the threshold, two types of feature points can be obtained. However, currently, it is still in the form of feature points and needs to be further aggregated so that the line feature points belonging to the feature line are aggregated together, and the surface feature points belonging to the feature surface are aggregated together. The aggregation process uses a breadth-first traversal method to generate a directed graph for the points. For each feature point, find 5 nearby feature points of the same type. Calculate the distance and angle between it and the parent node. If the distance is less than the specified distance threshold and the included angle is less than the angle threshold, it is considered to be the same type of point. In the experiment, the distance threshold is selected as 0.3m and the angle threshold is 5°. If it meets the conditions, add it to the directed graph; otherwise, discard the point. The judgment formula is as follows:

[0117]

[0118] Its principle is to match the line feature and the surface feature by the included angle of the line feature and the included angle of the surface feature normal vector. The two methods are slightly different, as follows:

[0119] (1) Line feature matching: The matching of line to line usually uses LBD (Line Band Discriptor). Its construction process is mainly divided into three parts: First, construct the scale space of the image. At each scale, extract line segments, and N groups of line features can be obtained. Then use LSR (Line Support Region) to construct LBD. Each LSR is divided into equally wide strips. Finally, calculate the descriptor of each strip, and all the descriptors are combined to form LBD. Assume that the line features in two consecutive frames of images are First, what needs to be judged is the angle θ 1 and θ 2 . If it satisfies:

[0120] |θ 1 -θ 2 |<α 1 .

[0121] When the absolute value of their difference satisfies less than a certain threshold, it can be considered that the two lines are in a parallel state. Then judge the length similarity of the two lines. Assume that the lengths of the two lines are l 1 and l 2 . If the proportion of the overlapping part is greater than the threshold, it is considered that the lengths of the two lines are similar.

[0122]

[0123] Finally, judge the distance between the centers of the two lines. The center point vector of line L 1 is m 1 =A - B. From this, L can be calculated1 The distance from the center point of 2 to L is d. If d is less than α 3 , then the line is considered to be matched.

[0124] d(m 1 ,L 2 ) < α 3 .

[0125] Among them, α 1 is a very small angle used to tolerate calculation errors, and α 2 is the length similarity threshold, which is between 0 and 1. And α 3 is the distance threshold, indicating whether it is close.

[0126] (2) Surface feature matching: Common methods for judging whether two planes are similar include the angle between the normal vectors of the two planes, the perpendicular distance between the normal vectors of the two planes, etc. The centroid and covariance matrix of the plane are mainly considered.

[0127] Suppose the centroid of a certain plane on the first frame of the image is m 1 , and the normal vector is n 1 . The centroid of a certain surface feature in the second frame is m 2 , and the normal vector is n 2 . Then the similarity measurement method of the two planes is as follows.

[0128]

[0129] Among them, S represents the number of points, n represents the normal vector, and S x→y represents the covariance.

[0130] The number of three-dimensional points inside plane x is represented by S 1 . The covariance of the centroid of plane x relative to plane y is represented by S x→y .

[0131]

[0132] Among them, p k is the three-dimensional coordinate of point k, and m y is the centroid coordinate of plane y. The sum of the calculated centroid offsets is the covariance matrix.

[0133] When performing plane matching, a two-stage method is used for matching. First, geometric features are used to screen out the items to be matched with higher credibility, and finally, the measurement formula is used for further matching calculation. When the measurement formulas V (x,y) and V (y,x) of the two planes are both less than the threshold, it is considered that the matching is successful.

[0134] S6: Obtain the fused image according to the line features and surface features.

[0135] S6 specifically includes:

[0136] In an exemplary embodiment, the overall algorithm is developed using the Robot Operating System (ROS). Utilizing the ROS platform and an improved deep learning framework, the extracted feature points are processed in real time, and the results are presented in a visual form for real-time perception and analysis of the surrounding environment.

[0137] When using the KITTI dataset for algorithm testing, compared with traditional algorithms, the improved algorithm is less affected by illumination changes and degraded environments, and has high accuracy in feature extraction and matching.

[0138] This application proposes a feature fusion method based on vision lidar to replace the feature point extraction part of the SLAM front end. Based on the feature point extraction method of visual feature and laser fusion and the point cloud curvature analysis algorithm, it can effectively identify and extract feature points in complex environments to cope with challenges such as illumination changes and environmental interference; adopt an unsupervised self-learning deep learning model to enhance the features of the point cloud data obtained by multiple sensors, improve the success rate and accuracy of feature point matching, thereby enhancing the overall performance of positioning and mapping; at the same time, utilize the improved deep learning framework under the ROS platform to process the extracted feature points in real time and present the results in a visual form for real-time perception and analysis of the surrounding environment, improving the accuracy, robustness and real-time performance of the system in complex environments.

[0139] In an exemplary embodiment, a feature fusion system based on vision lidar is provided, including:

[0140] A data acquisition module for acquiring raw data; the raw data includes: the image to be processed and lidar point cloud data.

[0141] A visual feature extraction module for extracting the visual features of the image to be processed using the superpoint network.

[0142] A point cloud curvature acquisition module for extracting the point cloud features of the lidar point cloud data using the PCPnet network and obtaining the point cloud curvature.

[0143] A feature fusion module for fusing the visual features and the lidar point cloud data using a feature fuser based on the point cloud curvature to obtain fused feature points; the fused feature points include: line feature points and surface feature points; the feature fuser constructs a spherical neighborhood in the lidar point cloud data with a set radius centered on the feature points corresponding to each visual feature, and performs feature point fitting within the spherical neighborhood to obtain fused feature points.

[0144] A feature point aggregation module, which is used to aggregate the fused feature points by using the least square method to obtain line features and surface features.

[0145] A fused image acquisition module, which is used to obtain a fused image according to the line features and surface features.

[0146] In an exemplary embodiment, in order to verify the progressiveness of the present application, relevant experiments were carried out. The experimental contents include:

[0147] (1) SuperPoint feature matching experiment.

[0148] In the SuperPoint feature matching experiment, the extraction and matching of SuperPoint feature points and ORB feature points were carried out on the HPatches dataset, and a comparative experiment was carried out. The HPathces dataset contains 59 sequences of images with perspective changes and 57 image sequences with illumination changes. It can verify the performance of the algorithm in different scenarios.

[0149] Comparison chart of the feature point extraction effects of the two algorithms under different illuminations, as Figure 8 shown, and the comparison chart of the matching effects, as Figure 9 shown.

[0150] Intuitively, it can be seen that compared with the distribution of feature points extracted by the traditional ORB algorithm, the result distribution obtained by using the SuperPoint algorithm is relatively uniform, and it still has a good performance in the case of poor illumination. The comparison experimental results of SuperPoint and ORB algorithm feature points are shown in Table 1.

[0151] Table 1 Comparison experimental results table of SuperPoint and ORB algorithm feature points

[0152]

[0153] Judging from the comparison results in Table 1, whether it is strong illumination, weak illumination or normal illumination, the performance of SuperPoint is better than that of the traditional algorithm. In the case of strong illumination, the proportion of SuperPoint matching numbers increased by 6.4 percentage points compared with the ORB algorithm. Under normal illumination, it increased by 6.5 percentage points. In the case of weak illumination, although the matching quantity and correct rate of both decreased significantly, SuperPoint was 32.8 percentage points higher than ORB.

[0154] (2) PcPNet network feature extraction experiment.

[0155] Separate the line feature extraction and surface feature extraction. In the network training of the experiment, for the PcPNet network, randomly select a point cloud area for training. For the point cloud neighborhood with obvious features, the size is set to 0.1, and for the relatively complex neighborhood, the size is set to 0.05. The fixed number of points in each neighborhood of the network is 1000. In each training cycle, each point cloud is iterated 10000 times. To improve the robustness of the PcPNet network, add random noise to the training data so that the training network can more accurately approximate the real data.

[0156] The experiment is based on the comparison between the LSD algorithm extraction and the LBD feature matching algorithm. Among them, the LSD line feature extraction algorithm obtains the pixel point set of the straight line through local analysis of the image, and then verifies and solves through the assumed parameters. Merge the pixel point set with the error control set, and then adaptively control the number of false detections. This algorithm is more commonly used in traditional vision SLAM based on line features, so this algorithm is used for comparison. The experimental results are evaluated using accuracy and recall. Accuracy refers to the number of correctly matched ones accounting for the number of correctly matched ones in the total data, while recall refers to the proportion of the number of correctly matched ones accounting for the total number of matches. For the selected experimental environment, the main body is distinguished by indoor and outdoor environments, and the randomly selected KITTI dataset environment is used for testing, while the local environment is used for testing in the indoor environment. The scene selection is as Figure 10 shown. By comparing the results of LSD+LBD and the fusion method of this application, statistics are carried out from two aspects of speed and accuracy, and the results are shown in Table 2.

[0157] Table 2 Comparison table of line feature extraction time (ms) and accuracy

[0158] Environment Number LSD + LBD Matching Accuracy Fusion Method Matching Accuracy Scenario 1 295.36 75.36% 305.13 85.36% Scenario 2 366.23 77.31% 370.54 88.49% Scenario 3 258.16 79.34% 255.47 87.69% Scenario 4 126.66 66.18% 134.32 89.55% Scenario 5 337.47 63.45% 315.69 91.06% Scenario 6 286.34 70.02% 278.53 90.57%

[0159] It can be seen from the results that the fusion algorithm takes approximately the same time as the traditional algorithm and can meet the real-time requirements. The accuracy of the fusion algorithm is on average 10% higher than that of the traditional algorithm, and the recall rate is on average increased by 17%.

[0160] To compare the feature extraction and matching under different illuminations in the same scene, select the local environment for the experiment, as Figure 11 shown.

[0161] By outputting the number of extracted features, under normal illumination, the number of line features extracted by the traditional algorithm is 127, while the number of line features of the fusion algorithm is 202; under weak illumination, the number of line features extracted by the traditional algorithm is 38, and the number of line features extracted by the fusion algorithm is 116; under normal illumination, the number of line features extracted by the traditional algorithm is 301, and the number of line features extracted by the fusion algorithm is 566. The number of features of the fusion algorithm is on average 50% higher than that of the traditional algorithm.

[0162] In previous experiments, an experimental comparison of line feature extraction was carried out, such as Figure 12 shown. For the extraction of surface features, the experimental session mainly demonstrated the algorithm effect and evaluated its own accuracy and recall rate. For the point cloud curvature obtained by PCPNet, the visualization result was presented in the form of labeled normal vectors, and finally the curvature visualization effect was achieved.

[0163] To demonstrate the effect of lines and surfaces after fusion extraction, as Figure 13 shown, the methods of point cloud visualization and image visualization were used to display the extracted line and surface features respectively. Figure 14 It shows the comparison of the accuracy and recall rate of surface feature extraction. Randomly select a scene for the visualization of surface features and the calculation of recall rate. The overall recall rate of surface features is 71.5%, and the accuracy is 79.68%.

[0164] This application proposes a new feature extraction method based on neural network fusion. By combining SuperPoint and PcPNet, the input of images and point clouds is processed separately and then fused, effectively reducing the impact of illumination changes and degraded point cloud environments. Experiments show that the accuracy and recall rate of feature extraction and matching of the fusion algorithm are higher than those of traditional extraction algorithms. At the same time, this application designs a least squares iterative fitting method for line and surface feature extraction. For each visual feature point, a spherical neighborhood with the visual feature point as the origin and a radius of R is established. Search all the lidar point clouds within the R neighborhood. In addition, this application sets new clustering parameters for the fitted features to make feature clustering more accurate.

[0165] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 15 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store visual features and lidar point cloud data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes a feature fusion system based on visual radar.

[0166] Those skilled in the art can understand that Figure 15 the structure shown in Figure 15 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in Figure 15 , or combine some components, or have different component arrangements. Figure 15 In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0167] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0168] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0169] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0170] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0171] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0172] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0173] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A feature fusion method based on visual radar, characterized in that: The feature fusion method based on visual radar includes: Acquire raw data; the raw data includes: images to be processed and laser radar point cloud data; Extracting visual features of the image to be processed using a superpoint network; The PCPnet network is used to extract point cloud features of the laser radar point cloud data and obtain point cloud curvature; Based on the point cloud curvature, a feature fuser is used to fuse the visual features with the laser radar point cloud data to obtain fused feature points; the fused feature points include: line feature points and surface feature points; the feature fuser takes the feature point corresponding to each visual feature as the center, constructs a spherical neighborhood with a set radius in the laser radar point cloud data, and performs feature point fitting in the spherical neighborhood to obtain fused feature points; The fused feature points are aggregated using the least square method to obtain line features and surface features; According to the line features and surface features, a fused image is obtained.

2. The feature fusion method based on visual radar according to claim 1 is characterized in that: The obtaining of original data specifically includes: Acquiring the image to be processed by using an industrial camera; The laser radar point cloud data is acquired using a laser radar.

3. The feature fusion method based on visual radar according to claim 1 is characterized in that: The method of extracting point cloud features of the laser radar point cloud data using the PCPnet network and obtaining point cloud curvature specifically includes: The laser radar point cloud data is standardized by using a spatial transformation network to obtain a standardized point cloud; based on the standardized point cloud, a symmetric function is used for processing to obtain a symmetric point cloud; The PCPnet network is used to extract the point cloud features of the symmetric point cloud and obtain the point cloud curvature.

4. The feature fusion method based on visual radar according to claim 1 is characterized in that: The PCPnet network is used to extract the point cloud features of the laser radar point cloud data and obtain the point cloud curvature, and then the following steps are further included: Using the formula Converting the laser radar point cloud data from a discrete state to a continuous state; Among them, P img is the three-dimensional coordinate of the visual feature; is the transformation matrix from the point cloud feature to the visual feature; Lidar is a homogeneous representation of the point cloud features in the lidar coordinate system.

5. The feature fusion method based on visual radar according to claim 1, characterized in that: The least square method is used to aggregate the fused feature points to obtain line features and surface features, specifically including: Obtaining the extreme points of the laser radar point cloud data in the spherical neighborhood by using the least squares method; the extreme points include: maximum points and minimum points; According to the distance between the extreme point and the fused feature point, the feature classification of the fused feature point is determined; the feature classification includes: line feature points, surface feature points and outliers; Aggregating the line feature points to obtain line features; Aggregating the surface features to obtain surface features; The outliers are removed.

6. The feature fusion method based on visual radar according to claim 5 is characterized in that: Judging the feature classification of the fused feature point according to the distance between the extreme point and the fused feature point specifically includes: The first threshold is obtained by using the formula ε1=ρ; the second threshold is obtained by using the formula ε2=2ρ; If the distance from the maximum point to the fused feature point is less than a first threshold, the current fused feature point is used as the line feature point; If the distance from the minimum point to the fused feature point is greater than a second threshold, the current fused feature point is used as the surface feature point; Otherwise, the current fused feature point is used as the outlier point; Among them, ε1 is the first threshold, ε2 is the second threshold, and ρ is the local density of the fused feature points.

7. A feature fusion system based on visual radar, characterized in that: The feature fusion system based on visual radar includes: A data acquisition module is used to acquire raw data; the raw data includes: images to be processed and laser radar point cloud data; A visual feature extraction module, used to extract visual features of the image to be processed using a superpoint network; A point cloud curvature acquisition module is used to extract point cloud features of the laser radar point cloud data using a PCPnet network and obtain point cloud curvature; A feature fusion module, for fusing the visual features with the laser radar point cloud data based on the point cloud curvature using a feature fuser to obtain fused feature points; the fused feature points include: line feature points and surface feature points; the feature fuser takes the feature point corresponding to each visual feature as the center, constructs a spherical neighborhood with a set radius in the laser radar point cloud data, and performs feature point fitting in the spherical neighborhood to obtain fused feature points; A feature point aggregation module, used to aggregate the fused feature points using the least square method to obtain line features and surface features; The fused image acquisition module is used to obtain a fused image based on line features and surface features.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the feature fusion method based on visual radar according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the feature fusion method based on visual radar described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the feature fusion method based on visual radar described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Electric power environment monitoring method and device based on joint calibration of camera and laser radar

    CN121883614A

  • Point cloud time-varying tracking modeling method, device, equipment and medium in complex dynamic environment

    CN122089785A

  • Methods, apparatus, equipment and media for time-varying point cloud tracking modeling in complex dynamic environments

    CN122089785B