Robot positioning method based on double-stage feature association
Through the robot positioning method of two-stage feature association, feature matching is combined with image semantics and point cloud feature information, the problem of insufficient feature matching error and robustness in the prior art is solved, and higher positioning accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510116640.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing feature matching methods have problems of insufficient matching error, robustness and accuracy in complex scenarios and dynamic environments, which affect the accuracy and stability of robot positioning.
The robot positioning method of two-stage feature association is adopted to obtain environmental data through fisheye cameras and lidars, and feature matching is performed by combining image semantics and point cloud feature information. The spatial commonality constraints and rigid body structural invariance are used for refined matching, reducing the influence of environmental factors and improving the robustness of the matching.
It improves the robustness and accuracy of feature matching and visual SLAM positioning, and can more effectively deal with the problem of light changes, texture loss and feature matching in dynamic environments, providing a stable and reliable robot positioning solution.
Smart Images

Figure CN120070574A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent robot positioning, and relates to a robot positioning method based on two-stage feature association. Background Art
[0002] Currently, with the booming development of robot technology, the SLAM (Simultaneous Localization and Mapping) technology, as one of the key technologies for autonomous navigation and environmental perception, has attracted much attention from researchers. Among them, visual feature matching, as a core component of the SLAM algorithm, helps the robot to perform real-time positioning and construct an environmental map by extracting and matching feature points or descriptors in images. However, the existing feature matching methods face a series of challenges in practical applications. When dealing with complex scenarios and dynamic environments, there are problems such as feature matching errors, insufficient robustness and accuracy, which restrict their wide application in actual scenarios.
[0003] In traditional visual SLAM methods, usually only a single feature description is used for feature matching between images. In classic visual algorithms such as ORB_SLAM and LSD-SLAM, the ORB and SIFT feature points they adopt are all matched using a single feature description. On the one hand, this method that only uses a single gray-scale feature is easily affected by environmental factors such as illumination changes and perspective changes. When the environmental illumination intensity changes, causing the gray-scale information of the feature points to change, the detection and matching of the feature points will become difficult, resulting in a decrease in matching accuracy. On the other hand, a single feature description cannot well describe and distinguish similar features. In the face of scenes with unclear features and repetitive texture regions, relying only on a single feature and local descriptions of very small regions often cannot well describe and distinguish similar features, and it is easy to have feature matching failures, resulting in poor positioning accuracy and robustness of the mobile robot. Therefore, it is necessary to solve the limitations of the current feature matching methods to provide more reliable support for the development of future intelligent robot navigation and environmental perception fields. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a robot positioning method based on two-stage feature association, mainly aiming at the problems of matching failures, reduced accuracy and robustness in scenarios such as illumination changes, texture loss, and texture duplication in current traditional feature matching methods. By using the idea of two-stage matching and combining multi-dimensional features such as image semantics and feature point spatial information, feature matching from coarse to fine is realized. At the same time, considering the influence of dynamic objects in the environment, dynamic feature points affecting robot positioning are screened out by combining regional average optical flow, providing a stable and robust robot positioning solution, which plays an important role in the future intelligent vehicle navigation and environmental perception fields and provides strong support for the development of intelligent vehicles.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A robot positioning method based on two-stage feature association, the method comprising the following steps:
[0007] S1: Obtain the image and point cloud data of the environment through a fisheye camera and a lidar, and construct a point cloud feature map by using a point cloud raster rendering method based on locally adaptive transformation control;
[0008] S2: Use a block space feature analysis method based on heterogeneous dynamic suppression fusion to extract the geometric description feature vector of the point cloud block;
[0009] S3: Adopt a method of block association based on spatial commonality constraints to establish the matching relationship between the semantic 3D region blocks of each point cloud block, and perform the matching association of the image blocks;
[0010] S4: Adopt a feature matching method with multi-branch constraints, use the rigid body structure invariance to construct a matching affinity factor, and perform refined matching of the feature points;
[0011] S5: Perform joint iterative optimization of region matching alignment and pose reconstruction to obtain the best robot pose estimation.
[0012] Further, step S1 includes the following steps:
[0013] S11: Install the IMU, fisheye camera and lidar at the preset positions of the intelligent robot respectively, and collect the IMU data, fisheye data and point cloud data respectively; use the OpenCalib toolbox to obtain the position transformation parameters between the camera and the radar, the camera and the IMU, and the internal parameters of the sensors;
[0014] S12: Use the longitude and latitude method to correct the distortion in the fisheye camera image, and input the corrected image into the EffcientNet neural network to extract the visual information F of the image feature points o , F o including the color, texture and semantic information of the image; use PCL to extract the point cloud feature F pi , F pi including the local density and normal information of the point cloud;
[0015] S13: Perform cross-modal diffusion of the visual features to the point cloud feature space to obtain the visually feature F after preliminary mapping v ; At the same time, design a point cloud raster rendering function based on adaptive local transformation control to dynamically adjust the propagation strength of the visual features in the point cloud feature space;
[0016] S14: Update the visual feature information corresponding to each point cloud using a diffusion rendering function controlled by adaptive transformation, and perform weighted fusion of the diffused visual features and the point cloud features to obtain the rendered point cloud fusion feature F p ′ i ; Then, perform DBSCAN clustering analysis on the rendered fusion features F p ′ i of all point clouds to construct the point cloud atlas F
[0017] Further, in step S13, use the CNN convolutional neural network model and the position calibration parameters of the sensor to obtain the mapping weight matrix W 1 from the visual feature space to the point cloud feature space, and obtain the preliminarily mapped visual feature F v , F v and F pi have the same dimension;
[0018] The point cloud raster rendering function based on adaptive local transformation control is related to the local point cloud geometric structure, local density, and visual feature differences. It dynamically adjusts the propagation strength of visual features in the point cloud feature space according to the complexity of the specific point cloud geometric structure. When more visual feature information is needed to make up for the deficiency of point cloud feature information, the diffusion rendering parameter is increased; otherwise, it is decreased; among them, the calculation formula for the diffusion rendering parameter controlled by adaptive transformation is:
[0019]
[0020] Among them, T diff (i,j) represents the diffusion intensity coefficient between point i and point j, N i represents the neighborhood of point i, n represents the total number of points in the neighborhood N i , D max , ρ max represents the maximum value of the difference in point cloud geometric features and density difference within N i for normalizing the difference, α geo , α dens respectively represent the influence weights of the point cloud geometric characteristics and local density on the diffusion intensity, ρ i , ρ j respectively represent the local densities of point i and point j, represents the visual feature vector of the visual feature information of point i after being diffused by adaptive transformation control;
[0021] In step S14, the calculation formula for the rendered point cloud fusion feature is:
[0022]
[0023] Among them, denotes the feature vector of the multi-attribute feature raster point cloud, p i represents the i-th point cloud to be rendered, β p , β v are respectively the hyperparameters that control the fusion of visual features and point cloud features.
[0024] Furthermore, step S2 includes the following sub-steps:
[0025] S21: Using the parameter fitting analysis method, obtain the object fitting parameter vectors of different models based on the point cloud data after raster rendering, and then construct the geometric anchor point feature vectors of different models in combination with the complexity of the internal structure of the point cloud;
[0026] S22: Dynamically adjust the fusion parameters between different models according to the differences in the fitting results between different models;
[0027] S23: Obtain the optimal block geometric anchor point feature vector according to the adjusted adaptive fusion parameters
[0028] Furthermore, in step S21, the parameter fitting analysis method first uses the point cloud semantic information to constrain the direction of point cloud clustering to obtain point cloud clusters with the same semantic information, and then uses the parameter fitting models based on density, spherical, and cylindrical shapes to obtain the object fitting parameter vectors of different models. At the same time, in combination with the complexity of the internal structure of the point cloud, construct the geometric anchor point feature vectors of different models. The calculation formula is as follows:
[0029]
[0030] Among them, q represents the tightness of the block point cloud distribution, and Q represents the internal complexity of the block point cloud, which is determined by the point cloud tightness q. denotes the block geometric anchor point feature vector obtained based on the density model. respectively denote the block geometric anchor point feature vectors obtained based on the spherical and cylindrical models, B a represents the point cloud data of the block in the current frame, and n represents the number of point clouds in the point cloud block B a in, d i represents S a the distance from the i-th point cloud in the block to the centroid point cloud, and μ represents the average distance from the points in the block to the centroid, C d represents the object fitting parameter vector based on the density model, C g represents the object fitting parameter vector based on the spherical model, C c represents the object fitting parameter vector based on the cylindrical model;
[0031] In step S22, the greater the difference between the fitting results based on the shape models, the less credible the fitting results based on the shape models are, and the contribution weight w is reduced. 1 and w 2 , and vice versa, the contribution weight is increased. The formula for calculating the fusion parameter is as follows:
[0032]
[0033] where w 1 represents the dynamic suppression fusion factor of the spherical fitting model, and w 2 represents the dynamic suppression fusion factor of the cylindrical fitting model. s g , s c respectively represent the similarity between the geometric feature vectors of the feature points and and . s d represents the similarity between the geometric feature vectors of the feature points and . λ g , λ c , λ d represent the fusion sensitivity coefficients of the spherical, cylindrical, and density fitting models. η g , η c , η d represents the similarity difference sensitivity coefficient of the geometric anchor point features based on the spherical, cylindrical, and density fitting models. The larger the value, the higher the sensitivity to distance;
[0034] In step S23, the optimal block geometric anchor point feature vector is obtained according to the adaptive fusion parameter The calculation formula is:
[0035]
[0036] where represents the optimal block geometric anchor point feature vector after fusion, respectively represent the block geometric anchor point feature vectors calculated by formula (3) in step S21, and w 1 represents the dynamic suppression fusion factor of the spherical fitting model, and w 2 represents the dynamic suppression fusion factor of the cylindrical fitting model.
[0037] Furthermore, step S3 includes the following sub-steps:
[0038] S31: Obtain the block geometric feature vectors of the current moment t and the previous moment t - 1 respectively and in the robot body coordinate system b at moment t tUsing the robot pose transformation matrix T obtained by the IMU as the reference coordinate system init Perform a spatial transformation to obtain the block geometric anchor vector in the same space-time dimension and
[0039] S32: Use the DBSCAN algorithm to perform clustering analysis on the block geometric anchor vectors in the same space-time dimension and to obtain the point cloud space region block sets at time t and time t-1 Then, comprehensively consider the spatial characteristics, texture features, and semantic information of each point cloud block, and calculate the commonality degree between the point cloud blocks at two adjacent times;
[0040] S33: According to the calculated spatial commonality degree of the point cloud blocks, select the pair with the highest similarity as the candidate best matching pair for each region, project it into the 2D pixel coordinate space of the image, transform the problem of 2D image semantic region block matching into the matching of 3D space blocks, and finally obtain the matching relationship of the 2D image region.
[0041] Furthermore, in step S32, the formula for calculating the commonality degree between the point cloud blocks at two adjacent times is expressed as:
[0042]
[0043] where represents the similarity between the i-th and j-th point cloud blocks in the point cloud block sets at two adjacent times , N and M are the numbers of point cloud blocks in the point cloud data at time t and time t-1 respectively, represents the point cloud block sets the i-th and j-th point cloud blocks in represents the geometric anchor feature vectors of the i-th and j-th point cloud blocks in the point cloud block sets at two times , θ ij represents the horizontal azimuth error between the i-th point cloud block and the j-th point cloud block, is expressed as the semantic similarity function of the point cloud blocks at two times and ; represents the average feature vector of the point cloud atlas inside the spatial block , represents the average feature vector of the point cloud atlas inside the spatial block , and ε is a constant greater than 1.
[0044] Furthermore, step S4 includes the following sub-steps:
[0045] S41: Gaussian blur is performed on the image regions that have completed association respectively, ORB feature points are extracted, SteerBRIEF feature descriptors are calculated, and preliminary feature point matching pairs are obtained by calculating the Hamming distance between the descriptors;
[0046] S42: Based on the invariance of grayscale, taking the preliminarily matched feature points as the center, considering the information of the local neighborhood, a square region with a preset size is taken, the optical flow vector is calculated using the optical flow equation, and the optical flow equation is solved by least squares to obtain the motion velocity V of the feature points in the u and v axis directions u and V v , and then the average predicted optical flow of the semantic region pixels is calculated Then calculate the difference between the measured optical flow of each feature point from the frame at time t - 1 to the frame at time t and the average predicted optical flow
[0047] S43: According to the invariance of the rigid body structure, that is, the corresponding optical flow changes of the feature points in the same rigid body region are consistent, the matching affinity factor is constructed with the regional average optical flow as the screening threshold;
[0048] S44: Improve the feature point matcher based on the matching affinity factor, calculate the matching accuracy between the feature points, and obtain the best matching point pairs according to the matching degree value M between the point pairs;
[0049] S45: Convert the camera pose solving problem into a non - linear optimization problem for solution to obtain the initially optimized robot pose.
[0050] Furthermore, in step S42, the average predicted optical flow of the semantic region pixels is calculated as:
[0051]
[0052] where, is denoted as the average predicted optical flow, n is denoted as the total number of feature points in the region, V u,k and V v,k respectively represent the motion velocity values of the k - th feature point in the u and v axis directions, and respectively represent the average motion velocity values of the semantic region in the u and v axis directions;
[0053] The difference between the measured optical flow of each feature point from the frame at time t - 1 to the frame at time t and the average predicted optical flow is calculated as:
[0054]
[0055] where, V i is denoted as the feature point The measured optical flow is obtained from the feature matching pairs obtained in step S41 and calculates the pixel change of the feature points. is the i-th feature point of the frame at time t-1;
[0056] In step S43, the matching affinity factor is defined as:
[0057]
[0058] where d represents the optical flow threshold of the feature points, and δ(i,j) is the matching affinity factor between the i-th and j-th feature points;
[0059] In step S44, the calculation formula for the best matching point pair is:
[0060]
[0061] where M i,j represents the matching accuracy between two feature points and ; represents the i-th feature point in region a of the image at time t-1, represents the j-th feature point in region b of the image at time t, f i and f j respectively represent the binary descriptors of two feature points and , n is the dimension of the descriptor, D f (·) represents the Hamming distance calculation function between two binary descriptors;
[0062] In step S45, the optimization problem is formulated as:
[0063]
[0064] where R * , t * are the robot rotation transformation matrix and translation transformation matrix to be solved respectively, u i represents the projection point coordinates of the observation of the camera at time t, P i represents the coordinates of the spatial point corresponding to the projection point u i at time t, s i represents the depth value of the 3D point observed by the camera at time t, T is the robot pose to be optimized, and K is the internal parameter matrix of the camera.
[0065] Furthermore, in step S5, after obtaining the initially optimized robot pose, joint iterative optimization is performed. First, update the robot pose T in step S3 initRecalculate the similarity between spatial blocks based on the values, solve and update the matching relationships of spatial region blocks, then execute the processes of S3 and S4 again to solve the pose of the robot. Iterate the above process several times, and take the transformation matrix with the minimum reprojection error as the optimal estimate of the robot's pose.
[0066] The beneficial effects of the present invention are as follows:
[0067] 1. Compared with the traditional feature matching visual SLAM positioning method, the robot positioning method based on two-stage feature association proposed by the present invention fully utilizes feature information such as feature point semantics, position change differences, local grayscale, and spatial relationships for feature matching and association, overcoming the limitation that traditional matching algorithms only use a single feature description for matching, thereby improving the robustness of feature matching and visual SLAM positioning.
[0068] 2. Compared with the existing feature matching methods, the present invention proposes a 3D-2D two-stage matching method, which introduces semantic and spatial constraints. By considering the semantic attributes and spatial position relationships of feature points, a rough matching of spatial regions is performed in the first stage to reduce the influence of environmental factors such as illumination changes; project the spatial regions onto the camera plane, and on the basis of plane region matching, perform refined point matching, reducing the feature search space from the entire image-based to region-based point matching, greatly reducing the redundancy of feature matching. Utilize the invariance of the rigid body structure of the features, combined with the velocity information of the feature points, to further screen out abnormal matches and improve the matching accuracy.
[0069] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0070] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0071] Figure 1 is the overall flowchart of the robot positioning method based on two-stage feature association of the present invention.
[0072] Figure 2 is the schematic diagram of the implementation scenario of the robot positioning method based on two-stage feature association of the present invention. Detailed Embodiments
[0073] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0074] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0075] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0076] Please refer to Figures 1 to 2 , which is a robot positioning method based on dual-stage feature association.
[0077] Embodiment
[0078] In this embodiment, first, sensors such as a fish-eye camera, an IMU, and a lidar are rigidly connected to the robot body. During the movement of the robot, the external transformation parameters between the camera, the IMU, and the central coordinate axis of the robot remain unchanged. A 32-line velodyne lidar is selected as the depth sensor, a Stek fish-eye camera is selected as the image data sensor, and a high-precision IMU of Epson is selected to collect the motion data of the intelligent vehicle. The algorithm is implemented in the Ubuntu system.
[0079] A robot positioning method based on dual-stage feature association designed by the present invention, its overall process is as Figure 1 shown, and specifically includes the following steps:
[0080] S1: Use the fisheye camera and lidar mounted on the robot to obtain the image and point cloud data of the environment; propose a point cloud raster rendering method based on local adaptive transformation control, and adaptively control the diffusion strength of visual feature information to the lidar point cloud feature space according to the geometric structure and distribution of the point cloud, and then construct a point cloud feature map.
[0081] Step S1 specifically includes the following steps:
[0082] S11: Use the IMU to collect the motion data of the robot and install it at the center position of the robot body; in order to make full use of the wide-angle characteristics of the fisheye camera, install the fisheye camera 0.4 meters directly in front of the center of the robot body and collect the image data of the environment in front of the robot at a frequency of 20 frames per second. The sensor installation schematic diagram is as Figure 2 shown. Use the OpenCalib toolbox to obtain the position transformation parameters between the camera and the radar, the camera and the IMU, and the internal parameters of the sensor.
[0083] S12: Use the longitude and latitude method to correct the distortion in the fisheye camera image to make it closer to the geometric shape of the real world, and input the corrected image into the EffcientNet neural network to extract the visual information F of the image feature points o , including the color, texture, semantic information, etc. of the image; use PCL to extract the point cloud feature F pi , including feature information such as the local density and normal of the point cloud.
[0084] S13: In order to achieve the cross-modal diffusion of visual features to the point cloud feature space, first use the CNN convolutional neural network model and the position calibration parameters of the sensor to obtain the mapping weight matrix W from the visual feature space to the point cloud feature space 1 , and obtain the preliminarily mapped visual feature F according to the mapping weight matrix v , at this time, F v and F pi have the same dimension; at the same time, a point cloud raster rendering function based on adaptive local transformation control is designed. This function is related to the local point cloud geometric structure, local density, and visual feature difference. The system will dynamically adjust the propagation strength of visual features in the point cloud feature space according to the complexity of the specific point cloud geometric structure. When the environmental point cloud geometric structure is simple and the point cloud density is small, more visual feature information is needed to make up for the lack of point cloud feature information, so the diffusion rendering parameter is increased; otherwise, the diffusion rendering parameter is reduced, so that the feature expression is more inclined to the point cloud feature; ensure better adaptation to the feature differences in different regions, avoid interference from irrelevant information, and improve the accuracy and reliability of cross-modal feature diffusion fusion. The calculation formula of the diffusion rendering parameter of adaptive transformation control is as follows:
[0085]
[0086] Among them, T diff (i, j) represents the diffusion intensity coefficient between point i and point j, and N i represents the neighborhood of point i, and n represents the total number of points in the neighborhood N i , D max , ρ max represents the maximum value of the difference in geometric features and density difference of the point cloud inside N, which is used to normalize the difference, and α i , α geo respectively represent the weights of the geometric characteristics of the point cloud and the local density on the diffusion intensity, and ρ dens , ρ i respectively represent the local densities of point i and point j. j represents the visual feature vector after the visual feature information of point i is controlled by adaptive transformation for diffusion.
[0087] S14: Update the visual feature information corresponding to each point cloud using the diffusion rendering function controlled by adaptive transformation, and perform weighted fusion of the diffused visual features and the point cloud features to obtain the fused feature of the rendered point cloud The formula is as follows:
[0088]
[0089] Among them, represents the feature vector of the raster point cloud with multiple attribute features, and p i represents the i-th point cloud to be rendered, and β p , β v are respectively hyperparameters for controlling the fusion of visual features and point cloud features; finally, through the rendered fusion features of all point clouds perform DBSCAN clustering analysis to construct the point cloud map F, where the point cloud map F is a vector matrix with m rows and 10 columns, and m is the number of point clouds completed in the current frame for rendering.
[0090] S2: Use the method of analyzing the spatial characteristics of the block based on heterogeneous dynamic suppression fusion to analyze the spatial characteristics of the block object, and obtain the point cloud block geometric description feature vector for characterizing the spatial characteristics of the block.
[0091] Step S2 specifically includes the following steps:
[0092] S21: Considering the influence of the irregularity of the point cloud cluster shape on the traditional center fitting method, the parameter fitting analysis method used is as follows: First, obtain the point cloud data after raster rendering from step S1, use the point cloud semantic information to constrain the direction of point cloud clustering to obtain point cloud clusters with the same semantic information, and then use the parameter fitting models based on density, sphere, and cylinder respectively to obtain the object fitting parameter vectors of different models. At the same time, combined with the complexity of the internal structure of the point cloud, construct the geometric anchor point feature vectors of different models. The calculation formula is as follows:
[0093]
[0094] Among them, q represents the tightness of the block point cloud distribution, and Q represents the internal complexity of the block point cloud, which is determined by the point cloud tightness q. represents the block geometric anchor point feature vector obtained based on the density model. respectively represent the block geometric anchor point feature vectors obtained based on the sphere and cylinder models, B a represents the point cloud data of the block in the current frame, and n represents the number of point clouds in the point cloud block B a in. i represents S a the distance from the i-th point cloud in the block to the centroid point cloud, and μ represents the average distance from the points in the block to the centroid, C d represents the object fitting parameter vector based on the density model, C g represents the object fitting parameter vector based on the sphere model, C c represents the object fitting parameter vector based on the cylinder model.
[0095] S22: According to the differences in the fitting results between different models, dynamically fuse the fitting results between different models. When the differences between the fitting results based on the shape models are greater, it means that the fitting results based on the shape models are less reliable, and reduce their contribution weights w 1 and w 2 . Conversely, increase their contribution weights to ensure the accuracy of the model's fitting of point cloud blocks with different shapes. The calculation formula for the fusion parameters is as follows:
[0096]
[0097] Among them, w 1 represents the dynamic suppression fusion factor of the sphere fitting model, and w 2 represents the dynamic suppression fusion factor of the cylinder fitting model, s g , s c respectively represent the geometric anchor point feature vectors and and the similarity between them, sd Expressed as a geometric anchor point feature vector and The similarity between, λ g , λ c , λ d Represents the spherical, cylindrical, density fitting model fusion sensitivity coefficient, η g , η c , η d Expressed as the similarity difference sensitivity coefficient of the geometric anchor point features based on the spherical, cylindrical, and density fitting models. The larger the value, the higher the sensitivity to distance, and its value is set according to the specific environment.
[0098] S23: Obtain the optimal block geometric anchor point feature vector from the adaptive fusion parameters The calculation formula is as follows:[[]]
[0099]
[0100] Among them, Represents the optimal block geometric anchor point feature vector after fusion, Respectively represent the block geometric anchor point feature vectors calculated by formula (3) in step S21, w 1 Represents the dynamic suppression fusion factor of the spherical fitting model, w 2 Represents the dynamic suppression fusion factor of the cylindrical fitting model.
[0101] S3: Based on S1 and S2, propose a method for block association based on spatial commonality constraints. Comprehensively consider the spatial characteristics, texture features, and semantic information of each point cloud block, calculate the commonality metric between adjacent point cloud blocks at two moments, establish the matching relationship between semantic 3D region blocks, and then complete the matching association of image blocks.
[0102] Step S3 specifically includes the following steps:[[]]
[0103] S31: Use the method in step S2 above to obtain the block geometric anchor point feature vectors of the current moment and the previous moment respectively and To unify the spatio-temporal dimensions of the block geometric anchor points before and after, take the robot body coordinate system b at time t t As the reference coordinate system, use the robot pose transformation matrix T obtained by imu init Perform spatial transformation to obtain the block geometric anchor point vectors in the same spatio-temporal dimension and
[0104] S32: According to the b at time t and t - 1 obtained in S31 tThe geometric anchor point feature information of the spatial block in the coordinate system is then clustered using the DBSCAN algorithm to obtain the point cloud spatial area block set at time t and time t-1 Taking into account the spatial characteristics, texture features and semantic information of each point cloud block, the commonality between point cloud blocks at two adjacent moments is calculated. The calculation formula is as follows:
[0105]
[0106] in, Represents a set of point cloud blocks at two adjacent moments The similarity between the i-th and j-th point cloud blocks in , N and M are the number of point cloud blocks in the point cloud data at time t and time t-1 respectively, Represents a collection of point cloud blocks The i-th and j-th point cloud blocks in Represents a set of point cloud blocks at two times The geometric anchor feature vectors of the i-th and j-th point cloud blocks in θ ij represents the horizontal orientation error between the i-th point cloud block and the j-th point cloud block, Represented as point cloud blocks at two times and The semantic similarity function is 1 when the semantic attributes are the same, and 0.4 when the semantic attributes are different. Represented as a spatial block The average eigenvector of the internal point cloud atlas, Represented as a spatial block The average eigenvector of the internal point cloud atlas, ε is a constant with a value greater than 1 to avoid division by zero errors when the vector distance difference is zero.
[0107] S33: According to the calculated spatial commonality of the point cloud blocks, the one with the highest similarity in each area is selected as the candidate best matching pair, which is reprojected into the 2D pixel coordinate space of the image, and the problem of 2D image semantic area block matching is converted into 3D space block matching, and finally the matching relationship of the 2D image area is obtained.
[0108] S4: A multi-branch constrained feature matching method is proposed. It uses the rigid body structure invariance, that is, the consistent movement trend of feature points in the same region, and calculates the average optical flow of the region as the screening threshold to construct a matching affinity factor to achieve further refined matching of feature points. Finally, the joint iterative optimization of regional matching alignment and pose reconstruction is performed to obtain the best robot pose estimation.
[0109] Step S4 specifically includes the following steps:
[0110] S41: Obtain the image block matching pairs according to the above S3 step. Gaussian blur is performed on the image regions that have completed association respectively. Extract ORB feature points, calculate the Steer BRIEF feature descriptors, and obtain the preliminary feature point matching pairs by calculating the Hamming distance between the descriptors.
[0111] S42: Based on the preliminary matching results obtained in S41 step, considering the gray invariance, taking a 6*6 square region centered on the feature points and considering the information of the local neighborhood, calculate the optical flow vector using the optical flow equation, and solve the optical flow equation by least squares to obtain the motion velocities V u and V v in the u and v axis directions of the feature points, and then calculate the average predicted optical flow of the pixels in the semantic region The calculation formula is as follows:
[0112]
[0113] where, represents the average predicted optical flow, n represents the total number of feature points in the region, V u,k and V v,k respectively represent the motion velocity values of the k-th feature point in the u and v axis directions, and respectively represent the average motion velocity values of the semantic region in the u and v axis directions; then, calculate the difference between the measured optical flow of each feature point from the frame at time t-1 to the frame at time t and the average predicted optical flow The calculation formula is as follows:
[0114]
[0115] where, V i represents the measured optical flow of the feature point which is obtained from the feature matching pairs obtained in S41 step and calculates the pixel change of the feature point, is the i-th feature point of the frame at time t-1, represents the average predicted optical flow, represents the difference between each optical flow vector and the average optical flow vector in the semantic region from the previous frame to the next frame.
[0116] S43: According to the invariance of the rigid body structure, that is, the corresponding optical flow changes of the feature points in the same rigid body region are consistent, construct a matching affinity factor with the regional average optical flow as the screening threshold, and its definition is as follows:
[0117]
[0118] where, d represents the optical flow threshold of the feature point, and δ(i,j) is the matching affinity factor between the i-th and j-th feature points.
[0119] S44: Improve the feature point matcher based on the matching affinity factor, calculate the matching accuracy between feature points, and obtain the best matching point pairs according to the matching degree value M between point pairs. The calculation formula is as follows:
[0120]
[0121] where M i,j represents the matching accuracy between two feature points and , represents the i-th feature point in region a of the image at time t-1, represents the j-th feature point in region b of the image at time t, f i and f j respectively represent the binary descriptors of two feature points and . n is the dimension of the descriptor, which is set to 256 here. D f (·) represents the Hamming distance calculation function between two binary descriptors.
[0122] S45: Convert the camera pose solution problem into a non-linear optimization problem, and the optimization problem is described as follows:
[0123]
[0124] where, R * , t * are the robot rotation transformation matrix and translation transformation matrix to be solved respectively. u i represents the projected point coordinates of the camera observation at time t. P i represents the coordinates of the spatial point corresponding to the projected point u i at time t. s i represents the depth value of the 3D point observed by the camera at time t. T is the robot pose to be optimized, and K is the internal parameter matrix of the camera.
[0125] S5: After obtaining the initially optimized robot pose, perform joint iterative optimization. First, update the value of the robot pose T init in step S3, recalculate the similarity between spatial blocks, solve and update the matching relationship of spatial region blocks, and then execute the S3 and S4 processes again to solve the robot pose. Iterate the above process 5 times, and take the transformation matrix with the minimum reprojection error as the optimal estimate of the robot pose.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A robot positioning method based on two-stage feature association, characterized in that: The method comprises the following steps: S1: The image and point cloud data of the environment are acquired through fisheye camera and lidar, and the point cloud raster rendering method based on local adaptive transformation control is used to construct the point cloud feature map; S2: Use the block spatial characteristic analysis method based on heterogeneous dynamic suppression fusion to extract the geometric feature vector of the point cloud block; S3: Using the block association method based on spatial commonality constraints, the matching relationship between the semantic 3D area blocks of each point cloud block is established to perform matching association of image blocks; S4: A multi-branch constrained feature matching method is used to construct a matching affinity factor using the rigid body structure invariance to perform refined matching of feature points; S5: Perform joint iterative optimization of region matching alignment and pose reconstruction to obtain the best robot pose estimation.
2. The robot positioning method based on two-stage feature association according to claim 1, characterized in that: Step S1 includes the following steps: S11: Install the IMU, fisheye camera and lidar at the preset positions of the intelligent robot, collect IMU data, fisheye data and point cloud data respectively; use the OpenCalib toolbox to obtain the position transformation parameters between the camera and the lidar, the camera and the IMU and the internal parameters of the sensor; S12: Use the latitude and longitude method to correct the distortion in the fisheye camera image, and input the corrected image into the EffcientNet neural network to extract the visual information of the image feature points F o , F o Including image color, texture, and semantic information; using PCL to extract point cloud features Includes local density and normal information of the point cloud; S13: Perform cross-modal diffusion of visual features to the point cloud feature space to obtain the visual features F after preliminary mapping v ; At the same time, a point cloud raster rendering function based on adaptive local transformation control is designed to dynamically adjust the propagation strength of visual features in the point cloud feature space; S14: Use the diffusion rendering function controlled by adaptive transformation to update the visual feature information corresponding to each point cloud, and perform weighted fusion of the diffused visual features and point cloud features to obtain the rendered point cloud fusion features. Then, by rendering and fusing features of all point clouds Perform DBSCAN clustering analysis and construct the point cloud atlas F.
3. The robot positioning method based on two-stage feature association according to claim 2 is characterized in that: In step S13, the CNN convolutional neural network model and the position calibration parameters of the sensor are used to obtain the mapping weight matrix W1 from the visual feature space to the point cloud feature space, and the visual feature F after preliminary mapping is obtained according to the mapping weight matrix v , F v and F pi have the same dimensions; The point cloud raster rendering function based on adaptive local transformation control is related to the local point cloud geometric structure, local density and visual feature differences. It dynamically adjusts the propagation strength of visual features in the point cloud feature space according to the complexity of the specific point cloud geometric structure. When more visual feature information is needed to make up for the lack of point cloud feature information, the diffusion rendering parameter is increased; otherwise, the diffusion rendering parameter is reduced. The calculation formula of the diffusion rendering parameter controlled by adaptive transformation is: Among them, T diff (i,j) represents the diffusion intensity coefficient between point i and point j, N i represents the domain of point i, and n represents the domain N i The total number of midpoints, D max ,ρ max N i The maximum value of the geometric feature difference and density difference of the inner point cloud is used to normalize the difference, α geo ,α dens Respectively represent the weights of the geometric characteristics of the point cloud and the local density on the diffusion intensity, ρ i ,ρ j denote the local density of point i and point j respectively, The visual feature vector representing the visual feature information of point i after diffusion through adaptive transformation control; In step S14, the rendered point cloud is fused with features The calculation formula is: in, Represented as the feature vector of the multi-attribute feature raster point cloud, p i represents the i-th point cloud to be rendered, β p ,β v are the hyperparameters that control the fusion of visual features and point cloud features respectively.
4. The robot positioning method based on two-stage feature association according to claim 2 is characterized in that: Step S2 includes the following sub-steps: S21: The object fitting parameter vectors of different models are obtained based on the point cloud data after raster rendering using the parameter fitting analysis method, and then the geometric anchor point feature vectors of different models are constructed based on the complexity of the internal structure of the point cloud; S22: Dynamically adjust the fusion parameters between different models according to the differences in fitting results between different models; S23: Obtain the optimal block geometry anchor feature vector based on the adjusted adaptive fusion parameters 5. The robot positioning method based on two-stage feature association according to claim 4 is characterized in that: In step S21, the parameter fitting analysis method first uses the semantic information of the point cloud to constrain the direction of point cloud clustering to obtain point cloud clusters with the same semantic information, and then uses density-based, spherical, and cylindrical parameter fitting models to obtain object fitting parameter vectors of different models. At the same time, combined with the complexity of the internal structure of the point cloud, the geometric anchor point feature vectors of different models are constructed. The calculation formula is as follows: Among them, q represents the density of the block point cloud distribution, Q represents the internal complexity of the block point cloud, which is determined by the density of the point cloud q. represents the block geometric anchor point feature vector obtained based on the density model, Denote the block geometric anchor feature vectors based on the spherical and cylindrical models, respectively, and B a Represents the point cloud data of the block in the current frame, n represents the point cloud block B a The number of point clouds in , d i Indicates S a The distance from the i-th point cloud to the centroid point cloud in the block, μ is represented as the average distance from the point in the block to the centroid, C d represents the object fitting parameter vector based on the density model, C g represents the object fitting parameter vector based on the spherical model, C c It is represented as a vector of object fitting parameters based on a cylindrical model; In step S22, when the difference between the fitting results based on the shape model is greater, it means that the result based on the shape model fitting is less credible, and its contribution weights w1 and w2 are reduced, otherwise, its contribution weights are increased. The calculation formula of the fusion parameter is as follows: Among them, w1 represents the dynamic suppression fusion factor of the spherical fitting model, w2 represents the dynamic suppression fusion factor of the cylindrical fitting model, and s g ,s c Respectively represented as geometric point feature vectors and and The similarity between d Represented as geometric point feature vector and The similarity between g ,λ c ,λ d Represents the fusion sensitivity coefficient of spherical, cylindrical and density fitting models, η g ,η c ,η d It is expressed as the similarity difference sensitivity coefficient of geometric anchor point features based on spherical, cylindrical, and density fitting models. The larger the value, the higher the sensitivity to distance. In step S23, the optimal block geometric anchor point feature vector is obtained according to the adaptive fusion parameters The calculation formula is: in, represents the best block geometric anchor feature vector after fusion, They are respectively represented as the block geometric anchor point feature vectors calculated by formula (3) in step S21, w1 is represented as the dynamic suppression fusion factor of the spherical fitting model, and w2 is represented as the dynamic suppression fusion factor of the cylindrical fitting model.
6. The robot positioning method based on two-stage feature association according to claim 4 is characterized in that: Step S3 includes the following sub-steps: S31: Obtain the block geometry point feature vectors at the current time t and the previous time t-1 respectively and The robot body coordinate system b at time t t As the reference coordinate system, the robot posture transformation matrix T obtained by IMU is init Perform spatial transformation to obtain the block geometric anchor vector of the same space-time dimension and S32: Use DBSCAN algorithm to find the geometric anchor vectors of blocks in the same space-time dimension and Perform cluster analysis to obtain the point cloud space area block set at time t and time t-1 Then, the spatial characteristics, texture features and semantic information of each point cloud block are comprehensively considered to calculate the commonality between the point cloud blocks at two adjacent moments; S33: According to the calculated spatial commonality of the point cloud blocks, the one with the highest similarity in each area is selected as the candidate best matching pair, which is reprojected into the 2D pixel coordinate space of the image, and the problem of 2D image semantic area block matching is converted into 3D space block matching, and finally the matching relationship of the 2D image area is obtained.
7. The robot positioning method based on two-stage feature association according to claim 6 is characterized in that: In step S32, the formula for calculating the commonality between point cloud blocks at two adjacent moments is expressed as: in, Represents a set of point cloud blocks at two adjacent moments The similarity between the i-th and j-th point cloud blocks in , N and M are the number of point cloud blocks in the point cloud data at time t and time t-1 respectively, Represents a collection of point cloud blocks The i-th and j-th point cloud blocks in Represents a set of point cloud blocks at two times The geometric anchor feature vectors of the i-th and j-th point cloud blocks in θ ij represents the horizontal orientation error between the i-th point cloud block and the j-th point cloud block, Represented as point cloud blocks at two times and The semantic similarity function of Represented as a spatial block The average eigenvector of the internal point cloud atlas, Represented as a spatial block The average eigenvector of the internal point cloud atlas, ε is a constant greater than 1.
8. The robot positioning method based on two-stage feature association according to claim 6, characterized in that: Step S4 includes the following sub-steps: S41: Gaussian blur is performed on the associated image regions respectively, ORB feature points are extracted, Steer BRIEF feature descriptors are calculated, and preliminary feature point matching pairs are obtained by calculating the Hamming distance between the descriptors; S42: Based on the grayscale invariance, taking the initially matched feature point as the center and considering the information of the local neighborhood, a square area of a preset size is taken, and the optical flow vector is calculated using the optical flow equation. The optical flow equation is solved by the least squares method to obtain the motion speed V of the feature point in the u and v axis directions. u and V v , and then calculate the average predicted optical flow of pixels in the semantic area Then calculate the difference between the measured optical flow and the average predicted optical flow of each feature point from the frame at time t-1 to the frame at time t S43: Based on the rigid body structure invariance, that is, the corresponding optical flow changes of feature points in the same rigid body region are consistent, the regional average optical flow is used as the screening threshold to construct a matching affinity factor; S44: improving the feature point matcher based on the matching affinity factor, calculating the matching accuracy between the feature points, and obtaining the best matching point pair according to the matching degree value M between the point pairs; S45: The camera pose solving problem is transformed into a nonlinear optimization problem for solving, and the initial optimized robot pose is obtained.
9. The robot positioning method based on two-stage feature association according to claim 8, characterized in that: In step S42, the average predicted optical flow of pixels in the semantic region The calculation method is: in, It is represented as the average predicted optical flow, n is the total number of feature points in the region, V u,k and V v,k Respectively represent the movement speed value of the kth feature point in the u and v axis directions, and Respectively represent the average movement speed of the semantic area in the u and v axis directions; The difference between the measured optical flow and the average predicted optical flow of each feature point from the frame at time t-1 to the frame at time t The calculation formula is: Among them, V i Represented as feature points The measured optical flow is obtained by calculating the pixel change of the feature point by the feature matching pair obtained in step S41. is the i-th feature point of the frame at time t-1; In step S43, the matching affinity factor is defined as: Where d represents the optical flow threshold of the feature point, and δ(i,j) is the matching affinity factor between the i-th and j-th feature points; In step S44, the calculation formula for the best matching point pair is: Among them, M i,j Represented as two feature points and The matching accuracy between represents the i-th feature point in region a of the image at time t-1, represents the jth feature point in region b of the image at time t, f i and f j Represented as two feature points and The binary descriptor of n is the dimension of the descriptor, D f (·) represents the Hamming distance calculation function between two binary descriptors; In step S45, the optimization problem is expressed as: Among them, R * ,t * are the robot rotation transformation matrix and translation transformation matrix to be solved, u i It is expressed as the coordinates of the projection point observed by the camera at time t, P i Represented as the projection point u at time t i The coordinates of the corresponding space point, s i It is expressed as the depth value of the 3D point observed by the camera at time t, T is the robot pose to be optimized, and K is the intrinsic parameter matrix of the camera.
10. The robot positioning method based on two-stage feature association according to claim 8, characterized in that: In step S5, after obtaining the initially optimized robot posture, a joint iterative optimization is performed to first update the robot posture T in step S3. init The value of is used to recalculate the similarity between the spatial blocks, solve and update the matching relationship of the spatial area blocks, and then execute S3 and S4 processes again to solve the robot's posture. The S3 and S4 processes are iterated several times, and the transformation matrix with the smallest reprojection error is taken as the optimal estimate of the robot's posture.
Citation Information
Cited By
Robot navigation method integrated with multi-modal perception and robot control equipment
CN120645215A
A robot navigation method and robot control device fusing multi-modal perception
CN120645215B