Deep Learning-Based Interaction Method and System
By adopting a deep learning-based interactive method in the autonomous driving system, combining Gaussian hybrid model and feature weight distribution function to optimize feature point matching, and using the compensation depth map to calibrate the pose parameters, the problem of pose estimation error accumulation in complex environments is solved, and higher matching stability and reconstruction accuracy are achieved.
Patent Information
- Application Number
- CN202510362969.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-26
AI Technical Summary
In complex environments such as autonomous driving, the vision-based pose estimation is affected by factors such as lighting changes, dynamic occlusion, motion blur and scene changes, resulting in unstable matching of feature points, accumulated errors in pose estimation, and deep learning methods have high computational complexity and insufficient generalization capabilities, making it difficult to effectively combine feature change information and depth information for error compensation.
A deep learning-based interaction method is adopted, by collecting continuous image sequences of on-board cameras, extracting feature points and constructing two-dimensional feature vectors, using Gaussian mixed model clustering to obtain stable feature classes, calculating feature weight distribution function to empower feature points, calculating feature similarity based on weighted feature points, filtering reliable matching feature pairs, and constructing optimization equations to solve pose transformation parameters. At the same time, a compensation depth map is generated by combining the feature change sequence with depth information, the scene is reconstructed using the compensation depth map and pose transformation parameters, and the pose parameters are calibrated according to the scene confidence score.
It improves the stability of feature matching and the accuracy of position estimation, enhances the quality of scene reconstruction and the adaptability of the system in complex environments, and is suitable for applications such as autonomous driving, intelligent transportation and high-precision map construction.
Smart Images

Figure CN119919749B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision technology, and particularly to an interaction method and system based on deep learning. Background Art
[0002] In applications such as autonomous driving, intelligent transportation, and robot navigation, vision-based pose estimation is one of the key technologies. Traditional methods usually rely on feature point detection and matching, combined with geometric constraints to solve pose transformation parameters. However, in the actual environment, factors such as light changes, dynamic occlusion, motion blur, and scene changes will affect the stability of feature point matching, resulting in the accumulation of pose estimation errors. In addition, traditional methods use a fixed matching threshold and cannot adapt to the changes of vehicles at different speeds, perspectives, and scenes, reducing the reliability of pose calculation.
[0003] The development of deep learning in the field of computer vision provides new ideas for feature extraction, matching optimization, and pose estimation. Although there are already deep learning methods applied to feature detection and matching, there are generally problems of high computational complexity and insufficient generalization ability. At the same time, the potential of depth information in pose estimation has not been fully utilized, and existing methods are difficult to effectively combine feature change information and depth information for error compensation, resulting in limited scene reconstruction accuracy. Therefore, there is an urgent need for an interaction technology that combines deep learning and traditional vision methods to improve feature matching stability, pose estimation accuracy, and scene reconstruction quality to meet the application requirements in complex environments such as autonomous driving. Summary of the Invention
[0004] Embodiments of the present invention provide an interaction method and system based on deep learning, which can solve the problems in the prior art.
[0005] In the first aspect of the embodiments of the present invention,
[0006] Provide an interaction method based on deep learning, including:
[0007] Collect a continuous image sequence obtained by an in-vehicle camera and extract feature points, combine the position change amount of feature points at adjacent moments and the gradient change amount of feature points at adjacent moments to construct a two-dimensional feature vector, use a Gaussian mixture model to cluster the two-dimensional feature vector to obtain a stable feature class, calculate a feature weight distribution function based on the distribution characteristics of feature vectors in the stable feature class, and assign weights to feature points according to the weight distribution function to obtain weighted feature points;
[0008] Calculate the feature similarity based on weighted feature points to obtain a similarity matrix, normalize the similarity matrix, construct a feature matching relationship using the normalized similarity matrix to obtain initial matching pairs, obtain the vehicle motion speed and calculate the dynamic parallax threshold, screen the parallax of the feature points in the initial matching pairs according to the dynamic parallax threshold to obtain reliable matching feature pairs, construct an optimization equation based on the reliable matching feature pairs and solve it to obtain the pose transformation parameters between adjacent images;
[0009] Input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate the feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with depth information to generate a compensated depth map, reconstruct the scene using the compensated depth map and pose transformation parameters and calculate the scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated sequence of pose parameters.
[0010] In an alternative embodiment,
[0011] Use a Gaussian mixture model to cluster two-dimensional feature vectors to obtain stable feature classes, calculate a feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature classes, and assign weights to the feature points according to the weight distribution function to obtain weighted feature points, including:
[0012] Use a Gaussian mixture model to calculate the mixture weights, mean vectors, and covariance matrices of the mixture Gaussian distribution for two-dimensional feature vectors, and cluster the two-dimensional feature vectors using the mixture weights, mean vectors, and covariance matrices to obtain multiple feature classes;
[0013] Calculate the average distance from the intra-class feature vectors of the feature classes to the class center to obtain the intra-class aggregation degree, calculate the minimum distance between the class centers of the feature classes to obtain the inter-class separation degree, screen the feature classes according to the ratio of the intra-class aggregation degree to the inter-class separation degree, and determine the feature classes that meet the preset feature threshold conditions as stable feature classes;
[0014] Calculate the intra-class covariance matrix of the two-dimensional feature vectors in the stable feature classes and perform eigenvalue decomposition to obtain the spatial distribution characteristics, calculate the Euclidean distance between the two-dimensional feature vectors at adjacent times and convert it into a temporal correlation score, and perform a weighted combination of the spatial distribution characteristics and the temporal correlation score to obtain a feature weight distribution function;
[0015] Use the feature weight distribution function to calculate the Mahalanobis distance of the two-dimensional feature vectors to obtain the spatial weight, combine the temporal correlation score to obtain the temporal weight, fuse the spatial weight and the temporal weight and normalize them within a preset neighborhood range, and assign the normalized weight values to the corresponding feature points to obtain weighted feature points.
[0016] In an alternative embodiment,
[0017] Calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function to obtain the spatial weight, combining the temporal correlation score to obtain the temporal weight, fusing the spatial weight and the temporal weight and normalizing within a preset neighborhood range, and assigning the normalized weight value to the corresponding feature point to obtain the weighted feature point includes:
[0018] Calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function, and obtaining the spatial weight through exponential mapping with the Mahalanobis distance;
[0019] Calculating the curvature of the motion trajectory of the two-dimensional feature vector as the time scale coefficient, non-linearly combining the time scale coefficient with the Euclidean distance of the two-dimensional feature vector at adjacent moments to obtain the temporal correlation score, and calculating the temporal weight based on the temporal correlation score;
[0020] Constructing the local variance distribution of the spatial weight, determining the adaptive kernel function based on the local variance distribution, and fusing the spatial weight and the temporal weight using the adaptive kernel function to obtain the fused weight;
[0021] Calculating the spatial distribution gradient of the fused weight, constructing the adaptive neighborhood radius based on the spatial distribution gradient, and obtaining the dynamic neighborhood range by density-weighted combination of the adaptive neighborhood radius and the preset neighborhood range;
[0022] Calculating the local covariance feature of the fused weight within the dynamic neighborhood range, constructing the anisotropic Gaussian kernel function, and spatially modulating the fused weight using the anisotropic Gaussian kernel function; normalizing the modulated fused weight within the dynamic neighborhood range to obtain the initial weight value, calculating the weight consistency constraint based on the initial weight value, and iteratively optimizing the weight consistency constraint and the initial weight value to obtain the final weight value, and assigning the final weight value to the corresponding two-dimensional feature vector to obtain the weighted feature point.
[0023] In an alternative embodiment,
[0024] Obtaining the vehicle motion speed and calculating the dynamic parallax threshold, screening the parallax of the feature points in the initial matching pairs according to the dynamic parallax threshold to obtain the reliable matching feature pairs, and constructing and solving the optimization equation according to the reliable matching feature pairs to obtain the pose transformation parameters between adjacent images includes:
[0025] Obtaining the vehicle motion speed and the adjacent frame time interval, statistically calculating the image entropy value as the scene structure complexity based on the feature point distribution in the image area, and calculating the dynamic parallax threshold according to the vehicle motion speed, the adjacent frame time interval and the scene structure complexity;
[0026] Calculate the image coordinate difference of the initial matching feature point pairs in adjacent image frames to obtain the feature point parallax. Take the initial matching feature point pairs with the feature point parallax less than the dynamic parallax threshold as candidate matching feature point pairs, and calculate the reference transformation matrix according to the image coordinate correspondence of the candidate matching feature point pairs;
[0027] Count other candidate matching feature point pairs within the neighborhood range of the candidate matching feature point pairs, calculate the difference score between the coordinate transformation result of other candidate matching feature point pairs and the reference transformation matrix, calculate the consistency support degree according to the difference score, and take the candidate matching feature point pairs with the consistency support degree greater than the preset support degree as reliable matching feature point pairs;
[0028] Calculate the local transformation matrix according to the image coordinate correspondence of the reliable matching feature point pairs, substitute the difference value between the local transformation matrix and the reference transformation matrix into the Gaussian kernel function to obtain the weight coefficient of the reliable matching feature point pairs, and construct a weighted reprojection error optimization equation by using the corresponding relationship between the three-dimensional space coordinates and the two-dimensional image coordinates of the reliable matching feature point pairs and the weight coefficient;
[0029] Construct a Jacobian matrix according to the reprojection error optimization equation, construct a weight matrix by using the weight coefficient, construct an increment equation according to the Jacobian matrix and the weight matrix, iteratively calculate the pose update amount of the increment equation and update the pose parameters, and obtain the pose transformation parameters between adjacent image frames when the pose update amount is less than the preset update amount threshold.
[0030] In an alternative embodiment,
[0031] Obtain the vehicle motion speed and the adjacent frame time interval, count the feature point distribution in the image area and calculate the image entropy value as the scene structure complexity. Calculating the dynamic parallax threshold according to the vehicle motion speed, the adjacent frame time interval and the scene structure complexity includes:
[0032] Obtain the vehicle motion speed and the adjacent frame time interval, calculate the acceleration of the vehicle motion, classify the vehicle motion state according to the acceleration, calculate the product of the vehicle motion speed and the acceleration, and obtain the motion curvature as the ratio of the product to the cube of the vehicle motion speed;
[0033] Construct a multi-scale pyramid for the image, divide the image at each scale into adaptive grid cells, count the number of feature points in each grid cell, and calculate the ratio of the number of feature points in each grid cell to the total number of image feature points to obtain the feature point distribution probability;
[0034] Calculate the information entropy of each scale based on the distribution probability of the feature points, calculate the average value of the image gradient directions within each grid cell, calculate the direction consistency based on the average value of the image gradient directions, and combine the information entropy of each scale with the direction consistency as the scene structure complexity of the current scale;
[0035] Determine the weight coefficient of each scale according to the difference between the scene structure complexity and the average value of the scene structure complexity, and perform cross-scale weighting on the scene structure complexity to obtain the comprehensive scene structure complexity;
[0036] Calculate the motion response value according to the vehicle motion speed and the motion curvature, and combine the comprehensive scene structure complexity with the motion response value to obtain the dynamic parallax threshold.
[0037] In an alternative embodiment,
[0038] Combining the motion feature mapping matrix with the depth information to generate a compensated depth map, and reconstructing the scene and calculating the scene confidence score using the compensated depth map and the pose transformation parameters includes:
[0039] Combine the motion feature mapping matrix with the depth information, perform propagation compensation on the depth information according to the motion information in the motion feature mapping matrix to obtain a compensated depth map, and perform three-dimensional back-projection on the compensated depth map to obtain a scene point cloud;
[0040] Extract plane structure elements and edge structure elements from the scene point cloud, calculate the spatial distribution characteristics of the plane structure elements and edge structure elements to obtain a structure distribution map, and construct a topological relationship map of the structure elements based on the structure distribution map;
[0041] Decompose the pose transformation parameters into a rotation component and a translation component, obtain the auxiliary pose information of the visual odometer, calculate the fusion weight of the pose transformation parameters and the auxiliary pose information according to the structure distribution map, and perform weighted fusion on the pose transformation parameters and the auxiliary pose information according to the fusion weight to obtain the fused pose parameters;
[0042] Perform coordinate transformation on the scene point cloud according to the fused pose parameters to obtain a reconstructed scene, calculate the local geometric consistency of the point cloud in the reconstructed scene to obtain the local confidence, calculate the structural similarity of the point cloud at adjacent times in the reconstructed scene to obtain the global confidence, and perform weighted combination on the local confidence and the global confidence to obtain the scene confidence score;
[0043] Calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated sequence of pose parameters.
[0044] In an alternative embodiment,
[0045] Calibrating the pose transformation parameters based on the scene confidence score to obtain a calibrated sequence of pose parameters, including:
[0046] Establish an evaluation vector for each pose according to the scene confidence score, compare the evaluation vector with a preset vector threshold to obtain a pose credibility index, and mark abnormal poses among the poses where the pose credibility index is lower than the preset vector threshold;
[0047] Construct an influence propagation matrix of abnormal poses using the topological relationship graph, analyze the structural association strength between abnormal poses and adjacent poses in the influence propagation matrix, determine the spatio-temporal influence range of abnormal poses according to the structural association strength, and use the spatio-temporal influence range as the calibration range;
[0048] Construct a pose optimization objective function, which includes a structure preservation term, a depth consistency term, and a temporal smoothing term. Iteratively optimize the abnormal poses. In each iteration, calculate the structure preservation score of the current pose based on the structure distribution map, calculate the depth consistency score based on the compensated depth map, calculate the temporal smoothing score based on adjacent poses, and substitute the structure preservation score, depth consistency score, and temporal smoothing score into the pose optimization objective function to calculate the optimization direction and step size;
[0049] After each round of iterative optimization, reapply the optimized calibrated pose to scene reconstruction, extract the structural features in the scene reconstruction to calculate the structural consistency, and stop the optimization when the structural consistency is greater than the preset convergence threshold;
[0050] Update the original pose sequence according to the finally optimized calibrated pose, and sort it in chronological order to obtain a calibrated sequence of pose parameters.
[0051] In the second aspect of the embodiments of the present invention,
[0052] Provide an interaction system based on deep learning, including:
[0053] A first unit for collecting a continuous image sequence obtained by an in-vehicle camera and extracting feature points, combining the position change amount of feature points at adjacent times with the gradient change amount of feature points at adjacent times to construct a two-dimensional feature vector, clustering the two-dimensional feature vector using a Gaussian mixture model to obtain stable feature classes, calculating a feature weight distribution function based on the distribution characteristics of feature vectors in the stable feature classes, and weighting the feature points according to the weight distribution function to obtain weighted feature points;
[0054] A second unit, configured to calculate a feature similarity based on weighted feature points to obtain a similarity matrix, normalize the similarity matrix, construct a feature matching relationship by using the normalized similarity matrix to obtain an initial matching pair, acquire a vehicle motion speed and calculate a dynamic parallax threshold, filter the parallax of the feature points in the initial matching pair according to the dynamic parallax threshold to obtain a reliable matching feature pair, construct an optimization equation according to the reliable matching feature pair and solve it to obtain pose transformation parameters between adjacent images;
[0055] A third unit, configured to input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate a feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with depth information to generate a compensated depth map, reconstruct a scene by using the compensated depth map and the pose transformation parameters and calculate a scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0056] In a third aspect of the embodiments of the present invention,
[0057] there is provided an electronic device, including:
[0058] a processor;
[0059] a memory for storing instructions executable by the processor;
[0060] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0061] In a fourth aspect of the embodiments of the present invention,
[0062] there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0063] In this embodiment, through the integration of deep learning and computer vision technologies, automatic calibration is achieved, improving the accuracy and robustness of pose estimation. The Gaussian mixture model is used to cluster feature points, screen stable feature classes, and optimize feature point matching based on the feature weight distribution, improving the matching accuracy and reducing the impact of environmental changes on feature matching. A dynamic parallax threshold calculation method is introduced, enabling the feature matching screening to adapt to the vehicle's motion state, effectively reducing false matches, and improving the calculation accuracy of pose transformation parameters. By combining deep learning to extract feature change information, constructing a motion feature mapping matrix, and fusing depth information to generate a compensated depth map, adaptive compensation for feature errors is achieved, enhancing the quality of scene reconstruction. The pose parameters are calibrated based on the scene confidence score, reducing cumulative errors and improving the stability of long-term operation. The present invention not only improves the accuracy of visual pose estimation but also enhances the adaptability of the system in complex environments, and is applicable to applications such as autonomous driving, intelligent transportation, and high-precision map construction, providing reliable technical support for high-precision positioning and environmental perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic flowchart of the interaction method based on deep learning according to an embodiment of the present invention;
[0065] Figure 2 is a relationship diagram between the dynamic parallax threshold and the scene complexity according to an embodiment of the present invention;
[0066] Figure 3 is a comparison diagram of 3D reconstruction accuracy under changes in scene complexity according to an embodiment of the present invention;
[0067] Figure 4 is a schematic structural diagram of the interaction system based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0070] Figure 1 is a schematic flowchart of the interaction method based on deep learning according to an embodiment of the present invention, as shown in Figure 1As shown, the method includes:
[0071] Collect a continuous image sequence obtained by an in-vehicle camera and extract feature points. Combine the position change amount of feature points at adjacent times and the gradient change amount of feature points at adjacent times to construct a two-dimensional feature vector. Use a Gaussian mixture model to cluster the two-dimensional feature vector to obtain stable feature classes. Calculate a feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature classes. Assign weights to the feature points according to the weight distribution function to obtain weighted feature points;
[0072] Calculate feature similarity based on the weighted feature points to obtain a similarity matrix. Perform normalization processing on the similarity matrix. Use the normalized similarity matrix to construct a feature matching relationship to obtain initial matching pairs. Obtain the vehicle motion speed and calculate a dynamic parallax threshold. Screen the parallax of the feature points in the initial matching pairs according to the dynamic parallax threshold to obtain reliable matching feature pairs. Construct an optimization equation based on the reliable matching feature pairs and solve it to obtain the pose transformation parameters between adjacent images;
[0073] Input the continuous image sequence into a feature detection network to obtain a feature change sequence. Calculate the feature change gradient between adjacent frames according to the feature change sequence. Construct a motion feature mapping matrix based on the feature change gradient. Combine the motion feature mapping matrix with depth information to generate a compensated depth map. Use the compensated depth map and the pose transformation parameters to reconstruct the scene and calculate the scene confidence score. Calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0074] Among them, a feature point refers to a key point in an image, usually having a strong local contrast, such as a corner point, an edge point, or a texture feature point. In this solution, feature points are used to calculate the position change amount and gradient change amount at adjacent times to construct subsequent feature vectors. A Gaussian mixture model (GMM) is a probability statistical model used to represent the mixture of multiple Gaussian distributions. In this solution, GMM is used to perform clustering analysis on the two-dimensional feature vector to distinguish stable feature classes from unstable feature classes, thereby improving the reliability of matching. A stable feature class refers to a set of feature points classified as stable after GMM clustering. These feature points have a low change rate between adjacent frames and can be used more reliably for pose estimation and matching optimization. The feature weight distribution function is a function calculated based on the distribution characteristics of the feature vectors in the stable feature classes. This function is used to evaluate the importance of feature points and assign weights to feature points during the feature matching process to improve the accuracy of matching. A weighted feature point refers to a feature point assigned a weight according to the feature weight distribution function. During the matching process, weighted feature points can enhance the influence of stable feature points, thereby improving the stability of pose estimation.
[0075] The dynamic parallax threshold is a threshold calculated based on the vehicle's motion speed and is used to screen feature matching pairs. At different speeds, the parallax range is different. The dynamic parallax threshold can adaptively adjust the matching strategy to improve the reliability of the matching pairs. The feature detection network is a neural network based on deep learning and is used to extract the feature change sequence from the image sequence. This network can adaptively learn the change law of feature points and improve the robustness of feature extraction. The feature change gradient refers to the change rate of feature points between adjacent frames and is used to measure the dynamics of feature points. In this solution, this gradient information is used to construct the motion feature mapping matrix.
[0076] The compensated depth map is generated by combining the motion feature mapping matrix and depth information and is used to correct the errors of the traditional depth map. In this solution, this compensated depth map can improve the accuracy of scene reconstruction and reduce the cumulative deviation caused by visual errors. The scene confidence score is used to evaluate the reliability of the reconstructed scene and is usually calculated based on factors such as depth information consistency and feature matching quality. In this solution, the scene confidence score is used to calibrate the pose parameters to improve the accuracy of pose estimation.
[0077] The pose transformation parameters are the parameters used to describe the camera motion between adjacent image frames and include the rotation matrix and translation vector. In this solution, these parameters are obtained through optimization and adjustment during the final calibration process to improve the positioning accuracy. The calibrated pose parameter sequence is a set of pose parameters calibrated by the scene confidence score and can more accurately describe the motion state of the camera in the continuous image sequence. This solution reduces the cumulative error through this calibration process and improves the stability of long-term pose estimation.
[0078] In summary, the present invention improves the robustness of feature matching, enhances the accuracy of pose estimation, and strengthens the reliability of scene reconstruction by introducing technologies such as adaptive parallax threshold calculation and depth compensation, and can achieve high-precision automatic calibration in complex environments, providing strong technical support for applications such as autonomous driving, intelligent transportation, and high-precision map construction.
[0079] In an alternative embodiment,
[0080] Using a Gaussian mixture model to cluster two-dimensional feature vectors to obtain stable feature classes, calculating a feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature classes, and weighting the feature points according to the weight distribution function, including:
[0081] Calculating the mixing weights, mean vectors, and covariance matrices of the mixture Gaussian distribution for the two-dimensional feature vectors using a Gaussian mixture model, and clustering the two-dimensional feature vectors using the mixing weights, mean vectors, and covariance matrices to obtain multiple feature classes;
[0082] The average distance from the intra-class feature vectors of the feature class to the class center is calculated to obtain the intra-class aggregation degree, the minimum distance between the class centers of the feature classes is calculated to obtain the inter-class separation degree, and the feature classes are screened according to the ratio of the intra-class aggregation degree to the inter-class separation degree, and the feature classes that meet the preset feature threshold conditions are determined as stable feature classes;
[0083] Calculate the intra-class covariance matrix of the two-dimensional feature vectors in the stable feature class and perform eigenvalue decomposition to obtain the spatial distribution features, calculate the Euclidean distance of the two-dimensional feature vectors at adjacent times and convert it into a temporal correlation score, and perform weighted combination of the spatial distribution features and the temporal correlation score to obtain a feature weight distribution function;
[0084] Use the feature weight distribution function to calculate the Mahalanobis distance of the two-dimensional feature vectors to obtain the spatial weight, combine the temporal correlation score to obtain the temporal weight, fuse the spatial weight and the temporal weight and normalize them within a preset neighborhood range, and assign the normalized weight value to the corresponding feature point to obtain weighted feature points.
[0085] Specifically, first, an image sequence continuously captured by an in-vehicle camera is collected, and key feature points (such as corner points and edge intersection points) are extracted from it. For the position change amount (displacement vector) and gradient change amount (change in pixel gradient direction and amplitude) of the same feature point in two adjacent frames of images, these two amounts are combined into a two-dimensional feature vector. The Gaussian mixture model is used to model all the feature vectors, and the parameters of the Gaussian mixture distribution are estimated through the expectation-maximization algorithm, including the mixing weight of each Gaussian component (representing the proportion of this component in the overall distribution), the mean vector (the central position of the component), and the covariance matrix (the shape and direction of the component). According to these parameters, the feature vectors are divided into multiple feature classes, and each class corresponds to a Gaussian component.
[0086] During the vehicle driving process, the in-vehicle camera captures the corner point features of the guardrail on the front road. Due to vehicle movement, the position of the guardrail corner point in the image undergoes translation (position change amount), and at the same time, the local gradient direction changes due to the change in the viewing angle (gradient change amount). After combining these change amounts into a two-dimensional vector, the Gaussian mixture model can distinguish the feature class belonging to the static guardrail from the feature class of the dynamic vehicle.
[0087] Traditional methods (such as K-means clustering) only rely on the feature space distance for partitioning and cannot model complex distributions. This solution captures the multi-modal distribution characteristics of the feature vectors through the Gaussian mixture model and more accurately distinguishes the feature classes in different motion states.
[0088] For each feature class, calculate its within-class aggregation degree (the average Euclidean distance from all feature vectors within the class to the class center) and between-class separation degree (the minimum distance from the current class center to other class centers). If the ratio of the within-class aggregation degree to the between-class separation degree of a certain feature class is lower than a preset threshold (such as 0.3), it is considered that the feature vectors within this class are closely distributed and significantly separated from other classes, and it is marked as a stable feature class. Subsequently, calculate the within-class covariance matrix for the feature vectors within the stable feature class, decompose its principal component directions (feature vectors) and variance ratios (eigenvalues) to obtain the spatial distribution characteristics (such as whether the distribution direction is consistent with the vehicle movement direction). At the same time, calculate the Euclidean distance of the same feature point in adjacent frames, and map it to a temporal correlation score ranging from 0 to 1 (the smaller the distance, the higher the score). Weight the spatial distribution characteristics (principal component direction consistency) and the temporal correlation score according to a preset ratio to generate a feature weight distribution function.
[0089] The change amount of the feature point positions of static road signs is small and the direction is consistent (along the vehicle movement direction), and its gradient change amount is stable (no obvious distortion of the texture). Therefore, the spatial distribution characteristics are manifested as the main direction of the covariance matrix being aligned with the vehicle movement direction, and the temporal correlation score is high. For dynamic pedestrians, the covariance direction is chaotic and the temporal score is low due to sudden position changes. Traditional fixed weight assignment methods do not combine the spatial distribution and temporal continuity of features. This solution dynamically adjusts the feature point weights by quantifying the consistency in the spatial and temporal dimensions to suppress the influence of dynamic interference features.
[0090] Further, based on the feature weight distribution function, calculate the Mahalanobis distance of each feature vector (measuring its degree of deviation from the class center, considering the anisotropy of the covariance matrix), and map it to a spatial weight through an exponential function (the smaller the deviation, the higher the weight). At the same time, directly use the temporal correlation score as the temporal weight. Fuse the spatial weight and the temporal weight according to an adaptive ratio (such as a higher proportion of the temporal weight in dynamic scenarios) to obtain a preliminary fused weight. Further, statistically analyze the distribution of the fused weight within the image neighborhood of the feature point (such as a 5×5 pixel range), eliminate the local density difference through normalization processing, and finally assign the normalized weight to the corresponding feature point to generate a set of weighted feature points.
[0091] In the vehicle turning scenario, the position change direction of the static building feature points is consistent with the vehicle movement trajectory, the Mahalanobis distance is small (high spatial weight), and the displacement between adjacent frames is continuous (high temporal score). Therefore, the fused weight is large; while the false feature points caused by road surface reflection have a large Mahalanobis distance and a low temporal score due to position jumps, and the weight is significantly suppressed.
[0092] This technical solution clusters the feature vectors through a Gaussian mixture model, weights the feature points by combining spatial and temporal information, and achieves more accurate and stable feature point matching. Compared with traditional clustering methods based on Euclidean distance (such as K-means), using a Gaussian mixture model can better capture the multimodal distribution of feature vectors, and then accurately distinguish different motion states, such as static and dynamic targets. By evaluating the within-class aggregation degree and between-class separation degree of feature classes, stable feature classes are screened out, effectively removing noise and unstable features, thereby improving the accuracy and robustness of feature matching. Through the weighted combination of spatial distribution features and temporal correlation scores, the calculated feature weight distribution function can dynamically adjust the weight of each feature point, enabling static feature points to obtain higher weights in matching, while feature points with large dynamic changes are effectively suppressed. Traditional methods usually use fixed weights and fail to fully consider the temporal continuity and spatial consistency of features, resulting in being easily interfered with in dynamic scenarios. However, this solution comprehensively considers spatial and temporal information, making the matching effect of feature points more accurate in complex and dynamic environments. It can improve the matching stability of feature points in dynamic and complex environments, especially in the case of perspective changes or dynamic interferences (such as pedestrians, vehicles) in the scene. Through refined feature point weighting and dynamic adjustment strategies, the matching accuracy of feature points in dynamic scenarios is finally significantly improved, reducing the risk of mis-matching, thereby enhancing the adaptability and robustness of the system in real complex environments.
[0093] In an alternative embodiment,
[0094] Calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function to obtain a spatial weight, combining the temporal correlation score to obtain a temporal weight, fusing the spatial weight and the temporal weight and normalizing within a preset neighborhood range, and assigning the normalized weight value to the corresponding feature point to obtain a weighted feature point includes:
[0095] Calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function, and obtaining the spatial weight through exponential mapping with the Mahalanobis distance;
[0096] Calculating the curvature of the motion trajectory of the two-dimensional feature vector as a time scale coefficient, non-linearly combining the time scale coefficient with the Euclidean distance of the two-dimensional feature vector at adjacent moments to obtain a temporal correlation score, and calculating the temporal weight based on the temporal correlation score;
[0097] Constructing the local variance distribution of the spatial weight, determining an adaptive kernel function based on the local variance distribution, and using the adaptive kernel function to fuse the spatial weight and the temporal weight to obtain a fused weight;
[0098] Calculating the spatial distribution gradient of the fusion weight, constructing an adaptive neighborhood radius based on the spatial distribution gradient, and combining the adaptive neighborhood radius with a preset neighborhood range through density weighting to obtain a dynamic neighborhood range;
[0099] Calculate the local covariance characteristics of the fusion weight within the dynamic neighborhood, construct an anisotropic Gaussian kernel function, and use the anisotropic Gaussian kernel function to spatially modulate the fusion weight; normalize the modulated fusion weight within the dynamic neighborhood to obtain an initial weight value, calculate the weight consistency constraint based on the initial weight value, iteratively optimize the weight consistency constraint and the initial weight value to obtain a final weight value, and assign the final weight value to the corresponding two-dimensional feature vector to obtain a weighted feature point.
[0100] This embodiment uses the feature weight distribution function to calculate the Mahalanobis distance for each two-dimensional feature vector. The Mahalanobis distance is a distance metric that measures the deviation between the feature vector and the average position (class center) of the class to which it belongs, and it takes into account the covariance structure of the feature space. By performing exponential mapping on the Mahalanobis distance, the spatial distance of the feature point can be converted into a spatial weight, indicating the importance of the feature point in space. The larger the spatial weight, the smaller the distance of the feature point from its class center, and the more stable its distribution in space. By calculating the curvature of the motion trajectory of each two-dimensional feature vector as the time scale coefficient, combined with the Euclidean distance between the feature points at adjacent moments, the temporal correlation score is obtained. Curvature is a measure of the degree of curvature of the motion trajectory, reflecting the law of motion change of the feature point. After the time scale coefficient is nonlinearly combined with the Euclidean distance, the temporal correlation score is obtained, which reflects the continuity of the feature point over time. The higher the score, the more stable and continuous the motion of the feature point. The temporal weight is calculated based on the temporal correlation score. The temporal weight is used to reflect the importance of the feature point in the time dimension. The feature point with a more stable temporal change will obtain a higher temporal weight.
[0101] Construct the local variance distribution of spatial weights. The local variance is a measure of the distribution change of feature points within their neighborhood, reflecting the stability and consistency of feature points in the local area. Through the local variance distribution, an adaptive kernel function can be determined, which is used to fuse spatial weights and temporal weights. The adaptive kernel function adjusts the degree of weight fusion according to the distribution of the local variance, so that in some regions, the fusion of spatial weights and temporal weights will be more refined, while in other regions, different fusion strategies may be adopted. Then, calculate the spatial distribution gradient of the fusion weights. The spatial distribution gradient is used to measure the change trend of the fusion weights in space, thus reflecting the spatial structure and change of feature points. Based on the spatial distribution gradient, construct an adaptive neighborhood radius. The adaptive neighborhood radius is dynamically adjusted according to the change of the local spatial distribution, and can more accurately determine the neighborhood range of feature points. Through the density weighted combination with the preset neighborhood range, a dynamic neighborhood range is obtained, which can more flexibly adapt to the distribution characteristics of feature points in different scenarios.
[0102] Calculate the local covariance features of the fusion weights within the dynamic neighborhood range. The local covariance reflects the change pattern of feature points within their neighborhood. By calculating the local covariance, the stability and change of feature points can be further evaluated. Then, based on the local covariance features, construct an anisotropic Gaussian kernel function. The anisotropic Gaussian kernel function helps to more accurately capture the distribution characteristics of feature points in different spatial directions by adjusting the expansion degree in different directions.
[0103] Use the anisotropic Gaussian kernel function to perform spatial modulation on the fusion weights. This process adjusts the weight values in different directions, making the spatial distribution more uniform and more in line with the distribution pattern in the actual scenario. After modulation, the fusion weights will be normalized within the dynamic neighborhood range to eliminate the influence caused by spatial density differences, ensuring that the weight values of each feature point can be reasonably adjusted within its neighborhood. Finally, calculate the weight consistency constraint based on the normalized fusion weights. The weight consistency constraint is used to ensure the consistency of feature point weights throughout the process, avoiding weight instability caused by external interference or abnormal changes of feature points. Through the iterative optimization process, combine the weight consistency constraint with the initial weight values to obtain the final weight values. The final weight values are assigned to the corresponding two-dimensional feature vectors, thus obtaining weighted feature points, ensuring the spatial and temporal stability of the weighted results of feature points.
[0104] In the autonomous driving scenario, in-vehicle cameras capture landmark feature points on the road edge (such as building corners). Due to the movement of the camera and the vehicle, the image positions of these feature points change. At the same time, due to different viewing angles, the local gradient directions of the feature points also change. By combining the change amounts of the feature point positions and the gradient changes in adjacent frames into a two-dimensional feature vector and then performing clustering through a Gaussian mixture model, it is possible to accurately distinguish static buildings from false feature points on the dynamic road surface. Further, based on the fusion of spatial weights and temporal weights, the feature points of static buildings will obtain higher weighted values, while dynamic false feature points caused by reflection and other reasons will be effectively suppressed. Finally, based on these weighted feature points, more accurate image feature matching and scene reconstruction results can be obtained.
[0105] This technical solution combines spatial weights and temporal weights, and uses Mahalanobis distance and temporal correlation scores to weight the feature points, significantly improving the accuracy and robustness of the feature points. Existing technologies usually rely on traditional clustering methods to process feature points, but these methods cannot effectively model complex distribution features, resulting in feature points with dynamic changes or interference may be misidentified as valid features. This solution finely models the two-dimensional feature vector through a Gaussian mixture model and weights the feature points based on the spatial distribution and temporal consistency of the features. This innovation significantly improves the ability to identify stable feature points, especially in dynamic scenarios, and can accurately distinguish static and dynamic features.
[0106] By dynamically adjusting the feature point weights, this solution can suppress the influence of false feature points caused by environmental changes (such as light changes, moving objects, etc.). Compared with traditional fixed-weight methods, this solution can flexibly adapt to different scenarios by quantifying the consistency in the spatial and temporal dimensions, improving the accuracy and robustness of feature matching. In complex environments, the dynamic neighborhood range and local covariance calculation further optimize the selection of feature points, ensuring high-precision image matching. The improved weighted feature points provide more reliable feature matching and scene reconstruction results, especially suitable for application scenarios such as autonomous driving and intelligent navigation that require high-precision image recognition. Compared with existing technologies, the adaptability and accuracy of this solution in dynamic scenarios have been significantly improved, reducing the risk of false matching and enhancing the robustness and stability of the system.
[0107] In an optional embodiment,
[0108] Obtain the vehicle movement speed and calculate the dynamic parallax threshold. Screen the parallax of the feature points in the initial matching pairs according to the dynamic parallax threshold to obtain reliable matching feature pairs. Construct an optimization equation based on the reliable matching feature pairs and solve to obtain the pose transformation parameters between adjacent images, including:
[0109] Obtain the vehicle motion speed and the time interval between adjacent frames, count the distribution of feature points in the image area, calculate the image entropy value as the scene structure complexity, and calculate the dynamic parallax threshold based on the vehicle motion speed, the time interval between adjacent frames, and the scene structure complexity;
[0110] Calculate the image coordinate difference of the initial matching feature point pairs in adjacent image frames to obtain the feature point parallax, use the initial matching feature point pairs with the feature point parallax less than the dynamic parallax threshold as candidate matching feature point pairs, and calculate the reference transformation matrix according to the image coordinate correspondence of the candidate matching feature point pairs;
[0111] Count other candidate matching feature point pairs within the neighborhood range of the candidate matching feature point pairs, calculate the difference score between the coordinate transformation result of other candidate matching feature point pairs and the reference transformation matrix, calculate the consistency support degree according to the difference score, and use the candidate matching feature point pairs with the consistency support degree greater than the preset support degree as reliable matching feature point pairs;
[0112] Calculate the local transformation matrix according to the image coordinate correspondence of the reliable matching feature point pairs, substitute the difference value between the local transformation matrix and the reference transformation matrix into the Gaussian kernel function to obtain the weight coefficient of the reliable matching feature point pairs, and construct a weighted reprojection error optimization equation using the corresponding relationship between the three-dimensional space coordinates and the two-dimensional image coordinates of the reliable matching feature point pairs and the weight coefficient;
[0113] Construct a Jacobian matrix according to the reprojection error optimization equation, construct a weight matrix using the weight coefficient, construct an increment equation according to the Jacobian matrix and the weight matrix, iteratively calculate the pose update amount of the increment equation and update the pose parameters, and obtain the pose transformation parameters between adjacent image frames when the pose update amount is less than the preset update amount threshold.
[0114] Exemplarily, obtain the vehicle motion speed and the time interval between adjacent frames, count the distribution of feature points in the image area, and calculate the image entropy value as the scene structure complexity. The vehicle motion speed represents the displacement size of the vehicle per unit time and reflects the motion state of the vehicle. The time interval between adjacent frames is the time interval between the acquisition of two frames of images. The feature point distribution refers to the key points that can be detected in the image, such as corner points or edge points. The image entropy value is used to measure the complexity of the image information. The higher the entropy value, the richer the effective information contained in the image. Based on the vehicle motion speed, the time interval between adjacent frames, and the scene structure complexity, calculate the dynamic parallax threshold. The parallax threshold determines the screening criteria for feature point matching, and its size is calculated comprehensively by factors such as vehicle motion speed, camera sampling frequency, and scene feature point density.
[0115] Calculate the image coordinate differences of the initial matching feature point pairs in adjacent image frames to obtain the feature point disparity. The feature point disparity refers to the pixel coordinate difference of the same feature point in two frames of images, reflecting the influence of the camera viewing angle change on the feature point position. Filter out the initial matching feature point pairs with disparities less than the dynamic disparity threshold to obtain candidate matching feature point pairs. The candidate matching feature point pairs refer to the feature point pairs that still have a high matching possibility after preliminary screening.
[0116] Calculate the reference transformation matrix based on the image coordinate correspondence of the candidate matching feature point pairs. The transformation matrix is used to describe the motion relationship of the camera between adjacent frames, including rotation and translation information. The reference transformation matrix is a preliminary camera pose transformation model estimated based on the candidate matching feature point pairs.
[0117] Count other candidate matching feature point pairs within the neighborhood range of the candidate matching feature point pairs, and calculate the difference score between their coordinate transformation results and the reference transformation matrix. The difference score is used to measure the deviation degree of the feature point pairs under the transformation matrix. The smaller the deviation, the greater the possibility that the matching points conform to the global transformation relationship. Calculate the consistency support degree according to the difference score. The consistency support degree refers to the credibility of a matching feature point pair among all matching points, reflecting the degree of conformity of this point with the overall motion trend. Filter out the matching point pairs with a consistency support degree greater than the preset support degree threshold as reliable matching feature point pairs.
[0118] Calculate the local transformation matrix based on the image coordinate correspondence of the reliable matching feature point pairs. The local transformation matrix describes the motion changes within the local range of the feature points. Calculate the difference value between the local transformation matrix and the reference transformation matrix, and calculate the weight coefficient of the reliable matching feature point pairs through the Gaussian kernel function. The weight coefficient is used to adjust the influence degree of different matching feature point pairs in the optimization calculation. The points with high credibility have large weights, and the points with low credibility have small weights.
[0119] Utilize the correspondence between the three-dimensional space coordinates and two-dimensional image coordinates of the reliable matching feature point pairs, as well as the calculated weight coefficients, to construct a weighted reprojection error optimization equation. The reprojection error refers to the deviation between the estimated position and the actual observed position of a three-dimensional point on the camera projection plane. This optimization equation is used to minimize this deviation to improve the accuracy of pose estimation.
[0120] Construct the Jacobian matrix, which is used to describe the rate of change of variables during the optimization process. Construct a weight matrix using weight coefficients, which is used to balance the influence of different feature point pairs. Based on the Jacobian matrix and the weight matrix, construct an incremental equation, and obtain the pose update amount by iteratively calculating the solution of the incremental equation, and continuously update the pose parameters. The pose update amount is used to describe the deviation between the currently calculated pose parameters and the true pose. When the deviation converges below a preset threshold, the optimization process terminates, and finally the pose transformation parameters between two adjacent frames of images are obtained.
[0121] In the prior art, during the pose estimation process, a fixed parallax threshold is usually used to screen feature point matches, without fully considering the influence of the vehicle motion state and scene structure, which easily leads to a decrease in matching accuracy in a dynamic environment. At the same time, when traditional methods optimize the pose, they mostly adopt a globally consistent weight assignment method, which cannot effectively distinguish between high-confidence feature points and low-confidence feature points, thus affecting the final estimation accuracy. This solution introduces a dynamic parallax threshold, adaptively adjusts the feature point screening criteria according to the vehicle motion speed, the time interval between adjacent frames, and the scene structure complexity, making the matching more robust. Compared with the traditional fixed-threshold method, this solution can adapt to different motion states and complex scenes, and improve the reliability of matching points. In addition, this solution screens reliable matching points through consistency support and calculates weights in combination with the Gaussian kernel function, so that different confidence feature points are given different influences during the optimization process, thereby improving the accuracy of pose estimation. The purpose of this solution is to enhance the robustness of feature matching and the adaptability of optimized weights to reduce the problem of false matching in a dynamic environment and improve the accuracy of the final pose estimation. Through the iterative calculation of the optimization equation, the pose update is made smoother and more accurate, thereby improving the robustness of the visual odometer or positioning system and providing a more reliable pose estimation ability for the autonomous driving or assisted driving system.
[0122] As Figure 2 shown, this figure shows the dynamic adjustment effect of the parallax threshold under different vehicle speeds and scene complexities. The curve represents the dynamic threshold change trend of this solution, and the dotted line represents the traditional fixed-threshold method. The horizontal axis represents the scene structure complexity (characterized by the image entropy value), and the vertical axis is the parallax threshold (unit: pixel). In a low-speed scenario (5 km / h), since the vehicle displacement is small, this solution sets the parallax threshold relatively large (about 8 pixels), which allows more potential matching points to participate in the calculation and improves the feature utilization rate. As the vehicle speed increases to medium speed (30 km / h), the parallax threshold dynamically decreases to about 5 pixels because stricter matching conditions are required to ensure accuracy at higher speeds. When the vehicle speed reaches high speed (60 km / h), the threshold further decreases to 3 pixels, effectively avoiding false matching caused by high-speed movement.
[0123] The curve shows that the threshold adjustment is also closely related to the scene complexity. In the area of complex scenes (high image entropy value), the parallax threshold tends to tighten because there are abundant feature points in complex environments and stricter screening criteria are needed. While in the area of simple scenes, the threshold is relatively relaxed to ensure obtaining sufficient matching points. This adaptive adjustment mechanism significantly improves the accuracy and robustness of feature matching.
[0124] In an alternative embodiment,
[0125] Obtain the vehicle motion speed and the time interval between adjacent frames, count the distribution of feature points in the image area, calculate the image entropy value as the scene structure complexity, and calculate the dynamic parallax threshold according to the vehicle motion speed, the time interval between adjacent frames and the scene structure complexity, including:
[0126] Obtain the vehicle motion speed and the time interval between adjacent frames, calculate the acceleration of the vehicle motion, classify the vehicle motion state according to the acceleration, calculate the ratio of the product of the vehicle motion speed and the acceleration to the cube of the vehicle motion speed to obtain the motion curvature;
[0127] Construct a multi-scale pyramid for the image, divide the image at each scale into adaptive grid cells, count the number of feature points in each grid cell, and calculate the ratio of the number of feature points in each grid cell to the total number of image feature points to obtain the feature point distribution probability;
[0128] Calculate the information entropy of each scale based on the feature point distribution probability, calculate the mean value of the image gradient direction in each grid cell, calculate the direction consistency based on the mean value of the image gradient direction, and combine the information entropy of each scale with the direction consistency as the scene structure complexity of the current scale;
[0129] Determine the weight coefficient of each scale according to the difference between the scene structure complexity and the mean value of the scene structure complexity, and perform cross-scale weighting on the scene structure complexity to obtain the comprehensive scene structure complexity;
[0130] Calculate the motion response value according to the vehicle motion speed and the motion curvature, and combine the comprehensive scene structure complexity with the motion response value to obtain the dynamic parallax threshold.
[0131] Exemplarily, after obtaining the vehicle motion speed and the time interval between adjacent frames, calculate the acceleration of the vehicle motion. Acceleration represents the rate of change of speed and can be used to determine whether the vehicle is moving at a constant speed, accelerating or decelerating. According to the numerical value and direction of the acceleration, determine the vehicle motion state, such as stationary, slowly accelerating, rapidly accelerating, slowly decelerating or rapidly decelerating.
[0132] Calculate the product of the vehicle's moving speed and acceleration, and perform a ratio operation with the cube of the vehicle's moving speed to obtain the motion curvature. The motion curvature is an important indicator to measure the degree of change in the vehicle's motion trajectory, and can reflect the smoothness and curvature characteristics of the vehicle's driving path.
[0133] Construct a multi-scale pyramid for the image. The pyramid structure refers to forming multiple image levels with different resolutions by repeatedly shrinking the original image, which helps to extract image features at different scales. In each scale of the image, divide the image into adaptive grid cells, and the size of each grid cell is dynamically adjusted according to the density of image feature points to ensure an even distribution of feature points in the local area.
[0134] Count the number of feature points in each grid cell and calculate the proportion of it to the total number of feature points in the entire image. This proportion is called the feature point distribution probability. The feature point distribution probability is used to measure the density of feature points in different regions of the image, thereby reflecting the local complexity of the image information.
[0135] Based on the feature point distribution probability, calculate the information entropy for each scale. The information entropy is an indicator to measure the complexity of the image content, representing the uniformity of the distribution of feature points in the image. The higher the information entropy, the more complex the distribution of feature points in the image. Further calculate the average value of the image gradient directions in each grid cell. The gradient direction represents the direction of the image brightness change, and the average value of the gradient direction can be used to measure the directional characteristics of the image edges and textures.
[0136] Based on the average value of the image gradient directions, calculate the direction consistency. The direction consistency indicates whether the edges and feature point directions in the local area of the image tend to be consistent. If the direction consistency is high, it means that the arrangement of feature points in this area is relatively regular, otherwise it means that the distribution of feature points in this area is relatively chaotic. Combine the information entropy and the direction consistency of each scale as the scene structure complexity of the current scale.
[0137] According to the difference between the scene structure complexity of each scale and the average value of the scene structure complexities of all scales, determine the weight coefficient for each scale. The weight coefficient is used to adjust the influence of different scales, so that the scale that can better represent the scene complexity obtains a higher weight. Perform a weighted calculation on the scene structure complexities of all scales to obtain the comprehensive scene structure complexity.
[0138] Calculate the motion response value based on the vehicle's moving speed and motion curvature. The motion response value is an important parameter to measure the change in the vehicle's current motion state and can reflect the dynamic characteristics of the vehicle. Finally, combine the comprehensive scene structure complexity and the motion response value to obtain the dynamic parallax threshold. The dynamic parallax threshold is used to adjust the feature point screening criteria in the image matching process, so that the feature matching can adapt to different motion states and scene complexities, improving the stability and accuracy of the matching.
[0139] In practical applications, assume that the vehicle is traveling at a high speed and the acceleration changes significantly within a short period of time. At this time, the calculated motion curvature is relatively large, indicating that the vehicle may be making a sharp turn or accelerating rapidly. In this case, in order to improve the accuracy of feature point matching, it is necessary to adjust the dynamic parallax threshold to adapt to the rapidly changing parallax features. At the same time, obtain the image captured by the front camera and construct a multi-scale pyramid, reducing the original image to different scales and dividing multiple grid cells at each scale. Count the number of feature points in each grid cell. For example, at a certain scale, the density of feature points in the left region is relatively high, while the feature points in the right region are relatively sparse, so the calculated probability of feature point distribution is larger on the left. Further calculate the information entropy of each grid cell. If the feature points in some regions of the image are dense and evenly distributed, the information entropy is relatively high, indicating that this region contains rich texture information. In addition, calculate the average value of the image gradient direction. If the gradient directions in a certain region tend to be consistent, the direction consistency of this region is relatively high.
[0140] Then, combine the information entropy and direction consistency at different scales, calculate the scene structure complexity at each scale, and adjust the influence of each scale through the weight coefficient to finally obtain the comprehensive scene structure complexity. Combine the comprehensive scene structure complexity with the motion response value to obtain the dynamic parallax threshold of the current frame image. For example, in the case of night or weak light, the overall number of feature points in the image is relatively small, the calculated information entropy is relatively low, and the direction consistency may be relatively high. Therefore, the calculated dynamic parallax threshold is adjusted accordingly so that the feature point matching can better adapt to the influence brought by the light change and ensure the accuracy of pose estimation. When the vehicle is traveling at a high speed or making a sharp turn, the dynamic parallax threshold will also be adjusted accordingly to adapt to the parallax change and improve the robustness of the matching.
[0141] When calculating the parallax threshold in the prior art, only the vehicle motion speed or simple image features are usually considered, without fully combining the vehicle motion state and the complexity of the scene structure, resulting in poor adaptability of the dynamic parallax threshold and affecting the matching accuracy. In this application, by introducing the vehicle motion acceleration and calculating the motion curvature, the classification of the vehicle motion state is further refined, so that the dynamic parallax threshold can more accurately adapt to different motion situations. At the same time, a multi-scale pyramid is used to construct the image hierarchical structure, and an adaptive grid is divided at different scales to calculate the distribution probability and information entropy of feature points, so as to more accurately describe the complexity of the scene structure. Compared with the traditional method that simply relies on single-scale features or global statistical information, this method can describe the image feature distribution in a finer granularity and improve the calculation accuracy of the parallax threshold. In addition, by combining the motion response value with the comprehensive scene structure complexity for threshold calculation, the dynamic parallax threshold can be adaptively adjusted under different motion states and scene changes, enhancing the robustness of the algorithm. Finally, this solution improves the screening quality of the matching feature points, improves the accuracy of camera pose estimation, and enables the system to obtain a more stable matching effect in complex scenes and different motion states.
[0142] In an alternative embodiment,
[0143] Combining the motion feature mapping matrix with the depth information to generate a compensated depth map, and reconstructing the scene and calculating the scene confidence score by using the compensated depth map and the pose transformation parameters includes:
[0144] Combining the motion feature mapping matrix with the depth information, performing propagation compensation on the depth information according to the motion information in the motion feature mapping matrix to obtain a compensated depth map, and performing three-dimensional back-projection on the compensated depth map to obtain a scene point cloud;
[0145] Extracting planar structure elements and edge structure elements from the scene point cloud, calculating the spatial distribution characteristics of the planar structure elements and the edge structure elements to obtain a structure distribution map, and constructing a topological relationship map of the structure elements based on the structure distribution map;
[0146] Decomposing the pose transformation parameters into a rotation component and a translation component, obtaining the auxiliary pose information of the visual odometer, calculating the fusion weight of the pose transformation parameters and the auxiliary pose information according to the structure distribution map, and performing weighted fusion on the pose transformation parameters and the auxiliary pose information according to the fusion weight to obtain a fused pose parameter;
[0147] Performing coordinate transformation on the scene point cloud according to the fused pose parameter to obtain a reconstructed scene, calculating the local geometric consistency of the point cloud in the reconstructed scene to obtain a local confidence, calculating the structural similarity of the point cloud at adjacent moments in the reconstructed scene to obtain a global confidence, and performing weighted combination on the local confidence and the global confidence to obtain a scene confidence score;
[0148] Calibrate the pose transformation parameters based on the scene confidence score to obtain the calibrated pose parameter sequence.
[0149] Exemplarily, combine the motion feature mapping matrix with depth information to obtain more accurate scene depth data. The motion feature mapping matrix is used to describe the motion changes between image frames and contains the displacement information of feature points in the image coordinate system. This matrix can be calculated by the optical flow method or feature matching method. The depth information comes from sensors (such as lidar, binocular cameras, or structured light depth cameras) and represents the distance from objects in the scene to the camera. By using the motion feature mapping matrix to perform propagation compensation on the depth information, correct the depth offset caused by camera motion, and obtain the compensated depth map.
[0150] Next, perform three-dimensional back-projection on the compensated depth map to generate the scene point cloud. Back-projection is the process of converting points in the image coordinate system to the world coordinate system. Specifically, use the internal and external parameters of the camera to convert the depth value of each pixel into a three-dimensional spatial coordinate to form the scene point cloud.
[0151] Extract plane structure elements and edge structure elements from the point cloud data to analyze the geometric features of the scene. The plane structure elements are a set of points formed by the plane parts of objects and can be detected by the RANSAC (Random Sample Consensus) method; the edge structure elements represent regions where the object contour or surface changes drastically and can be extracted by gradient change analysis or methods based on normal vector calculation. Calculate the spatial distribution characteristics of these elements to form a structure distribution map, which is used to represent the distribution of different structures in three-dimensional space.
[0152] Construct a topological relationship graph based on the structure distribution map to describe the connection relationship between different structure elements in the point cloud data. The nodes of the topological relationship graph represent plane or edge elements, and the edges represent their spatial relationships, such as adjacent or coplanar relationships.
[0153] Decompose the pose transformation parameters into a rotation component and a translation component to more precisely analyze the motion state of the camera. The rotation component describes the angular change of the camera, and the translation component represents the displacement of the camera. Obtain the auxiliary pose information of the visual odometer. The visual odometer estimates the relative motion of the camera by tracking feature points in the image. Common methods include optical flow tracking, feature point matching, and direct methods.
[0154] Calculate the fusion weight of the pose transformation parameters and the auxiliary pose information according to the structure distribution map to ensure that the fused pose data is more reliable. The fusion weight represents the credibility of different data sources and is usually determined by historical data, sensor characteristics, or adaptive optimization methods. Weightedly fuse the pose transformation parameters and the auxiliary pose information according to the fusion weight to obtain the fused pose parameters.
[0155] Perform coordinate transformation on the scene point cloud according to the fused pose parameters to obtain the reconstructed scene in a unified coordinate system. Coordinate transformation is to convert all point cloud data into the same reference coordinate system to eliminate the influence of camera movement on data alignment.
[0156] Calculate the local geometric consistency of the point cloud in the reconstructed scene to measure the matching degree between different frame data. Local geometric consistency is calculated based on the spatial distribution of neighborhood points and is used to measure whether the local structure of the point cloud is consistent. For example, it can be evaluated by calculating the similarity of the normal vectors of the point cloud or the change in the distance from the point to the plane.
[0157] Calculate the structural similarity of the point cloud at adjacent times in the reconstructed scene to evaluate the global matching situation. Structural similarity is usually based on the overall shape of the point cloud, such as histogram description, point cloud density distribution, or feature curvature analysis.
[0158] Perform weighted combination of the local geometric consistency and the global structural similarity to obtain the scene confidence score. The confidence score represents the reliability of point cloud alignment, and a high confidence means that the pose estimation is relatively accurate.
[0159] Calibrate the pose transformation parameters based on the scene confidence score to optimize the pose sequence between adjacent frames, and finally obtain the calibrated pose parameter sequence. Pose calibration improves the accuracy and stability of the system by adjusting the initial estimate to better conform to the actual motion trajectory.
[0160] Existing technologies usually use fixed thresholds or simple filtering methods for depth information compensation, which are difficult to accurately adapt to changes in different motion states and scene structures, resulting in unstable depth compensation effects and affecting the accuracy of three-dimensional reconstruction. At the same time, in the process of pose calculation, traditional methods often rely on single features or simple weighted fusion and are difficult to fully utilize structural information for pose optimization, resulting in large cumulative errors. This application constructs a motion feature mapping matrix and combines depth information for propagation compensation, enabling the depth information to adapt to the dynamic motion environment and improving the accuracy of compensation. By constructing the scene point cloud through three-dimensional back-projection and extracting plane and edge structure elements, the scene structure features can be more comprehensively characterized, and the geometric consistency can be enhanced through topological relationships to improve the quality of the point cloud.
[0161] In terms of pose calculation, the present application decomposes the pose transformation parameters, combines the auxiliary pose information, and optimizes them through an adaptive weight fusion method. It makes full use of the scene structure information to constrain the pose calculation and reduce the cumulative error. At the same time, it calculates the scene confidence based on the local geometric consistency and global structure similarity, and calibrates the pose parameters using this confidence, thereby improving the accuracy and stability of pose estimation. Compared with the prior art, the improvement starting point of the present application is to combine the motion characteristics, scene structure, and multi-source pose information to optimize the depth compensation and pose calculation, making the 3D reconstruction result more accurate, reducing the pose estimation error, and improving the adaptability in complex environments.
[0162] As Figure 3 shown, this figure shows the comparison results of the 3D reconstruction accuracy of different methods under varying scene complexities. This technical solution shows obvious advantages in all scene complexities: when the scene complexity is 0.1, the reconstruction accuracy reaches 0.42 mm, while the accuracies of the traditional optical flow method, feature point method, and ORB-SLAM are approximately 0.48 mm, 0.52 mm, and 0.56 mm respectively; as the scene complexity increases to 0.9, the reconstruction accuracy of this technical solution still remains at 0.89 mm, significantly better than other methods (0.82 mm, 0.78 mm, and 0.75 mm respectively). Especially in the medium-complexity scenes with a scene complexity of 0.3 - 0.7, the accuracy improvement of this technical solution is particularly significant, with an average improvement of approximately 25% in reconstruction accuracy compared to other methods, reflecting the robustness and adaptability of this solution in complex scenes.
[0163] In an optional embodiment,
[0164] Calibrating the pose transformation parameters based on the scene confidence score to obtain the calibrated pose parameter sequence includes:
[0165] Establish an evaluation vector for each pose according to the scene confidence score, compare the evaluation vector with the preset vector threshold to obtain the pose credibility index, and mark the abnormal poses among the poses where the pose credibility index is lower than the preset vector threshold;
[0166] Use the topological relationship graph to construct the influence propagation matrix of the abnormal poses, analyze the structural association strength between the abnormal poses and adjacent poses in the influence propagation matrix, determine the spatio-temporal influence range of the abnormal poses according to the structural association strength, and use the spatio-temporal influence range as the calibration range;
[0167] Construct a pose optimization objective function, which includes a structure preservation term, a depth consistency term, and a temporal smoothness term. Iteratively optimize the abnormal pose. In each iteration, calculate the structure preservation score of the current pose based on the structure distribution map, calculate the depth consistency score based on the compensated depth map, calculate the temporal smoothness score based on adjacent poses, and substitute the structure preservation score, depth consistency score, and temporal smoothness score into the pose optimization objective function to calculate the optimization direction and step size;
[0168] After each round of iterative optimization, reapply the calibrated pose obtained by optimization to scene reconstruction, extract the structural features in the scene reconstruction to calculate the structural consistency, and stop the optimization when the structural consistency is greater than the preset convergence threshold;
[0169] Update the original pose sequence according to the finally optimized calibrated pose, and organize the calibrated pose parameter sequence in chronological order.
[0170] Exemplarily, first construct a pose evaluation index, and evaluate the reliability of each pose through the scene confidence score. The scene confidence score includes two aspects: the scene reconstruction quality and the feature matching reliability. The scene reconstruction quality is evaluated by the density and distribution uniformity of the reconstructed point cloud; the feature matching reliability is evaluated by the reprojection error of the matching feature pairs and the similarity of the feature descriptors. Convert the scene confidence score into a multi-dimensional evaluation vector, and the dimensions of the evaluation vector include the reconstructed point cloud density, distribution uniformity, reprojection error, and descriptor similarity. The preset vector threshold is obtained based on a large amount of experimental data statistics, indicating the normal range of each dimension index. When any dimension of the evaluation vector is lower than the corresponding threshold, mark this pose as an abnormal pose.
[0171] The topological relationship graph is used to describe the spatio-temporal correlation relationship between poses, where nodes represent poses and edges represent the connection relationship between poses. The influence propagation matrix describes the influence degree of an abnormal pose on other poses, and the matrix element values are jointly determined by the time distance and space distance between poses. The structural association strength is calculated by the number of co-visible feature points and the structural overlap degree of the reconstructed scene. The method for determining the spatio-temporal influence range is: starting from the abnormal pose, gradually expand to adjacent poses, and stop expanding when the structural association strength is lower than the preset threshold. The finally obtained range is the pose range that needs to be calibrated.
[0172] The pose optimization objective function contains three constraint terms: the structure preservation term ensures the consistency of the scene structure features, the depth consistency term guarantees the continuity of the scene depth information, and the temporal smoothness term constrains the smooth change of adjacent poses. The structure preservation term is calculated through the structure distribution map, which records the main structure lines and plane features in the scene. The depth consistency term is calculated based on the compensated depth map, and the compensated depth map is generated through the feature change sequence and depth information. The temporal smoothness term is obtained by calculating the relative motion of adjacent poses.
[0173] The iterative optimization process includes: First, reconstruct the scene based on the current pose, extract the structural features in the scene and compare them with the structural distribution map, and calculate the structural retention score. Then, calculate the depth consistency score using the compensated depth map, and calculate the temporal smoothness score based on the relative motion of adjacent poses. Substitute the three scores into the objective function, and use the gradient descent method to determine the optimization direction and step size. The optimization direction is the opposite direction of the gradient direction of the objective function at the current pose, and the step size is determined by line search.
[0174] After each round of iteration, use the optimized pose to reconstruct the scene and extract the structural features. The structural consistency is calculated by comparing the structural features between adjacent frames, including the direction consistency of the structural lines and the normal vector consistency of the plane features. When the structural consistency is higher than the preset convergence threshold for multiple consecutive rounds, it is considered that the optimization reaches the convergence state and the iterative process stops.
[0175] Replace the corresponding abnormal poses in the original pose sequence with the finally optimized calibration poses in chronological order, keep other poses unchanged, and finally obtain the calibrated complete pose parameter sequence.
[0176] In a continuous image sequence of an urban road scene, a large number of building contour lines and ground marking features are extracted through a feature detection network. At a certain moment, due to building occlusion, the feature matching is abnormal. At this pose, it is detected that the density of the reconstructed point cloud is reduced to one-third of the normal value, and the reprojection error of the feature points is increased to twice the normal value. Through the analysis of the topological relationship graph, it is found that this abnormal pose has a strong structural association with the poses of the two frames before and after, and multiple building edge lines and ground markings are jointly observed. During the optimization process, the continuity of these structural lines is mainly maintained, and at the same time, the smooth transition of the depth map is ensured. After multiple rounds of iterative optimization, the direction consistency of the structural lines is significantly improved, and the jumps in the depth map are significantly reduced, and finally a calibrated continuous pose sequence is obtained.
[0177] Figure 4 It is a schematic structural diagram of the interaction system based on deep learning according to an embodiment of the present invention, as Figure 4 shown, the system includes:
[0178] The first unit is used to collect the continuous image sequence obtained by the vehicle-mounted camera and extract feature points, combine the position change amount of the feature points at adjacent moments and the gradient change amount of the feature points at adjacent moments to construct a two-dimensional feature vector, use the Gaussian mixture model to cluster the two-dimensional feature vector to obtain a stable feature class, calculate the feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature class, and assign weights to the feature points according to the weight distribution function to obtain weighted feature points;
[0179] A second unit, configured to calculate feature similarity based on weighted feature points to obtain a similarity matrix, normalize the similarity matrix, construct a feature matching relationship using the normalized similarity matrix to obtain initial matching pairs, acquire the vehicle movement speed and calculate a dynamic parallax threshold, screen the parallax of feature points in the initial matching pairs according to the dynamic parallax threshold to obtain reliable matching feature pairs, construct an optimization equation based on the reliable matching feature pairs and solve it to obtain pose transformation parameters between adjacent images;
[0180] A third unit, configured to input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate the feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with depth information to generate a compensated depth map, reconstruct a scene using the compensated depth map and the pose transformation parameters and calculate a scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0181] In a third aspect of the embodiments of the present invention,
[0182] there is provided an electronic device, comprising:
[0183] a processor;
[0184] a memory for storing instructions executable by the processor;
[0185] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0186] In a fourth aspect of the embodiments of the present invention,
[0187] there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0188] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.
[0189] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The interactive method based on deep learning is characterized by: include: The continuous image sequence obtained by the vehicle-mounted camera is collected and feature points are extracted. The position change of the feature points at adjacent moments and the gradient change of the feature points at adjacent moments are combined to construct a two-dimensional feature vector. The two-dimensional feature vector is clustered using a Gaussian mixture model to obtain a stable feature class. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The feature points are weighted according to the weight distribution function to obtain weighted feature points. The similarity matrix is obtained by calculating the feature similarity based on the weighted feature points, and the similarity matrix is normalized. The feature matching relationship is constructed using the normalized similarity matrix to obtain the initial matching pair, the vehicle movement speed is obtained and the dynamic disparity threshold is calculated, and the disparity of the feature points in the initial matching pair is screened according to the dynamic disparity threshold to obtain reliable matching feature point pairs, and the optimization equation is constructed according to the reliable matching feature point pairs and solved to obtain the posture transformation parameters between adjacent images; Input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate the feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with depth information to generate a compensated depth map, use the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence; Obtaining the vehicle's speed and calculating the dynamic parallax threshold include: Obtaining the vehicle speed and the time interval between adjacent frames, calculating the acceleration of the vehicle, classifying the vehicle motion state according to the acceleration, calculating the product of the vehicle speed and the acceleration, and obtaining the motion curvature by the ratio of the product to the cube of the vehicle speed; Construct a multi-scale pyramid for the image, divide the image of each scale into adaptive grid units, count the number of feature points in each grid unit, and calculate the ratio of the number of feature points in each grid unit to the total number of feature points in the image to obtain the distribution probability of feature points; Calculating the information entropy of each scale based on the distribution probability of the feature points, calculating the image gradient direction mean in each grid unit, calculating the directional consistency based on the image gradient direction mean, and combining the information entropy of each scale with the directional consistency as the scene structure complexity of the current scale; Determining a weight coefficient for each scale according to the difference between the scene structure complexity and the mean value of the scene structure complexity, and weighting the scene structure complexity across scales to obtain a comprehensive scene structure complexity; A motion response value is calculated according to the vehicle motion speed and the motion curvature, and a dynamic parallax threshold is obtained by combining the comprehensive scene structure complexity with the motion response value.
2. The method according to claim 1, characterized in that The Gaussian mixture model is used to cluster the two-dimensional feature vectors to obtain stable feature classes. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The weighted feature points obtained by weighting the feature points according to the weight distribution function include: Using a Gaussian mixture model to calculate a mixture weight, a mean vector and a covariance matrix of a mixed Gaussian distribution for a two-dimensional feature vector, and using the mixture weight, the mean vector and the covariance matrix to cluster the two-dimensional feature vector to obtain a plurality of feature classes; The average distance from the intra-class feature vector to the class center of the feature class is calculated to obtain the intra-class aggregation degree, the minimum distance between the class centers of the feature classes is calculated to obtain the inter-class separation degree, the feature classes are screened according to the ratio of the intra-class aggregation degree to the inter-class separation degree, and the feature classes that meet the preset feature threshold conditions are determined as stable feature classes; Calculate the intra-class covariance matrix of the two-dimensional feature vector in the stable feature class and perform eigenvalue decomposition to obtain spatial distribution features, calculate the Euclidean distance of the two-dimensional feature vector at adjacent moments and convert it into a time series correlation score, and perform weighted combination of the spatial distribution feature and the time series correlation score to obtain a feature weight distribution function; The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector to obtain the spatial weight, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are fused and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain a weighted feature point.
3. The method according to claim 2, characterized in that The spatial weight is obtained by calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are merged and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain the weighted feature point, which includes: The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector, and the spatial weight is obtained by exponential mapping with the Mahalanobis distance; Calculate the curvature of the motion trajectory of the two-dimensional feature vector as a time scale coefficient, perform nonlinear combination of the time scale coefficient and the Euclidean distance of the two-dimensional feature vector at adjacent moments to obtain a time series correlation score, and calculate a time series weight based on the time series correlation score; Constructing a local variance distribution of the spatial weight, determining an adaptive kernel function based on the local variance distribution, and using the adaptive kernel function to fuse the spatial weight with the temporal weight to obtain a fused weight; Calculating the spatial distribution gradient of the fusion weight, constructing an adaptive neighborhood radius based on the spatial distribution gradient, and combining the adaptive neighborhood radius with a preset neighborhood range through density weighting to obtain a dynamic neighborhood range; Calculate the local covariance characteristics of the fusion weight within the dynamic neighborhood, construct an anisotropic Gaussian kernel function, and use the anisotropic Gaussian kernel function to spatially modulate the fusion weight; normalize the modulated fusion weight within the dynamic neighborhood to obtain an initial weight value, calculate the weight consistency constraint based on the initial weight value, iteratively optimize the weight consistency constraint and the initial weight value to obtain a final weight value, and assign the final weight value to the corresponding two-dimensional feature vector to obtain a weighted feature point.
4. The method according to claim 1, characterized in that: The vehicle speed is obtained and the dynamic parallax threshold is calculated. The parallax of the feature points in the initial matching pair is screened according to the dynamic parallax threshold to obtain reliable matching feature point pairs. The optimization equation is constructed based on the reliable matching feature point pairs and the pose transformation parameters between adjacent images are obtained by solving the equations. The vehicle speed and the time interval between adjacent frames are obtained, the distribution of feature points in the image area is counted to calculate the image entropy value as the scene structure complexity, and the dynamic parallax threshold is calculated according to the vehicle speed, the time interval between adjacent frames and the scene structure complexity; Calculating the image coordinate difference of the initial matching feature point pairs in adjacent image frames to obtain feature point disparity, taking the initial matching feature point pairs whose feature point disparity is less than the dynamic disparity threshold as candidate matching feature point pairs, and calculating the reference transformation matrix according to the image coordinate correspondence of the candidate matching feature point pairs; Counting other candidate matching feature point pairs within the neighborhood of the candidate matching feature point pair, calculating the difference scores between the coordinate transformation results of other candidate matching feature point pairs and the reference transformation matrix, calculating the consistency support according to the difference scores, and taking the candidate matching feature point pairs whose consistency support is greater than the preset support as reliable matching feature point pairs; A local transformation matrix is calculated according to the image coordinate correspondence of the reliable matching feature point pair, a difference value between the local transformation matrix and the reference transformation matrix is substituted into a Gaussian kernel function to obtain a weight coefficient of the reliable matching feature point pair, and a weighted reprojection error optimization equation is constructed using the correspondence between the three-dimensional space coordinates and the two-dimensional image coordinates of the reliable matching feature point pair and the weight coefficient; A Jacobian matrix is constructed according to the reprojection error optimization equation, a weight matrix is constructed using the weight coefficients, an incremental equation is constructed according to the Jacobian matrix and the weight matrix, the posture update amount of the incremental equation is iteratively calculated and the posture parameters are updated, and when the posture update amount is less than a preset update amount threshold, the posture transformation parameters between adjacent image frames are obtained.
5. The method according to claim 1, characterized in that Combining the motion feature map matrix with the depth information to generate a compensated depth map, using the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score includes: Combining the motion feature mapping matrix with the depth information, performing propagation compensation on the depth information according to the motion information in the motion feature mapping matrix to obtain a compensated depth map, and performing three-dimensional back-projection on the compensated depth map to obtain a scene point cloud; Extracting plane structural elements and edge structural elements from the scene point cloud, calculating spatial distribution characteristics of the plane structural elements and edge structural elements to obtain a structural distribution map, and constructing a topological relationship map of the structural elements based on the structural distribution map; Decomposing the posture transformation parameters into rotation components and translation components, obtaining auxiliary posture information of the visual odometer, calculating the fusion weights of the posture transformation parameters and the auxiliary posture information according to the structural distribution map, and weightedly fusing the posture transformation parameters and the auxiliary posture information according to the fusion weights to obtain fused posture parameters; Performing coordinate transformation on the scene point cloud according to the fusion pose parameters to obtain a reconstructed scene, calculating the local geometric consistency of the point cloud in the reconstructed scene to obtain a local confidence, calculating the structural similarity of the point clouds at adjacent moments in the reconstructed scene to obtain a global confidence, and performing a weighted combination of the local confidence and the global confidence to obtain a scene confidence score; The pose transformation parameters are calibrated based on the scene confidence score to obtain a calibrated pose parameter sequence.
6. The method according to claim 5, characterized in that The pose transformation parameters are calibrated based on the scene confidence score to obtain the calibrated pose parameter sequence including: Establishing an evaluation vector for each posture according to the scene confidence score, comparing the evaluation vector with a preset vector threshold to obtain a posture credibility index, and marking abnormal postures in postures whose posture credibility index is lower than the preset vector threshold; An influence propagation matrix of abnormal posture is constructed by using the topological relationship graph, the structural correlation strength between the abnormal posture and the adjacent posture is analyzed in the influence propagation matrix, the spatiotemporal influence range of the abnormal posture is determined according to the structural correlation strength, and the spatiotemporal influence range is used as the calibration range; Constructing a pose optimization objective function, the pose optimization objective function includes a structure preservation term, a depth consistency term and a temporal smoothing term, iteratively optimizing the abnormal pose, calculating a structure preservation score of the current pose based on a structure distribution map, calculating a depth consistency score based on a compensated depth map, and calculating a temporal smoothing score based on adjacent poses in each iteration, substituting the structure preservation score, the depth consistency score and the temporal smoothing score into the pose optimization objective function to calculate an optimization direction and a step size; After each round of iterative optimization, the optimized calibration pose is reapplied to the scene reconstruction, the structural features in the scene reconstruction are extracted to calculate the structural consistency, and the optimization is stopped when the structural consistency is greater than a preset convergence threshold; The original pose sequence is updated according to the calibration pose obtained by the final optimization, and the calibrated pose parameter sequence is obtained by chronological order.
7. An interactive system based on deep learning, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to collect a continuous image sequence obtained by the vehicle-mounted camera and extract feature points, combine the position change of the feature points at adjacent moments with the gradient change of the feature points at adjacent moments to construct a two-dimensional feature vector, cluster the two-dimensional feature vector using a Gaussian mixture model to obtain a stable feature class, calculate a feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature class, and assign weights to the feature points according to the weight distribution function to obtain weighted feature points; The second unit is used to calculate feature similarity based on weighted feature points to obtain a similarity matrix, normalize the similarity matrix, use the normalized similarity matrix to construct a feature matching relationship to obtain an initial matching pair, obtain the vehicle movement speed and calculate the dynamic disparity threshold, screen the disparity of the feature points in the initial matching pair according to the dynamic disparity threshold to obtain a reliable matching feature point pair, and construct an optimization equation based on the reliable matching feature point pair and solve it to obtain the posture transformation parameters between adjacent images; The third unit is used to input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate the feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with the depth information to generate a compensated depth map, use the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Visual positioning method for complex road environment in automatic driving scene
CN117710930A
Road scene point cloud identification method and system suitable for inspection robot
CN118887641A