Interaction method and system based on deep learning
By adopting deep learning-based interaction methods in the autonomous driving system, using Gaussian hybrid model and feature weight distribution function to optimize feature point matching, and combining depth information for scene reconstruction and pose calibration, the problems of instability in position estimation and high computational complexity in the existing technology are solved, and more efficient and accurate pose estimation and scene reconstruction are achieved.
Patent Information
- Application Number
- CN202510362969.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-03-26
AI Technical Summary
In complex environments such as autonomous driving, the vision-based pose estimation is affected by factors such as lighting changes, dynamic occlusion, motion blur and scene changes, resulting in unstable matching of feature points, accumulated errors in pose estimation, and deep learning methods have high computational complexity and insufficient generalization capabilities, making it difficult to effectively combine feature change information and depth information for error compensation.
A deep learning-based interaction method is adopted, by collecting continuous image sequences of on-board cameras, extracting feature points and constructing two-dimensional feature vectors, using Gaussian mixed model clustering to obtain stable feature classes, computed feature weight distribution function to empower feature points, computed feature similarity and dynamic parallax threshold based on weighted feature points, filtered reliable matching feature pairs, constructed optimization equations to solve pose transformation parameters, and generated compensation depth maps based on depth information, reconstructed the scene and calculated the scene confidence score for pose calibration.
It improves the stability of feature matching and the accuracy of position estimation, enhances the quality of scene reconstruction and the adaptability of the system in complex environments, and is suitable for applications such as autonomous driving, intelligent transportation and high-precision map construction.
Smart Images

Figure CN119919749A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision technology, and in particular to an interaction method and system based on deep learning. Background Art
[0002] In applications such as autonomous driving, intelligent transportation, and robot navigation, vision-based pose estimation is one of the key technologies. Traditional methods usually rely on feature point detection and matching, combined with geometric constraints to solve pose transformation parameters. However, in real environments, factors such as illumination changes, dynamic occlusion, motion blur, and scene changes can affect the stability of feature point matching, leading to the accumulation of pose estimation errors. In addition, traditional methods use a fixed matching threshold, which cannot adapt to changes in vehicle speeds, viewing angles, and scenes, reducing the reliability of pose calculations.
[0003] The development of deep learning in the field of computer vision has provided new ideas for feature extraction, matching optimization and pose estimation. Although deep learning methods have been applied to feature detection and matching, they generally have the problems of high computational complexity and insufficient generalization ability. At the same time, the potential of depth information in pose estimation has not been fully utilized. Existing methods are difficult to effectively combine feature change information with depth information for error compensation, resulting in limited scene reconstruction accuracy. Therefore, there is an urgent need for an interactive technology that integrates deep learning and traditional visual methods to improve feature matching stability, pose estimation accuracy and scene reconstruction quality to meet the application requirements in complex environments such as autonomous driving. Summary of the invention
[0004] The embodiments of the present invention provide an interactive method and system based on deep learning, which can solve the problems in the prior art.
[0005] According to a first aspect of the embodiments of the present invention, Provides an interactive method based on deep learning, including: The continuous image sequence obtained by the vehicle-mounted camera is collected and feature points are extracted. The position change of the feature points at adjacent moments and the gradient change of the feature points at adjacent moments are combined to construct a two-dimensional feature vector. The two-dimensional feature vector is clustered using a Gaussian mixture model to obtain a stable feature class. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The feature points are weighted according to the weight distribution function to obtain weighted feature points. The similarity matrix is obtained by calculating the feature similarity based on the weighted feature points, and the similarity matrix is normalized. The feature matching relationship is constructed using the normalized similarity matrix to obtain an initial matching pair, the vehicle movement speed is obtained and the dynamic disparity threshold is calculated, and the disparity of the feature points in the initial matching pair is screened according to the dynamic disparity threshold to obtain a reliable matching feature pair, and an optimization equation is constructed based on the reliable matching feature pair and solved to obtain the posture transformation parameters between adjacent images; A continuous image sequence is input into the feature detection network to obtain a feature change sequence. The feature change gradient between adjacent frames is calculated according to the feature change sequence. A motion feature mapping matrix is constructed based on the feature change gradient. The motion feature mapping matrix is combined with the depth information to generate a compensated depth map. The compensated depth map and pose transformation parameters are used to reconstruct the scene and calculate the scene confidence score. The pose transformation parameters are calibrated based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0006] In an optional embodiment, The Gaussian mixture model is used to cluster the two-dimensional feature vectors to obtain stable feature classes. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The weighted feature points obtained by weighting the feature points according to the weight distribution function include: Using a Gaussian mixture model to calculate a mixture weight, a mean vector and a covariance matrix of a mixed Gaussian distribution for a two-dimensional feature vector, and using the mixture weight, the mean vector and the covariance matrix to cluster the two-dimensional feature vector to obtain a plurality of feature classes; The average distance from the intra-class feature vector to the class center of the feature class is calculated to obtain the intra-class aggregation degree, the minimum distance between the class centers of the feature classes is calculated to obtain the inter-class separation degree, the feature classes are screened according to the ratio of the intra-class aggregation degree to the inter-class separation degree, and the feature classes that meet the preset feature threshold conditions are determined as stable feature classes; Calculate the intra-class covariance matrix of the two-dimensional feature vector in the stable feature class and perform eigenvalue decomposition to obtain spatial distribution features, calculate the Euclidean distance of the two-dimensional feature vector at adjacent moments and convert it into a time series correlation score, and perform weighted combination of the spatial distribution feature and the time series correlation score to obtain a feature weight distribution function; The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector to obtain the spatial weight, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are fused and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain a weighted feature point.
[0007] In an optional embodiment, The spatial weight is obtained by calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are merged and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain the weighted feature point, which includes: The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector, and the spatial weight is obtained by exponential mapping with the Mahalanobis distance; Calculate the curvature of the motion trajectory of the two-dimensional feature vector as a time scale coefficient, perform nonlinear combination of the time scale coefficient and the Euclidean distance of the two-dimensional feature vector at adjacent moments to obtain a time series correlation score, and calculate a time series weight based on the time series correlation score; Constructing a local variance distribution of the spatial weight, determining an adaptive kernel function based on the local variance distribution, and using the adaptive kernel function to fuse the spatial weight with the temporal weight to obtain a fused weight; Calculating the spatial distribution gradient of the fusion weight, constructing an adaptive neighborhood radius based on the spatial distribution gradient, and combining the adaptive neighborhood radius with a preset neighborhood range through density weighting to obtain a dynamic neighborhood range; Calculate the local covariance characteristics of the fusion weight within the dynamic neighborhood, construct an anisotropic Gaussian kernel function, and use the anisotropic Gaussian kernel function to spatially modulate the fusion weight; normalize the modulated fusion weight within the dynamic neighborhood to obtain an initial weight value, calculate the weight consistency constraint based on the initial weight value, iteratively optimize the weight consistency constraint and the initial weight value to obtain a final weight value, and assign the final weight value to the corresponding two-dimensional feature vector to obtain a weighted feature point.
[0008] In an optional embodiment, The vehicle speed is obtained and the dynamic parallax threshold is calculated. The parallax of the feature points in the initial matching pair is screened according to the dynamic parallax threshold to obtain reliable matching feature pairs. The optimization equation is constructed based on the reliable matching feature pairs and the pose transformation parameters between adjacent images are obtained by solving the equations. The vehicle speed and the time interval between adjacent frames are obtained, the distribution of feature points in the image area is counted to calculate the image entropy value as the scene structure complexity, and the dynamic parallax threshold is calculated according to the vehicle speed, the time interval between adjacent frames and the scene structure complexity; Calculating the image coordinate difference of the initial matching feature point pairs in adjacent image frames to obtain feature point disparity, taking the initial matching feature point pairs whose feature point disparity is less than the dynamic disparity threshold as candidate matching feature point pairs, and calculating the reference transformation matrix according to the image coordinate correspondence of the candidate matching feature point pairs; Counting other candidate matching feature point pairs within the neighborhood of the candidate matching feature point pair, calculating the difference scores between the coordinate transformation results of other candidate matching feature point pairs and the reference transformation matrix, calculating the consistency support according to the difference scores, and taking the candidate matching feature point pairs whose consistency support is greater than the preset support as reliable matching feature point pairs; A local transformation matrix is calculated according to the image coordinate correspondence of the reliable matching feature point pair, a difference value between the local transformation matrix and the reference transformation matrix is substituted into a Gaussian kernel function to obtain a weight coefficient of the reliable matching feature point pair, and a weighted reprojection error optimization equation is constructed using the correspondence between the three-dimensional space coordinates and the two-dimensional image coordinates of the reliable matching feature point pair and the weight coefficient; A Jacobian matrix is constructed according to the reprojection error optimization equation, a weight matrix is constructed using the weight coefficients, an incremental equation is constructed according to the Jacobian matrix and the weight matrix, the posture update amount of the incremental equation is iteratively calculated and the posture parameters are updated, and when the posture update amount is less than a preset update amount threshold, the posture transformation parameters between adjacent image frames are obtained.
[0009] In an optional embodiment, The vehicle speed and the time interval between adjacent frames are obtained, and the distribution of feature points in the image area is counted to calculate the image entropy value as the scene structure complexity. The dynamic parallax threshold is calculated according to the vehicle speed, the time interval between adjacent frames and the scene structure complexity, including: Obtaining the vehicle speed and the time interval between adjacent frames, calculating the acceleration of the vehicle, classifying the vehicle motion state according to the acceleration, calculating the product of the vehicle speed and the acceleration, and obtaining the motion curvature by the ratio of the product to the cube of the vehicle speed; Construct a multi-scale pyramid for the image, divide the image of each scale into adaptive grid units, count the number of feature points in each grid unit, and calculate the ratio of the number of feature points in each grid unit to the total number of feature points in the image to obtain the distribution probability of feature points; Calculating the information entropy of each scale based on the distribution probability of the feature points, calculating the image gradient direction mean in each grid unit, calculating the directional consistency based on the image gradient direction mean, and combining the information entropy of each scale with the directional consistency as the scene structure complexity of the current scale; Determining a weight coefficient for each scale according to the difference between the scene structure complexity and the mean value of the scene structure complexity, and weighting the scene structure complexity across scales to obtain a comprehensive scene structure complexity; A motion response value is calculated according to the vehicle motion speed and the motion curvature, and a dynamic parallax threshold is obtained by combining the comprehensive scene structure complexity with the motion response value.
[0010] In an optional embodiment, Combining the motion feature map matrix with the depth information to generate a compensated depth map, using the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score includes: Combining the motion feature mapping matrix with the depth information, performing propagation compensation on the depth information according to the motion information in the motion feature mapping matrix to obtain a compensated depth map, and performing three-dimensional back-projection on the compensated depth map to obtain a scene point cloud; Extracting plane structural elements and edge structural elements from the scene point cloud, calculating spatial distribution characteristics of the plane structural elements and edge structural elements to obtain a structural distribution map, and constructing a topological relationship map of the structural elements based on the structural distribution map; Decomposing the posture transformation parameters into rotation components and translation components, obtaining auxiliary posture information of the visual odometer, calculating the fusion weights of the posture transformation parameters and the auxiliary posture information according to the structural distribution map, and weightedly fusing the posture transformation parameters and the auxiliary posture information according to the fusion weights to obtain fused posture parameters; Performing coordinate transformation on the scene point cloud according to the fusion pose parameters to obtain a reconstructed scene, calculating the local geometric consistency of the point cloud in the reconstructed scene to obtain a local confidence, calculating the structural similarity of the point clouds at adjacent moments in the reconstructed scene to obtain a global confidence, and performing a weighted combination of the local confidence and the global confidence to obtain a scene confidence score; The pose transformation parameters are calibrated based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0011] In an optional embodiment, The pose transformation parameters are calibrated based on the scene confidence score to obtain the calibrated pose parameter sequence including: Establishing an evaluation vector for each posture according to the scene confidence score, comparing the evaluation vector with a preset vector threshold to obtain a posture credibility index, and marking abnormal postures in postures whose posture credibility index is lower than the preset vector threshold; An influence propagation matrix of abnormal posture is constructed by using the topological relationship graph, the structural correlation strength between the abnormal posture and the adjacent posture is analyzed in the influence propagation matrix, the spatiotemporal influence range of the abnormal posture is determined according to the structural correlation strength, and the spatiotemporal influence range is used as the calibration range; Constructing a pose optimization objective function, the pose optimization objective function includes a structure preservation term, a depth consistency term and a temporal smoothing term, iteratively optimizing the abnormal pose, calculating a structure preservation score of the current pose based on a structure distribution map, calculating a depth consistency score based on a compensated depth map, and calculating a temporal smoothing score based on adjacent poses in each iteration, substituting the structure preservation score, the depth consistency score and the temporal smoothing score into the pose optimization objective function to calculate an optimization direction and a step size; After each round of iterative optimization, the optimized calibration pose is reapplied to the scene reconstruction, the structural features in the scene reconstruction are extracted to calculate the structural consistency, and the optimization is stopped when the structural consistency is greater than a preset convergence threshold; The original pose sequence is updated according to the calibration pose obtained by the final optimization, and the calibrated pose parameter sequence is obtained by chronological order.
[0012] According to a second aspect of the embodiments of the present invention, Provided is an interactive system based on deep learning, comprising: The first unit is used to collect a continuous image sequence obtained by the vehicle-mounted camera and extract feature points, combine the position change of the feature points at adjacent moments with the gradient change of the feature points at adjacent moments to construct a two-dimensional feature vector, cluster the two-dimensional feature vector using a Gaussian mixture model to obtain a stable feature class, calculate a feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature class, and assign weights to the feature points according to the weight distribution function to obtain weighted feature points; The second unit is used to calculate feature similarity based on weighted feature points to obtain a similarity matrix, normalize the similarity matrix, use the normalized similarity matrix to construct a feature matching relationship to obtain an initial matching pair, obtain the vehicle movement speed and calculate the dynamic disparity threshold, screen the disparity of the feature points in the initial matching pair according to the dynamic disparity threshold to obtain a reliable matching feature pair, and construct an optimization equation based on the reliable matching feature pair and solve it to obtain the posture transformation parameters between adjacent images; The third unit is used to input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate the feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with the depth information to generate a compensated depth map, use the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0013] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0014] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0015] In this embodiment, by integrating deep learning and computer vision technology, automatic calibration is achieved, and the accuracy and robustness of pose estimation are improved. A Gaussian mixture model is used to cluster feature points, screen stable feature classes, and optimize feature point matching based on feature weight distribution, improve matching accuracy, and reduce the impact of environmental changes on feature matching. A dynamic parallax threshold calculation method is introduced to enable feature matching screening to adapt to the vehicle motion state, effectively reduce mismatching, and improve the calculation accuracy of pose transformation parameters. Combined with deep learning to extract feature change information, construct a motion feature mapping matrix, and fuse depth information to generate a compensated depth map, achieve adaptive compensation of feature errors, and improve the quality of scene reconstruction. Calibrate pose parameters based on scene confidence scores to reduce cumulative errors and improve long-term operation stability. The present invention not only improves the accuracy of visual pose estimation, but also enhances the adaptability of the system in complex environments. It is suitable for applications such as autonomous driving, intelligent transportation, and high-precision map construction, and provides reliable technical support for high-precision positioning and environmental perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flowchart of an interactive method based on deep learning according to an embodiment of the present invention; Figure 2 A diagram showing the relationship between a dynamic disparity threshold and scene complexity according to an embodiment of the present invention; Figure 3 A comparison chart of 3D reconstruction accuracy under changing scene complexity in an embodiment of the present invention; Figure 4 Schematic diagram of the structure of an interactive system based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0019] Figure 1 Schematic diagram of the flow of the interactive method based on deep learning in an embodiment of the present invention. Figure 1 As shown, the method includes: The continuous image sequence obtained by the vehicle-mounted camera is collected and feature points are extracted. The position change of the feature points at adjacent moments and the gradient change of the feature points at adjacent moments are combined to construct a two-dimensional feature vector. The two-dimensional feature vector is clustered using a Gaussian mixture model to obtain a stable feature class. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The feature points are weighted according to the weight distribution function to obtain weighted feature points. The similarity matrix is obtained by calculating the feature similarity based on the weighted feature points, and the similarity matrix is normalized. The feature matching relationship is constructed using the normalized similarity matrix to obtain an initial matching pair, the vehicle movement speed is obtained and the dynamic disparity threshold is calculated, and the disparity of the feature points in the initial matching pair is screened according to the dynamic disparity threshold to obtain a reliable matching feature pair, and an optimization equation is constructed based on the reliable matching feature pair and solved to obtain the posture transformation parameters between adjacent images; A continuous image sequence is input into the feature detection network to obtain a feature change sequence. The feature change gradient between adjacent frames is calculated according to the feature change sequence. A motion feature mapping matrix is constructed based on the feature change gradient. The motion feature mapping matrix is combined with the depth information to generate a compensated depth map. The compensated depth map and pose transformation parameters are used to reconstruct the scene and calculate the scene confidence score. The pose transformation parameters are calibrated based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0020] Among them, feature points refer to key points in an image, which usually have strong local contrast, such as corner points, edge points or texture feature points. In this scheme, feature points are used to calculate the position change and gradient change at adjacent moments to construct subsequent feature vectors. Gaussian Mixture Model (GMM) is a probabilistic statistical model used to represent the mixture of multiple Gaussian distributions. In this scheme, GMM is used to perform cluster analysis on two-dimensional feature vectors to distinguish between stable feature classes and unstable feature classes, thereby improving the reliability of matching. Stable feature classes refer to feature points that are classified as stable after GMM clustering. These feature points have a low rate of change between adjacent frames and can be more reliably used for pose estimation and matching optimization. The feature weight distribution function is a function calculated based on the distribution characteristics of feature vectors in the stable feature class. This function is used to evaluate the importance of feature points and weight feature points in the feature matching process to improve the accuracy of matching. Weighted feature points refer to feature points that are weighted according to the feature weight distribution function. During the matching process, weighted feature points can enhance the influence of stable feature points, thereby improving the stability of pose estimation.
[0021] The dynamic disparity threshold is a threshold calculated based on the vehicle's speed and is used to screen feature matching pairs. The disparity range is different at different speeds, and the dynamic disparity threshold can adaptively adjust the matching strategy to improve the reliability of matching pairs. The feature detection network is a neural network based on deep learning, which is used to extract feature change sequences from image sequences. The network can adaptively learn the change rules of feature points and improve the robustness of feature extraction. The feature change gradient refers to the rate of change of feature points between adjacent frames, which is used to measure the dynamics of feature points. In this scheme, the gradient information is used to construct the motion feature mapping matrix.
[0022] The compensated depth map is generated by combining the motion feature map matrix and the depth information, and is used to correct the error of the traditional depth map. In this scheme, the compensated depth map can improve the accuracy of scene reconstruction and reduce the cumulative deviation caused by visual errors. The scene confidence score is used to evaluate the reliability of the reconstructed scene, and is usually calculated based on factors such as the consistency of depth information and the quality of feature matching. In this scheme, the scene confidence score is used to calibrate the pose parameters to improve the accuracy of pose estimation.
[0023] The pose transformation parameters are used to describe the camera motion between adjacent image frames, including the rotation matrix and the translation vector. In this scheme, these parameters are obtained by optimization and adjusted in the final calibration process to improve positioning accuracy. The calibrated pose parameter sequence is a set of pose parameters calibrated by the scene confidence score, which can more accurately describe the motion state of the camera in a continuous image sequence. This scheme reduces the cumulative error through this calibration process and improves the stability of long-term pose estimation.
[0024] In summary, the present invention improves the robustness of feature matching, improves the accuracy of pose estimation, and enhances the reliability of scene reconstruction by introducing technologies such as adaptive disparity threshold calculation and depth compensation. It can achieve high-precision automatic calibration in complex environments, and provides strong technical support for applications such as autonomous driving, intelligent transportation, and high-precision map construction.
[0025] In an optional embodiment, The Gaussian mixture model is used to cluster the two-dimensional feature vectors to obtain stable feature classes. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The weighted feature points obtained by weighting the feature points according to the weight distribution function include: Using a Gaussian mixture model to calculate a mixture weight, a mean vector and a covariance matrix of a mixed Gaussian distribution for a two-dimensional feature vector, and using the mixture weight, the mean vector and the covariance matrix to cluster the two-dimensional feature vector to obtain a plurality of feature classes; The average distance from the intra-class feature vector to the class center of the feature class is calculated to obtain the intra-class aggregation degree, the minimum distance between the class centers of the feature classes is calculated to obtain the inter-class separation degree, the feature classes are screened according to the ratio of the intra-class aggregation degree to the inter-class separation degree, and the feature classes that meet the preset feature threshold conditions are determined as stable feature classes; Calculate the intra-class covariance matrix of the two-dimensional feature vector in the stable feature class and perform eigenvalue decomposition to obtain spatial distribution features, calculate the Euclidean distance of the two-dimensional feature vector at adjacent moments and convert it into a time series correlation score, and perform weighted combination of the spatial distribution feature and the time series correlation score to obtain a feature weight distribution function; The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector to obtain the spatial weight, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are fused and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain a weighted feature point.
[0026] Specifically, first collect the image sequence continuously taken by the on-board camera, and extract the key feature points (such as corner points and edge intersections) from them. For the position change (displacement vector) and gradient change (change in pixel gradient direction and amplitude) of the same feature point in two adjacent frames, these two quantities are combined into a two-dimensional feature vector. The Gaussian mixture model is used to model all feature vectors, and the parameters of the mixed Gaussian distribution are estimated by the expectation maximization algorithm, including the mixing weight of each Gaussian component (characterizing the proportion of the component in the overall distribution), the mean vector (the center position of the component), and the covariance matrix (the shape and direction of the component). According to these parameters, the feature vector is divided into multiple feature classes, each class corresponds to a Gaussian component.
[0027] During the driving process, the on-board camera captures the features of the guardrail corner points on the road ahead. Due to the movement of the vehicle, the position of the guardrail corner points in the image is translated (position change), and the local gradient direction changes (gradient change) due to the change in perspective. After combining these changes into a two-dimensional vector, the Gaussian mixture model can be used to distinguish the feature class belonging to the static guardrail and the feature class belonging to the dynamic vehicle.
[0028] Traditional methods (such as K-means clustering) only rely on feature space distance partitioning and cannot model complex distributions. This solution captures the multimodal distribution characteristics of feature vectors through a Gaussian mixture model to more accurately distinguish feature classes of different motion states.
[0029] For each feature class, calculate its intra-class aggregation (the average Euclidean distance from all feature vectors in the class to the class center) and inter-class separation (the minimum distance from the current class center to the center of other classes). If the ratio of intra-class aggregation to inter-class separation of a feature class is lower than the preset threshold (such as 0.3), it is considered that the feature vectors in the class are closely distributed and clearly separated from other classes, and it is marked as a stable feature class. Subsequently, the intra-class covariance matrix is calculated for the feature vectors in the stable feature class, and its principal component direction (eigenvector) and variance ratio (eigenvalue) are decomposed to obtain spatial distribution characteristics (such as whether the distribution direction is consistent with the vehicle movement direction). At the same time, the Euclidean distance of the same feature point in adjacent frames is calculated and mapped to a temporal correlation score from 0 to 1 (the smaller the distance, the higher the score). The spatial distribution characteristics (consistency of the principal component direction) and the temporal correlation score are weighted according to the preset ratio to generate a feature weight distribution function.
[0030] The position change of the feature points of static road signs is small and consistent in direction (along the direction of vehicle movement), and the gradient change is stable (no obvious texture distortion). Therefore, the spatial distribution characteristics are characterized by the main direction of the covariance matrix being aligned with the direction of vehicle movement, and the temporal correlation score is high. The dynamic pedestrian feature points have chaotic covariance directions and low temporal scores due to sudden changes in position. The traditional fixed weight allocation method does not combine the spatial distribution and temporal continuity of features. This scheme dynamically adjusts the weights of feature points by quantifying the consistency of spatial and temporal dimensions to suppress the influence of dynamic interference features.
[0031] Further, based on the feature weight distribution function, the Mahalanobis distance of each feature vector is calculated (to measure its degree of deviation from the class center, taking into account the anisotropy of the covariance matrix), and mapped to a spatial weight through an exponential function (the smaller the deviation, the higher the weight). At the same time, the temporal correlation score is used directly as the temporal weight. The spatial weight and the temporal weight are fused in an adaptive proportion (for example, the temporal weight accounts for a higher proportion in dynamic scenes) to obtain a preliminary fusion weight. Further, the distribution of the fusion weight is statistically analyzed in the image neighborhood of the feature point (such as a 5×5 pixel range), and the local density difference is eliminated through normalization. Finally, the normalized weight is assigned to the corresponding feature point to generate a weighted feature point set.
[0032] In the vehicle turning scene, the position change direction of the static building feature points is consistent with the vehicle's motion trajectory, the Mahalanobis distance is small (high spatial weight), and the displacement of adjacent frames is continuous (high temporal score), so the fusion weight is large; while the false feature points caused by road reflections have a large Mahalanobis distance and low temporal score due to position jumps, and the weight is significantly suppressed.
[0033] This technical solution clusters feature vectors through a Gaussian mixture model and weights feature points in combination with spatial and temporal information, achieving more accurate and stable feature point matching. Compared with traditional clustering methods based on Euclidean distance (such as K-means), the Gaussian mixture model can better capture the multimodal distribution of feature vectors, and then accurately distinguish different motion states, such as static and dynamic targets. By evaluating the intra-class aggregation and inter-class separation of feature classes, stable feature classes are screened out, and noise and unstable features are effectively eliminated, thereby improving the accuracy and robustness of feature matching. Through the weighted combination of spatial distribution features and temporal correlation scores, the calculated feature weight distribution function can dynamically adjust the weight of each feature point, so that static feature points obtain higher weights in matching, while feature points with large dynamic changes are effectively suppressed. Traditional methods usually use fixed weights and fail to fully consider the temporal continuity and spatial consistency of features, resulting in susceptibility to interference in dynamic scenes. This solution comprehensively considers spatial and temporal information to make the matching effect of feature points more accurate in complex and dynamic environments. It can improve the matching stability of feature points in dynamic and complex environments, especially when the perspective changes or there are dynamic interferences in the scene (such as pedestrians and vehicles). Through refined feature point weighting and dynamic adjustment strategies, it ultimately achieves a significant improvement in the accuracy of feature point matching in dynamic scenes, reduces the risk of mismatching, and thus enhances the system's adaptability and robustness in real complex environments.
[0034] In an optional embodiment, The spatial weight is obtained by calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are merged and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain the weighted feature point, which includes: The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector, and the spatial weight is obtained by exponential mapping with the Mahalanobis distance; Calculate the curvature of the motion trajectory of the two-dimensional feature vector as a time scale coefficient, perform nonlinear combination of the time scale coefficient and the Euclidean distance of the two-dimensional feature vector at adjacent moments to obtain a time series correlation score, and calculate a time series weight based on the time series correlation score; Constructing a local variance distribution of the spatial weight, determining an adaptive kernel function based on the local variance distribution, and using the adaptive kernel function to fuse the spatial weight with the temporal weight to obtain a fused weight; Calculating the spatial distribution gradient of the fusion weight, constructing an adaptive neighborhood radius based on the spatial distribution gradient, and combining the adaptive neighborhood radius with a preset neighborhood range through density weighting to obtain a dynamic neighborhood range; Calculate the local covariance characteristics of the fusion weight within the dynamic neighborhood, construct an anisotropic Gaussian kernel function, and use the anisotropic Gaussian kernel function to spatially modulate the fusion weight; normalize the modulated fusion weight within the dynamic neighborhood to obtain an initial weight value, calculate the weight consistency constraint based on the initial weight value, iteratively optimize the weight consistency constraint and the initial weight value to obtain a final weight value, and assign the final weight value to the corresponding two-dimensional feature vector to obtain a weighted feature point.
[0035] This embodiment uses the feature weight distribution function to calculate the Mahalanobis distance for each two-dimensional feature vector. The Mahalanobis distance is a distance metric that measures the deviation between the feature vector and the average position (class center) of the class to which it belongs, and it takes into account the covariance structure of the feature space. By performing exponential mapping on the Mahalanobis distance, the spatial distance of the feature point can be converted into a spatial weight, indicating the importance of the feature point in space. The larger the spatial weight, the smaller the distance of the feature point from its class center, and the more stable its distribution in space. By calculating the curvature of the motion trajectory of each two-dimensional feature vector as the time scale coefficient, combined with the Euclidean distance between the feature points at adjacent moments, the temporal correlation score is obtained. Curvature is a measure of the degree of curvature of the motion trajectory, reflecting the law of motion change of the feature point. After the time scale coefficient is nonlinearly combined with the Euclidean distance, the temporal correlation score is obtained, which reflects the continuity of the feature point over time. The higher the score, the more stable and continuous the motion of the feature point. The temporal weight is calculated based on the temporal correlation score. The temporal weight is used to reflect the importance of the feature point in the time dimension. The feature point with a more stable temporal change will obtain a higher temporal weight.
[0036] Construct the local variance distribution of spatial weights. Local variance is a measure of the distribution change of feature points in their neighborhood, reflecting the stability and consistency of feature points in the local area. Through the local variance distribution, an adaptive kernel function can be determined, which is used to fuse spatial weights and temporal weights. The adaptive kernel function adjusts the degree of weight fusion according to the distribution of local variance, so that in some areas, the fusion of spatial weights and temporal weights will be more refined, while in other areas, different fusion strategies may be adopted. Then, the spatial distribution gradient of the fusion weight is calculated. The spatial distribution gradient is used to measure the change trend of the fusion weight in space, thereby reflecting the spatial structure and change of the feature points. Based on the spatial distribution gradient, an adaptive neighborhood radius is constructed. The adaptive neighborhood radius is dynamically adjusted according to the change of the local spatial distribution, and the neighborhood range of the feature point can be determined more accurately. Through the density-weighted combination with the preset neighborhood range, the dynamic neighborhood range is obtained, which can more flexibly adapt to the distribution characteristics of feature points in different scenarios.
[0037] The local covariance feature of the fusion weight is calculated within the dynamic neighborhood. The local covariance reflects the change pattern of the feature point in its neighborhood. By calculating the local covariance, the stability and change of the feature point can be further evaluated. Then, based on the local covariance feature, an anisotropic Gaussian kernel function is constructed. The anisotropic Gaussian kernel function helps to more accurately capture the distribution characteristics of the feature points in different spatial directions by adjusting the degree of expansion in different directions.
[0038] The fusion weights are spatially modulated using anisotropic Gaussian kernel functions. This process adjusts the weight values in different directions to make the spatial distribution more uniform and more consistent with the distribution pattern in the actual scene. After modulation, the fusion weights will be normalized within the dynamic neighborhood to eliminate the influence caused by spatial density differences and ensure that the weight value of each feature point can be reasonably adjusted within its neighborhood. Finally, the weight consistency constraint is calculated based on the normalized fusion weights. The weight consistency constraint is used to ensure the consistency of the feature point weights throughout the process to avoid weight instability caused by external interference or abnormal changes in feature points. Through the iterative optimization process, the weight consistency constraint is combined with the initial weight value to obtain the final weight value. The final weight value is assigned to the corresponding two-dimensional feature vector to obtain the weighted feature point, ensuring that the weighted result of the feature point has spatial and temporal stability.
[0039] In the autonomous driving scenario, the on-board camera captures landmark feature points on the edge of the road (such as the corners of buildings). Due to the movement of the camera and the vehicle, the image position of these feature points will change, and due to the different viewing angles, the local gradient direction of the feature points will also change. By combining the position change and gradient change of feature points in adjacent frames into a two-dimensional feature vector, and then clustering them through a Gaussian mixture model, it is possible to accurately distinguish between static buildings and false feature points on dynamic roads. Furthermore, based on the fusion of spatial weights and temporal weights, the feature points of static buildings will obtain higher weighted values, while dynamic false feature points caused by reflections and other reasons will be effectively suppressed. Finally, based on these weighted feature points, more accurate image feature matching and scene reconstruction results can be obtained.
[0040] This technical solution combines spatial weights with temporal weights, and uses Mahalanobis distance and temporal correlation scores to weight feature points, which significantly improves the accuracy and robustness of feature points. Existing technologies usually rely on traditional clustering methods to process feature points, but these methods cannot effectively model complex distribution features, resulting in dynamically changing or interfering feature points that may be mistakenly identified as valid features. This solution uses a Gaussian mixture model to fine-tune the modeling of two-dimensional feature vectors, and weights feature points based on the spatial distribution and temporal consistency of the features. This innovation significantly improves the ability to recognize stable feature points, especially in dynamic scenes, and can accurately distinguish between static and dynamic features.
[0041] By dynamically adjusting the weights of feature points, this solution can suppress the influence of false feature points caused by environmental changes (such as lighting changes, moving objects, etc.). Compared with the traditional fixed weight method, this solution can flexibly adapt to different scenarios by quantifying the consistency of spatial and temporal dimensions, thereby improving the accuracy and robustness of feature matching. In complex environments, the dynamic neighborhood range and local covariance calculation further optimize the selection of feature points and ensure high accuracy of image matching. The improved weighted feature points provide more reliable feature matching and scene reconstruction results, which are especially suitable for application scenarios such as autonomous driving and intelligent navigation that require high-precision image recognition. Compared with the existing technology, the adaptability and accuracy of this solution in dynamic scenes have been significantly improved, reducing the risk of false matching and improving the robustness and stability of the system.
[0042] In an optional embodiment, The vehicle speed is obtained and the dynamic parallax threshold is calculated. The parallax of the feature points in the initial matching pair is screened according to the dynamic parallax threshold to obtain reliable matching feature pairs. The optimization equation is constructed based on the reliable matching feature pairs and the pose transformation parameters between adjacent images are obtained by solving the equations. The vehicle speed and the time interval between adjacent frames are obtained, the distribution of feature points in the image area is counted to calculate the image entropy value as the scene structure complexity, and the dynamic parallax threshold is calculated according to the vehicle speed, the time interval between adjacent frames and the scene structure complexity; Calculating the image coordinate difference of the initial matching feature point pairs in adjacent image frames to obtain feature point disparity, taking the initial matching feature point pairs whose feature point disparity is less than the dynamic disparity threshold as candidate matching feature point pairs, and calculating the reference transformation matrix according to the image coordinate correspondence of the candidate matching feature point pairs; Counting other candidate matching feature point pairs within the neighborhood of the candidate matching feature point pair, calculating the difference scores between the coordinate transformation results of other candidate matching feature point pairs and the reference transformation matrix, calculating the consistency support according to the difference scores, and taking the candidate matching feature point pairs whose consistency support is greater than the preset support as reliable matching feature point pairs; A local transformation matrix is calculated according to the image coordinate correspondence of the reliable matching feature point pair, a difference value between the local transformation matrix and the reference transformation matrix is substituted into a Gaussian kernel function to obtain a weight coefficient of the reliable matching feature point pair, and a weighted reprojection error optimization equation is constructed using the correspondence between the three-dimensional space coordinates and the two-dimensional image coordinates of the reliable matching feature point pair and the weight coefficient; A Jacobian matrix is constructed according to the reprojection error optimization equation, a weight matrix is constructed using the weight coefficients, an incremental equation is constructed according to the Jacobian matrix and the weight matrix, the posture update amount of the incremental equation is iteratively calculated and the posture parameters are updated, and when the posture update amount is less than a preset update amount threshold, the posture transformation parameters between adjacent image frames are obtained.
[0043] Exemplarily, the vehicle speed and the time interval between adjacent frames are obtained, and the distribution of feature points in the image area is statistically calculated to calculate the image entropy value as the scene structure complexity. The vehicle speed indicates the displacement of the vehicle per unit time, reflecting the motion state of the vehicle. The time interval between adjacent frames is the time interval between the acquisition of two frames of images. The distribution of feature points refers to key points that can be detected in the image, such as corner points or edge points. The image entropy value is used to measure the complexity of image information. The higher the entropy value, the richer the effective information contained in the image. Based on the vehicle speed, the time interval between adjacent frames and the scene structure complexity, the dynamic parallax threshold is calculated. The parallax threshold determines the screening criteria for feature point matching, and its size is calculated comprehensively by factors such as vehicle speed, camera sampling frequency and scene feature point density.
[0044] The feature point disparity is obtained by calculating the image coordinate difference of the initial matching feature point pairs in adjacent image frames. Feature point disparity refers to the difference in pixel coordinates of the same feature point in two frames, reflecting the impact of camera perspective changes on the position of feature points. Initial matching feature point pairs with disparity less than the dynamic disparity threshold are screened out to obtain candidate matching feature point pairs. Candidate matching feature point pairs refer to feature point pairs that still have a high matching probability after preliminary screening.
[0045] The reference transformation matrix is calculated based on the image coordinate correspondence of the candidate matching feature point pairs. The transformation matrix is used to describe the motion relationship of the camera between adjacent frames, including rotation and translation information. The reference transformation matrix is a preliminary camera pose transformation model estimated based on the candidate matching feature point pairs.
[0046] Count other candidate matching feature point pairs within the neighborhood of the candidate matching feature point pair, and calculate the difference score between their coordinate transformation results and the reference transformation matrix. The difference score is used to measure the degree of deviation of the feature point pair under the transformation matrix. The smaller the deviation, the greater the possibility that the matching point conforms to the global transformation relationship. Calculate the consistency support based on the difference score. The consistency support refers to the credibility of a matching feature point pair among all matching points, reflecting the degree of conformity of the point with the overall motion trend. Filter out the matching point pairs whose consistency support is greater than the preset support threshold as reliable matching feature point pairs.
[0047] The local transformation matrix is calculated based on the image coordinate correspondence of the reliable matching feature point pairs. The local transformation matrix describes the movement changes within the local range of the feature points. The difference between the local transformation matrix and the reference transformation matrix is calculated, and the weight coefficient of the reliable matching feature point pairs is calculated through the Gaussian kernel function. The weight coefficient is used to adjust the influence of different matching feature point pairs in the optimization calculation. The weight of points with high credibility is large, and the weight of points with low credibility is small.
[0048] The correspondence between the 3D space coordinates and the 2D image coordinates of the reliable matching feature point pairs and the calculated weight coefficients are used to construct a weighted reprojection error optimization equation. The reprojection error refers to the deviation between the estimated position of a 3D point on the camera projection plane and the actual observed position. The optimization equation is used to minimize this deviation to improve the accuracy of pose estimation.
[0049] Construct the Jacobian matrix, which is used to describe the rate of change of variables during the optimization process. Use the weight coefficients to construct the weight matrix, which is used to balance the influence of different feature point pairs. Construct the incremental equation based on the Jacobian matrix and the weight matrix, and obtain the pose update amount by iteratively calculating the solution of the incremental equation, and continuously update the pose parameters. The pose update amount is used to describe the deviation between the currently calculated pose parameters and the true pose. When the deviation converges below the preset threshold, the optimization process terminates, and finally the pose transformation parameters between two adjacent frames are obtained.
[0050] In the process of pose estimation, the existing technology usually uses a fixed disparity threshold to screen feature point matching, which does not fully consider the influence of vehicle motion state and scene structure, and easily leads to reduced matching accuracy in dynamic environments. At the same time, when optimizing pose, traditional methods often use a globally consistent weight distribution method, which cannot effectively distinguish high-confidence feature points from low-confidence feature points, thus affecting the final estimation accuracy. This solution introduces a dynamic disparity threshold to adaptively adjust the feature point screening criteria according to the vehicle motion speed, the time interval between adjacent frames and the complexity of the scene structure, making the matching more robust. Compared with the traditional fixed threshold method, this solution can adapt to different motion states and complex scenes and improve the reliability of matching points. In addition, this solution selects reliable matching points by consistent support, and combines the Gaussian kernel function to calculate weights, so that different influences are given to feature points with different confidences during the optimization process, thereby improving the accuracy of pose estimation. This solution aims to enhance the robustness of feature matching and the adaptability of optimization weights to reduce the problem of mismatching in dynamic environments and improve the accuracy of final pose estimation. By optimizing the iterative calculation of the equation, the pose update is made smoother and more accurate, thereby improving the robustness of the visual odometer or positioning system and providing more reliable pose estimation capabilities for autonomous driving or assisted driving systems.
[0051] like Figure 2 As shown in the figure, the dynamic adjustment effect of the disparity threshold under different vehicle speeds and scene complexities is demonstrated. The curve represents the dynamic threshold change trend of this scheme, and the dotted line represents the traditional fixed threshold method. The horizontal axis represents the scene structure complexity (represented by the image entropy value), and the vertical axis is the disparity threshold (unit: pixel). In low-speed scenes (5km / h), due to the small displacement of the vehicle, this scheme sets the disparity threshold to a larger value (about 8 pixels), which allows more potential matching points to participate in the calculation and improves feature utilization. As the vehicle speed increases to medium speed (30km / h), the disparity threshold dynamically decreases to about 5 pixels, because higher speeds require more stringent matching conditions to ensure accuracy. When the vehicle speed reaches high speed (60km / h), the threshold is further reduced to 3 pixels, effectively avoiding mismatching caused by high-speed motion.
[0052] The curve shows that the threshold adjustment is also closely related to the scene complexity. In complex scene areas (high image entropy), the disparity threshold tends to be tightened. This is because complex environments have abundant feature points and require stricter screening criteria. In simple scene areas, the threshold is relatively relaxed to ensure that enough matching points are obtained. This adaptive adjustment mechanism significantly improves the accuracy and robustness of feature matching.
[0053] In an optional embodiment, The vehicle speed and the time interval between adjacent frames are obtained, and the distribution of feature points in the image area is counted to calculate the image entropy value as the scene structure complexity. The dynamic parallax threshold is calculated according to the vehicle speed, the time interval between adjacent frames and the scene structure complexity, including: Obtaining the vehicle speed and the time interval between adjacent frames, calculating the acceleration of the vehicle, classifying the vehicle motion state according to the acceleration, calculating the product of the vehicle speed and the acceleration, and obtaining the motion curvature by the ratio of the product to the cube of the vehicle speed; Construct a multi-scale pyramid for the image, divide the image of each scale into adaptive grid units, count the number of feature points in each grid unit, and calculate the ratio of the number of feature points in each grid unit to the total number of feature points in the image to obtain the distribution probability of feature points; Calculating the information entropy of each scale based on the distribution probability of the feature points, calculating the image gradient direction mean in each grid unit, calculating the directional consistency based on the image gradient direction mean, and combining the information entropy of each scale with the directional consistency as the scene structure complexity of the current scale; Determining a weight coefficient for each scale according to the difference between the scene structure complexity and the mean value of the scene structure complexity, and weighting the scene structure complexity across scales to obtain a comprehensive scene structure complexity; A motion response value is calculated according to the vehicle motion speed and the motion curvature, and a dynamic parallax threshold is obtained by combining the comprehensive scene structure complexity with the motion response value.
[0054] For example, after obtaining the vehicle's speed and the time interval between adjacent frames, the acceleration of the vehicle is calculated. Acceleration represents the rate of change of speed and can be used to determine whether the vehicle is moving at a constant speed, accelerating or decelerating. According to the magnitude and direction of the acceleration, the vehicle's motion state is determined, such as stationary, slow acceleration, rapid acceleration, slow deceleration or rapid deceleration.
[0055] The product of the vehicle's speed and acceleration is calculated, and the ratio is calculated with the cube of the vehicle's speed to obtain the curvature of motion. Curvature of motion is an important indicator to measure the degree of change in the vehicle's motion trajectory, and can reflect the smoothness and curvature characteristics of the vehicle's driving path.
[0056] Construct a multi-scale pyramid for the image. The pyramid structure means that the original image is reduced multiple times to form multiple image levels with different resolutions, which helps to extract image features at different scales. In each scale of the image, the image is divided into adaptive grid units, and the size of each grid unit is dynamically adjusted according to the density of the image feature points to ensure that the feature points in the local area are evenly distributed.
[0057] The number of feature points in each grid unit is counted, and the ratio of the feature points in the total number of feature points in the entire image is calculated. This ratio is called the feature point distribution probability. The feature point distribution probability is used to measure the density of feature points in different areas of the image, thereby reflecting the local complexity of the image information.
[0058] Based on the probability of feature point distribution, the information entropy of each scale is calculated. Information entropy is an indicator to measure the complexity of image content, indicating the uniformity of the distribution of feature points in the image. The higher the information entropy, the more complex the distribution of image feature points. The mean value of the image gradient direction in each grid cell is further calculated. The gradient direction indicates the direction of image brightness change. The mean value of the gradient direction can be used to measure the directional characteristics of image edges and textures.
[0059] Based on the mean value of the image gradient direction, the directional consistency is calculated. The directional consistency indicates whether the directions of the edges and feature points in the local area of the image tend to be consistent. If the directional consistency is high, it means that the feature points in the area are arranged more regularly, otherwise it means that the distribution of the feature points in the area is more chaotic. The information entropy of each scale is combined with the directional consistency as the scene structure complexity of the current scale.
[0060] The weight coefficient of each scale is determined based on the difference between the scene structure complexity of each scale and the mean of the scene structure complexity of all scales. The weight coefficient is used to adjust the influence of different scales so that the scale that better represents the scene complexity gets a higher weight. The scene structure complexity of all scales is weighted and calculated to obtain the comprehensive scene structure complexity.
[0061] The motion response value is calculated based on the vehicle's motion speed and motion curvature. The motion response value is an important parameter to measure the change in the vehicle's current motion state and can reflect the dynamic characteristics of the vehicle. Finally, the comprehensive scene structure complexity is combined with the motion response value to obtain the dynamic disparity threshold. The dynamic disparity threshold is used to adjust the feature point screening criteria in the image matching process, so that feature matching can adapt to different motion states and scene complexities, and improve the stability and accuracy of matching.
[0062] In practical applications, assuming that the vehicle is traveling at a high speed and the acceleration changes greatly in a short period of time, the calculated motion curvature is large, indicating that the vehicle may be making a sharp turn or accelerating sharply. In this case, in order to improve the accuracy of feature point matching, it is necessary to adjust the dynamic disparity threshold to adapt it to the rapidly changing disparity features. At the same time, the image taken by the front camera is obtained, and a multi-scale pyramid is constructed to reduce the original image to different scales and divide multiple grid cells at each scale. The number of feature points in each grid cell is counted. For example, at a certain scale, the density of feature points in the left area is higher, while the feature points in the right area are sparse, then the calculated feature point distribution probability is larger on the left. The information entropy of each grid cell is further calculated. If the feature points in some areas of the image are dense and evenly distributed, the information entropy is high, indicating that the area contains rich texture information. In addition, the mean of the image gradient direction is calculated. If the gradient direction of a certain area tends to be consistent, the direction consistency of the area is high.
[0063] Then, the information entropy and directional consistency of different scales are combined to calculate the scene structure complexity of each scale, and the influence of each scale is adjusted by the weight coefficient to finally obtain the comprehensive scene structure complexity. The comprehensive scene structure complexity is combined with the motion response value to obtain the dynamic disparity threshold of the current frame image. For example, at night or in low light conditions, the overall feature points of the image are fewer, the calculated information entropy is lower, and the directional consistency may be higher. Therefore, the calculated dynamic disparity threshold is adjusted accordingly, so that the feature point matching can better adapt to the impact of lighting changes and ensure the accuracy of pose estimation. When the vehicle is driving at high speed or making sharp turns, the dynamic disparity threshold will also be adjusted accordingly to adapt to the parallax changes and improve the robustness of the matching.
[0064] When calculating the disparity threshold, the prior art usually only considers the vehicle's motion speed or simple image features, but does not fully combine the vehicle's motion state and the complexity of the scene structure, resulting in poor adaptability of the dynamic disparity threshold, which affects the matching accuracy. The present application introduces the vehicle's motion acceleration and calculates the motion curvature to further refine the classification of the vehicle's motion state, so that the dynamic disparity threshold can more accurately adapt to different motion situations. At the same time, a multi-scale pyramid is used to construct an image hierarchy, and adaptive grids are divided at different scales to calculate the distribution probability and information entropy of feature points, thereby more accurately describing the complexity of the scene structure. Compared with the traditional method that simply relies on single-scale features or global statistical information, this method can more finely characterize the distribution of image features and improve the calculation accuracy of the disparity threshold. In addition, by combining the motion response value with the comprehensive scene structure complexity for threshold calculation, the dynamic disparity threshold can be adaptively adjusted under different motion states and scene changes, thereby enhancing the robustness of the algorithm. Finally, the scheme improves the screening quality of matching feature points and the accuracy of camera pose estimation, so that the system can obtain a more stable matching effect in complex scenes and different motion states.
[0065] In an optional embodiment, Combining the motion feature map matrix with the depth information to generate a compensated depth map, using the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score includes: Combining the motion feature mapping matrix with the depth information, performing propagation compensation on the depth information according to the motion information in the motion feature mapping matrix to obtain a compensated depth map, and performing three-dimensional back-projection on the compensated depth map to obtain a scene point cloud; Extracting plane structural elements and edge structural elements from the scene point cloud, calculating spatial distribution characteristics of the plane structural elements and edge structural elements to obtain a structural distribution map, and constructing a topological relationship map of the structural elements based on the structural distribution map; Decomposing the posture transformation parameters into rotation components and translation components, obtaining auxiliary posture information of the visual odometer, calculating the fusion weights of the posture transformation parameters and the auxiliary posture information according to the structural distribution map, and weightedly fusing the posture transformation parameters and the auxiliary posture information according to the fusion weights to obtain fused posture parameters; Performing coordinate transformation on the scene point cloud according to the fusion pose parameters to obtain a reconstructed scene, calculating the local geometric consistency of the point cloud in the reconstructed scene to obtain a local confidence, calculating the structural similarity of the point clouds at adjacent moments in the reconstructed scene to obtain a global confidence, and performing a weighted combination of the local confidence and the global confidence to obtain a scene confidence score; The pose transformation parameters are calibrated based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0066] Exemplarily, a motion feature mapping matrix is combined with depth information to obtain more accurate scene depth data. The motion feature mapping matrix is used to describe the motion changes between image frames, and contains the displacement information of feature points in the image coordinate system. The matrix can be calculated by the optical flow method or the feature matching method. The depth information comes from sensors (such as lidar, binocular cameras, or structured light depth cameras), which indicates the distance from the object in the scene to the camera. By using the motion feature mapping matrix to propagate and compensate for the depth information, the depth offset caused by camera motion is corrected to obtain a compensated depth map.
[0067] Next, the compensated depth map is back-projected in three dimensions to generate a scene point cloud. Back-projection is the process of transforming points in the image coordinate system to the world coordinate system. The specific method is to use the intrinsic and extrinsic parameters of the camera to convert the depth value of each pixel into a three-dimensional space coordinate to form a scene point cloud.
[0068] Extract plane structural elements and edge structural elements from point cloud data to analyze the geometric characteristics of the scene. Plane structural elements are a set of points formed by the plane part of an object, which can be detected by the RANSAC (random sampling consensus) method; edge structural elements represent areas where the contour or surface of an object changes dramatically, and can be extracted by gradient change analysis or methods based on normal vector calculation. Calculate the spatial distribution characteristics of these elements to form a structural distribution map, which is used to represent the distribution of different structures in three-dimensional space.
[0069] A topological relationship graph is constructed based on the structural distribution graph to describe the connection relationship between different structural elements in the point cloud data. The nodes of the topological relationship graph represent plane or edge elements, and the edges represent the spatial relationship between them, such as adjacent or coplanar relationships.
[0070] Decompose the pose transformation parameters into rotation and translation components to more accurately analyze the camera's motion state. The rotation component describes the angle change of the camera, and the translation component represents the displacement of the camera. Obtain auxiliary pose information for the visual odometer. The visual odometer estimates the relative motion of the camera by tracking feature points in the image. Common methods include optical flow tracking, feature point matching, and direct methods.
[0071] The fusion weights of the pose transformation parameters and the auxiliary pose information are calculated according to the structural distribution graph to ensure that the fused pose data is more reliable. The fusion weight represents the credibility of different data sources and is usually determined by historical data, sensor characteristics, or adaptive optimization methods. The pose transformation parameters and the auxiliary pose information are weightedly fused according to the fusion weight to obtain the fused pose parameters.
[0072] The scene point cloud is transformed according to the fusion pose parameters to obtain the reconstructed scene in a unified coordinate system. Coordinate transformation is to convert all point cloud data into the same reference coordinate system to eliminate the impact of camera motion on data alignment.
[0073] Calculate the local geometric consistency of the point cloud in the reconstructed scene to measure the degree of match between different frame data. Local geometric consistency is calculated based on the spatial distribution of neighborhood points and is used to measure whether the local structure of the point cloud is consistent. For example, it can be evaluated by calculating the normal vector similarity of the point cloud or the change in the distance from the point to the surface.
[0074] The structural similarity of point clouds at adjacent moments in the reconstructed scene is calculated to evaluate the global matching. Structural similarity is usually based on the overall morphology of the point cloud, such as histogram description, point cloud density distribution, or characteristic curvature analysis.
[0075] The local geometric consistency and global structural similarity are weighted together to obtain the scene confidence score. The confidence score indicates the reliability of the point cloud alignment, and a high confidence score means that the pose estimation is more accurate.
[0076] The pose transformation parameters are calibrated based on the scene confidence score to optimize the pose sequence between adjacent frames, and finally a calibrated pose parameter sequence is obtained. The pose calibration improves the accuracy and stability of the system by adjusting the initial estimate to make it more consistent with the actual motion trajectory.
[0077] The existing technology usually uses a fixed threshold or a simple filtering method to compensate for depth information, which is difficult to accurately adapt to changes in different motion states and scene structures, resulting in unstable depth compensation effects and affecting the accuracy of three-dimensional reconstruction. At the same time, in the pose calculation process, traditional methods often rely on a single feature or simple weighted fusion, which makes it difficult to fully utilize structural information for pose optimization, resulting in large cumulative errors. This application constructs a motion feature mapping matrix and combines it with depth information for propagation compensation, so that the depth information can adapt to dynamic motion environments and improve the accuracy of compensation. By constructing a scene point cloud through three-dimensional back projection and extracting plane and edge structure elements, the scene structure features can be more comprehensively characterized, and the geometric consistency can be enhanced through topological relationships to improve the quality of the point cloud.
[0078] In terms of pose calculation, this application decomposes the pose transformation parameters, combines them with auxiliary pose information, optimizes them through adaptive weight fusion, and fully utilizes the scene structure information to constrain the pose calculation and reduce the cumulative error. At the same time, the scene confidence is calculated based on local geometric consistency and global structural similarity, and the pose parameters are calibrated using this confidence, thereby improving the accuracy and stability of pose estimation. Compared with the prior art, the starting point of the improvement of this application is to combine motion features, scene structure and multi-source pose information, optimize depth compensation and pose calculation, make the three-dimensional reconstruction results more accurate, reduce pose estimation errors, and improve adaptability in complex environments.
[0079] like Figure 3 As shown in the figure, the comparison results of 3D reconstruction accuracy of different methods under the condition of changing scene complexity are shown. This technical solution shows obvious advantages under each scene complexity: when the scene complexity is 0.1, the reconstruction accuracy reaches 0.42mm, while the accuracy of traditional optical flow method, feature point method and ORB-SLAM are approximately 0.48mm, 0.52mm and 0.56mm respectively; as the scene complexity increases to 0.9, the reconstruction accuracy of this technical solution remains at 0.89mm, which is significantly better than other methods (0.82mm, 0.78mm and 0.75mm respectively). Especially in medium-complex scenes with a scene complexity of 0.3-0.7, the accuracy of this technical solution is particularly significant, and the reconstruction accuracy is improved by about 25% on average compared with other methods, which reflects the robustness and adaptability of this solution in complex scenes.
[0080] In an optional embodiment, The pose transformation parameters are calibrated based on the scene confidence score to obtain the calibrated pose parameter sequence including: Establishing an evaluation vector for each posture according to the scene confidence score, comparing the evaluation vector with a preset vector threshold to obtain a posture credibility index, and marking abnormal postures in postures whose posture credibility index is lower than the preset vector threshold; An influence propagation matrix of abnormal posture is constructed by using the topological relationship graph, the structural correlation strength between the abnormal posture and the adjacent posture is analyzed in the influence propagation matrix, the spatiotemporal influence range of the abnormal posture is determined according to the structural correlation strength, and the spatiotemporal influence range is used as the calibration range; Constructing a pose optimization objective function, the pose optimization objective function includes a structure preservation term, a depth consistency term and a temporal smoothing term, iteratively optimizing the abnormal pose, calculating a structure preservation score of the current pose based on a structure distribution map, calculating a depth consistency score based on a compensated depth map, and calculating a temporal smoothing score based on adjacent poses in each iteration, substituting the structure preservation score, the depth consistency score and the temporal smoothing score into the pose optimization objective function to calculate an optimization direction and a step size; After each round of iterative optimization, the optimized calibration pose is reapplied to the scene reconstruction, the structural features in the scene reconstruction are extracted to calculate the structural consistency, and the optimization is stopped when the structural consistency is greater than a preset convergence threshold; The original pose sequence is updated according to the calibration pose obtained by the final optimization, and the calibrated pose parameter sequence is obtained by chronological order.
[0081] Exemplarily, a pose evaluation index is first constructed, and the reliability of each pose is evaluated by the scene confidence score. The scene confidence score includes two aspects: scene reconstruction quality and feature matching reliability: the scene reconstruction quality is evaluated by the density and distribution uniformity of the reconstructed point cloud; the feature matching reliability is evaluated by the reprojection error of the matching feature pairs and the similarity of the feature descriptors. The scene confidence score is converted into a multidimensional evaluation vector, and the dimensions of the evaluation vector include the density of the reconstructed point cloud, the distribution uniformity, the reprojection error, and the descriptor similarity. The preset vector threshold is obtained based on a large amount of experimental data statistics, which represents the normal range of indicators in each dimension. When any dimension of the evaluation vector is lower than the corresponding threshold, the pose is marked as an abnormal pose.
[0082] The topological relationship graph is used to describe the spatiotemporal relationship between postures, where nodes represent postures and edges represent the connection relationship between postures. The influence propagation matrix describes the degree of influence of abnormal postures on other postures, and the values of matrix elements are determined by the temporal distance and spatial distance between postures. The structural association strength is calculated by the number of common view feature points and the structural overlap of the reconstructed scene. The method for determining the spatiotemporal influence range is: starting from the abnormal posture, gradually expanding to the adjacent posture, and stopping the expansion when the structural association strength is lower than the preset threshold. The final range is the posture range that needs to be calibrated.
[0083] The pose optimization objective function contains three constraints: the structure preservation term ensures the consistency of the scene structure features, the depth consistency term ensures the continuity of the scene depth information, and the temporal smoothness term constrains the smooth changes of adjacent poses. The structure preservation term is calculated through the structure distribution map, which records the main structural lines and plane features in the scene. The depth consistency term is calculated based on the compensated depth map, which is generated by the feature change sequence and depth information. The temporal smoothness term is calculated by the relative motion of adjacent poses.
[0084] The iterative optimization process includes: first, reconstructing the scene according to the current pose, extracting the structural features in the scene and comparing them with the structural distribution map, and calculating the structure preservation score. Then, the depth consistency score is calculated using the compensated depth map, and the temporal smoothness score is calculated based on the relative motion of adjacent poses. The three scores are substituted into the objective function, and the optimization direction and step size are determined by the gradient descent method. The optimization direction is the opposite direction of the gradient direction of the objective function at the current pose, and the step size is determined by linear search.
[0085] After each iteration, the scene is reconstructed using the optimized pose and structural features are extracted. Structural consistency is calculated by comparing structural features between adjacent frames, including the directional consistency of structural lines and the normal vector consistency of planar features. When the structural consistency is higher than the preset convergence threshold for multiple consecutive rounds, the optimization is considered to have reached a convergence state and the iteration process is stopped.
[0086] The final optimized calibration pose is used to replace the corresponding abnormal poses in the original pose sequence in chronological order, while keeping other poses unchanged, and finally obtaining a complete pose parameter sequence after calibration.
[0087] In a continuous image sequence of an urban road scene, a large number of building contours and ground marking features were extracted through the feature detection network. At a certain moment, due to building occlusion, the feature matching was abnormal. At this pose, the density of the reconstructed point cloud was detected to be reduced to one-third of the normal value, and the feature point reprojection error increased to twice the normal value. Through topological relationship graph analysis, it was found that this abnormal pose had a strong structural correlation with the poses of the two frames before and after, and multiple building edge lines and ground markings were observed together. During the optimization process, the focus was on maintaining the continuity of these structural lines while ensuring a smooth transition of the depth map. After multiple rounds of iterative optimization, the directional consistency of the structural lines was significantly improved, and the jumps in the depth map were significantly reduced, and finally a calibrated continuous pose sequence was obtained.
[0088] Figure 4 Schematic diagram of the structure of the interactive system based on deep learning according to an embodiment of the present invention. Figure 4 As shown, the system comprises: The first unit is used to collect a continuous image sequence obtained by the vehicle-mounted camera and extract feature points, combine the position change of the feature points at adjacent moments with the gradient change of the feature points at adjacent moments to construct a two-dimensional feature vector, cluster the two-dimensional feature vector using a Gaussian mixture model to obtain a stable feature class, calculate a feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature class, and assign weights to the feature points according to the weight distribution function to obtain weighted feature points; The second unit is used to calculate feature similarity based on weighted feature points to obtain a similarity matrix, normalize the similarity matrix, use the normalized similarity matrix to construct a feature matching relationship to obtain an initial matching pair, obtain the vehicle movement speed and calculate the dynamic disparity threshold, screen the disparity of the feature points in the initial matching pair according to the dynamic disparity threshold to obtain a reliable matching feature pair, and construct an optimization equation based on the reliable matching feature pair and solve it to obtain the posture transformation parameters between adjacent images; The third unit is used to input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate the feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with the depth information to generate a compensated depth map, use the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence.
[0089] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0090] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0091] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The interactive method based on deep learning is characterized by: include: The continuous image sequence obtained by the vehicle-mounted camera is collected and feature points are extracted. The position change of the feature points at adjacent moments and the gradient change of the feature points at adjacent moments are combined to construct a two-dimensional feature vector. The two-dimensional feature vector is clustered using a Gaussian mixture model to obtain a stable feature class. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The feature points are weighted according to the weight distribution function to obtain weighted feature points. The similarity matrix is obtained by calculating the feature similarity based on the weighted feature points, and the similarity matrix is normalized. The feature matching relationship is constructed using the normalized similarity matrix to obtain an initial matching pair, the vehicle movement speed is obtained and the dynamic disparity threshold is calculated, and the disparity of the feature points in the initial matching pair is screened according to the dynamic disparity threshold to obtain a reliable matching feature pair, and an optimization equation is constructed based on the reliable matching feature pair and solved to obtain the posture transformation parameters between adjacent images; A continuous image sequence is input into the feature detection network to obtain a feature change sequence. The feature change gradient between adjacent frames is calculated according to the feature change sequence. A motion feature mapping matrix is constructed based on the feature change gradient. The motion feature mapping matrix is combined with the depth information to generate a compensated depth map. The compensated depth map and pose transformation parameters are used to reconstruct the scene and calculate the scene confidence score. The pose transformation parameters are calibrated based on the scene confidence score to obtain a calibrated pose parameter sequence.
2. The method according to claim 1, characterized in that The Gaussian mixture model is used to cluster the two-dimensional feature vectors to obtain stable feature classes. The feature weight distribution function is calculated based on the distribution characteristics of the feature vectors in the stable feature class. The weighted feature points obtained by weighting the feature points according to the weight distribution function include: Using a Gaussian mixture model to calculate a mixture weight, a mean vector and a covariance matrix of a mixed Gaussian distribution for a two-dimensional feature vector, and using the mixture weight, the mean vector and the covariance matrix to cluster the two-dimensional feature vector to obtain a plurality of feature classes; The average distance from the intra-class feature vector to the class center of the feature class is calculated to obtain the intra-class aggregation degree, the minimum distance between the class centers of the feature classes is calculated to obtain the inter-class separation degree, the feature classes are screened according to the ratio of the intra-class aggregation degree to the inter-class separation degree, and the feature classes that meet the preset feature threshold conditions are determined as stable feature classes; Calculate the intra-class covariance matrix of the two-dimensional feature vector in the stable feature class and perform eigenvalue decomposition to obtain spatial distribution features, calculate the Euclidean distance of the two-dimensional feature vector at adjacent moments and convert it into a time series correlation score, and perform weighted combination of the spatial distribution feature and the time series correlation score to obtain a feature weight distribution function; The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector to obtain the spatial weight, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are fused and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain a weighted feature point.
3. The method according to claim 2, characterized in that The spatial weight is obtained by calculating the Mahalanobis distance of the two-dimensional feature vector using the feature weight distribution function, and the temporal weight is obtained by combining the temporal correlation score. The spatial weight and the temporal weight are merged and normalized within a preset neighborhood range, and the normalized weight value is assigned to the corresponding feature point to obtain the weighted feature point, which includes: The feature weight distribution function is used to calculate the Mahalanobis distance of the two-dimensional feature vector, and the spatial weight is obtained by exponential mapping with the Mahalanobis distance; Calculate the curvature of the motion trajectory of the two-dimensional feature vector as a time scale coefficient, perform nonlinear combination of the time scale coefficient and the Euclidean distance of the two-dimensional feature vector at adjacent moments to obtain a time series correlation score, and calculate a time series weight based on the time series correlation score; Constructing a local variance distribution of the spatial weight, determining an adaptive kernel function based on the local variance distribution, and using the adaptive kernel function to fuse the spatial weight with the temporal weight to obtain a fused weight; Calculating the spatial distribution gradient of the fusion weight, constructing an adaptive neighborhood radius based on the spatial distribution gradient, and combining the adaptive neighborhood radius with a preset neighborhood range through density weighting to obtain a dynamic neighborhood range; Calculate the local covariance characteristics of the fusion weight within the dynamic neighborhood, construct an anisotropic Gaussian kernel function, and use the anisotropic Gaussian kernel function to spatially modulate the fusion weight; normalize the modulated fusion weight within the dynamic neighborhood to obtain an initial weight value, calculate the weight consistency constraint based on the initial weight value, iteratively optimize the weight consistency constraint and the initial weight value to obtain a final weight value, and assign the final weight value to the corresponding two-dimensional feature vector to obtain a weighted feature point.
4. The method according to claim 1, characterized in that: The vehicle speed is obtained and the dynamic parallax threshold is calculated. The parallax of the feature points in the initial matching pair is screened according to the dynamic parallax threshold to obtain reliable matching feature pairs. The optimization equation is constructed based on the reliable matching feature pairs and the pose transformation parameters between adjacent images are obtained by solving the equations. The vehicle speed and the time interval between adjacent frames are obtained, the distribution of feature points in the image area is counted to calculate the image entropy value as the scene structure complexity, and the dynamic parallax threshold is calculated according to the vehicle speed, the time interval between adjacent frames and the scene structure complexity; Calculating the image coordinate difference of the initial matching feature point pairs in adjacent image frames to obtain feature point disparity, taking the initial matching feature point pairs whose feature point disparity is less than the dynamic disparity threshold as candidate matching feature point pairs, and calculating the reference transformation matrix according to the image coordinate correspondence of the candidate matching feature point pairs; Counting other candidate matching feature point pairs within the neighborhood of the candidate matching feature point pair, calculating the difference scores between the coordinate transformation results of other candidate matching feature point pairs and the reference transformation matrix, calculating the consistency support according to the difference scores, and taking the candidate matching feature point pairs whose consistency support is greater than the preset support as reliable matching feature point pairs; A local transformation matrix is calculated according to the image coordinate correspondence of the reliable matching feature point pair, a difference value between the local transformation matrix and the reference transformation matrix is substituted into a Gaussian kernel function to obtain a weight coefficient of the reliable matching feature point pair, and a weighted reprojection error optimization equation is constructed using the correspondence between the three-dimensional space coordinates and the two-dimensional image coordinates of the reliable matching feature point pair and the weight coefficient; A Jacobian matrix is constructed according to the reprojection error optimization equation, a weight matrix is constructed using the weight coefficients, an incremental equation is constructed according to the Jacobian matrix and the weight matrix, the posture update amount of the incremental equation is iteratively calculated and the posture parameters are updated, and when the posture update amount is less than a preset update amount threshold, the posture transformation parameters between adjacent image frames are obtained.
5. The method according to claim 4, characterized in that The vehicle speed and the time interval between adjacent frames are obtained, and the distribution of feature points in the image area is counted to calculate the image entropy value as the scene structure complexity. The dynamic parallax threshold is calculated according to the vehicle speed, the time interval between adjacent frames and the scene structure complexity, including: Obtaining the vehicle speed and the time interval between adjacent frames, calculating the acceleration of the vehicle, classifying the vehicle motion state according to the acceleration, calculating the product of the vehicle speed and the acceleration, and obtaining the motion curvature by the ratio of the product to the cube of the vehicle speed; Construct a multi-scale pyramid for the image, divide the image of each scale into adaptive grid units, count the number of feature points in each grid unit, and calculate the ratio of the number of feature points in each grid unit to the total number of feature points in the image to obtain the distribution probability of feature points; Calculating the information entropy of each scale based on the distribution probability of the feature points, calculating the image gradient direction mean in each grid unit, calculating the directional consistency based on the image gradient direction mean, and combining the information entropy of each scale with the directional consistency as the scene structure complexity of the current scale; Determining a weight coefficient for each scale according to the difference between the scene structure complexity and the mean value of the scene structure complexity, and weighting the scene structure complexity across scales to obtain a comprehensive scene structure complexity; A motion response value is calculated according to the vehicle motion speed and the motion curvature, and a dynamic parallax threshold is obtained by combining the comprehensive scene structure complexity with the motion response value.
6. The method according to claim 1, characterized in that Combining the motion feature map matrix with the depth information to generate a compensated depth map, using the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score includes: Combining the motion feature mapping matrix with the depth information, performing propagation compensation on the depth information according to the motion information in the motion feature mapping matrix to obtain a compensated depth map, and performing three-dimensional back-projection on the compensated depth map to obtain a scene point cloud; Extracting plane structural elements and edge structural elements from the scene point cloud, calculating spatial distribution characteristics of the plane structural elements and edge structural elements to obtain a structural distribution map, and constructing a topological relationship map of the structural elements based on the structural distribution map; Decomposing the posture transformation parameters into rotation components and translation components, obtaining auxiliary posture information of the visual odometer, calculating the fusion weights of the posture transformation parameters and the auxiliary posture information according to the structural distribution map, and weightedly fusing the posture transformation parameters and the auxiliary posture information according to the fusion weights to obtain fused posture parameters; Performing coordinate transformation on the scene point cloud according to the fusion pose parameters to obtain a reconstructed scene, calculating the local geometric consistency of the point cloud in the reconstructed scene to obtain a local confidence, calculating the structural similarity of the point clouds at adjacent moments in the reconstructed scene to obtain a global confidence, and performing a weighted combination of the local confidence and the global confidence to obtain a scene confidence score; The pose transformation parameters are calibrated based on the scene confidence score to obtain a calibrated pose parameter sequence.
7. The method according to claim 6, characterized in that The pose transformation parameters are calibrated based on the scene confidence score to obtain the calibrated pose parameter sequence including: Establishing an evaluation vector for each posture according to the scene confidence score, comparing the evaluation vector with a preset vector threshold to obtain a posture credibility index, and marking abnormal postures in postures whose posture credibility index is lower than the preset vector threshold; An influence propagation matrix of abnormal posture is constructed by using the topological relationship graph, the structural correlation strength between the abnormal posture and the adjacent posture is analyzed in the influence propagation matrix, the spatiotemporal influence range of the abnormal posture is determined according to the structural correlation strength, and the spatiotemporal influence range is used as the calibration range; Constructing a pose optimization objective function, the pose optimization objective function includes a structure preservation term, a depth consistency term and a temporal smoothing term, iteratively optimizing the abnormal pose, calculating a structure preservation score of the current pose based on a structure distribution map, calculating a depth consistency score based on a compensated depth map, and calculating a temporal smoothing score based on adjacent poses in each iteration, substituting the structure preservation score, the depth consistency score and the temporal smoothing score into the pose optimization objective function to calculate an optimization direction and a step size; After each round of iterative optimization, the optimized calibration pose is reapplied to the scene reconstruction, the structural features in the scene reconstruction are extracted to calculate the structural consistency, and the optimization is stopped when the structural consistency is greater than a preset convergence threshold; The original pose sequence is updated according to the calibration pose obtained by the final optimization, and the calibrated pose parameter sequence is obtained by chronological order.
8. An interactive system based on deep learning, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to collect a continuous image sequence obtained by the vehicle-mounted camera and extract feature points, combine the position change of the feature points at adjacent moments with the gradient change of the feature points at adjacent moments to construct a two-dimensional feature vector, cluster the two-dimensional feature vector using a Gaussian mixture model to obtain a stable feature class, calculate a feature weight distribution function based on the distribution characteristics of the feature vectors in the stable feature class, and assign weights to the feature points according to the weight distribution function to obtain weighted feature points; The second unit is used to calculate feature similarity based on weighted feature points to obtain a similarity matrix, normalize the similarity matrix, use the normalized similarity matrix to construct a feature matching relationship to obtain an initial matching pair, obtain the vehicle movement speed and calculate the dynamic disparity threshold, screen the disparity of the feature points in the initial matching pair according to the dynamic disparity threshold to obtain a reliable matching feature pair, and construct an optimization equation based on the reliable matching feature pair and solve it to obtain the posture transformation parameters between adjacent images; The third unit is used to input a continuous image sequence into a feature detection network to obtain a feature change sequence, calculate the feature change gradient between adjacent frames according to the feature change sequence, construct a motion feature mapping matrix based on the feature change gradient, combine the motion feature mapping matrix with the depth information to generate a compensated depth map, use the compensated depth map and pose transformation parameters to reconstruct the scene and calculate the scene confidence score, and calibrate the pose transformation parameters based on the scene confidence score to obtain a calibrated pose parameter sequence.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Visual positioning method for complex road environment in automatic driving scene
CN117710930A
Road scene point cloud identification method and system suitable for inspection robot
CN118887641A
Visual-inertial odometry method and apparatus, electronic device, storage medium and computer program
WO2023051019A1
Cited By
Adaptive matrix block weighted distributed optical fiber strain measurement method
CN120234548A
Target cooperative control method and system based on deep learning
CN120370719A
Mixed shooting image correction method and system based on deep learning
CN120765486A
Intelligent analysis method for test data of intelligent network connection automobile test bench
CN120832538A
An intelligent analysis method for intelligent networked vehicle test bench test data
CN120832538B