Visual feature detection and matching method fusing ORB and SIFT
By integrating the visual feature detection and matching methods of ORB and SIFT, combining the FAST operator and smoothing weight function, dynamically adjusting the descriptor ratio, and adaptively adjusting the matching threshold, the matching accuracy and real-time performance problems of the ORB algorithm in complex scenes are solved, and efficient and stable feature matching is achieved.
Patent Information
- Application Number
- CN202511203473.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-27
AI Technical Summary
In existing technologies, the ORB algorithm's matching accuracy decreases when processing large scale changes, obvious lighting changes, or image parallax, making it difficult to find a balance between real-time performance and stability. Traditional single-feature matching algorithms are difficult to balance speed and accuracy in binocular vision systems.
A visual feature detection and matching method that integrates ORB and SIFT is proposed. DoG extreme point detection is performed by combining the response significance of the FAST operator. The ratio of SIFT and ORB descriptors is dynamically adjusted using a smoothly changing weight function. The matching threshold is adaptively adjusted based on local texture complexity. Geometric consistency check and sliding window filter are introduced to optimize key point trajectories.
It improves the efficiency and stability of feature detection, and enhances the accuracy and robustness of matching. It is suitable for binocular vision systems in complex scenes, adapts to changes in image complexity, and maintains real-time performance without increasing system latency.
Smart Images

Figure CN120689373A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision technology, and in particular to a visual feature detection and matching method integrating ORB and SIFT. Background Art
[0002] In the field of machine vision, in order to identify objects, features need to be extracted from image pixels. With the rapid development of computer vision and artificial intelligence technologies, the extraction and matching of image features has become the foundation of many visual tasks, playing a key role in tasks such as 3D reconstruction, target recognition, stereo matching, and visual navigation. Currently, widely used image feature detection and description algorithms include SIFT, SURF, ORB, BRIEF, AKAZE, etc. Among them, SIFT is considered a high-precision solution for image matching due to its excellent scale invariance, rotation invariance, and robustness to illumination changes. However, the algorithm requires Gaussian pyramid construction, gradient direction calculation, and descriptor generation during the key point detection and description process, resulting in a high overall computational complexity and unsuitable for embedded or mobile platforms with high real-time requirements.
[0003] In contrast, the ORB algorithm achieves fast and efficient feature point extraction and matching by using an improved version of the FAST detector and the BRIEF descriptor. Its high computational efficiency makes it suitable for mobile devices and real-time systems. However, the ORB algorithm is prone to reduced matching accuracy when dealing with large scale changes, significant illumination changes, or images with severe parallax. In binocular vision systems, due to parallax, contrast differences, and certain image distortions between the left and right views, traditional single-feature matching algorithms struggle to balance matching speed and accuracy under all conditions. Therefore, improving matching stability and accuracy while ensuring operational efficiency has become a major challenge in current research and applications. Summary of the Invention
[0004] The purpose of the present invention is to provide a visual feature detection and matching method that integrates ORB and SIFT. By combining the response significance of the FAST operator, the detection of DoG extreme points in the scale space is guided. At the same time, according to the scale of the feature points, a smoothly changing weight function is used to dynamically adjust the ratio of SIFT and ORB descriptors, retaining the robustness of SIFT in large-scale areas and giving full play to the efficiency of ORB in small-scale areas, so that the descriptor has stronger scale adaptability and matching stability, so as to solve the problems raised in the above background technology.
[0005] The specific technical solution provided by the present invention is as follows: a visual feature detection and matching method integrating ORB and SIFT, comprising the following steps: Step 1: Initialize system parameters and set relevant parameters of ORB and SIFT algorithms, including the number of feature points, number of scale levels, and descriptor dimensions; Preferably, configure the binocular camera internal parameters , where the parameter Indicates that the camera is and The focal length of the direction, Indicates the coordinates of the image center in the pixel coordinate system, which is used for subsequent accurate geometric transformation and depth calculation of the image. By centralizing the origin of the image coordinate system, moving the origin to the center of the image, and using the intrinsic matrix model, the world point To pixel coordinates The radial distortion parameters of the left and right cameras can be obtained based on the intrinsic parameter matrix of the left camera and the intrinsic parameter matrix of the right camera.
[0006] Step 2: Use the FAST operator to detect DoG extreme points in the scale space; Preferably, performing DoG extreme point detection in the scale space specifically includes: Step 201: Align the binocular images As input, preprocessing of illumination correction and image alignment is performed; Step 202: constructing a joint multi-scale space including SIFT and ORB pyramids; Step 203: Detect the DoG extreme points by calculating only the mask area and using the constructed FAST response value responding to the mask; Step 204: using Taylor expansion to refine the generated candidate key point set, performing quadratic function fitting at the current position of the key point, and calculating the extreme point offset; Step 205: Fusing the key point directions calculated by SIFT and ORB to obtain the final key point directions, and forming a key point set containing key point position, scale and direction information for output.
[0007] Step 3: Fuse the hybrid descriptor of ORB and SIFT to generate a matching strategy; Preferably, generating a matching strategy specifically includes: Step 301: dynamically adjust the fusion weights of SIFT and ORB descriptors using scale space information to generate a hybrid descriptor; Step 302: adaptively adjusting the matching threshold according to the local texture complexity of the area around the feature point; Step 303: Eliminate false matches using geometric constraints between consecutive frames; Step 304: Optimize and output the key point trajectory within the sliding window.
[0008] Step 4: Output matching point pairs according to the matching strategy, and cache the key points of the current frame and their ORB / SIFT descriptors as the initial values for the next frame matching.
[0009] Preferably, the final matching point pair containing the left and right eye image coordinates and three-dimensional space coordinate information is output, the key points of the current frame and their ORB / SIFT descriptors are cached, and the initial values are provided for the next frame matching to achieve cyclic fusion detection and matching.
[0010] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention combines the response significance of the FAST operator to guide the detection of DoG extreme points in the scale space, uses a dynamic threshold to construct a response mask, and performs multi-scale differential operations only in the significant area, effectively avoiding large-scale invalid calculations and improving detection efficiency. It completes response-guided detection in the scale space, which not only takes into account both computational efficiency and feature stability, but also has high practical value and engineering promotion potential. (2) The present invention uses a smoothly varying weight function to dynamically adjust the ratio of SIFT and ORB descriptors based on the scale of the feature points, retaining the robustness of SIFT in large-scale areas and leveraging the efficiency of ORB in small-scale areas. This fusion mechanism enables the descriptor to have stronger scale adaptability and matching stability, which is a more innovative improvement to the traditional multi-descriptor fusion idea. In addition, the present invention addresses the problem that the threshold setting in the matching strategy is difficult to adapt to different image complexities, further introduces local entropy as a measure of image texture richness, and constructs an adaptive matching threshold mapping function. By linking the Hamming distance and the ratio test threshold, the matching strategy can automatically adjust the matching strength according to the regional complexity, significantly improving the matching accuracy and robustness. (3) The present invention jointly models the verification methods of reprojection error or optical flow consistency and introduces geometric consistency verification in the front-end matching stage, which effectively improves the spatiotemporal consistency of matching points. At the same time, combined with the sliding window filter, the motion trajectory of key points is smoothly optimized without increasing the system delay, further improving the stability of front-end feature tracking. At the same time, the sliding window filtering process takes into account the rigid body motion constraints, making the optimization results physically interpretable and applicable to a variety of dynamic scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 1 is a flowchart of the steps of the visual feature detection and matching method integrating ORB and SIFT provided in an embodiment of the present invention; Figure 21 is a schematic diagram of the execution logic of DoG extreme point detection in scale space provided in an embodiment of the present invention; Figure 3 It is a diagram illustrating that only the mask area is calculated when detecting DoG extreme points provided in an embodiment of the present invention; Figure 4 Schematic diagram of a fusion decision tree of ORB and SIFT provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention are within the scope of protection of the present invention.
[0013] Example 1: like Figure 1-Figure 4 As shown, the visual feature detection and matching method integrating ORB and SIFT described in this embodiment is mainly performed by a robot dog and a biped robot equipped with a laser radar and an optical camera including a binocular camera or a depth camera. The binocular camera of the robot dog is located at the head position and faces straight ahead; the camera of the biped robot is located slightly below the front and faces the ground in front at a downward angle of 45 degrees. The specific operating steps of the method are as follows: Step 1: Initialize system parameters and set relevant parameters of ORB and SIFT algorithms, including the number of feature points, number of scale levels, and descriptor dimensions; In this embodiment, because the accuracy of existing laser radars cannot support the analysis requirements, in order to improve the navigation and obstacle avoidance accuracy of robot dogs and robots, the present invention also equips the robot dogs and bipedal robots with optical cameras including binocular cameras or depth cameras in addition to laser radars. That is, the high pixels of the optical camera are used to identify more things around them for positioning, so as to improve the positioning accuracy. Therefore, the present invention configures binocular camera parameters including focal length, baseline, intrinsic parameter matrix, distortion parameters, etc. according to the application scenario; initializes the image preprocessing module, feature detection module and matching module, and sets the relevant parameters of the ORB and SIFT algorithms including the number of feature points, number of scale levels, and descriptor dimension. For example, configure the binocular camera internal parameters , where the parameter Indicates that the camera is and The focal length of the direction, Indicates the coordinates of the image center in the pixel coordinate system, which is used for subsequent accurate geometric transformation and depth calculation of the image. The origin of the image coordinate system is centralized and moved to the center of the image. Then, the world point is converted to the pixel coordinate system using the intrinsic matrix model. To pixel coordinates After transformation, the internal parameter matrix model is:
[0014] in, , Represent the original pixel coordinates, 、 Represents the rotation matrix and translation vector. Based on the intrinsic parameter matrix model, the intrinsic parameter matrix of the point in the camera coordinate system projected to the image coordinate system is calculated, that is, the intrinsic parameter matrix of the left camera Expressed as:
[0015] Intrinsic parameter matrix of the right camera Expressed as:
[0016] Based on the intrinsic parameter matrix of the left camera and the intrinsic parameter matrix of the right camera, the radial distortion parameters of the left and right cameras can be obtained. , the distortion coefficients include ,in, represents the radial distortion parameter, is the tangential distortion coefficient, which is used to correct the radial distortion and tangential distortion generated during the camera imaging process to restore the image to a more accurate geometric shape. In this embodiment, the Brown-Conrady model is used for distortion correction. The correction formula is:
[0017] in, represents the image coordinates after correction and normalization, Represents the normalized image coordinates after removing the influence of internal reference, 、 、 Both represent radial distortion coefficients, 、 represents the tangential distortion coefficient, represents the normalized radius, , represents radial distortion correction, represents the tangential distortion correction, represents the tangential distortion cross term, and They represent the tangential distortion correction terms for the x-coordinate and the y-coordinate respectively. By iteratively optimizing the reverse mapping, a lookup table is constructed to achieve real-time correction.
[0018] Furthermore, configure the baseline at the same time , the adaptive strategy based on near view mode and far view mode is set using the depth calculation model, where the depth calculation model is: ,in, is the depth value, which represents the vertical distance from the target point to the camera plane. Indicates the focal length of the camera, which is defined as the focal length of the two cameras in binocular vision same, represents the baseline length, that is, the distance between the optical centers of the two cameras, represents the disparity, which is the difference between the pixel coordinates of the same spatial point projected on the left and right camera images. It is calculated as follows: ,in is the horizontal pixel coordinate of the target point on the left camera image, is the horizontal pixel coordinate of the target point on the right camera image; baseline Determines the accuracy and range of depth calculations.
[0019] In this embodiment, when constructing a Gaussian pyramid, the scale space is typically divided into several groups, each of which is further divided into several layers. Feature detection parameters, including the number of pyramid groups, the number of layers in each group, and a reference scale, are initialized. The number of pyramid groups determines the number of layers in the constructed image pyramid. Different layers can detect features of different scales, and increasing the number of groups can expand the scale range of feature detection. Each group of layers further refines the image representation at each scale level, enabling the capture of richer detailed features within the same scale range. More layers facilitate more accurate location and description of feature points at different scales. The reference scale is the basic scale for constructing the scale space. Subsequent scale changes are calculated based on this reference scale. It determines the scale of initial feature detection. An appropriate reference scale can ensure that representative features are detected in the initial stage. Furthermore, initialization includes FAST corner threshold , BRIEF descriptor dimension and non-maximum suppression radius ORB key parameters, including The key IFT parameters of extreme value detection threshold and edge response threshold are used to fuse heterogeneous features and generate ORB-SIFT hybrid strategy: generate candidate pairs through ORB fast initial matching, use SIFT to verify candidate pairs, and fuse descriptors. , , using a weighted voting mechanism for matching; improving system adaptability through parameter dynamic optimization model and heterogeneous feature fusion, combining polar constraint acceleration and real-time distortion correction to ensure real-time performance and accuracy, suitable for binocular vision systems in complex scenarios.
[0020] Step 2: Use the response significance of the FAST operator to detect DoG extreme points in the scale space; Step 201: Align the binocular images As input, preprocessing of illumination correction and image alignment is performed; In this embodiment, the binocular image pair It is the basic data for subsequent feature detection and matching. The left and right images contain information about the same scene from different perspectives. The photometric correction model is used to eliminate the illumination difference between the left and right images. for:
[0021] in, represents the original image pixels, represents the local mean of the image, represents the local standard deviation of the image, represents the target standard deviation, Represents the target mean, calculates local illumination statistics by block, generates global illumination field by bilinear interpolation, realizes pixel-level normalized mapping, pre-processes the two images with photometric correction, and then uses histogram to match feature points. Through stereo correction, the matching points are constrained to the same horizontal line for epipolar alignment. .
[0022] Step 202: Construct a joint multi-scale space including SIFT and ORB pyramids: In this embodiment, a Gaussian pyramid is constructed and the Gaussian kernel variance is calculated. :
[0023] in, represents the pyramid index, Presentation layer index, Indicates the number of layers in each group, Indicates the reference scale. In this embodiment, , , , compensate for the blur caused by sampling, and further apply Gaussian blur to the preprocessed image to generate Group Scale-space image of the layer
[0024]
[0025] in, is a two-dimensional Gaussian kernel function, is the preprocessed input image by using a Gaussian kernel Perform convolution, Convolve the image with a one-dimensional Gaussian kernel in the direction The result of the previous step is convolved with a one-dimensional Gaussian kernel to reduce the computational complexity. Within each group, it is constructed through continuous Gaussian blurring and downsampling. After completing all layers of a group, the last layer is downsampled and used as the initial image of the next group. Multiple levels are constructed by downsampling the image to different degrees. The scale space is obtained by convolving the image with Gaussian kernels of different scales. This formula is used to determine the scale of the Gaussian kernel in different levels and groups. Gaussian kernels of different scales can smooth the image to different degrees, thereby detecting features at different scales. The joint scale space constructed in this way can perform unified feature analysis on binocular images at multiple scales.
[0026] Furthermore, a FAST response mask is constructed to calculate the FAST response value of the entire image. and dynamic thresholds , extract the quantiles of the response distribution:
[0027] in, is a binary mask matrix, , used to identify potential feature areas, Is an indicator function that outputs 1 if the condition is true, otherwise it outputs 0. is the FAST response value, indicating the pixel The corner strength, is a dynamic threshold, which is an adaptive threshold that changes with the scale, where the response value The calculation formula is as follows:
[0028] in, is the gray value of the center pixel, is the set pixel value, is the basic threshold, is the indicator function; Dynamic Threshold The calculation formula is as follows:
[0029] in, Indicates the 70% quantile, which is the quantile of the FAST response value of the entire map. is a natural constant, is the current scale, is the maximum scale; by collecting all Values, arranged in ascending order to form a sequence , and calculate the position index to get the corresponding value. FAST is a fast feature point detection algorithm that determines whether it is a feature point by calculating the FAST response value of the pixel point. The response mask is based on the FAST response value and the dynamic threshold. To screen potential feature point areas, the dynamic threshold is calculated by combining the 70% quantile of the FAST response value. As well as the relationship between the current scale and the maximum scale, this method can adaptively determine the feature point detection area according to the image content and scale, reducing invalid calculations.
[0030] Step 203: Detect the DoG extreme points by calculating only the mask area and using the constructed FAST response value responding to the mask; In this embodiment, two adjacent scales are calculated by only taking the pixel positions marked as 1 in the mask. and The difference of the Gaussian blurred image is used to obtain the DoG response map. The DoG response is directly set to 0 at the position where the mask is 0. The image set obtained by Gaussian blurring the original image at multiple scales is fused with the binary image of the same size as the original image to mark the area that may contain feature points. 1 represents the area of interest and 0 represents the background or non-feature area. For each pair of adjacent scales, their difference image is calculated and the above difference image is combined with the mask. Multiply pixel by pixel, that is, only calculate the mask area, and assign 0 to other positions to get the scale Response graph, output Response Plot , whose non-zero value only appears in the area marked as 1 in the mask. By detecting local extreme values in the set neighborhood, a set of candidate key points is generated; For example, the calculation formula for calculating only the mask area using the DoG calculation model is:
[0031] in, is the DoG response value, is the scale space image, is the current scale, is the next scale parameter, scale parameter and are two adjacent scales, and the proportional relationship between them is fixed. For the response mask, is the pixel coordinate. DoG is a commonly used feature point detection method that highlights potential feature points by calculating the difference between Gaussian blurred images of different scales. Here, only the mask area is calculated, that is, For areas where the value is 1, a large amount of invalid calculations can be reduced by utilizing the previously constructed FAST response mask, which greatly improves the detection efficiency. At the same time, the potential feature point areas are calculated centrally, which improves the accuracy of feature point detection.
[0032] Step 204: using Taylor expansion to refine the generated candidate key point set, performing quadratic function fitting at the current position of the key point, and calculating the extreme point offset; In this embodiment, the position of each candidate key point is refined iteratively, and the central difference method is used to calculate the gradient of the current position using the DoG image. and the Hessian matrix And the offset of the key point position , the specific calculation formula is as follows:
[0033]
[0034]
[0035] in, is the gradient vector of the DoG function at the key point, is the Hessian matrix of the DoG function at the key point, is the partial derivative, is the inverse Hessian matrix, and Indicates that the DoG function is Direction and The first-order partial derivative in direction; Furthermore, the position is updated and iterative relocation is used. The convergence condition is set when the offset is less than the set pixel threshold or the maximum number of iterations is reached. It is iterated according to the size of any component of the calculated offset and the set pixel threshold. When any component of the offset is less than the set pixel threshold, the current key point position is output and iterated again. When the maximum number of iterations is exceeded or the offset position exceeds the neighborhood, the point is discarded. The final refined key point position is obtained according to the iterative result. , ,in, The key point position of the discrete pixel coordinates of the current iteration. If the key point shifts to an adjacent pixel, the Hessian and gradient of the adjacent position are recalculated and the adaptive threshold is performed. Adjustment, the adaptive threshold formula is: , is the basic threshold, is the adjustment coefficient, To estimate the image noise level, the threshold is dynamically adjusted according to the image noise level. When the noise is large, the threshold is increased, and vice versa. The noise level is estimated by the statistical characteristics of the image gradient amplitude, and the Hessian matrix eigenvalue check is performed simultaneously to distinguish corners and edges. The check formula is: ,in and is the eigenvalue of the Hessian matrix, is a larger eigenvalue, is a smaller eigenvalue, For example, based on the second-order derivative matrix and first-order derivative of the DoG function, the key point position is fine-tuned to locate it at the characteristic position in the image, and the Hessian matrix is used to describe the second-order derivative information of the function at a certain point. By verifying the proportional relationship of its eigenvalues, some unstable key points caused by image noise or non-feature structure can be removed. In this embodiment, As the judgment threshold, only key points that meet this condition are retained, further improving the quality of key points.
[0036] Step 205: Fusing the key point directions calculated by SIFT and ORB to obtain the final key point directions, and forming a key point set containing key point position, scale, and direction information for output; In this embodiment, the direction calculation model is used to calculate SIFT direction and ORB direction. and ORB direction calculation They are:
[0037] in, For candidate directions, is the Gaussian weight, is the gradient direction histogram, is the histogram bucket offset, and They are and Directional moment of direction; Use fusion strategy to calculate the direction difference between SIFT direction calculation and ORB direction calculation calculate, The calculation formula is:
[0038] Set dynamic tolerance threshold, when SIFT and ORB direction difference When the value is less than or equal to the set threshold, it is not rejected directly, nor is a certain direction adopted directly. Instead, a weighted fusion of directions is performed. The two direction values are weighted averaged by the confidence level. The calculation formula for the weighted fusion direction is:
[0039] in, is the weight assigned to the SIFT direction, As for the weight assigned to the ORB direction, for the SIFT direction, if the main peak of its direction histogram is very prominent, it means that the direction of the key point is clearly defined, and a higher weight is given. At the same time, if the gradient amplitude at the key point is large, the SIFT weight is also increased. Because SIFT is based on gradients, its direction estimation is more reliable if the gradient is large. For the ORB direction, if the local contrast is high, it means that the grayscale changes in the area are obvious, and the ORB centroid method direction estimation may be more accurate, so it is given a higher weight. At the same time, because the ORB direction may be more stable at larger scales, and the contrast is normalized, the two directions are weighted averaged according to the weight to obtain the final direction. This dual-algorithm direction fusion method combines the advantages of the two algorithms in direction calculation and improves the accuracy of key point direction assignment.
[0040] Step 3: Fuse the hybrid descriptor of ORB and SIFT to generate a matching strategy; Step 301: dynamically adjust the fusion weights of SIFT and ORB descriptors using scale space information to generate a hybrid descriptor; In this embodiment, the SIFT descriptor is normalized, and the ORB descriptor is converted into a floating-point vector and normalized as well. Then, the SIFT weight and the ORB weight are calculated according to the key point scale, and the normalized descriptors are weighted fused. Then, the fused descriptors are normalized again. The calculation formula for weighted fusion of descriptors is:
[0041]
[0042] in, is the fused descriptor vector, is the scale-aware weight function, is the normalized SIFT descriptor, is the ORB descriptor converted to floating point and normalized, Convert binary to floating point. is the parameter that controls the rate of change of weights. is the scale of the key points constructed based on the scale space, The SIFT descriptor has good scale invariance and rotation invariance, but the computational complexity is large; the ORB descriptor has high computational efficiency but its performance is not as good as SIFT in some aspects. In this embodiment, a weight function that changes with scale is used. The SIFT and ORB descriptors are fused and their contribution ratio is dynamically adjusted at different scales. The generated scale-aware descriptor can combine the advantages of the two descriptors and has good description capabilities at different scales.
[0043] Step 302: adaptively adjusting the matching threshold according to the local texture complexity of the area around the feature point; In this embodiment, for each query descriptor, two candidate descriptors with the smallest Hamming distance are found in the target image. and ,like <Adaptive Hamming distance threshold and <Adaptive ratio test threshold of feature point position When , a match is accepted, dynamic adjustment is achieved by relaxing the Hamming distance threshold and tightening the ratio test in regions with high local entropy, while the opposite is true in simple regions.
[0044] Exemplarily, the dynamic threshold calculation formula is:
[0045] in, is the adaptive Hamming distance threshold of the feature point position, is the adaptive ratio test threshold of the feature point position, is the basic threshold set, and is the set adjustment coefficient, is the image entropy of the local area where the feature point is located, and the calculation formula is:
[0046] in, In the circular area centered on the feature point, the gray value The probability of occurrence, in the feature matching process, the selection of the matching threshold is crucial to the accuracy and efficiency of the matching results. Traditional methods usually use a fixed threshold. In this invention, the Hamming distance threshold is dynamically calculated based on the local entropy. and ratio test threshold ,Local entropy reflects the information entropy of the local area of the image,,that is, the complexity of the area.,Adaptively adjusting the matching threshold according to the different characteristics of the,local area can improve the matching accuracy and adapt to the needs of different,scenes.
[0047] Step 303: Eliminate false matches using geometric constraints between consecutive frames; In this embodiment, a spatiotemporal consistency check is performed to verify the geometric constraints of the matching points by calculating the optical flow error and the reprojection error. The optical flow error measures the motion consistency of the feature points between adjacent frames, and the reprojection error considers the error of projecting the three-dimensional point onto the image plane. The two are combined by weighted summation, and the weight is determined by the variance of the reprojection error. When the geometric constraint is greater than the set geometric error threshold, it is judged as an erroneous matching point and the match is eliminated. This spatiotemporal consistency check can remove erroneous matching points and improve the reliability of matching.
[0048] Step 304: Optimize the key point trajectory within the sliding window; In this embodiment, sliding window filtering is used to optimize the trajectory of key points. This filtering takes a weighted average of the key points within a certain time window and, for key points on the same object, determines the optimal translational velocity and angular velocity of the reference point. This minimizes the residual between the observed velocity of all key points and the velocity predicted by the rigid body motion model. The weight decays exponentially with increasing time distance, smoothing the trajectory of the key points and removing the influence of noise and outliers. At the same time, the rigid body motion constraints are satisfied. These rigid body motion constraints ensure that the motion of the key points conforms to the rigid body motion model. Specifically, the motion of the key points can be decomposed into the motion caused by the translational velocity and angular velocity. This further optimizes the key point trajectory, making it more consistent with the motion patterns of objects in real-world scenes.
[0049] Step 4: Output matching point pairs containing key points in the left and right images and the corresponding 3D point information, build a key point motion trajectory library, and cache the key points of the current frame and their ORB / SIFT descriptors as the initial values for the next frame matching; In this embodiment, the final matching point pair containing the left and right eye image coordinates and the three-dimensional space coordinate information is output, the key points of the current frame and their ORB / SIFT descriptors are cached, and the initial values are provided for the next frame matching, and the efficiency of temporal inter-frame matching is improved in conjunction with the optical flow.
[0050] For example, the final matching point pair contains the corresponding key points in the left and right images and their corresponding three-dimensional point information. It is the final result of visual feature detection and matching and can be used for further tasks such as depth estimation and three-dimensional reconstruction. A key point motion trajectory library is constructed to record and track the position information of key points at different times. By tracking the motion trajectory of key points in multiple frames, the motion state and behavior of the object can be further analyzed. It also helps to discover abnormal key point motion and improve the stability and accuracy of the system. Finally, through the cross-frame prediction model, based on the rigid body motion model, the position, translation speed and rotational angular velocity of the current key point are used to predict the position of the key point at the next moment, and the position of the key point in the next frame is estimated in advance, so that the search range is narrowed in the feature matching of subsequent frames, the matching efficiency is improved, and the matching results can also be verified and corrected, further improving the performance of visual feature detection and matching.
[0051] It should be noted that, in the present invention, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0052] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A visual feature detection and matching method integrating ORB and SIFT, characterized by: The following steps are included: Step 1: Initialize system parameters and set relevant parameters of ORB and SIFT algorithms, including the number of feature points, number of scale levels, and descriptor dimensions; Step 2: Use the FAST operator to detect DoG extreme points in the scale space; Step 3: Fuse the hybrid descriptor of ORB and SIFT to generate a matching strategy; Step 4: Output matching point pairs according to the matching strategy, and cache the key points of the current frame and their ORB / SIFT descriptors as the initial values for the next frame matching, and perform cyclic fusion detection and matching.
2. The visual feature detection and matching method integrating ORB and SIFT according to claim 1, characterized in that: The initialization system parameters specifically include: Configure the binocular camera internal parameters; The origin of the image coordinate system is centralized. After the origin is moved to the center of the image, the world point is transformed into pixel coordinates using the intrinsic parameter matrix model. Based on the intrinsic parameter matrix model, the intrinsic parameter matrix of the point in the camera coordinate system projected to the image coordinate system is calculated to obtain the radial distortion parameters of the left and right cameras; At the same time, the baseline length is configured, and the adaptive strategy based on the near view mode and the far view mode is set using the depth calculation model; Finally, the ORB key parameters including FAST corner threshold, BRIEF descriptor dimension and non-maximum suppression radius are initialized, as well as The key parameters of extreme value detection threshold and edge response threshold are used to fuse heterogeneous features and generate ORB-SIFT hybrid strategy.
3. The visual feature detection and matching method integrating ORB and SIFT according to claim 2, characterized in that: The step 2 specifically includes: Step 201: taking the binocular image pair as input, performing preprocessing of illumination correction and image alignment; Step 202: constructing a joint multi-scale space including SIFT and ORB pyramids; Step 203: Detect the DoG extreme points by calculating only the mask area and using the constructed FAST response value responding to the mask; Step 204: using Taylor expansion to refine the generated candidate key point set, performing quadratic function fitting at the current position of the key point, and calculating the extreme point offset; Step 205: Fusing the key point directions calculated by SIFT and ORB to obtain the final key point directions, and forming a key point set containing key point position, scale and direction information for output.
4. The method for visual feature detection and matching by integrating ORB and SIFT according to claim 3, wherein: The constructing of the joint multi-scale space includes: Construct a Gaussian pyramid and calculate the Gaussian kernel variance; Apply Gaussian blur to the preprocessed image to generate scale space images at each level of each group in the pyramid; Within each group, the layers are constructed by continuous Gaussian blurring and downsampling. After all layers of a group are completed, the last layer is downsampled and used as the initial image for the next group. Multiple layers are constructed by downsampling the images to varying degrees. Finally, the FAST response mask is constructed, the full-image FAST response value and dynamic threshold are calculated, and the quantiles of the response value distribution are extracted.
5. The method for visual feature detection and matching by integrating ORB and SIFT according to claim 4, characterized in that: The performing calculation on only the mask area includes: using the DoG calculation model to calculate only the mask area.
6. The method for visual feature detection and matching by integrating ORB and SIFT according to claim 5, characterized in that: The key point refining comprises: The position of each candidate key point is refined iteratively, and the central difference method is used to calculate the gradient and Hessian matrix of the current position and the offset of the key point position using the DoG image; Update the position and use iterative relocalization, setting the convergence condition when the offset is less than the set pixel threshold or the maximum number of iterations is reached; Iterate based on the size of any component of the calculated offset and the set pixel threshold. When any component of the offset is less than the set pixel threshold, output the current key point position and iterate again. When the maximum number of iterations is exceeded or the offset position is out of the neighborhood, the point is discarded and the final refined key point position is obtained based on the iteration result. If the key point is offset to an adjacent pixel, the Hessian and gradient of the adjacent position are recalculated and the adaptive threshold is adjusted. The set threshold is dynamically adjusted according to the image noise level. When the noise is large, the set threshold is increased, and vice versa. The noise level is estimated through the statistical characteristics of the image gradient amplitude, and the Hessian matrix eigenvalue check is performed simultaneously to distinguish corners and edges.
7. The method for visual feature detection and matching by integrating ORB and SIFT according to claim 6, characterized in that: Use the direction calculation model to calculate SIFT direction and ORB direction; Use the fusion strategy to calculate the direction difference between SIFT direction calculation and ORB direction calculation; A dynamic tolerance threshold is set. When the difference between the SIFT and ORB directions is less than or equal to the set threshold, a weighted fusion of directions is performed. The two direction values are weighted averaged by confidence, and the final direction is obtained by weighted averaging the two directions according to the weights.
8. The method for visual feature detection and matching by integrating ORB and SIFT according to claim 7, characterized in that: The generation of the matching strategy specifically includes: Step 301: dynamically adjust the fusion weights of SIFT and ORB descriptors using scale space information to generate a hybrid descriptor; Step 302: adaptively adjusting the matching threshold according to the local texture complexity of the area around the feature point; Step 303: Eliminate false matches using geometric constraints between consecutive frames; Step 304: Optimize and output the key point trajectory within the sliding window.
9. The method for visual feature detection and matching by integrating ORB and SIFT according to claim 8, characterized in that: The generation of the hybrid descriptor includes: normalizing the SIFT descriptor, converting the ORB descriptor into a floating-point vector and normalizing it, calculating the SIFT weight and the ORB weight according to the key point scale, performing weighted fusion on the normalized descriptors, and normalizing the fused descriptors again before outputting them.
10. The visual feature detection and matching method integrating ORB and SIFT according to claim 9, characterized in that: The step 4 specifically includes: outputting the final matching point pair containing the left and right eye image coordinates and the three-dimensional space coordinate information, caching the key points of the current frame and their ORB / SIFT descriptors, providing initial values for the next frame matching, and realizing cyclic fusion detection and matching.
Citation Information
Patent Citations
Method and device for detecting feature points in Fast approximated SIFT algorithm
CN103413326A
Image feature extraction method and system based on binocular camera, and intelligent terminal
CN113792752A
Method and device for registering optical image and infrared image of circuit board
CN116433733A
Optical element surface defect three-dimensional splicing method based on modulation degree and SIFT
CN117541467A
Methods and Systems for Vision-Based Motion Estimation
US20160063330A1
Cited By
Automatic correction method for angles of wine bottles
CN122066770A