Binocular depth accurate solving method based on region segmentation

By employing a binocular depth-accurate solution method based on region segmentation, utilizing histogram equalization, ORB feature extraction, and improved SAD sliding window matching, combined with parabolic fitting and K-means clustering, the problem of depth solution error under repetitive textures and low lighting conditions is solved, achieving high-precision scene depth acquisition and navigation positioning.

CN120976289APending Publication Date: 2025-11-18BEIJING AUTOMATION CONTROL EQUIP INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510978860.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In scenarios with repetitive textures and low lighting, existing binocular vision positioning technology suffers from large depth calculation errors and difficulties, making it hard to achieve high-precision scene depth acquisition.

Method used

A binocular depth-accurate solution method based on region segmentation is adopted. Image preprocessing is performed through histogram equalization and epipolar correction. Feature matching is performed by combining ORB feature extraction and an improved SAD sliding window matching method. Feature points are selected by parabolic fitting and K-means clustering. Finally, the scene depth is obtained by triangulation.

Benefits of technology

It improves the accuracy of feature matching in repetitive textures and low-light scenes, eliminates depth calculation errors, achieves higher-precision scene depth acquisition, and provides the necessary foundation for navigation and positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976289A_ABST
    Figure CN120976289A_ABST
Patent Text Reader

Abstract

The invention provides a binocular depth accurate solving method based on region segmentation, and the method comprises the steps: carrying out the preprocessing of a left eye image and a right eye image, carrying out the feature extraction, and obtaining the feature points of the left eye image and the right eye image; matching the feature points of the left eye image and the right eye image by adopting an SAD sliding window matching method according to the obtained feature points to obtain a matching pair; optimizing the obtained matching pair by using a parabola fitting method to obtain a matching pair of sub-pixel-level matching; according to the second-level depth elimination method based on pixels and planes, first-level feature point elimination of the scene depth is completed through pixel distribution, second-level feature point elimination of the scene depth is completed through a regional plane, after feature points which do not meet requirements are eliminated, a final matching pair is obtained, and finally a scene depth value is obtained through the triangularization technology based on the final matching pair. According to the invention, the problem of depth solving in scenes such as repeated textures and low illumination can be solved, so that high-precision scene depth acquisition is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of visual navigation, and relates to a binocular depth precise solving method based on region segmentation. BACKGROUND

[0002] With the rapid development of visual sensor technology, computer technology and artificial intelligence technology, as a new navigation method, visual navigation begins to occupy a place in the field of autonomous navigation as a new navigation method based on computer vision recognition and positioning technology. Visual sensor has the characteristics of concealment, lightness, low power consumption and low price, so positioning by using visual sensor has great advantages.

[0003] Binocular positioning simulates the way people watch the world, and the real world scene appears to move on the image plane through left and right two visual sensors imaging respectively. The difference between the positions of the pixel points constitutes the parallax, and the real world scene depth can be calculated reversely, which is a common depth positioning method. However, in the repeated texture scene, the feature extraction is misjudged, the feature matching is easy to be wrong, and the depth solving error is caused. In the scene with dim light and few features, the scene features are not rich enough, and the feature point observation information is weak, which also causes difficulty in depth solving. SUMMARY

[0004] The purpose of the application is to solve at least one of the problems existing in the prior art.

[0005] To this end, the application provides a binocular depth precise solving method based on region segmentation, which mainly faces the low-altitude navigation scene and the ground navigation scene of unmanned equipment, and can solve the problem of depth solving in repeated texture and low light scenes, and then realize high-precision scene depth acquisition.

[0006] The technical solution of the application is:

[0007] A binocular depth precise solving method based on region segmentation, the steps of the method are:

[0008] Step 1: After pre-processing the left eye image and the right eye image, the feature points of the left eye image and the right eye image are extracted by feature extraction, and the feature points of the left eye image and the right eye image are obtained;

[0009] Step 2: According to the obtained feature points of the left eye image and the right eye image, the SAD sliding window matching method is used to match the feature points of the left eye image and the right eye image, and the matching pairs are obtained;

[0010] Step 3: The obtained matching pairs are optimized by using the parabolic fitting method, and the sub-pixel level matching pairs are obtained;

[0011] Step 4: A two-stage depth rejection method based on pixels and planes is used to complete the scene depth one-stage rejection feature points using pixel distribution, and the scene depth two-stage rejection feature points using regional planes. After the feature points that do not meet the requirements are rejected, the final matching pair is obtained. Finally, the scene depth value is obtained based on the final matching pair using the triangulation technique.

[0012] Further, in step 1, the method of pre-processing the left eye image and the right eye image includes using histogram equalization and polar correction.

[0013] Further, the process of pre-processing the image using histogram equalization is as follows:

[0014] First, the gray scale distribution of the left eye image or the right eye image at the same time t is calculated to determine whether to perform histogram equalization. The maximum pixel gray value is subtracted from the minimum pixel gray value in the left eye image or the right eye image. If the gray scale difference is within 1 / 4 of the image gray scale value range, it is determined that histogram equalization needs to be performed; otherwise, it is determined that histogram equalization does not need to be performed.

[0015] Then, the gray scale mapping function of the left eye image at time t is calculated, and the right eye image is processed using the same gray scale mapping function to complete histogram equalization, so that the left eye image and the right eye image at the same time can be processed on the same basis.

[0016] Further, in step 1, the ORB feature extraction method is used to extract features from the pre-processed left eye image and right eye image.

[0017] Further, in step 2,

[0018] Step 2-1, the initial matching pair of the left and right eye images is constructed, and the initial matching pair is: the feature point (u l ,v l ) on the left eye image and the matching point (u r ,v r ) on the right eye image.

[0019] Step 2-2, the SAD sliding window matching method is used to adjust the initial matching pair obtained in step 2-1:

[0020] First, a pixel block is taken out in the left eye image, and the pixel block is centered on the feature point (u l ,v l ) with a pixel region size of 11*11. Then, a pixel block is also established in the right eye image, and the pixel block is centered on the x-axis coordinate u r of the feature point and slides along the x-axis of the feature point of the right eye image. Finally, the sum r of the absolute values of the differences between all pixel gray values of the pixel block of the left eye image and the sliding pixel block of the right eye image is calculated.i :

[0021]

[0022] wherein, represents the gray value of the i-th pixel in the pixel block of the left eye image, represents the gray value of the i-th pixel in the pixel block of the right eye image;

[0023] record the two smallest r i values, let the minimum error corresponding to the minimum value of the calculated r i value be Dr1, and the right eye deviation amount corresponding to the second minimum error be Dr2; then the adjusted matching pair is obtained, and the coordinates of the matching points on the right eye image of the adjusted matching pair are (u r +Dr1,v r ) and (u r +Dr2,v r ), respectively.

[0024] Further, in step 2-1, the process of constructing the initial matching pair is as follows:

[0025] Let the x-axis coordinate of the feature point on the left eye image be u l , and the y-axis coordinate be v l , traverse all the feature points in the horizontal search band of the right eye image, and calculate the descriptor distance between all the feature points and the feature point (u l ,v l ) on the left eye image, respectively, record the feature point coordinates (u r ,v r ) of the right eye image with the minimum descriptor distance, and the feature point is the matching point on the right eye image; therefore, the feature point (u l ,v l ) on the left eye image and the matching point (u r ,v r ) on the right eye image form an initial matching pair.

[0026] Further, after step 2-2, the matching points on the right eye image of the adjusted matching pair are screened:

[0027] calculate the descriptor distance d1 between the feature point (u l ,v l ) on the left eye image and the matching point (u r +Dr1,v r ) on the right eye image, and the descriptor distance d2 between the feature point (u l ,v l ) on the left eye image and the matching point (u r +Dr2,v rthe descriptor distance d2 of the right eye image;

[0028] If the difference between d1 and d2 is less than a preset threshold, the matching point of the right eye image retains two coordinates, i.e. retains (u r +Dr1,v r ) and (u r +Dr2,v r ); otherwise, only the matching point (u r +Dr1,v r ) of the right eye image is retained.

[0029] Further, in step 3, the process of obtaining the matching pair of sub-pixel level matching is as follows:

[0030] Since there is a best correction amount between the matching point of the right eye image of the currently selected step 2-2 and the best matching point of the right eye image, an error parabola can occur in the sliding window matching process of the feature point of the right eye image, and the position of the minimum value of the error parabola is the position of the best matching point of the right eye image.

[0031] If the minimum value is at the boundary of the error parabola, it means that there is no inflection point, and the best matching point is abandoned, and the matching point of the right eye image of step 2-2 is still selected; if the minimum value is not at the boundary of the error parabola, the best matching point corresponding to the minimum value position is selected as the matching point of the right eye image.

[0032] Further, the process of selecting the best matching point corresponding to the minimum value position of the error parabola as the matching point of the right eye image is as follows:

[0033] If there is a matching point (u r +Dr1,v r ) on the right eye image, the error parabola is constructed with the matching points (u r +Dr1), (u r +Dr1-1), (u r +Dr1+1) of the right eye image as the horizontal axis and the descriptor distance of the feature point (u l ,v l ) of the left eye image as the vertical axis; by solving the minimum value of the error parabola, the sub-pixel level correction amount Dr1¢ can be obtained, and the final matching point coordinate of the right eye image after parabola fitting is (u r +Dr1+Dr1¢,v r ).

[0034] Similarly, if there are two matching points on the right eye image, the final matching point coordinate of the right eye image after parabola fitting of the other matching point (u r +Dr2,v r ) is (u r+Dr2+Dr2¢,v r );

[0035] Thus, the sub-pixel level matching pairs are obtained as: the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image; or the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image; or the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr2+Dr2¢,v r ) of the right eye image.

[0036] Further, in step 4,

[0037] When the feature points are removed at the pixel level, the r i values are sorted first; then the median of the r i values is taken as the average matching error of the scene, and finally the depth values greater than 2 times the average matching error and the corresponding feature points in the current scene are removed;

[0038] When the feature points are removed at the plane level, the left eye image and the right eye image are divided into different regions by using the image segmentation method based on K-means clustering; the depth of the image in each region is determined again: if the depth value of a pixel point in the same region is depth, the depth value depth of the pixel point should be within the depth envelope range of the 8-neighborhood within the region; if the depth value depth of the pixel point is greater than 1.5 times of the depth envelope range or less than 0.5 times of the depth envelope range, the matching result and the depth solving result corresponding to the pixel point are removed; otherwise, they are not removed;

[0039] Finally, when the secondary removal of the image is completed, all the feature points in the right eye image are traversed, if the sub-pixel level matching pairs are only one, i.e. the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image, the remaining one sub-pixel level matching pair is the final accurate matching pair; if the sub-pixel level matching pairs are still two, i.e. the matching points of the right eye image still exist two, the matching point (u r +Dr1+Dr1¢,vr ) and the matching point (u r +Dr2+Dr2¢,v r ) of the right eye image, then the matching point (u r +Dr2+Dr2¢,v r ) of the right eye image is removed, only the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image is reserved, then the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image are the final accurate matching pair;

[0040] Therefore, the final matching pair after the removal is the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image; after the matching of the feature points of the left and right eye images, the scene depth value can be obtained by using the triangulation technology, so as to realize the accurate binocular depth solving.

[0041] By applying the technical scheme, the present application has the following beneficial effects:

[0042] (1) The present application realizes image preprocessing by histogram equalization, and gives a feature matching candidate set by using an improved SAD sliding window matching method, further improves the correctness of feature matching under repeated textures, and can solve the problems of weak light and repeated textures.

[0043] (2) The present application proposes two-layer mechanisms of pixel-level depth screening and scene plane-level depth screening to remove depth solving errors, can remove the wrong matching pixel pairs in a large range of scenes, obtain more accurate scene depth solving results, improve the depth solving precision in application scenarios, and provides a necessary technical basis for realizing scene composition and navigation positioning. BRIEF DESCRIPTION OF DRAWINGS

[0044] The accompanying drawings, which are included to provide a further understanding of the embodiments of the application and constitute a part of the specification, illustrate the embodiments of the application and together with the text description serve to explain the principles of the application. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0045] Figure 1 is a flowchart of the present application. DETAILED DESCRIPTION

[0046] It should be noted that the embodiments and features of the embodiments in the present application can be combined with each other in the case of no conflict. The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The description of the at least one example embodiment is actually only illustrative, but not as any limitation on the present application and its application or use. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0047] It should be noted that the terms used herein are only intended to describe specific embodiments, and are not intended to limit the example embodiments according to the present application. As used herein, the singular form is intended to include the plural form, unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in the specification, there is a reference to the presence of a feature, step, operation, device, component and / or combinations thereof.

[0048] Unless specifically stated otherwise, the relative arrangement of components and steps, numerical expressions, and numerical values set forth in the examples provided herein are not intended to limit the scope of the application. It should also be understood that the size of the various parts shown in the drawings can not be to scale, and that the drawings are intended to be illustrative, not definitive. Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but should be considered as part of the authorized description. In all examples shown and discussed herein, any specific value should be interpreted as merely exemplary, not as a limitation. Therefore, other examples of the example embodiments can have different values. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0049] Embodiment 1:

[0050] The present embodiment provides a binocular depth precision solving method based on region segmentation, referring to the attached Figure 1 The steps of the method are as follows:

[0051] Step 1: Histogram equalization and polar correction are used to pre-process the images of the left and right cameras, and ORB feature extraction method is used to extract features from the pre-processed images to obtain feature points of the images of the left and right cameras;

[0052] Step 2: According to the feature points of the images of the left and right cameras obtained, the feature points of the images of the left and right cameras are matched by using a modified SAD sliding window matching method to obtain a matching pair, and a matching candidate set is constructed;

[0053] Step 3: The matching pair obtained is optimized by using a parabolic fitting method to obtain a sub-pixel level matching pair;

[0054] Step 4: Based on a two-level depth rejection method of pixels and planes, pixel distribution is used to complete scene depth one-level rejection feature points, and regional planes are used to complete scene depth two-level rejection feature points. After the feature points that do not meet the requirements are rejected, the final matching pair is obtained. Finally, based on the final matching pair, the scene depth value can be obtained by using a triangulation technique.

[0055] Embodiment 2

[0056] Based on the embodiment 1, the process of each step is provided, which is as follows:

[0057] In step 1, the image is preprocessed by using histogram equalization as follows:

[0058] The histogram equalization is a simple and effective image enhancement technique, which changes the gray scale of each pixel in the image by changing the histogram of the image, and is mainly used to enhance the contrast of the image with a small dynamic range. Due to the overexposure or overexposure of light, the gray scale of the original image may be concentrated in a narrow interval, such as the gray scale of the overexposed image being concentrated in the high brightness range, and the underexposed image will make the image gray scale concentrated in the low brightness range. By using histogram equalization, the histogram of the original image can be transformed into a uniformly distributed form, achieving the effect of enhancing the overall contrast of the image.

[0059] Therefore, the process of pre-processing the image by using histogram equalization in this embodiment is as follows:

[0060] Firstly, the gray scale distribution of the left and right camera imaging at the same time (assuming t time) is judged. Since the judgment of whether the images of the left and right cameras need to develop histogram equalization is consistent, that is, if the image of the left camera needs to develop histogram equalization, the image of the right camera must also develop histogram equalization, and if the image of the left camera does not need to develop histogram equalization, the image of the right camera also does not need to develop histogram equalization. Therefore, only the left camera image or the right camera image needs to be judged. In this embodiment, the left camera image is taken as an example. If the maximum pixel gray scale in the image is subtracted from the minimum pixel gray scale, and the gray scale difference is within 1 / 4 of the image gray scale value range (such as the commonly used light image gray scale range of 0-255), it is determined that histogram equalization needs to be developed. Otherwise, it is determined that histogram equalization does not need to be developed.

[0061] Then, given the assumption that the light conditions of the left and right cameras at the same time t are similar, the gray map mapping function is calculated using the imaging result of the left camera at time t, and the imaging result of the right camera is processed by the same gray map mapping function to complete histogram equalization, so that the images of the left and right cameras at the same time can be processed on the same basis subsequently. Thus, the histogram equalization of the imaging results of the left and right cameras is completed.

[0062] In step 1, the reason for performing polar correction on the image is that binoculars can measure depth through left and right camera imaging point matching pairs, which is actually based on an assumption that the image planes of the left and right cameras are accurately located on the same plane, achieving line alignment, and the two optical axes are strictly parallel. However, due to installation errors, the coordinate axes of the two cameras are difficult to achieve ideal conditions, so polar correction is needed for the left and right cameras.

[0063] In step 1, when performing feature extraction on the preprocessed image using the ORB feature extraction method, the feature points and feature descriptors of the preprocessed images of the left and right cameras are extracted to complete the feature extraction of the corresponding generated images of the left and right cameras at the same time.

[0064] The specific steps of step 2 are as follows:

[0065] Step 2-1, construct the initial matching pairs of the images of the left and right cameras (hereinafter referred to as left and right eye images):

[0066] Since the left and right cameras after polar correction have feature points on the same horizontal line extension, when matching the pixels of the left and right eye images, the feature points of the left eye image are fixed, and the feature points of the right eye image corresponding to the feature points of the left eye image are searched on the horizontal line extension to solve the closest feature points of the right eye image as matching points. The pixel matching of the left and right eye images is achieved. However, considering the small errors in the algorithm, when performing preliminary matching of the feature points of the left and right eye images, the feature points of the left eye image are fixed, and the matching points of the corresponding right eye image are not actually searched on the horizontal line extension, but on a horizontal search band. The width of the horizontal search band is generally given as 5 rows, i.e. adding 2 rows (positive and negative 2 rows) above and below the horizontal line to form a 5-row horizontal search band.

[0067] Let the x-axis coordinate of the feature point on the left eye image be u l , and the y-axis coordinate be v l , and all feature points of the right camera in the horizontal search band are called candidate feature point set. The feature points of the candidate feature point set are calculated respectively with the feature points (u l , v l) and record the feature point coordinates of the right eye image with the minimum descriptor distance as (u r ,v r ), which is the matching point on the right eye image; thus, the feature point (u l ,v l ) on the left eye image and the matching point (u r ,v r ) on the right eye image form an initial matching pair;

[0068] Step 2-2, the initial matching pair obtained in step 2-1 is refined and adjusted by using the improved SAD sliding window matching method:

[0069] First, a pixel block is taken out from the left eye image, which is centered at the feature point (u l ,v l ) and has a size of 11*11 pixels; then, a pixel block is also established in the right eye image, which is centered at the x-axis coordinate u r of the feature point and slides along the x-axis of the feature point of the right eye image; finally, the sum r i of the absolute values of the differences between all pixel gray values of the pixel block of the left eye image and the sliding pixel block of the right eye image (i.e., the sum of the gray value differences of the two pixel blocks) is calculated:

[0070]

[0071] In the formula, Ii represents the i-th pixel gray value in the pixel block of the left eye image, and Ii represents the i-th pixel gray value in the pixel block of the right eye image.

[0072] Considering that the feature point descriptor has similarity in a repetitive texture scene, the left and right eye image matching "misalignment" phenomenon is prone to occur, therefore, the improved SAD matching method of the embodiment records two minimum r i values, and the minimum error right eye deviation amount corresponding to the minimum value in the calculated r i value is Dr1, and the next minimum error right eye deviation amount corresponding to the next minimum value is Dr2; thus, the adjusted matching pair is obtained, and the coordinates of the matching points on the right eye image of the adjusted matching pair are (u r +Dr1,v r ) and (u r +Dr2,v r ), respectively.

[0073] Step 2-3, the matching points on the right eye image of the adjusted matching pair are screened:

[0074] The feature point (u l ,vl Matching points (u) with the right eye image r +Dr1,v r The descriptor distance d1 and the feature points (u) of the left eye image l ,v l Matching points (u) with the right eye image r +Dr2,v r The descriptor distance d2 of );

[0075] If the difference between d1 and d2 is less than a preset threshold, then the matching point of the right eye image retains two coordinates, that is, it retains (u... r +Dr1,v r ) and (u r +Dr2,v r Conversely, only the matching points of the right eye image are retained. r +Dr1,v r ).

[0076] Therefore, the feature points (u) of the left eye image l ,v l Matching points (u) of the right eye image and the right eye image r +Dr1,v r The matching candidate set is composed of the feature points (u) of the left eye image. l ,v l Matching points (u) of the right eye image and the right eye image r +Dr1,v r ), feature points of the left eye image (u l ,v l Matching points (u) of the right eye image and the right eye image r +Dr2,v r The matching candidate set is composed of the following.

[0077] In step 3, the method for obtaining sub-pixel level matching pairs is as follows:

[0078] To obtain more accurate sub-pixel level matching pairs, further pixel fitting is needed for the matching points in the right eye image. This is because there is an optimal correction between the currently selected matching points in the right eye image (i.e., the matching points in step 2) and the optimal matching points in the right eye image. The closer the currently selected matching points are to the optimal matching points in the right eye image, the smaller the difference in pixel grayscale values ​​(i.e., the smaller the difference in grayscale values ​​between pixels in the left and right eye images). Conversely, the further away from the optimal matching points in the right eye image, the greater the difference in pixel grayscale values. Therefore, an error parabola can appear during the sliding window matching process of the feature points in the right eye image. The location where the minimum value of the error parabola appears is the location of the optimal matching point in the right eye image.

[0079] If the minimum value is not on the boundary of the error parabola, it means that no inflection point appears, then the matching point of the right eye image is still selected as the matching point of the present right eye image (i.e. the matching point of step 2); if the minimum value is not on the boundary of the error parabola, the best matching point corresponding to the minimum value is selected as the matching point of the right eye image, and the specific selection process is as follows:

[0080] The x coordinate corresponding to the bottom of the error parabola (minimum value) is the most accurate matching coordinate, i.e. the x coordinate of the best matching point, since the minimum value is necessarily in the vicinity of the best correction amount obtained in step 2, then the matching point (u r +Dr1,v r ) of the right eye image is taken as an example, the error parabola is constructed by using the coordinates of the three points before and after the best correction amount, i.e. (u r +Dr1), (u r +Dr1-1), (u r +Dr1+1) are taken as the horizontal axis, and the descriptor distance of the feature point (u l ,v l ) of the left eye image is taken as the vertical axis, and the error parabola equation is solved by bringing it into the parabola curve; by solving the minimum value of the error parabola, the sub-pixel level correction amount Dr1¢ can be obtained, and the final matching point coordinate of the right eye image after parabola fitting is (u r +Dr1+Dr1¢,v r );

[0081] Similarly, if there are two matching points on the right eye image, then the final matching point coordinate of the right eye image after parabola fitting of the other matching point (u r +Dr2,v r ) is (u r +Dr2+Dr2¢,v r );

[0082] Therefore, in this step, the matching pairs of sub-pixel matching are: the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image; or the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r ) of the right eye image; or the feature point (u l ,v l ) of the left eye image and the matching point (u r +Dr2+Dr2¢,v r ) of the right eye image.

[0083] In step 4, according to the binocular depth solving principle, the real depth of a pixel point can be obtained by formula (2):

[0084]

[0085] In the formula, Z is the real depth of the pixel point, the disparity d represents the coordinate difference between the two pixels, f is the focal length, and b is the pixel size.

[0086] The accuracy of the disparity d directly determines the accuracy of the real depth solving, and the depth in the scene also has a structural relationship. In combination with the characteristics of the visual sensor, a two-level depth rejection method based on pixels and planes is proposed.

[0087] When the feature points are rejected at the pixel level, first, all the scene depths solved are sorted according to the SAD matching error value, that is, the r i values are sorted; then the middle value (i.e. the median) of the r i values is taken as the average value of the matching error of the scene, and finally, according to the error characteristics of the visual sensor, the depth values greater than 2 times the average value of the matching error and the corresponding feature points in the current scene are rejected, so as to avoid introducing large errors to the algorithm by false depth matching;

[0088] When the feature points are rejected at the plane level, the image segmentation method based on K-means clustering is used to divide the image (including the left eye image and the right eye image) into different regions. The purpose of region segmentation is to divide the region range of different objects in the image, and the depths of different regions should have similarity or progression. Therefore, the image depth in each region is determined twice: if the depth value of a pixel point (i.e. a feature point) in the same region is depth, the depth value of the pixel point depth should be within the depth envelope range of the 8-neighborhood within its region. If it exceeds the depth envelope range, it means that the pixel point depth is a noise point, a false depth or a depth similar point. Therefore, if the depth value of the pixel point depth is greater than 1.5 times the depth envelope range or less than 0.5 times the depth envelope range, the matching result and the depth solving result corresponding to the pixel point are removed; otherwise, they do not need to be removed.

[0089] Finally, after completing the two-level rejection of the image, all the feature points in the right eye image are traversed. If the matching pair of the sub-pixel level matching only remains one, that is, the feature point (u l , v l ) of the left eye image and the matching point (u r +Dr1+Dr1¢,v r), the remaining one of the sub-pixel level matching matching pair is the final accurate matching pair; if the sub-pixel level matching matching pair is still two, the matching point of the right eye image still exists two, the matching point of the right eye image (u r +Dr1+Dr1¢,v r ) and the matching point of the right eye image (u r +Dr2+Dr2¢,v r ), the matching point of the right eye image (u r +Dr2+Dr2¢,v r ) is removed, and only the matching point of the right eye image (u r +Dr1+Dr1¢,v r ), the feature point of the left eye image (u l ,v l ) and the matching point of the right eye image (u r +Dr1+Dr1¢,v r ) are the final accurate matching pair;

[0090] Therefore, the final accurate matching pair after the removal is the feature point of the left eye image (u l ,v l ) and the matching point of the right eye image (u r +Dr1+Dr1¢,v r ); after obtaining the accurate matching of the feature points of the left and right eye images, the scene depth value can be obtained by using the triangulation technology, and the binocular depth accurate solution is realized.

[0091] For the sake of description, spatial relative terms, such as "above", "upper", "on top", "top", and the like, can be used herein to describe a spatial relationship between one device or feature and another device or feature as shown in the drawings. It should be understood that the spatial relative terms are intended to include different orientations of the device in use or operation in addition to the orientation depicted in the drawings. For example, if the device in the drawings is inverted, the device described as "above" or "on" the other device or structure will be positioned "below" or "under" the other device or structure. Thus, the exemplary term "above" can include both the "above" and "below" orientations. The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatial relative descriptions used herein are interpreted accordingly.

[0092] In addition, it should be noted that the use of the terms "first", "second" and the like to describe various components is merely intended to facilitate differentiation of the corresponding components, and the above terms do not have special meanings unless otherwise stated. Therefore, it cannot be understood as a limitation on the scope of protection of the present application.

[0093] The above only is the preferred embodiment of the present application, and is not used to limit the present application, for those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for binocular depth precision solving based on region segmentation, characterized in that, The steps of the method are: Step 1: After preprocessing the left-eye image and the right-eye image, feature extraction is performed on the left-eye image and the right-eye image to obtain feature points of the left-eye image and the right-eye image; Step 2: According to the obtained feature points of the left-eye image and the right-eye image, SAD sliding window matching method is used to match the feature points of the left-eye image and the right-eye image to obtain a matching pair; Step 3: The obtained matching pair is optimized by using a parabolic fitting method to obtain a sub-pixel level matching pair; Step 4: Based on a two-level depth rejection method of pixels and planes, pixel distribution is used to complete scene depth one-level rejection feature points, and region planes are used to complete scene depth two-level rejection feature points. After rejecting the feature points that do not meet the requirements, the final matching pair is obtained. Finally, based on the final matching pair, the scene depth value is obtained by using triangulation technology.

2. The method of claim 1, wherein, In step 1, the method of preprocessing the left-eye image and the right-eye image includes using histogram equalization and polar correction.

3. The method of claim 2, wherein, The process of using histogram equalization for image preprocessing is as follows: Firstly, the gray scale distribution of the left-eye image or the right-eye image at the same time t is calculated to determine whether to perform histogram equalization. The maximum pixel gray value is subtracted from the minimum pixel gray value in the left-eye image or the right-eye image. If the gray scale difference is within 1 / 4 of the image gray scale value range, it is determined that histogram equalization needs to be performed; otherwise, it is determined that histogram equalization does not need to be performed. Then, the gray scale mapping function of the left-eye image at time t is calculated, and the right-eye image is processed by the same gray scale mapping function to complete histogram equalization, so that the left-eye image and the right-eye image at the same time can be processed on the same basis.

4. The method of claim 1, wherein, In step 1, ORB feature extraction method is used to extract features from the preprocessed left-eye image and right-eye image.

5. The method of claim 1, wherein, In step 2, Step 2-1, constructing initial matching pairs for left and right eye images, the initial matching pairs being a feature point (u l ,v l ) on the left eye image and a matching point (u r ,v r ) on the right eye image; Step 2-2, the initial matching pair obtained in step 2-1 is adjusted by using SAD sliding window matching method: First, a pixel block is taken out in the left eye image, the pixel block is centered at the feature point (u l ,v l ), and the size of the pixel region is 11*11; then, a pixel block is also established in the right eye image, the pixel block is centered at the x-axis coordinate u r of the feature point, and slides along the x-axis of the feature point of the right eye image; finally, the sum r i of the absolute values of the differences of all pixel gray values between the pixel block of the left eye image and the sliding pixel block of the right eye image is calculated. wherein represents the gray value of the i-th pixel in the pixel block of the left eye image, represents the gray value of the i-th pixel in the pixel block of the right eye image; Record two minimum r i values, let the minimum value in the calculated r i value corresponds to the minimum error of the right eye deviation amount of Δr1, the second minimum value corresponds to the second minimum error of the right eye deviation amount of Δr2; then get the adjusted matching pair, the coordinates of the matching points on the right eye image of the adjusted matching pair are (u r +Δr1,v r ) and (u r +Δr2,v r ) respectively.

6. The binocular depth precision solving method based on region segmentation according to claim 5, wherein, In step 2-1, the process of constructing the initial matching pair is as follows: Let the x-axis coordinate of the feature point on the left eye image be u l , and the y-axis coordinate be v l , traverse all the feature points in the transverse search band of the right eye image, and calculate the descriptor distance between all the feature points and the feature point (u l , v l ) of the left eye image respectively, record the feature point coordinates (u r , v r ) of the right eye image with the minimum descriptor distance as the matching point on the right eye image; therefore, the feature point (u l , v l ) on the left eye image and the matching point (u r , v r ) on the right eye image form an initial matching pair.

7. The method of claim 5, wherein, After step 2-2, the matching points on the right-eye image of the adjusted matching pair are screened: compute a descriptor distance d1 between the feature point (u l ,v l ) of the left eye image and the matching point (u r +Δr1,v r ) of the right eye image and a descriptor distance d2 between the feature point (u l ,v l ) of the left eye image and the matching point (u r +Δr2,v r ) of the right eye image; If the difference between d1 and d2 is less than a predetermined threshold, the matching point of the right eye image retains both coordinates, i.e. (u r +Δr1,v r ) and (u r +Δr2,v r ); otherwise, only the matching point of the right eye image (u r +Δr1,v r ) is retained.

8. The method of claim 7, wherein, In step 3, the process of obtaining a sub-pixel level matching pair is as follows: Since there is a best correction amount between the matching points of the right-eye image in step 2-2 and the best matching points of the right-eye image, an error parabola can occur in the sliding window matching process of the feature points of the right-eye image. The position of the minimum value of the error parabola is the position of the best matching point of the right-eye image. If the minimum value is not at the boundary of the error parabola, it means that there is no inflection point, so the best matching point is discarded and the matching point of the right-eye image in step 2-2 is selected. If the minimum value is not at the boundary of the error parabola, the best matching point corresponding to the minimum value position is selected as the matching point of the right-eye image.

9. The binocular depth precision solving method based on region segmentation according to claim 8, wherein, The process of selecting the best matching point corresponding to the minimum value position of the error parabola as the matching point of the right-eye image is as follows: If there is a matching point (u r +Δr1,v r ) in the right eye image, then construct an error parabola with the matching points (u r +Δr1), (u r +Δr1-1), (u r +Δr1+1) of the right eye image as the horizontal axis and the descriptor distance of the feature point (u l ,v l ) of the left eye image as the vertical axis; by solving the minimum value of the error parabola, the sub-pixel level correction amount Δr′1 can be obtained, and the final matching point coordinates of the right eye image after parabola fitting are (u r +Δr1+Δr′1,v r ). Similarly, if there are two matching points on the right eye image, the other matching point (u r + Δr2, v r ) is obtained after parabolic fitting. The final matching point coordinates of the right eye image are (u r + Δr2+ Δr′2, v r ). Thus, the matching pairs of sub-pixel level matching are: the feature point (u l ,v l ) of the left eye image and the matching point (u r +Δr1+Δr′1,v r ) of the right eye image; or the feature point (u l ,v l ) of the left eye image and the matching point (u r +Δr1+Δr′1,v r ) of the right eye image; or the feature point (u l ,v l ) of the left eye image and the matching point (u r +Δr2+Δr′2,v r ) of the right eye image.

10. The binocular depth precision solving method based on region segmentation according to claim 9, wherein, In step 4, When the feature points are removed at the pixel level, first, the solved r i values are sorted; then the median of the r i values is taken as the average matching error of the scene, and finally the depth values greater than 2 times the average matching error of the current scene and the corresponding feature points are removed; In the plane level, the left and right eye images are divided into different regions by using the K-means clustering based image segmentation method; the image depth in each region is determined in two levels: if the depth value of a pixel in the same region is depth, the depth value of the pixel depth should be within the depth envelope of the 8-neighborhood in its region; if the depth value of the pixel depth is greater than 1.5 times of the depth envelope or less than 0.5 times of the depth envelope, the matching result and the depth solving result corresponding to the pixel are removed; otherwise, the matching result and the depth solving result corresponding to the pixel are not removed. Finally, when the secondary culling of the image is completed, all feature points in the right eye image are traversed, if the sub-pixel level matching matching pair only remains one, that is, the feature point (u l ,v l ) of the left eye image and the matching point (u r +Δr1+Δr′1,v r ) of the right eye image, then the remaining one sub-pixel level matching matching pair is the final accurate matching pair; if the sub-pixel level matching matching pair is still two, that is, the matching point (u r +Δr1+Δr′1,v r ) of the right eye image and the matching point (u r +Δr2+Δr′2,v r ) of the right eye image, then the matching point (u r +Δr2+Δr′2,v r ) of the right eye image is culled, only the matching point (u r +Δr1+Δr′1,v r ) of the right eye image is retained, then the feature point (u l ,v l ) of the left eye image and the matching point (u r +Δr1+Δr′1,v r ) of the right eye image are the final accurate matching pair. Thus, the final matched pair after the elimination is the feature point of the left eye image (u l ,v l ) and the matched point of the right eye image (u r +Δr1+Δr′1,v r ); after the matching of the feature points of the left and right eye images, the scene depth value can be obtained by using the triangulation technique, thereby realizing the accurate solution of the binocular depth.