Depth map optimization method and system based on binocular vision
By employing a depth map optimization method based on binocular vision, and by selecting pixel pairs with consistent structural orientations, statistically matching density confidence, and performing depth value replacement, the method addresses the insufficient accuracy of depth maps in edge regions and parallax abrupt change regions in existing technologies, thereby achieving more accurate parallax field reconstruction and depth map optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-27
AI Technical Summary
Existing depth map optimization techniques are prone to causing blurring of contours and loss of structural details when processing image edge regions. They are particularly difficult to accurately recover true depth information in areas with complex image textures or abrupt parallax changes. Furthermore, they are prone to introducing erroneous depth propagation under high-frequency noise interference. They also lack fine modeling of the directional structural consistency between matching pixels, resulting in insufficient depth maps in terms of boundary representation, transition smoothness, and structural integrity.
By acquiring left and right view image frames from binocular images, extracting the gradient direction dominant frequency value, filtering pixel pairs with consistent structural orientation, calculating the matching density confidence, performing pixel-level directional consistency connectivity expansion, calculating the absolute depth difference, filtering jump points for depth value replacement, generating a boundary connectivity correction region image matrix, and finally reconstructing the disparity field image.
It improves the overall performance of depth maps in terms of preserving boundary details, repairing abrupt changes, and expressing disparity map accuracy, thus enhancing the overall edge detail preservation and disparity map accuracy of depth maps.
Smart Images

Figure CN121746207A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image analysis technology, and in particular to a depth map optimization method and system based on binocular vision. Background Technology
[0002] Image analysis technology primarily involves methods for recognizing, understanding, and processing image content, including image acquisition, preprocessing, feature extraction, matching and recognition, and higher-order information reasoning and reconstruction. It is widely applied in scenarios such as security monitoring, medical imaging, autonomous driving, and industrial inspection. Typically, it uses image acquisition, image enhancement, edge detection, region segmentation, target recognition and tracking, and 3D reconstruction to perform quantitative or qualitative analysis of image information. Traditional depth map optimization methods refer to a series of post-processing techniques employed to improve the accuracy and integrity of the initial depth map generated by a binocular camera or other multi-view devices. Common methods include edge-preserving denoising based on bilateral filtering, region filling based on disparity confidence, weighted averaging strategies using guided images for joint optimization, path optimization using cost aggregation mechanisms to finely adjust pixel correspondences, window adaptive interpolation techniques based on local structural consistency, and constrained smoothing incorporating image gradient information. These methods typically aim to improve the accuracy of depth map edges, reduce hole areas, and lower matching errors.
[0003] Existing depth map optimization techniques, such as bilateral filtering, weighted averaging, cost aggregation, or guided image optimization, often suffer from blurred contours and missing structural details when processing image edge regions. This is especially true in areas with complex textures or abrupt disparity changes, where it is difficult to accurately recover true depth information. Confidence assessment mechanisms relying on global or local statistics suffer from inaccurate matching point selection in sparse feature regions, resulting in the inability to effectively fill in holes in the depth map. Methods based on window interpolation or gradient smoothing are prone to introducing erroneous depth propagation under high-frequency noise interference. Furthermore, the lack of refined modeling methods for the directional structural consistency between matched pixels makes it impossible to guarantee the continuity of connectivity between pixels. This can easily lead to discontinuities and error accumulation at disparity jump locations, resulting in deficiencies in the final depth map in terms of boundary representation, transition smoothness, and structural integrity. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a depth map optimization method based on binocular vision.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a depth map optimization method based on binocular vision, comprising the following steps: S1: Obtain left and right view image frames from the binocular image, extract the gradient direction main frequency value of each pixel to be matched within the image block, filter all pixel pairs whose gradient direction angle is lower than the direction tolerance angle, and generate a list of pixel pairs with consistent structural directions. S2: Statistically analyze the spatial distribution of all pixel pairs in the list of pixel pairs with consistent structural orientation, construct a sliding window with each pair of matching points as the center, calculate the confidence of pixel matching density with the same disparity value as the matching point within the sliding window, determine the concentrated matching points and retain them as reliable references, and generate a list of distributed aggregated matching points. S3: Based on the pixel positions marked in the left figure of the distributed aggregation matching point list, extract the main gradient direction starting from the edge protruding pixel, perform pixel-level directional consistent connectivity expansion, retain pixels with included angles within the tolerance range to construct a boundary structure sequence, and generate a set of boundary coherent structure pixel paths. S4: Based on the initial depth map values of all pixel regions in the set of pixel paths of the boundary coherent structure, calculate the absolute depth difference between adjacent points in the continuous pixel group, filter and mark the jump points, count the points in the path where the adjacent pixels of the jump points satisfy the direction consistency and continuity, construct the reference set of the region connectivity graph, perform value replacement processing on all jump points, and generate the boundary connectivity correction region image matrix. S5: Extract the depth information of the location points in the boundary connectivity correction region image matrix and the distribution aggregation matching point list, perform disparity field image reconstruction, and obtain the binocular visual depth map optimization results.
[0006] As a further aspect of the present invention, the concentrated matching point specifically refers to the matching point where the pixel matching density confidence level is greater than the confidence density reference frequency standard. The jump point specifically refers to the location point where the absolute depth difference is greater than the set jump reference difference.
[0007] As a further aspect of the present invention, the list of structurally consistent pixel pairs includes pixel pairs consistent in the main frequency direction and structurally matched pairs filtered by the included angle tolerance; the list of distributed aggregated matching points includes confidence matching points and spatially consistent matching region points; the set of boundary-connected structural pixel paths includes gradient main direction continuous paths and direction-consistent pixel chains; the boundary connectivity correction region image matrix includes trend-adjusted jump point depth values, depth continuity optimized regions, and boundary consistency enhanced image patches; and the binocular vision depth map optimization results include reconstructed disparity field images, encapsulated optimized image formats, and archived resource datasets.
[0008] As a further aspect of the present invention, the step of obtaining the list of pixel pairs with consistent structural orientation specifically comprises: S111: Obtain the left-view image frame and the right-view image frame in the binocular image, and extract the gradient direction values of all pixels in the image block for each pixel to be matched, construct the gradient direction set of the image block, perform frequency statistics operation on each set, count the number of times each gradient direction value appears in the image block, and determine the dominant frequency direction value based on the statistical results, and generate a list of dominant frequency values of gradient direction. S112: Based on the gradient direction main frequency value list, read the main frequency direction value of the image block corresponding to the pixel to be matched at the same position in the left and right view image frames, calculate the direction angle and compare the difference with the structural direction tolerance angle. When the direction angle is less than the structural direction tolerance angle, record the pixel index position of the corresponding pixel pair to be matched and obtain the direction consistent index pair list. S113: Based on the list of direction-consistent index pairs, extract the pixel position parameters of each index pair in the left and right view image frames, and combine and aggregate each pair of pixel pairs that meet the direction consistency condition to generate a list of structural direction-consistent pixel pairs.
[0009] As a further aspect of the present invention, the step of obtaining the distributed aggregation matching point list specifically includes: S211: Based on the list of pixel pairs with consistent structural orientation, a sliding window is constructed in the left and right views of the binocular image, with the left view pixel of each pixel pair as the center. The pixel index of each position in the window is extracted one by one and the existence of a corresponding pixel pair is checked. If it exists, the disparity value of the corresponding matching pixel pair in the list is obtained and recorded to generate a sliding window disparity information matrix. S212: Based on the disparity value data of all records in each window centered on the matching point in the sliding window disparity information matrix, compare the disparity value corresponding to the center pixel with the disparity value corresponding to other pixels in the same window item by item, count the number of matching points with consistent disparity, calculate and obtain the pixel matching density confidence value based on the coverage frequency of the matching point in the window area, if it is greater than the confidence density reference frequency standard threshold, then the current matching point is determined as a reliable matching point, and obtain the reliable matching point index list; S213: Based on the list of reliable matching points, extract the coordinate values in the left and right views corresponding to all pixel pairs that meet the density confidence screening conditions, and organize and output them in a pixel-pair alignment manner to establish a list of distributed aggregated matching points.
[0010] As a further aspect of the present invention, the formula for calculating the pixel matching density confidence value is as follows: ; in, Indicates the first Density confidence of each matching point Represents the disparity value of the matching point itself, Indicates the first in the same sliding window disparity value per pixel This represents the total number of pixels within the sliding window that have matching points.
[0011] As a further aspect of the present invention, the step of obtaining the set of pixel paths of the boundary coherent structure specifically includes: S311: Based on the left image pixel positions marked in the distribution aggregation matching point list, detect the edge positions where there is a gradient change in the gray-scale distribution of each pixel, extract the gray-scale gradient direction angle of each pixel in the corresponding image block, and count the frequency distribution of each direction angle. Determine the dominant direction angle of the current pixel based on the item with the maximum frequency in the direction angle interval, and generate a list of dominant direction angles of edge pixels. S312: According to the list of main direction angles of edge pixels, for each main direction angle, a step-by-step pixel extension search is performed with the corresponding direction as the axis. The local gradient direction angle value of each pixel on the search path is detected, and it is determined whether the angle with the starting direction falls within the direction consistency tolerance range. When the angle meets the tolerance condition, the corresponding pixel position is recorded and the connection expansion continues to obtain the direction consistency connected pixel set. S313: Based on all pixel indices in the directionally consistent connected pixel set, extract the position sequence in the image, perform path structure aggregation processing according to the geometric order on the connected chain, encode all path segments in sequence according to the connection order, establish a complete boundary trajectory data structure, and generate a boundary coherent structure pixel path set.
[0012] As a further aspect of the present invention, the step of obtaining the boundary connectivity correction region image matrix specifically comprises: S411: Based on each group of pixel paths recorded in the set of pixel paths of the boundary coherent structure, read the grayscale depth value of each pixel in the original image depth map, construct continuous pixel groups according to the path order, calculate the absolute depth difference between two adjacent pixels in sequence, count the gradient change rate values of all pixels in the horizontal and vertical directions in the entire original depth map, calculate the average value of the change rate, take twice as the jump reference difference, mark the position of the points in each group of paths whose depth difference is greater than the jump reference difference, and generate a jump pixel index mark matrix; S412: Based on the path index position of all jump points in the jump pixel index marker matrix, retrieve the adjacent pixels in each path, and determine whether the direction between the adjacent point and the jump point is consistent with the main direction of the path. If the direction difference is within the set tolerance threshold, then aggregate the corresponding jump point and its adjacent points into a direction-continuous point group, and record all point groups that meet the direction continuity condition to generate a direction-continuous jump reference point set. S413: Based on the set of reference points for continuous directional jumps, extract the depth values of adjacent reference points corresponding to each jump point, and use the adjacent points as a benchmark to determine the trend of depth change before and after. If the trend is continuous and the depth value falls within the reference point change range, then replace the depth value of the jump point, update it by using the linear weighted average of the depth values of adjacent reference points, and establish the boundary connectivity correction region image matrix.
[0013] As a further aspect of the present invention, the steps for obtaining the binocular vision depth map optimization results are specifically as follows: S511: Based on the image matrix of the boundary connectivity correction region and the position index information in the distribution aggregation matching point list, extract the depth value of each matching point in the image matrix, and perform point-by-point aggregation processing in combination with the disparity value of the corresponding point to construct a disparity-depth bidirectional mapping structure. According to the image coordinate system rules, rearrange each mapping pair in a two-dimensional matrix format and fill it into the disparity mapping frame in the complete spatial range to generate a spatial structure disparity reconstruction matrix. S512: Obtain the spatial structure disparity reconstruction matrix, package all pixels in column-major order, perform three-channel pseudo-color encoding according to the image output specification, map the disparity information into an RGB distribution image, compress and encapsulate the encoded image to obtain a standard encapsulated disparity image data frame. S513: Based on the spatial location index corresponding to the standard encapsulated parallax image data frame, construct an image index dictionary and merge it with the source image number mapping structure to form an archiveable multi-source image index tuple, establish a structured data archive record at the system level, and obtain the binocular vision depth map optimization results.
[0014] A depth map optimization system based on binocular vision, comprising: The orientation consistency screening module is used to perform S1: acquire left and right view image frames in the binocular image, extract the gradient direction main frequency value in the image block where each pixel to be matched is located, filter all pixel pairs whose gradient direction angle is lower than the orientation tolerance angle, and generate a list of pixel pairs with consistent structural orientation. The confidence density detection module is used to perform S2: statistically analyze the spatial distribution of all pixel pairs in the list of pixel pairs with consistent structural orientation, construct a sliding window with each pair of matching points as the center, statistically analyze the matching density confidence of pixels with the same disparity value as the matching points within the sliding window, determine the concentrated matching points and retain them as reliable references, and generate a list of distributed aggregated matching points. The boundary structure tracking module is used to perform S3: based on the pixel positions marked in the left figure of the distributed aggregation matching point list, extract the main gradient direction starting from the edge protruding pixel, perform pixel-level direction consistency connectivity expansion, retain pixels with included angles within the tolerance range to construct a boundary structure sequence, and generate a set of boundary coherent structure pixel paths; The jump region correction module is used to perform S4: based on the initial depth map values of all pixel regions in the set of pixel paths of the boundary coherent structure, calculate the absolute depth difference between adjacent points in the continuous pixel group, filter and mark jump points, count the points that satisfy the direction consistency and continuity of adjacent pixels on the path of the jump points to construct a reference set of the region connectivity graph, perform value replacement processing on all jump points, and generate the boundary connectivity correction region image matrix. The depth map integration and output module is used to perform S5: extract the depth information of the location points in the boundary connectivity correction region image matrix and the distribution aggregation matching point list, perform disparity field image reconstruction, and obtain the binocular vision depth map optimization results.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, the structural consistency of matched pixel pairs is improved by constructing gradient direction main frequency screening combined with directional angle tolerance rules; low confidence distribution points are eliminated by sliding window frequency statistics to enhance matching reliability; a complete boundary structure is established by directional consistency connectivity extension starting from edge pixels; the structural correlation of jump regions is strengthened by depth difference jump detection and directional continuity point group construction; depth value replacement is performed by combining the trend relationship within the path to enhance the naturalness of image boundary transition; and more accurate disparity field reconstruction is achieved by fusing the information of the corrected region with the reliable distribution points. Overall, the comprehensive performance of the depth map in terms of boundary detail preservation, jump region repair, and disparity map accuracy expression is improved. Attached Figure Description
[0016] Figure 1 This is a flowchart of the main steps of the present invention; Figure 2 This is a flowchart of the process for obtaining the list of pixel pairs with consistent structural orientation in this invention. Figure 3 This is a flowchart illustrating the process of obtaining the distributed aggregation matching point list in this invention. Figure 4 This is a flowchart of the process for obtaining the pixel path set of the boundary coherent structure in this invention; Figure 5 This is a flowchart of the process for obtaining the image matrix of the boundary connectivity correction region in this invention; Figure 6 This is a flowchart of the process for obtaining the optimization results of the binocular vision depth map in this invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0018] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0019] Please see Figure 1 A depth map optimization method based on binocular vision includes the following steps: S1: Obtain left and right view image frames from the binocular image, extract the gradient direction main frequency value within the image block where each pixel to be matched is located, and perform orientation screening on all pixel pairs according to the screening rule that the gradient direction angle is lower than the orientation tolerance angle, and generate a list of pixel pairs with consistent structural orientation. S2: Analyze the spatial distribution of all pixel pairs in the list of pixel pairs with consistent structural orientation. Construct a 5×5 pixel sliding window centered on each pair of matching points in the binocular image. Calculate the confidence level of pixel matching density within the sliding window. Compare the confidence level of pixel matching density with the confidence density benchmark. When the confidence level of pixel matching density is greater than the confidence density benchmark frequency standard, it is determined to be a concentrated matching point. Only such matching points are retained as reliable references, and a list of distributed aggregated matching points is generated. S3: Based on the pixel positions marked in the left image of the distributed aggregation matching point list, extract the main gradient direction starting from the edge protruding pixel, and perform pixel-level directional consistent connectivity expansion in the image with the corresponding direction as a reference. Retain pixels with included angles within the tolerance range to construct a boundary structure sequence and generate a set of boundary coherent structure pixel paths. S4: Based on the initial depth map values of the regions where all pixels are located in the pixel path set of the boundary coherent structure, calculate the absolute depth difference between adjacent points in the continuous pixel group, select the points whose absolute difference is greater than the set jump benchmark difference (the setting is based on twice the average depth gradient change value of the whole image) and mark them as jump points, count the points in the path where the jump points are adjacent pixels that satisfy the direction consistency continuity, construct the region connectivity map reference set, combine the depth trend information of adjacent reference pixels, perform value replacement processing on all jump points, and generate the boundary connectivity correction region image matrix; S5: Extract the depth information of the location points in the image matrix of the boundary connectivity correction region and the distribution aggregation matching point list, spatially integrate the disparity information, reconstruct the disparity field image, and complete the image format encapsulation and resource archiving operations to obtain the binocular vision depth map optimization results.
[0020] The list of structurally consistent pixel pairs includes pixel pairs consistent in the main frequency direction and structurally matched pairs filtered by the included angle tolerance. The list of distributed aggregated matching points includes confidence matching points and spatially consistent matching region points. The set of boundary-connected structural pixel paths includes gradient main direction coherent paths and direction-consistent pixel chains. The boundary connectivity correction region image matrix includes trend-adjusted jump point depth values, depth continuity optimized regions, and boundary consistency enhanced image patches. The binocular vision depth map optimization results include reconstructed disparity field images, encapsulated optimized image formats, and archived resource datasets.
[0021] Please see Figure 2 Step S1 is as follows: S111: Obtain the left-view image frame and the right-view image frame in the binocular image, and extract the gradient direction values of all pixels in the image block for each pixel to be matched, construct the gradient direction set of the image block, perform frequency statistics operation on each set, count the number of times each gradient direction value appears in the image block, and determine the dominant frequency direction value based on the statistical results, and generate a list of dominant frequency values of gradient direction. To acquire the left-view and right-view image frames from the binocular images, it is necessary to synchronously acquire data from the two image channels output by the binocular camera system at the same time point, ensuring that the spatial correspondence between the two image frames is consistent. The left and right image frames are denoted as follows: and After acquiring the image frame, a sliding window method needs to be used to determine the image block region for each pixel to be matched in the image. The image block size is set to [size missing]. Pixels, using this size can reduce boundary interference while ensuring the ability to represent local features. For example, for the first pixel in the left image frame... Pixel, extract the center of that pixel. Image Patch Extract the image patch that is parallax-aligned with the right image frame. Then, gradient orientation angle extraction is performed on each pixel within each image block. The orientation angle is calculated based on the Sobel operator to perform horizontal and vertical gradient calculations. The orientation angle is defined as: ; in, and These represent the gradient values of the image in the horizontal and vertical directions, respectively, and can be obtained through convolution operations. For example, the Sobel horizontal gradient of image pixel (i, j) is: ; The orientation angle of each pixel in the image patch is statistically analyzed and divided into 8 orientation intervals: 0° to 22.5°, 22.5° to 45°, ..., 157.5° to 180°. The frequency count for each interval is incremented by 1, forming an orientation angle frequency distribution vector. For example, in the image patch... The frequency vector obtained after extracting the orientation angle of the inner pixel is: ; The value 7 appears in the second interval, corresponding to a dominant frequency direction angle of 45°. The dominant frequency direction angle is defined as the center angle of the direction interval with the highest frequency in the frequency vector. If multiple directions have the same frequency, the angle interval closest to the line connecting the image center is selected. Finally, a list of dominant frequency direction angles for all image blocks in the image frame is constructed. Each record contains the center pixel coordinates of the image block and the dominant frequency direction angle information.
[0022] To verify the stability of the directional angular frequency statistics, 1000 image blocks from the stereo image frames were sampled for directional dominant frequency analysis. The proportion of each dominant frequency direction is shown in the table below: Table 1. Statistics of dominant frequency direction of image blocks; As shown in Table 1, the distribution of the dominant frequency direction angles of different image blocks is different, which verifies the diversity of the spatial distribution of the internal structure of the image blocks. This table structure is also used for subsequent direction consistency screening of reference angle values, and finally a list of gradient direction dominant frequency values is obtained.
[0023] S112: Based on the gradient direction main frequency value list, read the main frequency direction value of the image block corresponding to the pixel to be matched at the same position in the left and right view image frames, calculate the direction angle and compare it with the difference with the structural direction tolerance angle. When the direction angle is less than the structural direction tolerance angle, record the pixel index position of the corresponding pixel pair to be matched and obtain the direction consistent index pair list. Based on the list of gradient direction dominant frequency values, the dominant frequency direction values of the image blocks corresponding to the pixels to be matched at the same position in the left and right view image frames are called, and denoted as follows: and ,in The current assumed disparity value is typically initialized to 4 pixels. The direction angle is calculated as follows: ; To determine structural orientation consistency, a structural orientation tolerance angle needs to be set. This value is used to determine whether two direction angles can be considered approximately the same. Based on experiments, the angle tolerance value is selected as 15°. That is, if the angle between directions is less than or equal to 15°, it is determined that the directions are the same. This value is based on the statistical results of direction perturbation within the image patch, and the setting is as follows: The statistical range of orientation angle differences in the sampled image blocks is [0°~45°]. Approximately 82% of pixel pairs have similar structural orientations within a 15° range. Therefore, 15° is used as the criterion. A specific example is used to illustrate the calculation of the orientation angle. , ,but Record the coordinates of this pixel. and Enter the list of pixel pairs with the same orientation.
[0024] After processing all candidate pixel pairs in the image, the directional angle value is compared with the tolerance angle value to construct an index set that meets the screening criteria, i.e., a list of index records with consistent orientation, in the form of a two-dimensional index matrix: ; To better illustrate the index record situation, the first 5 pairs of pixels with consistent orientation are extracted and indexed as follows: Table 2 Index table of pixel pairs with consistent orientation; As shown in Table 2 above, the final list of index pairs with consistent directions is generated.
[0025] S113: Based on the list of index pairs with consistent orientation, extract the pixel position parameters of each index pair in the left and right view image frames, and combine and aggregate each pair of pixel pairs that meet the orientation consistency condition to generate a list of pixel pairs with consistent orientation. Based on the list of indexed pairs with consistent orientation, the coordinate positions of each pixel pair that meets the filtering criteria are extracted in the left and right view image frames. These are combined to form structural matching pairs, each consisting of two items: the left pixel position and the right pixel position. The aggregation operation is performed sequentially according to the index order, importing the index pair content into the matching pair structure, in the following form: ; in This represents the number of pixel pairs that meet the orientation consistency condition. For the first The corresponding disparity values for pixel pairs provide a structural matching basis for subsequent stereo matching construction. The results of the first 3 sampled combinations are as follows: Table 3 shows the combination results. As can be seen from Table 3, structural orientation consistency is maintained, and all matching pairs meet the set angle tolerance standard, indicating that they have similar structural attributes. Thus, a list of pixel pairs with consistent structural orientation is finally generated.
[0026] Please see Figure 3 Step S2 is as follows: S211: Based on the list of pixel pairs with consistent structural orientation, a sliding window is constructed in the left and right views of the binocular image, with the left view pixel of each pixel pair as the center. The pixel index of each position in the window is extracted one by one and the existence of the corresponding pixel pair is checked. If it exists, the disparity value of the corresponding matching pixel pair in the list is obtained and recorded to generate the sliding window disparity information matrix. Based on the list of pixel pairs with consistent structural orientation, it is necessary to explicitly extract the center pixel coordinates of the left view image in each pixel pair of the binocular image. When constructing a 5×5 sliding window, it is essential to ensure that the window coverage is fully embedded within the image boundary to avoid edge truncation. For example, when the pixel coordinates of the left image are (120, 80), the boundary range of the sliding window is x∈[118, 122], y∈[78, 82], consisting of 25 pixels. Next, using all pixels within this window as the center, the list of pixel pairs with consistent structural orientation should be called to search for whether there is a corresponding matching point. If it exists, obtain the x-coordinate of the corresponding right image matching pixel and calculate the difference between it and the left image coordinate. This difference is the disparity value of the pixel pair. Record the disparity value at the corresponding position in the current window, and finally form a 25-bit two-dimensional array matrix. This matrix is used to record the disparity information of each pixel in the sliding window. To avoid interference from null values, the disparity value at the position corresponding to the unmatched pixel is marked as -1. This process needs to be performed once for all center matching points, and the result of each sliding window operation is added to the cache matrix list in sequence for subsequent calculation of window density frequency.
[0027] S212: Based on the disparity value data of all records within each window centered on the matching point in the sliding window disparity information matrix, compare the disparity value corresponding to the center pixel with the disparity values corresponding to other pixels within the same window item by item, count the number of matching points with consistent disparity, and use the formula based on the coverage frequency of the matching point in the window area: ; The pixel matching density confidence value is calculated. If it is greater than the confidence density reference frequency standard threshold, the current matching point is determined as a reliable matching point, and a list of reliable matching points is obtained. Indicates the first Density confidence of each matching point Represents the disparity value of the matching point itself, Indicates the first in the same sliding window disparity value per pixel This represents the total number of pixels within the sliding window that have matching points. Obtain a 5×5 sliding window matrix centered on the matching point from the disparity information matrix of the sliding window. Perform a difference operation on the disparity value of the center pixel and the disparity values of the other 24 pixels. Before performing the operation, invalid values of -1 must be removed from the sliding window disparity matrix. Let the disparity value of the center pixel be... Other effective parallax points are If there are a total of n items, then each item needs to be calculated individually. Compare each difference with The difference is normalized, and the average value is obtained by summation. Then, a density index is constructed by reverse mapping. This operation is used to measure the degree of disparity aggregation within the window. This density index is the pixel matching density confidence value, which is calculated using a formula.
[0028] This reflects the degree of parallax deviation from the center point; the absolute value is used to avoid the cancellation of positive and negative differences. To avoid division by zero errors caused by parallax of 0, the normalization baseline is set to be no less than 1. The calculation example is as follows: Assuming a central pixel has a disparity of 8, and the number of other valid matching points within the sliding window is n=4, with disparity values of 7, 9, 8, and 6 respectively, then: ; The confidence level of the matching point density is 0.875. If the current set confidence density reference frequency standard is 0.80, then the point is determined to be a reliable matching point.
[0029] To clarify the rationality of the confidence screening criteria, a density index was calculated for 500 randomly matched points in the sample images. The results are shown in the table below: Table 4. Confidence values for matching point density; As shown in Table 4, the density confidence value directly affects the screening effect of reliable points. Further, the above screening rules were applied to all matching points to establish a list of reliable matching points. The results show that matching points with a confidence value greater than 0.80 will be retained for subsequent matching result aggregation.
[0030] The formula for calculating the density confidence index used in the formula is as follows: The overall calculation logic of this formula is based on the measurement of relative deviation by considering the disparity consistency of each matching point within the sliding window, and by using the disparity value of the center matching point. disparity value with all other valid pixels in the same window Item-by-item difference calculations are performed, and the absolute values are taken to reflect the degree of consistency with the center point disparity. The absolute value operation ensures that all differences are non-negative, and is used to measure the magnitude of deviation; whether greater than or less than the center disparity, it is considered a deviation. Subsequently, a normalization factor is applied. Normalizing the differences involves introducing a minimum value of "1" in the normalization factor to avoid division by zero when the central disparity is 0, ensuring numerical stability during the calculation process. The normalized value is subtracted from 1, meaning the closer to the central disparity, the smaller the difference, the higher the confidence value, and consequently, the greater the overall density contribution. Summing all difference contributions accumulates the support for the current central pixel's disparity consistency within the entire sliding window, and then dividing by the number of effective matching points. By obtaining the average value, the influence of window size or the number of matching points on the results is eliminated. Therefore, this formula can effectively measure the disparity density concentration of a matching point within its local window. The advantage of the formula lies in the fact that by introducing multiple operations such as normalized bias measurement, difference consistency mapping, and local averaging, a confidence index that is sensitive to disparity consistency and has a stable numerical range is constructed, thereby giving the selection of reliable matching points a strong discriminative ability.
[0031] The pixel matching density confidence score measures the degree of disparity consistency of a pixel matching pair within its local spatial range. Its value reflects whether the matching point is located in a region with a relatively stable and convergent disparity structure within the current sliding window. Specifically, this confidence score measures the relative difference between the disparity value of the central pixel and the disparity values of other valid matching points within the same window. When the disparities of surrounding pixels are similar to those of the central pixel, it indicates that the depth variation within the region is relatively smooth, and the distribution of matching points has spatial coherence, thus resulting in a high confidence score. Conversely, if the disparities of surrounding pixels fluctuate greatly or differ significantly from the central pixel, the confidence score is low, indicating that the matching structure of this point may be unstable or contain errors. The value of this index typically ranges from 0 to 1. The closer the value is to 1, the stronger the disparity consistency and the higher the reliability of the match. Therefore, the pixel matching density confidence score serves as an important criterion for selecting reliable pixel matching points in this scheme.
[0032] S213: Based on the list of reliable matching points, extract the coordinate values in the left and right views corresponding to all pixel pairs that meet the density confidence screening conditions, and organize and output them in a pixel-pair alignment manner to establish a list of distributed aggregated matching points. Based on the coordinate information of the trusted matching points recorded in the trusted matching point index list, the corresponding pixel coordinate pairs in the left view image frame and the right view image frame are extracted. The left and right pixels are stored in the form of structural matching pairs and arranged in order to be output as a list structure. The matching coordinate combination is represented in the form of a two-dimensional array. For each pair of trusted matching points, the left view index and right view index information are recorded in the list. For example, the corresponding coordinate combinations of trusted index pairs (120, 80)-(124, 80), (75, 95)-(79, 95) are added to the output set in turn, and finally a distributed aggregated matching point list is established.
[0033] Please see Figure 4 Step S3 is as follows: S311: Based on the left image pixel positions marked in the distribution aggregation matching point list, detect the edge positions where there are gradient abrupt changes in the gray-level distribution of each pixel, extract the gray-level gradient direction angle of each pixel in the corresponding image block, and count the frequency distribution of each direction angle. Determine the dominant direction angle of the current pixel based on the maximum frequency item in the direction angle interval, and generate a list of dominant direction angles of edge pixels. Based on the pixel positions marked on the left-view image in the distributed aggregation matching point list, the grayscale value of each pixel is first obtained, and a 3×3 image block region centered on that pixel is constructed around it. The grayscale change rate of each pixel within this image block is calculated to obtain the gradient values in the horizontal and vertical directions. The gradient calculation uses a fixed Sobel operator for convolution operation, with the horizontal gradient being Gx and the vertical gradient being Gy. Furthermore, the gradient direction angle value is calculated for each pixel. The results are normalized to the range of 0~180 degrees and recorded in the orientation angle array. Then, the distribution frequency of all orientation angle values is counted, and they are sorted according to the number of times each orientation angle appears. The orientation angle with the highest frequency is taken as the dominant orientation angle of the current center pixel. For example, if a pixel is located at (145, 88), the orientation angle distribution in its corresponding image block is as follows: [0°: 1 time, 45°: 3 times, 90°: 7 times, 135°: 4 times]. Then the dominant orientation angle is 90°. After performing the above processing operation on all matching points, a list of dominant orientation angles corresponding to their coordinate positions is established. The list structure is a two-dimensional array, and each element contains pixel coordinates and corresponding dominant orientation angle information, thus generating the list of dominant orientation angles of edge pixels.
[0034] S312: Based on the list of main direction angles of edge pixels, for each main direction angle, perform a step-by-step pixel extension search with the corresponding direction as the axis, detect the local gradient direction angle value of each pixel on the search path, and determine whether the angle with the starting direction falls within the direction consistency tolerance range. When the angle meets the tolerance condition, record the corresponding pixel position and continue to connect and expand to obtain the set of direction-consistent connected pixels. The system calculates the principal orientation angle values for each pixel in the list of principal orientation angles for edge pixels. For each principal orientation angle, a search vector direction is defined. A step-by-step search is performed in the original image along the path indicated by the orientation angle, with a step size of 1 pixel and a maximum extension distance of 15 pixels. After each step, the gradient orientation angle value of the current pixel is obtained and its angle with the principal orientation angle of the starting pixel is calculated. The angle is defined as the absolute value of the difference between the two orientation angles. A directional consistency tolerance threshold of 20 degrees is set; if the angle is less than or equal to 20 degrees, the pixel is considered to be consistent with the principal orientation direction and is retained. If the index position does not meet the condition, the path direction expansion stops. For example, if the initial main direction is 90°, and the expanded pixel direction angle is 108°, then the included angle is 18°, which meets the tolerance condition, and the expansion continues. If at a certain step the pixel direction angle is 67°, then the included angle is 23°, which does not meet the threshold condition, and the direction expansion terminates. Following this logic, the connectivity operation is expanded pixel by pixel, and each pixel path that meets the direction consistency is structurally recorded. Finally, all reachable pixel indices are obtained and a two-dimensional dataset is formed, which can obtain the set of direction-consistent connected pixels.
[0035] S313: Based on all pixel indices in the directionally consistent connected pixel set, extract the position sequence in the image, perform path structure aggregation processing according to the geometric order on the connected chain, encode all path segments in sequence according to the connection order, establish a complete boundary trajectory data structure, and generate a set of boundary coherent structure pixel paths. Based on the pixel index information recorded in the directionally consistent connected pixel set, each pixel sequence is grouped and aggregated according to the connectivity of the path structure. First, the pixel coordinates in each path set are sorted according to the access order during the search. Then, they are connected into a directed sequence using a chain structure encoding method. The path number and the start and end index constitute the path segment description information, and parameters such as path length, start and end point coordinates, and average direction angle are recorded to form a structured path data frame. Furthermore, all path frame sets are merged to construct the overall image edge connectivity structure path graph. The list form is an index mapping set consisting of the path number and the list of pixel coordinates it contains. Each path corresponds to a set of connected pixels and their physical positions in the image, thereby establishing a boundary coherent structure pixel path set.
[0036] Please see Figure 5 Step S4 is as follows: S411: Based on each set of pixel paths recorded in the boundary coherent structure pixel path set, read the grayscale depth value of each pixel in the original image depth map, construct continuous pixel groups according to the path order, calculate the absolute depth difference between two adjacent pixels in sequence, count the gradient change rate values of all pixels in the horizontal and vertical directions in the entire original depth map, calculate the average value of the change rate, take twice as the jump reference difference, mark the position of the points in each set of paths whose depth difference is greater than the jump reference difference, and generate a jump pixel index mark matrix. Based on the continuous pixel sequence within each path of the boundary-coherent pixel path set, the grayscale depth values of each pixel in the path are retrieved from the original depth map, and a depth value sequence is constructed for each path. The absolute difference between the depth values of any two adjacent pixels in the sequence is used as a local jump metric. Simultaneously, the horizontal and vertical gradient magnitudes are calculated for all valid pixels in the entire depth map and summed. The average of the horizontal and vertical gradients is then calculated to obtain the average gradient value for the entire image, defined as Δ. The jump baseline difference is set to 2×Δ. For example, if the image resolution is 640×480, the average horizontal gradient in the depth map is 4.3. If the vertical distance is 5.1, then Δ = (4.3 + 5.1) / 2 = 4.7, and the jump reference difference is 9.4. The depth difference of each pair of adjacent pixels in the path is compared with the jump reference value. If the difference is greater than 9.4, the pixel position is marked as 1 in the jump marker matrix. For example, in path 1, the depth values of pixel pair (52, 118) - (53, 118) are 22.5 and 33.6 respectively, and the difference is 11.1 > 9.4, so a jump is marked. All pixels that meet the jump condition are counted, and their path index positions, coordinates, and differences are recorded to generate a jump pixel index marker matrix. Key example data is shown in Table 5. Table 5. Pixel depth difference and marking information for jump transitions; As shown in Table 5, some paths have depth jumps exceeding the baseline difference, which constitute a marker record and provide jump location information for subsequent construction of directional consistency references.
[0037] S412: Based on the path index position of all jump points in the jump pixel index marker matrix, retrieve the adjacent pixels in each path and determine whether the direction between the adjacent point and the jump point is consistent with the main direction of the path. If the direction difference is within the set tolerance threshold, aggregate the corresponding jump point and its adjacent points into a direction-continuous point group, and record all point groups that meet the direction continuity condition to generate a direction-continuous jump reference point set. Based on the position index of the marked jump point in each path in the jump pixel index marking matrix, the index positions of the two adjacent pixels before and after the corresponding point in the path are retrieved, and the angle between the main direction angle value and the jump point direction angle is calculated. If the angle is less than the tolerance angle threshold of 20 degrees, it is determined that the jump point and the adjacent point have the same direction. The jump point and the reference points before and after it are formed into a direction continuous point group and effectively recorded. For example, if the jump point direction is 135°, and the directions of the points before and after it are 128° and 142° respectively, the angle between them is 7° and 13°, both less than the threshold, then a reference point group is formed. The structure of this group includes the jump point index, path number, and direction angle difference information. At the same time, all jump point groups that meet the direction consistency are summarized to form a spatial structure set for correction. Finally, the direction continuous jump reference point set is obtained. The key examples are shown in Table 6 below: Table 6. Record of reference points with consistent jump point directions; Referring to Table 6, all transition points have valid reference point pairs with consistent orientations within the tolerance range, and depth value replacement operations can be performed based on this point group subsequently.
[0038] S413: Based on the set of reference points with continuous directional jumps, extract the depth values of the adjacent reference points corresponding to each jump point, and use the adjacent points as a benchmark to judge the trend of depth change before and after. If the trend is continuous and the depth value falls within the reference point change range, then replace the depth value of the jump point, and update it by using the linear weighted average of the depth values of the adjacent reference points to establish the boundary connectivity correction region image matrix. The coordinates of the preceding and following reference points for each group of reference points in the set of reference points with continuous directional jumps are read. The original depth map values of the reference points are extracted and archived according to group number. For each group of jump points, the depth trend is determined. If the increasing or decreasing trend of the reference point depth values is consistent, it is considered to be a continuous trend and a replacement operation is allowed. The replacement adopts a linear weighted average method, that is, the new value of the jump point is the arithmetic mean of the reference point depth values. For example, if the jump point is (52, 118), its preceding and following reference values are 22.5 and 33.6, then the replacement value is (22.5+33.6) / 2=28.05. This value is updated to the position of the corresponding jump point in the depth map matrix. This process is repeated for all jump points that meet the conditions. Finally, the image matrix structure with corrected depth values is constructed. The original values of the non-jumping areas are retained, and the corresponding values of the jump points are replaced. The boundary connectivity correction region image matrix is established. Relevant replacement examples are shown in Table 7. Table 7. Calculation of Jump Point Correction Depth Values; As shown in Table 7, the corrected depth values of all replaced transition points are calculated using reference points with consistent orientations, and the replacement is written into the image matrix, thereby constructing an image matrix of the boundary connectivity correction region that can be used for subsequent processing.
[0039] Please see Figure 6 The S5 steps are as follows: S511: Based on the location index information in the image matrix of the boundary connectivity correction region and the distribution aggregation matching point list, extract the depth value of each matching point in the image matrix, and perform point-by-point aggregation processing in combination with the disparity value of the corresponding point to construct a disparity-depth bidirectional mapping structure. According to the image coordinate system rules, rearrange each mapping pair in a two-dimensional matrix format and fill it into the disparity mapping frame in the complete spatial range to generate a spatial structure disparity reconstruction matrix. Based on the pixel coordinate index information recorded in the image matrix of the boundary connectivity correction region and the list of distribution aggregation matching points, the depth value of each pair of matching points in the image matrix is extracted as a disparity field to construct a reference dataset. Then, a two-dimensional coordinate index of the matching points on the image plane is established and bound to the corresponding depth value, generating a structured data unit containing position coordinates, disparity value, and corresponding depth. During the integration process, it is necessary to ensure that the matching points are within the image boundary range and that the depth value is within the valid grayscale range, i.e., between 0 and 255. Outliers are masked to remove invalid points. Finally, the data is processed using the top left corner of the image... Starting from the first point, all valid matching points are rearranged according to row and column order. Each pair of disparity-depth structures is mapped to the corresponding matrix unit according to coordinates. The missing areas are filled by neighborhood interpolation. The interpolation method can be edge extension, using the depth value in the neighboring direction as an approximate filling result. For example, if there is no match at a certain coordinate (128, 64), the average value of the two points (127, 64) and (129, 64) can be used for filling. Finally, the mapping and rearrangement of all matching points and the structure filling are completed, and a matrix structure with complete spatial distribution and depth coverage characteristics is constructed, generating a spatial structure disparity reconstruction matrix.
[0040] S512: Obtain the spatial structure disparity reconstruction matrix, pack all pixels in column-major order, perform three-channel pseudo-color encoding according to the image output specification, map the disparity information into an RGB distribution image, compress and encapsulate the encoded image to obtain a standard encapsulated disparity image data frame. The process involves reading the disparity information of all pixels in the spatial structure disparity reconstruction matrix. First, the two-dimensional matrix data is expanded into a one-dimensional vector sequence row by row. Based on this sequence, image pixel channel conversion processing is performed to convert the original grayscale disparity values into a pseudo-color encoding format. The pseudo-color conversion rule is based on a fixed color spectrum mapping method, such as mapping the disparity value range of 0~255 to the HSV color gamut and then converting it to the RGB color space to ensure that the image has significant visual distribution differences on the display end. Next, the original pixel structure is reconstructed in the form of a three-channel color matrix to generate a color frame with image display capabilities. Then, according to the image encapsulation standard, such as 8-bit deep compression, image header description information and pixel data segment encoding structure are allocated and integrated into a complete image data block. Finally, the compressed and encapsulated frame is written into the image data resource pool, and the encapsulated image path and corresponding tag information are attached to the resource description index to form a standard image resource unit that can be read and stored by the system. The encapsulation operation is completed, and a standard encapsulated disparity image data frame is obtained.
[0041] S513: Based on the spatial location index corresponding to the standard encapsulated parallax image data frame, construct an image index dictionary and merge it with the source image number mapping structure to form an archiveable multi-source image index tuple, establish a structured data archive record at the system level, and obtain the binocular vision depth map optimization results; Based on the path structure and numbering information of the encapsulated images in the standard encapsulated disparity image data frame, the corresponding image index number, image generation timestamp, image size metadata, etc., are extracted to construct an image data dictionary structure, which serves as the basic content of the image archiving index. This index structure is then bound and combined with the source image data structure to form a one-to-one mapping entry between images and disparity data. Next, a number mapping table is established in the system image data archiving area, using the unique image number as the retrieval key. The source image number, depth matrix position index, and encapsulation path are aggregated into an image resource archiving unit. This structure records the spatial index position and original path mapping relationship corresponding to each encapsulated image. The archiving structure is stored in a serialized storage structure, and data mapping key-value binding is implemented in the file system or database structure to form a consistent archiving system between the binocular image frame and its optimized disparity map. Finally, a diversified image storage structure containing image spatial structure and depth fusion content is established to obtain the optimized binocular visual depth map results.
[0042] A depth map optimization system based on binocular vision, comprising: The orientation consistency screening module is used to perform S1: acquire left and right view image frames in the binocular image, extract the gradient direction main frequency value in the image block where each pixel to be matched is located, filter all pixel pairs whose gradient direction angle is lower than the orientation tolerance angle, and generate a list of pixel pairs with consistent structural orientation. The confidence density detection module is used to perform S2: statistically analyze the spatial distribution of all pixel pairs in the list of pixel pairs with consistent structural orientation, construct a sliding window with each pair of matching points as the center, statistically analyze the matching density confidence of pixels with the same disparity value as the matching points within the sliding window, determine the concentrated matching points and retain them as reliable references, and generate a list of distributed aggregated matching points. The boundary structure tracking module is used to perform S3: based on the pixel positions marked in the distribution aggregation matching point list in the left figure, extract the main gradient direction starting from the edge protruding pixel, perform pixel-level direction consistent connectivity expansion, retain pixels with included angles within the tolerance range to construct a boundary structure sequence, and generate a set of boundary coherent structure pixel paths; The jump region correction module is used to perform S4: based on the initial depth map values of all pixel regions in the set of pixel paths with boundary continuity structure, calculate the absolute depth difference between adjacent points in a continuous pixel group, filter and mark jump points, count the points that satisfy the direction consistency continuity of adjacent pixels on the path of the jump points to construct a reference set of region connectivity graphs, perform value replacement processing on all jump points, and generate a boundary connectivity correction region image matrix. The depth map integration output module is used to perform S5: extract the depth information of the location points in the image matrix of the boundary connectivity correction region and the distribution aggregation matching point list, perform disparity field image reconstruction, and obtain the binocular vision depth map optimization results.
[0043] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for optimizing a depth map based on binocular vision, characterized in that, The method comprises the following steps: S1: obtaining left and right view image frames in binocular images, extracting gradient direction dominant frequency values of each pair of to-be-matched pixels in the image blocks in which the pixels are located, screening all pixel pairs with an included angle of gradient direction lower than a direction tolerance angle, and generating a list of structure direction consistent pixel pairs; S2: statistically analyzing spatial distribution of all pixel pairs in the list of structure direction consistent pixel pairs, constructing a sliding window with each pair of matching points as the center, statistically analyzing pixel matching density confidence in the sliding window, determining concentrated matching points and reserving the concentrated matching points as reliable references, and generating a list of distribution aggregated matching points; S3: based on pixel positions marked in the list of distribution aggregated matching points in the left image, extracting gradient dominant directions from edge prominent pixels, performing pixel-level direction consistency connected extension, reserving pixels with an included angle within a tolerance range to construct a boundary structure sequence, and generating a boundary coherent structure pixel path set; S4: based on initial depth map values of all pixel regions in the boundary coherent structure pixel path set, calculating absolute depth difference values of adjacent points in a continuous pixel group, screening marked jump points, statistically analyzing point groups of the jump points that satisfy direction consistency continuity of adjacent pixels on the path to construct a region connected graph reference set, performing value replacement processing on all jump points, and generating a boundary connected modified region image matrix; S5: extracting depth information of position points in the boundary connected modified region image matrix and the list of distribution aggregated matching points, performing disparity field image reconstruction, and obtaining binocular vision depth map optimization results.
2. The binocular vision-based depth map optimization method of claim 1, wherein: The concentrated matching points refer to matching points with a pixel matching density confidence greater than a confidence density reference frequency standard. The jump points refer to position points with a depth absolute difference value greater than a set jump reference difference value.
3. The binocular vision based depth map optimization method of claim 1, wherein: The list of structure direction consistent pixel pairs comprises dominant direction consistent pixel pairs and structure matching pairs screened through an included angle tolerance, the list of distribution aggregated matching points comprises confidence matching points and spatial consistency matching region points, the boundary coherent structure pixel path set comprises gradient dominant direction coherent paths and direction consistent pixel chains, the boundary connected modified region image matrix comprises jump point depth values adjusted through a trend, depth continuity optimized regions, and boundary consistency enhanced image blocks, and the binocular vision depth map optimization results comprise a reconstructed disparity field image, an encapsulated optimized image format, and an archived resource dataset.
4. The binocular vision based depth map optimization method of claim 1, wherein, The acquisition step of the list of structure direction consistent pixel pairs is specifically as follows: S111: obtaining left view image frames and right view image frames in binocular images, extracting gradient direction values of all pixels in image blocks in which each pair of to-be-matched pixels is located, constructing an image block gradient direction set, statistically analyzing frequency of each gradient direction value in the image block, and determining a dominant frequency value based on a statistical result, and generating a gradient direction dominant frequency value list. S112: Based on the gradient direction main frequency value list, read the main frequency direction value of the image block corresponding to the to-be-matched pixels at the same position in the left view and right view image frames, calculate the direction included angle and difference compare with the structure direction tolerance angle, when the direction included angle is less than the structure direction tolerance angle, record the pixel index position corresponding to the to-be-matched pixel pair, and obtain the direction consistent index pair list; S113: According to the direction consistent index pair list, extract the pixel position parameters of each index pair in the left view and right view image frames, and combine and aggregate each pair of pixel pairs that meet the direction consistency condition to generate a structure direction consistent pixel pair list.
5. The binocular vision-based depth map optimization method of claim 1, wherein, The obtaining step of the distribution aggregation matching point list is specifically: S211: Based on the structure direction consistent pixel pair list, a sliding window is constructed in the left view and right view of the binocular image respectively with the left view pixel in each pixel pair as the center, the pixel index of each position in the window is extracted one by one, and whether there is a corresponding pixel pair is retrieved, if there is, the disparity value of the corresponding matching pixel pair in the list is obtained and recorded, and a sliding window disparity information matrix is generated; S212: According to all the recorded disparity value data in each window centered on the matching point in the sliding window disparity information matrix, the corresponding disparity value of the center pixel is compared with the corresponding disparity value of other pixels in the window, the number of matching points with consistent disparity is counted, and the pixel matching density confidence value is calculated according to the coverage frequency of the matching point in the window area, if it is greater than the confidence density reference frequency standard threshold, the current matching point is determined as a reliable matching point, and a reliable matching point index list is obtained; S213: According to the reliable matching point index list, the coordinate values of all pixel pairs that meet the density confidence screening condition in the left view and right view are extracted, and are arranged and output in a one-to-one alignment manner, and a distribution aggregation matching point list is established.
6. The binocular vision-based depth map optimization method of claim 5, wherein, The calculation formula of the pixel matching density confidence value is: ; wherein, denotes the density confidence of the th matching point, denotes the disparity value of the matching point itself, denotes the disparity value of the th pixel point in the sliding window, denotes the number of all matching points in the sliding window.
7. The binocular vision-based depth map optimization method of claim 1, wherein, The obtaining step of the boundary continuous structure pixel path set is specifically: S311: Based on the left image pixel position marked in the distribution aggregation matching point list, detect the edge position where each pixel exists in the image gray scale distribution gradient mutation, extract the gray scale gradient direction angle of each pixel in the corresponding image block, and count the direction angle frequency distribution, determine the dominant direction angle of the current pixel according to the maximum item of the direction angle interval frequency, and generate an edge pixel main direction angle list; S312: According to the edge pixel main direction angle list, for each main direction angle, stepwise pixel extension search is performed with the corresponding direction as the axis, the local gradient direction angle value of each pixel on the search path is detected, and whether the included angle with the starting direction falls within the direction consistency tolerance range is judged item by item, when the included angle meets the tolerance condition, the corresponding pixel position is recorded and continuous expansion is continued, and a direction consistent pixel set is obtained; S313: According to the direction consistency of all pixel indexes in the connected pixel set, a sequence of positions in the image is extracted, and a path structure aggregation process is performed according to the geometric order on the connected chain. All path segments are sequentially encoded in the connection order, a complete boundary track data structure is established, and a boundary coherent structure pixel path set is generated.
8. The binocular vision based depth map optimization method of claim 1, wherein, The acquisition step of the boundary connected modified region image matrix is specifically: S411: Based on each group of pixel paths recorded in the boundary coherent structure pixel path set, the gray depth value of each pixel point in the original image depth map is read, and a continuous pixel group is constructed in the path order. The absolute difference value between adjacent two pixel points is calculated in sequence, the gradient change rate value of all pixel points in the original depth map in the horizontal and vertical directions is calculated, and the average value of the change rate is calculated. Take twice as the jump reference difference value. The points in each path whose depth difference value is greater than the jump reference difference value are marked in position, and a jump pixel index marking matrix is generated. S412: According to the path index position of all jump points in the jump pixel index marking matrix, the front and rear adjacent pixel points in each path are searched, and it is judged whether the direction between the adjacent points and the jump points is consistent with the main direction of the path. If the direction difference is within the set tolerance threshold, the corresponding jump point and its adjacent point are aggregated as a direction continuous point group, and all point groups satisfying the direction continuity condition are recorded to generate a direction continuous jump reference point set. S413: According to the direction continuous jump reference point set, the depth value of each adjacent reference point corresponding to the jump point is extracted, and the trend of the front and rear depth is judged based on the adjacent point as a reference. If the trend is continuous and the depth value falls within the reference point change interval, the jump point depth value is replaced, the linear weighted average method of the adjacent reference point depth value is used for updating, and a boundary connected modified region image matrix is established.
9. The binocular vision based depth map optimization method of claim 1, wherein, The acquisition step of the binocular vision depth map optimization result is specifically: S511: Based on the boundary connected modified region image matrix and the position index information in the distribution aggregation matching point list, the depth value of each matching point in the image matrix is extracted, and point-by-point aggregation processing is performed combined with the corresponding point disparity value. A disparity-depth bidirectional mapping structure is constructed, each mapping pair is rearranged according to the image coordinate system rule in the form of a two-dimensional matrix, and is filled into the disparity mapping frame in the complete space range to generate a spatial structure disparity reconstruction matrix. S512: Obtain the spatial structure disparity reconstruction matrix, pack all pixel points in column order, perform three-channel pseudo-color coding processing according to the image output specification, map the disparity information into an RGB distribution image, compress and package the coded image to obtain a standard packaged disparity image data frame. S513: According to the spatial position index corresponding to the standard packaged disparity image data frame, an image index dictionary is constructed and merged with the source image number mapping structure to form an archivable multi-source image index tuple, a structured data archiving record is established at the system level, and a binocular vision depth map optimization result is obtained.
10. A binocular vision based depth map optimization system, characterized in that, The system is used to realize the binocular vision based depth map optimization method of any one of claims 1-9, comprising: The direction consistent screening module is used to perform S1: obtaining left and right view image frames in binocular images, extracting gradient direction dominant value in image blocks where each pair of to-be-matched pixels is located, screening all pixel pairs with gradient direction included angle lower than the direction tolerance angle, and generating a list of structure direction consistent pixel pairs; The confidence density detection module is used to perform S2: counting the spatial distribution of all pixel pairs in the list of structure direction consistent pixel pairs, constructing a sliding window centered on each pair of matching points, counting the pixel matching density confidence in the sliding window, determining centralized matching points and keeping them as trusted references, and generating a list of distribution aggregated matching points; The boundary structure tracking module is used to perform S3: based on the labeled pixel positions in the left image in the list of distribution aggregated matching points, extracting gradient dominant direction from edge prominent pixels, performing pixel-level direction consistent connectivity expansion, keeping pixels with included angle within the tolerance range to construct a boundary structure sequence, and generating a set of boundary coherent structure pixel paths; The jump region correction module is used to perform S4: based on the initial depth map values of all pixel regions in the set of boundary coherent structure pixel paths, calculating the depth absolute difference values of adjacent points in continuous pixel groups, screening labeled jump points, counting point groups on the path where adjacent pixels of the jump points satisfy the direction consistency continuity to construct a region connectivity graph reference set, performing value replacement processing on all jump points, and generating a boundary connectivity correction region image matrix; The depth map integration output module is used to perform S5: extracting the depth information of position points in the boundary connectivity correction region image matrix and the list of distribution aggregated matching points, performing disparity field image reconstruction, and obtaining binocular vision depth map optimization results.