Image binocular stereo matching method and device, electronic equipment and storage medium
By performing grayscale image processing and local matching on the binocular stereo matching algorithm, the matching accuracy and speed issues in dark environments and weak texture conditions are solved, and real-time depth information acquisition is achieved in the fields of computer vision and robot vision.
Patent Information
- Application Number
- CN202410255855.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-09
AI Technical Summary
The existing binocular stereo matching algorithms have insufficient matching accuracy and speed in dark environments and weak texture conditions, making it difficult to meet real-time processing requirements.
The image format is converted into a grayscale image, the operator size and local matching range are set, and the disparity result array is obtained through local matching and disparity function processing. The disparity map is calculated based on the minimum value of the disparity result array.
It improves the matching accuracy and speed in dark environments and weak texture conditions, can output disparity maps in real time, and supports the acquisition of depth information in fields such as computer vision and robotic vision.
Smart Images

Figure CN120612356A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a binocular stereo matching method and device for images, an electronic device, and a storage medium. Background Art
[0002] Binocular stereo vision, a research topic in computer vision, primarily simulates the human eye to acquire three-dimensional information. Compared to other methods for acquiring 3D information, binocular vision offers advantages such as simple equipment and superior performance. It has been widely used in a variety of fields, including scene reconstruction, autonomous navigation, industrial measurement, and security monitoring. Binocular vision utilizes biomimetic principles, using two cameras positioned at different locations on the same horizontal line to capture two images. A disparity map is generated through a stereo matching algorithm, which then determines the image's depth information. Most stereo matching systems primarily consist of three components: image acquisition, image preprocessing, and stereo matching. Image acquisition primarily involves capturing valid information about the scene using two calibrated cameras, preparing for image preprocessing. Image preprocessing involves filtering and transforming the initial image, extracting features, and performing pre-matching processing. Finally, improving the matching accuracy of stereo matching algorithms has long been a key research focus and challenge. These algorithms are time-consuming and computationally intensive, and both matching accuracy and speed require significant improvement. Summary of the Invention
[0003] In view of the above problems, this application is proposed to provide a binocular stereo matching method and device, electronic device and storage medium for images that overcome the above problems or at least partially solve the above problems. The technical solution is as follows:
[0004] In a first aspect, a binocular stereo matching method for images is provided, comprising:
[0005] Acquire a first initial image and a second initial image of a target scene captured by two cameras of a binocular device at the same time;
[0006] Performing format conversion on the collected first initial image and the second initial image respectively, converting them into a first grayscale image and a second grayscale image of set bit numbers;
[0007] According to the basic image data of the first grayscale image and the second grayscale image, an operator size, a local matching range, and a segmentation threshold are set;
[0008] The first grayscale image and the second grayscale image are divided into blocks according to the operator size, and the grayscale value of each pixel in the block is remapped and normalized based on the segmentation threshold to obtain a first processed block image and a second processed block image;
[0009] Performing local matching on the first processed block image and the second processed block image according to a pre-constructed disparity function including a local matching range to obtain a disparity result array, wherein elements in the disparity result array represent correlations between the block images;
[0010] The minimum value of the elements in the disparity result array is found, and the disparity map is obtained based on the offset corresponding to the minimum value of the elements in the disparity result array.
[0011] In a second aspect, a binocular stereo matching device for images is provided, comprising:
[0012] An acquisition unit, configured to acquire a first initial image and a second initial image respectively captured by two cameras of a binocular device at the same time of a target scene;
[0013] A format conversion unit, configured to convert the acquired first initial image and the second initial image into a first grayscale image and a second grayscale image of set bit numbers;
[0014] A setting unit, configured to set an operator size, a local matching range, and a segmentation threshold according to basic image data of the first grayscale image and the second grayscale image;
[0015] a remapping and normalization processing unit, configured to divide the first grayscale image and the second grayscale image into blocks according to the operator size, and remap and normalize the grayscale value of each pixel in the block based on the segmentation threshold to obtain a first processed block image and a second processed block image;
[0016] a local matching unit, configured to perform local matching on the first processed block image and the second processed block image according to a pre-constructed disparity function including a local matching range, to obtain a disparity result array, wherein elements in the disparity result array represent correlations between the block images;
[0017] The disparity unit is used to find the minimum value of the elements in the disparity result array and obtain the disparity map based on the offset corresponding to the minimum value of the elements in the disparity result array.
[0018] According to a third aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the binocular stereo matching method of images described in any one of the above items.
[0019] In a fourth aspect, a storage medium is provided, wherein the storage medium stores a computer program, wherein the computer program is configured to execute any one of the above-mentioned binocular stereo matching methods for images when running.
[0020] By means of the above technical solution, the embodiments of the present application provide a binocular stereo matching method and apparatus, an electronic device, and a storage medium for images. The method can obtain a first initial image and a second initial image captured simultaneously by two cameras of a binocular device of a target scene; convert the captured first and second initial images into a first grayscale image and a second grayscale image of a set number of bits; set an operator size, a local matching range, and a segmentation threshold based on basic image data of the first and second grayscale images; divide the first and second grayscale images into blocks according to the operator size, and remap and normalize the grayscale values of each pixel in the block based on the segmentation threshold to obtain a first processed block image and a second processed block image; perform local matching on the first and second processed block images according to a pre-constructed disparity function including a local matching range to obtain a disparity result array, wherein elements in the disparity result array represent correlations between the block images; find the minimum value of the elements in the disparity result array, and obtain a disparity map based on the offset corresponding to the minimum value of the elements in the disparity result array. It can be seen that this embodiment converts the format of the collected first initial image and the second initial image respectively into a first grayscale image and a second grayscale image of a set number of bits, which can be used for binocular matching in weak texture, pure color environment or dark environment, and improves the matching accuracy and speed through local matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0022] Figure 1 A flow chart of a binocular stereo matching method for images provided in an embodiment of the present application is shown;
[0023] Figure 2 A flowchart of a binocular stereo matching method for images provided in another embodiment of the present application is shown;
[0024] Figure 3 A schematic diagram showing the disparity function values provided by an embodiment of the present application is shown;
[0025] Figure 4 A schematic diagram of a front-end model design provided in an embodiment of the present application is shown;
[0026] Figure 5 A schematic diagram of calibrating two cameras of a binocular device provided in an embodiment of the present application is shown;
[0027] Figure 6 A structural diagram of a binocular stereo matching device for images provided in an embodiment of the present application is shown;
[0028] Figure 7A structural diagram of a binocular stereo matching device for images provided in another embodiment of the present application is shown;
[0029] Figure 8 A structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0030] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0031] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such usage is interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "including" and its variations are to be interpreted as open-ended terms meaning "including but not limited to."
[0032] The common stereo matching algorithm process is as follows: 1. Matching cost calculation; 2. Cost aggregation; 3. Disparity calculation or optimization; 4. Disparity improvement. Stereo matching algorithms are categorized as local, global, and semi-global based on their principles, with local and semi-global algorithms being the most widely used. Currently, SGBM (Semi-Global Block Matching) is widely used in binocular matching due to its excellent efficiency and good matching results.
[0033] SGBM algorithm: It is an algorithm for stereo matching, mainly used to calculate the disparity between corresponding points in the image, thereby obtaining three-dimensional information of the scene. Its principle is based on the block matching method. SGBM uses the idea of block matching to divide the image into small blocks and then search for similar blocks in the left and right views. By comparing the similarity of pixels within the blocks, the algorithm determines the disparity of corresponding blocks in the two views. For each block, SGBM calculates the matching cost, that is, the difference between corresponding pixels in the two blocks. Common cost metrics include absolute difference and squared difference. The idea of global energy minimization is introduced. By aggregating the cost of each pixel and considering the relationship between adjacent pixels, a globally consistent disparity map is obtained. Dynamic programming is used to find the global minimum energy path to ensure a smooth and accurate disparity map.
[0034] The SAD (Sum of Absolute Differences) algorithm is an image matching algorithm whose basic principle is to calculate the sum of the absolute differences. This algorithm is commonly used for image block matching. It sums the absolute differences between the corresponding values of each pixel to evaluate the similarity between two image blocks. Cost aggregation often uses a fixed window to calculate the sum of all disparities within the window. The most intuitive way to calculate disparity is to use the WTA (Winner Takes All) method, directly selecting the disparity value that minimizes the aggregation cost.
[0035] Census (statistical) algorithm: is a classic algorithm for stereo matching, which is used to calculate the disparity between corresponding points in an image. The Census algorithm first selects a window, usually a small square area, surrounding each pixel. For each pixel in the selected window, the Census algorithm compares it with the surrounding pixels and encodes the comparison results. This usually involves comparing the grayscale values of the pixels to generate a binary code that represents the relative relationship between adjacent pixels. For each pair of matching windows, the Census algorithm calculates the Hamming distance between them, that is, the number of different bits in the binary code. This distance is considered the cost of the match, that is, the degree to which the two windows are different. By comparing all possible matching costs, the Census algorithm selects the match with the minimum cost, that is, the disparity corresponding to the minimum Hamming distance.
[0036] Existing binocular matching algorithms suffer from various shortcomings, making them unsuitable for binocular matching in dark environments. For example, the SAD algorithm only calculates the sum of absolute differences between pixels and does not consider image feature information. This results in a high false positive rate and is easily affected by factors such as image noise and texture. The Census algorithm, which requires matching of certain pixel blocks, cannot match pure colors or repetitive textures. Its matching performance is poor in areas with depth discontinuities or weak textures.
[0037] This embodiment designs a new algorithm that can perform binocular matching under dark conditions. The algorithm can perform feature matching on the collected pseudo-random dot patterns under dark conditions and has good parallelization effect. It is very suitable for fast calculation on hardware acceleration platforms such as DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), achieving real-time display effects, and provides a detailed hardware acceleration solution.
[0038] Figure 1 FIG. 4 shows a flow chart of a binocular stereo matching method for images provided in an embodiment of the present application. Figure 1 As shown, the binocular stereo matching method of the image may include the following steps S101 to S106:
[0039] Step S101 : obtaining a first initial image and a second initial image of a target scene respectively captured by two cameras of a binocular device at the same time.
[0040] In this step, the two cameras of the binocular device respectively capture a first initial image and a second initial image, where the first initial image may also be referred to as a left image, and the second initial image may also be referred to as a right image.
[0041] Step S102 : converting the formats of the collected first initial image and the second initial image into first grayscale images and second grayscale images of set bit numbers.
[0042] In this step, the number of bits can be set according to actual needs, such as 8 bits, which is not limited in this embodiment. Taking an 8-bit grayscale image as an example, an 8-bit grayscale image is an image in which each pixel is represented by 8 bits, that is, the grayscale value of each pixel ranges from 0 to 255. Images in this format do not contain color information, but instead represent the grayscale of the image through a single brightness value.
[0043] Step S103 : setting an operator size, a local matching range, and a segmentation threshold according to basic image data of the first grayscale image and the second grayscale image.
[0044] Step S104 : The first grayscale image and the second grayscale image are divided into blocks according to the operator size, and the grayscale value of each pixel in the block is remapped and normalized based on the segmentation threshold to obtain a first processed block image and a second processed block image.
[0045] Step S105 , performing local matching on the first processed block image and the second processed block image according to a pre-built disparity function including a local matching range, to obtain a disparity result array, wherein the elements in the disparity result array represent the correlation between the block images.
[0046] Step S106 , searching for the minimum value of the elements in the disparity result array, and obtaining a disparity map based on the offset corresponding to the minimum value of the elements in the disparity result array.
[0047] This embodiment converts the formats of the collected first initial image and the second initial image into a first grayscale image and a second grayscale image of a set number of bits, which can be used for binocular matching in weak texture, pure color environment or dark environment, and improves the matching accuracy and speed through local matching.
[0048] An embodiment of the present application provides a possible implementation method, which may further include the following step A after obtaining the disparity map in step S106:
[0049] Step A: Calculate the depth map through the disparity map based on the calibration parameters of the binocular device.
[0050] In this step, the input of binocular stereo matching is a pair of left and right images captured at the same time and corrected by the epipolar line, and the output is a disparity map composed of the disparity values corresponding to each pixel in the reference image (generally the left image is used as the reference image). Disparity is the pixel-level difference between the positions of corresponding points in the left and right images of a certain point in a three-dimensional scene. Then, based on the calibration parameters of the binocular device and the relationship between disparity and depth, the depth value of each pixel is calculated to obtain a depth map. The relationship here can be derived based on the geometric relationship between the camera of the binocular device and the scene. Common methods include triangulation, epipolar constraint method, etc.
[0051] This embodiment calculates a depth map through a disparity map. The depth map can provide rich depth information for fields such as computer vision, image processing, and robot vision, providing support for more application scenarios.
[0052] In an embodiment of the present application, a possible implementation method is provided. The basic image data of the first grayscale image and the second grayscale image mentioned in step S103 above may include the width and height of the first grayscale image, the width and height of the second grayscale image, and the environmental complexity of the target scene; the width of the first grayscale image is equal to the width of the second grayscale image, and the height of the first grayscale image is equal to the height of the second grayscale image. Then, step S103 sets the operator size, local matching range, and segmentation threshold according to the basic image data of the first grayscale image and the second grayscale image, and may specifically include the following steps B1 to B3:
[0053] Step B1: setting the operator size according to the heights of the first grayscale image and the second grayscale image.
[0054] In this step, the operator size can be set according to the size of the image, for example, the operator size is set to 1 / 50 of the height of the image, represented by n. It should be noted that the examples listed here are only illustrative and do not limit this embodiment.
[0055] Step B2: setting the height of the local matching range according to the operator size, and setting the width of the local matching range according to the widths of the first grayscale image and the second grayscale image.
[0056] In this step, the width of the local matching range can be set to 1 / 10 of the width of the image, and the height of the local matching range can be set to 1.2 to 1.5 times the operator. H and W represent the height and width of the image respectively, and H represents the height and width of the image respectively. T 、W T It should be noted that the enumeration here is only for illustration and does not limit the present embodiment.
[0057] Step B3: setting a segmentation threshold according to the environmental complexity of the target scene.
[0058] In this step, 4 gradients may be used for segmentation, and the segmentation threshold may be increased or decreased accordingly according to the complexity of the environment, which is not limited in this embodiment.
[0059] This embodiment sets the height and width of the local matching range according to the operator size, which can limit the search range during matching and improve matching efficiency; and sets the segmentation threshold according to the environmental complexity of the target scene, which can control the accuracy of the matching result.
[0060] In an embodiment of the present application, a possible implementation method is provided. In the above step S104, the grayscale value of each pixel in the block is remapped and normalized based on the segmentation threshold to obtain a first processed block image and a second processed block image. Specifically, the following steps C may be included:
[0061] Step C, remapping and normalizing the grayscale value of each pixel in the block based on the segmentation threshold according to the following formula to obtain a first processed block image and a second processed block image;
[0062]
[0063]
[0064] Where P(i,j) is the grayscale value of the pixel at position (i,j), (i,j)∈O, O is the block plane; thres is the normalized threshold interval obtained based on the segmentation threshold; Dev Min is the lower limit of the threshold interval after mapping; Dev Max is the upper limit of the threshold interval after mapping.
[0065] In this embodiment, the image can be blurred, and the grayscale value of each pixel in the block can be remapped and normalized based on the segmentation threshold to obtain a first processed block image and a second processed block image. This helps reduce interference, highlight features, and improve data consistency, thereby improving the quality and accuracy of image segmentation.
[0066] The embodiment of the present application provides a possible implementation method. The pre-built disparity function containing the local matching range mentioned in step S105 above is as follows:
[0067]
[0068] Among them, C(x,y,d) is the disparity function; S is the local matching range of the selected block; I R (x, y) is the gray value matrix of the pixel points in the block image after the first processing; I T (x, y) is the grayscale value matrix of the pixel points in the block image after the second processing; (x, y) is the position of the pixel point; d is the offset.
[0069] In this embodiment, local matching is performed on the first processed block image and the second processed block image according to a pre-established disparity function that includes a local matching range, thereby obtaining a disparity result array. This disparity result array includes the disparity value of each pixel, that is, the displacement or distance between the two blocks.
[0070] The present application provides a possible implementation method. The above step A calculates the depth map through the disparity map based on the calibration parameters of the binocular device, which may specifically include the following step D:
[0071] In step D, the depth map is calculated from the disparity map using the following formula based on the calibration parameters of the binocular device.
[0072]
[0073] Among them, DeepMap is the depth map; deltaPI is the disparity map; L is the baseline distance between the two cameras of the binocular device; PN is the number of horizontal pixels of the two cameras of the binocular device; FVA is the horizontal viewing angle of the two cameras of the binocular device.
[0074] This embodiment can accurately and efficiently calculate the depth map through the disparity map. The depth map can provide rich depth information for fields such as computer vision, image processing, and robot vision, and provide support for more application scenarios.
[0075] In an embodiment of the present application, a possible implementation method is provided. In the above step S101, the first initial image and the second initial image of the target scene respectively captured by two cameras of the binocular device at the same time are obtained. Specifically, the following steps E1 and E2 may be included:
[0076] Step E1, projecting pseudo-random dot matrix structured light on the target scene;
[0077] Step E2: obtaining a first initial image and a second initial image respectively captured by two cameras of the binocular device at the same time of the target scene, wherein the first initial image and the second initial image have pseudo-random dot matrices.
[0078] This embodiment uses structured light technology to project a pseudo-random dot matrix onto a target scene. The two cameras in the binocular device capture the resulting image, which then undergoes depth inference and calculation. By comparing the offset of the pseudo-random dot matrix in the two images, the distance and depth of the object can be determined, enabling depth perception of the target scene.
[0079] The above introduces Figure 1 There are multiple implementation methods for each link of the embodiment shown. The binocular stereo matching method of the image in the embodiment of the present application will be further explained below through specific examples.
[0080] In this specific embodiment, matching algorithm design and hardware acceleration solution design may be included, which will be described in detail below.
[0081] 1. Matching algorithm design.
[0082] Figure 2 A flowchart of a binocular stereo matching method for images provided by another embodiment of the present application is shown as follows:
[0083] Part to be processed: Image data acquisition, using the two cameras of the binocular device to collect two pictures (i.e., the first initial image and the second initial image) of the target scene at the same time. The two pictures contain random dot matrix and brightness information.
[0084] Known parameters: Basic information of the binocular device, such as horizontal viewing angle, vertical viewing angle, horizontal number of pixels, vertical number of pixels, baseline distance, etc.
[0085] Step 1: Convert the collected image into an 8-bit grayscale image to facilitate subsequent data processing and obtain basic image format information such as length and width.
[0086] At this point, obtain the number of image channels, image length and width, and basic device data.
[0087] Step 2: 1. Set the operator size. The operator size can be set according to the size of the image. Generally, it is set to 1 / 50 of the image height, which is represented by n. 2. Set the range of local matching. Generally, the width is set to 1 / 10 of the image width and the height is set to 1.2 to 1.5 times of the operator. H and W represent the height and width of the image respectively. T 、W T Represent the height and width of the local matching range respectively; 3. Set the segmentation threshold. Here, 4 gradients are used for segmentation. The segmentation threshold can be increased or decreased accordingly according to the complexity of the environment.
[0088] Step 3: Blur processing is performed. Since the image contains a large number of pseudo-random dots, this can be used as part of the depth information. The image is divided into blocks according to the operator size, and the grayscale value of each pixel in the block is remapped and normalized according to the following formula. This image is called a pseudo-brightness image.
[0089]
[0090]
[0091] Where P(i,j) is the grayscale value of the pixel at position (i,j), (i,j)∈O, O is the block plane; thres is the normalized threshold interval obtained based on the segmentation threshold; Dev Min is the lower limit of the threshold interval after mapping; Dev Max is the upper limit of the threshold interval after mapping.
[0092] Step 4: Perform local matching and perform shift calculation according to the following formula. Fill the edges of the image with 0.
[0093]
[0094] Among them, C(x,y,d) is the disparity function; S is the local matching range of the selected block; I R (x, y) is the gray value matrix of the pixel points in the block image after the first processing; I T(x, y) is the grayscale value matrix of the pixel points in the block image after the second processing; (x, y) is the position of the pixel point; d is the offset.
[0095] For each pair of blocks, the result is a (H T —n+1, W T —n+1) matrix, the matrix value is as follows Figure 3 As shown, the horizontal coordinate d and the vertical coordinate are matrix values. Find the minimum value among them, and the horizontal coordinate d of the minimum value is the parallax result. Figure 3 The Winner d* shown in FIG is the horizontal coordinate of the minimum value, which is the disparity result.
[0096] In this step, the minimum value can be compared with a preset threshold. If the minimum value is smaller than the preset threshold, the horizontal coordinate d of the minimum value is the disparity result; if the minimum value is greater than or equal to the preset threshold, the block is eliminated.
[0097] Step 5: Calculate the depth map from the disparity map. Use the following formula based on the calibration parameters of the binocular device to calculate the depth map from the disparity map.
[0098]
[0099] Among them, DeepMap is the depth map; deltaPI is the disparity map; L is the baseline distance between the two cameras of the binocular device; PN is the number of horizontal pixels of the two cameras of the binocular device; FVA is the horizontal viewing angle of the two cameras of the binocular device.
[0100] 2. Hardware acceleration solution design.
[0101] Currently, the most common method for acquiring depth information in three-dimensional scenes is binocular stereo vision. The disparity map generated by the stereo matching algorithm can roughly restore the three-dimensional information. However, this method is limited by the baseline length and the matching accuracy of pixels between the left and right images, and requires processing a large amount of data. With the advent of the high-definition era, the public's demand for visual enjoyment is increasing. 720P and 1080P HD video are becoming mainstream, and 4K resolution ultra-high-definition video is expected to be the future trend. This poses a greater challenge to processing speed. Traditional software platforms and serial processing methods are no longer able to meet the processing speed and matching accuracy requirements.
[0102] 720P typically refers to a resolution of 1280x720 pixels, with the "P" standing for progressive scan. 1080P refers to a resolution of 1920x1080 pixels, also with progressive scan. 4K refers to a resolution of 3840x2160 pixels, also expressed as 2160P. It has more pixels than 1080P, providing a clearer image and a higher visual experience. Meanwhile, the higher resolution 4K (4096x2160 pixels) is commonly used in filmmaking and digital cinema projection.
[0103] Hardware acceleration plays an important role in binocular vision, improving the speed and efficiency of algorithms. The following is a detailed explanation of some common hardware acceleration aspects.
[0104] Image sensor: The image sensor is the foundation of the binocular system, responsible for capturing images of the scene. For hardware acceleration, a high-resolution, low-noise, and high-sensitivity image sensor is crucial, as these characteristics directly impact the ability to obtain accurate depth information. For example, a CMOS (Complementary Metal Oxide Semiconductor) sensor or a dedicated binocular camera module can be used.
[0105] Hardware accelerators: Dedicated hardware accelerators are hardware components designed to perform specific tasks and can significantly increase the execution speed of algorithms. In binocular vision, some common hardware accelerators include:
[0106] Image Processing Unit (IPU): The IPU focuses on image processing tasks and can accelerate operations such as image preprocessing, feature extraction, and image matching.
[0107] Vision Processing Unit (VPU): The VPU is specifically designed to handle computer vision tasks, including binocular vision. It can accelerate the generation of depth maps and other depth perception-related calculations.
[0108] Hardware-accelerated FPGA: FPGA is a flexible hardware platform that can customize the hardware structure according to specific applications, thereby implementing highly optimized binocular vision algorithms.
[0109] Graphics Processing Unit (GPU): General-purpose computing graphics processing units have been widely used to accelerate computer vision tasks, including binocular vision. GPUs provide high performance on large-scale data through their parallel computing capabilities.
[0110] Custom chips: Some companies and research institutions may design custom chips specifically for binocular vision applications. These chips integrate specific hardware accelerators to provide higher performance and lower power consumption.
[0111] In practical applications, hardware acceleration is often combined with software algorithms to create a synergistic effect, providing faster and more efficient binocular vision solutions. Choosing the right hardware accelerator depends on application needs, budget, and performance requirements. As technology continues to advance, the application of hardware acceleration in binocular vision will continue to evolve, supporting a wider range of application scenarios.
[0112] (1) Model design: A binocular model is used to project pseudo-random dot matrix structured light to generate characteristic patterns. The front-end model design is as follows: Figure 4 As shown, in Figure 4 In the figure, Camera1 is the left camera, Camera2 is the right camera, and Projector is the projector. The pseudo-random dot matrix is generated according to a certain pattern, making it easier to decode and analyze the projected light spot. This pseudo-randomness reduces mutual interference between the light spots, improving system reliability. Pseudo-random dot matrix structured light is suitable for a variety of surface shapes, including complex curves and textures. This makes it widely applicable in applications such as industrial inspection, 3D reconstruction, and virtual reality.
[0113] (2) Binocular calibration: Binocular camera calibration is an important preprocessing step used to adjust the image geometry of the binocular camera to a specific standard form to simplify the matching and depth estimation process in stereo vision. Figure 5 As shown, xyz three-dimensional axis, O l and O r Here, we have the left and right cameras, and P is the target point. Binocular camera correction has the following main functions: I. Eliminate distortion; II. Improve depth estimation accuracy; III. Align corresponding points.
[0114] For the left and right cameras, there are the following formulas:
[0115]
[0116] It can be expressed as P by transformation w form:
[0117]
[0118] Subtracting two expressions:
[0119] (R l ) -1 ·(P l -T l )-(R r ) -1 ·(P r -T r )=0
[0120] Multiply by R r :
[0121] R r (R l ) -1 ·(P l -T l )-(P r -T r )=0
[0122] P r Putting this on one side of the equation gives:
[0123] P r =R r (R l ) -1 ·(P l -T l )+(T r )
[0124] Continuing to process, we get:
[0125] P r =R r (R l ) -1 P l +T r -R r (R l ) -1 T l
[0126] In addition, the right camera's principal point relative to the left camera's principal point obviously has:
[0127] P r =R·P l +T
[0128] So we have:
[0129]
[0130] Among them, P w is the coordinate of a point on the calibration plate in the world coordinate system; R and T are the rotation matrix and translation matrix of the right camera relative to the left camera; P l 、P r is the coordinate of the left and right cameras in the world coordinate system; R l 、T l is the rotation matrix and translation matrix of the point relative to the left camera principal point; R r 、T r is the rotation matrix and translation matrix of the point relative to the right camera principal point.
[0131] Substituting into the above formula, since multiple pictures are taken, the least squares method, or singular value decomposition, is used to minimize the error to obtain the best estimated matrix.
[0132] (3) Data transmission: The host computer transmits image data to the acceleration platform via UDP (User Data Protocol). The 8-bit data of the left and right images is summed into 16 bits and sent. The width of the image is used as a packet of data, and the number of packets sent is equal to the image height. When the acceleration platform receives the image frame header, it begins writing pixel data to DDR3 (Double Data Rate 3).
[0133] (4) Hardware acceleration of the algorithm: 8 bits of data are read from DDR3, each of which is divided into high and low bits. The image data is cached in FIFO (First In First Out) buffers, operators are extracted, and correlation calculations are performed. The threshold information is also calculated. The obtained data is used to find the minimum value and corresponding position through a least binary tree, which is used as the disparity.
[0134] (5) Output the disparity map via HDMI (High Definition Multimedia Interface).
[0135] This embodiment belongs to a local matching algorithm. The basic idea of the matching algorithm is to obtain corresponding correlation based on two-dimensional phase relationship matching, and use a two-dimensional random dot matrix as a structured light pattern, and a hardware acceleration device designed based on the algorithm.
[0136] This embodiment can perform binocular feature matching in weak texture and pure color environments, with better results than the SAD and Census algorithms. It can achieve 360P, 10FPS (frames per second) depth map output, and has been verified to have good parallel effects on the FPGA platform. It can output disparity maps in real time to achieve the purpose of depth extraction.
[0137] It should be noted that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In practical applications, all possible implementation methods described above can be combined in any manner to form possible embodiments of the present application, and will not be described in detail here.
[0138] Based on the binocular stereo matching methods of images provided in the above embodiments, and based on the same inventive concept, an embodiment of the present application further provides a binocular stereo matching device for images.
[0139] Figure 6This is a structural diagram of the binocular stereo matching device for images provided in the embodiment of the present application. Figure 6 As shown, the binocular stereo matching device for the image may specifically include an acquisition unit 610 , a format conversion unit 620 , a setting unit 630 , a remapping and normalization processing unit 640 , a local matching unit 650 and a disparity unit 660 .
[0140] An acquisition unit 610 is configured to acquire a first initial image and a second initial image of a target scene captured by two cameras of a binocular device at the same time;
[0141] A format conversion unit 620 is used to convert the acquired first initial image and the second initial image into a first grayscale image and a second grayscale image of a set number of bits respectively;
[0142] A setting unit 630 is used to set an operator size, a local matching range, and a segmentation threshold according to basic image data of the first grayscale image and the second grayscale image;
[0143] a remapping and normalization processing unit 640 for dividing the first grayscale image and the second grayscale image into blocks according to the operator size, and remapping and normalizing the grayscale value of each pixel in the block based on the segmentation threshold to obtain a first processed block image and a second processed block image;
[0144] A local matching unit 650 is configured to perform local matching on the first processed block image and the second processed block image according to a pre-constructed disparity function including a local matching range, to obtain a disparity result array, wherein elements in the disparity result array represent correlations between the block images;
[0145] The disparity unit 660 is configured to find a minimum value of elements in the disparity result array and obtain a disparity map based on an offset corresponding to the minimum value of the elements in the disparity result array.
[0146] A possible implementation method is provided in the embodiment of the present application, such as Figure 7 As shown above Figure 6 The device shown may further include a depth unit 710 for calculating a depth map from the disparity map based on calibration parameters of the binocular device.
[0147] In an embodiment of the present application, a possible implementation is provided, wherein the basic image data of the first grayscale image and the second grayscale image include the width and height of the first grayscale image, the width and height of the second grayscale image, and the environmental complexity of the target scene; the width of the first grayscale image is equal to the width of the second grayscale image, and the height of the first grayscale image is equal to the height of the second grayscale image; and the setting unit 630 is further configured to:
[0148] Setting the operator size according to the heights of the first grayscale image and the second grayscale image;
[0149] Setting a height of the local matching range according to the operator size, and setting a width of the local matching range according to the widths of the first grayscale image and the second grayscale image;
[0150] The segmentation threshold is set according to the environmental complexity of the target scene.
[0151] In an embodiment of the present application, a possible implementation is provided, wherein the remapping and normalization processing unit 640 is further configured to:
[0152] Remapping and normalizing the grayscale value of each pixel in the block based on the segmentation threshold according to the following formula to obtain a first processed block image and a second processed block image;
[0153]
[0154]
[0155] Where P(i,j) is the grayscale value of the pixel at position (i,j), (i,j)∈O, O is the block plane; thres is the normalized threshold interval obtained based on the segmentation threshold; Dev Min is the lower limit of the threshold interval after mapping; Dev Max is the upper limit of the threshold interval after mapping.
[0156] An embodiment of the present application provides a possible implementation method, wherein the pre-built disparity function including the local matching range is as follows:
[0157]
[0158] Among them, C(x,y,d) is the disparity function; S is the local matching range of the selected block; I R (x, y) is the gray value matrix of the pixel points in the block image after the first processing; I T (x, y) is the grayscale value matrix of the pixel points in the block image after the second processing; (x, y) is the position of the pixel point; d is the offset.
[0159] An embodiment of the present application provides a possible implementation method, wherein the depth unit 710 is further configured to:
[0160] The following formula is used to calculate the depth map from the disparity map based on the calibration parameters of the binocular device;
[0161]
[0162] Among them, DeepMap is the depth map; deltaPI is the disparity map; L is the baseline distance between the two cameras of the binocular device; PN is the number of horizontal pixels of the two cameras of the binocular device; FVA is the horizontal viewing angle of the two cameras of the binocular device.
[0163] A possible implementation is provided in an embodiment of the present application, where the acquiring unit 610 is further configured to:
[0164] Project pseudo-random dot matrix structured light on the target scene;
[0165] A first initial image and a second initial image of a target scene respectively captured by two cameras of a binocular device at the same time are obtained, wherein the first initial image and the second initial image have pseudo-random dot matrices.
[0166] Based on the same inventive concept, an embodiment of the present application also provides an electronic device, including a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the binocular stereo matching method of images of any one of the above embodiments.
[0167] In an exemplary embodiment, an electronic device is provided, such as Figure 8 As shown, Figure 8 The electronic device 800 shown includes a processor 801 and a memory 803. The processor 801 and the memory 803 are connected, for example, via a bus 802. Optionally, the electronic device 800 may further include a transceiver 804. It should be noted that in actual applications, the number of transceivers 804 is not limited to one, and the structure of the electronic device 800 does not constitute a limitation on the embodiments of the present application.
[0168] The processor 801 may be a CPU (Central Processing Unit), a GPU, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor 801 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0169] The bus 802 may include a path for transmitting information between the above components. The bus 802 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 802 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0170] The memory 803 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0171] The memory 803 is used to store computer program codes for executing the solution of the present application, and the execution is controlled by the processor 801. The processor 801 is used to execute the computer program codes stored in the memory 803 to implement the contents shown in the above method embodiment.
[0172] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0173] Based on the same inventive concept, an embodiment of the present application further provides a storage medium storing a computer program, wherein the computer program is configured to execute the binocular stereo matching method of images of any of the above embodiments when running.
[0174] Those skilled in the art will clearly understand that the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the aforementioned method embodiments, and for the sake of brevity, they will not be further described here.
[0175] Those skilled in the art will appreciate that the technical solution of the present application, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of program instructions for causing an electronic device (e.g., a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application when the program instructions are executed. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0176] Alternatively, all or part of the steps of implementing the aforementioned method embodiments may be accomplished by hardware related to program instructions (such as electronic devices such as personal computers, servers, or network devices), and the program instructions may be stored in a computer-readable storage medium. When the program instructions are executed by a processor of an electronic device, the electronic device executes all or part of the steps of the methods described in the various embodiments of the present application.
[0177] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that, within the spirit and principles of the present application, they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate from the protection scope of the present application.
Claims
1. A binocular stereo matching method for images, characterized in that: include: Acquire a first initial image and a second initial image of a target scene captured by two cameras of a binocular device at the same time; Performing format conversion on the collected first initial image and the second initial image respectively, converting them into a first grayscale image and a second grayscale image of set bit numbers; According to the basic image data of the first grayscale image and the second grayscale image, an operator size, a local matching range, and a segmentation threshold are set; The first grayscale image and the second grayscale image are divided into blocks according to the operator size, and the grayscale value of each pixel in the block is remapped and normalized based on the segmentation threshold to obtain a first processed block image and a second processed block image; Performing local matching on the first processed block image and the second processed block image according to a pre-constructed disparity function including a local matching range to obtain a disparity result array, wherein elements in the disparity result array represent correlations between the block images; The minimum value of the elements in the disparity result array is found, and the disparity map is obtained based on the offset corresponding to the minimum value of the elements in the disparity result array.
2. The method according to claim 1, characterized in that After obtaining the disparity map, the method further includes: Based on the calibration parameters of the binocular device, the depth map is calculated through the disparity map.
3. The method according to claim 1 or 2, characterized in that The basic image data of the first grayscale image and the second grayscale image include the width and height of the first grayscale image, the width and height of the second grayscale image, and the environmental complexity of the target scene; the width of the first grayscale image is equal to the width of the second grayscale image, and the height of the first grayscale image is equal to the height of the second grayscale image; According to the basic image data of the first grayscale image and the second grayscale image, the operator size, local matching range and segmentation threshold are set, including: Setting the operator size according to the heights of the first grayscale image and the second grayscale image; Setting a height of the local matching range according to the operator size, and setting a width of the local matching range according to the widths of the first grayscale image and the second grayscale image; The segmentation threshold is set according to the environmental complexity of the target scene.
4. The method according to claim 1 or 2, characterized in that Remapping and normalizing the grayscale value of each pixel in the block based on the segmentation threshold to obtain a first processed block image and a second processed block image, including: Remapping and normalizing the grayscale value of each pixel in the block based on the segmentation threshold according to the following formula to obtain a first processed block image and a second processed block image; Where P(i,j) is the grayscale value of the pixel at position (i,j), (i,j)∈O, O is the block plane; thres is the normalized threshold interval obtained based on the segmentation threshold; Dev Min is the lower limit of the threshold interval after mapping; Dev Max is the upper limit of the threshold interval after mapping.
5. The method according to claim 1 or 2, characterized in that The pre-built disparity function containing the local matching range is as follows: Among them, C(x,y,d) is the disparity function; S is the local matching range of the selected block; I R (x, y) is the gray value matrix of the pixel points in the block image after the first processing; I T (x, y) is the grayscale value matrix of the pixel points in the block image after the second processing; (x, y) is the position of the pixel point; d is the offset.
6. The method according to claim 2, characterized in that Based on the calibration parameters of the binocular device, the depth map is calculated through the disparity map, including: The following formula is used to calculate the depth map from the disparity map based on the calibration parameters of the binocular device; Among them, DeepMap is the depth map; deltaPI is the disparity map; L is the baseline distance between the two cameras of the binocular device; PN is the number of horizontal pixels of the two cameras of the binocular device; FVA is the horizontal viewing angle of the two cameras of the binocular device.
7. The method according to any one of claims 1 to 6, characterized in that Acquiring a first initial image and a second initial image of a target scene respectively captured by two cameras of a binocular device at the same time, including: Project pseudo-random dot matrix structured light on the target scene; A first initial image and a second initial image of a target scene respectively captured by two cameras of a binocular device at the same time are obtained, wherein the first initial image and the second initial image have pseudo-random dot matrices.
8. A binocular stereo matching device for images, characterized in that: include: An acquisition unit, configured to acquire a first initial image and a second initial image respectively captured by two cameras of a binocular device at the same time of a target scene; A format conversion unit, configured to convert the acquired first initial image and the second initial image into a first grayscale image and a second grayscale image of set bit numbers; A setting unit, configured to set an operator size, a local matching range, and a segmentation threshold according to basic image data of the first grayscale image and the second grayscale image; a remapping and normalization processing unit, configured to divide the first grayscale image and the second grayscale image into blocks according to the operator size, and remap and normalize the grayscale value of each pixel in the block based on the segmentation threshold to obtain a first processed block image and a second processed block image; a local matching unit, configured to perform local matching on the first processed block image and the second processed block image according to a pre-constructed disparity function including a local matching range, to obtain a disparity result array, wherein elements in the disparity result array represent correlations between the block images; The disparity unit is used to find the minimum value of the elements in the disparity result array and obtain the disparity map based on the offset corresponding to the minimum value of the elements in the disparity result array.
9. An electronic device, characterized in that: The system comprises a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the binocular stereo matching method of images according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the binocular stereo matching method for images according to any one of claims 1 to 7 when running.