Binocular stereo vision matching method and device based on pooling
By pooling the cost space in binocular stereoscopic visual matching, the problem of high storage resources requirements of the semi-global matching algorithm is solved, and the effect of reducing storage resource occupation is achieved.
Patent Information
- Application Number
- CN202510133447.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, semi-global matching algorithms require storage of cost space, resulting in high requirements for storage resources.
By pooling the cost space, it is reduced to a preset size to form an updated cost space, thereby reducing the consumption of storage resources.
It realizes the reduction of storage resource usage, reduces the requirements for storage resource, and maintains matching accuracy.
Smart Images

Figure CN119992142A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a binocular stereo vision matching method and device based on pooling. Background Art
[0002] Binocular stereo vision refers to two cameras capturing images of the same scene from different perspectives, and calculating the three-dimensional coordinates of each point in the scene by comparing and matching the features in the two images. This method is widely used in robot navigation, autonomous driving, virtual reality and other fields.
[0003] Binocular stereo vision matching methods mainly include pixel-based matching methods, feature-based matching methods, global optimization methods, and deep learning methods. The pixel-based method is simple to calculate, but is sensitive to texture loss and noise; the feature-based method can extract meaningful matching features and is suitable for complex scenes; the global optimization method can effectively solve problems such as discontinuity of parallax, but the amount of calculation is large. The deep learning method consumes a lot of on-chip resources and is not conducive to hardware implementation. The semi-global matching algorithm (SGM, Semi-Global Matching) combines the high precision of global matching and the high efficiency of local matching and is widely used.
[0004] However, although the semi-global matching algorithm has higher accuracy, it has higher requirements on storage resources because it needs to store the cost space.
[0005] Therefore, the prior art has defects and needs to be improved and developed. Summary of the invention
[0006] The technical problem to be solved by the present invention is that, in view of the above-mentioned defects of the prior art, a binocular stereo vision matching method and device based on pooling is provided, aiming to solve the problem that the semi-global matching algorithm in the prior art needs to store the cost space and has high requirements for storage resources.
[0007] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0008] A binocular stereo vision matching method based on pooling, wherein the method comprises:
[0009] Acquire a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image;
[0010] Determine a cost space of a current pixel on the first image on the second image;
[0011] Performing a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space;
[0012] A pooling value for each window in the updated cost space is calculated, and a pixel in the second image that matches a current pixel in the first image is determined based on each of the pooling values.
[0013] In one embodiment of the present application, before determining the cost space of the current pixel on the first image on the second image, the method further includes:
[0014] After performing epipolar correction on the first image and the second image, extracting a first image feature of the first image and extracting a second image feature of the second image;
[0015] A similarity between the first image feature and the second image feature is calculated.
[0016] In one embodiment of the present application, determining a cost space of a current pixel on a first image on the second image includes:
[0017] Get the preset disparity range and cost data width;
[0018] A cost space of a current pixel on the first image on the second image is determined according to the similarity, the image heights and image widths of the first image and the second image, the disparity range, and the cost data width.
[0019] In one embodiment of the present application, performing a pooling operation on the cost space to form an updated cost space includes:
[0020] Perform a 2×2 pooling operation on the cost space, reduce the image widths of the first image and the second image to half of the original, and reduce the disparity range to a quarter of the original, so as to reduce the cost space to a preset size, and divide the cost space into multiple 2×2 windows to form an updated cost space.
[0021] In one embodiment of the present application, calculating the pooling value of each window in the updated cost space includes:
[0022] The pooling value of each window in the updated cost space is calculated using a preset pooling strategy, where the preset pooling strategy is a minimum pooling strategy, a maximum pooling strategy, a median pooling strategy, or a mean pooling strategy.
[0023] In one embodiment of the present application, the pooled value is position information represented by a 2-bit sign bit.
[0024] In one embodiment of the present application, determining a pixel in the second image that matches a current pixel in the first image based on each of the pooling values includes:
[0025] When up-sampling the pooled value using a bilinear interpolation method, activating a pre-stored enable signal once every two clock cycles so that an output result changes once every two clock cycles;
[0026] A pixel in the second image that matches a current pixel in the first image is determined based on the output result.
[0027] The present application also provides a binocular stereo vision matching device based on pooling, wherein the device comprises:
[0028] An acquisition module, used for acquiring a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image;
[0029] A determination module, used to determine a cost space of a current pixel on the first image on the second image;
[0030] A pooling module, used for performing a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space;
[0031] An aggregation module is used to calculate a pooling value for each window in the updated cost space, and determine a pixel in the second image that matches a current pixel on the first image based on each of the pooling values.
[0032] The present application also provides a terminal, which includes: a memory, a processor, and a binocular stereo vision matching program based on pooling stored in the memory and executable on the processor, wherein the binocular stereo vision matching program based on pooling, when executed by the processor, implements the steps of the binocular stereo vision matching method based on pooling as described above.
[0033] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the binocular stereo vision matching method based on pooling as described above.
[0034] The present invention provides a binocular stereo vision matching method and device based on pooling, and the binocular stereo vision matching method based on pooling includes: obtaining a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image; determining the cost space of the current pixel on the first image on the second image; performing a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space; calculating the pooling value of each window in the updated cost space, and determining the pixel in the second image that matches the current pixel on the first image based on each of the pooling values. The present application reduces the cost space to a preset size by performing a pooling operation on the original cost space, thereby reducing the occupation of storage resources and thus reducing the requirements for storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flow chart of a preferred embodiment of the binocular stereo vision matching method based on pooling in the present invention;
[0036] Figure 2 It is the overall architecture of the binocular stereo vision matching processor with 2×2 pooling added in the present invention;
[0037] Figure 3 It is a schematic diagram of the working principle of the enable signal in the present invention;
[0038] Figure 4 It is a functional principle block diagram of a preferred embodiment of a binocular stereo vision matching device based on pooling in the present invention;
[0039] Figure 5 It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] Traditional semi-global matching algorithms require the complete cost space of the binocular image. Each pair of inputs requires cost data of H×W×D×C, where H and W are the image height and width, D is the disparity range, and C is the cost data width. Larger disparity ranges and data widths improve matching accuracy, but also significantly increase memory and energy consumption, especially in high-resolution and high-precision scenes. In the aggregation process, the huge amount of data requires too much memory and computing resources, so optimization of cost storage is necessary.
[0042] See also Figure 1 , Figure 1 Flowchart of binocular stereo vision matching method based on pooling in the present invention. Figure 1 As shown, the binocular stereo vision matching method based on pooling described in the embodiment of the present invention includes:
[0043] Step S100: Acquire a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image.
[0044] Specifically, the first image and the second image are a left-eye image and a right-eye image.
[0045] like Figure 1 As shown, the binocular stereo vision matching method based on pooling described in this embodiment also includes:
[0046] Step S200: Determine a cost space of a current pixel on the first image on the second image.
[0047] In the embodiment of the present application, before determining the cost space of the current pixel on the first image on the second image, the method further includes:
[0048] After performing epipolar correction on the first image and the second image, extracting a first image feature of the first image and extracting a second image feature of the second image;
[0049] A similarity between the first image feature and the second image feature is calculated.
[0050] Specifically, epipolar correction is performed on the first image and the second image to ensure that the corresponding points are on the same horizontal line (or close to the same horizontal line) in the two images, which greatly simplifies the subsequent matching process. When extracting the features of the first image, a feature extraction algorithm (such as SIFT, SURF, ORB, etc.) is used to process the first image to extract the key points (feature points) and their descriptors (feature vectors) in the image. These descriptors contain local information around the key points, that is, the first image features include the first key points and the corresponding first descriptors. The second image features include the second key points and the corresponding second descriptors.
[0051] When calculating the similarity, a certain distance metric between descriptors (such as Euclidean distance, Hamming distance, etc.) is used to compare the first descriptor of each first key point in the first image with the second descriptor of all second key points in the second image, and several comparison pairs are obtained. According to the distance metric result, a similarity score is assigned to each comparison pair. Generally, the smaller the distance, the more similar the two descriptors are, so the higher the similarity score. Setting a threshold and only retaining matching pairs with a similarity score higher than the threshold helps to remove false matches and improve the accuracy of matching.
[0052] By extracting image features and performing descriptor matching, the present application can more accurately find corresponding points between two images; compared with the method of directly comparing the pixels of the entire image, the feature-based matching method greatly reduces the amount of data that needs to be compared and improves the efficiency of the matching process, especially when processing high-resolution images.
[0053] In one embodiment of the present application, step S200 specifically includes:
[0054] Step S210, obtaining a preset disparity range and cost data width;
[0055] Step S220 , determining a cost space of a current pixel on the first image on the second image according to the similarity, the image heights and image widths of the first image and the second image, the disparity range, and the cost data width.
[0056] Specifically, for each current pixel on the first image, a cost space is constructed based on all corresponding points that may exist in the disparity range on the second image. The cost space is a multidimensional array (or matrix). At each position in the cost space, the similarity between the current pixel of the first image and its corresponding point on the second image is stored. These values can be similarity scores obtained by direct calculation.
[0057] By constructing a cost space, the embodiment of the present application can comprehensively consider the similarities between the current pixel and all possible corresponding points on the second image, thereby selecting the most matching pixel point, which helps to reduce the occurrence of mismatches and improve the accuracy of matching.
[0058] like Figure 1 As shown, the binocular stereo vision matching method based on pooling described in this embodiment also includes:
[0059] Step S300: performing a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space.
[0060] like Figure 2 As shown, the embodiment of the present application adds a pooling module for performing pooling operations, and the initial cost calculation before pooling can include multiple calculation methods such as gradient cost and Census cost. Specifically, the cost calculation includes AD, SAD and other methods, and the feature extraction includes Census, Gradient and other methods. The cost calculation and feature extraction can be combined arbitrarily.
[0061] The cost calculation method is used to evaluate the similarity between two image blocks or pixels. The AD algorithm, or Absolute Differences, is a matching cost method based on single pixel calculation. It continuously compares the grayscale values (or color values) of corresponding points in the left and right images (or image blocks), and calculates the absolute value of the difference between them as the matching cost. The advantages of the AD algorithm are simple and intuitive calculation, and good matching effect for texture-rich areas. The SAD algorithm, or Sum of Absolute Differences, is an extended form of the AD algorithm. It considers the matching cost of multiple pixels in an image block. It evaluates the similarity between two image blocks by calculating the sum of the absolute values of the grayscale values (or color values) of all pixels in a fixed window (or image block). This window can slide on the image to calculate the matching cost at different positions. The advantages of the SAD algorithm are relatively simple and fast calculation, and it is suitable for occasions with high real-time requirements.
[0062] The Census feature extraction method, also known as Census transform or Census encoding, is a non-parametric local feature description method. Census transform generates a binary string as the feature description of a pixel by comparing the grayscale values of a pixel with other pixels in its neighborhood. The Census feature is simple to calculate, easy to implement, and insensitive to image noise. The Gradient feature extraction method, also known as gradient feature extraction, is a feature description method based on image gradient information. Gradient feature extraction captures the edge and texture information of the image by calculating the gradient value (including gradient size and direction) of each pixel in the image. The gradient feature is sensitive to the edge and texture information of the image and can effectively describe the local structure in the image. The gradient feature has a certain robustness to illumination changes and image rotation.
[0063] In an embodiment of the present application, step S300 specifically comprises: performing a 2×2 pooling operation on the cost space, reducing the image widths of the first image and the second image to half of the original size, and reducing the disparity range to a quarter of the original size, so as to reduce the cost space to a preset size, and dividing the cost space into multiple 2×2 windows to form an updated cost space.
[0064] Specifically, the pooling operation of the present application reduces the cost space to H×W×D×C / 2 by adjusting the image width to W / 2, which greatly saves the size of on-chip storage, that is, saves the space of the memory located on the processor chip, especially in the case of high resolution. At the same time, the regional optimization method can be applied to the disparity range D. The regional optimization method can retain the pooled values of four distances, and the actual distance is restored after aggregation. The regional optimization method reduces the cost size of the cost space to H×W×D×C / 8 by reducing the disparity range to D / 4. Since the pooling method and the regional optimization method operate on different dimensions, they do not interfere with each other. The embodiment of the present application adds a part of parallel computing resources, but after calculation, it still has an advantage over the original method in terms of overall area. After pooling, the regional optimization of the present application further reduces the memory size, and then the winner-takes-all (WTA) method can be used to summarize and refine the optimization results.
[0065] The embodiment of the present application adds a pooling module for performing pooling operations, which reduces the demand for cost storage resources on the hardware to one-half with little impact on accuracy, while reducing the amount of calculation.
[0066] like Figure 1 As shown, the binocular stereo vision matching method based on pooling described in this embodiment also includes:
[0067] Step S400: Calculate the pooling value of each window in the updated cost space, and determine the pixel in the second image that matches the current pixel on the first image based on each of the pooling values.
[0068] In an embodiment of the present application, the pooling value of each window in the updated cost space is calculated, including: using a preset pooling strategy to calculate the pooling value of each window in the updated cost space, the preset pooling strategy is a minimum pooling strategy, a maximum pooling strategy, a median pooling strategy or a mean pooling strategy.
[0069] Specifically, the minimum pooling strategy refers to performing a pooling operation on the cost space of the current pixel on the first image on the second image, and after forming an updated cost space, in the updated cost space, using a preset window size, calculating the minimum value in each window as the pooling value of the window. The maximum pooling strategy refers to performing a pooling operation on the cost space of the current pixel on the first image on the second image, and after forming an updated cost space, calculating the maximum value in each window in the updated cost space as the pooling value of the window. The median pooling strategy refers to performing a pooling operation on the cost space of the current pixel on the first image on the second image, and after forming an updated cost space, calculating the median in each window in the updated cost space as the pooling value of the window. The mean pooling strategy refers to performing a pooling operation on the cost space of the current pixel on the first image on the second image, and after forming an updated cost space, calculating the mean in each window in the updated cost space as the pooling value of the window.
[0070] This application uses the minimum pooling strategy to find the corresponding point that is most similar to the current pixel, thereby improving the matching accuracy; in certain specific scenarios, the maximum pooling strategy can reflect a special matching relationship or feature; the median pooling strategy can reduce the impact of noise and outliers while maintaining good matching accuracy; the mean pooling strategy can smooth fluctuations in the cost space and reduce the impact of noise, and the mean calculation is relatively simple, which can reduce the computational complexity and improve matching efficiency.
[0071] In one embodiment of the present application, the pooling value is position information represented by a 2-bit sign bit. Specifically, the 2-bit sign bit is such as positions (0,0), (0,1), (1,0) and (1,1). Taking the minimum aggregation strategy as an example, the minimum value of position (0,0) and position (1,0) can be obtained by comparison, and the minimum value of (0,1) and (1,1) is determined by register transfer. Comparing these results can obtain the final pooling value of the 2×2 window.
[0072] The embodiment of the present application uses a 2-bit sign bit to represent position information, which can greatly simplify the calculation process of the pooling operation. Since the pooling value is represented by only a 2-bit sign bit, the memory requirement for storing the cost space can be significantly reduced. This is especially important when processing high-resolution images or large-scale image data sets, because the size of the cost space is usually proportional to the image resolution.
[0073] In the embodiment of the present application, the step S400 of “determining a pixel in the second image that matches a current pixel in the first image based on each of the pooling values” specifically includes:
[0074] When up-sampling the pooled value using a bilinear interpolation method, activating a pre-stored enable signal once every two clock cycles so that an output result changes once every two clock cycles;
[0075] A pixel in the second image that matches a current pixel in the first image is determined based on the output result.
[0076] Specifically, the embodiments of the present application can apply upsampling and other post-processing techniques to improve the results. The embodiments of the present application implement a pooling hardware architecture without timing conversion by storing an enable control signal and a bilinear interpolation method. Unlike software, the calculation of the hardware pipeline is continuous. Therefore, Figure 3 As shown, an embodiment of the present application uses an enable signal to discard unused pixels, which is activated once every two clock cycles to ensure that each module runs only at the right time, thereby saving dynamic energy consumption. An embodiment of the present application saves the necessary results to the memory and ensures that the read address increases once every two clocks by controlling the write enable signal of the memory, thereby generating a "pseudo low-frequency" output, that is, the output result changes once every two clock cycles. During the upsampling process, the initial value of every two clock cycles is output directly, and the subsequent values are delayed by the register for the next two cycles. The value generated thereafter is the average of the current and delayed data. This method effectively restores the "pseudo low-frequency" data to the original clock cycle without the need for additional time domain operations or storage resources.
[0077] The embodiment of the present application stores the enable signal to complete the storage of the effective cost, and restores to the normal timing through a bilinear interpolation method. No additional timing conversion module is required, and no multiple clock domains need to be introduced, which increases reliability and reduces power consumption.
[0078] In one embodiment, if Figure 4 As shown, based on the binocular stereo vision matching method based on pooling, the present invention also provides a binocular stereo vision matching device based on pooling, including:
[0079] An acquisition module 100 is used to acquire a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image;
[0080] A determination module 200, configured to determine a cost space of a current pixel on a first image on the second image;
[0081] A pooling module 300, configured to perform a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space;
[0082] The aggregation module 400 is used to calculate the pooling value of each window in the updated cost space, and determine the pixel in the second image that matches the current pixel on the first image based on each of the pooling values.
[0083] Figure 5 A schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include:
[0084] A memory 501 , a processor 502 , and a computer program stored in the memory 501 and executable on the processor 502 .
[0085] When the processor 502 executes the program, the binocular stereo vision matching method based on pooling provided in the above embodiment is implemented.
[0086] Furthermore, the terminal further includes:
[0087] The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0088] The memory 501 is used to store computer programs that can be executed on the processor 502 .
[0089] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0090] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0091] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0092] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0093] This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the binocular stereo vision matching method based on pooling as described above is implemented.
[0094] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0095] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0096] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, in which the order shown or discussed may not be followed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0097] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can read instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary and then storing it in a computer memory.
[0098] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0099] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0100] In addition, each functional unit in each embodiment of the present application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above-mentioned embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above-mentioned embodiments within the scope of the present application.
[0101] In summary, the present invention discloses a binocular stereo vision matching method and device based on pooling, and the binocular stereo vision matching method based on pooling includes: obtaining a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image; determining the cost space of the current pixel on the first image on the second image; performing a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space; calculating the pooling value of each window in the updated cost space, and determining the pixel in the second image that matches the current pixel on the first image based on each of the pooling values. The present application reduces the cost space to a preset size by performing a pooling operation on the original cost space, thereby reducing the occupation of storage resources and thus reducing the requirements for storage resources.
[0102] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A binocular stereo vision matching method based on pooling, characterized in that: The method comprises: Acquire a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image; Determine a cost space of a current pixel on the first image on the second image; Performing a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space; A pooling value for each window in the updated cost space is calculated, and a pixel in the second image that matches a current pixel in the first image is determined based on each of the pooling values.
2. The binocular stereo vision matching method based on pooling according to claim 1, characterized in that: Before determining the cost space of the current pixel on the first image on the second image, the method further includes: After performing epipolar correction on the first image and the second image, extracting a first image feature of the first image and extracting a second image feature of the second image; A similarity between the first image feature and the second image feature is calculated.
3. The binocular stereo vision matching method based on pooling according to claim 2, characterized in that: Determining a cost space of a current pixel on the first image on the second image includes: Get the preset disparity range and cost data width; A cost space of a current pixel on the first image on the second image is determined according to the similarity, the image heights and image widths of the first image and the second image, the disparity range, and the cost data width.
4. The binocular stereo vision matching method based on pooling according to claim 3, characterized in that: Performing a pooling operation on the cost space to form an updated cost space includes: Perform a 2×2 pooling operation on the cost space, reduce the image widths of the first image and the second image to half of the original, and reduce the disparity range to a quarter of the original, so as to reduce the cost space to a preset size, and divide the cost space into multiple 2×2 windows to form an updated cost space.
5. The binocular stereo vision matching method based on pooling according to claim 1, characterized in that: Calculate the pooling value of each window in the updated cost space, including: The pooling value of each window in the updated cost space is calculated using a preset pooling strategy, where the preset pooling strategy is a minimum pooling strategy, a maximum pooling strategy, a median pooling strategy, or a mean pooling strategy.
6. The binocular stereo vision matching method based on pooling according to claim 1, characterized in that: The pooled value is position information represented by a 2-bit sign bit.
7. The binocular stereo vision matching method based on pooling according to claim 1, characterized in that: Determining a pixel in the second image that matches a current pixel in the first image based on each of the pooling values includes: When up-sampling the pooled value using a bilinear interpolation method, activating a pre-stored enable signal once every two clock cycles so that an output result changes once every two clock cycles; A pixel in the second image that matches a current pixel in the first image is determined based on the output result.
8. A binocular stereo vision matching device based on pooling, characterized in that: The device comprises: An acquisition module, used for acquiring a binocular image pair to be processed, wherein the binocular image pair includes a first image and a second image; A determination module, used to determine a cost space of a current pixel on the first image on the second image; A pooling module, used for performing a pooling operation on the cost space to reduce the cost space to a preset size to form an updated cost space; An aggregation module is used to calculate a pooling value for each window in the updated cost space, and determine a pixel in the second image that matches a current pixel on the first image based on each of the pooling values.
9. A terminal, characterized in that: include: A memory, a processor, and a binocular stereo vision matching program based on pooling stored in the memory and executable on the processor, wherein the binocular stereo vision matching program based on pooling, when executed by the processor, implements the steps of the binocular stereo vision matching method based on pooling as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the binocular stereo vision matching method based on pooling as described in any one of claims 1 to 7.