Binocular semi-global matching hardware acceleration method and device based on FPGA
By implementing hardware acceleration of binocular semi-global matching algorithm on FPGA, the problem of high CPU computing resources is solved, the operation efficiency and positioning accuracy of SLAM systems are improved, and it is suitable for embedded systems with resource-constrained resources.
Patent Information
- Application Number
- CN202510223949.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
When the binocular semi-global matching algorithm is run on the CPU platform, it causes serious consumption of CPU computing resources, affecting the operation efficiency and positioning accuracy of the SLAM system.
Using the FPGA-based hardware acceleration method, the calculation-intensive tasks such as Census transformation and cost aggregation are transferred from the CPU to the FPGA for processing, and the high parallel computing power and low power consumption characteristics are used to achieve efficient semi-global matching calculation of binocular images.
It significantly reduces the CPU load, improves the running speed and positioning accuracy of the SLAM algorithm, reduces system power consumption, and is suitable for embedded systems or mobile devices with resource-constrained.
Smart Images

Figure CN120164084A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a binocular semi-global matching hardware acceleration method and device based on FPGA. Background Art
[0002] In SLAM (Simultaneous Localization and Mapping) technology, the stereo matching of binocular images is usually the core link of the visual odometry module, and its output result provides a key initial estimate for the subsequent pose optimization algorithm. However, this algorithm needs to process all pixel points in the image frame, and when running on a CPU platform, it needs to wait for the entire image data to be received, and at the same time, it needs to perform multiple rounds of traversal calculations on the image data. This calculation mode causes serious consumption of CPU computing resources and affects the overall operation efficiency and positioning accuracy of the SLAM system. Summary of the Invention
[0003] In order to solve the problem that the binocular semi-global algorithm completely depends on the computing resources of the CPU and the frequent reading and writing operations of the memory increase the overhead of CPU resources, affecting the reconstruction efficiency of the algorithm, this application provides a binocular semi-global matching hardware acceleration method and device based on FPGA.
[0004] In a first aspect, this application provides a binocular stereo vision hardware acceleration method based on FPGA, which is applied to FPGA. The method includes:
[0005] Receiving left-view data and right-view data of the same image, and respectively converting the left-view data and the right-view data into first left-view data and first right-view data in grayscale format;
[0006] Respectively inputting the first left-view data and the first right-view data into the first-in-first-out memory FIFO of the line buffer module, and reading second left-view data and second right-view data from the line buffer module;
[0007] Respectively converting the second left-view data and the second right-view data into a left-view Census statistical sequence and a right-view Census statistical sequence;
[0008] Based on the left-view Census sequence and the right-view Census sequence, calculating the initial cost value of the matching point;
[0009] Performing cost aggregation on the initial cost value to obtain an aggregated cost value;
[0010] According to the aggregated cost value and the sub-pixel interpolation formula, obtaining an optimized disparity value.
[0011] Further, the step of converting the left view data and the right view data into first left view data and first right view data in grayscale format respectively specifically includes:
[0012] The FPGA receives the left view data and the right view data, and performs clock synchronization on the left view data and the right view data through the FIFO hardware module;
[0013] Convert the left view data and the right view data from RGB format data to grayscale format data, and shift the converted data 8 bits to the right respectively to obtain the first left view data and the first right view data.
[0014] Further, the step of respectively inputting the first left view data and the first right view data into the first-in-first-out memory (FIFO) of the line buffer module and reading second left view data and second right view data from the line buffer module specifically includes:
[0015] Input the first left view data and the first right view data into the FIFO queues in the line buffer module respectively. When the first FIFO is full of a row of data, the data will be output from the first FIFO and input into the second FIFO, while new row data continues to be input into the first FIFO, and is passed to the (n - 1)-th FIFO in turn. Eventually, each of the (n - 1) FIFOs stores a row of image data;
[0016] Read one pixel data from each of the (n - 1) FIFOs simultaneously, as well as the pixel data newly input into the first FIFO, to generate pixel data of consecutive n rows and located in the same column, thereby obtaining the second left view data and the second right view data.
[0017] Further, the step of respectively converting the second left view data and the second right view data into a left view Census sequence and a right view Census sequence specifically includes:
[0018] Obtain the second left view data and the second right view data respectively through line buffering, wherein the line buffering is used to store the pixel data of the current processing row and its adjacent rows;
[0019] Use a sliding window to slide on the second left view data and the second right view data respectively, and process the pixel data within one window each time;
[0020] For each window, the central pixel is used as the reference pixel, and other pixels are compared with the central pixel in terms of grayscale value. For each pixel within the window, if its grayscale value is lower than or equal to the grayscale value of the central pixel, it is marked as 0, and if its grayscale value is higher than the grayscale value of the central pixel, it is marked as 1;
[0021] Thus, for each pixel within each window, a binary bit of data is generated for each pixel except the central pixel, obtaining the left-view Census sequence and the right-view Census sequence.
[0022] Further, based on the left-view Census sequence and the right-view Census sequence, calculating an initial cost value column for matching points specifically includes:
[0023] Performing real-time data storage on the input right-view Census sequence, and the storage depth is equal to the set maximum disparity distance d max ;
[0024] Within one clock cycle, performing parallel Hamming distance calculation on the data of d max right-view Census sequences stored and the left-view Census sequence of the matching point in the left view, obtaining all the initial cost values of the matching points in the left view within the disparity range from 0 to d max The Hamming distance calculation formula is C(p, d) = hamming(T(p), T(p d ));
[0025] In the formula, T(p) is the Census sequence obtained by the point p in the left view, and T(p d ) is the Census sequence of the corresponding point p in the right view at a disparity of d after the Census transform, and C(p, d) is the initial cost value obtained after the Hamming distance calculation.
[0026] Further, performing cost aggregation on the initial cost values to obtain aggregated cost values specifically includes:
[0027] Reading the initial cost values from the cache FIFO and performing storage;
[0028] Reading the historical matching cost values from the dual-port RAM, that is, d max matching cost values corresponding to p - r, where p - r represents the previous matching pixel point of the current pixel point p in the direction r, finding the minimum value from the d max matching cost values of p - r, and storing it;
[0029] Using the single-direction matching cost value calculation formula, calculating the matching point p at a disparity of d along;
[0030] The single-direction matching cost value calculation formula is as follows:
[0031]
[0032] In the formula, L r(p, d) represents the matching cost of the matching pixel point p with a parallax of d in the direction r, and C(p, d) represents the initial cost value of the pixel point p at the same parallax. It is used to represent the minimum matching cost of the previous pixel point with a parallax of d in the same direction. This is to avoid excessive accumulation of the matching cost, where P1 is the penalty for small parallax changes and P2 is the penalty for large parallax changes.
[0033] Add the matching cost results in four directions, and the aggregated cost value of the matching point p at the parallax d can be obtained.
[0034] Furthermore, according to the aggregated cost value and the sub-pixel interpolation formula, the optimized parallax value is obtained, specifically including:
[0035] Read the aggregated cost value from the upper-level cache FIFO, and use the register to latch d max such aggregated cost values.
[0036] Adopt the winner-takes-all strategy to calculate the minimum aggregated cost value C0, so as to obtain the minimum parallax d0 corresponding to the matching point p. At the same time, search the data latched by the register, read the parallax values C1 and C2 corresponding to the matching point p at the parallax of d0 - 1 and d0 + 1, and substitute the corresponding values into the calculation using the sub-pixel interpolation formula, and the optimized best parallax d of the right-view target point corresponding to the left-view matching point p can be obtained. sub The sub-pixel interpolation formula is expressed as In the formula, d is the parallax before optimization, and d sub is the parallax after optimization.
[0037] In the second aspect, the present application also provides a binocular stereo vision hardware acceleration device based on FPGA, which is applied to FPGA. The method includes:
[0038] The first processing module is used to receive the left-view data and the right-view data of the same image, and convert the left-view data and the right-view data into the first left-view data and the first right-view data in grayscale format respectively.
[0039] The second processing module is used to input the first left-view data and the first right-view data into the first-in-first-out memory FIFO of the line buffer module respectively, and read the second left-view data and the second right-view data from the line buffer module.
[0040] The third processing module is used to convert the second left-view data and the second right-view data into the left-view Census statistical sequence and the right-view Census statistical sequence respectively.
[0041] The fourth processing module is used to calculate the initial cost value of the matching points based on the left view Census sequence and the right view Census sequence;
[0042] The fifth processing module is used to perform cost aggregation on the initial cost value to obtain the aggregated cost value;
[0043] The sixth processing module is used to obtain the optimized disparity value according to the aggregated cost value and the sub-pixel interpolation formula.
[0044] Furthermore, the first processing module is specifically configured to receive the left view data and the right view data by the FPGA, and perform clock synchronization on the left view data and the right view data through the FIFO hardware module;
[0045] Convert the left view data and the right view data from RGB format data to grayscale format data, and shift the converted data 8 bits to the right respectively to obtain the first left view data and the first right view data.
[0046] Furthermore, the second processing module is specifically configured to input the first left view data and the first right view data into the FIFO queues in the line buffer module respectively. When the first FIFO is full of one line of data, the data will be output from the first FIFO and input into the second FIFO, and at the same time, new line data continues to be input into the first FIFO, and is passed to the (n - 1)th FIFO in turn. Finally, one line of image data is stored in the (n - 1)th FIFOs;
[0047] Read one pixel data from the (n - 1)th FIFOs simultaneously, and the pixel data newly input into the first FIFO to generate continuous n rows of pixel data located in the same column, so as to obtain the second left view data and the second right view data.
[0048] Furthermore, the third processing module is specifically configured to obtain the second left view data and the second right view data respectively by means of line buffering, wherein the line buffering is used to store the pixel data of the current processing line and its adjacent lines;
[0049] Use a sliding window to slide on the second left view data and the second right view data respectively, and process the pixel data within one window each time;
[0050] For each window, the central pixel is used as the reference pixel, and other pixels are compared with the central pixel in terms of grayscale value. For each pixel within the window, if its grayscale value is lower than or equal to the grayscale value of the central pixel, it is marked as 0, and if its grayscale value is higher than the grayscale value of the central pixel, it is marked as 1;
[0051] Thus, each pixel within each window generates one bit of data except for the central pixel, resulting in the left-view Census sequence and the right-view Census sequence.
[0052] Further, the fourth processing module is specifically configured to perform real-time data storage on the input right-view Census sequence, and the storage depth is equal to the set maximum disparity distance d max ;
[0053] Within one clock cycle, perform parallel Hamming distance calculation on the data of d max right-view Census sequences stored and the left-view Census sequence of the matching point in the left view, to obtain all the initial cost values of the matching point in the left view within the disparity range from 0 to d max The Hamming distance calculation formula is C(p, d) = hamming(T(p), T(p d ));
[0054] In the formula, T(p) is the Census sequence obtained by the point p in the left view, and T(p d ) is the Census sequence corresponding to the point p in the right view after the Census transform at a disparity of d, and C(p, d) is the initial cost value obtained after the Hamming distance calculation.
[0055] Further, the fifth processing module is specifically configured to read out the initial cost value from the cache FIFO and perform storage;
[0056] Read the historical matching cost values from the dual-port RAM, that is, the d max matching cost values corresponding to p - r. The point p - r represents the previous matching pixel point of the current pixel point p in the direction r. Find the minimum value among the dmax matching cost values of p - r and store it;
[0057] Use the unidirectional matching cost value calculation formula to calculate the matching cost of the matching point p at a disparity of d along;
[0058] The unidirectional matching cost value calculation formula is as follows:
[0059]
[0060] In the formula, L r (p, d) represents the matching cost of the matching pixel point p at a disparity of d in the direction r, C(p, d) represents the initial cost value of the pixel point p at the same disparity, which is used to represent the minimum matching cost of the previous pixel point at a disparity of d in the same direction, This is to avoid excessive accumulation of matching costs, where P1 is the penalty for small parallax changes and P2 is the penalty for large parallax changes;
[0061] Add the matching cost results in four directions to obtain the aggregated cost value of the matching point p at the parallax d.
[0062] Further, the sixth processing module is specifically configured to read the aggregated cost value from the upper-level cache FIFO and latch d max of the aggregated cost values using a register;
[0063] Adopt the winner-takes-all strategy to calculate the minimum aggregated cost value C0, thereby obtaining the minimum parallax d0 corresponding to the matching point p. At the same time, search the data latched by the register, read the parallax values C1 and C2 corresponding to the matching point p at the parallax of d0 - 1 and d0 + 1, and substitute the corresponding values into the sub-pixel interpolation formula for calculation, that is, obtain the optimized best parallax d of the target point in the right view corresponding to the matching point p in the left view. sub The sub-pixel interpolation formula is expressed as In the formula, d is the parallax before optimization, and d sub is the parallax after optimization.
[0064] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the FPGA-based binocular semi-global matching hardware acceleration method described in any one of the first aspects.
[0065] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the score calibration method for the network alarm prediction model described in any one of the first aspects.
[0066] By hardware-accelerating the binocular semi-global matching algorithm, the present invention effectively solves the problems that when the SLAM algorithm performs stereo matching calculations at the front end, the calculation speed of the algorithm decreases and the system power consumption increases due to over-reliance on the computing resources of the CPU. Utilizing the high parallel computing ability and low power consumption characteristics of FPGA (Field Programmable Gate Array), it transfers computationally intensive tasks such as Census transform and cost aggregation from the CPU to the FPGA for processing, significantly reducing the CPU load. By adopting pipeline and parallel processing technologies, it performs efficient semi-global matching calculations on binocular images, achieving real-time and high-precision depth information generation. In addition, by adjusting the direction of the cost set, in the disparity calculation stage, a binomial-based sub-pixel interpolation technology is introduced, enabling the disparity calculation and optimization processes to be synchronized, reducing the CPU calculation delay, further reducing resource consumption, and improving the reconstruction efficiency of the SLAM algorithm. Compared with traditional pure software implementations, the hardware acceleration method of the present invention not only significantly improves the running speed of the algorithm but also significantly reduces the power consumption, and is applicable to resource-constrained embedded systems or mobile devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0068] Figure 1 It is a flowchart of the binocular semi-global matching hardware acceleration method based on FPGA in an embodiment of the present invention;
[0069] Figure 2 It is a flowchart of the operation of the row Census transform module in an embodiment of the present invention;
[0070] Figure 3 It is a calculation flowchart of the Hamming distance calculation module in an embodiment of the present invention;
[0071] Figure 4 It is a calculation flowchart of unidirectional cost aggregation in an embodiment of the present invention;
[0072] Figure 5 It is a schematic diagram of the four-path cost aggregation direction in an embodiment of the present invention;
[0073] Figure 6 It is a calculation flowchart of disparity calculation and disparity optimization in an embodiment of the present invention;
[0074] Figure 7 It is a system framework diagram of the experimental platform in an embodiment of the present invention;
[0075] Figure 8 It is an actual image data diagram collected by a binocular camera in an embodiment of the present invention;
[0076] Figure 9 These are the disparity maps of the binocular image matching results and the matching results of the traditional method in the embodiments of the present invention;
[0077] Figure 10 It is a schematic flowchart of a binocular semi-global matching hardware acceleration method based on FPGA provided in another embodiment of the present application;
[0078] Figure 11 It is a schematic diagram of a binocular semi-global matching hardware acceleration device based on FPGA provided in another embodiment of the present application. Detailed implementation manners
[0079] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0080] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed descriptions of the present invention. Unless otherwise specified, all technical terms adopted by the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the present invention are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention.
[0081] Binocular stereo vision is a perceptual ability based on the principle of binocular parallax. It has always been an important topic in the field of computer vision research and is widely used in fields such as 3D modeling, autonomous vehicle driving, robot obstacle avoidance, and UAV navigation. The key to binocular stereo vision lies in the stereo matching algorithm. According to different matching point search strategies, traditional stereo matching algorithms are usually divided into three types. The first type is the local matching algorithm. This type of algorithm mainly establishes a matching window on the image and uses the local information of the matching window for disparity calculation. However, since the local matching algorithm uses the local information of the pixel neighborhood for calculation, it has a low computational complexity and is suitable for the application of real-time systems. But this method only relies on local information and ignores the correlation between pixels outside the window, resulting in low matching accuracy in some special regions of the image. The second type is the global matching algorithm. This type of algorithm is based on Terzopoulos' energy minimization theory and introduces a smoothing term when calculating the energy function. The global matching algorithm considers both local information and global constraints when processing the entire image, so it performs better in dealing with occlusion and weak texture problems. However, its large computational amount makes the matching speed slow and it is difficult to meet the requirements of embedded real-time processing. The third type is the semi-global matching algorithm (SGM). This type of algorithm uses mutual information (MI) as the basis for cost calculation. The semi-global matching algorithm combines the advantages of local and global matching algorithms. This method draws on the efficiency of local methods when calculating the initial matching cost and introduces global smoothing and edge constraints during the cost aggregation process, thus taking into account both the calculation speed while improving the matching accuracy.
[0082] However, compared with general algorithms, the binocular semi-global matching algorithm has a large computational amount. The complex processing algorithm makes the serial processing operation speed slow using the CPU instruction set and cannot meet the highly parallel computing requirements of binocular semi-global matching. Usually, GPU is used for hardware acceleration. However, in SLAM (Simultaneous Localization and Mapping) technology, the stereo matching of binocular images is usually the core link of the visual odometry module, and its output result provides a key initial estimate for the subsequent pose optimization algorithm. Since this algorithm needs to process all pixel points in the image frame and waits for the entire image data to be received when running on the CPU platform, and also needs to perform multiple rounds of traversal calculations on the image data, this calculation mode leads to serious consumption of CPU computing resources, affecting the overall operation efficiency and positioning accuracy of the SLAM system. In addition, the CPU is good at processing parallel computing tasks, especially image processing and deep learning, and is also suitable for executing binocular stereo matching calculation tasks. However, its high power consumption and large volume limit its application in the embedded field.
[0083] To solve the problem that the binocular semi-global algorithm completely relies on the computing resources of the CPU and the frequent read and write operations on the memory increase the overhead of CPU resources, affecting the reconstruction efficiency of the algorithm, this application proposes a hardware acceleration method for binocular semi-global matching based on FPGA. This method is applied to FPGA and includes the following steps:
[0084] Step S1: In each camera clock, the binocular camera respectively acquires the left view and the right view, and sends the left view data and the right view data in 24-bit RGB format to the FPGA.
[0085] After receiving the left view data and the right view data, the FPGA synchronizes the data through the FIFO hardware module, and then sends the data to the grayscale conversion module. The grayscale conversion module converts the left view data and the right view data in RGB format into the left view data and the right view data in GRAY format respectively, and then sends the converted left view data and right view data in GRAY format to the line buffer module.
[0086] Specifically, the conversion formula for converting RGB format data into GRAY format data is:
[0087] GRAY = 0.299×R + 0.584×G + 0.114×B
[0088] Where, R is the value of the red channel, G is the value of the green channel, B is the value of the blue channel, and GRAY is the calculated grayscale value.
[0089] In addition, to avoid the resource consumption and performance bottleneck caused by floating-point operations, all coefficients are multiplied by 2 during specific implementation 8 to convert them into integers, and the integer result is right-shifted by 8 bits to obtain the final result.
[0090] Step S2: Use the line buffer module for the left view data and the right view data in GRAY format calculated in S1 to respectively obtain multiple first-in-first-out queues, which can be used for subsequent Census transforms.
[0091] The line buffer module consists of several queues based on the first-in-first-out FIFO structure, and each queue is used to store single-row image data. Through the processing of this module, the subsequent algorithm can synchronously obtain the pixel values of multiple consecutive rows in the same column at the rising edge of the same clock cycle, providing the necessary data support for the subsequent module to extract a sliding window of a specific size in the image.
[0092] In this example, the line buffer module is designed to have 4 rows. Such a design enables each clock cycle to simultaneously generate pixel data of 5 consecutive rows and in the same column, providing a data basis for the subsequent 5×5 Census transform calculation window.
[0093] During the processing of the line buffer module, in each clock cycle, the image data is input row by row into the first FIFO. When the first FIFO is full of one row of data, the data will be output from the first FIFO and input into the second FIFO. At the same time, new row data continues to be input into the first FIFO. This process will be sequentially passed to the (n - 1)-th FIFO. Eventually, each FIFO stores one row of image data.
[0094] Similarly, during the reading process, in each clock cycle, one pixel data is read from each FIFO simultaneously. In this way, it is possible to generate consecutive n rows of pixel data located in the same column. These data will be sent to the calculation window of the subsequent module for Census transform.
[0095] S3. For the left-view data and right-view data that have undergone line buffering in S2, respectively, use the data caching method to intercept them with a sliding window to obtain the left-view window image and the right-view window image. Convert the left-view window image and the right-view window image into Census sequences respectively, and send the obtained Census sequences of the left view and the right view to the subsequent calculation module.
[0096] Census transform is a local feature description method for image processing. This method uses the local gray-scale differences around the pixel points to convert the gray-scale values into binary strings.
[0097] In the specific hardware acceleration algorithm of Census transform, as Figure 2 shown, first obtain a 5×5 sliding calculation window through line buffering. Compare the central pixel in the window with other pixels. Those lower than or equal to the reference value are marked as 0, and those higher are marked as 1. Thus, a binary Census sequence with a length of 24 bits is obtained.
[0098] Assume a 5×5 sliding window is used, and the central pixel of the window is the reference pixel. Each pixel in the window except the central pixel will be compared with the central pixel to generate a binary bit. 0 and 1 indicate whether the pixel is greater than the central pixel. The generated 24-bit binary sequence is used as the Census sequence.
[0099] S4. For the Census sequences of the left-view data and the right-view data generated by the transform in S3, calculate the Hamming distance between the Census sequences of two matching points to obtain the initial cost value of the matching points.
[0100] Hamming distance is a measure to evaluate the difference between two binary strings of the same length. In the calculation of the initial matching cost value, it is used to calculate the similarity between the neighborhoods of two pixels. Perform a bitwise exclusive OR operation on the Census sequences of the matching points of the left image and the right image, and count the number of 1s in the obtained binary bit string, which is denoted as the Hamming distance between the two matching points.
[0101] In the specific hardware acceleration algorithm for calculating the Hamming distance, such as Figure 3 shown, the Census sequence of the input right view is stored in real-time data, and the storage depth is equal to the set maximum disparity distance d max , within one clock cycle, the data of the d max Census sequences stored are used to perform parallel Hamming distance calculation with the Census sequence of the matching points of the left view, and the initial cost values of all the matching points of the left view within the disparity range from 0 to d max are obtained. The Hamming distance calculation formula is C(p, d) = hamming(T(p), T(p d ))
[0102] where T(p) is the Census sequence obtained by the point p of the left view, and T(p d ) is the Census sequence of the corresponding point p of the right view at the disparity d after the Census transform. C(p, d) is the initial cost value obtained after the Hamming distance calculation.
[0103] By calculating the Hamming distance between the Census sequences, the initial cost values of the matching points can be obtained. This matching method based on binary sequences has high calculation efficiency, is suitable for hardware implementation, and has good robustness under illumination changes and noise interference. The calculation of the Hamming distance is an important step in stereo matching and can provide reliable basic data for subsequent disparity optimization.
[0104] S5. Select a 4-direction scanning architecture that is the same as the FPGA image data flow direction, and perform cost aggregation on the initial cost values obtained in S4.
[0105] In the specific cost aggregation acceleration algorithm, such as Figure 4 shown, take the cost aggregation of a single method as an example. First, read the initial cost values from the cache FIFO and store them, and then read the historical matching cost values from the dual-port RAM, that is, the d max matching cost values corresponding to the point p - r, find the minimum value among them and store it. Using the single-direction matching cost value calculation formula, calculate the matching cost value of the matching point p at the disparity d along, such as Figure 5The matching cost calculation results in the lr0 direction shown. Add the matching cost results in the four directions to obtain the total aggregation cost value of the matching point p at the disparity d. The calculation formula for the matching cost value in a single direction is as follows,
[0106] In the formula, L r (p, d) represents the matching cost of the matching pixel point p at the disparity d in the direction r. p - r represents the previous matching pixel point of the matching pixel point p in the direction r. C(p, d) represents the initial cost value of this pixel point at the same disparity. The second term on the right side of the formula is used to represent the minimum matching cost of the previous pixel point at the disparity d in the same direction. The role of the third term is to avoid excessive accumulation of the matching cost. Among them, P1 is the penalty for small disparity changes, and P2 is the penalty for large disparity changes. Aggregate the matching costs in each decomposed direction, and the total aggregation cost calculation formula is S(p, d) = ∑ r L r (p, d).
[0107] Perform cost aggregation along the following 4 directions respectively: horizontal direction (from left to right), vertical direction (from top to bottom), diagonal direction (from top left to bottom right), and diagonal direction (from top right to bottom left). For each direction, calculate the single - direction matching cost L r (p, d), and aggregate the matching costs in each decomposed direction.
[0108] In step S5, cost aggregation is performed on the initial cost value through a 4 - direction scanning architecture, which can effectively smooth the disparity map and improve the matching accuracy. The calculation of the single - direction matching cost takes into account the smoothness of the disparity change and controls the amplitude of the disparity change through the penalty coefficients P1 and P2. Finally, aggregate the matching costs in all directions to obtain the total aggregation cost S(p, d) for subsequent disparity selection. This method of multi - direction cost aggregation can significantly improve the effect of stereo matching while ensuring the calculation efficiency.
[0109] S6. Receive the aggregation cost value calculated in S5. According to the d max aggregation cost values of the received matching points, calculate the initial disparity of the matching points, and then use the calculated initial disparity and the aggregation cost value in S5 for disparity optimization.
[0110] In the specific hardware acceleration algorithm for disparity calculation and disparity optimization, as Figure 6 shown, the disparity calculation and optimization module reads the aggregation cost value from the upper - level cache FIFO and uses registers for d maxLatch the aggregated cost value, calculate the minimum aggregated cost value C0 using the winner-takes-all strategy, so as to obtain the minimum disparity d0 corresponding to the matching point p. At the same time, search for the data latched in the register, read the disparity values C1 and C2 corresponding to the matching point p at disparities d0 - 1 and d0 + 1 respectively, and substitute the corresponding values into the sub-pixel interpolation formula for calculation, that is, obtain the optimized best disparity d of the target point in the right view corresponding to the matching point p in the left view sub , and the sub-pixel interpolation formula is expressed as In the formula, d is the disparity before optimization, and d sub is the disparity after optimization.
[0111] Select the disparity that minimizes the aggregated cost value as the initial disparity using the winner-takes-all strategy. Based on the initial disparity, optimize the disparity using the sub-pixel interpolation method to obtain a sub-pixel level disparity with higher precision. The sub-pixel interpolation method can significantly improve the resolution of the disparity map, especially in areas where the disparity changes smoothly, and can effectively reduce the jagged effect of the disparity map, thereby improving the accuracy of stereo matching.
[0112] The embodiment of the present application also provides a binocular stereo vision hardware acceleration method based on FPGA, which is applied to FPGA, as Figure 9 shown, and includes the following steps:
[0113] 110. Receive the left view data and the right view data of the same image, and convert the left view data and the right view data into the first left view data and the first right view data in grayscale format respectively;
[0114] 120. Input the first left view data and the first right view data into the first-in first-out memory FIFO of the line buffer module respectively, and read the second left view data and the second right view data from the line buffer module;
[0115] 130. Convert the second left view data and the second right view data into the left view Census statistical sequence and the right view Census statistical sequence respectively;
[0116] 140. Calculate the initial cost value of the matching point based on the left view Census sequence and the right view Census sequence;
[0117] 150. Aggregate the initial cost value to obtain the aggregated cost value;
[0118] 160. Obtain the optimized disparity value according to the aggregated cost value and the sub-pixel interpolation formula.
[0119] Further, step 110 specifically includes:
[0120] The FPGA receives the left view data and the right view data, and performs clock synchronization on the left view data and the right view data through a FIFO hardware module;
[0121] The left view data and the right view data are converted from RGB format data to grayscale format data, and the converted data are respectively right-shifted by 8 bits to obtain the first left view data and the first right view data.
[0122] Furthermore, step 120 specifically includes:
[0123] Input the first left view data and the first right view data into the FIFO queues in the line buffer module respectively. When the first FIFO is full of one line of data, the data is output from the first FIFO and input into the second FIFO. Meanwhile, new line data continues to be input into the first FIFO and is sequentially transmitted to the n-1th FIFO. Finally, one line of image data is stored in each of the n-1 FIFOs.
[0124] One pixel data is read from n-1 FIFOs simultaneously, as well as the pixel data newly input into the first FIFO, to generate n consecutive rows of pixel data located in the same column, thereby obtaining the second left view data and the second right view data.
[0125] Furthermore, step 130 specifically includes:
[0126] Respectively acquiring the second left view data and the second right view data by means of a line buffer, wherein the line buffer is used to store pixel data of a currently processed line and its adjacent lines;
[0127] Using a sliding window, sliding on the second left view data and the second right view data respectively, processing pixel data within one window each time;
[0128] For each window, the central pixel is used as the reference pixel, and the grayscale values of other pixels are compared with the central pixel. For each pixel in the window, if its grayscale value is lower than or equal to the grayscale value of the central pixel, it is marked as 0, and if its grayscale value is higher than the grayscale value of the central pixel, it is marked as 1;
[0129] Thus, each pixel in each window except the central pixel generates a binary bit data, and the left view Census sequence and the right view Census sequence are obtained.
[0130] Furthermore, step 140 specifically includes:
[0131] The input right view Census sequence is stored in real time, and the storage depth is equal to the set maximum parallax distance d max;
[0132] Within one clock cycle, perform parallel Hamming distance calculation on the data of the d max right view Census sequences stored and the left view Census sequences of the matching points in the left view, and obtain all the initial cost values of the matching points in the left view within the parallax range from 0 to d max The Hamming distance calculation formula is C(p, d) = hamming(T(p), T(p d ));
[0133] In the formula, T(p) is the Census sequence obtained by the point p in the left view, and T(p d ) is the Census sequence of the corresponding point p in the right view at a parallax of d after the Census transform. C(p, d) is the initial cost value obtained after the Hamming distance calculation.
[0134] Furthermore, step 150 specifically includes:
[0135] Read the initial cost value from the cache FIFO and store it;
[0136] Read the historical matching cost values from the dual-port RAM, that is, the d max matching cost values corresponding to p - r. p - r represents the previous matching pixel point of the current pixel point p in the direction r. Find the minimum value among the d max matching cost values of p - r and store it;
[0137] Use the unidirectional matching cost value calculation formula to calculate the matching cost of the matching point p at a parallax of d along;
[0138] The unidirectional matching cost value calculation formula is as follows:
[0139]
[0140] In the formula, L r (p, d) represents the matching cost of the matching pixel point p at a parallax of d in the direction r. C(p, d) represents the initial cost value of the pixel point p at the same parallax. is used to represent the minimum matching cost of the previous pixel point at a parallax of d in the same direction. is to avoid excessive accumulation of the matching cost;
[0141] Add the matching cost results in the four directions to obtain the aggregated cost value of the matching point p at a parallax of d.
[0142] Furthermore, step 160 specifically includes:
[0143] Read the aggregated cost value from the upper-level cache FIFO, and latch d max aggregated cost values using a register;
[0144] Adopt the winner-takes-all strategy to calculate the minimum aggregated cost value C0, so as to obtain the minimum disparity d0 corresponding to the matching point p. At the same time, search the data latched by the register, read the disparity values C1 and C2 corresponding to the matching point p at the corresponding disparities of d0 - 1 and d0 + 1, and substitute the corresponding values into the calculation using the sub-pixel interpolation formula, that is, obtain the optimized best disparity d of the right-view target point corresponding to the left-view matching point p sub , and the sub-pixel interpolation formula is expressed as In the formula, d is the disparity before optimization, and d sub is the disparity after optimization.
[0145] In this embodiment, by hardware-accelerating the binocular semi-global matching algorithm, the problem that the SLAM algorithm has a decreased calculation speed and an increased system power consumption due to over-reliance on the CPU's computing resources when implementing stereo matching calculations at the front end is effectively solved. Utilizing the high parallel computing ability and low power consumption characteristics of FPGA (Field Programmable Gate Array), computing-intensive tasks such as Census transform and cost aggregation are transferred from the CPU to the FPGA for processing, significantly reducing the CPU load. By adopting pipeline and parallel processing technologies, efficient semi-global matching calculations are performed on binocular images, realizing real-time and high-precision depth information generation. In addition, by adjusting the direction of the cost set, in the disparity calculation stage, the binomial-based sub-pixel interpolation technology is introduced, enabling the disparity calculation and optimization processes to be synchronized, reducing the CPU calculation delay, further reducing resource consumption, and improving the reconstruction efficiency of the SLAM algorithm. Compared with the traditional pure software implementation, the hardware acceleration method of the present invention not only greatly improves the running speed of the algorithm, but also significantly reduces the power consumption, and is applicable to resource-constrained embedded systems or mobile devices.
[0146] In addition, the embodiment of the present application also provides a binocular stereo vision hardware acceleration device based on FPGA, which is applied to FPGA, as Figure 10 shown, including:
[0147] The first processing module is used to receive the left-view data and the right-view data of the same image, and convert the left-view data and the right-view data into the first left-view data and the first right-view data in grayscale format respectively;
[0148] The second processing module is used to input the first left-view data and the first right-view data into the first-in-first-out memory FIFO of the line buffer module respectively, and read the second left-view data and the second right-view data from the line buffer module;
[0149] A third processing module for separately converting the second left view data and the second right view data into a left view Census statistical sequence and a right view Census statistical sequence;
[0150] A fourth processing module for calculating an initial cost value of matching points based on the left view Census sequence and the right view Census sequence;
[0151] A fifth processing module for performing cost aggregation on the initial cost value to obtain an aggregated cost value;
[0152] A sixth processing module for obtaining an optimized disparity value according to the aggregated cost value and a sub-pixel interpolation formula.
[0153] Further, the first processing module is specifically configured to: the FPGA receives the left view data and the right view data, and performs clock synchronization on the left view data and the right view data through a FIFO hardware module;
[0154] Convert the left view data and the right view data from RGB format data to grayscale format data, and shift the converted data 8 bits to the right respectively to obtain the first left view data and the first right view data.
[0155] Further, the second processing module is specifically configured to separately input the first left view data and the first right view data into FIFO queues in the row buffer module. When the first FIFO is full of one row of data, the data will be output from the first FIFO and input into the second FIFO, while new row data continues to be input into the first FIFO, and is sequentially passed to the (n - 1)th FIFO. Finally, one row of image data is stored in each of the (n - 1) FIFOs;
[0156] Read one pixel data from each of the (n - 1) FIFOs simultaneously, as well as the pixel data newly input into the first FIFO, to generate pixel data of consecutive n rows and in the same column, thereby obtaining the second left view data and the second right view data.
[0157] Further, the third processing module is specifically configured to separately obtain the second left view data and the second right view data through row buffering, wherein the row buffering is used to store pixel data of the current processing row and its adjacent rows;
[0158] Use a sliding window to slide on the second left view data and the second right view data respectively, and process pixel data within one window each time;
[0159] For each window, the central pixel serves as the reference pixel, and the gray values of other pixels are compared with that of the central pixel. For each pixel within the window, if its gray value is lower than or equal to that of the central pixel, it is marked as 0; if its gray value is higher than that of the central pixel, it is marked as 1;
[0160] Thus, for each pixel within each window except the central pixel, a binary bit of data is generated, obtaining the left-view Census sequence and the right-view Census sequence.
[0161] Further, the fourth processing module is specifically configured to perform real-time data storage on the input right-view Census sequence, and the storage depth is equal to the set maximum disparity distance d max ;
[0162] Within one clock cycle, perform parallel Hamming distance calculation on the data of d max right-view Census sequences stored and the left-view Census sequence of the matching point in the left view, obtaining all the initial cost values of the matching point in the left view within the disparity range from 0 to d max The Hamming distance calculation formula is C(p, d) = hamming(T(p), T(p d ));
[0163] In the formula, T(p) is the Census sequence obtained by the point p in the left view, T(p d ) is the Census sequence of the corresponding point p in the right view at the disparity of d after the Census transform, and C(p, d) is the initial cost value obtained after the Hamming distance calculation.
[0164] Further, the fifth processing module is specifically configured to read out the initial cost value from the cache FIFO and perform storage;
[0165] Read the historical matching cost values from the dual-port RAM, that is, the d max matching cost values corresponding to p - r. p - r represents the previous matching pixel point of the current pixel point p in the direction r. Find the minimum value among the d max matching cost values of p - r and perform storage;
[0166] Use the unidirectional matching cost value calculation formula to calculate the matching cost value of the matching point p at the disparity of d along;
[0167] The unidirectional matching cost value calculation formula is as follows:
[0168]
[0169] In the formula, L r(p, d) represents the matching cost of the matching pixel point p with a parallax of d in the direction r, and C(p, d) represents the initial cost value of the pixel point p at the same parallax. It is used to represent the minimum matching cost of the previous pixel point with a parallax of d in the same direction. This is to avoid excessive accumulation of the matching cost, where P1 is the penalty for small parallax changes and P2 is the penalty for large parallax changes.
[0170] Add the matching cost results in four directions to obtain the aggregated cost value of the matching point p at the parallax d.
[0171] Furthermore, the sixth processing module is specifically configured to read the aggregated cost value from the upper-level cache FIFO and latch d max of these aggregated cost values using a register.
[0172] Adopt the winner-takes-all strategy to calculate the minimum aggregated cost value C0, thereby obtaining the minimum parallax d0 corresponding to the matching point p. At the same time, search the data latched by the register, read the parallax values C1 and C2 corresponding to the matching point p at the parallax of d0 - 1 and d0 + 1, and substitute the corresponding values into the sub-pixel interpolation formula for calculation, that is, obtain the optimized best parallax d of the right-view target point corresponding to the left-view matching point p. sub The sub-pixel interpolation formula is expressed as In the formula, d is the parallax before optimization, and d sub is the parallax after optimization.
[0173] Specific implementation method 1: As Figures 1 to 9 shown, the present invention provides a binocular semi-global matching hardware acceleration method based on FPGA, including the following steps:
[0174] S1. Grayscale conversion
[0175] In each camera clock, the binocular camera sends 24-bit RGB format data to the FPGA. After receiving the RGB data, the FPGA first synchronizes the data through the FIFO, then sends the data to the grayscale conversion module to convert the RGB format data into GRAY format data, and then sends the converted GRAY format data to the line buffer module.
[0176] The Census transform of the image is based on the grayscale value of the image. Since the data received by the FPGA is in RGB format, the RGB format data is first converted into GRAY format data, and its conversion formula is as follows.
[0177] GRAY = 0.299×R + 0.584×G + 0.114×B
[0178] Wherein, R is the value of the red channel, G is the value of the green channel, B is the value of the blue channel, and GRAY is the calculated grayscale value.
[0179] To avoid resource consumption and performance bottlenecks caused by floating-point operations, in the specific implementation, all coefficients are multiplied by 28 to be converted into integers, and the integer result is right-shifted by 8 bits to obtain the final result.
[0180] S2. Row buffer
[0181] The row buffer module consists of several queues based on the first-in-first-out (FIFO) structure, and each queue is used to store single-row image data. Through the processing of this module, the subsequent algorithm can synchronously obtain the pixel values of multiple consecutive rows in the same column at the rising edge of the same clock cycle, providing the necessary data support for the subsequent module to extract a sliding window of a specific size in the image.
[0182] In this example, the row buffer module is designed to have 4 rows. Such a design enables the generation of pixel data of 5 consecutive rows in the same column in each clock cycle, providing a data basis for the subsequent 5×5 Census transform calculation window.
[0183] S3. Census transform
[0184] The Census transform module receives the image data input by the row buffer module in step S2, intercepts a 5×5 sliding image calculation window through data caching, converts the image data within the sliding window into a Census sequence, and finally sends the Census sequence to the subsequent calculation module.
[0185] The Census transform is a local feature description method for image processing. This method uses the local gray-scale differences around the pixel points to convert the gray-scale values into binary strings.
[0186] In the specific hardware acceleration algorithm of the Census transform, as Figure 2 shown, first, a 5×5 sliding calculation window is obtained through the row buffer, the central pixel within the window is compared with other pixels, those lower than or equal to the reference value are marked as 0, and those higher are marked as 1. Thus, a 24-bit binary Census sequence is obtained.
[0187] S4. Hamming distance calculation
[0188] The Hamming distance calculation module receives the Census sequence generated by the Census transform module, calculates the Hamming distance between the Census sequences of two matching points, and obtains the initial cost value of the matching points.
[0189] Hamming distance is a measure to evaluate the difference between two binary strings of the same length and is used to calculate the similarity between the neighborhoods of two pixels in the calculation of the initial matching cost value. Perform a bitwise exclusive OR operation on the Census sequences of the matching points of the left image and the right image, and count the number of 1s in the obtained binary bit string, which is recorded as the Hamming distance between the two matching points.
[0190] In a specific hardware acceleration algorithm for calculating Hamming distance, such as Figure 3 shown, the input Census sequence of the right image is stored in real-time data, and the storage depth is equal to the set maximum disparity distance d max , and within one clock cycle, the d max stored Census sequences are used to perform parallel Hamming distance calculations with the Census sequences of the matching points of the left image, and the initial cost values of all the matching points of the left image within the disparity range from 0 to d max are obtained. The Hamming distance calculation formula is as follows.
[0191] C(p, d) = hamming(T(p), T(p d )) (2)
[0192] In the formula, T(p) is the Census sequence obtained by the Census transform of point p in the left image, and T(p d ) is the Census sequence of the corresponding point p in the right image at a disparity of d after the Census transform. C(p, d) is the initial cost value obtained after the Hamming distance calculation.
[0193] S5. Cost aggregation
[0194] The cost aggregation module selects a 4-direction scanning architecture that is the same as the FPGA image data flow direction, and receives the initial cost values calculated in S4 for four-path cost aggregation.
[0195] The purpose of cost aggregation is to optimize the initial cost by combining local and global information. The initial matching cost often includes noise and unreliable matching results, so cost aggregation is needed to improve the robustness and accuracy of the matching.
[0196] In a specific cost aggregation acceleration algorithm, such as Figure 4 shown, take the cost aggregation of a single method as an example. First, the initial cost values are read from the cache FIFO and stored, and then the historical matching cost values, that is, the d max matching cost values corresponding to point p - r, are read from the dual-port RAM, the minimum value among them is found and stored, and using the single-direction matching cost value calculation formula, the matching cost value of point p at a disparity of d along as Figure 5The matching cost calculation results in the lr0 direction shown. By adding the matching cost results in the four directions, the total aggregated cost value of the matching point p at the disparity d can be obtained. The calculation formula for the matching cost value in a single direction is as follows,
[0197]
[0198] In the formula, L r (p, d) represents the matching cost of the matching pixel point p at the disparity d in the direction r. p - r represents the previous matching pixel point of the matching pixel point p in the direction r. C(p, d) represents the initial cost value of this pixel point at the same disparity. The second term on the right side of the formula is used to represent the minimum matching cost of the previous pixel point at the disparity d in the same direction. The role of the third term is to avoid excessive accumulation of the matching cost. By aggregating the matching costs in each decomposed direction, the calculation formula for the total aggregated cost is as follows,
[0199] s(p, d) = ∑ r L r (p, d) (4)
[0200] S6. Disparity Calculation and Disparity Optimization
[0201] S6. Disparity calculation and disparity optimization module. First, it receives the aggregated cost value calculated by S5. Based on the d max aggregated cost values of the received matching points, it calculates the initial disparity of the matching points, and then uses the calculated initial disparity and the aggregated cost value of S5 for disparity optimization.
[0202] In the specific hardware acceleration algorithm for disparity calculation and disparity optimization, as Figure 6 shown, the disparity calculation and disparity optimization module reads the aggregated cost value from the upper-level cache FIFO, uses the register to latch the d max aggregated cost values, adopts the winner-takes-all strategy to calculate the minimum aggregated cost value C0, thereby obtaining the minimum disparity d0 corresponding to the matching point p. At the same time, it searches the data latched by the register, reads the disparity values C1 and C2 corresponding to the matching point p at the disparities d0 - 1 and d0 + 1, and substitutes the corresponding values into the sub-pixel interpolation formula for calculation, that is, the optimized best disparity d of the target point corresponding to the left image matching point P is obtained sub , and the sub-pixel interpolation formula is expressed as:
[0203]
[0204] In the formula, d is the disparity before optimization, and d sub is the disparity after optimization.
[0205] Specific Embodiment 2: The present invention provides a binocular semi-global matching hardware acceleration method and system based on FPGA, including,
[0206] A grayscale conversion module for converting RGB format data transmitted by a binocular camera to the FPGA into GRAY format data;
[0207] A line buffer module for storing grayscale image data as multiple FIFO queues;
[0208] A Census transform module for converting pixel data stored in the line buffer module into a binary Census sequence;
[0209] A Hamming distance calculation module for calculating the Hamming distance between binary Census sequences to obtain an initial cost value;
[0210] A cost aggregation module for performing 4-path cost aggregation calculation on the initial cost value.
[0211] A disparity calculation and disparity optimization module for calculating an initial disparity based on the aggregated cost value, and then optimizing the initial disparity according to the initial disparity and the aggregated cost value to finally obtain an optimized disparity.
[0212] Other combinations and connection relationships in this embodiment are the same as those in the above embodiment.
[0213] Next, using the Cones image of the Middlebury binocular stereo matching dataset as a test image, as Figure 8 shown, an experimental device as Figure 7 was built. In the debugging environment, the computer program was downloaded into the experimental device through an EDA tool. The program reads the left and right binocular images, and the test output matching result is as Figure 9 shown. The optimized binocular semi-global matching algorithm provides more accurate matching results in areas with continuous disparity, effectively reducing the ripple phenomenon and improving the accuracy of disparity matching.
[0214] In this experiment, a hardware acceleration scheme for the algorithm was designed using a hardware description language, and a dedicated experimental platform was built to verify its effectiveness. The experimental results show that while moderately improving the calculation accuracy, this scheme significantly improves the real-time performance of the calculation, reduces the CPU load, and thus overall optimizes the execution efficiency of the SLAM algorithm.
[0215] In addition, an embodiment of the present application includes a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements a method for calibrating scores of a network alarm prediction model as described in any one of the above technical solutions.
[0216] The embodiment of the present application further includes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for calibrating scores of a network alarm prediction model according to any one of the above technical solutions.
[0217] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still modifications or equivalent replacements can be made to the specific implementation manners of the present invention, and any modifications or equivalent replacements without departing from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A binocular stereo vision hardware acceleration method based on FPGA, characterized in that: Applied to FPGA, the method comprises: Receiving left view data and right view data of the same image, and converting the left view data and the right view data into first left view data and first right view data in grayscale format, respectively; Inputting the first left view data and the first right view data into a first-in first-out memory FIFO of a line buffer module respectively, and reading second left view data and second right view data from the line buffer module; Convert the second left view data and the second right view data into a left view Census statistical sequence and a right view Census statistical sequence respectively; Calculating an initial cost value of a matching point based on the left view Census sequence and the right view Census sequence; Performing cost aggregation on the initial cost value to obtain an aggregated cost value; According to the aggregation cost value and the sub-pixel interpolation formula, an optimized disparity value is obtained.
2. The method according to claim 1, characterized in that: The converting the left view data and the right view data into first left view data and first right view data in grayscale format respectively specifically includes: The FPGA receives the left view data and the right view data, and performs clock synchronization on the left view data and the right view data through a FIFO hardware module; The left view data and the right view data are converted from RGB format data to grayscale format data, and the converted data are respectively right-shifted by 8 bits to obtain the first left view data and the first right view data.
3. The method according to claim 1, characterized in that: The step of inputting the first left view data and the first right view data into a first-in first-out memory FIFO of a line buffer module respectively, and reading the second left view data and the second right view data from the line buffer module specifically includes: Input the first left view data and the first right view data into the FIFO queues in the line buffer module respectively. When the first FIFO is full of one line of data, the data is output from the first FIFO and input into the second FIFO. Meanwhile, new line data continues to be input into the first FIFO and is sequentially transmitted to the n-1th FIFO. Finally, one line of image data is stored in each of the n-1 FIFOs. One pixel data is read from n-1 FIFOs simultaneously, as well as the pixel data newly input into the first FIFO, to generate n consecutive rows of pixel data located in the same column, thereby obtaining the second left view data and the second right view data.
4. The method according to claim 1, characterized in that: The converting the second left view data and the second right view data into a left view Census sequence and a right view Census sequence respectively specifically includes: Respectively acquiring the second left view data and the second right view data by means of a line buffer, wherein the line buffer is used to store pixel data of a currently processed line and its adjacent lines; Using a sliding window, sliding on the second left view data and the second right view data respectively, processing pixel data within one window each time; For each window, the central pixel is used as the reference pixel, and the grayscale values of other pixels are compared with the central pixel. For each pixel in the window, if its grayscale value is lower than or equal to the grayscale value of the central pixel, it is marked as 0, and if its grayscale value is higher than the grayscale value of the central pixel, it is marked as 1; Thus, each pixel in each window except the central pixel generates a binary bit data, and the left view Census sequence and the right view Census sequence are obtained.
5. The method according to claim 1, characterized in that: The calculating the initial cost value sequence of the matching points based on the left view Census sequence and the right view Census sequence specifically includes: The input right view Census sequence is stored in real time, and the storage depth is equal to the set maximum parallax distance d max ; In one clock cycle, the stored d max The data of the right view Census sequence and the left view Census sequence of the matching points of the left view are parallelly calculated to obtain the matching points of the left view from 0 to d. max The total initial cost value of the parallax range, the Hamming distance calculation formula C(p,d) = hamming(T(p),T(p d )), Where, T(p) is the Census sequence obtained by passing through point p in the left view, T(p d ) is the Census sequence of the point p corresponding to the right view after Census transformation when the disparity is d, and C(p,d) is the initial cost value obtained after Hamming distance calculation.
6. The method according to claim 1, characterized in that: The cost aggregation of the initial cost value to obtain the aggregated cost value specifically includes: Reading the initial cost value from the cache FIFO and storing it; Read the historical matching cost value from the dual-port RAM, that is, the d corresponding to the point pr max The point pr represents the previous matching pixel of the current pixel p in the direction r, from pr's d max Find the minimum value among the matching cost values and store it; Using the calculation formula of the matching cost value in one direction, the matching point p is calculated when the parallax is d; The calculation formula for the one-way matching cost is as follows: Where, L r (p, d) represents the matching cost of the matching pixel point p with a disparity of d in direction r, and C(p, d) represents the initial cost value of the pixel point p under the same disparity. It is used to express the minimum matching cost of the previous pixel in the same direction when the disparity is d. It is to avoid excessive accumulation of matching costs, where P1 is the penalty for small parallax changes and P2 is the penalty for large parallax changes; Add the matching cost results in the four directions to obtain the aggregate cost value of the matching point p under disparity d.
7. The method according to claim 1, characterized in that The step of obtaining an optimized disparity value according to the aggregation cost value and the sub-pixel interpolation formula specifically includes: Read the aggregate cost value from the upper cache FIFO and use the register to convert d max The aggregate cost is locked up; The winner-takes-all strategy is used to calculate the minimum aggregation cost C0, thereby obtaining the minimum disparity d0 corresponding to the matching point p. At the same time, the data latched in the register is searched to read the disparity values C1 and C2 corresponding to the matching point p at the corresponding disparity d0-1 and d0+1. The corresponding values are substituted into the calculation using the sub-pixel interpolation formula, and the optimal disparity d after optimization of the left view matching point p corresponding to the right view target point is obtained. sub , the sub-pixel interpolation formula is expressed as Where d is the parallax before optimization, d sub This is the optimized parallax.
8. A binocular stereo vision hardware acceleration device based on FPGA, characterized in that: Applied to FPGA, the method comprises: A first processing module, configured to receive left view data and right view data of the same image, and convert the left view data and the right view data into first left view data and first right view data in grayscale format, respectively; A second processing module, configured to input the first left view data and the first right view data into a first-in first-out memory FIFO of a line buffer module respectively, and read second left view data and second right view data from the line buffer module; A third processing module, configured to convert the second left view data and the second right view data into a left view Census statistical sequence and a right view Census statistical sequence respectively; A fourth processing module, configured to calculate an initial cost value of a matching point based on the left view Census sequence and the right view Census sequence; A fifth processing module, configured to perform cost aggregation on the initial cost value to obtain an aggregated cost value; The sixth processing module is used to obtain an optimized disparity value according to the aggregation cost value and the sub-pixel interpolation formula.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the FPGA-based binocular stereo vision hardware acceleration method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the FPGA-based binocular stereo vision hardware acceleration method according to any one of claims 1 to 7 is implemented.