High-performance real-time stereo image processing system based on software and hardware fusion
By adopting the software and hardware fusion architecture of the ZYNQ heterogeneous platform in the stereo image processing system, the problem of insufficient real-time and flexibility in the prior art is solved, high-performance, real-time stereo image processing is achieved, and rich information output is provided.
Patent Information
- Application Number
- CN202510034921.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-09
AI Technical Summary
The existing stereoscopic image processing solutions have shortcomings in real-time and flexibility, and it is difficult to meet the scenario requirements of high-demand, high-performance, and real-time processing.
Adopting a software and hardware fusion architecture based on the heterogeneous platform ZYNQ, the software and hardware fusion architecture of the dual-core ARM Cortex-A9 software processing unit PS and the FPGA programmable logic unit PL are deeply integrated to achieve efficient collaborative processing tasks and provide rich information output.
It realizes seamless connection with various types of image acquisition equipment, improving the universality and scalability of the system; ensuring efficient image processing performance and reducing response delay; it has dynamic parameter adjustment function to adapt to different scenes and image characteristics; and provides more comprehensive and intuitive visual information.
Smart Images

Figure CN119967145A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to a high-performance real-time stereoscopic image processing system, which can be used for robot navigation, automatic driving and three-dimensional reconstruction. Background Art
[0002] The core tasks of stereo image processing include image acquisition, correction, matching, optimization and output. With the development of hardware technology, FPGA and heterogeneous computing platforms are increasingly used in this field, but existing stereo image processing solutions still face the problems of insufficient real-time performance and flexibility. For example, solutions based on CPU or GPU use software libraries to process images. Although they are highly flexible, their computing performance cannot meet real-time requirements. Although solutions based on FPGA can improve performance, they lack flexibility and have a long development cycle and high threshold. Some heterogeneous platform solutions do not optimize task allocation and module collaboration, and do not give full play to the advantages of combining software and hardware.
[0003] The patent document with publication number CN111445380B discloses a method and device for achieving real-time binocular stereo matching based on ZYNQ. It is based on the heterogeneous platform ZYNQ, including a software processing unit PS and a programmable logic unit PL. It uses the programmable logic PL plug-in storage module and takes advantage of the FPGA hardware logic resource to process data in parallel, and achieves real-time binocular stereo matching calculation in a pipeline and parallel processing manner. PS is responsible for driving the camera to trigger, extract images, interact and control data with the programmable logic PL, and read the disparity calculation results and intermediate calculation results. However, this patent has the following shortcomings: First, there are certain restrictions on input devices, and it is not easy to seamlessly connect with various types of image acquisition devices, which limits the diversity and flexibility of the system in acquiring image data; second, when processing large-scale image data, it may not be able to meet application scenarios with high real-time requirements, resulting in system response delays and affecting real-time processing capabilities; third, since the optimization parameters cannot be adjusted dynamically, the image processing process cannot be flexibly optimized under different environments or image conditions, which reduces the system's adaptability to complex scenes; fourth, the information output is not rich enough, only focusing on the output of parallax data, lacking a variety of information fusion output methods, and unable to provide users with more comprehensive and intuitive visual information, limiting the application of the system in scenarios with high requirements for information display.
[0004] Patent document with publication number CN114757985A discloses a binocular depth perception device and image processing method based on the ZYNQ improved algorithm. It uses the PS and PL ends of the heterogeneous platform ZYNQ to exchange data through the AXI bus to achieve high-quality output of the depth map. Although this method achieves binocular depth perception through a combination of software and hardware, it still has many shortcomings: 1) Since it is only adapted to specific types of image sensors, it lacks broad compatibility with multiple input devices and convenient configuration capabilities, which limits the application of the system in different image acquisition scenarios and cannot quickly adapt to diverse image source requirements, reducing the versatility of the system; 2) Since the optimization of task allocation and module collaboration in terms of module collaboration is not sophisticated enough, the overall flexibility of the system is limited, and it is difficult to flexibly adjust the working mode and resource allocation of each module according to actual application needs; 3) Since it is impossible to dynamically adjust the optimization parameters, when facing different scenes or image characteristics, the processing process cannot be optimized in real time to obtain the best effect, and it lacks adaptability and self-adaptability; 4) Since the information output is relatively single, it only provides basic disparity calculation results and cannot provide rich information output; 5) Due to insufficient flexibility, it is difficult to quickly integrate new functional modules and cannot keep up with the pace of technological development and changes in market demand in a timely manner, limiting the long-term application potential of the system. Summary of the invention
[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a high-performance real-time stereoscopic image processing system based on software and hardware fusion, which can give full play to the advantages of software and hardware by deeply integrating the software processing unit PS and the programmable logic unit PL in the heterogeneous platform ZYNQ, efficiently and collaboratively process tasks and provide rich information output, meet high-demand, high-performance, real-time processing scenarios, and promote the widespread application of stereoscopic image processing technology.
[0006] To achieve the above object, the technical solution of the present invention includes:
[0007] Technical solution 1:
[0008] 1. A high-performance real-time stereoscopic image processing method based on software and hardware fusion is implemented based on the heterogeneous platform ZYNQ, which includes a software processing unit PS with a dual-core ARM Cortex-A9 as the core and an FPGA programmable logic unit PL, characterized in that:
[0009] The resolution and frame rate of the camera are set in the processing unit PS to collect image data; the image is corrected using the OpenCV software algorithm, and the corrected image is transferred to the system memory through the VDMA high-speed interface to establish a high-speed channel between the PS end and the PL end;
[0010] Performing hardware acceleration of a stereo matching algorithm in the programmable logic unit PL by means of hardware logic resources, which includes:
[0011] In the cost calculation phase, the calculation of Census transformation and Hamming distance is split into multiple parallel subtasks for synchronous processing;
[0012] In the cost aggregation stage, the combinational logic is combined with the sequential logic to increase the aggregation speed and reduce the number of logic levels;
[0013] In the parallax optimization phase, shift registers are used to implement left-right consistency detection and median filtering, and penalty parameters can be changed in real time through system buttons;
[0014] In the image stitching and output stage, the OSD IP core is called to stitch the left view and the disparity map, and the fused image is output through the HDMI interface.
[0015] Technical solution 2:
[0016] A high-performance real-time stereoscopic image processing system based on software and hardware fusion, characterized by comprising:
[0017] Image acquisition module, used to connect to the USB binocular camera, obtain the image stream of the left and right views in real time and transfer it to the system memory;
[0018] The image correction module is used to perform distortion and epipolar correction operations on the left and right original images to eliminate lens distortion and color deviation, ensure that the images are geometrically consistent in the same viewing plane, and improve the overall image quality;
[0019] Left and right eye alignment module, used to align the corrected left and right images at the pixel level to speed up image processing and reduce latency;
[0020] The cost calculation module is used to split and parallelize the cost calculation tasks based on Census transformation and Hamming distance calculation, and then generate a cost matrix to optimize resource allocation, reduce data transmission delay, and improve calculation efficiency while ensuring accuracy;
[0021] The cost aggregation module is used to insert timing logic into the combinational logic with the help of DSP and LUT hardware to accelerate the aggregation of the disparity map cost, so as to reduce the processing time of pixels;
[0022] The disparity optimization module is used to further optimize the disparity map after cost aggregation, remove pseudo matching points through left-right consistency detection, and use median filtering to smooth and remove noise;
[0023] The image stitching and display module is used to efficiently stitch the processed left view and the disparity map, and output a three-dimensional view that integrates two-dimensional texture and three-dimensional depth information.
[0024] Compared with the prior art, the present invention has the following advantages:
[0025] Firstly, the present invention adopts a USB camera as an image input source and utilizes the versatility of the USB interface to achieve seamless connection with various types of USB cameras, thereby enhancing compatibility with diverse image input requirements and improving versatility and expansibility;
[0026] Secondly, the present invention uses the software processing unit PS to run the Linux system and combines the OpenCV library to perform image correction, which not only simplifies the development process of image correction, but also ensures a higher correction accuracy;
[0027] Thirdly, the present invention uses an image transmission unit VDMA that can adjust the transmission bandwidth according to the image resolution, so that the image data can be quickly transmitted from the system memory to the FPGA for subsequent processing, thereby avoiding the backlog and delay of image data transmission;
[0028] Fourthly, since the functions of each submodule of the present invention are highly independent, with clear interfaces and functional boundaries, each module can be upgraded or replaced separately according to the needs, so as to adapt to the needs of different application scenarios, and it is easy to integrate new functional modules;
[0029] Fifthly, the present invention adopts asynchronous storage FIFO in cost aggregation, which effectively solves the problem of data synchronization in different clock domains, ensures smooth transmission of data at high frame rate and high resolution images, and avoids blockage and loss; and through the cooperation of combinational logic and sequential logic, the cost aggregation time of a single pixel can be reduced from 12 clock cycles to 3 clock cycles, realizing efficient pipeline processing of image data, and the maximum number of logic levels can be 14, ensuring the processing speed and stability of cost aggregation;
[0030] Sixth, since the present invention designs a dynamic parameter adjustment function for the parallax optimization module, the optimization parameters can be adjusted in real time through system buttons, so that the system can flexibly optimize the image processing flow according to the actual application scenario and image characteristics;
[0031] Seventhly, the present invention can provide users with more comprehensive and intuitive visual information by splicing the processed left and right view images with the depth map, fusing the two-dimensional texture and three-dimensional depth information;
[0032] Eighth, the present invention uses Shift RAM logic to accurately align the pixels of the left and right images. This hardware design can achieve pixel-level accuracy, ensuring that the left and right images have a highly consistent basis in subsequent parallax calculation and optimization, directly improving the accuracy of stereoscopic vision processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a structural block diagram of Embodiment 1 of the present invention;
[0034] Figure 2 This is a structural block diagram of Embodiment 2 of the present invention;
[0035] Figure 3 It is a schematic diagram of the process of calculating the cost value using the image cost calculation in Embodiment 1 of the present invention;
[0036] Figure 4 A schematic diagram of a process of performing cost value aggregation for image cost aggregation in Embodiment 1 of the present invention;
[0037] Figure 5 It is a schematic diagram of the process of performing aggregated disparity optimization for disparity optimization in Embodiment 1 of the present invention;
[0038] Figure 6 A before-and-after comparison diagram of using the present invention to correct a binocular image;
[0039] Figure 7 The figure is a simulation result diagram of aligning and synchronizing the left and right eye images using the present invention;
[0040] Figure 8 This is a simulation result diagram after the image cost is calculated using the present invention;
[0041] Fig. 9 This is a simulation result diagram after cost aggregation using the present invention;
[0042] Fig.10 This is a simulation result diagram after parallax optimization using the present invention;
[0043] Fig.11 This is a simulation result diagram of modifying the P1 and P2 parameters of a key using the present invention;
[0044] Fig.12 This is a simulation result diagram of splicing the left view and the disparity map using the present invention. DETAILED DESCRIPTION
[0045] The embodiments and effects of the present invention are further described in detail below with reference to the accompanying drawings.
[0046] The present invention deeply integrates the software processing unit PS and the programmable logic unit PL in the heterogeneous platform ZYNQ to complement the advantages of software and hardware and realize a high-performance real-time stereoscopic image processing system.
[0047] The software processing unit PS is a processor core integrated with ARM Cortex-A9, which has rich components and interfaces, including various memory interfaces, communication interfaces and interrupt controllers, etc., and has high flexibility and scalability. In the system of the present invention, PS is mainly responsible for running the Linux system and completing image acquisition and correction with the help of the OpenCV library.
[0048] The programmable logic unit PL implements the hardware acceleration function through custom programming logic, has the advantages of high performance, low power consumption, high flexibility and easy integration, and can be dynamically adjusted according to different application scenarios and computing requirements. In the system of the present invention, PL is used to implement the hardware acceleration work of stereo image processing, while ensuring high-performance computing, taking into account the flexibility of the system, optimizing the overall performance, and adapting to the diverse stereo matching computing requirements.
[0049] Embodiment 1: A high-performance real-time stereoscopic image processing method based on software and hardware fusion.
[0050] Reference Figure 1 This example is based on the heterogeneous platform ZYNQ, which includes a software processing unit PS based on dual-core ARM Cortex-A9 and an FPGA programmable logic unit PL. Its implementation includes the following two parts:
[0051] 1. Software processing unit PS part
[0052] Step 1: Setting the resolution and frame rate of the camera in the processing unit PS to collect image data.
[0053] 1.1) Obtain the function control interface of the USB camera according to the video class standard specification;
[0054] 1.2) The resolution of the camera is adjusted by sending a resolution adjustment command to the camera driver, that is, the sampling parameters of the image sensor are changed to achieve adjustment of different resolutions of 640×480 and 1280×720;
[0055] 1.3) The number of image frames collected per second is controlled by adjusting the camera's clock frequency and data transmission rate. Different image frame rates are used according to the requirements of different application scenarios to ensure high-quality image acquisition.
[0056] Step 2: perform image correction on the binocular image data.
[0057] Commonly used image correction algorithms include OpenCV image correction, Matlab image correction, and deep learning image correction. Matlab image correction and deep learning image correction have poor adaptability to the heterogeneous platform ZYNQ and are difficult to guarantee real-time performance. The present invention uses but is not limited to the OpenCV image correction algorithm to fully utilize the advantages of rapid development and flexible programming of the Linux system running on the processing unit PS. The implementation steps include the following:
[0058] 2.1) Use the cv::calibrateCamera function in the calibration tool of OpenCV to set the specification parameters of the chessboard. As an example, assume that the number of corner points in the chessboard is 11*8 and the size of the chessboard is 20mm. By collecting multiple chessboard images of this specification at different angles and positions, the corner points are detected and analyzed, and the intrinsic and extrinsic matrix of the camera is calculated to accurately obtain the internal and external parameter data of the camera;
[0059] 2.2) Use the undistortion function cv::undistort in OpenCV to correct radial and tangential distortions. By remapping each pixel in the image, the distorted pixel coordinates are converted into corrected coordinates, thereby removing the radial and tangential distortions caused by the lens, restoring the true geometric structure of the image, and improving the image quality.
[0060] 2.3) Use the epipolar correction cv::stereoRectify function in OpenCV to calculate the rotation and translation relationship between the left and right cameras, and reproject the left and right images so that the same object points in the two images are located on the same epipolar line, which helps the accuracy and efficiency of disparity calculation in the subsequent stereo matching algorithm and improves the accuracy of disparity matching.
[0061] The effect of the binocular image correction in this example is as follows: Figure 6 As shown, Figure 6 a is the binocular image before correction, Figure 6 b is the corrected binocular image; Figure 6 As can be seen in b, the epipolar lines of the corrected binocular image are almost parallel, which indicates that the correction process successfully makes the epipolar lines parallel to the row direction of the image, which means that when looking for corresponding points in the left and right images, it is only necessary to search on the same row, simplifying the two-dimensional search problem to a one-dimensional search problem, thereby greatly improving the matching efficiency and accuracy.
[0062] Step 3: Establish a high-speed channel between the PS and PL ends through the video direct memory access (VDMA) interface.
[0063] The VDMA interface is a high-speed interface for fast and efficient transmission of video image data between different storage areas or processing units. It can automatically handle operations such as address mapping and data transfer during data transmission without CPU intervention, thereby greatly reducing the CPU burden and improving system performance.
[0064] For example, the transmission mode of the VDMA high-speed interface is set to burst transmission, and the data transmission width is set to 24 bits;
[0065] The corrected image data is read out from the system memory through the VDMA interface, and after waiting for the successful handshake between the sending valid signal vaild and the receiving ready signal ready, the image data is sent to the PL end, thereby establishing an efficient and stable high-speed channel between the PS end and the PL end.
[0066] 2. Programmable logic unit PL part
[0067] Step 4: In the cost calculation phase, the calculation of Census transformation and Hamming distance is split into multiple parallel subtasks for synchronous processing.
[0068] refer to Figure 3 , the implementation of this step includes the following:
[0069] 4.1) Census transformation of binocular pixels:
[0070] The Census calculation of each pixel is split into multiple 5x5 pixel windows, and the pixel value in each window is compared with the center value at the same time through parallel operation: if the pixel value is greater than the center value, 1 is output, and if the pixel value is less than the center value, 0 is output;
[0071] Use the lookup table LUT to calculate the Census transformation result for each pair of pixels to avoid performing an XOR operation each time and save calculation time;
[0072] 4.2) Calculate the Hamming distance of the result of Census transformation:
[0073] XOR and accumulate the Census transformation results of each pair of pixels to obtain a Hamming distance calculation subtask;
[0074] By instantiating multiple Hamming distance calculation subtasks in parallel, the Hamming distances of multiple pixel pairs are calculated in one clock cycle to shorten the cost calculation time.
[0075] Step 5: In the cost aggregation phase, the combinational logic is coordinated with the sequential logic to increase the aggregation speed and reduce the number of logic levels;
[0076] refer to Figure 4, the implementation of this step includes the following:
[0077] 5.1) Through combinatorial logic, the aggregation operation steps that originally need to be executed sequentially are integrated into a combinatorial circuit at one time, eliminating the delay of multi-cycle calculation and improving the calculation efficiency;
[0078] 5.2) Cost image aggregation through temporal logic:
[0079] In the first clock cycle, the cost image data is read from the output buffer of the combinational logic by sampling at the rising edge of the clock signal;
[0080] In the second clock cycle, the digital signal processor DSP and the lookup table LUT in the PL are used to sum and compare the cost image data to complete the local aggregation calculation;
[0081] In the third clock cycle, all local aggregation results are added through the accumulator, and finally the precise aggregation result is output.
[0082] Step 6: In the disparity optimization phase, shift registers are used to implement left-right consistency detection and median filtering, and penalty parameters are changed in real time through system buttons.
[0083] refer to Figure 5 , the implementation of this step includes the following:
[0084] 6.1) Store the disparity value after cost aggregation into the register, swap the left and right images, perform cost aggregation again, get a new disparity value, and then compare the current disparity value in the register with the newly calculated disparity value:
[0085] If the difference between them is less than the preset left-right consistency detection threshold, the current disparity value is retained;
[0086] Otherwise, the current disparity value is regarded as a mismatch point and deleted;
[0087] 6.2) Use a shift register to store a 5x5 window of data. By continuously moving new data into the register, the window data is updated. Each time data is moved in, the median of the data in the window is calculated to complete the median filtering operation.
[0088] 6.3) Use system buttons to change penalty parameters:
[0089] The system keys, in the example of the present invention, refer to physical keys, which include a first key key1 and a second key key2;
[0090] The penalty parameters include a smoothness penalty parameter P1 and a parallax consistency penalty parameter P2:
[0091] In the example, by pressing the first button key1 to change the smoothness penalty parameter P1, the path smoothness of the disparity map is controlled to reduce noise and inconsistency in the disparity map; by pressing the second button key2 to change the disparity consistency penalty parameter P2, the degree of disparity consistency is controlled to enhance consistency and reduce false matching;
[0092] Step 7: In the image stitching and output link, the left view and the disparity map are stitched, and the image is output through the HDMI interface.
[0093] 7.1) Read the storage address of the left view and the disparity map through the OSD IP core, and splice the left view and the disparity map according to the left and right layout;
[0094] 7.2) The spliced left view and disparity map are sent to an external display device through the transmission circuit of the HDMI interface to achieve image output.
[0095] Embodiment 2: A high-performance real-time stereoscopic image processing system based on software and hardware fusion.
[0096] Reference Figure 2 This example is built on the heterogeneous platform ZYNQ, which includes image acquisition module 1, image correction module 2, left and right eye alignment module 3, cost calculation module 4, cost aggregation module 5, parallax optimization module 6, image stitching and display module 7. These modules cooperate with each other to complete the real-time processing of stereo images, among which:
[0097] The image acquisition module 1 is connected to a USB binocular camera and is used to obtain image streams of left and right views in real time and transmit them to the system memory;
[0098] The image correction module 2 is used to read out the left and right original images in the system memory to perform distortion and epipolar correction, thereby eliminating lens distortion and color deviation, ensuring that the images achieve geometric consistency in the same viewing plane, so as to improve the overall quality of the image;
[0099] The left-right eye alignment module 3 uses Shift RAM logic to accurately align the pixels of the corrected left and right images to achieve pixel-level accuracy, ensuring that the left and right images can be processed faster and have less delay in subsequent parallax calculations;
[0100] The cost calculation module 4 performs Census transformation and Hamming distance calculation on the aligned left and right images, splits and processes the cost calculation tasks in parallel, and then generates a cost matrix to reduce data transmission delay and improve calculation efficiency while ensuring accuracy;
[0101] The cost aggregation module 5: adopts an asynchronous FIFO design, and is used to perform cost aggregation on the cost calculation data. It includes a data synchronization submodule 51 and an aggregation operation splitting submodule 52. The data synchronization submodule 51 solves the data transmission delay problem between different clock domains through Gray code pointer and synchronization processing, and improves the transmission flexibility and reliability; the aggregation operation splitting submodule 52 decomposes complex operations into simple logic units, combines combinational logic and sequential logic parallel processing, and improves the aggregation speed and stability, that is, the single pixel cost aggregation time is reduced from 12 clock cycles to 3, and the number of logic levels is up to 14 levels;
[0102] The disparity optimization module 6 is used to optimize the aggregated disparity data, and includes: a disparity data image processing submodule 61 and a system key dynamic parameter adjustment submodule 62. The disparity data image processing submodule 61 is used to process the aggregated disparity data, remove pseudo matching points in the image and remove image noise, improve the quality of the disparity map, and provide more reliable data support for subsequent depth calculation and three-dimensional reconstruction; the system key dynamic parameter adjustment submodule 62 is used to adjust the optimization parameters in real time according to the actual scene and image characteristics, enhance the flexibility and adaptability of the system, and obtain better image processing effects in different scenes.
[0103] The image stitching and display module 7 is used to efficiently stitch the left view after parallax optimization with the parallax map, and output a three-dimensional view that integrates two-dimensional texture and three-dimensional depth information.
[0104] The effect of the present invention can be further illustrated by the following simulation results:
[0105] A simulation condition
[0106] This simulation is performed in the Vivado environment. The left and right eye images taken by the binocular camera are selected, and the image stereo matching process and results are verified through the testbench script.
[0107] 2. Simulation content
[0108] Simulation 1: Under the above conditions, the left and right images are aligned using the method of the present invention. The results are as follows: Figure 7 shown.
[0109] from Figure 7 It can be seen that after the left-right eye alignment process, the left-right eye images are accurately aligned at the pixel level, providing a good foundation for subsequent disparity calculation.
[0110] Simulation 2: Under the above conditions, the cost of the left and right images is calculated using the method of the present invention. The results are as follows: Figure 8 shown.
[0111] from Figure 8It can be seen that there are various noises and texture changes in the disparity map before cost aggregation, which will affect the accuracy of stereo matching.
[0112] Simulation 3: Under the above conditions, the method of the present invention is used to Figure 8 The disparity map shown in the figure is subjected to cost aggregation, and the result is as follows Fig. 9 shown.
[0113] from Fig. 9 It can be seen that after cost aggregation, the disparity image is smoother, and the noise and discontinuous areas are significantly reduced.
[0114] Simulation 4: Under the above conditions, the method of the present invention is used to optimize the parallax of the aggregated image. The results are as follows: Fig.10 shown.
[0115] from Fig.10 It can be seen that after parallax optimization, many noise points are successfully filtered, the mismatched areas are marked as invalid black, the depth information is more continuous and uniform, the texture artifacts in the background area and the surface of the object are basically eliminated, and the effect is closer to the real depth map.
[0116] Simulation 5: Under the above conditions, the method of the present invention is used to modify the parameters P1 and P2 of the left and right eye images. The results are as follows: Fig.11 shown.
[0117] from Fig.11 It can be seen that under the combination of P1 = 20 and P2 = 80, a good balance is achieved between image detail preservation and overall effect, which is a better parameter combination choice. Fig.11 As shown in d.
[0118] Simulation 6: Under the above conditions, the left view and the disparity map of the left and right eye images are spliced using the method of the present invention. The results are as follows: Fig.12 shown.
[0119] from Fig.12 It can be seen that the left view is on the left side of the stitched image, and the disparity map is on the right side of the stitched image, realizing the fusion of the two-dimensional information of the left view and the three-dimensional information of the disparity map, which can provide users with more comprehensive and intuitive visual information.
[0120] The above descriptions are only two specific examples of the present invention and do not constitute any limitation to the present invention. Obviously, for professionals in this field, after understanding the content and principles of the present invention, it is possible to make various modifications and changes in form and details without departing from the principles and structures of the present invention. However, these modifications and changes based on the ideas of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A high-performance real-time stereoscopic image processing method based on software and hardware fusion is implemented based on the heterogeneous platform ZYNQ, which includes a software processing unit PS with a dual-core ARM Cortex-A9 as the core and an FPGA programmable logic unit PL, characterized in that: The resolution and frame rate of the camera are set in the processing unit PS to collect image data; the image is corrected using the OpenCV software algorithm, and the corrected image is transferred to the system memory through the VDMA high-speed interface to establish a high-speed channel between the PS end and the PL end; Performing hardware acceleration of a stereo matching algorithm in the programmable logic unit PL by means of hardware logic resources, which includes: In the cost calculation phase, the calculation of Census transformation and Hamming distance is split into multiple parallel subtasks for synchronous processing; In the cost aggregation stage, the combinational logic is combined with the sequential logic to increase the aggregation speed and reduce the number of logic levels; In the parallax optimization phase, shift registers are used to implement left-right consistency detection and median filtering, and penalty parameters can be changed in real time through system buttons; In the image stitching and output stage, the OSD IP core is called to stitch the left view and the disparity map, and the fused image is output through the HDMI interface.
2. The method according to claim 1, characterized in that The camera driver is set in the processing unit PS, which is used to control the functional operation of the camera according to the standard specification of the USB camera according to its video class, that is, to adjust the resolution, frame rate, and exposure time of the camera to ensure high-quality image acquisition.
3. The method according to claim 1, characterized in that: The image correction using OpenCV software algorithm includes: 3a) Using OpenCV’s calibration tool, collect multiple images of a checkerboard or other calibration target and calculate the camera’s internal and external parameters, including focal length, principal point, and distortion coefficients; 3b) Using the calibrated camera internal and external parts to participate in the dedistortion algorithm in OpenCV, the image is distorted, that is, the radial distortion and tangential distortion caused by the lens are removed to restore the true geometric structure of the image; 3c) Use the calibrated camera internal and external parts to participate in the epipolar correction algorithm in OpenCV to perform epipolar correction on the image, that is, the same object points in the two images are located on the same epipolar line, improving the accuracy of parallax matching.
4. The method according to claim 1, characterized in that The video direct memory access VDMA high-speed interface transmits video data by using a fast channel in the advanced scalable protocol AXI, that is, no CPU intervention is required, which can greatly reduce the burden of the CPU and improve system performance.
5. The method according to claim 1, characterized in that In the cost calculation phase, the calculation of Census and Hamming distance is split into multiple parallel subtasks for synchronous processing, including: 5a) Split the Census calculation of each pixel into multiple 5x5 pixel windows, and compare each window with its center value in parallel: the window with a value greater than the center value outputs 1, and the window with a value less than the center value outputs 0. After all windows are compared, the Census transformation result is obtained; 5b) The Census transformation result is stored in the hardware register of the FPGA, and the DSP and LUT are used to perform parallel XOR operations to obtain the Hamming distance between each pair of pixels.
6. The method according to claim 1, characterized in that In the cost aggregation stage, the combinational logic is coordinated with the sequential logic, including: 6a) by using a multi-input adder, a multiplier and a shift operation in a combinational logic circuit to parallelly calculate the weighted sum of multiple disparity costs, multiple operation steps in the cost aggregation process are merged into a more compact logic structure to reduce the delay of intermediate operations; 6b) Input the processed data of the combinational logic into the sequential logic and aggregate them in order: In the first clock cycle, input data is read and distributed to each computing unit; In the second clock cycle, the computation unit completes the local aggregation calculation and passes the result to the integration unit; In the third clock cycle, the integration unit completes the processing and outputs the final aggregation result.
7. The method according to claim 1, characterized in that In the disparity optimization link, the shift register is used to realize left-right consistency detection and median filtering, including: 7a) Use a shift register to perform a shift operation on the left and right disparity images, that is, compare the two disparity images pixel by pixel to determine whether the left and right disparities are consistent: If the disparity of a certain pixel is inconsistent in the left and right images, it is considered a wrong match and is removed or corrected; If the left and right images are consistent, the result is directly output; 7b) Use a shift register to store a 5x5 window of data. By continuously shifting new data into the register, the window data is updated. Each time data is shifted in, the median of the data in the window is calculated to complete the median filtering operation.
8. The method according to claim 1, characterized in that: The change penalty parameter includes a change to the smoothness penalty item P1 and a change to the disparity consistency penalty item P2, which controls the path smoothness of the disparity map and reduces noise and inconsistency in the disparity map by changing P1; controls the degree of disparity consistency by changing P2, enhances consistency, and reduces false matches; In the image stitching and output link, the OSD IP core is called to stitch the left view and the disparity map, and the storage range of the left view and the disparity map in the memory address is first sent to the OSD IP core, and the image or video stitching function of the OSD IP is used to automatically read the left view and the disparity map to complete the stitching of the two and output the result.
9. A high-performance real-time stereoscopic image processing system based on software and hardware fusion, characterized in that: include: Image acquisition module, used to connect to the USB binocular camera, obtain the image stream of the left and right views in real time and transfer it to the system memory; The image correction module is used to perform distortion and epipolar correction operations on the left and right original images to eliminate lens distortion and color deviation, ensure that the images are geometrically consistent in the same viewing plane, and improve the overall image quality; Left and right eye alignment module, used to align the corrected left and right images at the pixel level to speed up image processing and reduce latency; The cost calculation module is used to split and parallelize the cost calculation tasks based on Census transformation and Hamming distance calculation, and then generate a cost matrix to optimize resource allocation, reduce data transmission delay, and improve computing efficiency while ensuring accuracy. The cost aggregation module is used to insert timing logic into the combinational logic with the help of DSP and LUT hardware to accelerate the aggregation of the disparity map cost, so as to reduce the processing time of pixels; The disparity optimization module is used to further optimize the disparity map after cost aggregation, remove pseudo matching points through left-right consistency detection, and use median filtering to smooth and remove noise; The image stitching and display module is used to efficiently stitch the processed left view and the disparity map, and output a three-dimensional view that combines two-dimensional texture and three-dimensional depth information.
10. The system according to claim 9, characterized in that: The cost aggregation module adopts an asynchronous FIFO design, which includes: The data input synchronization submodule is used to transfer data between different clock domains in the cost aggregation, and through the Gray code pointer design and synchronization processing, it avoids the influence of parallax data transmission delay on the cost aggregation, and improves the flexibility and reliability of data transmission; The aggregation operation splitting submodule is used to coordinate the combinational logic with the sequential logic to split the complex aggregation operation into multiple simple logic units, so that multiple logic operations can be performed simultaneously, thereby improving the speed and stability of cost aggregation.
11. The system according to claim 9, characterized in that: The parallax optimization module comprises: The disparity data image processing submodule is used to process the aggregated disparity data, remove pseudo matching points in the image and remove image noise, improve the quality of the disparity map, and provide more reliable data support for subsequent depth calculation and 3D reconstruction; The system button dynamic parameter adjustment sub-module is used to adjust the optimization parameters in real time according to the actual scene and image characteristics, enhance the flexibility and adaptability of the system, and obtain better image processing effects in different scenes.
Citation Information
Patent Citations
A method and apparatus for real-time binocular stereo matching based on ZYNQ
CN111445380B
Binocular depth perception device based on ZYNQ improved algorithm and image processing method
CN114757985A
Visual odometer based on ZYNQ hardware acceleration
CN108827340A
Software and hardware collaborative design method of binocular stereoscopic vision system
CN110276110A
Semi-global stereo matching method adopting cost fusion and hierarchical matching strategy
CN114299132A
Cited By
FPGA (Field Programmable Gate Array) implementation method for 16-path CAN (Controller Area Network) and 16-path serial port data transceiving
CN120560131A
Real-time polar line correction FPGA (field programmable gate array) pipeline architecture of uncompressed LUT (lookup table)
CN121810477A
A Real-Time Pole Line Correction FPGA Pipeline Architecture with Uncompressed LUTs
CN121810477B
PMSM-FOC motor model heterogeneous simulation method, device, equipment, storage medium and product
CN122450030A