3D image acquisition and processing method based on binocular vision and storage medium
By processing and correcting binocular vision images, and using ISP neural networks and Gaussian difference pyramid techniques for feature matching, the problem of mismatch between left and right images is solved, improving the accuracy and realism of 3D pose estimation.
Patent Information
- Application Number
- CN202511629975.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-08
- Publication Date
- 2026-02-10
AI Technical Summary
In existing technologies, mismatches are common in the left and right images captured by binocular cameras, resulting in inaccurate 3D poses of the target object in the recovered image.
The ISP neural network model is used to process and correct the binocular vision images. The scale transformation and Gaussian difference scale space are constructed by Gaussian convolution kernel to generate a feature vector set. Similarity matching is performed and the intersection is calculated to achieve accurate matching of the left and right images. Finally, the image is restored by inverse discrete wavelet transform.
It improves the accuracy and realism of 3D pose estimation of target objects in images, and ensures the matching accuracy of the same point in the left and right images.
Smart Images

Figure CN121504776A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer vision, in particular to a 3D image acquisition and processing method based on binocular vision and a storage medium. BACKGROUND
[0002] Stereovision is widely used in automatic driving, virtual reality, three-dimensional reconstruction and other fields by simulating human eyes to obtain depth information of the environment.
[0003] In the prior art, left and right images of different angles are collected by a binocular camera, and the left and right images obtained often appear mismatched at the same point, resulting in that the three-dimensional pose of the target object in the recovered image cannot accurately reflect the real pose. SUMMARY
[0004] To solve the technical problems in the background art, the present application provides a 3D image acquisition and processing method based on binocular vision and a storage medium, which helps to match the left and right images collected by the binocular camera at the same point, thereby improving the accuracy and authenticity of the three-dimensional pose estimation of the target object in the image.
[0005] To achieve the above technical solutions, in a first aspect, the present application provides a 3D image acquisition and processing method based on binocular vision, comprising: Step 1: collecting binocular vision images of different angles by a binocular camera; Step 2: processing the collected binocular vision images by an ISP neural network model to obtain processed binocular vision images; Step 3: correcting the processed binocular vision images; Step 4: matching the corrected binocular vision images; Step 5: optimizing the matched binocular vision images to obtain final binocular vision images.
[0006] Further, the step 4 comprises: performing scale transformation on the binocular vision images using a Gaussian convolution kernel to obtain a scale space of the binocular vision images; convolving the function difference between two adjacent Gaussian scale spaces with the vision images to obtain a Gaussian difference scale space; determining the sampling feature points in the binocular vision images; constructing a Gaussian difference pyramid based on the obtained Gaussian difference scale space and determining the position of each sampling feature point in the binocular vision images and the scale space in which the sampling feature point is located based on the constructed Gaussian difference pyramid; Generate the feature vector of each sampled feature point in the binocular vision image based on the position and the scale space where each sampled feature point in the determined binocular vision image is located; Construct the feature vector set of the left vision image and the feature vector set of the right vision image in the binocular vision image based on the feature vector of each sampled feature point in the binocular vision image; Take the left vision image in the binocular vision image as the reference image, and perform similarity matching with the right vision image to obtain the matching point pair set M1; Take the right vision image in the binocular vision image as the reference image, and perform similarity matching with the left vision image to obtain the matching point pair set M2; Find the intersection M of the matching point pair set M1 and the matching point pair set M2 as the final matching point pair set; Update the binocular vision image based on the obtained final matching point pair set to obtain the updated binocular vision image.
[0007] Further, the fifth step includes: Assume that L(x + d1, y) and L(x + d2, y) are two points on a certain epipolar line in the left vision image, where d1 < d2, and R(x, y) is the matching point of L(x + d1, y) and L(x + d2, y) on the right vision image. E(x, y, d1) and E(x, y, d2) respectively represent the matching cost functions; If E(x + d2, x, y) < E(x + d1, x, y), then L(x + d2, y) replaces L(x + d1, y) as the only matching point of R(x, y), and re-search for its best matching point within the disparity range belonging to L(x + d1, y); If E(x + d2, x, y) ≥ E(x + d1, x, y), then L(x + d1, y) still serves as the only matching point of R(x, y); re-search for its best matching point within the disparity range belonging to L(x + d2, y).
[0008] Further, the second step includes: Split the binocular vision image into multiple channels according to pixel arrangement to obtain a preprocessed image; Perform multi-level operations on the obtained preprocessed image to decompose the preprocessed image into images with different frequency bands; Recombine the images with different frequency bands, and perform multi-level inverse operations using the inverse discrete wavelet transform IDWT to restore the recombined image to the size of the original image and output it to obtain the processed binocular vision image.
[0009] Secondly, the present invention provides a computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the computer-readable storage medium is located to execute the above-described method for 3D image acquisition and processing based on binocular vision.
[0010] The beneficial effects of this invention are as follows: This invention processes and corrects binocular vision images using an ISP neural network model before matching them, and further optimizes the matched binocular vision images. This helps to match the same point in the left and right images, thereby improving the accuracy and realism of the 3D pose estimation of the target object in the image. Attached Figure Description
[0011] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0012] Figure 1 This is a flowchart of a 3D image acquisition and processing method based on binocular vision according to the present invention. Detailed Implementation
[0013] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0014] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, each technical and scientific term used in these embodiments has the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0015] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0016] In this invention, terms such as "upper," "lower," "left," "right," "front," "back," "vertical," "horizontal," "side," and "bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used only to facilitate the description of the structural relationships of the various components or elements of this invention and do not specifically refer to any component or element in this invention. They should not be construed as limiting the invention.
[0017] In this invention, terms such as "fixed connection," "connected," and "linked" should be interpreted broadly, indicating a fixed connection, an integral connection, or a detachable connection; a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can determine the specific meaning of these terms in this invention based on the specific circumstances, and they should not be construed as limitations on the invention.
[0018] Example 1: like Figure 1 As shown, this embodiment provides a 3D image acquisition and processing method based on binocular vision, including the following steps: S1: Acquire binocular visual images from different angles using a binocular camera, wherein the binocular visual images include left visual images and right visual images.
[0019] S2: The acquired binocular vision images are processed using the ISP neural network model to obtain the processed binocular vision images.
[0020] Specifically, it includes the following steps: S1-1: The binocular vision image is split into multiple channels according to the pixel arrangement to obtain a preprocessed image; S1-2: Perform multi-level operations on the obtained preprocessed image to decompose the preprocessed image into images with different frequency bands; S1-3: Reconstruct images with different frequency bands and perform multi-level inverse operations using inverse discrete wavelet transform (IDWT) to restore the reconstructed image to the size of the original image and output it, thus obtaining the processed binocular vision image.
[0021] S3: Correct the processed binocular vision image.
[0022] By correcting the processed binocular vision images, any point on the left vision image has the same row number as its corresponding point on the right vision image, so that a one-dimensional search can be performed on that row to match the corresponding point.
[0023] S4: Match the corrected binocular vision images.
[0024] Specifically, assuming the left visual image is P1 and the right visual image is P2, the specific steps include: S4-1: Apply a Gaussian convolution kernel to the binocular vision image P(x, y) to perform a scale transformation, thereby obtaining the scale space of the binocular vision image; the formula is as follows: (1) in, Scale space factor; The scale space representing the binocular vision image P(x,y); Let be a function representation in Gaussian scale space, which is expressed as: (2); S4-2: The difference between two adjacent Gaussian scale spaces is convolved with the binocular vision image P(x, y) to obtain the Gaussian difference scale space. The specific formula is as follows: (3) in, Represents the Gaussian difference scale space. Indicates and Functions in adjacent Gaussian scale spaces.
[0025] S4-3: Based on the obtained Gaussian difference scale space, construct a Gaussian difference pyramid, and determine the position of each sampled feature point and its scale space based on the constructed Gaussian difference pyramid.
[0026] S4-4: Filter and locate the obtained sampling feature points and their scale space, specifically including the following steps: A: Eliminating contrast points: Assuming the offset distance between candidate feature point X and the sampling point is... Expanding equation (3) using Taylor's formula yields the following equation: (4); in, .
[0027] B: By taking the derivative of formula (4) and setting the derivative to zero, the positions of the candidate feature points can be obtained. ,and Substituting into formula (4), we get ; C: The result obtained The precise location of the feature point is obtained by comparing it with a preset threshold T.
[0028] This point needs to satisfy the following equation: .
[0029] S4-5: Based on the positions of the acquired feature points, generate a feature vector for each feature point; where the feature set of the left visual image is represented as: The feature set of the right visual image is represented as: ; S4-6: Using the right visual image as the reference image, perform similarity matching between each feature vector in the feature set of the right visual image and each feature vector in the feature set of the left visual image to obtain the matching point pair set M1. The specific steps are as follows: A: Select any eigenvector from the feature set of the right visual image, and calculate the Euclidean distance between this eigenvector and the eigenvectors in the feature set of the left visual image. The comparison formula is as follows:
[0030] Among them, represents the eigenvector and is the Euclidean distance; B: Based on the calculated Euclidean distance, determine the eigenvector closest to this eigenvector and the second-closest eigenvector .
[0031] C: Based on the Euclidean distance between two points, determine whether the ratio of the two is less than a preset threshold. If it is less than, the eigenvector is a matching point.
[0032] D: Repeat the above steps A to C until all matching point pairs M1 in the right visual image are found; S4-7: Take the left visual image as the reference image, and perform similarity matching between each eigenvector in the feature set of the left visual image and each eigenvector in the feature set of the right visual image to obtain the matching point pair set M2; It should be noted that the steps are the same as those in S4-6.
[0033] S4-8: Find the intersection of the matching point pair sets M1 and M2, denoted as , then M is the final matching point pair set.
[0034] S4-9: Based on the obtained final matching point pair set, update the binocular visual image to obtain the updated binocular visual image.
[0035] S5: Optimize the updated binocular visual image to obtain the final binocular visual image.
[0036] Specifically, it includes the following steps: Assume that L(x + d1, y) and L(x + d2, y) are two points on a certain epipolar line of the left visual image, where d1 < d2, and R(x, y) is the matching point of L(x + d1, y) and L(x + d2, y) on the right visual image. E(x, y, d1) and E(x, y, d2) respectively represent the matching cost functions.
[0037] If E(x + d2, x, y) < E(x + d1, x, y), then L(x + d2, y) replaces L(x + d1, y) as the only matching point of R(x, y), and the best matching point of L(x + d1, y) is searched again within the range of the disparity value belonging to L(x + d1, y).
[0038] If E(x + d2, x, y) ≥ E(x + d1, x, y), then L(x + d1, y) still serves as the only matching point of R(x, y); the best matching point of L(x + d2, y) is searched again within the range of the disparity value belonging to L(x + d2, y), and the final binocular vision image is obtained in this way.
[0039] For the same and similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the descriptions in the method embodiments.
[0040] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the systems or units can be in electrical, mechanical or other forms.
[0041] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0042] In addition, it should be noted that the flowcharts in the drawings show the methods of the embodiments of the present disclosure. In the corresponding descriptions in the flowcharts or block diagrams in the drawings, the operations or steps corresponding to different blocks may also occur in a different order from that disclosed in the description. Sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and sometimes they can be executed in the reverse order, which depends on the functions involved. Each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that executes the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0043] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for 3D image acquisition and processing based on binocular vision, characterized in that, Including: Step 1: Collect binocular vision images at different angles through a binocular camera; Step 2: Process the collected binocular vision images through an ISP neural network model to obtain processed binocular vision images; Step 3: Correct the processed binocular vision images; Step 4: Match the corrected binocular vision images; Step 5: Optimize the matched binocular vision images to obtain the final binocular vision images.
2. The 3D image acquisition and processing method based on binocular vision according to claim 1, characterized in that, The said Step 4 includes: Perform scale transformation on the binocular vision images using a Gaussian convolution kernel to obtain the scale space of the binocular vision images; Perform convolution on the function difference between two adjacent Gaussian scale spaces and the vision images to obtain the Gaussian difference scale space; Determine the sampling feature points in the binocular vision images; According to the obtained Gaussian difference scale space, construct a Gaussian difference pyramid, and based on the constructed Gaussian difference pyramid, determine the position and the scale space where each sampling feature point in the binocular vision images is located; Based on the position and the scale space where each sampling feature point in the binocular vision images is located, generate the feature vectors of each sampling feature point in the binocular vision images; Based on the feature vectors of each sampling feature point in the binocular vision images, construct the feature vector set of the left vision image and the feature vector set of the right vision image in the binocular vision images; Take the left vision image in the binocular vision images as the reference image and perform similarity matching with the right vision image to obtain the matching point pair set M1; Take the right vision image in the binocular vision images as the reference image and perform similarity matching with the left vision image to obtain the matching point pair set M2; Find the intersection M of the matching point pair set M1 and the matching point pair set M2 as the final matching point pair set; Based on the obtained final matching point pair set, update the binocular vision images to obtain the updated binocular vision images.
3. The 3D image acquisition and processing method based on binocular vision according to claim 1, characterized in that, The said Step 5 includes: Assume that L(x + d1, y) and L(x + d2, y) are two points on a certain epipolar line of the left vision image, where d1 < d2, and R(x, y) is the matching point of L(x + d1, y) and L(x + d2, y) on the right vision image, and E(x, y, d1) and E(x, y, d2) respectively represent the matching cost functions; If E(x + d2, x, y) < E(x + d1, x, y), then L(x + d2, y) replaces L(x + d1, y) as the only matching point of R(x, y), and re-search for its best matching point within the disparity range belonging to L(x + d1, y); If E(x + d2, x, y) ≥ E(x + d1, x, y), then L(x + d1,y) still serves as the only matching point of R(x, y); re-search for its best matching point within the disparity range belonging to L(x + d2, y).
4. The 3D image acquisition and processing method based on binocular vision according to claim 1, characterized in that, The said Step 2 includes: Split the binocular vision images into multiple channels according to pixel arrangement to obtain preprocessed images; Perform multi-level operations on the obtained preprocessed images to decompose the preprocessed images into images with different frequency bands; Images with different frequency bands are reconstructed, and multi-level inverse operations are performed using inverse discrete wavelet transform (IDWT) to restore the reconstructed image to the size of the original image and output it, thus obtaining the processed binocular vision image.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the 3D image acquisition and processing method based on binocular vision as described in claims 1 to 4.