Parallax optimization method and multi-view camera

By integrating a monocular camera with dual-camera systems to correct and enhance depth estimation, the system addresses structural degradation and light sensitivity issues while maintaining efficient computation.

CN120318288APending Publication Date: 2025-07-15元橡科技(北京)有限公司 +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510360983.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the case of light changes and camera structure degradation, the accuracy of parallax calculation is affected. The calculation volume and cost of triple-eye cameras are large, which affects their popularity.

Method used

A monocular camera module is added to the binocular camera. Through the overlapping area of the field of view and the deep learning network, it is determined whether the parallax image meets the preset conditions and optimizes it to generate an optimized parallax image.

Benefits of technology

In the case of improving part of the computing power, the reliability of the parallax map is improved, the impact of light changes on the parallax map is reduced, and the impact of camera structure degradation is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318288A_ABST
    Figure CN120318288A_ABST
Patent Text Reader

Abstract

The invention discloses a parallax optimization method and a multi-view camera, and the method comprises the steps: judging whether a binocular parallax image accords with a preset condition, and if yes, optimizing the binocular parallax image based on a monocular image obtained by a monocular camera module, and generating an optimized parallax image, judging whether the binocular parallax image meets the preset condition or not at least comprises one of the following judging conditions: judging whether the parallax sparsity in the binocular parallax image is lower than a sparse threshold value or not, and judging whether the binocular parallax image is distorted or not. Through the technical scheme in the invention, under the condition of improving part of computing power, the reliability of outputting the disparity map is improved, and the influence of illumination variation on the disparity map is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer vision. Specifically, it relates to a parallax optimization method and a multi-camera. Background Art

[0002] With the continuous development of computer vision technology, pure vision positioning technology is gradually being applied to vehicle autonomous driving, mobile robots, and other related fields.

[0003] For existing binocular cameras, the technology is relatively mature. Usually, two camera modules are fixedly installed on a crossbeam module, and then the binocular parallax algorithm is used to calculate the parallax map corresponding to the reference image (left-eye or right-eye image), and then combined with the binocular camera parameters to obtain the spatial information of each target within the field of view.

[0004] However, during the use of binocular cameras, especially for the crossbeam module, there will inevitably be a phenomenon of camera structure degradation. At this time, it is necessary to cooperate with the binocular parallax correction algorithm to correct / optimize the calculated parallax map. In addition, binocular cameras usually use the texture in the image for parallax matching calculation, and this texture is relatively easily affected by changes in illumination. Therefore, the influence of illumination changes on the accuracy of the parallax calculation results is relatively large.

[0005] For existing trinocular cameras, after obtaining images within the front field of view through three independently operating camera modules, a deep learning network is used to fuse and predict the depth / parallax of the three obtained images to provide the parallax / distance / size information of each object in the scene. Since the installation positions of the three camera modules are different, the influence of illumination changes on parallax prediction can be reduced, and since they are three independent camera modules, the influence of camera structure degradation on depth / parallax prediction can be alleviated to a certain extent.

[0006] However, the existing trinocular camera technology is not very mature. Usually, it is an improvement based on the binocular parallax calculation framework, and it has problems such as a large amount of calculation, high requirements for the data processing ability of the chip, and high production and manufacturing costs, which affect the popularization of trinocular camera technology. Summary of the Invention

[0007] The purpose of this application is to optimize the parallax map obtained by a binocular camera by adding an additional monocular camera module on the basis of the binocular camera, improve the reliability of the output parallax map with only a partial increase in computing power, and reduce the influence of illumination changes on the parallax map.

[0008] The technical solution of the first aspect of this application is: to provide a parallax optimization method, which is applicable to a multi-camera. The multi-camera includes at least one binocular camera module and at least one monocular camera module. There is a field of view overlapping area between the binocular camera module and the monocular camera module. The binocular camera module is used to generate a binocular parallax image based on the acquired left-eye image and right-eye image. The parallax optimization method includes: judging whether the binocular parallax image meets a preset condition. If so, based on the monocular image acquired by the monocular camera module, the binocular parallax image is optimized to generate an optimized parallax image. Among them, judging whether the binocular parallax image meets the preset condition includes at least one of the following judgment conditions: judging whether the parallax sparsity in the binocular parallax image is lower than the sparsity threshold, and judging whether the binocular parallax image is distorted.

[0009] In any of the above technical solutions, further, judging whether the binocular parallax image is distorted specifically includes: based on the internal parameters of the monocular camera and the binocular parallax image, calculating the first image position of any clustering target in the monocular image; performing target recognition on the monocular image and calculating the second image position of any target recognition result in the monocular image; matching any clustering target with any target recognition result, and based on the matching result, calculating the deviation between the first image position and the second image position, and judging whether the binocular parallax image is distorted based on the deviation and the position deviation threshold.

[0010] In any of the above technical solutions, further, calculating the second image position of any target recognition result in the monocular image specifically includes: extracting the key points of any target recognition result in the monocular image, where the key points include at least any one of feature points, corner points, and edge points; based on the key points, calculating the positioning point corresponding to the target recognition result, and denoting the positioning point as the second image position, where the positioning point is the center point or the centroid.

[0011] In any of the above technical solutions, further, judging whether the binocular parallax image is distorted specifically further includes: extracting the first key points of any target recognition result in the monocular image, where the first key points include at least any one of feature points, corner points, and edge points; extracting the second key points of any clustering target in the left-eye image and / or the right-eye image, where the second key points are of the same type as the first key points; based on a preset intercept frame, intercepting the images around the first key points and the second key points, and respectively calculating the first local image feature and the second local image feature; calculating the feature difference between the first local image feature and the second local image feature, and judging whether the binocular parallax image is distorted based on the size of the feature difference and the feature deviation threshold.

[0012] In any of the above technical solutions, further, based on the monocular image obtained by the monocular camera module, the binocular disparity image is optimized to generate an optimized disparity image, which specifically includes: based on the monocular image and the reference image, a disparity weighted image is generated by means of disparity cost matching, where the reference image is one of the left-eye image and the right-eye image; based on the disparity weighted image, the binocular disparity image is weighted by means of weighted calculation to generate an optimized disparity image, where the weight value in the weighted calculation process is determined by the difference between the feature difference and the feature deviation threshold.

[0013] In any of the above technical solutions, further, the disparity optimization method further includes: calibrating the binocular camera based on the optimized disparity image.

[0014] In any of the above technical solutions, further, when it is determined that the disparity sparsity in the binocular disparity image is lower than the sparsity threshold, based on the monocular image obtained by the monocular camera module, the binocular disparity image is optimized to generate an optimized disparity image, which specifically includes: performing object recognition on the reference image and the monocular image, and performing object matching, and marking the successfully matched object as the reference object, and marking the object that fails to be matched in the monocular image as the object to be mapped; based on the image information of the reference object and the object to be mapped in the monocular image and the disparity of the reference object, the disparity of the object to be mapped is calculated by means of deep learning network prediction; based on the calculated disparity of the object to be mapped, the binocular disparity image is optimized to generate an optimized disparity image, where the reference image is any one of the left-eye image and the right-eye image, and the image information includes at least the object type information, image coordinate information, and image size information.

[0015] The technical solution of the second aspect of the present application is: a multi-camera is provided, and the multi-camera includes: at least one binocular camera module, and the binocular camera module is used to generate a binocular disparity image according to the obtained left-eye image and right-eye image; at least one monocular camera module, where there is a field-of-view overlapping area between the binocular camera module and the monocular camera module; a disparity optimization unit, and the disparity optimization unit is configured to determine whether the binocular disparity image meets a preset condition, and if so, based on the monocular image obtained by the monocular camera module, the binocular disparity image is optimized to generate an optimized disparity image, where determining whether the binocular disparity image meets the preset condition includes at least one of the following judgment conditions: determining whether the disparity sparsity in the binocular disparity image is lower than the sparsity threshold, and determining whether the binocular disparity image is distorted.

[0016] In any of the above technical solutions, further, to determine whether the binocular disparity image is distorted, it specifically includes: based on the internal parameters of the monocular camera and the binocular disparity image, calculating the first image position of any clustering target in the binocular disparity image in the monocular image; performing target recognition on the monocular image, and calculating the second image position of any target recognition result in the monocular image; matching any clustering target with any target recognition result, and based on the matching result, calculating the deviation between the first image position and the second image position, and determining whether the binocular disparity image is distorted based on the deviation and the position deviation threshold.

[0017] In any of the above technical solutions, further, when it is determined that the disparity sparsity in the binocular disparity image is lower than the sparsity threshold, based on the monocular image obtained by the monocular camera module, optimizing the binocular disparity image to generate an optimized disparity image, which specifically includes: performing target recognition on the reference image and the monocular image, and performing target matching, marking the successfully matched target as the reference target, and marking the target that is not successfully matched in the monocular image as the target to be mapped; based on the image information of the reference target and the target to be mapped in the monocular image and the disparity of the reference target, calculating the disparity of the target to be mapped by means of prediction by a deep learning network; based on the calculated disparity of the target to be mapped, optimizing the binocular disparity image to generate an optimized disparity image, where the reference image is any one of the left-eye image and the right-eye image, and the image information at least includes the type information, image coordinate information, and image size information of the target.

[0018] The beneficial effects of this application are:

[0019] In the technical solution of this application, on the basis of a binocular camera, by adding an additional monocular camera module, the disparity map obtained by the binocular camera is corrected and optimized with only a partial increase in computing power, improving the reliability of the output disparity map and helping to reduce the influence of light changes on the disparity map.

[0020] In this application, considering that if only the position deviation of the target in the image is used as the basis for judging whether the binocular disparity image is distorted, there may be certain errors. Therefore, to ensure the reliability of the judgment, the image information of the image is introduced as the judgment basis, calculating the image features at the corresponding positions in the two images through a deep learning network, and then judging whether the binocular disparity image is distorted based on the difference between the features to improve the reliability of the output disparity map. Description of the Drawings

[0021] The above and / or additional advantages of this application will become obvious and easy to understand when combined with the description of the embodiments with the following drawings, where:

[0022] Figure 1 is a schematic flowchart of a disparity optimization method according to an embodiment of this application;

[0023] Figure 2 It is a schematic diagram of a multi-camera according to an embodiment of the present application. Specific embodiments

[0024] In order to be able to more clearly understand the above objects, features, and advantages of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0025] In the following description, many specific details are set forth in order to fully understand the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below.

[0026] Those skilled in the art can understand that for a binocular camera, usually two cameras (camera modules) are placed on a crossbeam module. When the baseline distance is fixed, it mimics the human eye and calculates the three-dimensional coordinates and distances of each target within the field of view based on the parallax principle.

[0027] For a trinocular camera, it usually adds a camera on the basis of the binocular camera structure, and then fuses the data of the three cameras through a deep learning network, and obtains the corresponding parallax map through feature matching to complete ranging.

[0028] Although both are based on the triangulation ranging principle to achieve parallax calculation, obviously, the trinocular camera has a large amount of calculation and high requirements for the data processing ability of the chip. However, in an environment with complex lighting, due to the different viewing angles of the three cameras, the influence of lighting changes on parallax calculation can be reduced.

[0029] Therefore, on the basis of the binocular camera, by adding an additional monocular camera module, the parallax map obtained by the binocular camera can be optimized. Without only increasing part of the computing power, the reliability of the output parallax map can be improved, and the influence of lighting changes on the parallax map can be reduced.

[0030] Embodiment 1:

[0031] As Figure 1 shown, this embodiment provides a parallax optimization method, which is applicable to a multi-camera. The multi-camera includes at least one binocular camera module and at least one monocular camera module. There is a field of view overlap area between the binocular camera module and the monocular camera module. The binocular camera module is used to generate a binocular parallax image according to the acquired left-eye image and right-eye image. The parallax optimization method includes:

[0032] Determine whether the binocular disparity image meets the preset conditions.

[0033] If so, optimize the binocular disparity image based on the monocular image obtained by the monocular camera module, generate an optimized disparity image and output the optimized disparity image; if not, output the binocular disparity image.

[0034] Among them, determining whether the binocular disparity image meets the preset conditions includes at least one of the following judgment conditions:

[0035] Determine whether the disparity sparsity in the binocular disparity image is lower than the sparsity threshold;

[0036] Determine whether the illumination is abnormal, such as a sudden increase or decrease in illumination;

[0037] Determine whether the binocular disparity image is distorted;

[0038] Determine whether the binocular disparity image needs to perform self-check, where the self-check conditions can be set as needed, such as when the binocular camera module runs continuously for 5000 hours, is subjected to severe collision / shaking, etc.

[0039] Now, it is illustrated in the form of assembling a binocular camera module and a monocular camera module into a three-eye camera.

[0040] In this embodiment, there is a field of view overlapping area between the binocular camera module and the monocular camera module, that is, for any target, it can correspond to the monocular image captured by the monocular camera module in the reference image and the binocular disparity image output by the binocular camera module.

[0041] In this embodiment, the left-eye camera module is selected as the reference module, and the left-eye image it captures is the reference image, and the reference image corresponds to the calculated binocular disparity image. It should be noted that the right-eye camera module can also be selected as the reference module.

[0042] The above monocular camera module can be installed on the left or right side of the binocular camera module, and can be a monocular camera module with the same lens parameters as those in the binocular camera module, or a conventional camera module with different lens parameters, or a wide-angle camera module.

[0043] In this embodiment, both the binocular camera module and the monocular camera module are calibrated camera modules, and their corresponding internal and external parameters (focal length f, baseline distance b, RT matrix, etc.), the optical center distance between the monocular camera module and the reference module, and other parameters used to calculate disparity / distance / three-dimensional coordinates are all known.

[0044] The multi-eye camera in this embodiment can be used in vehicles with autonomous driving functions, or can also be applied to indoor / outdoor low-speed moving robots (such as food delivery robots, park tour cars, lawn mowers, etc.).

[0045] In this embodiment, the detection distance of the multi-camera can be set according to actual requirements, and the specific distance is not limited. For example, for a low-speed mobile robot, its detection distance can be set to 10 meters or 15 meters.

[0046] Under normal circumstances, the multi-camera operates in the working mode of a binocular camera. Based on the principle of parallax calculation, according to the acquired left and right eye images, the corresponding parallax map is calculated, and then information such as the position and size of each target within the field of view is obtained.

[0047] However, during use, on the one hand, when using a binocular camera, especially the crossbeam module, it will be affected by temperature, vibration, etc., and there will be inevitable degradation (for example, if the crossbeam module is distorted due to uneven heating, the baseline b will deviate, affecting the calculation of parallax); on the other hand, when the vehicle / robot is moving, it will be affected by the light in the environment, such as the reflection of sunlight / lamplight. These factors will all affect the accuracy of the binocular camera in calculating the parallax map. Therefore, when it is determined that the binocular parallax image has a large error or parallax is missing, the extra monocular camera module can be used to optimize the binocular parallax image.

[0048] For the case of camera degradation, the existing method is usually to use the binocular camera in the current state to capture images of a standard object (such as a target) at a fixed position, calculate the corresponding parallax map, and then compare it with the standard parallax map obtained during the calibration process to obtain the parallax offset Δd, and perform parallax compensation (calibration) on the degraded binocular camera.

[0049] In this embodiment, since there is an extra monocular camera module, without considering the aging of camera components and loosening of the structure, it can be considered that there is no degradation problem, that is, it can provide stable and accurate images.

[0050] Therefore, when it is determined that the binocular parallax image is distorted or parallax self-check is required, the monocular image captured by the monocular camera module can be used to perform parallax / depth prediction based on the deep learning network, predict the distance between each obstacle within the field of view and the multi-camera or the vehicle, and then obtain the parallax corresponding to the binocular camera module. Since the shooting conditions and positions of this monocular image are the same as those of the left and right eye images, this parallax can be used as the true value / check value to obtain the required parallax offset Δd for optimizing the binocular parallax map.

[0051] It should be noted that the parallax offset Δd will be saved for subsequent calculation of the binocular parallax image.

[0052] Regarding the influence of light, those skilled in the art can understand that for a moving object, within a certain period of time (such as 1 hour), such problems only exist within a specific area (such as within 5 meters). When it leaves this area, the influence of light will be eliminated.

[0053] In this embodiment, since there is an additional monocular camera module, the method of temporarily increasing the GPU occupancy rate can be adopted. Using the monocular image and the left and right monocular images, the framework for calculating the disparity image by the three-eye camera is called to calculate the three-eye disparity image at the current moment, optimize the binocular disparity map or directly use the three-eye disparity image as the output. After the influence of light is eliminated, the algorithm for calculating the disparity by the binocular camera is used again to release the GPU computing power.

[0054] It should be noted that regarding the influence of light, since the local images in the left and right monocular images cannot match the corresponding pixel points, it will lead to the lack of disparity, that is, the disparity value in this area is 0. Therefore, in this embodiment, the concept of disparity sparsity is introduced, and the ratio between the number of effective disparity pixels and the total number of pixels in the image is calculated. When this ratio is small, it is considered that there is a lack of disparity, and when it is greater than a certain threshold, it is considered that there is no lack of disparity.

[0055] It should be noted that when there are few targets within the field of view, it may be misjudged as a lack of disparity. At this time, the above method can still be used to calculate the three-eye disparity image. Without considering the calculation error, it can be understood that this three-eye disparity image should be similar to the binocular disparity image. Therefore, the above binocular disparity image does not need to be optimized. At this time, it can be set that after a period of time (such as 5 minutes), the above judgment is made again to avoid occupying the GPU.

[0056] It should be noted that the current light change can also be determined by setting a light sensor, and the specific implementation process will not be elaborated here.

[0057] Through the disparity optimization method in this embodiment, on the basis of the binocular camera, by adding an additional monocular camera module, the optimization of the disparity map obtained by the binocular camera is realized with only a partial increase in computing power, the reliability of the output disparity map is improved, and the influence of light change on the disparity map is reduced.

[0058] In any of the above embodiments, further, in order to timely detect the degradation of the binocular camera and ensure the reliability of the binocular disparity image, this embodiment also proposes a method for judging whether the binocular disparity image is distorted, and this method specifically includes:

[0059] Step 101: Based on the internal parameters of the monocular camera and the binocular disparity image, calculate the first image position of any clustering target in the binocular disparity image in the monocular image;

[0060] Specifically, the parallax calculation formula is as follows:

[0061]

[0062] In the formula, Z is the distance between the target and the binocular camera, b is the baseline (distance), f is the focal length, and d is the parallax.

[0063] For any clustered target, the parallax within its area is the same value, and its geometric center, centroid, or the center of its maximum circumscribed rectangle can be obtained. These positions can be used as the basis for calculating the position in the first image.

[0064] Based on the corresponding parallax, in the case where other parameters such as the internal parameters of the monocular camera and the optical center distance between the monocular camera module and the reference module are known, the position of the above-mentioned position in the monocular image can be calculated by means of inverse parallax operation, denoted as the first image position, and this first image position is a theoretical value.

[0065] Step 102: Perform target recognition on the monocular image and calculate the second image position of any target recognition result in the monocular image; among them, target recognition can be carried out using a deep learning network to obtain the second image position of the target in the image (such as the upper left corner coordinates of the target box). In addition, the type of the target, the length and width of the target box, etc. can also be obtained. Among them, the type of the target can be set artificially, such as pedestrians, vehicles, ground, traffic signs, etc.

[0066] It should be noted that for the binocular camera module, it is also necessary to recognize the targets in the left-eye image, and the recognized targets correspond one by one to the clustered targets.

[0067] Step 103: Match any clustered target with any target recognition result, and based on the matching result, calculate the deviation between the first image position and the second image position, and judge whether the binocular parallax image is distorted based on the deviation and the position deviation threshold.

[0068] Specifically, when performing the matching, it is necessary to unify the first and second image positions, such as selecting the center point coordinates. The matching can be carried out according to the above-mentioned image positions and the types of the targets. For example: when matching, it should be ensured that the types of the targets are the same and the difference in the image positions is the smallest, that is, there can be a deviation between the first and second image positions when the types are the same.

[0069] After the matching is completed, calculate the deviation between any pair of the first image position and the second image position after matching. It should be noted that the above-mentioned multiple deviations can be subjected to operations such as mean and variance. Then compare the operation result with the set position deviation threshold. If the difference between the two is large, it is considered that the binocular parallax image is distorted; otherwise, it is considered that the binocular parallax image is normal and does not need to be optimized.

[0070] In any of the above embodiments, further, considering that there will inevitably be a deviation between the target box selected by the deep learning network during target recognition and the actual size of the target in the monocular image. Therefore, in order to improve the accuracy of the second image position calculation, calculating the second image position of any target recognition result in the monocular image specifically includes:

[0071] Step 1021: Extract the key points of any target recognition result in the monocular image, where the key points at least include any one of feature points, corner points, and edge points;

[0072] Step 1022: Based on the key points, calculate the positioning point corresponding to the target recognition result, and record the positioning point as the second image position, where the positioning point is the center point or the centroid.

[0073] In any of the above embodiments, further, considering that there may be a certain error if only the position deviation of the target in the image is used as the judgment basis for whether the binocular disparity image is distorted. Therefore, in order to ensure the reliability of the judgment, the image feature information of the image is introduced as the judgment basis. The image features at the corresponding positions in the two images are calculated through the deep learning network, and then based on the difference between the features, it is judged whether the binocular disparity image is distorted. Specifically, it further includes:

[0074] Step 201: Extract the first key points of any target recognition result in the monocular image, where the first key points at least include any one of feature points, corner points, and edge points;

[0075] Step 202: Extract the second key points of any clustering target in the left-eye image and / or the right-eye image, where the types of the second key points are the same as those of the first key points;

[0076] Step 203: Based on the preset cropping frame, crop the images around the first key points and the second key points, and calculate the first image local feature and the second image local feature respectively; the size of the preset cropping frame can be set to 5×5 pixel sizes, that is, a piece of image around the key points is cropped, and convolutional calculations are performed using the deep learning network, and the corresponding image features are obtained through convolutional layers, pooling layers, etc.

[0077] It should be noted that when the number of the first and second key points is not equal, based on the group of key points with fewer numbers, in a traversing manner, the key points in the other group with the smallest (Euclidean) distance from it are selected and retained, and the rest are discarded, and then the feature calculations are performed on these retained key points.

[0078] In this embodiment, the above key points can be joint points on a pedestrian, such as shoulders, feet, hands, heads, etc., or can be corner points and edge points of a license plate or a traffic sign on a vehicle.

[0079] Step 204: Calculate the feature difference between the local features of the first image and the local features of the second image, and determine whether the binocular disparity image is distorted based on the magnitude relationship between the feature difference and the feature deviation threshold.

[0080] Specifically, this feature difference is used as the basis for measuring whether the first key point and the second key point are the same / similar. That is, when the feature difference is less than or equal to the feature deviation threshold, it is considered that the features of the first and second key points in the image are the same / similar. Furthermore, it is considered that the target recognized in the monocular image and the target in the clustering target in the left-eye image and / or the right-eye image are the same target. Therefore, it is determined that the binocular disparity image is not distorted.

[0081] Correspondingly, if the feature difference is greater than the feature deviation threshold, it is considered that there is a deviation in the positions of the above two targets in the image, and it is determined that the binocular disparity image is distorted.

[0082] In any of the above embodiments, further, based on the monocular image obtained by the monocular camera module, the binocular disparity image is optimized to generate an optimized disparity image, which specifically includes:

[0083] Step 301: Based on the monocular image and the reference image, a disparity weighted image is generated by means of disparity cost matching, where the reference image is one of the left-eye image and the right-eye image;

[0084] Step 302: Based on the disparity weighted image, the binocular disparity image is weighted by means of weighted calculation to generate an optimized disparity image, where the weight value in the weighted calculation process is determined by the magnitude of the difference between the feature difference and the feature deviation threshold.

[0085] Specifically, when optimizing the binocular disparity image, in addition to using the framework of a trinocular camera to calculate the disparity image and using the monocular image and the left and right eye images to calculate the trinocular disparity, the binocular disparity calculation method can still be used. A virtual binocular camera is formed by using the monocular camera module and the left-eye camera module (the reference module), and then the disparity weighted image is obtained. Furthermore, the binocular disparity image is weighted by means of weighting to generate an optimized disparity image.

[0086] In this embodiment, the above weights can be implemented using a piecewise function. For ease of explanation, the feature deviation threshold can be set to 0.5. When the feature difference is between 0.5 and 0.6, the weight of the binocular disparity image can be set to be greater than the weight of the disparity weighted image. For example, the former is 0.7 and the latter is 0.3. When the feature difference is between 0.6 and 0.7, the weight of the binocular disparity image can be set to be less than the weight of the disparity weighted image. For example, the former is 0.3 and the latter is 0.7.

[0087] In any of the above embodiments, further, the disparity optimization method further includes: based on the optimized disparity image, calibrating the binocular camera, that is, calculating the disparity offset △d. Accordingly, the disparity calculation formula should be adjusted to:

[0088]

[0089] In any of the above embodiments, further, when it is determined that the disparity sparsity in the binocular disparity image is lower than the sparsity threshold,

[0090] optimize the binocular disparity image based on the monocular image obtained by the monocular camera module to generate an optimized disparity image, specifically including:

[0091] Step 311: Perform object recognition on the reference image and the monocular image, and perform object matching. Mark the successfully matched object as the reference object, and mark the object that fails to match successfully in the monocular image as the object to be mapped;

[0092] Step 312: Based on the image information of the reference object and the object to be mapped in the monocular image and the disparity of the reference object, calculate the disparity of the object to be mapped by means of prediction by a deep learning network;

[0093] Step 313: Optimize the binocular disparity image based on the calculated disparity of the object to be mapped to generate an optimized disparity image, where the reference image is any one of the left-eye image and the right-eye image, and the image information includes at least the type information, image coordinate information, and image size information of the object.

[0094] Specifically, when it is determined that the disparity sparsity in the binocular disparity image is lower than the sparsity threshold, it is considered that the binocular camera module is affected by external light at this time. For example, when light reflection occurs on the surface of an object, in the corresponding binocular disparity image, the disparity of the area where the object is located is 0. Therefore, it is necessary to fill in / supplement the disparity here.

[0095] At a certain moment, for any object, its position and size in space are fixed, and the spatial relationship with other objects in the surrounding environment is also fixed. Therefore, a deep learning network can be used to learn this characteristic.

[0096] When a certain target exists in a monocular image but not in a binocular disparity image, it can be considered as a disparity missing. Therefore, the above deep learning network can be used to predict information such as the position, size, and disparity of the target in the monocular image, and then map the obtained disparity of the target to be mapped to the binocular disparity image for optimized filling / completion to improve the reliability of the output disparity map.

[0097] Embodiment 2:

[0098] As Figure 2 shown, this embodiment provides a multi-camera, which includes:

[0099] At least one binocular camera module, which is used to generate a binocular disparity image according to the acquired left-eye image and right-eye image;

[0100] At least one monocular camera module, where there is a field-of-view overlapping area between the binocular camera module and the monocular camera module;

[0101] A disparity optimization unit, which is configured to determine whether the binocular disparity image meets a preset condition. If so, based on the monocular image acquired by the monocular camera module, optimize the binocular disparity image to generate an optimized disparity image. Among them, determining whether the binocular disparity image meets the preset condition includes at least one of the following judgment conditions:

[0102] Determining whether the disparity sparsity in the binocular disparity image is lower than a sparsity threshold, and determining whether the binocular disparity image is distorted.

[0103] In any of the above embodiments, further, determining whether the binocular disparity image is distorted specifically includes: calculating the first image position of any clustering target in the monocular image based on the internal parameters of the monocular camera and the binocular disparity image; performing target recognition on the monocular image and calculating the second image position of any target recognition result in the monocular image; matching any clustering target with any target recognition result, and based on the matching result, calculating the deviation between the first image position and the second image position, and determining whether the binocular disparity image is distorted based on the deviation and the position deviation threshold.

[0104] In any of the above embodiments, further, calculating the second image position of any target recognition result in the monocular image specifically includes: extracting the key points of any target recognition result in the monocular image, where the key points include at least any one of feature points, corner points, and edge points; based on the key points, calculating the positioning point corresponding to the target recognition result, and denoting the positioning point as the second image position, where the positioning point is the center point or the centroid.

[0105] In any of the above embodiments, further, to determine whether the binocular disparity image is distorted, it specifically further includes: extracting a first key point of any target recognition result in the monocular image, where the first key point includes at least any one of feature points, corner points, and edge points; extracting a second key point of any clustering target in the left-eye image and / or the right-eye image, where the type of the second key point is the same as that of the first key point; based on a preset cropping frame, cropping the images around the first key point and the second key point, and respectively calculating the local features of the first image and the local features of the second image; calculating the feature difference between the local features of the first image and the local features of the second image, and determining whether the binocular disparity image is distorted based on the magnitude relationship between the feature difference and the feature deviation threshold.

[0106] In any of the above embodiments, further, based on the monocular image obtained by the monocular camera module, optimizing the binocular disparity image to generate an optimized disparity image, which specifically includes: based on the monocular image and a reference image, generating a disparity weighted image through disparity cost matching, where the reference image is one of the left-eye image and the right-eye image; based on the disparity weighted image, weighting the binocular disparity image through weighted calculation to generate an optimized disparity image, where the weight value in the weighted calculation process is determined by the magnitude of the difference between the feature difference and the feature deviation threshold.

[0107] In any of the above embodiments, further, the disparity optimization method further includes: calibrating the binocular camera based on the optimized disparity image.

[0108] In any of the above embodiments, further, when it is determined that the disparity sparsity in the binocular disparity image is lower than the sparsity threshold, based on the monocular image obtained by the monocular camera module, optimizing the binocular disparity image to generate an optimized disparity image, which specifically includes: performing target recognition on the reference image and the monocular image, and performing target matching, marking the successfully matched target as the reference target, and marking the target that is not successfully matched in the monocular image as the target to be mapped; calculating the disparity of the target to be mapped through deep learning network prediction based on the image information of the reference target and the target to be mapped in the monocular image and the disparity of the reference target; based on the calculated disparity of the target to be mapped, optimizing the binocular disparity image to generate an optimized disparity image, where the reference image is any one of the left-eye image and the right-eye image, and the image information includes at least the type information, image coordinate information, and image size information of the target.

[0109] So far, the embodiments of the present application have been described in detail. To avoid obscuring the concept of the present application, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0110] Although some specific embodiments of the present application have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the present application.

[0111] The steps in the present application can be adjusted, combined, and deleted according to actual needs.

[0112] Although the present application has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and are not intended to limit the application of the present application. The protection scope of the present application is defined by the appended claims and may include various modifications, adaptations, and equivalent solutions made to the invention without departing from the protection scope and spirit of the present application.

Claims

1. A parallax optimization method, which is applicable to a multi-camera, the multi-camera includes at least one binocular camera module and at least one monocular camera module, there is a field of view overlapping area between the binocular camera module and the monocular camera module, the binocular camera module is used to generate a binocular parallax image according to the acquired left-eye image and right-eye image, and is characterized in that, The parallax optimization method includes: judging whether the binocular parallax image meets a preset condition; if so, optimizing the binocular parallax image based on the monocular image obtained by the monocular camera module to generate an optimized parallax image; wherein, judging whether the binocular parallax image meets the preset condition includes at least one of the following judgment conditions: judging whether the parallax sparsity in the binocular parallax image is lower than a sparsity threshold; judging whether the binocular parallax image is distorted.

2. The parallax optimization method according to claim 1, characterized in that, The judging whether the binocular parallax image is distorted specifically includes: calculating a first image position of any clustering target in the monocular image in the binocular parallax image based on the internal parameters of the monocular camera and the binocular parallax image; performing target recognition on the monocular image and calculating a second image position of any target recognition result in the monocular image; matching any clustering target with any target recognition result, and calculating a deviation between the first image position and the second image position based on the matching result, and judging whether the binocular parallax image is distorted based on the deviation and a position deviation threshold.

3. The parallax optimization method according to claim 2, wherein The calculating the second image position of any target recognition result in the monocular image specifically includes: extracting key points of any target recognition result in the monocular image, where the key points include at least any one of feature points, corner points, and edge points; calculating a positioning point corresponding to the target recognition result based on the key points, and denoting the positioning point as the second image position, where the positioning point is a center point or a centroid.

4. The parallax optimization method according to claim 2, wherein The judging whether the binocular parallax image is distorted specifically further includes: extracting first key points of any target recognition result in the monocular image, where the first key points include at least any one of feature points, corner points, and edge points; extracting second key points of any clustering target in the left-eye image and / or the right-eye image, where the types of the second key points are the same as those of the first key points; intercepting the images around the first key points and the second key points based on a preset intercepting frame, and respectively calculating a first local image feature and a second local image feature; calculating a feature difference between the first local image feature and the second local image feature, and judging whether the binocular parallax image is distorted based on the magnitude relationship between the feature difference and a feature deviation threshold.

5. The parallax optimization method according to any one of claims 2 to 4, characterized in that, The optimizing the binocular parallax image based on the monocular image obtained by the monocular camera module to generate an optimized parallax image specifically includes: generating a disparity weighted image based on the monocular image and a reference image by means of disparity cost matching, where the reference image is one of the left-eye image and the right-eye image; weighting the binocular parallax image by means of weighted calculation based on the disparity weighted image to generate the optimized parallax image; wherein, the weight value in the weighted calculation process is determined by the magnitude of the difference between the feature difference and the feature deviation threshold.

6. The parallax optimization method according to claim 5, wherein The parallax optimization method further includes: correcting the binocular camera based on the optimized parallax image.

7. The parallax optimization method according to claim 1, wherein When it is determined that the disparity sparsity in the binocular disparity image is lower than the sparse threshold, optimizing the binocular disparity image based on the monocular image obtained by the monocular camera module to generate an optimized disparity image, specifically including: performing object recognition on the reference image and the monocular image, and performing object matching, denoting the successfully matched object as the reference object, and denoting the object that fails to be matched in the monocular image as the object to be mapped; calculating the disparity of the object to be mapped by means of prediction by a deep learning network based on the image information of the reference object and the object to be mapped in the monocular image and the disparity of the reference object; optimizing the binocular disparity image based on the calculated disparity of the object to be mapped to generate the optimized disparity image, wherein the reference image is any one of the left-eye image and the right-eye image, and the image information at least includes the type information, image coordinate information, and image size information of the object.

8. A multi-camera, characterized in that, The multi-camera includes: at least one binocular camera module for generating a binocular disparity image according to the obtained left-eye image and right-eye image; at least one monocular camera module, wherein there is a field-of-view overlapping area between the binocular camera module and the monocular camera module; a disparity optimization unit configured to determine whether the binocular disparity image meets a preset condition, if so, optimizing the binocular disparity image based on the monocular image obtained by the monocular camera module to generate an optimized disparity image, wherein determining whether the binocular disparity image meets the preset condition at least includes one of the following determination conditions: determining whether the disparity sparsity in the binocular disparity image is lower than the sparse threshold, determining whether the binocular disparity image is distorted.

9. The multi-view camera according to claim 8, characterized in that, The determination of whether the binocular disparity image is distorted specifically includes: calculating the first image position of any clustered object in the binocular disparity image in the monocular image based on the internal parameters of the monocular camera and the binocular disparity image; performing object recognition on the monocular image and calculating the second image position of any object recognition result in the monocular image; matching any clustered object with any object recognition result, and calculating the deviation between the first image position and the second image position based on the matching result, and determining whether the binocular disparity image is distorted based on the deviation and the position deviation threshold.

10. The multi-camera according to claim 8, characterized in that, When it is determined that the disparity sparsity in the binocular disparity image is lower than the sparse threshold, optimizing the binocular disparity image based on the monocular image obtained by the monocular camera module to generate an optimized disparity image, specifically including: performing object recognition on the reference image and the monocular image, and performing object matching, denoting the successfully matched object as the reference object, and denoting the object that fails to be matched in the monocular image as the object to be mapped; calculating the disparity of the object to be mapped by means of prediction by a deep learning network based on the image information of the reference object and the object to be mapped in the monocular image and the disparity of the reference object; Optimize the binocular disparity image based on the calculated disparity of the target to be mapped, and generate the optimized disparity image. Wherein, the reference image is any one of the left-eye image and the right-eye image, and the image information at least includes the type information of the target, the image coordinate information, and the image size information.

Citation Information

Cited By

  • Lattice tower rod piece missing detection method and program product

    CN120932131A