Camera parameter automatic optimization system and method for binocular depth estimation, medium and equipment
By applying an actor-criticist-based deep reinforcement learning framework in the field of binocular depth estimation, automatically optimizing camera parameters, solving the challenges of camera parameter selection and optimization in the existing technology, and achieving more efficient and more general depth estimation effects.
Patent Information
- Application Number
- CN202510027768.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In the field of binocular depth estimation, there are many challenges in the selection and optimization of camera parameters, including the difficulty in obtaining optimal results in all cases, and the application of deep learning models has limitations and is difficult to apply in different scenarios and tasks.
Using a deep reinforcement learning framework based on actor-criticist, through the Auto-CAM learning module, backpropagation learning is performed according to the current camera configuration, state vector and depth error, and better camera parameter configuration is output. The system includes truth acquisition module, Auto-CAM learning module, Auto-CAM deployment module, camera calibration module, image acquisition module, depth map prediction module and visual rendering module.
It realizes the optimal camera parameter configuration that automatically outputs the optimal camera parameter configuration in specific scenarios, improves the accuracy and efficiency of binocular depth maps, overcomes the limitations of traditional manual adjustment and experience-based presets, and has a wide range of application prospects.
Smart Images

Figure CN120147432A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a camera parameter automatic optimization system, method, medium, and device for binocular depth estimation. Background Art
[0002] In the field of binocular depth estimation, the selection and optimization of camera parameters have a crucial impact on the accuracy and robustness of depth estimation. In-depth research on camera parameters is a key link to improving the performance of depth estimation algorithms and promoting the development of computer vision applications.
[0003] As one of the basic parameters of a camera, the focal length determines the size of the object projection on the imaging plane, and its optimization is crucial for improving the accuracy of depth estimation. Although researchers explore the optimal focal length setting through experimental adjustment and multi-scene testing, this method is not only time-consuming and laborious but also difficult to ensure obtaining the optimal depth estimation result in all cases. In addition, the adjustment of the focal length may be restricted by camera hardware limitations and cost factors, further limiting its optimization space.
[0004] The baseline is the distance between two cameras in a binocular camera system and is also one of the key factors affecting depth estimation. Different baseline settings directly affect the density and range of the depth map. Researchers explore the possibility of obtaining better depth estimation in a specific scene by adjusting the length of the baseline. An overly long baseline may result in too small a parallax for distant objects, making it difficult to accurately estimate the depth; while an overly short baseline may reduce the depth resolution and affect the depth estimation accuracy of nearby objects. In addition, the adjustment of the baseline may also be restricted by factors such as device volume, weight, and cost, making it difficult to find the optimal baseline configuration in practical applications.
[0005] Multi-camera systems can provide more depth information compared to binocular systems but also face more technical challenges. By optimizing the camera layout and configuration, researchers hope to achieve higher-quality depth estimation. This aspect of research involves aspects such as the spatial layout and angle selection between cameras, aiming to find the optimal configuration plan. Although the rise of deep learning has brought a new perspective to the research of camera parameters, its application in binocular depth estimation still has limitations. Existing deep learning models often rely on large-scale datasets for training, and the acquisition and processing costs of these datasets are high. In addition, the sensitivity and generalization ability of deep learning models to camera parameters still need to be improved, which limits their applicability in different scenarios and tasks. At the same time, the interpretability of deep learning models is poor, making it difficult for researchers to deeply understand the complex relationship between camera parameters and depth estimation.
[0006] Therefore, the selection and optimization of camera parameters in the current binocular depth estimation field still face many challenges and deficiencies. To overcome these shortcomings, more efficient, general, and adaptable optimization strategies and methods need to be explored. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides an automatic camera parameter optimization system and method for binocular depth estimation.
[0008] The technical solution of the present invention is as follows:
[0009] An automatic camera parameter optimization system for binocular depth estimation, characterized by comprising:
[0010] A ground truth acquisition module, used to obtain a depth ground truth map in a simulation environment;
[0011] An Auto-CAM learning module, adopting a deep reinforcement learning framework based on actor-critic, performing backpropagation learning according to the current camera configuration, state vector, and depth error, wherein the state vector is generated by an encoder with shared weights from the left and right view images, the depth error is calculated from the depth map generated by a pre-trained stereo matching model and the depth ground truth map, and the camera configuration includes focal length, baseline distance, and distortion coefficient;
[0012] An Auto-CAM deployment module, used to output the optimal camera parameter configuration according to the trained reinforcement learning agent in a specific scenario;
[0013] A camera calibration module, used to obtain accurate camera internal and external parameters by corner scanning after changing the camera parameters;
[0014] An image acquisition module, used to obtain RGB binocular images;
[0015] A depth map prediction module, outputting a predicted disparity map according to the deployed stereo matching network algorithm, and converting the disparity map into a depth map by using the internal and external parameters obtained by the camera calibration module;
[0016] A visualization rendering module, used to filter and color-process the obtained depth map to present a binocular depth map.
[0017] Furthermore, the Auto-CAM learning module adopts a reward function, including:
[0018] A first reward term, calculating the difference between the depth estimation error under the current camera configuration and the target depth error set manually;
[0019] A second reward term, by analyzing the error estimated by the model; wherein, the analysis model is as follows:
[0020]
[0021] Among them, ΔD is the accuracy of binocular depth, and Δd px is the accuracy of the disparity map in pixels. b and f are the baseline and focal length of the camera respectively, and D gt is the true value of binocular depth, and r hor is the pixel density in the horizontal direction, and w sen is the width of the sensor.
[0022] The third reward item, RoI exclusion penalty, is used to ensure that the pixels of the RoI are always within the field of view FoV of the binocular camera view during the update of the camera configuration.
[0023] The specific process of the Auto-CAM deployment module includes:
[0024] - The binocular camera captures the left view and the right view under the current camera configuration;
[0025] - The left image and the right image are respectively fed into an encoder with shared weights to generate two feature vectors, which are integrated into a state vector through a specific fusion strategy;
[0026] - The reinforcement learning agent takes the state vector as input and generates an action vector;
[0027] - The binocular camera adjusts its own configuration according to the action vector and captures new image data.
[0028] Furthermore, it also includes an embedded system, which includes a deep learning edge computing device and a binocular camera system with a variable baseline, capable of obtaining high-precision RGB binocular images according to a specific scenario, inputting the images into the encoder to obtain a state vector, then the reinforcement learning agent will generate an action vector, and finally change the parameters of the binocular camera according to the output better configuration and perform camera calibration. Subsequently, the embedded system will perform depth map prediction and visualization rendering according to the loaded deep neural network.
[0029] On the other hand, the present invention also provides a method for automatically optimizing the parameters of a binocular depth estimation camera, characterized by including the following steps:
[0030] S1. Obtain a depth ground truth map in a simulation environment;
[0031] S2. Adopt a deep reinforcement learning framework based on actor-critic, perform backpropagation learning according to the current camera configuration, state vector, and depth error, and output a better camera configuration;
[0032] S3. In a specific scenario, use the trained reinforcement learning agent to output the optimal camera parameter configuration;
[0033] S4. After changing the camera parameters, obtain accurate intrinsic and extrinsic camera parameters through corner scanning;
[0034] S5. Obtain RGB binocular images and output a predicted disparity map according to the deployed stereo matching network algorithm;
[0035] S6. Use the intrinsic and extrinsic parameters obtained by the camera calibration module to convert the disparity map into a depth map;
[0036] S7. Filter and color-process the obtained depth map to finally present a binocular depth map with higher quality.
[0037] A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed, the steps of the method are implemented.
[0038] A computer device, including a processor, a memory, and a computer program stored on the memory, characterized in that when the processor executes the computer program, the steps of the method are implemented.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] 1) For the exploration scenario with ground truth assistance, the present invention is superior to all existing technologies in terms of performance and speed; while for the deployment scenario without ground truth assistance, the present invention, as the only existing method, can output better camera parameters based only on the input binocular images to improve the accuracy of the binocular depth map. The present invention is a unified solution for two different scenarios.
[0041] 2) The present invention applies deep reinforcement learning to the automatic optimization of camera parameters, breaking through the limitations of traditional manual adjustment or empirical preset, and can automatically output the optimal camera parameter configuration in a specific scenario, with broad application prospects and market demands. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 : Flowchart of the automatic optimization system for camera parameter configuration
[0043] Figure 2 : Flowchart of the learning module of the Auto-CAM framework
[0044] Figure 3 : Flowchart of the deployment module of the Auto-CAM framework
[0045] Figure 4 : Example diagram of the result of depth optimization in the exploration stage of the Auto-CAM framework
[0046] Figure 5: Depth example diagrams before and after optimization using Auto-CAM for the emulator deployment scenario of the experiment
[0047] Figure 6 : Schematic diagram of the optimization path in the exploration phase of the Auto-CAM framework
[0048] Figure 7 : Depth example diagrams before and after optimization using Auto-CAM for the actual setup deployment scenario of the experiment. Detailed implementation manners
[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, but the protection scope of the present invention should not be limited thereby.
[0050] Construct an analysis model that takes into account the effects of factors such as focal length, baseline, pixel density, and stereo matching model accuracy on binocular depth estimation accuracy. The specific formula is as follows:
[0051]
[0052] Among them, ΔD is the accuracy of binocular depth, and Δd px is the accuracy of the disparity map (in pixels), b and f are the baseline and focal length of the camera respectively, D gt is the true value of binocular depth, r hor is the pixel density in the horizontal direction, and w sen is the width of the sensor.
[0053] The mathematical model in the formula provides the following trends for improving the accuracy of BDE: 1) Camera parameters: The depth error can be reduced by expanding the product of the baseline and focal length (b·f); 2) Camera specifications: Selecting an image sensor with a higher resolution does not reduce the depth error, but a higher horizontal pixel density (r hor / w sen ) will be helpful. 3) Stereo matching algorithm: Optimizing the stereo matching algorithm (i.e., reducing Δd px ) can reduce the depth error. 4) Target distance: The depth error increases with the distance from the target to the binocular vision system.
[0054] However, the above summarized trends are not always applicable in the simulation scenario deployment because the disparity error (Δd px ) is not a fixed value and it varies with the camera parameters (i.e., b and f). In Figure 1 , the top subgraph shows the variation of the disparity error with b and FoV. Since FoV is inversely proportional to f, in Figure 2In the middle sub - figure, we can see that when the FoV is 60° and b is 2 (i.e., the maximum b·f), the best setting (minimum log|ΔD AM|) is adopted. However, the actual depth error obtained from experimental measurements confirms that the optimization of camera configuration is a more complex problem, where the error distribution is different from the mathematical model.
[0055] Therefore, the Auto - CAM framework is proposed, which uses deep reinforcement learning to automatically obtain ideal camera parameters quickly for different scenarios. Auto - CAM consists of three core components: an actor - critic agent, an encoder, and an environment, and is divided into a learning stage and a deployment stage.
[0056] Actor - critic agent: Responsible for generating the action vector of camera parameter configuration and evaluating the quality of this action vector. It is trained and optimized through deep reinforcement learning algorithms (such as the actor - critic algorithm).
[0057] Encoder: Uses a pre - trained Transformer model to process the input binocular images and generate a high - dimensional state vector. This state vector reflects the geometric and texture information in the images and provides input data for the reinforcement learning agent.
[0058] Environment: Provides an input interface for binocular images and an adjustment interface for camera parameters. Receives the action vector (camera parameter configuration) generated by the reinforcement learning agent and returns the corresponding reward signal and a new state vector. The environment part is also responsible for camera calibration and the prediction and visual rendering of depth maps.
[0059] In the learning stage, Auto - CAM optimizes the reinforcement learning agent through continuous trial - and - error and reward feedback. In the deployment stage, the trained agent is used to quickly obtain the ideal camera parameter configuration and apply it to the actual scenario.
[0060] Finally, an embedded system implementing the above framework is built, which includes a deep - learning edge - computing device of Jetson Orin Nano and a binocular camera system with a variable baseline. High - precision rgb binocular images can be obtained according to a specific scenario and input into the encoder to obtain the state vector. Then, the reinforcement learning agent will generate the action vector. Finally, we change the parameters of the binocular camera according to the output of the better configuration and perform camera calibration. Subsequently, the embedded system will perform depth - map prediction and visual rendering according to the embedded deep - neural network.
[0061] Camera parameter optimization in binocular depth estimation plays a crucial role in multiple high-tech fields, especially in two application scenarios with extremely high precision requirements, namely autonomous driving and defect detection. Binocular depth estimation generates a disparity map through a pair of synchronously captured cameras to calculate the depth information of the target scene. The precise optimization of camera parameters (such as intrinsic parameters, extrinsic parameters, distortion coefficients, baseline distance, etc.) directly affects the accuracy and robustness of the depth estimation results. In the field of autonomous driving, vehicles need to perceive the distance and position of obstacles, road structures, and dynamic targets in the three-dimensional environment in real time to complete core tasks such as path planning, obstacle avoidance, and lane keeping. Unoptimized camera parameters may lead to biases in the depth estimation results, posing serious safety hazards. For example, a small error in the baseline distance may cause significant deviations in depth measurement, especially in the detection of distant targets. Through camera parameter optimization, the autonomous driving system can greatly improve the accuracy of depth estimation, ensuring that the vehicle can make reliable decisions in diverse scenarios (such as urban congestion, rural roads, and highways). Especially in complex dynamic environments, such as detecting pedestrians, cyclists, or other vehicles, accurate depth information can significantly enhance the vehicle's reaction speed and safety. In the field of industrial defect detection, binocular depth estimation generates high-resolution three-dimensional point cloud data for identifying tiny defects (such as cracks, dents, wear, etc.) on the surface of workpieces. Optimized camera parameters ensure the stability of disparity calculation and the accuracy of the depth map, thus providing more reliable detection results. In some key manufacturing fields (such as semiconductor, aerospace, and automotive manufacturing), the detection accuracy directly determines the product quality and safety. For the detection of irregular surfaces or highly reflective materials, the optimization of the camera's intrinsic parameters (such as focal length and principal point offset) and extrinsic parameters (such as the baseline distance and relative angle between cameras) is particularly important. In addition, for the harsh conditions in some industrial scenarios (such as uneven lighting and dusty environments), camera parameter optimization combined with advanced calibration algorithms can also effectively improve the anti-interference ability and adaptability of the system. Given the importance of camera parameter optimization in binocular depth estimation in the fields of autonomous driving and industrial defect detection, the present invention is implemented in two edge applications. By integrating multiple modules such as ground truth acquisition, Auto-CAM learning, deployment, camera calibration, image acquisition, depth map prediction, and visualization rendering, the present invention realizes the automatic optimization of camera parameters and the precise calculation of depth estimation. This not only improves the overall performance and efficiency of the system but also provides more reliable technical support and innovative solutions for fields such as autonomous driving and industrial defect detection.
[0062] Such as Figure 1As shown in the figure, it is a flowchart of an automatic optimization system for camera parameter configuration applied to an emulator. It includes a ground truth acquisition module, an Auto-CAM learning module, an Auto-CAM deployment module, a camera calibration module, an image acquisition module, a depth map prediction module, and a visualization rendering module. Among them: The ground truth acquisition module uses the emulator to obtain the depth ground truth map. The Auto-CAM learning module performs backpropagation learning of reinforcement learning based on the current camera configuration, state vector, and depth error. The Auto-CAM deployment module stops learning and outputs a better camera configuration using the current state vector. The camera calibration module needs to obtain accurate internal and external camera parameters through corner scanning after changing the camera parameters. The image acquisition module obtains high-precision rgb binocular images. The depth map prediction module outputs a predicted disparity map according to the deployed stereo matching network algorithm, and then performs the conversion of the depth map using the internal and external camera parameters obtained by the camera calibration module. The visualization rendering module filters and processes the obtained depth map, and finally presents a binocular depth map with higher quality.
[0063] Example 1 forms a closed-loop automatic optimization process. From ground truth acquisition to depth map prediction and then to visualization rendering, each module plays an indispensable role. Through advanced technologies such as reinforcement learning and camera calibration, the system can automatically optimize the camera parameter configuration, thereby improving the accuracy and efficiency of depth estimation. Such an automatic optimization system has broad application prospects and important practical value in high-tech fields such as autonomous driving and industrial defect detection.
[0064] Figure 2 It is a flowchart of the Auto-CAM framework learning module. As shown in the figure, it aims to automatically optimize the parameter configuration of a binocular camera through a reinforcement learning framework. The following are the steps of the iterative learning phase:
[0065] (1) Initialize the camera configuration: The binocular camera captures the left view and the right view under the current camera configuration. For initialization, the baseline and focal length are randomly selected from the value range respectively.
[0066] (2) Image capture and feature extraction: The binocular camera captures the left view and the right view under the current camera configuration. The left image and the right image are fed into an encoder with shared weights, and two feature vectors are output.
[0067] (3) Feature fusion: The feature vectors of the left image and the right image are fused into a state vector through an element-wise addition fusion method.
[0068] (4) Reinforcement learning agent decision: The reinforcement learning agent (RL agent) processes the state vector and generates an action vector, which describes how to adjust the camera parameters to optimize the accuracy of depth estimation.
[0069] (5) Camera configuration update and image re-capture: Configure the stereo camera with the new settings and re-capture the left and right view images and the ground truth depth map.
[0070] (6) Depth map estimation and reward calculation:
[0071] Generate the estimated depth map using a pre-trained stereo matching model (such as HITNet). During the reinforcement learning process, the parameters of the stereo matching model are frozen to ensure the stability and consistency of reward calculation.
[0072] Calculate the reward based on the difference between the estimated depth map and the ground truth depth map. The reward function consists of three terms: depth estimation error reward, model estimation error reward, and RoI exclusion penalty.
[0073] (7) Update the model parameters of the actor and critic networks: Use the encoder to convert the pixels of the left and right view images into state vectors (embeddings). The encoder inherits the encoder part of the masked autoencoder (MAE) pre-trained on the ImageNet dataset. The optimization objective can be re-decomposed into our reward function to maximize:
[0074]
[0075] where the first reward term calculates the depth estimation error under the current camera configuration, which is the target depth error set manually. The second reward term represents the error we analyze from the model estimation. The third reward term is the RoI exclusion penalty to ensure that the pixels in the RoI are always within the FoV of the stereo camera views during the update of the camera configuration. Alpha and beta are reward scaling coefficients, which return 1 if the condition is met, otherwise 0. Among them, the second reward term greatly helps the DRL to converge, which is inspired by reward shaping. It uses the AM estimation error as a potential energy function to avoid sparse rewards during the spatial exploration process.
[0076] In the described Auto-CAM deployment phase, the specific process is as Figure 3 shown. The deployment phase is the core application phase of the Auto-CAM system, which directly uses the trained reinforcement learning agent to output the optimal camera parameter configuration according to a specific scenario. This dynamic adjustment method can significantly improve the accuracy of stereo depth estimation and provide reliable support for various practical application scenarios. The specific process is as follows:
[0077] (1) First, the binocular camera captures the left view and the right view under the current camera configuration. These images record the basic visual information of the current scene, laying a data foundation for subsequent depth estimation. (2) Then, the left image and the right image are respectively fed into an encoder with shared weights. The encoder extracts important features in the images through a convolutional network, generating two feature vectors respectively. These feature vectors are integrated into a state vector through a specific fusion strategy, characterizing the current camera configuration and environmental conditions. (3) Subsequently, the reinforcement learning agent takes the state vector as input, analyzes and makes decisions, generating an action vector. The action vector describes how to adjust the camera parameters to optimize the configuration, including focal length, baseline distance, distortion coefficient, etc. (4) Finally, the binocular camera adjusts its own configuration according to the action vector and captures new image data, realizing closed-loop optimization.
[0078] To evaluate the actual effect of Auto-CAM, Auto-CAM was deployed and verified in two edge applications: autonomous driving and defect detection. On an efficient simulator CARLA that can obtain the depth ground truth map in real time, we comprehensively evaluated the performance of Auto-CAM through experiments. In the experimental design, we selected multiple classic stereo matching models (such as PSMNet, HITNet, etc.) as benchmark testing tools, and conducted a comparative analysis with the fixed camera parameter configuration widely adopted in the current industry and other optimization methods. Table 1 shows the specific experimental results. It can be clearly seen from the data that Auto-CAM performs excellently in optimizing the camera configuration, significantly improving the accuracy of depth estimation. The depth map generated by Auto-CAM is more accurate in target distance estimation, reducing the errors caused by improper camera parameters. This result proves that the dynamic camera configuration method based on reinforcement learning can more flexibly adapt to different scenarios, providing high-quality depth information support for the perception module in autonomous driving.
[0079] Furthermore, we conducted a quantitative analysis of the exploration efficiency of Auto-CAM. In reinforcement learning methods, the exploration convergence speed is an important indicator to measure the algorithm efficiency. Figure 4 The exploration results of six methods including random search (RS), simulated annealing (SA), Bayesian optimization (BO), etc. are shown. It can be concluded that while significantly improving the performance, Auto-CAM also achieves a remarkable optimization of the exploration efficiency. In specific applications, Auto-CAM can quickly converge to the optimal parameter configuration with fewer iterations, thus reducing the time cost.
[0080] To further analyze the exploration path of Auto-CAM when optimizing the camera parameter configuration, we recorded its exploration process on multiple stereo matching models and visualized it through a heat map, as Figure 6As shown. In the heatmap, the red area represents a larger depth error generated by the configuration, while the blue area indicates a smaller depth error. It can be clearly seen from the heatmap that the Auto-CAM method exhibits a stable and efficient optimization process. Different from the possible wandering or detouring of other methods in the parameter space, Auto-CAM can directly and accurately approach the optimal area, demonstrating the efficiency and reliability of its decision-making path. This intuitive expression of the optimization process further proves the powerful adaptability of the reinforcement learning agent in the complex parameter space.
[0081] To evaluate the applicability of Auto-CAM in actual deployment, we designed a special experimental scenario to verify the accuracy level it can still maintain without a ground truth depth map. The experimental results are shown in Table 2. In this scenario, due to the lack of a real ground truth depth map, traditional optimization methods (such as RS, SA, BO, etc.) become infeasible because they cannot define the objective function. However, Auto-CAM does not rely on the ground truth depth map and can be adjusted dynamically entirely through the reinforcement learning mechanism, so it can operate normally. In this experiment, we compared Auto-CAM with the expert setting (ES) method. The results show that even under the condition of no ground truth map, Auto-CAM is still significantly better than ES. For example, in the test of the PSMNet model, the effect of the camera configuration optimized by Auto-CAM is 2.18 times better than the fixed configuration of ES. This result indicates that Auto-CAM has higher robustness and wide applicability, providing strong guarantee for the popularization of depth estimation in actual scenarios.
[0082] In addition to quantitative analysis, we also conducted a qualitative evaluation of Auto-CAM and provided the visualization results of the depth maps in these two edge applications in Figure 5 and Figure 7 . Through comparison, it can be found that the quality of the depth map generated by the camera configuration optimized by Auto-CAM has been significantly improved. Specifically, the optimized depth map not only reduces noise and artifacts, but also is clearer in terms of detail performance, especially in scenes with sparse texture or complex lighting. This improvement is particularly important for scenarios such as autonomous driving and defect detection. For example, in autonomous driving, an accurate depth map helps the vehicle correctly judge the position and distance of the obstacles ahead, improving path planning and obstacle avoidance capabilities. In defect detection, a clear depth map can help the detection system better identify the details of surface defects, improving the accuracy of product quality control.
[0083] In summary, Auto-CAM dynamically adjusts camera parameters through a reinforcement learning agent, demonstrating excellent performance advantages in multiple edge application scenarios. As can be seen from the experimental data, Auto-CAM is significantly superior to existing methods in terms of depth estimation accuracy, exploration efficiency, and adaptability. At the same time, its characteristic of not relying on depth ground truth maps makes it more practical and feasible in actual deployment. Whether it is the real-time perception requirements in autonomous driving or the high-precision requirements in industrial defect detection, Auto-CAM provides a reliable solution for depth estimation and camera configuration optimization.
[0084]
[0085] Table 1: Depth error △D of Auto CAM and different methods when optimizing in the learning phase.
[0086]
[0087] Table 2: Depth error △D of Auto CAM and different methods when optimizing in the deployment phase.
Claims
1. A camera parameter automatic optimization system for binocular depth estimation, characterized in that: include: A truth value acquisition module is used to obtain a depth truth map in a simulation environment; The Auto-CAM learning module uses an actor-critic based deep reinforcement learning framework to perform back-propagation learning based on the current camera configuration, state vector, and depth error, where the state vector is generated by the left and right view images through a weight-sharing encoder, and the depth error is calculated from the depth map generated by the pre-trained stereo matching model and the depth truth map. The camera configuration includes focal length, baseline distance, and distortion coefficient. The Auto-CAM deployment module is used to output the optimal camera parameter configuration based on the trained reinforcement learning agent in a specific scenario; The camera calibration module is used to obtain accurate camera intrinsic and extrinsic parameters through corner point scanning after changing the camera parameters; Image acquisition module, used to obtain RGB binocular images; The depth map prediction module outputs the predicted disparity map according to the deployed stereo matching network algorithm, and converts the disparity map into a depth map using the intrinsic and extrinsic parameters obtained by the camera calibration module; The visualization rendering module is used to filter and color the obtained depth map to present the binocular depth map.
2. The camera parameter automatic optimization system for binocular depth estimation according to claim 1, characterized in that: The Auto-CAM learning module adopts a reward function, including: The first reward item calculates the difference between the depth estimation error under the current camera configuration and the manually set target depth error; The second reward item is the error estimated by the analytical model; wherein the analytical model is as follows: Among them, ΔD is the accuracy of binocular depth, Δd px is the accuracy of the disparity map in pixels, b and f are the baseline and focal length of the camera, respectively, and D gt is the true value of binocular depth, r hor is the pixel density in the horizontal direction, w sen is the width of the sensor. The third bonus term, RoI exclusion penalty, is used to ensure that the pixels of the RoI are always within the field of view (FoV) of the stereo camera view during the update of the camera configuration.
3. The camera parameter automatic optimization system for binocular depth estimation according to claim 1, characterized in that: The specific process of the Auto-CAM deployment module includes: -The stereo camera captures the left view and the right view in the current camera configuration; -The left and right images are fed into the shared weight encoders respectively to generate two feature vectors, which are then integrated into a state vector through a specific fusion strategy; -The reinforcement learning agent takes the state vector as input and generates an action vector; -The stereo camera adjusts its configuration based on the motion vector and captures new image data.
4. The camera parameter automatic optimization system for binocular depth estimation according to claim 1, characterized in that: It also includes an embedded system, which includes a deep learning edge computing device and a variable baseline binocular camera system, which can obtain high-precision RGB binocular images according to specific scenes, and input the images into the encoder to obtain a state vector. The reinforcement learning agent then generates an action vector, and finally changes the parameters of the binocular camera and performs camera calibration based on the output of the better configuration. The embedded system then predicts the depth map and performs visual rendering based on the onboard deep neural network.
5. A method for automatically optimizing camera parameters for binocular depth estimation, characterized in that: The steps include: S1. Obtaining a depth truth map in a simulation environment; S2. Adopt the actor-critic based deep reinforcement learning framework to perform back propagation learning based on the current camera configuration, state vector and depth error, and output a better camera configuration; S3. Use the trained reinforcement learning agent to output the optimal camera parameter configuration in a specific scenario; S4. After changing the camera parameters, obtain accurate camera intrinsic and extrinsic parameters by corner point scanning; S5. Obtain an RGB binocular image and output a predicted disparity map according to the deployed stereo matching network algorithm; S6. Convert the disparity map into a depth map using the intrinsic and extrinsic parameters obtained by the camera calibration module; S7. Filter and color process the obtained depth map to finally present a binocular depth map with higher quality.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the method according to claim 5 are implemented.
7. A computer device comprising a processor, a memory and a computer program stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to claim 5 are implemented.
Citation Information
Patent Citations
Binocular stereo vision camera high-precision calibration algorithm based on deep reinforcement learning
CN116309869A
Mechanical arm visual depth estimation method based on deep reinforcement learning
CN119131104A
Apparatuses and methods for machine vision systems including creation of a point cloud model and / or three dimensional model based on multiple images from different perspectives and combination of depth CUES from camera motion and defocus with various applications including navigation systems, and pattern matching systems as well as estimating relative blur between images for use in depth from defocus or autofocusing applications
US20190122378A1
Method and system for depth estimation using gated stereo imaging
US20240420356A1