Binocular vision depth estimation method and system for complex weather environment
By introducing a binocular visual depth estimation method with depth control and consistency constraints, binocular image data under high-quality complex weather conditions are generated, which solves the problem of low stereo matching accuracy under severe weather conditions in the prior art, achieves higher robustness and accuracy, and improves data preparation efficiency and model training efficiency.
Patent Information
- Application Number
- CN202510022510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-23
AI Technical Summary
The existing stereo matching methods perform poorly in bad weather conditions, mainly due to the lack of coverage of the training data for complex weather scenarios, resulting in inaccurate parallax prediction and unreliable depth estimation.
The binocular visual depth estimation method is adopted with depth control and consistency constraints. By generating binocular image data under high-quality complex weather conditions, the structural consistency and content fidelity of the generated data are ensured, thereby improving the robustness and accuracy of the stereo matching model.
It significantly improves the stereo matching accuracy under severe weather conditions, improves the robustness and accuracy of the model in complex environments, avoids the tedious process of manually collecting high-quality labeled data, and reduces the training time and computing resources consumption.
Smart Images

Figure CN120031940A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a binocular vision depth estimation method and system for complex weather environments, and belongs to the technical field of computer vision and image processing. Background Art
[0002] As a core task in the field of computer vision, stereo matching has been widely used in robotics, autonomous driving, augmented reality and other fields. By estimating the disparity between the left and right views, stereo matching can generate accurate depth information, providing a solid foundation for scene understanding and spatial perception. Especially in autonomous driving, the application of stereo matching is crucial, and vehicles need to rely on accurate depth information for obstacle detection, path planning and decision-making.
[0003] With the rapid development of deep learning technology, learning-based stereo matching methods have achieved remarkable results. Advanced models such as IGEV and StereoBase have demonstrated excellent performance in the KITTI benchmark, greatly promoting the development of this field.
[0004] Despite many breakthroughs, existing stereo matching methods still perform poorly in the face of adverse weather conditions (such as rainy and foggy days). This is mainly because the performance of stereo matching models is heavily dependent on the diversity and quality of training data. Currently, most existing stereo datasets are collected under normal lighting conditions and lack sufficient coverage of complex weather scenes, which results in the trained models being unable to effectively cope with the challenges of complex environments such as low visibility and high reflectivity in practical applications, manifested as inaccurate disparity prediction or unreliable depth estimation. In addition, high-quality annotated datasets for severe weather are very scarce. This is not only because data collection is technically challenging, but also because the performance of existing active sensors (such as LiDAR and ToF) is greatly degraded in complex environments, further limiting the acquisition of accurate training data.
[0005] To solve these problems, some researchers have tried to generate stereoscopic datasets under severe weather conditions through graphical modeling or physics-based simulation methods. Although these methods can partially simulate complex scenes, the generated data have a significant field gap with the real world due to the difficulty in accurately reproducing the complex light reflections and interactions in severe weather. In addition, some researchers have expanded the dataset by actually collecting data from severe weather scenes. Although this has improved the diversity of the data to a certain extent, the number of extreme scene samples it contains is still insufficient due to equipment conditions and collection costs. Especially in dynamic environments such as rainy and snowy days, it is even more difficult to collect high-precision and wide-coverage data.
[0006] Therefore, how to construct an efficient and flexible data generation method to supplement and enhance the diversity of existing datasets becomes the key to solving the above problems. Summary of the invention
[0007] The purpose of the present invention is to overcome the defects of the prior art and creatively propose a binocular vision depth estimation method and system for complex weather environments. The present invention can not only realistically reproduce the complex environmental characteristics in bad weather, but also take into account the structural consistency and content fidelity of the generated data, ensuring that the generated data can be used to train and optimize the existing stereo matching model, and improve its robustness and accuracy in actual scenes.
[0008] The present invention introduces depth control and consistency constraints, improves the content fidelity and visual diversity of generated data, and enhances the robustness of the depth estimation model in severe weather.
[0009] The present invention is implemented by adopting the following technical solutions.
[0010] A binocular vision depth estimation method for complex weather environment includes the following steps:
[0011] Step 1: Based on the definition and acquisition method of binocular images, collect paired data sets under normal environment, weather and lighting with real parallax.
[0012] Specifically, a calibrated binocular camera is used to collect binocular image pairs in a normal environment, and the real depth data of the scene is obtained at the same time (which can be achieved using a LiDAR device). The obtained depth data is converted into a disparity map through the conversion formula between disparity and depth, thereby constructing a high-quality training dataset containing disparity labels.
[0013] The training data set includes: pairs of binocular images and their corresponding disparity data. Each pair of images is equipped with the acquired real depth, and the disparity value is calculated by the disparity and depth conversion formula. The disparity and depth conversion formula is uniformly expressed as:
[0014]
[0015] Among them, d is the disparity, f is the focal length of the camera, B is the camera baseline (that is, the distance between the two cameras), and Z is the depth value.
[0016] Through this conversion method, accurate disparity labels are obtained, providing reliable training data for the stereo matching model.
[0017] Step 2: Construct a complex weather binocular data generation model, introduce depth control and consistency constraint mechanisms, and use the normal weather dataset to generate a complex weather dataset.
[0018] Specifically, a complex weather binocular data generation model is constructed, which uses the binocular dataset and depth data under normal weather conditions to generate binocular image data under complex weather conditions with high authenticity through depth control and consistency constraints.
[0019] The generation model includes: depth control and consistency constraints.
[0020] Step 2.1: Construct a depth-controlled stereo image generation network.
[0021] Specifically, the depth-controlled binocular image generation network is used to use depth information to guide the generation of binocular images in complex weather conditions. The depth image is generated by using a depth estimation method, and a control network is used to ensure that the generated complex weather image pairs have the correct depth structure.
[0022] First, a depth estimation network is used to generate the depth information of the target image, and this depth information is used to guide the image generation process to ensure that the depth of the generated image under bad weather conditions is consistent with the depth of the normal weather image. Specifically, the depth estimation network is expressed as:
[0023] D pred =D(I R ,I L ) (2)
[0024] Among them, D pred (·) is the depth of the generated complex weather image, I R ,I L is the input image under normal weather conditions. In this way, the depth generated by the depth estimation network will be used to control the generation of binocular images under complex weather conditions and ensure that it is consistent with the real depth.
[0025] Then, a diffusion model control network is used to control the diffusion model in combination with the generated depth information to generate complex environment images that conform to the depth information. Specifically, the control network is used to guide the diffusion model to generate images using depth information as an additional condition. The control network is a network architecture that can guide a generative model (such as a diffusion model) to generate images that meet the requirements according to specific conditions. In practical applications, the role of the control network is to combine the depth information with the target weather conditions to ensure that the generated image not only visually conforms to the characteristics of severe weather, but also maintains consistency in depth.
[0026] The control network is implemented by taking the depth information as one of the inputs and inputting it into the diffusion model with the target weather characteristics (such as rain, fog, snow, etc.), thereby controlling the depth structure of the generated image. Specifically, the control network maps the depth information and weather conditions into a latent space, in which it guides the diffusion model to generate an image consistent with the input depth information. This is achieved by:
[0027] z t =Encoder(I normal ) (3)
[0028] c=F CtrlNet (D pred ,T weather ) (4)
[0029] I gen =F SD (z t ,T weather |c) (5)
[0030] Among them, I gen is the generated complex weather image, I nomal is the input image under normal weather conditions, D pred is the depth information generated by the depth estimation network, T weather is a textual prompt describing the target weather conditions, F SD (·) represents the diffusion generation model, F CtrlNet (·) depth control network, Encoder (·) represents the image encoder, z t represents the potential code corresponding to the image, and c represents the specific control information.
[0031] In this way, the control network ensures that the deep structure of the generated image is consistent with that of the normal weather image, while also giving the image the visual effect of the target severe weather.
[0032] The core idea of the control network is to introduce depth information guidance in the image generation process to maintain the depth consistency and realism of the generated image. In this way, the generated complex weather images not only conform to the conditions of bad weather in appearance, but also can truly reflect the spatial structure of objects, avoiding the loss or distortion of depth information caused by weather changes, thereby improving the usability of generated images in tasks such as stereo matching.
[0033] Step 2.2: Construct consistency constraints.
[0034] In order to ensure that the generated stereo image pair can remain consistent under adverse weather conditions, a consistency constraint is introduced. It uses a variety of techniques to enhance the matching between the left and right views of the generated image, ensuring that the generated stereo image pair has the correct geometric consistency.
[0035] The consistency of the left and right images is optimized by using a similarity calculation method based on image features and combined with depth information. In this way, the consistency constraint ensures that the generated images are not only diverse in visual content but also maintain parallax consistency.
[0036] Specifically, a feature-based matching method is adopted, and a self-attention mechanism is used to model the long-distance dependencies between image features, helping the generation network to ensure the consistency between the generated left and right views while maintaining the diversity of image content.
[0037] In addition, the vector fusion method is used to further increase the consistency of the left and right views. First, the generated image is converted into multiple vectors, where each vector represents a part of the image features. Then, for the vectors of the left and right views, the corresponding vector pairs are matched by calculating their similarity. During the matching process, the cosine similarity is used to evaluate the similarity between the vectors in the left and right images, and the most similar n vector pairs are selected for merging.
[0038] Specifically, the core process of vector fusion is as follows:
[0039] The first step is vector assignment and similarity calculation. First, the image features are converted into vectors and divided into source set and target set. Then the vector similarity between the source set and the target set is calculated.
[0040] The second step is to merge the most similar vector pairs. The n most similar vector pairs are selected, merged, and passed as new vectors to the generative model.
[0041] The third step is to generate a consistency-enhanced image. Through vector fusion, the features of the left and right images are finely adjusted to make the two images more visually consistent and meet the consistency requirements of the deep structure.
[0042] The specific expressions are as follows:
[0043] E = Match (src, dst, n)
[0044]
[0045] in, and They represent the merging and unmerging process of vectors respectively, E is the vector matching mapping obtained by similarity calculation, Match(·) represents the vector similarity calculation function, and the cosine similarity function is used for vector matching. src represents the source vector group, dst represents the target vector group, n represents the number of matches, and T m Indicates the intermediate result, T u Represents the result vector array.
[0046] Step 3: Establish supervision constraints and use the loss function to optimize the network parameters.
[0047] Specifically, the generated images under bad weather conditions and the real disparity data are used as supervisory signals to train the binocular matching network to improve its robustness in complex environments. Using the image pairs (including left and right views) generated by the generative model and their corresponding real disparity values as input, an optimization loss function is constructed, and the network parameters are adjusted by minimizing the loss function to optimize the model performance.
[0048] Specifically, firstly, the generated image is and Input into the binocular matching network and predict the disparity map D pred The goal of the network is to generate the predicted disparity D of the image pair pred and the real disparity map D gt Align to ensure the accuracy of the generated image in the depth structure. This loss is used to measure the generated disparity map D pred and the true disparity map D gt The difference between.
[0049] Specifically, L 1 The loss is expressed as:
[0050]
[0051] Where N is the total number of pixels in the image, D pred (i) and D gt (i) represents the disparity value of the generated disparity map and the real disparity map at pixel position i respectively.
[0052] By minimizing L 1 The network parameters are optimized so that the generated images can still accurately predict disparity and maintain depth consistency under adverse weather conditions.
[0053] Step 4: Save the training parameters, generate binocular matching disparity images based on the input data in complex weather environments, and complete reasoning and indicator evaluation.
[0054] Specifically, in order to objectively evaluate the effect of the generated natural image, objective evaluation indicators can be generated based on EPE (End-Point Error) and D1 (Disparity One Pixel Error).
[0055] Based on the above method, the present invention further proposes a binocular visual depth estimation system for complex weather environments, including a raw data collection subsystem, a weather data generation subsystem, a binocular matching subsystem, and a supervised optimization and result evaluation subsystem.
[0056] The connection relationship between the above components is:
[0057] The output end of the original data collection subsystem is connected to the input end of the weather data generation subsystem, the output end of the weather data generation subsystem is connected to the input end of the binocular matching subsystem, and the output end of the binocular matching subsystem is connected to the input end of the supervision optimization and result evaluation subsystem.
[0058] Beneficial Effects
[0059] Compared with the prior art, the present invention has the following advantages:
[0060] 1. The present invention can generate high-quality binocular image data under adverse weather conditions (such as rainy days, foggy days, etc.) by introducing an image generation method based on depth control, significantly improving the accuracy of stereo matching under these conditions. By combining the generated complex environment data with the real disparity data, the trained stereo matching network can better adapt to different weather changes, improving the robustness of the model in practical applications.
[0061] 2. The present invention ensures the consistency of depth and geometric structure of the generated left and right views through the depth guidance module and vector fusion method, solving the problem of inconsistent matching caused by diversity. The generated data under severe weather conditions is combined with real disparity data for training, which effectively enhances the reliability of the generated data and makes the network more stable in various complex environments.
[0062] 3. The method of the present invention avoids the tedious process of manually collecting a large amount of high-quality annotated data in different environments by generating high-quality data under complex weather conditions, greatly improving the efficiency of data preparation. At the same time, by using a strategy combining a diffusion model with depth estimation, the training time and consumption of computing resources are reduced, enabling the model to achieve higher accuracy and robustness in a shorter time. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a flow chart of the method of the present invention.
[0064] Figure 2 Schematic diagram of the neural network structure described in the method of the present invention.
[0065] Figure 3 It is a schematic diagram of the composition of the system of the present invention. DETAILED DESCRIPTION
[0066] In order to better illustrate the purpose and advantages of the present invention, the inventive method is further described below with reference to the accompanying drawings and examples.
[0067] Example
[0068] like Figure 1 As shown, a binocular vision depth estimation method for complex weather environment includes the following steps:
[0069] Step 1: Based on the definition and acquisition method of binocular images, collect paired data sets under normal environment, bad weather and lighting conditions with real parallax to ensure the diversity and comprehensiveness of the data.
[0070] Step 2: Construct a complex weather binocular data generation model, introduce a depth control module and a consistency constraint module, and use the normal weather dataset to generate a complex weather dataset, thereby improving the model's adaptability to various environmental changes.
[0071] Step 2.1: Construct a depth-controlled binocular image generation network, use depth information to guide the generation process, and ensure that the generated image is consistent with the real data in the depth dimension.
[0072] Step 2.2: Construct a consistency constraint module to ensure that the generated left and right images maintain consistency in geometry and depth information, and avoid inconsistencies caused by noise in the generation process.
[0073] Step 3: Optimize the network using a loss function to ensure that the network can accurately predict disparity in bad weather by calculating the difference between the generated image and the real data.
[0074] Step 4: Save the training parameters, generate binocular matching disparity images based on the input data in complex weather environments, and complete reasoning and indicator evaluation
[0075] Among them, the binocular visual depth estimation method for complex weather environments in the embodiment of the present application collects paired data sets under normal environments, bad weather and lighting conditions with real parallax through the definition and acquisition method of binocular images; constructs a complex weather binocular data generation model (including a depth control module and a consistency constraint module), and uses the normal weather data set to generate a complex weather data set; uses the generated data set to train the depth estimation network, and optimizes the network parameters using the loss function through supervision constraints; saves the training parameters, generates binocular matching disparity images under complex weather environments based on the input data, and completes reasoning and index evaluation. As a result, the accuracy problem of stereo matching under bad weather conditions in the prior art can be solved, and the robustness and accuracy of the depth estimation model in complex environments can be effectively improved.
[0076] Further, in one embodiment of the present application, specifically, the present invention uses a calibrated binocular camera to collect binocular image pairs in a normal environment, and uses a LiDAR device to obtain real depth data of the scene. Through the conversion formula between disparity and depth, the obtained depth data is converted into a disparity map to construct a high-quality training data set containing disparity labels.
[0077] The training data set includes: pairs of binocular images and their corresponding disparity data, where each pair of images is equipped with the real depth obtained by LiDAR, and the disparity value is calculated by the disparity and depth conversion formula. The disparity and depth conversion formula is uniformly expressed as:
[0078]
[0079] Among them, d is the disparity, f is the focal length of the camera, B is the camera baseline (i.e. the distance between the two cameras), and Z is the depth value. Through this conversion method, accurate disparity labels can be obtained, providing reliable training data for the stereo matching model.
[0080] Further, in one embodiment of the present application, a depth-controlled binocular image generation network is used to use depth information to guide the generation of binocular images in complex weather. The module generates depth images by using a depth estimation method and uses a control network to ensure that the generated complex weather image pair has the correct depth structure.
[0081] Specifically, a depth estimation network is constructed to generate the depth information of the target image, and the depth information is used to guide the image generation process to ensure that the depth of the generated image under severe weather conditions is consistent with the depth of the normal weather image. Specifically, the depth estimation network is expressed as:
[0082] D pred =D(I R ,I L ) (2)
[0083] Among them, D pred (·) is the depth of the generated complex weather image, I R I L It is the input image under normal weather conditions. In this way, the depth generated by the depth estimation network will be used later to control the generation of binocular images under complex weather conditions and ensure that it is consistent with the actual depth.
[0084] Then, a diffusion model control network is used to combine the generated depth information to control the diffusion model to generate complex environment images that conform to the depth information. Specifically, the control network is used to guide the diffusion model to generate images through depth information as an additional condition. The control network is a network architecture that can guide a generative model (such as a diffusion model) to generate images that meet the requirements according to specific conditions. In practical applications, the main role of the control network is to combine the depth information with the target weather conditions to ensure that the generated image not only visually conforms to the characteristics of severe weather, but also maintains consistency in depth.
[0085] The implementation principle of the control network is to input the depth information as one of the inputs into the diffusion model together with the target weather characteristics (such as rain, fog, snow, etc.), thereby controlling the depth structure of the generated image. Specifically, the control network maps the depth information and weather conditions into a latent space, in which the diffusion model is guided to generate an image consistent with the input depth information. This can be achieved in the following ways:
[0086] z t =Encoder(I normal ) (3)
[0087] c=F CtrlNet (D pred ,T weather ) (4)
[0088] I gen =F SD (z t ,T|c) (5)
[0089] Among them, I gen is the generated complex weather image, I nomal is the input image under normal weather conditions, D pred is the depth information generated by the depth estimation network, T weather is a textual prompt describing the target weather conditions, F SD (·) represents the diffusion generation model, F CtrlNet (·) depth control network, Encoder (·) represents the image encoder, z t represents the potential code corresponding to the image, and c represents the specific control information. In this way, the control network ensures that the deep structure of the generated image is consistent with the deep structure of the normal weather image, and also makes the image have the visual effect of the target bad weather.
[0090] The core idea of the control network is to introduce depth information guidance in the image generation process to maintain the depth consistency and realism of the generated image. In this way, the generated complex weather images not only conform to the conditions of bad weather in appearance, but also can truly reflect the spatial structure of objects, avoiding the loss or distortion of depth information caused by weather changes, thereby improving the usability of generated images in tasks such as stereo matching.
[0091] Furthermore, in one embodiment of the present application, a consistency constraint mechanism is introduced, which uses a variety of techniques to enhance the matching degree between the left and right views of the generated image, ensuring that the generated binocular image pair has the correct geometric consistency. Specifically, a similarity calculation method based on image features is used to optimize the consistency of the left and right images in combination with depth information. In this way, the consistency constraint module can ensure that the generated image is not only diverse in visual content, but also maintains parallax consistency.
[0092] Specifically, a feature-based matching method is adopted, and a self-attention mechanism is used to model the long-distance dependencies between image features, helping the generation network to ensure the consistency between the generated left and right views while maintaining the diversity of image content.
[0093] In addition, the consistency of the left and right views is further increased by using a vector fusion method. First, the generated image is converted into multiple vectors, where each vector represents a part of the image features. Then, for the vectors of the left and right views, the corresponding vector pairs are matched by calculating their similarity. During the matching process, the cosine similarity is used to evaluate the similarity between the vectors in the left and right images, and the most similar n vector pairs are selected for merging.
[0094] The core process of vector fusion is as follows:
[0095] 1. Vector assignment and similarity calculation: First, the image features are converted into vectors and divided into source and target sets. Then the vector similarity between the source and target sets is calculated.
[0096] 2. Merge the most similar vector pairs: Select the n most similar vector pairs, merge them, and pass them as new vectors to the generative model.
[0097] 3. Generate consistency-enhanced images: Through vector fusion, the features of the left and right images are finely adjusted, making the two images more visually consistent and meeting the consistency requirements of the deep structure.
[0098] The specific formula is as follows:
[0099] E = Match (src, dst, n)
[0100]
[0101] in, and They represent the merging and unmerging processes of the vectors respectively, E is the vector matching mapping obtained by similarity calculation, Match(·) represents the vector similarity calculation function, and the cosine similarity function is used for the vector matching module.
[0102] Furthermore, in one embodiment of the present application, the generated images under severe weather conditions and the real disparity data are used as supervisory signals to train the binocular matching network to improve its robustness in complex environments. The image pairs (including left and right views) generated by the generative model and their corresponding real disparity values are used as input to construct an optimization loss function, and the network parameters are adjusted by minimizing the loss function to optimize the model performance.
[0103] Specifically, firstly, the generated image and Input into the binocular matching network and predict the disparity map D pred The goal of the network is to generate the predicted disparity D of the image pair pred and the real disparity map D gt Align to ensure the accuracy of the generated image in the depth structure. This loss is used to measure the generated disparity map D pred and the true disparity map D gt Specifically, L 1 The loss is expressed as:
[0104]
[0105] Where N is the total number of pixels in the image, D pred (i) and D gt (i) represents the disparity value of the generated disparity map and the real disparity map at pixel position i respectively.
[0106] By minimizing L 1 The network parameters are optimized so that the generated images can still accurately predict disparity and maintain depth consistency under adverse weather conditions.
[0107] Furthermore, in one embodiment of the present application, in order to objectively evaluate the effect of the generated natural image, an objective evaluation index may be generated based on EPE (End-Point Error) and D1 (Disparity One Pixel Error) and the like.
[0108] Figure 2 A schematic diagram of the neural network structure used in the embodiment of the present application, which generates a binocular image data set under normal weather and complex weather conditions through a set depth control model construction; learns depth information through a depth control module to guide the depth consistency of the generated image, ensuring that the depth of the generated image under complex weather conditions is consistent with the depth of the normal weather image; learns the geometric consistency between the left and right images through a consistency constraint module to enhance the matching degree of the depth and geometric structure of the generated image, and finally solves the challenge of binocular visual depth estimation in complex weather environments.
[0109] Figure 3A schematic diagram of the composition of a binocular visual depth estimation system for complex weather environments provided in an embodiment of the present application includes a raw data collection subsystem 10, a severe weather data generation subsystem 20, a binocular matching subsystem 30, and a supervision optimization and result evaluation subsystem 40:
[0110] The raw data collection subsystem 10 is used to collect binocular image data in normal weather and complex weather environments, and generate a training data set by pairing it with the acquired real depth data.
[0111] The severe weather data generation subsystem 20 includes a depth control module and a consistency constraint module, which generates binocular image data under complex weather conditions containing real depth information.
[0112] The binocular matching subsystem 30 uses the generated data for training, guides the depth consistency of the generated image through the depth control network, enhances the geometric consistency between the left and right views through the consistency constraint module, and optimizes the generated depth map.
[0113] The supervised optimization and result evaluation subsystem 40 is used to establish a loss function, optimize the training of the aforementioned network, further save the trained network parameters and generate the final disparity map, and use built-in evaluation indicators such as EPE (End-Point Error) and D1 (Disparity One Pixel Error) to evaluate the results.
[0114] The connection relationship between the above-mentioned component systems is:
[0115] The output end of the raw data collection subsystem 10 is connected to the input end of the severe weather data generation subsystem 20 , the output end of the severe weather data generation subsystem 20 is connected to the input end of the binocular matching subsystem 30 , and the output end of the binocular matching subsystem 30 is connected to the input end of the supervision optimization and result evaluation subsystem 40 .
[0116] The explanation of the binocular vision depth estimation method for complex weather environments in the aforementioned embodiment is also applicable to the system for performing depth estimation based on the generated binocular images in this embodiment, and will not be repeated here.
[0117] The specific description above further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A binocular vision depth estimation method for complex weather environments, characterized in that: The following steps are involved: Step 1: Collect paired datasets under normal environment, weather and lighting with real parallax; Step 2: Construct a complex weather binocular data generation model, use the binocular dataset and depth data under normal weather conditions, introduce depth control and consistency constraint mechanisms, and generate binocular image data under complex weather conditions with high authenticity; Among them, the generation model includes depth control and consistency constraints; Step 2.1: Construct a depth-controlled binocular image generation network; First, a depth estimation network is used to generate the depth information of the target image, and this depth information is used to guide the image generation process to ensure that the depth of the generated image under bad weather conditions is consistent with that of the normal weather image; the depth estimation network is expressed as: D pred =D(I R ,I L ) Among them, D pred (·) is the depth of the generated complex weather image, I R ,I L is the input image under normal weather conditions; Then, a diffusion model control network is used to combine the generated depth information to control the diffusion model to generate a complex environment image that conforms to the depth information; the control network is used to guide the diffusion model to generate images through the depth information as an additional condition. The control network is a network architecture that can guide the generation model to generate images that meet the requirements according to specific conditions; The control network maps the depth information and weather conditions into a latent space where the diffusion model is guided to generate images consistent with the input depth information, by: z t =Encoder(I normal ) c=F CtrlNet (D pred ,T weather ) I gen =F SD (z t ,T weather |c) Among them, I gen is the generated complex weather image, I nomal is the input image under normal weather conditions, D pred is the depth information generated by the depth estimation network, T weather is a textual prompt describing the target weather conditions, F SD (·) represents the diffusion generation model, F CtrlNet (·) depth control network, Encoder (·) represents the image encoder, z t represents the potential code corresponding to the image, and c represents the specific control information; Step 2.2: Construct consistency constraints; Use a similarity calculation method based on image features and combine depth information to optimize the consistency of left and right images; A feature-based matching method is adopted, using a self-attention mechanism to model the long-distance dependencies between image features, helping the generative network to ensure the consistency between the generated left and right views while maintaining the diversity of image content; The vector fusion method is used to further increase the consistency of the left and right views. First, the generated image is converted into multiple vectors, where each vector represents a part of the image features. Then, for the vectors of the left and right views, the corresponding vector pairs are matched by calculating their similarity. During the matching process, the cosine similarity is used to evaluate the similarity between the vectors in the left and right images, and the most similar n vector pairs are selected for merging. The first step is vector assignment and similarity calculation. First, the image features are converted into vectors and divided into source set and target set. Then, the vector similarity between the source set and the target set is calculated. The second step is to merge the most similar vector pairs. Select the n most similar vector pairs, merge them, and pass them as new vectors to the generative model. The third step is to generate a consistency-enhanced image. Through vector fusion, the features of the left and right images are finely adjusted to make the two images more visually consistent and meet the consistency requirements of the deep structure. It is expressed as follows: E = Match (src, dst, n) in, and They represent the merging and unmerging processes of vectors respectively. E is the vector matching mapping obtained by similarity calculation. Match(·) represents the vector similarity calculation function. The cosine similarity function is used for vector matching. src represents the source vector group, dst represents the target vector group, n represents the number of matches, and T m Indicates the intermediate result, T u Represents the result vector group; Step 3: Establish supervision constraints and use loss function to optimize network parameters; The generated images under bad weather conditions and the real disparity data are used as supervisory signals to train the binocular matching network to improve its robustness in complex environments. An optimization loss function is constructed using the image pairs generated by the generative model and their corresponding real disparity values as input. The network parameters are adjusted by minimizing the loss function to optimize the model performance. Step 4: Save the training parameters, generate binocular matching disparity images based on the input data in complex weather environments, and complete reasoning and indicator evaluation.
2. The binocular vision depth estimation method for complex weather environment as claimed in claim 1, characterized in that: In step 1, a calibrated binocular camera is used to collect binocular image pairs in a normal environment, and the real depth data of the scene is obtained at the same time; the obtained depth data is converted into a disparity map through the conversion formula between disparity and depth, and a high-quality training data set containing disparity labels is constructed; The training data set includes: pairs of binocular images and their corresponding disparity data. Each pair of images is equipped with the acquired real depth, and the disparity value is calculated by the disparity and depth conversion formula. The disparity and depth conversion formula is uniformly expressed as: Where d is the disparity, f is the focal length of the camera, B is the camera baseline, and Z is the depth value; Through this conversion method, accurate disparity labels are obtained, providing reliable training data for the stereo matching model.
3. The binocular vision depth estimation method for complex weather environment as claimed in claim 1, characterized in that: In step 3, the generated image is first and Input into the binocular matching network and predict the disparity map D pred ; The goal of the network is to generate the predicted disparity D of the image pair pred and the real disparity map D gt Alignment to ensure the accuracy of the generated image in the depth structure; this loss is used to measure the generated disparity map D pred and the true disparity map D gt The difference between Among them, L1 loss is expressed as: Where N is the total number of pixels in the image, D pred (i) and D gt (i) represents the disparity value of the generated disparity map and the real disparity map at pixel position i respectively; By minimizing L1 and optimizing the network parameters, the generated images can still accurately predict disparity and maintain depth consistency under adverse weather conditions.
4. The binocular vision depth estimation method for complex weather environment as claimed in claim 1, characterized in that: In step 1 and step 4, objective evaluation indicators are generated based on End-Point Error and Disparity One Pixel Error.
5. A binocular vision depth estimation system for complex weather environments implementing the method as claimed in claim 1, characterized in that: It includes a raw data collection subsystem 10, a severe weather data generation subsystem 20, a binocular matching subsystem 30 and a supervision optimization and result evaluation subsystem 40: The raw data collection subsystem 10 is used to collect binocular image data in normal weather and complex weather environments, and generate a training data set by pairing it with the acquired real depth data; The severe weather data generation subsystem 20 includes a depth control module and a consistency constraint module to generate binocular image data under complex weather conditions containing real depth information; The binocular matching subsystem 30 uses the generated data for training, guides the depth consistency of the generated image through the depth control network, enhances the geometric consistency between the left and right views through the consistency constraint module, and optimizes the generated depth map; The supervised optimization and result evaluation subsystem 40 is used to establish a loss function, optimize the training of the aforementioned network, further save the trained network parameters and generate the final disparity map, and evaluate the results using evaluation indicators; The connection relationship between the above-mentioned component systems is: The output end of the raw data collection subsystem 10 is connected to the input end of the severe weather data generation subsystem 20 , the output end of the severe weather data generation subsystem 20 is connected to the input end of the binocular matching subsystem 30 , and the output end of the binocular matching subsystem 30 is connected to the input end of the supervision optimization and result evaluation subsystem 40 .