Foggy day image data set forming method and device based on depth estimation
Through the combination of depth estimation and optical fog plus model, high-quality, detailed label fog-day image data sets are generated, which solves the problem of unreality and insufficient universality of the fog-day image data set generation effect in the prior art, and improves the application effect of image processing and computer vision models in fog-day environments.
Patent Information
- Application Number
- CN202510948609.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-10
AI Technical Summary
The number of existing foggy day image data sets is limited and the quality is uneven. The existing synthetic fog technology lacks modeling of natural optical laws, the generation effect is unreal, and the authenticity and universality of the generated image data of virtual scene simulation fog technology affects the quality and practical application effect of synthetic fog data sets.
The depth estimation model and optical fog plus model are combined to obtain the depth information of the target image and perform semantic deepening. The image is atomized in combination with the optical fog plus model to generate images with natural fog effect, and the fog day visual task matching algorithm is used to determine the fog day visual task matching algorithm to generate the target fog day image data set.
The generation accuracy and efficiency of synthetic fog data sets are improved, and high-quality fog-day image data sets with detailed labels are generated, which improves the application effect of image processing and computer vision models in fog-day environments, and supports object detection and fusion tasks in fog-day scenarios.
Smart Images

Figure CN120451727A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fog data synthesis, and in particular to a method and device for synthesizing foggy image data based on depth estimation. Background Art
[0002] Artificial intelligence is increasingly integrated into our lives. To achieve full development, industries like autonomous driving and surveillance require adaptability in diverse weather conditions. Fog, a common natural phenomenon, is particularly prevalent in the early mornings of autumn and winter, and in the southern mountains. However, existing fog datasets are limited and of varying quality due to lighting and environmental constraints. Inconsistent fog concentration and lighting conditions result in poor usability. With the introduction of multimodal fusion technology, infrared and visible imagery have garnered significant attention due to their complementary nature. Given the scarcity and difficulty of collecting paired fog datasets, synthetic fog datasets offer significant advantages, better supporting advanced vision task models such as fusion and detection.
[0003] When collecting foggy datasets, obtaining real data is difficult and label updates are expensive. To address the scarcity of open-source data, overlaying virtual fog on images is often used to construct foggy image datasets. However, existing methods have several shortcomings in generating fog: First, existing RGB channel-based synthetic fog techniques only apply a uniform density fog layer to the image, lacking modeling and application of natural optical laws, resulting in unrealistic results. Second, although the center point synthetic fog method based on the standard optical model applies an atmospheric scattering model, its applicable range is limited due to the random selection of center points, making it difficult to apply to complex outdoor environments. Furthermore, while virtual scene simulation fog technology can generate paired images on a simulation platform, the authenticity and versatility of the generated image data still need to be verified. Furthermore, the lack of labels limits its application in fusion and detection models. These shortcomings affect the quality and practical application of synthetic fog datasets. Summary of the Invention
[0004] Based on this, it is necessary to provide a method and device for synthesizing foggy image datasets based on depth estimation, which can improve the accuracy of generating synthetic fog datasets and enhance the efficiency of constructing large-scale datasets, in order to address the above technical problems.
[0005] A method for synthesizing foggy image data sets based on depth estimation, the method comprising: Construct a foggy image data synthesis model. The foggy image data synthesis model includes: depth estimation model, optical fog model, and synthetic fog image model.
[0006] After optimizing and training the depth estimation model using a monocular depth estimation method based on a public dataset, the target image scene data is input into the optimized depth estimation model for semantic deepening, and a synthetic fog effect of the depth map corresponding to the target image scene data is output.
[0007] According to the synthetic fog effect, the depth map is atomized by the optical fog model to generate a fogged image.
[0008] The foggy visual task matching algorithm of the synthetic foggy image model is determined according to the foggy image and the preset target task, and the target foggy image dataset is generated according to the foggy visual task matching algorithm.
[0009] A device for synthesizing foggy image data based on depth estimation, the device comprising: The synthesis model building module is used to build a synthesis model for foggy image data. The synthesis model for foggy image data includes: a depth estimation model, an optical fogging model, and a synthetic fog image model.
[0010] The synthetic fog effect generation module is used to optimize and train the depth estimation model using a monocular depth estimation method based on a public dataset, input the target image scene data into the optimized depth estimation model for semantic deepening, and output the synthetic fog effect of the depth map corresponding to the target image scene data.
[0011] The fog image generation module is used to generate a fog image after performing fog processing on the depth map through an optical fog model according to the synthetic fog effect.
[0012] The synthetic fog data module is used to determine the foggy visual task matching algorithm of the synthetic fog image model based on the fogged image and the preset target task, and generate the target foggy image dataset based on the foggy visual task matching algorithm.
[0013] The above-mentioned method and device for synthesizing foggy image datasets based on depth estimation firstly combines a depth estimation model with an optical fogging model to more accurately simulate the optical phenomena in foggy environments. After optimizing and training the depth estimation model, it is used to obtain depth information of the target image and perform semantic deepening to generate a depth map that more closely resembles real-world fog effects. Next, the optical fogging model performs fogging based on the scene's depth information, producing an image with a natural fog effect. Secondly, the depth map information guides the fogging process of the optical fogging model, achieving a more refined fogging effect that is more adaptable to complex outdoor environments. Specifically, the depth map provides spatial position information of different objects and backgrounds in the scene, allowing the optical fogging model to generate fog layers of varying densities and levels based on the scene's actual structure, simulating a more realistic foggy effect. Furthermore, this technical solution addresses the issues of virtual scene fog simulation technology regarding the authenticity and versatility of generated image data. By combining depth estimation with the optical fogging model, foggy images can be synthesized in real scenes, avoiding the limitations of virtual scenes. Furthermore, the synthetic fog image model also determines the fog vision task matching algorithm based on the characteristics of the original public dataset by presetting the target task. This generates a target fog image dataset with the original detailed labels, further improving the dataset's quality and versatility. This successfully constructs a high-quality, detailed fog image dataset. This dataset not only improves the application of image processing and computer vision models in foggy environments, but also provides more realistic and efficient data support for tasks such as object detection and fusion in foggy scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of a fog map synthesis framework of a fog image dataset synthesis method based on depth estimation in one embodiment; Figure 2 1 is a flow chart of a method for synthesizing foggy image data based on depth estimation in one embodiment; Figure 3 Schematic diagram of a depth estimation model framework based on ResNet in one embodiment; Figure 4 A schematic diagram of a process for synthesizing a fogged image using a standard optical model in one embodiment; Figure 5 Schematic diagram showing the comparison of synthesis results of dense fog data from the MFNet public dataset in a daytime scene in one embodiment; Figure 6 Schematic diagram showing the comparison of synthesis results of dense fog data from the MFNet public dataset for a night scene in one embodiment; Figure 7Schematic diagram showing the comparison of synthesis results of dense fog data from the RoadScene public dataset in a daytime scene in one embodiment; Figure 8 Schematic diagram showing the comparison of synthesis results of dense fog data from the RoadScene public dataset in a night scene in one embodiment; Figure 9 A schematic diagram of constructing a classification for a data set in one embodiment; Figure 10 A schematic diagram showing a comparison of synthetic fog images generated with different fog concentration parameters in one embodiment; Figure 11 is a structural block diagram of a device for synthesizing foggy image data based on depth estimation in one embodiment; Figure 12 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0016] The present invention provides a method for synthesizing foggy image data based on depth estimation, which can be applied to Figure 1 The fog image synthesis framework shown includes three modules: depth estimation model, optical fog model and synthetic fog dataset, which respectively realize image scene depth estimation, optical model fog processing, and synthetic fog data pairing and organization.
[0017] In one embodiment, Figure 2 As shown in the figure, a method for synthesizing foggy image dataset based on depth estimation is provided. Figure 1 The fog map synthesis framework in
[15] is used as an example to illustrate the following steps: Step 202: Construct a foggy image data synthesis model.
[0018] The foggy image data synthesis model includes: depth estimation model, optical fogging model and synthetic fog image model.
[0019] In step 204 , after optimizing and training the depth estimation model using a monocular depth estimation method based on a public dataset, the target image scene data is input into the optimized depth estimation model for semantic deepening, and a synthetic fog effect of the depth map corresponding to the target image scene data is output.
[0020] Step 206 : After performing a fogging process on the depth map using an optical fogging model according to the synthetic fog effect, a fogged image is generated.
[0021] Step 208 : determining a foggy visual task matching algorithm for the synthetic foggy image model according to the foggy image and the preset target task, and generating a target foggy image dataset according to the foggy visual task matching algorithm.
[0022] The above-mentioned method for synthesizing foggy image datasets based on depth estimation first introduces a combination of a depth estimation model and an optical fogging model to more accurately simulate the optical phenomena in foggy environments. After optimizing and training the depth estimation model, it is used to obtain depth information of the target image and perform semantic deepening to generate a depth map that more closely resembles real-world fog effects. Next, the optical fogging model performs fogging based on the scene's depth information, producing an image with a natural fog effect. Secondly, the depth map information guides the optical fogging model's fogging process, achieving a more refined fogging effect that is more adaptable to complex outdoor environments. Specifically, the depth map provides spatial position information of different objects and backgrounds in the scene, allowing the optical fogging model to generate fog layers of varying density and depth based on the scene's actual structure, simulating a more realistic foggy effect. Furthermore, this technical solution addresses the issues of virtual scene fog simulation technology regarding the authenticity and versatility of generated image data. By combining depth estimation with the optical fogging model, foggy images can be synthesized in real scenes, avoiding the limitations of virtual scenes. Furthermore, the synthetic fog image model also determines the fog vision task matching algorithm based on the characteristics of the original public dataset by presetting the target task. This generates a target fog image dataset with the original detailed labels, further improving the dataset's quality and versatility. This successfully constructs a high-quality, detailed fog image dataset. This dataset not only improves the application of image processing and computer vision models in foggy environments, but also provides more realistic and efficient data support for tasks such as object detection and fusion in foggy scenes.
[0023] In one embodiment, the depth estimation model is built based on ResNet, wherein the depth estimation model includes: a feature fusion module, a residual convolution module, and an adaptive output module.
[0024] It is worth noting that the depth estimation model is used to generate the depth map of the target image scene. Figure 3 As shown in the figure, this model is built on the ResNet framework and consists of a feature fusion module, a residual convolution module, and an adaptive output module. It can convert input images into corresponding depth maps. It uses monocular depth estimation methods and is trained using large-scale data from multiple classic datasets. Specifically, training for each dataset is treated as an independent task, aiming to find a near-Pareto optimal solution for that dataset.
[0025] In one embodiment, a monocular depth estimation method is used according to different training data sets in the public data set to define the loss function of each training data set in the depth estimation model training: ; ; in, is the size of the training dataset l, is the scale and translation invariance loss, is the multi-scale and scale-invariant gradient matching term, M is the number of pixels in the corresponding image of the training dataset, is the disparity prediction value, is the true value of the label, is the type of loss function, is a hyperparameter. The least squares method is used to minimize the loss function to obtain the scale and translation invariant loss function: ; ; in, s is the resolution parameter, t is a positional parameter, is the difference between the normalized disparity prediction value and the label value, For scale k The difference in the disparity map at In the x-axis direction The gradient, In the y-axis direction The gradient, It is a scale level with multiple scale changes.
[0026] By introducing a 3D dataset augmentation training task into each of the training datasets, if the scale and translation invariance loss functions of the augmented task optimization training are reduced, the minimum value of all scale and translation invariance loss functions after training is taken as the model optimization parameter: ; in, is the augmented task for current training, are model parameters. Otherwise, the scale- and translation-invariant loss function is retained, the augmented task currently being trained is discarded, and the task dataset to be optimized is obtained. The depth estimation model is optimized based on the scale- and translation-invariant loss function and the task dataset to be optimized, and the target image is trained using the optimized depth estimation model.
[0027] In one embodiment, target image scene data and the target image are input into an optimized depth estimation model for semantic deepening, and a predicted depth map is output. The predicted depth map is grayscale converted and then brightness inverted to generate a first grayscale map. The scene depth information of the first grayscale image is extracted using a standard optical algorithm: ; ; in, For the fog image, is the target image, is the transmission diagram, is atmospheric light, is the pixel in the image, is the scene depth, is the atmospheric scattering coefficient. According to the scene depth information, a synthetic fog effect of the depth map corresponding to the target image scene data is generated.
[0028] It is worth noting that based on the standard optical model, the fog image is synthesized by combining the depth map and the original clear image. The specific process is as follows Figure 4 First, the depth map obtained in the previous stage is grayscale converted and brightness inverted to generate the corresponding grayscale image. Then, the optical formula is used: ; in, For the fog image, is the target image, is the transmission diagram, is atmospheric light, is the pixel point in the image. The transmission map can be obtained by the formula ,in Indicates the scene depth, is the atmospheric scattering coefficient.
[0029] In one embodiment, the visible light image of the depth map is subjected to superimposed fogging processing by the optical fogging model according to the synthetic fog effect, thereby generating a fogged image corresponding to the depth map.
[0030] In one embodiment, the target tasks include: constructing a foggy weather vision task matching algorithm performance test dataset, constructing a foggy weather infrared-visible light pixel-level pairing fusion dataset, and constructing a foggy weather target detection dataset.
[0031] In one of the embodiments, a performance test dataset for foggy visual task matching algorithms is constructed, and the specific steps are as follows: the semantic segmentation part of the MFNet public dataset is selected, and its 4-channel data is first extracted and reconstructed to obtain a 3-channel RGB image. Subsequently, the image is classified according to the original classification label and divided into daytime and nighttime scenes. These classified data are input into the depth estimation model to match the corresponding fog brightness, and foggy image datasets of different concentrations during the day and at night are generated by adjusting the fog concentration parameters. This dataset retains the original image information and pairs the depth map, which is suitable for the effect test and comparative evaluation of foggy visual task matching algorithms, and supports subjective visual comparison and objective indicator evaluation. Some examples of datasets are as follows: Figure 5 and Figure 6 shown.
[0032] It is worth noting that this method only needs to input clear images to quickly generate synthetic fog images of multiple scenes and different concentrations and the corresponding depth maps.
[0033] In one of the embodiments, a foggy infrared-visible light pixel-level pairing fusion dataset is constructed, and the specific steps are as follows: experiments are conducted based on the RoadScene public dataset. This dataset divides infrared and visible light images independently. First, 221 visible light images are classified into daytime scenes under strong lighting conditions and nighttime scenes under weak lighting conditions, and the corresponding depth maps are generated using a depth estimation model. In view of the fact that the daytime scenes in the dataset are sunny and the night scenes are bright, we selected a higher fog brightness value to match the scene conditions. Finally, the fogged visible light image is paired with the original infrared image to construct a dataset containing pixel-level paired visible light images, infrared images, depth maps and fogged images, which is suitable for image fusion tasks. At the same time, retaining the original image is helpful for evaluating fusion algorithms based on subjective visuality and objective indicators, such as comparative analysis of indicators such as visual information fidelity (VIF) and mutual information (MI). The foggy infrared-visible light pixel-level pairing fusion dataset constructed based on the RoadScene dataset is shown as follows: Figure 7 、 Figure 8 As shown in the figure, the dataset realizes pixel-level pairing, from left to right they are visible light image, depth map, fog image, and infrared image.
[0034] It is worth noting that, based on an existing visible-infrared paired dataset, synthetic fog was overlaid on the visible light images. Research has shown that light fog has little impact on infrared imaging. Therefore, the original infrared image data and labels were retained to construct a pixel-level visible-infrared paired dataset for foggy conditions. This dataset effectively fills a gap in multimodal research on foggy conditions and provides a large amount of reliable test data for related image fusion algorithms.
[0035] In one embodiment, a foggy target detection dataset is constructed. The specific steps are as follows: the detection part in the MFNet public dataset is selected, and then it is classified according to the lighting conditions, such as Figure 9 As shown in Figure 2, the images are divided into three categories: strong light, insufficient light, and weak light. The visible light image data in these categories are input into the depth estimation model, and after matching the corresponding fog brightness, different fog concentration parameters are selected to generate the images in Figure 2. Figure 10 The foggy visible light image data under the shown lighting conditions and with different fog concentrations are paired with the infrared images and detection labels in the original dataset to obtain the foggy target detection dataset. Figure 10 This is a comparison display of some image data of this dataset: from top to bottom, they are thick fog and light fog scenes during the day, thick fog and light fog scenes at night.
[0036] It is worth noting that fog was added to the publicly available annotated dataset to generate clearly labeled foggy images to support the training and generalization of the foggy object detection algorithm. Furthermore, because infrared images are less susceptible to fog, a paired infrared and visible light dataset was used to evaluate the detection algorithm's performance on different image data, thereby improving its overall accuracy.
[0037] It should be understood that although Figure 2 、 Figure 4 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 2 、 Figure 4 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0038] In one embodiment, Figure 11 As shown, a device for synthesizing foggy image data based on depth estimation is provided, comprising: a synthesis model building module 1102, a synthetic fog effect generation module 1104, a fogged image generation module 1106, and a synthetic fog data module 1108, wherein: The synthesis model building module 1102 is used to build a synthesis model for foggy image data. The synthesis model for foggy image data includes: a depth estimation model, an optical fogging model, and a synthetic fog image model.
[0039] The synthetic fog effect generation module 1104 is used to optimize and train the depth estimation model using a monocular depth estimation method based on a public data set, input the target image scene data into the optimized depth estimation model for semantic deepening, and output the synthetic fog effect of the depth map corresponding to the target image scene data.
[0040] The fog image generation module 1106 is configured to generate a fog image by performing fog processing on the depth map using an optical fog model according to a synthetic fog effect.
[0041] The synthetic fog data module 1108 is used to determine the foggy visual task matching algorithm of the synthetic fog image model according to the foggy image and the preset target task, and generate the target foggy image dataset according to the foggy visual task matching algorithm.
[0042] For the specific definition of a foggy weather image data synthesis device based on depth estimation, please refer to the definition of a foggy weather image data synthesis method based on depth estimation above, which will not be repeated here. The various modules in the above-mentioned foggy weather image data synthesis device based on depth estimation can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0043] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 12 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for synthesizing foggy image data based on depth estimation is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.
[0044] Those skilled in the art will understand that Figure 11-12The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0045] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Construct a foggy image data synthesis model. The foggy image data synthesis model includes: depth estimation model, optical fog model, and synthetic fog image model.
[0046] After optimizing and training the depth estimation model using a monocular depth estimation method based on a public dataset, the target image scene data is input into the optimized depth estimation model for semantic deepening, and a synthetic fog effect of the depth map corresponding to the target image scene data is output.
[0047] According to the synthetic fog effect, the depth map is atomized by the optical fog model to generate a fogged image.
[0048] Determine the foggy visual task matching algorithm of the synthetic foggy image model based on the foggy image and the preset target task, and generate the target foggy image dataset based on the foggy visual task matching algorithm Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0049] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0050] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A method for synthesizing foggy image datasets based on depth estimation, characterized in that: The method comprises: Constructing a foggy image data synthesis model; the foggy image data synthesis model includes: a depth estimation model, an optical fogging model, and a synthetic fog image model; After optimizing and training the depth estimation model using a monocular depth estimation method based on a public dataset, the target image scene data is input into the optimized depth estimation model for semantic deepening, and a synthetic fog effect corresponding to the depth map of the target image scene data is output; After performing a fogging process on the depth map using the optical fogging model according to the synthetic fog effect, a fogged image is generated; A foggy weather vision task matching algorithm of the synthetic foggy image model is determined according to the foggy image and a preset target task, and a target foggy weather image dataset is generated according to the foggy weather vision task matching algorithm.
2. The method for synthesizing foggy image data based on depth estimation according to claim 1, characterized in that: The depth estimation model is built based on ResNet, wherein the depth estimation model includes: a feature fusion module, a residual convolution module and an adaptive output module.
3. The method for synthesizing foggy image data based on depth estimation according to claim 1, characterized in that: The depth estimation model is optimized and trained using a monocular depth estimation method based on a public dataset, including: According to different training data sets in the public data set, a monocular depth estimation method is used to define the loss function of each training data set in the depth estimation model training: in, For training data set The size of is the scale and translation invariance loss, is the multi-scale and scale-invariant gradient matching term, M is the number of pixels in the corresponding image of the training dataset, is the disparity prediction value, is the true value of the label, is the type of loss function, is a hyperparameter; The least squares method is used to minimize the loss function to obtain the scale and translation invariant loss function: in, s is the resolution parameter, t is a positional parameter, is the difference between the normalized disparity prediction value and the label value, For scale k The difference in the disparity map at In the x-axis direction The gradient, In the y-axis direction The gradient, is a scale level with multiple scale changes; By introducing a 3D dataset enhancement training task into each of the training datasets, if the scale and translation invariance loss function of the enhanced task optimization training is reduced, the minimum value of all scale and translation invariance loss functions after training is taken as the model optimization parameter: in, is the augmented task for current training, are model parameters; Otherwise, the scale and translation invariant loss function is retained, the enhanced task currently being trained is discarded, and the task dataset to be optimized is obtained; The depth estimation model is optimized according to the scale and translation invariant loss function and the task dataset to be optimized, and the target image is trained using the optimized depth estimation model.
4. The method for synthesizing foggy image data based on depth estimation according to claim 3, characterized in that: Input the target image scene data into the optimized depth estimation model for semantic deepening, and output a synthetic fog effect corresponding to the depth map of the target image scene data, including: Inputting target image scene data, i.e., the target image, into an optimized depth estimation model for semantic deepening, outputting a predicted depth map, and performing brightness inversion on the predicted depth map through grayscale conversion to generate a first grayscale map; The scene depth information of the first grayscale image is extracted using a standard optical algorithm: in, For the fog image, is the target image, is the transmission diagram, is atmospheric light, is the pixel in the image, is the scene depth, is the atmospheric scattering coefficient; A synthetic fog effect of a depth map corresponding to the target image scene data is generated according to the scene depth information using professional optical theory.
5. The method for synthesizing foggy image data based on depth estimation according to claim 4, characterized in that: After performing a fogging process on the depth map using the optical fogging model according to the synthetic fog effect, a fogged image corresponding to the original target image is generated, including: After the visible light image of the depth map is subjected to superimposed fogging processing through the optical fogging model according to the synthetic fog effect, a fogged image corresponding to the original target image is generated.
6. The method for synthesizing foggy image data based on depth estimation according to claim 5, characterized in that: The target tasks include: constructing a foggy visual task matching algorithm performance test dataset, constructing a foggy infrared-visible light pixel-level pairing fusion dataset, and constructing a foggy target detection dataset.
7. A device for synthesizing foggy image data based on depth estimation, characterized in that: The device comprises: A synthesis model construction module is used to construct a foggy image data synthesis model; the foggy image data synthesis model includes: a depth estimation model, an optical fogging model and a synthetic fog image model; A synthetic fog effect generation module is used to optimize and train the depth estimation model using a monocular depth estimation method based on a public dataset, input the target image scene data into the optimized depth estimation model for semantic deepening, and output a synthetic fog effect corresponding to the depth map of the target image scene data; a fog image generation module, configured to generate a fog image by performing fogging processing on the depth map using the optical fog model according to the synthetic fog effect; The synthetic fog data module is used to determine the foggy visual task matching algorithm of the synthetic fog image model according to the fogged image and the preset target task, and generate the target foggy image dataset according to the foggy visual task matching algorithm.
Citation Information
Patent Citations
Image defogging method and device based on deep learning staged training
CN116452470A
Mist synthesis method based on scattering model and depth estimation
CN117078552A
Improved YOLOv8n-based unmanned aerial vehicle target detection method in foggy environment
CN118570449A
Foggy day data set synthesis method and device based on depth estimation model in power transmission scene
CN120182751A
Method and device for learning fog-invariant feature
US20230419654A1