Method for restoring image in fog environment by using artificial intelligence model, computing device therefor, and recording medium therefor
The neural network-based fog removal method segments images by depth and combines depth-specific defogged segments to enhance image clarity and object recognition in foggy conditions, addressing the limitations of existing technologies.
Patent Information
- Application Number
- PCT/KR2025/006868
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-05-21
- Publication Date
- 2026-01-22
AI Technical Summary
Existing fog removal technologies fail to accurately restore objects in images with extreme optical haze due to varying fog conditions at different depths, leading to loss of important details in foreground or background objects.
An artificial neural network model segments fog images based on depth information, generates depth-specific defog images, and combines these segments to produce a high-quality defogged image, using a depth adversarial generator, autoencoder, and defog adversarial generator to enhance image clarity.
The method effectively improves object recognition and image restoration in foggy environments by preserving detailed features and maintaining global consistency, reducing blurring, and enhancing visual quality.
Smart Images

Figure KR2025006868_22012026_PF_FP_ABST
Abstract
Description
Image restoration method in a foggy environment using an artificial intelligence model, a computational device therefor, and a recording medium thereof
[0001] The present invention relates to a deep learning-based image processing technology for improving object recognition in a foggy environment.
[0002]
[0003] Fog removal technology, based on an atmospheric scattering model, reduces optical haze caused by fog and improves image clarity. This technology improves object visibility and is utilized in a variety of applications, including traffic management, navigation, aviation, and security systems.
[0004] Most classical techniques focus on reconstructing contrast and color to restore image details, and recently, many image technologies are being developed that remove fog with high accuracy even under complex fog conditions using deep learning models.
[0005] However, extreme optical haze severely reduces the visibility of objects in images, making object recognition and reconstruction impossible. In particular, in extremely dense fog, portions of the image appear nearly white, completely losing object details and making reconstruction difficult.
[0006] Thus, fog removal technologies proposed to date have problems in accurately identifying and restoring objects under these extreme conditions, specifically as follows.
[0007] Existing models apply the same defog method to the entire image. However, this approach has a technical limitation: it fails to provide optimal defog for both the foreground and background because it does not account for varying fog conditions at different depths.
[0008] Such a single depth-based approach is likely to lose important details in complex scenes, especially details in the foreground or distant objects in the background.
[0009]
[0010] In fields such as traffic surveillance, aviation, security, and disaster relief, where the goal is to restore images or videos similar to the original, solving the above-mentioned problems is of utmost importance.
[0011] The present invention was created against this technical background, and aims to improve the accuracy of object recognition and image restoration in an environment with high optical turbidity due to thick fog.
[0012]
[0013] In order to solve the above technical problem, the image restoration method in a foggy environment of the embodiment includes a first step of receiving a foggy image as input in an artificial neural network model, a second step of generating a segmentation image representing a distribution of segmentation for each depth from a depth image created based on depth information of the fog, a third step of generating a defog image for each depth level of fog from the input fog image based on the segmentation image, and a fourth step of generating a defog image of the fog image by selectively combining only partial images corresponding to segmentation for each depth from the defog images for each depth level.
[0014] The artificial neural network model includes a depth adversarial generator that receives the fog image as input, generates a depth image representing the distribution of fog at each depth, and operates to determine the authenticity of the generated depth image, and in the second step, the depth image is generated based on the fog image input by the depth adversarial generator.
[0015] The depth information of the above fog can be obtained through a 3D camera.
[0016] The loss function of the above-mentioned deep adversarial generator ( )Is, And, = predicted depth image, = This is the actual depth image.
[0017] The second step is executed by an autoencoder, an unsupervised learning neural network that compresses (encodes) and then restores (decodes) the input data.
[0018] The loss function of the above autoencoder ( )Is, , x = input depth image, = segmentation image generated based on x, = This is the correct answer for the generated segmentation image.
[0019] The fourth step above comprises: i) generating a binary map for each depth level based on the segmentation image; ii) applying the generated binary map for each depth level to a defog image for a corresponding depth level among the defog images for each depth level generated in the second step to generate a partial image for each depth; and iii) combining the partial images for each depth generated in the process ii) to generate the defog image.
[0020] The artificial neural network model includes a defog adversarial generator that receives depth array information of the segmentation image and the fog image as input to generate the defog image and determines the authenticity of the generated defog image, and the third and fourth steps are performed by the defog adversarial generator.
[0021] The above defog adversarial generator includes a generator that generates the defog image, and a discriminator that compares the defog image generated by the generator with an actual image to determine whether it is genuine, and the loss function of the generator ( )Is, And, = Output image of the defog adversarial generator, = actual value of defog image, λ = loss weight, r: loss weight complement (1 - λ).
[0022] The above loss weight is a value greater than the above loss weight compensation.
[0023] The above artificial neural network model further includes an image enhancer that improves the visual quality of the generated defog image, and the output of the image enhancer is input to the discriminator.
[0024] The loss function of the above discriminator ( )Is, And, = Defog image for patch i, = is the actual image corresponding to patch i.
[0025] Another embodiment of the present invention also includes a computing device and a recording medium for implementing the above-described method.
[0026]
[0027] According to the embodiment, by realizing sophisticated and natural fog removal in various environments through depth information-based segmentation and customized fog removal technology, and by improving object recognition performance and providing high-quality image restoration technology applicable to a wide range of fields such as military, automotive, and security by using a deep learning model capable of real-time processing, it can contribute to improving visual quality and safety.
[0028]
[0029] Figure 1 shows the overall framework of the embodiment.
[0030] Figure 2 shows the overall flow of an image restoration method according to an embodiment.
[0031] Figure 3 illustrates the functional configuration of an artificial neural network model in a modular manner.
[0032] Figure 4 schematically shows the structure of data used to train an artificial neural network model.
[0033] Figure 5 is a functional block of a computational device that implements an image restoration method of an embodiment.
[0034]
[0035] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, detailed descriptions of well-known functions or components that may obscure the gist of the present invention will be omitted in the following description and the attached drawings. Additionally, throughout the specification, the term "including" a component does not exclude other components, unless specifically stated otherwise, but rather implies the inclusion of other components.
[0036] Additionally, while terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms may be used to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.
[0037] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0038] Unless specifically defined otherwise, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning within the context of the relevant technology, and shall not be construed in an idealized or overly formal sense unless explicitly defined herein.
[0039]
[0040] Previously proposed fog removal techniques have had limitations in image restoration because they restore images based on a single depth, even though fog has multiple depths. However, the present invention segments a fogged image input to an artificial neural network model based on the depth of the fog, defogs the segmented portion, and combines the defogged segments according to depth to predict the defogged image, thereby restoring the defogged image more accurately.
[0041] Figure 1 shows the overall framework of the embodiment, and Figure 2 shows the overall flow of the image restoration method according to the embodiment.
[0042] The image restoration method in a foggy environment of the embodiment includes a first step (S10) of receiving a foggy image as input in an artificial neural network model, a second step (S20) of generating a segmentation image representing a distribution of segmentation for each depth from a depth image created based on depth information of the fog, a third step (S30) of generating a defog image for each depth level of fog from the input fog image based on the segmentation image, and a fourth step (S40) of generating a defog image of the fog image by selectively combining only partial images corresponding to segmentation for each depth from the defog images for each depth level.
[0043]
[0044] Meanwhile, in the embodiment, the artificial neural network model is configured to predict fog depth information from an input image, but this is not necessarily limited to this. For example, if an image is acquired through a 3D camera that records depth information along with the image, the depth information provided together with the image can be used. In this case, there is the advantage of reducing the cost and size of training the artificial neural network model.
[0045] Here, the artificial neural network model can be configured as shown in Fig. 3 for the above data processing. Fig. 3 merely illustrates the division by functional modules, and the artificial neural network model can be implemented in various forms depending on the network environment. For example, the image restoration method in a foggy environment of the embodiment comprises steps 1 through 4, and the artificial intelligence for data operations in each step can be implemented separately, and the artificial neural network model can be implemented by connecting these artificial intelligences in a network form.
[0046] As illustrated in FIG. 3, the artificial neural network model (830) can be configured to include a depth adversarial generator (831), a defog adversarial generator (833), and an autoencoder (835), and each module can be implemented as a model pre-learned through a training data set.
[0047] Here, the deep adversarial generator (831) and the defog adversarial generator (833) can be implemented based on a generative adversarial network (GAN), which consists of two neural networks (a generator and a discriminator) that learn competitively. Here, the generator generates fake data, the discriminator distinguishes between real and fake data, and through this process, the generator generates increasingly realistic data.
[0048] Although not shown in the drawing, the depth adversarial generator (831) and the defog adversarial generator are configured to include a generator and a discriminator, respectively.
[0049] Generator (Generator_G) in Deep Adversarial Generator (831) depth ) generates a depth image with depth information about the perspective of the fog from the input fog image, and the discriminator (Discriminator_D) depth ) receives the generated depth image and the actual image (if available) as input and determines whether they are authentic. The depth adversarial generator (831) repeats this process, and the two networks develop competitively with each other.
[0050] Meanwhile, Fig. 4 schematically shows the configuration of data for training an artificial neural network model (830).
[0051] Creating a high-quality training data set requires 3D spatial data containing various fog shapes. To achieve this, in the embodiment, training images are extracted from random points in a 3D image (900), and points where the shape of the object appears different are identified and acquired as training images to form a training data set.
[0052] An artificial neural network model (830) can be pre-trained using a training data set acquired in this manner. At this time, the training data may comprise a defog image, a fog image, and a depth image, which constitute one training data set.
[0053] The depth adversarial generator (831) is a model that infers depth information from an image, and through training, it has a loss function as shown in Equation 1 ( ) is trained to minimize the value.
[0054]
[0055] In mathematical expression 1, = Output (depth image), = is the correct answer for the output.
[0056] This loss function evaluates the accuracy of depth information generation and trains the model to minimize the difference between the predicted depth image and the actual one.
[0057] The defog adversarial generator (833) receives the depth array information of the segmentation image (SI) output from the auto-encoder (835) and the fog image as input, generates a defog image (DEI) for each depth level of fog from the input fog image (RI), and generates a defog image (DI) by combining the partial images (PI) for each depth obtained by applying a depth-specific binary map (BM).
[0058] The created defog image (DI) can be compared with the actual image (RI) to determine its authenticity and output as the final model output value.
[0059] For this purpose, the defog adversarial generator (833) is a generator (Generator_G) that generates a defog image (DI). defog ) and a discriminator (Discriminator_D) that determines the authenticity of the generated defog image (DI). defog ) can be configured, and mathematical expression 2 is a double generator (Generator_G defog ) loss function ( ) is shown.
[0060]
[0061] In mathematical expression 2, = Output image of the defog adversarial generator, = actual value of defog image, λ = loss weight, r: loss weight complement (1 - λ).
[0062] This loss function weight-updates the loss in a way that gives more confidence to the correct answer in shallower fog, and the reason for this weighted update is that defogging to different degrees depending on the depth of the fog can better retain information about the object. To this end, in Equation 1, the loss weight (λ) is defined as a value greater than the loss weight compensation (r), for example, the loss weight (λ) is 0.6, and the loss weight compensation is 1-λ, which is 0.4, but this is not intended to be limited thereto, and these values may be appropriately adjusted during the learning and inference process of the artificial neural network model.
[0063] And, Discriminator_D defog ) loss function ( ) is as shown in mathematical formula 3.
[0064]
[0065] In mathematical expression 3, = Defog image for patch i, = is the actual image corresponding to patch i.
[0066] Here, a patch is a small subregion of an image, which can be thought of as dividing the entire image into multiple smaller regions. For example, dividing a 224x224 image into 16x16 patches results in a total of 196 patches.
[0067] The patch-based operation process is as follows:
[0068] The discriminator divides the original image and the generated image (defog image) into patches of the same size, and compares the patches of the original image and the patches of the generated image corresponding to each location.
[0069] The discriminator computes the difference pixel-wise for each patch pair, and then sums or averages the differences across all patches to compute the overall Patch Loss.
[0070] This loss function can be particularly useful for tasks such as haze removal, as it preserves detailed features of the image while improving its overall quality.
[0071] Let's look at this part in more detail:
[0072] Preservation of local details: By comparing small regions of the image, detailed features are well preserved.
[0073] Maintain global consistency: Applied uniformly across the entire image, improving overall quality.
[0074] Texture and structure improvements: By comparing small areas, texture and structural features can be better captured.
[0075] Blur reduction: Prevents excessive blurring that can occur when using only the average loss over the entire image.
[0076]
[0077] Meanwhile, if the artificial neural network model (830) is configured to include an image enhancement, = is the output of the image enhancement. The image enhancement is a generator (Generator_G defog )) and Discriminator (Discriminator_D defog ) is positioned between the generator and the de-fog image, thereby improving the visual quality of the defog image before it is passed to the discriminator.
[0078] This image enhancer (IE) corrects artifacts or unnatural parts that may occur during the fog removal process, and adjusts the defog image (DI) created by combining partial images (PI) of fog at different depths so that it looks like a single natural image.
[0079] The Image Enhancer (IE) receives an image processed by the generator as input and adjusts contrast, brightness, and color to improve overall image quality. It also blends the depth-processed subimages seamlessly. This allows the IE to enhance image clarity, improve color balance, and smooth out edges resulting from depth-processing.
[0080] Such image enhancers can be implemented using models based on convolutional neural networks (CNNs), models utilizing attention mechanisms, generative adversarial networks (GANs), reinforcement learning-based models, and multi-scale processing models.
[0081] An autoencoder (835) generates a segmentation image (SI) by segmenting the input depth image (FI) with reference to depth information. This autoencoder is an unsupervised learning neural network that compresses (encodes) and then restores (decodes) input data.
[0082] The autoencoder (835) receives a depth image as input and outputs a segmentation image divided according to each depth. The process is described in detail as follows.
[0083] The encoder receives a depth image as input and extracts features. For example, it learns high-dimensional features of the image through a multi-layer convolutional neural network (CNN). Based on the extracted features, the decoder generates a segmentation image of the original image size. For example, it can perform segmentation while gradually increasing the resolution using deconvolution layers. Through this process, each pixel of the input image is classified into one of several depth levels, and depth array information is created.
[0084] This segmentation result is then used to generate a depth-specific binary map, which in turn plays a crucial role in the depth-specific dehazing process. Segmentation allows us to accurately determine the depth at which each part of the image belongs, allowing us to apply appropriate dehazing intensity to each region.
[0085]
[0086] Hereinafter, each step of the image restoration method in a foggy environment using an artificial neural network model is described. The detailed description is as follows, focusing on the operation of the artificial neural network model described above.
[0087]
[0088] S10 stage
[0089] At this stage, there is no direct neural network processing. The fog image is fed into an artificial neural network model and prepared as input data for the neural networks in the next stage.
[0090] Here, fog images may or may not include depth information. For example, if a fog image is acquired through a 2D camera, there is no depth information and only an RGB image exists. However, if it is acquired through a 3D camera, the fog image includes depth information and is input to the artificial neural network model.
[0091]
[0092] S20 stage
[0093] This step is a process of generating a depth-based segmentation image, and is a process performed by a depth adversarial generator (831) and an autoencoder (835) among artificial neural network models.
[0094] The depth adversarial generator (831) predicts depth information of fog based on pre-learning from the input fog image when the input fog image does not contain depth information.
[0095] In addition, the depth adversarial generator (831) creates a depth image (FI) reflecting depth information from an input fog image (RI) based on fog depth information. In the embodiment, rather than simply applying uniform processing to the entire image, customized fog removal that takes into account the structure of the scene is possible.
[0096] The depth image (FI) generated by the depth adversarial generator (831) is input to the auto-encoder (835) and serves as the basis for segmenting the fog image based on depth. When the depth image is input, the auto-encoder (835) extracts features of the image through the encoder, and segments the image according to the depth of the fog through the decoder to generate a segmentation image (SI). The auto-encoder operates to determine segmentation by selecting the depth level with the highest probability for each pixel.
[0097]
[0098] S30 stage
[0099] This step is the process of generating depth-level defog images (DEI). Here, depth-level defog images (DEI) are created for the number of depth levels contained in the input image. For example, if the input fog image (RI) contains four levels of fog depth information, a total of four depth-level defog images (DEI) are created, one for each depth level.
[0100] When an artificial neural network model predicts a defog image, the part (segmentation) corresponding to each depth level is the most reliable. For example, in Fig. 1, the first level-based defog image (DEI) exemplifies that the segmentation corresponding to green is the most reliable among the segmentations, and the second level-based defog image (DEI) exemplifies that the segmentation corresponding to blue is the most reliable. In this case, in the embodiment, only the restored partial images corresponding to green are selected in the first level-based defog image (DEI), and only the restored partial images corresponding to blue are selected in the second level-based defog image (DEI), and the defog image (DI) is generated by combining the partially selected partial images, so that more robust fog removal is possible.
[0101] In the embodiment, a binary map (BM) is used to selectively select only a desired partial image (belonging to a specific level). This binary map can be created using a segmentation image, where pixels corresponding to a specific depth range are represented as 1 and the remaining pixels as 0, thereby creating a binary map for a specific level. Such a binary map is created separately for each depth level and applied to the defog image (DEI) for each depth level to selectively create only the partial image (PI) corresponding to the segmentation of each depth.
[0102]
[0103] S40 stage
[0104] This step is the process of generating the final defog image (DI) through selective combination.
[0105] This process can be done primarily with computational logic rather than direct neural network processing, and uses segmentation images to mask and combine the defog results for each depth to generate a defog image (DI).
[0106] This step may, more specifically, include the steps of: i) generating a binary map for each depth level based on the segmentation image, ii) applying the generated binary map for each depth level to a defog image of a corresponding depth level among the defog images for each depth level generated in the second step to generate a partial image for each depth, and iii) combining the partial images for each depth generated in the step ii) to generate the defog image.
[0107]
[0108] Finally, the defog image (DI) created by the depth of the fog is processed by the Discriminator D defog ) is transmitted and compared with the actual input image to determine whether it is true or false and output it.
[0109]
[0110] Figure 5 is a block diagram illustrating a computing device (800) that executes the image restoration method described above in a foggy environment. It reconstructs a series of processing steps according to the restoration method described above from the perspective of hardware configuration. Therefore, to avoid redundancy in explanation, only an outline of the functions and operations of each component will be provided.
[0111] The computing device (800) is configured to include a memory (810) that stores an artificial intelligence model (830) trained to segment a fog image by depth level and selectively combine only partial images of the segmented depth levels to generate a defog image, and a processor (820) that performs a series of operations to remove fog based on the artificial intelligence model, wherein the artificial intelligence model receives a fog image as input, generates a segmentation image representing the distribution of segmentation for each depth from a depth image generated based on depth information of the fog, generates a defog image for each depth level of the fog in the input fog image based on the segmentation image, and selectively combines only partial images corresponding to the segmentation of each depth from the defog image for each depth level to generate a defog image of the fog image.
[0112]
[0113] Meanwhile, the image restoration method in a foggy environment of the above-described embodiment can be implemented as computer-readable code on a computer-readable recording medium. A computer-readable recording medium includes any type of recording device that stores data that can be read by a computer system.
[0114] Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Furthermore, computer-readable recording media can be distributed across network-connected computer systems, allowing computer-readable code to be stored and executed in a distributed manner. Furthermore, functional programs, codes, and code segments for implementing the present invention can be readily inferred by programmers in the technical field to which the present invention pertains.
[0115] The present invention has been described above, focusing on various embodiments thereof. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.
Claims
1. In the artificial neural network model, Step 1: Inputting a fog image; A second step of generating a segmentation image representing the distribution of segmentation for each depth from a depth image created based on the depth information of the above fog; A third step of generating a defog image from which fog has been removed for each depth level of fog in the input fog image based on the segmentation image; A fourth step of generating a defog image of the fog image by selectively combining only partial images corresponding to segmentation of each depth from the defog images for each depth level; A method for image restoration in a foggy environment including .
2. In paragraph 1, The above artificial neural network model is, A depth adversarial generator is provided that receives the fog image as input, generates a depth image representing the distribution of fog at each depth, and determines the authenticity of the generated depth image. In the second step, the depth image is generated based on the fog image input by the depth adversarial generator. Image restoration method in foggy environments.
3. In paragraph 1, A method for restoring images in a foggy environment, wherein the depth information of the fog is obtained through a 3D camera.
4. In paragraph 2, The loss function of the above-mentioned deep adversarial generator ( )Is, And, = predicted depth image, = A method for restoring images in foggy environments, which are real depth images.
5. In paragraph 1, The second step is executed by an autoencoder, which is an unsupervised learning neural network that compresses (encodes) input data and then restores (decodes) it. Image restoration method in foggy environments.
6. In paragraph 5, The loss function of the above autoencoder ( )Is, , x = input depth image, = segmentation image generated based on x, = The correct answer for the generated segmentation image, Image restoration method in foggy environments.
7. In paragraph 1, The fourth step above is, i) Generate a binary map for each depth level based on the above segmentation image, ii) Applying the binary map generated for each depth level to the defog image of the corresponding depth level among the defog images generated for each depth level in the second step to generate a partial image for each depth, iii) Creating the defog image by combining the partial images for each depth generated in the above ii) process, Image restoration method in foggy environments.
8. In paragraph 1, The above artificial neural network model is, A defog adversarial generator is provided that receives the depth array information of the segmentation image and the fog image as input to generate the defog image and determines the authenticity of the generated defog image. The third and fourth steps are performed by the defog adversarial generator. Image restoration method in foggy environments.
9. In paragraph 8, The above defog adversarial generator is, It includes a generator that generates the defog image, and a discriminator that compares the defog image generated by the generator with an actual image to determine whether it is genuine. The loss function of the above generator ( )Is, And, = Output image of the defog adversarial generator, = actual value of defog image, λ = loss weight, r : loss weight complement (1 - λ), Image restoration method in foggy environments.
10. In paragraph 9, A method for restoring an image in a foggy environment, wherein the above loss weight is a value greater than the above loss weight compensation.
11. In paragraph 9, The above artificial neural network model is, Further comprising an image enhancer for improving the visual quality of the generated defog image, The output of the image enhancer is input to the discriminator, Image restoration method in foggy environments.
12. In paragraph 11, The loss function of the above discriminator ( )Is, And, = Defog image for patch i, = A method for restoring images in a foggy environment, which is an actual image corresponding to patch i.
13. A memory storing an artificial intelligence model that implements an image restoration method in a foggy environment as described in any one of paragraphs 1 to 12; and A processor that drives the above artificial intelligence model; A computing device including a .
14. A recording medium having recorded thereon a computer-readable program coded to perform the image restoration method in a foggy environment as described in any one of paragraphs 1 to 12.
Citation Information
Patent Citations
Apparatus for Removing Fog and Improving Visibility in Unit Frame Image and Driving Method Thereof
KR102151750B1
KR20220088075A