Generative adversarial network for transformer-based defogging

By adopting a transformer-based cyclic generative adversarial network in the image defog model, using paired and unpaired image samples to decompose the defog path, the problems of inaccurate prior estimation and scarcity of samples in the prior art are solved, and more efficient image defog and re-atomization effects are achieved.

CN120106145APending Publication Date: 2025-06-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411778577.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-09
Filing Date
2024-12-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems with inaccuracy of priori estimating the transmission pattern and scarcity of paired image samples in the image defogging process, resulting in a defogging performance.

Method used

Using a transformer-based cyclic generation adversarial network (GAN), the defogging and re-atomization model is trained, and the defogging path is decomposed into depth and density calculations using paired and unpaired image samples, thereby improving the remote spatial dependence and depth estimation performance of the model.

Benefits of technology

Reduced dependence on paired image samples, improved performance of the defogging model, especially in the re-atomization path, achieving better visual reconstruction and detail recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106145A_ABST
    Figure CN120106145A_ABST
Patent Text Reader

Abstract

A method and system for performing image defogging, comprising: obtaining an input image; estimating the transmission map by providing the input image to a defogging converter model trained by performing a training process on a cyclically generative adversarial network (GAN) including the defogging converter model; and generating an output image based on the transmission map, in which an amount of haze included in the output image is smaller than an amount of haze included in the input image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 606,851, filed on December 6, 2023, the disclosure of which is incorporated herein by reference in its entirety as if fully set forth herein. Technical Field

[0003] The present disclosure relates generally to image dehazing. More specifically, the subject matter disclosed herein relates to a transformer-based dehazing model that is trained as part of a recurrent generative adversarial network based on paired and unpaired image samples. Background Art

[0004] When an image is captured by a camera (e.g., a camera of an electronic device such as a mobile device), the image may be affected by haze. Haze may refer to the result of light scattering by particles in the atmosphere. Therefore, the electronic device may perform dehazing to estimate a haze-free image. The haze-free image may have improved visual quality, which may be beneficial for computer vision tasks such as image segmentation and object detection.

[0005] To perform dehazing, some prior-based methods involve estimating the transmission map by investigating the prior. However, in practice, the prior may be easily violated, which may lead to degradation of dehazing performance.

[0006] Other approaches may involve deep neural network (DNN) dehazing models that rely on datasets that include hazy image samples paired with clean image samples. However, due to the scarcity of paired clean hazy image samples, some DNN dehazing models have focused on using synthetic hazy image samples for training, which may result in a domain gap between the images output by the DNN dehazing models and the input images on which they are based.

[0007] Other approaches may involve convolutional neural network (CNN) encoder / decoder architectures. However, these architectures may have limited receptive fields, which may lead to performance disadvantages as the dehazing task may require long-range spatial dependencies. Summary of the invention

[0008] To overcome these problems, the systems and methods described herein involve a recurrent generative model for training dehazing and rehazing models based on both paired and unpaired image samples. In addition, embodiments may use visual transformers as image generators in the dehazing and rehazing paths of the training process, and may decompose the dehazing path into depth and density calculations.

[0009] The above method improves upon previous methods because the recurrent generative model can reduce the number of paired clean and foggy image samples required for training. Furthermore, the visual transformer used in the above method can provide improved long-range spatial dependencies between the paired clean and foggy image samples used, and can provide improved depth estimation performance, which can result in better performance in the rehazing path. Furthermore, the above method can provide spatially consistent reconstructions due to the decomposition of the dehazing path into depth and density. Furthermore, by using both paired and unpaired samples, the above method can provide improved reconstructions of details while providing strong constraints on dehazing.

[0010] In an embodiment, a method for performing image dehazing includes: obtaining an input image; estimating a transmission map by providing the input image to a dehazing transformer model, wherein the dehazing transformer model is trained by performing a training process on a recurrent generative adversarial network (GAN) including the dehazing transformer model; and generating an output image based on the transmission map, wherein an amount of haze included in the output image is less than an amount of haze included in the input image.

[0011] In one embodiment, a system for performing image dehazing includes: a training module configured to perform a training process on a recurrent generative adversarial network (GAN) including a dehazing transformer model; and an electronic device configured to: obtain an input image; estimate a transmission map by providing the input image to the dehazing transformer model; and generate an output image based on the transmission map, wherein an amount of haze included in the output image is less than an amount of haze included in the input image. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In the following sections, various aspects of the subject matter disclosed herein will be described with reference to exemplary embodiments shown in the accompanying drawings, in which:

[0013] Figure 1 is a block diagram illustrating an image defogging system according to an embodiment.

[0014] Figure 2 is a block diagram illustrating a training module for training a recurrent generative adversarial network including a dehazing transformer model, according to an embodiment.

[0015] Figure 3is a block diagram illustrating a defogging network included in a defogging-refogging path according to an embodiment.

[0016] Figure 4 is a block diagram illustrating a defogging network included in a refogging-defogging path according to an embodiment.

[0017] Figure 5 is a block diagram illustrating a re-atomization network included in a re-atomization-desmearing path according to an embodiment.

[0018] Figure 6 is a block diagram illustrating a defogging network included in a refogging-defogging path according to an embodiment.

[0019] Figure 7 is a flow chart of a process for performing image dehazing according to an embodiment.

[0020] Figure 8 is a flow chart of a process for training a recurrent generative adversarial network including a dehazing transformer model, according to an embodiment.

[0021] Fig. 9 is a block diagram of an electronic device in a network environment according to an embodiment.

[0022] Fig.10 A system including a UE and a gNB communicating with each other is shown. DETAILED DESCRIPTION

[0023] In the following detailed description, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, those skilled in the art will appreciate that the disclosed aspects may be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits are not described in detail to avoid obscuring the subject matter disclosed herein.

[0024] References throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment may be included in at least one embodiment disclosed herein. Therefore, the phrases "in one embodiment" or "in an embodiment" or "according to an embodiment" (or other phrases with similar meanings) that appear in various places throughout this specification may not necessarily all refer to the same embodiment. In addition, the particular features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In this regard, as used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" should not be interpreted as necessarily being preferred or advantageous over other embodiments. In addition, the particular features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In addition, depending on the context discussed herein, a singular term may include a corresponding plural form, and a plural term may include a corresponding singular form. Similarly, hyphenated terms (e.g., "two-dimensional", "pre-determined", "pixel-specific", etc.) may occasionally be used interchangeably with corresponding non-hyphenated versions (e.g., "two dimensional", "predetermined", "pixelspecific", etc.), and capitalized terms (e.g., "Counter Clock", "Row Select", "PIXOUT", etc.) may be used interchangeably with corresponding non-capitalized versions (e.g., "counterclock", "row select", "pixout", etc.). Such occasional interchangeable usages should not be considered inconsistent with each other.

[0025] In addition, depending on the context of the discussion herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. It should also be noted that the various drawings (including component drawings) shown and discussed herein are for illustrative purposes only and are not drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. In addition, if deemed appropriate, reference numerals are repeated in the drawings to indicate corresponding and / or similar elements.

[0026] The terms used herein are only used for the purpose of describing some example embodiments and are not intended to limit the claimed subject matter. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include" and / or "comprise" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] It should be understood that when an element or layer is referred to as being on, "connected to" or "coupled to" another element or layer, it can be directly on, connected to or coupled to another element or layer, or there can be intermediate elements or layers. In contrast, when an element is referred to as being "directly on," "directly connected to" or "directly coupled to" another element or layer, there are no intermediate elements or layers. The same reference numerals always represent the same elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0028] As used herein, the terms "first," "second," and the like are used as labels for the nouns that follow them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless explicitly so defined. In addition, the same reference numerals may be used across two or more figures to refer to parts, components, blocks, circuits, units, or modules having the same or similar functions. However, this usage is only for simplicity of illustration and ease of discussion; it does not mean that the construction or architectural details of such components or units are the same in all embodiments, or that such commonly referenced parts / modules are the only way to implement some example embodiments disclosed herein.

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter belongs. It will be further understood that terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and will not be interpreted in an idealized or overly formal sense unless explicitly so defined herein.

[0030] As used herein, the term "module" refers to any combination of software, firmware, and / or hardware configured to provide the functionality described herein in conjunction with the module. For example, software may be embodied as a software package, code, and / or instruction set or instructions, and the term "hardware" as used in any embodiment described herein may include, for example (alone or in any combination) components, hardwired circuits, programmable circuits, state machine circuits, and / or firmware storing instructions executed by programmable circuits. Modules may be collectively or individually embodied as circuits that form part of a larger system, such as, but not limited to, an integrated circuit (IC), a system on a chip (SoC), a component, etc.

[0031] Figure 1 is a block diagram showing an image defogging system according to an embodiment. Figure 1 As shown, the image defogging system 10 may include an electronic device 100 and a training module 200 .

[0032] The electronic device 100 may include a processor 101, a memory 102, a camera module 103, and a defogging module 104, which may include a defogging transformer model 1041 and an image generation module 1042. The processor 101 may control the overall operation of the electronic device 100, for example, based on instructions stored in the memory 102. As discussed above, when an image is captured, for example, using the camera module 103, the image quality of the captured image may be degraded due to the presence of fog, which may be the result of scattering of light by particles in the atmosphere. The presence of fog may produce visually unappealing images, and it may also be harmful to computer vision tasks such as image segmentation and object detection. Therefore, in some embodiments, the electronic device 100 may perform image defogging to improve the image quality of the captured image. For example, the electronic device 100 may use the defogging module 104 to perform an image defogging process, which may include estimating a clean image with improved visual quality. In an embodiment, a clean image may refer to a fog-free image, or an image that reduces or eliminates the effects of fog.

[0033] Image quality degradation caused by the influence of haze can be expressed by Koschmieder's law, which can be expressed according to the following equations 1 and 2:

[0034] Equation 1

[0035] Equation 2

[0036] In Equation 1 and Equation 2 above, can represent the real foggy image, which can be called the observed foggy image, can represent global atmospheric light, It can represent the transmission diagram, can represent the true clean image, which can be called the observed clean image or scene radiance, can represent a depth map, and can represent the density coefficient, which can be called the medium attenuation coefficient. In addition, can be called direct attenuation, and It can be called airlight.

[0037] Therefore, for real foggy images , the electronic device 100 can use the defogging module 104 to estimate the global atmospheric light , and estimate the transmission function using the dehazing transformer model 1041 , and can be based on the estimated global atmospheric light and the estimated transmission function The estimated true clean image is generated using the image generation module 1042 using equation 1 above .

[0038] In some embodiments, the dark channel Can be used to estimate global atmospheric light For example, in some embodiments, the dark channel The top 0.1% brightest pixels in can be selected. Among these pixels, the real foggy image The pixel with the highest brightness in the image can be considered as the global atmospheric light. Dark Channel It can be defined according to the following equation 3:

[0039] (Equation 3)

[0040] In equation 3 above, Can represent real foggy images The color channels and Can be expressed as The local patch centered on the

[0041] In some embodiments, the real foggy image The image may be an image captured using the camera module 103, but the embodiment is not limited thereto. For example, in some embodiments, a real foggy image The image may be an image received by the electronic device 100 from another device and stored in the memory 102 , for example.

[0042] The training module 200 may include a defogging network 201, a refogging network 202, and a discriminator 203. In an embodiment, at least one of the defogging network 201, the refogging network 202, and the discriminator 203 may be or may include a visual transformer, but the embodiment is not limited thereto. In some embodiments, the discriminator 203 may include a plurality of discriminators (e.g., Figure 2 In an embodiment, the defogging network 201 may include a defogging transformer model 1041, or, for example, a defogging module 104 including a defogging transformer model 1041. The defogging network 201 may be used to apply a defogging transformation to transform a foggy image H (e.g., a real foggy image I) into a clean image. , which can be expressed as The rehazing network 202 can be used to apply a rehazing transformation to transform the clean image (For example, real clean images ) is transformed into a foggy image , which can be expressed as .

[0043] In an embodiment, the training module 200 may be used to train the re-fogging network 202 and the defogging network 201 including the defogging transformer model 1041. For example, the training module 200 may train the re-fogging network 202 and the defogging network 201 by, for example, calculating or otherwise determining one or more losses using the discriminator 203, and may adjust the re-fogging network 202 and the defogging network 201 based on the calculated losses, for example, by adjusting the weights of the re-fogging network 202 and the defogging network 201. Although the training module 200 is shown as being separated from the electronic device 100, the embodiment is not limited thereto. For example, in some embodiments, the training module 200 may be included in the electronic device 100, or may be included in a separate device (e.g., a server device with which the electronic device 100 can communicate).

[0044] The dehazing problem can be an underdetermined problem because it may involve, for example, equation and unknown number, among which Indicates the size of the image. Due to the scarcity of paired clean and foggy samples, some image dehazing methods (e.g., deep dehazing models) may be trained using synthetic foggy data, which may lead to overfitting.

[0045] In contrast, embodiments relate to a recurrent generative adversarial network (GAN) for training the defogging network 201 and the refogging network 202, which may allow training to be performed using both paired image samples and unpaired image samples. Thus, embodiments may avoid overfitting while using a relatively small number of paired foggy and clean image samples.

[0046] Figure 2 2 is a block diagram illustrating a training module for training a dehazing transformer model according to an embodiment. In an embodiment, the training module 200 may be arranged as or may otherwise include a cycle GAN, which may receive paired hazy and clean image samples ( ) to improve training by providing strong constraints for defogging and reduce the possibility of under-constrained problems. The training module 200 can perform training according to two branches or paths, which can be respectively referred to as the defogging-refogging path and the refogging-defogging path.

[0047] For ease of description, Figure 2 The two paths are shown as comprising separate elements. For example, Figure 2 The defogging-refogging path is illustrated as including a defogging network 201A, a refogging network 202A, and a discriminator 203A, and the refogging-defogging path is illustrated as including a refogging network 202B, a defogging network 201B, and a discriminator 203B. However, the embodiments are not limited thereto. For example, the defogging network 201A may represent a first instance of the defogging network 201, or a first use of the defogging network 201 (e.g., when the defogging-refogging path is executed), and the defogging network 201B may represent a second instance of the defogging network 201, or a second use of the defogging network 201 (e.g., when the refogging-defogging path is executed). Similarly, the refogging networks 202A and 202B may represent different instances or uses of the refogging network 202, and the discriminators 203A and 203B may represent different instances or uses of the discriminator 203.

[0048] Therefore, if Figure 2 As shown, in the defogging-refogging path, the defogging network 201A can be used to transform the real foggy image Transformed into the first synthetic clean image , and the first synthesized clean image can be transformed using the re-haze network 202A Transformed into the first synthetic foggy image In the rehazing-dehazing path, the rehazing network 202B can be used to transform the real clean image (It can be compared with the real foggy image Pairing) is transformed into the second synthetic foggy image , and the defogging network 201B can be used to transform the second synthetic foggy image Transformed into the second synthetic clean image .

[0049] In the dehazing-rehazing path, the discriminator 203A can be based on the real clean image Generating discriminator output , and the clean image can be synthesized based on the first Generating discriminator output In the rehazing-dehazing path, the discriminator 203B can be based on the real hazy image Generating discriminator output And based on the second synthetic foggy image Generating discriminator output .

[0050] In an embodiment, the training module 200 may perform training by calculating and minimizing a training loss L. The training loss L may be determined based on one or more different component losses. For example, the training loss L may be based at least in part on a cycle consistency loss can be determined according to the following equation 4:

[0051] Equation 4

[0052] In an embodiment, the cycle consistency loss Can be used to impose conditions on how intermediate images transmitted from one domain can be transmitted back.

[0053] The training loss L can also be based at least in part on the pairing loss can be determined according to the following equation 5:

[0054] Equation 5

[0055] The training loss L may also be determined based at least in part on the adversarial loss corresponding to the discriminator 203 (or the discriminators 203A and 203B). For example, the mapping function and and their corresponding discriminator outputs to determine the adversarial loss.

[0056] As an example, corresponding to the mapping function The first confrontation loss It can be expressed according to the following equation 6:

[0057] Equation 6

[0058] As another example, corresponding to the mapping function The second confrontation loss It can be expressed according to the following equation 7:

[0059] Equation 7

[0060] In an embodiment, the training loss L may be based at least in part on the first adversarial loss and the second adversarial loss One or two of the above should be used to determine the answer.

[0061] Figure 3 is a block diagram showing a defogging network included in a defogging-refogging branch according to an embodiment. Figure 3 As shown, the defogging network 201A may include the defogging module 104 discussed above, and the defogging module 104 includes a defogging transformer model 1041 and an image generation module 1042. The defogging network 201A may also include a density estimation module 301 and a depth estimation module 302.

[0062] In an embodiment, the dehaze transformer model 1041A may be or may include a UNet pixel-by-pixel visual transformer (UNet-ViT) model, and the density estimation module may include a convolutional neural network (CNN), but the embodiment is not limited thereto.

[0063] In an embodiment, the defogging network 201 may receive a real foggy image And the real foggy image can be Provided to the defogging module 104 and the density estimation module 301. The defogging module 104 can generate an estimated global atmospheric light , and the dehazing transformer model 1041 can be used to generate an estimated transmission function The defogging module 104 can also use the image generation module 1042 to generate a first synthetic clean image The density estimation module 301 can generate a first estimated density The depth estimation module 302 may be based on the first estimated density and the estimated transmission function Generate estimated dehazed depth map .

[0064] Figure 4 is a block diagram showing a defogging network included in a refogging-defogging branch according to an embodiment. Figure 4 As shown, the re-haze network 202A may include a transformer-based depth estimation module 401 and an image generation module 402. The transformer-based depth estimation module 401 may be based on the first synthesized clean image Generate estimated rehazed depth map , the image generation module 402 can be based on the estimated re-fogging depth map and the first estimated density Generate the first synthetic foggy image .

[0065] Figure 5 is a block diagram showing a re-atomization network included in a re-atomization-de-atomization branch according to an embodiment. Figure 5 As shown, the re-haze network 202B may include a transformer-based depth estimation module 401 and an image generation module 402. The transformer-based depth estimation module 401 may be based on a real clean image. Generate estimated rehazed depth map , the image generation module 402 can be based on the estimated re-fogging depth map and randomly generated densities Generate a second synthetic foggy image .

[0066] Figure 6 is a block diagram showing a defogging network included in a refogging-defogging branch according to an embodiment. Figure 6 As shown, the defogging network 201B may include the defogging module 104 discussed above, including a defogging transformer model 1041 and an image generation module 1042. The defogging network 201B may also include a density estimation module 301 and a depth estimation module 302.

[0067] In an embodiment, the defogging network 201B may receive a second synthetic foggy image And the second synthetic foggy image can be Provided to the defogging module 104 and the density estimation module 301. The defogging module 104 can generate an estimated global atmospheric light , and the dehazing transformer model 1041 can be used to generate an estimated transmission function The defogging module 104 can also use the image generation module 1042 to generate a second synthetic clean image The density estimation module 301 can generate a second estimated density The depth estimation module 302 may be based on the second estimated density and the estimated transmission function Generate estimated dehazed depth map .

[0068] In an embodiment, the training loss L may be based at least in part on the density loss can be determined according to the following equation 8:

[0069] Equation 8

[0070] In an embodiment, the density loss Can be used to apply a second estimated density Should be consistent with random density The matching condition.

[0071] In an embodiment, the training loss L may be based at least in part on the depth loss and can be determined according to the following equations 9 and 10:

[0072] Equation 9

[0073] Equation 10

[0074] In an embodiment, the depth loss and Can be used to impose conditions on which depth maps along each path should match.

[0075] In an embodiment, the training loss L may be calculated as a weighted sum of the four loss functions discussed above, as shown in Equation 11 below:

[0076] + + Equation 11

[0077] In Equation 11 above, the weight , and may represent training hyperparameters that may be used to tune the training performed by the training module 200 .

[0078] Figure 7 is a flow chart of a process for performing image dehazing according to an embodiment. In some implementations, Figure 7 One or more processing blocks of may be performed by one or more of the elements discussed above, such as one or more of the image defogging system 10 and any elements included therein, such as the electronic device 100 .

[0079] like Figure 7 As shown, at operation 701, process 700 may include obtaining an input image.

[0080] like Figure 7As further shown in , at operation 702, the process 700 may include estimating a transmission map by providing the input image to a dehazing transformer model. In an embodiment, the dehazing transformer model may correspond to the dehazing transformer model 1041 included in the dehazing module 104 discussed above. In an embodiment, the dehazing transformer model may be trained by performing a training process on a cyclic GAN including the dehazing transformer model, for example using the training module 200 discussed above. In an embodiment, the training process may be performed using a plurality of unpaired samples and a plurality of paired samples, and the plurality of paired samples may include a plurality of real hazy images paired with a plurality of real clean images.

[0081] like Figure 7 As further shown in FIG. 7 , at operation 703, process 700 may include generating an output image based on the transmission map, wherein the amount of haze included in the output image is less than the amount of haze included in the input image. In an embodiment, the output image may be generated by the defogging module 104 using the image generation module 1042 based on Equation 1 discussed above.

[0082] Figure 8 is a flowchart of a process for training a recurrent GAN including a dehazing transformer model according to an embodiment. In some embodiments, Figure 8 One or more processing blocks of may be performed by one or more of the elements discussed above, such as one or more of the image dehazing system 10 and any elements included therein, such as the training module 200 .

[0083] like Figure 8 As shown, at operation 801 , process 800 may include obtaining a real foggy image from a plurality of real foggy images.

[0084] like Figure 8 As further shown in FIG. 8 , at operation 802 , process 800 may include providing the real foggy image to a defogging network to obtain a first synthetic clean image. In an embodiment, the defogging network may correspond to the defogging network 201A discussed above.

[0085] like Figure 8 As further shown in FIG. 8 , at operation 803 , process 800 may include providing the first synthetic clean image to a rehazing network to obtain a first synthetic foggy image. In an embodiment, the rehazing network may correspond to the rehazing network 202A discussed above.

[0086] like Figure 8 As further shown in FIG. 8 , at operation 804 , process 800 may include providing the real clean image to a re-hazing network to obtain a second synthetic hazy image. In an embodiment, the re-hazing network may correspond to the re-hazing network 202B discussed above.

[0087] like Figure 8 As further shown in FIG. 8 , at operation 805 , process 800 may include providing the second synthetic foggy image to a defogging network to obtain a second synthetic clean image. In an embodiment, the defogging network may correspond to the defogging network 201B discussed above.

[0088] like Figure 8 As further shown in FIG. 8 , at operation 806, process 800 may include determining a training loss based on the real foggy image, the first synthetic clean image, the first synthetic foggy image, the real clean image, the second synthetic foggy image, and the second synthetic clean image. In an embodiment, the training loss may be determined based on one or more of equations 3 to 11 discussed above.

[0089] like Figure 8 As further shown in FIG. 8 , at operation 807 , process 800 may include adjusting at least one of the dehazing network and the rehazing network based on the training loss. In an embodiment, the adjusting may include adjusting a weight of at least one of the dehazing network and the rehazing network.

[0090] although Figure 7-8 Example blocks of processes 700 and 800 are shown, but in some implementations, processes 700 and 800 may include Figure 7-8 The blocks in the process 700 and 800 may be additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in the process 700 and 800. Additionally or alternatively, two or more blocks in the blocks of the process 700 and 800 may be performed in parallel.

[0091] Thus, embodiments relate to a dehazing transformer model that can be trained using a transformer-based cycle GAN, which can use a visual transformer in both the dehazing path and the rehazing path, such as in at least one of a dehazing network, a rehazing network, and a discriminator included in the cycle GAN. In addition to unpaired image samples, paired image samples can be used to train the transformer-based cycle GAN.

[0092] Thus, due to the relatively large receptive field provided by the visual transformer, embodiments can provide improved long-range spatial dependencies on clean and foggy images. Furthermore, due to the training using both paired and unpaired samples, embodiments can provide improved clean image reconstruction in terms of vision, as well as improved detail reconstruction.

[0093] Fig. 9 is a block diagram of electronic devices in a network environment 900 according to an embodiment.

[0094] refer to Fig. 9, an electronic device 901 in a network environment 900 may communicate with an electronic device 902 via a first network 998 (e.g., a short-range wireless communication network), or communicate with an electronic device 904 or a server 908 via a second network 999 (e.g., a long-range wireless communication network). The electronic device 901 may communicate with the electronic device 904 via the server 908. The electronic device 901 may include a processor 920, a memory 930, an input device 950, a sound output device 955, a display device 960, an audio module 970, a sensor module 976, an interface 977, a haptic module 979, a camera module 980, a power management module 988, a battery 989, a communication module 990, a subscriber identification module (SIM) card 996, or an antenna module 997. In one embodiment, at least one of the components (e.g., the display device 960 or the camera module 980) may be omitted from the electronic device 901, or one or more other components may be added to the electronic device 901. Some of the components may be implemented as a single integrated circuit (IC). For example, the sensor module 976 (e.g., a fingerprint sensor, an iris sensor, or an illumination sensor) may be embedded in the display device 960 (e.g., a display). In some embodiments, the electronic device 901 may correspond to the electronic device 100, the processor 920 may correspond to the processor 101, the memory 930 may correspond to the memory 102, and the camera module 980 may correspond to the camera module 103 discussed above, however, the embodiment is not limited thereto.

[0095] The processor 920 may execute software (eg, program 940 ) to control at least one other component (eg, hardware or software component) of the electronic device 901 coupled to the processor 920 , and may perform various data processing or calculations.

[0096] As at least part of data processing or calculation, the processor 920 may load commands or data received from another component (e.g., the sensor module 976 or the communication module 990) into the volatile memory 932, process the commands or data stored in the volatile memory 932, and store the resulting data in the non-volatile memory 934. The processor 920 may include a main processor 921 (e.g., a central processing unit (CPU) or an application processor (AP)) and an auxiliary processor 923 (e.g., a graphics processing unit (GPU), an image signal processor (ISP), a sensor hub processor, or a communication processor (CP)) that may operate independently of the main processor 921 or in conjunction with the main processor 921. Additionally or alternatively, the auxiliary processor 923 may be adapted to consume less power than the main processor 921, or to perform a specific function. The auxiliary processor 923 may be implemented to be separate from the main processor 921 or to be a part of the main processor 921.

[0097] The auxiliary processor 923 may replace the main processor 921 when the main processor 921 is in an inactive (e.g., sleep) state, or control at least some of the functions or states related to at least one of the components of the electronic device 901 (e.g., the display device 960, the sensor module 976, or the communication module 990) together with the main processor 921 when the main processor 921 is in an active state (e.g., executing an application). The auxiliary processor 923 (e.g., an image signal processor or a communication processor) may be implemented as a part of another component (e.g., a camera module 980 or a communication module 990) that is functionally related to the auxiliary processor 923.

[0098] The memory 930 may store various data used by at least one component of the electronic device 901 (e.g., the processor 920 or the sensor module 976). The various data may include, for example, input data or output data of software (e.g., program 940) and commands related thereto. The memory 930 may include a volatile memory 932 or a non-volatile memory 934. The non-volatile memory 934 may include an internal memory 936 and / or an external memory 938.

[0099] The program 940 may be stored as software in the memory 930 , and may include, for example, an operating system (OS) 942 , middleware 944 , or an application 946 .

[0100] The input device 950 may receive a command or data to be used by another component (eg, the processor 920) of the electronic device 901 from outside (eg, a user) of the electronic device 901. The input device 950 may include, for example, a microphone, a mouse, or a keyboard.

[0101] The sound output device 955 can output sound signals to the outside of the electronic device 901. The sound output device 955 may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as playing multimedia or recording, and the receiver may be used to receive incoming calls. The receiver may be implemented as being separate from the speaker or as a part of the speaker.

[0102] The display device 960 can visually provide information to the outside of the electronic device 901 (e.g., a user). The display device 960 may include, for example, a display, a hologram device, or a projector, and a control circuit for controlling a corresponding one of the display, the hologram device, and the projector. The display device 960 may include a touch circuit adapted to detect a touch or a sensor circuit (e.g., a pressure sensor) adapted to measure the strength of a force caused by a touch.

[0103] The audio module 970 can convert sound into an electrical signal and vice versa. The audio module 970 can obtain sound via the input device 950 or output sound via the sound output device 955 or an earphone of an external electronic device 902 directly (eg, wired) or wirelessly coupled to the electronic device 901.

[0104] The sensor module 976 may detect an operating state (e.g., power or temperature) of the electronic device 901 or an environmental state (e.g., a user's state) outside the electronic device 901, and then generate an electrical signal or data value corresponding to the detected state. The sensor module 976 may include, for example, a gesture sensor, a gyroscope sensor, an atmospheric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illumination sensor.

[0105] The interface 977 may support one or more designated protocols for the electronic device 901 to be directly (e.g., wired) or wirelessly coupled with the external electronic device 902. The interface 977 may include, for example, a high-definition multimedia interface (HDMI), a universal serial bus (USB) interface, a secure digital (SD) card interface, or an audio interface.

[0106] The connection terminal 978 may include a connector via which the electronic device 901 can be physically connected to the external electronic device 902. The connection terminal 978 may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (eg, a headphone connector).

[0107] The haptic module 979 may convert an electrical signal into a mechanical stimulus (eg, vibration or movement) or an electrical stimulus that can be recognized by a user via a tactile sense or a kinesthetic sense. The haptic module 979 may include, for example, a motor, a piezoelectric element, or an electrical stimulator.

[0108] The camera module 980 may capture still images or moving images. The camera module 980 may include one or more lenses, an image sensor, an image signal processor, or a flash. The power management module 988 may manage the power supplied to the electronic device 901. The power management module 988 may be implemented as at least a portion of a power management integrated circuit (PMIC), for example.

[0109] The battery 989 may supply power to at least one component of the electronic device 901. The battery 989 may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0110] The communication module 990 may support establishing a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device 901 and an external electronic device (e.g., electronic device 902, electronic device 904, or server 908), and perform communication via the established communication channel. The communication module 990 may include one or more communication processors that may operate independently of the processor 920 (e.g., AP) and support direct (e.g., wired) communication or wireless communication. The communication module 990 may include a wireless communication module 992 (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module 994 (e.g., a local area network (LAN) communication module or a power line communication (PLC) module). A corresponding one of these communication modules can communicate with an external electronic device via a first network 998 (e.g., a short-range communication network such as BLUETOOTH™, Wireless Fidelity (Wi-Fi) Direct, or Infrared Data Association (IrDA) standards) or a second network 999 (e.g., a long-range communication network such as a cellular network, the Internet, or a computer network (e.g., a LAN or a Wide Area Network (WAN)). These various types of communication modules may be implemented as a single component (e.g., a single IC), or may be implemented as multiple components (e.g., multiple ICs) separated from each other. The wireless communication module 992 may identify and authenticate the electronic device 901 in a communication network (such as the first network 998 or the second network 999) using subscriber information (e.g., an International Mobile Subscriber Identity (IMSI)) stored in the subscriber identification module 996.

[0111] The antenna module 997 may transmit or receive a signal or power to or from the outside of the electronic device 901 (e.g., an external electronic device). The antenna module 997 may include one or more antennas, and at least one antenna suitable for a communication scheme used in a communication network such as the first network 998 or the second network 999 may be selected, for example, by the communication module 990 (e.g., the wireless communication module 992). Then, a signal or power may be transmitted or received between the communication module 990 and the external electronic device via the selected at least one antenna.

[0112] Commands or data may be sent or received between the electronic device 901 and the external electronic device 904 via a server 908 coupled to the second network 999. Each of the electronic devices 902 and 904 may be a device of the same type or a different type as the electronic device 901. All or some operations to be performed at the electronic device 901 may be performed at one or more of the external electronic devices 902, 904, or 908. For example, if the electronic device 901 should automatically or in response to a request from a user or another device to perform a function or service, the electronic device 901 may request one or more external electronic devices to perform at least a portion of the function or service instead of performing the function or service, or in addition to performing the function or service, the electronic device 901 may request one or more external electronic devices to perform at least a portion of the function or service. The one or more external electronic devices receiving the request may perform at least a portion of the requested function or service, or an additional function or additional service related to the request, and transmit the result of the execution to the electronic device 901. The electronic device 901 may provide the result as at least a part of the reply to the request with or without further processing the result. To this end, for example, cloud computing, distributed computing, or client-server computing technology may be used.

[0113] Fig.10 A system including a UE 1005 and a gNB 1010 in communication with each other is shown. The UE may include a radio 815 and a processing circuit (or means for processing) 1020, which may perform various methods disclosed herein, for example, Figure 2 and 8 For example, the processing circuit 1020 may receive a transmission from the network node (gNB) 1010 via the radio 1015, and the processing circuit 1020 may send a signal to the gNB 1010 via the radio 1015.

[0114] The embodiments of the subject matter and operations described in this specification may be implemented in digital electronic circuits, or in computer software, firmware or hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, which are encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or additionally, program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or may be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. In addition, although a computer storage medium is not a propagation signal, a computer storage medium may be a source or destination of computer program instructions encoded in an artificially generated propagation signal. The computer storage medium may also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or be included therein. In addition, the operations described in this specification may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0115] Although this specification may include many specific implementation details, the implementation details should not be interpreted as a limitation on the scope of any claimed subject matter, but rather as a description of features specific to a particular embodiment. Certain features described in the context of a separate embodiment in this specification may also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. In addition, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from the claimed combination may be deleted from the combination in some cases, and the claimed combination may be directed to a sub-combination or a variation of the sub-combination.

[0116] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that the operations be performed in the particular order shown or in sequence, or that all of the operations shown be performed, to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0117] Thus, specific embodiments of the subject matter have been described herein. Other embodiments are within the scope of the following claims. In some cases, the actions set forth in the claims can be performed in a different order and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.

[0118] As those skilled in the art will recognize, the innovative concepts described herein may be modified and varied over a wide range of applications. Accordingly, the scope of claimed subject matter should not be limited to any of the specific exemplary teachings discussed above, but is instead defined by the appended claims.

Claims

1. A method for performing image defogging, the method comprising: Get the input image; estimating a transmission map by providing the input image to a dehazing transformer model, wherein the dehazing transformer model is trained by performing a training process on a recurrent generative adversarial network (GAN) including the dehazing transformer model; and generating an output image based on the transmission map, The amount of haze included in the output image is smaller than the amount of haze included in the input image.

2. The method according to claim 1, wherein: The training process is performed using a plurality of unpaired samples and a plurality of paired samples, The plurality of paired samples include a plurality of real foggy images paired with a plurality of real clean images, and The training loss of the training process is determined based on at least one of cycle loss, pairing loss, GAN loss, density loss and depth loss.

3. The method according to claim 2, wherein: The cycle GAN includes a transformer-based dehazing network and a transformer-based rehazing network, and Wherein, the defogging transformer model is included in the defogging network.

4. The method according to claim 3, wherein: The training process includes: Obtain a real foggy image from the multiple real foggy images; providing the real foggy image to the dehazing network to obtain a first synthetic clean image; and The first synthetic clean image is provided to the re-haze network to obtain a first synthetic hazy image.

5. The method according to claim 4, wherein: The circulation loss is determined based on a difference between the first synthetic foggy image and the real foggy image, wherein the pairing loss is determined based on a difference between the first synthetic clean image and a real clean image among the plurality of real clean images corresponding to the real foggy image, and The GAN loss is determined based on the first synthetic foggy image and the real foggy image.

6. The method according to claim 5, wherein: The training process also includes: providing the true clean image to the rehazing network to obtain a second synthetic hazy image; and The second synthetic hazy image is provided to the dehazing network to obtain a second synthetic clean image.

7. The method according to claim 6, wherein: The cycle loss is further determined based on a difference between the second synthesized clean image and the true clean image, wherein the pairing loss is further determined based on a difference between the second synthetic foggy image and the real foggy image, and The GAN loss is further determined based on the second synthesized clean image and the real clean image.

8. The method according to claim 6, wherein: The defogging network also includes a convolutional neural network (CNN). The second synthetic foggy image is generated by the re-fogging network based on a random density coefficient. The training process further includes generating an estimated density coefficient corresponding to the second synthetic foggy image using the CNN, and The density loss is determined based on a difference between the estimated density coefficient and the random density coefficient.

9. The method according to claim 6, in, The training process also includes: generating an estimated transmission map based on at least one of the real foggy image and the second synthetic foggy image using the dehaze transformer model; generating an estimated dehazed depth map based on the estimated transmission map; and generating an estimated rehazed depth map based on at least one of the real clean image and the first synthetic clean image using the rehazed network, and The depth loss is determined based on a difference between the estimated dehazed depth map and the estimated rehazed depth map.

10. A system for performing image dehazing, the system comprising: A training module configured to perform a training process on a recurrent generative adversarial network (GAN) including a dehazing transformer model; and An electronic device configured to: Get the input image; estimating a transmission map by providing the input image to the dehaze transformer model; and generating an output image based on the transmission map, The amount of haze included in the output image is smaller than the amount of haze included in the input image.

11. The system according to claim 10, wherein: The training process is performed using a plurality of unpaired samples and a plurality of paired samples, The plurality of paired samples include a plurality of real foggy images paired with a plurality of real clean images, and The training loss of the training process is determined based on at least one of cycle loss, pairing loss, GAN loss, density loss and depth loss.

12. The system according to claim 11, wherein: The cycle GAN includes a transformer-based defogging network and a transformer-based refogging network. Wherein, the defogging transformer model is included in the defogging network.

13. The system according to claim 12, wherein: In order to perform the training process, the training module is further configured to: Obtain a real foggy image from the multiple real foggy images; Providing the real foggy image to the defogging network to obtain a first synthetic clean image; and The first synthetic clean image is provided to the re-hazy network included in the cycle GAN to obtain a first synthetic hazy image.

14. The system according to claim 13, wherein: The circulation loss is determined based on a difference between the first synthetic foggy image and the real foggy image, wherein the pairing loss is determined based on a difference between the first synthetic clean image and a real clean image among the plurality of real clean images corresponding to the real foggy image, and The GAN loss is determined based on the first synthetic foggy image and the real foggy image.

15. The system of claim 14, wherein: In order to perform the training process, the training module is further configured to: providing the true clean image to the rehazing network to obtain a second synthetic hazy image; and The second synthetic hazy image is provided to the dehazing network to obtain a second synthetic clean image.

16. The system of claim 15, wherein: The cycle loss is further determined based on a difference between the second synthesized clean image and the true clean image, wherein the pairing loss is further determined based on a difference between the second synthetic foggy image and the real foggy image, and The GAN loss is determined based on the second synthesized clean image and the real clean image.

17. The system of claim 15, wherein: The defogging network also includes a convolutional neural network (CNN). The second synthetic foggy image is generated by the re-fogging network based on a random density coefficient. Wherein, in order to perform the training process, the training module is further configured to generate an estimated density coefficient corresponding to the second synthetic foggy image using the CNN, and The density loss is determined based on a difference between the estimated density coefficient and the random density coefficient.

18. The system according to claim 15, in, In order to perform the training process, the training module is further configured to: generating an estimated transmission map based on at least one of the real foggy image and the second synthetic foggy image using the dehaze transformer model; generating an estimated dehazed depth map based on the estimated transmission map; and generating an estimated rehazed depth map based on at least one of the real clean image and the first synthetic clean image using the rehazed network, and The depth loss is determined based on a difference between the estimated dehazed depth map and the estimated rehazed depth map.

19. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of a device for performing image dehazing, cause the at least one processor to: Get the input image; The transmission map is estimated by feeding the input image to a dehazing transformer model, where training the dehazing transformer model by performing a training process on a recurrent generative adversarial network (GAN) including the dehazing transformer model; and generating an output image based on the transmission map, The amount of haze included in the output image is smaller than the amount of haze included in the input image.

20. The non-transitory computer readable medium of claim 19, wherein: The training process is performed using a plurality of unpaired samples and a plurality of paired samples, The plurality of paired samples include a plurality of real foggy images paired with a plurality of real clean images, and The training loss of the training process is determined based on at least one of cycle loss, pairing loss, GAN loss, density loss and depth loss.