Method and apparatus for processing image data, storage medium, and electronic device

Through unsupervised learning of the target restore model, combined with background light and transmitted light data, the image processing model is optimized, which solves the problem of low accuracy in transmission noise removal and achieves efficient image clarity improvement in complex environments.

CN119850460BActive Publication Date: 2025-07-08ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510318725.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

In the prior art, the image processing model lacks generalization ability when removing transmitted noise, resulting in poor practical application effects and limited improvement in image clarity.

Method used

The target reduction model is used to train unsupervised learning, and uses background light extraction branches, transmitted light extraction branches and target output branches to generate clear sample images, optimize model parameters, and improve the accuracy of transmittance noise removal.

Benefits of technology

Without pairing clean images, efficient and accurate image transmission noise is removed, which significantly improves image clarity and information extraction capabilities, and is suitable for image processing in complex scattering environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850460B_ABST
    Figure CN119850460B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for processing image data, a storage medium, and an electronic device. Among them, the method includes: obtaining a degraded image; inputting the degraded image into a target restoration model to obtain a clear image, where the target restoration model is obtained by training an initial restoration model using a first sample image. The initial restoration model includes a background light extraction branch, a transmitted light extraction branch, and a target output branch. The target output branch is used to remove the transmitted noise of the first sample image, generate initial image data, and remove the transmitted noise of a second sample image determined by the background light data, the transmitted light data, and the initial image data to obtain a clear sample image. The clear sample image is used to adjust the model parameters of the initial restoration model to obtain the target restoration model. The present application solves the technical problem of low accuracy in removing the transmitted noise of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and more particularly, to a method and apparatus for processing image data, a storage medium, and an electronic device. Background Art

[0002] In the process of image processing, supervised learning methods are often used to train an image processing model. Then, the image processing model is used to remove the transmission noise of the original acquired image. Due to the difference between the sample images used in the training process and the real image acquisition environment, the generalization ability of the model is insufficient, and the actual application effect is limited. Therefore, in the actual application scenario, the effect of removing the transmission noise of the original image is not good, and the improvement of the image clarity is limited.

[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present application provide a method and apparatus for processing image data, a storage medium, and an electronic device to at least solve the technical problem of low accuracy in removing the transmission noise of an image.

[0005] According to one aspect of the embodiments of the present application, a method for processing image data is provided, including: obtaining a degraded image, where the degraded image represents an image with transmission noise; inputting the degraded image into a target restoration model to obtain a clear image, where the target restoration model is obtained by training an initial restoration model using a first sample image, the initial restoration model includes a background light extraction branch, a transmitted light extraction branch, and a target output branch, the background light extraction branch is used to extract background light data of the first sample image, the transmitted light extraction branch is used to extract transmitted light data of the first sample image, the target output branch is used to remove the transmission noise of the first sample image, generate initial image data, and remove the transmission noise of a second sample image determined by the background light data, the transmitted light data, and the initial image data to obtain a clear sample image, and the clear sample image is used to adjust model parameters of the initial restoration model to obtain the target restoration model.

[0006] According to another aspect of the embodiments of the present application, there is also provided an image data processing device, including: an acquisition module, configured to acquire a degraded image, where the degraded image represents an image with transmission noise; a processing module, configured to input the degraded image into a target restoration model to obtain a clear image, where the target restoration model is obtained by training an initial restoration model using a first sample image, the initial restoration model includes a background light extraction branch, a transmission light extraction branch, and a target output branch, the background light extraction branch is configured to extract background light data of the first sample image, the transmission light extraction branch is configured to extract transmission light data of the first sample image, the target output branch is configured to remove the transmission noise of the first sample image, generate initial image data, and remove the transmission noise of a second sample image determined by the background light data, the transmission light data, and the initial image data to obtain a clear sample image, and the clear sample image is used to adjust the model parameters of the initial restoration model to obtain the target restoration model.

[0007] Optionally, the device is further configured to: acquire the first sample image; input the first sample image into the initial restoration model to obtain the background light data, the transmission light data, and the initial image data, where the transmission light data represents image data with noise added to the depth image data of the first sample image, the transmission light data is used to indicate the light transmission state of the first sample image, and the background light data is used to indicate the image brightness of the first sample image; generate a second sample image according to the background light data, the transmission light data, and the initial image data; use the target output branch to remove the transmission noise of the second sample image to obtain the clear sample image, where the target output branch includes a plurality of convolutional layers; calculate a training loss value based on the clear sample image, and determine the initial restoration model as the target restoration model when the training loss value is less than or equal to a loss threshold.

[0008] Optionally, the device generates the second sample image according to the background light data, the transmission light data, and the initial image data in the following manner: determine the product of the initial image data and the transmission light data as the first data; determine the product of the difference between the transmission light data and the constant 1 and the background light data as the second data; and determine the second sample image according to the first data and the second data.

[0009] Optionally, the device is further configured to: perform a depth information extraction operation on the first sample image using a depth extraction module in the transmitted light extraction branch to obtain depth image data, and determine initial transmitted light data according to the depth image data and the scattering coefficient; add Gaussian noise to the initial transmitted light data using an adapter module in the transmitted light extraction branch to obtain the transmitted light data; and generate the second sample image based on the transmitted light data.

[0010] Optionally, the device is further configured to: perform a Fourier transform operation on the first sample image using the background light extraction branch to obtain the background light data, where the background light extraction branch includes a plurality of convolutional layers; and generate the second sample image based on the background light data.

[0011] Optionally, the device is further configured to: establish a target loss function based on the first sample image, the second sample image, and the clear sample image; determine the training loss value using the target loss function; and determine the initial restoration model as the target restoration model when the training loss value is less than or equal to a loss threshold.

[0012] Optionally, the device is configured to establish a target loss function based on the background light data, the transmitted light data, and the initial image data in at least one of the following manners: establish a first loss function based on the mean square error of the image data of the first sample image and the image data of the second sample image, where the target loss function includes the first loss function; establish a second loss function based on the mean square error of the image data of the clear sample image and the image data obtained after performing a gradient stopping operation on the initial image data, where the target loss function includes the second loss function; establish a third loss function using the partial derivative of the initial transmitted light data, where the initial transmitted light data is determined by the depth image data and the scattering coefficient of the first sample image, and the depth image data is obtained by performing a depth information extraction operation on the first sample image using a depth extraction module in the transmitted light extraction branch, and the target loss function includes the third loss function.

[0013] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the above-mentioned method for processing image data when running.

[0014] According to another aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for processing image data as described above.

[0015] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the method for processing the above-mentioned image data through the computer program.

[0016] In the embodiments of the present application, a degraded image is first obtained, and then the degraded image is input into a target restoration model to obtain a clear image, and the target restoration model is obtained by performing unsupervised learning training on an initial restoration model. Specifically, in the training process of the target restoration model, a first sample image is first used to generate a second sample image, and transmission noise is removed from the second sample image to obtain a clear sample image. In this process, the generation of the second sample image synthesizes the background light data, transmitted light data, and initial image data of the first sample image. The background light data reflects the brightness of the image, and the transmitted light data is obtained by adding Gaussian noise to the depth image data of the first sample image to simulate the light transmission situation in a real environment, while the initial image data is data obtained by directly performing a convolution operation on the first sample image to preliminarily remove transmission noise.

[0017] Next, by calculating the training loss value between the clear sample image and the first sample image, the parameters of the initial restoration model are optimized, so that the target restoration model can more accurately restore a clear image without transmission noise from the degraded image. That is, by designing a reasonable unsupervised training process and loss function, the purpose of improving the accuracy of removing transmission noise is achieved, so that the technical effect of efficiently and accurately removing image transmission noise can be realized without paired clean images, and further the technical problem of low accuracy of removing transmission noise in the field of image enhancement is solved.

[0018] In summary, the target restoration model in the embodiments of the present application significantly improves the accuracy of removing transmission noise, providing more powerful technical support for applications such as defogging of foggy images and enhancement of underwater images. Description of the Drawings

[0019] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0020] Figure 1 is a schematic diagram of an application environment of an optional method for processing image data according to an embodiment of the present application;

[0021] Figure 2 is a schematic flowchart of an optional method for processing image data according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of an optional method for processing image data according to an embodiment of the present application;

[0023] Figure 4 is a schematic structural diagram of an optional device for processing image data according to an embodiment of the present application;

[0024] Figure 5 is a schematic structural diagram of an optional product for processing image data according to an embodiment of the present application;

[0025] Figure 6 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0026] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] The present application will be described below with reference to the embodiments:

[0029] According to one aspect of the embodiments of the present application, a method for processing image data is provided. Optionally, in this embodiment, the above method for processing image data can be applied to, for exampleFigure 1 In the hardware environment composed of the server 101 and the terminal device 103 as shown. As Figure 1 shown, the server 101 is connected to the terminal device 103 through a network and can be used to provide services for the terminal device or the application 107 installed on the terminal device. The application can be a video application, an instant messaging application, a browser application, an educational application, a game application, etc. A database 105 can be set up on the server or independently of the server to provide data storage services for the server 101. For example, a game data storage server. The above network can include but is not limited to: a wired network, a wireless network. Among them, the wired network includes: a local area network, a metropolitan area network, and a wide area network. The wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The terminal device 103 can be a terminal configured with an application and can include but is not limited to at least one of the following: a mobile phone (such as an Android mobile phone, an iOS mobile phone, etc.), a laptop computer, a tablet computer, a handheld computer, a MID (Mobile Internet Devices, mobile Internet device), a PAD, a desktop computer, a smart TV, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a virtual reality (Virtual Reality, VR for short) terminal, an augmented reality (Augmented Reality, AR for short) terminal, a mixed reality (Mixed Reality, MR for short) terminal, etc. computer devices. The above server can be a single server, or a server cluster composed of multiple servers, or a cloud server.

[0030] Combined with Figure 1 shown, the above method for processing image data can be executed by an electronic device, and the electronic device can be a terminal device or a server. The above method for processing image data can be implemented by the terminal device or the server respectively, or jointly implemented by the terminal device and the server.

[0031] The above is only an example, and this embodiment does not make specific limitations.

[0032] Optionally, as an alternative implementation, as Figure 2 shown, the above method for processing image data includes:

[0033] S202, obtaining a degraded image, where the degraded image represents an image with transmission noise;

[0034] Optionally, in the embodiments of the present application, the above degraded images include but are not limited to foggy day surveillance images, underwater exploration images, endoscopic medical images, or any other images obtained under harsh lighting conditions. That is, the quality of the degraded images is damaged due to the intervention of scattering media such as fog, water, or dust.

[0035] It should be noted that the above degraded images can be sourced in various ways. They can be directly from a camera or sensor, or generated through post - processing or simulation. The specific acquisition method is not limited to a single approach. For example, they can be obtained by software - simulating the scattering effect, adding artificial noise, or through a specific image degradation process. This application does not limit it.

[0036] It should also be noted that the embodiments of this application do not limit the types of the above - mentioned degraded images either. They can include images in natural environments, and can also cover images in special scenarios such as industrial and medical scenarios. As long as the image is affected by transmission noise, it can be used as input and denoised using the method in the embodiments of this application.

[0037] In addition, there are also various possibilities for the size, format, resolution, etc. of the degraded images. As long as they can be accepted by the target restoration model, they are also not restricted.

[0038] S204. Input the degraded image into the target restoration model to obtain a clear image. Here, the target restoration model is obtained by training an initial restoration model using a first sample image. The initial restoration model includes a background light extraction branch, a transmitted light extraction branch, and a target output branch. The background light extraction branch is used to extract the background light data of the first sample image. The transmitted light extraction branch is used to extract the transmitted light data of the first sample image. The target output branch is used to remove the transmission noise of the first sample image, generate initial image data, and remove the transmission noise of a second sample image determined by the background light data, transmitted light data, and initial image data to obtain a clear sample image. The clear sample image is used to adjust the model parameters of the initial restoration model to obtain the target restoration model.

[0039] Optionally, in the embodiments of this application, the above - mentioned target restoration model refers to an optimized deep - learning model that can effectively restore clear images from images with transmission noise, including but not limited to models based on convolutional neural networks (CNNs), generative adversarial networks (GANs), or self - supervised learning frameworks. By unsupervised - learning the intrinsic features of the degraded images, especially the relationships between the transmitted light data, background light data, and initial image data, the ability to remove transmission noise is improved.

[0040] It should be noted that the above - mentioned first sample image can be sourced from different scenarios, including but not limited to foggy environments, underwater environments, dusty environments, or any other imaging conditions with scattering effects, or generated through post - processing or simulation. For example, it can be obtained by software - simulating the scattering effect, adding artificial noise, or through a specific image degradation process. This application does not limit it.

[0041] Furthermore, the above-mentioned background light extraction branch can be used to perform a background light extraction operation on the first sample image to obtain background light data, the transmitted light extraction branch can be used to perform a transmitted light extraction operation on the first sample image to obtain transmitted light data, and the target output branch can be used to perform a direct transmitted noise removal operation on the first sample image to obtain initial image data.

[0042] Among them, the above-mentioned background light extraction branch and target output branch can both be composed of simple convolutional layers, Norm, and the activation function Sigmoid, while the transmitted light extraction branch can include a DepthModel (pre-trained large model, depth extraction module) and an Adapter (adapter module). Specifically, open-source models such as DINO and DepthAnything can be used. The initial transmitted light data can be determined based on the depth image data of the first sample image and the scattering coefficient. Furthermore, the adapter module can be used to perform a data augmentation operation on the initial transmitted light data, such as adding noise, etc., so that the data augmentation operation of the transmitted light data can more flexibly adapt to different scattering environments and enhance the robustness and generalization ability of the target restoration model.

[0043] That is to say, the background light extraction branch is used to directly output the scattered light data of the first sample image, while the transmitted light extraction branch, through the combination of the depth extraction module and the adapter, not only uses depth information to estimate the transmitted light, but also can increase the flexibility and adaptability of the target restoration model through data augmentation operations, enabling it to effectively perform image restoration in the face of various scattering environments.

[0044] Optionally, in the embodiment of the present application, the above-mentioned initial image data refers to the image data directly extracted from the first sample image using the target data branch during the training stage of the initial restoration model and having part of the transmitted noise removed through a simplified convolutional operation. In other words, the above-mentioned initial image data refers to the image data obtained after the initial restoration model makes a preliminary denoising attempt on the first sample image without the guidance of depth information and background light data, reflecting the preliminary estimation of the target restoration model for removing image transmitted noise in the initial stage of unsupervised training.

[0045] Similarly, the embodiment of the present application does not limit the specific calculation methods of the background light data, transmitted light data, and initial image data involved in the generation process of the above-mentioned second sample image. For example, the extraction of the background light data is based on the statistics of local brightest pixels, the noise addition of the transmitted light data follows a specific Gaussian distribution, and the generation of the initial image data can adopt a combination of various convolutional layers and activation functions. The present application also does not limit this, to ensure the diversity and adaptability of the method.

[0046] Specifically, the above-mentioned second sample image can be determined according to the following formula:

[0047] I''(x) = J(x)T'(x) + (1 - T'(x))A(x);

[0048] Among them, I''(x) represents the second sample image, J(x) represents the initial image data, T'(x) represents the transmitted light data, A(x) represents the background light data. The second sample image represents a re-degraded image obtained based on the initial image data, the transmitted light data, and the background light data. At the end of model training, this re-degraded image satisfies the similarity condition with the first sample image. For example, the similarity between the re-degraded image and the first sample image is as high as 98%. Or, calculate the mean absolute error or mean square error of the pixel values between the re-degraded image and the first sample image to ensure that the two are numerically close to a preset threshold. Or use the structural similarity index, peak signal-to-noise ratio, feature vector similarity, etc. to evaluate the similarity between the two.

[0049] Next, after obtaining the second sample image, input the second sample image into the above-mentioned target output branch again to remove the transmitted noise from the second sample image using the above-mentioned target output branch, thereby obtaining a clear sample image without noise and with high clarity. This clear sample image represents the training result of the target restoration model in an unsupervised learning framework, which is not only visually closer to the true state of the original scene, that is, a clear image captured without the interference of a scattering medium, but also performs better in terms of image structure information, color fidelity, and detail retention.

[0050] It should also be noted that in the embodiments of the present application, multiple loss functions at different levels can be established for joint calculation to constrain model learning from different angles, so that the finally trained model can not only remove the scattering effect, but also maintain the naturalness and detail integrity of the image. Specifically:

[0051] S1, the re-degradation loss, that is, the second sample image reconstructed according to the atmospheric degradation scattering equation should be consistent with the first sample degraded image, and the first loss function L reson is designed to achieve this constraint. The formula is as follows:

[0052] ; where, I(x) represents the input first sample image, and I''(x) represents the second sample image.

[0053] S2, image secondary augmentation. Based on the image J(x) (the above-mentioned initial image data) processed for the first time, a new clean image J'(x) (the above-mentioned clear sample image) is regenerated. This component should be consistent, and the above-mentioned second loss function L sim is designed to achieve this constraint. In addition, considering the need to avoid falling into a degenerate solution, sg, that is, gradient stopping operation, can be adopted. The formula is as follows:

[0054] ; where J(x) represents the above-mentioned initial image data, and J'(x) represents the above-mentioned clear sample image.

[0055] S3. According to the power operation law of the Transmission-Map and considering the gradient consistency loss, the above-mentioned third loss function L is designed accordingly. tmap To implement this constraint, the formula is expressed as follows:

[0056] L tmap = ; where T(x) represents the initial transmitted light data directly processed by the above-mentioned target output branch, and the corresponding T'(x) represents the above-mentioned transmitted light data.

[0057] That is to say, the scattering coefficient is a constant, the partial derivative of the transmitted light data with respect to the depth image data is a constant, and after the training of the initial restoration model ends, its second derivative is 0.

[0058] Furthermore, according to the first loss function, the second loss function, and the third loss function, the target loss function is determined to be L total = L sim + L reson + L tmap .

[0059] Further, when the target loss function satisfies the preset loss value output condition, the training of the initial restoration model is successful. Here, the preset loss value output condition refers to that during the model training process, the value of the target loss function reaches or is lower than a preset threshold or standard, which usually means that the training of the model has converged, and the model parameters can effectively process the input data and achieve the expected performance indicators. For example, at least one of the loss function threshold, the loss function change rate, the validation set performance, the number of iterations, or the training time.

[0060] In an exemplary embodiment, during the unsupervised training process of the initial restoration model, the initial restoration model receives a first sample image, which also comes from a foggy environment. The initial restoration model generates a second sample image according to the first sample image. This generation process can be regarded as a preliminary defogging attempt of the model on the first sample image. The generated second sample image retains a similar structure to the first sample image, but there may be deficiencies or deviations in removing the transmitted noise.

[0061] Next, in order to obtain a clear sample image, the initial restoration model further processes the second sample image and inputs the second sample image into the target output branch again to remove the transmission noise therein. The second sample image is generated depending on the background light data, the transmission light data, and the initial image data. Specifically, the background light data reflects the atmospheric background brightness in the first sample image, the transmission light data is simulated by adding specific noise to the depth image data of the first sample image, and the initial image data is obtained by removing part of the transmission noise from the first sample image through preliminary convolutional operations. This processing provides preliminary guidance and clues for subsequent more in-depth defogging operations.

[0062] Through the above processing, the initial restoration model can finally generate the clear sample image, i.e., the above second sample image. The second sample image, the clear sample image, is visually closer to the real scene in a fog-free state and has been significantly improved in terms of structural information, color accuracy, and detail retention.

[0063] After obtaining the second sample image, the clear sample image, calculate the training loss value of the initial restoration model using this second sample image, the clear sample image. This loss value reflects the gap between the defogging effect of the model and the ideal effect. Adjust its parameters according to this feedback to gradually improve the denoising performance, and finally evolve into the target restoration model, which can effectively recover clear images from degraded images in foggy days or other scattering effect scenarios, greatly improving the efficiency and quality of image processing, and having significant application value in fields such as traffic monitoring, environmental monitoring, and outdoor photography.

[0064] In an exemplary embodiment, taking the unsupervised image defogging scenario based on deep learning as an example, first, the system obtains a degraded image from a surveillance camera in a foggy environment, inputs the input degraded image I(x) into the target restoration model, and directly processes it by the target output branch to output the above clear image, that is, the final defogged or underwater enhanced image.

[0065] In other words, in the actual application process, the target restoration model can only retain one of the above target output branches, while the background light extraction branch and the transmission light extraction branch in the initial restoration model are only used in the model training stage. That is, in the actual application process, the target restoration model only retains the target output branch and omits the background light extraction branch and the transmission light extraction branch. At this time, the data extraction capabilities of the background light extraction branch and the transmission light extraction branch have been internalized into the target output branch during the training stage of the initial restoration model. Therefore, there is no need to retain the background light extraction branch and the transmission light extraction branch in the target training model, thereby simplifying the calculation process of the actual application, improving the processing speed, and reducing the resource requirements at the same time.

[0066] Specifically, during the training phase of the initial restoration model, the background light extraction branch and the transmitted light extraction branch are used to establish interactive feedback with the target output branch, facilitating the learning and generalization of the initial restoration model on various foggy images. Through different paths and loss functions, the initial restoration model learns how to handle the fog effect in images without an explicit supervision signal, so that it can effectively remove the fog effect from unknown foggy images in practical applications. Once the initial restoration model is trained, the target output branch can independently process the degraded image and restore a clear image without explicitly using the background light extraction branch and the transmitted light extraction branch.

[0067] Therefore, in practical applications, the target restoration model includes the target output branch, while the background light extraction branch and the transmitted light extraction branch are optional, optimizing resource usage and reducing computational latency, making the target restoration model more suitable for running on real-time processing or resource-constrained devices.

[0068] Generally speaking, in the embodiment of this application, during the process of training the initial restoration model to determine the target restoration model, through mechanisms such as secondary augmentation and consistency loss, the target restoration model learns how to effectively remove the scattering effect between the given degraded image and component estimation, that is, the target output branch, the background light extraction branch, and the transmitted light extraction branch are collaboratively optimized to directly output a clear image. That is to say, by designing a reasonable unsupervised training process and loss function, the purpose of improving the accuracy of transmitted noise removal is achieved, thus realizing the technical effect of efficiently and accurately removing the transmitted noise of the image without paired clean images, and further solving the technical problem of low accuracy of transmitted noise removal in the field of image enhancement.

[0069] In summary, the target restoration model in the embodiment of this application significantly improves the accuracy of removing transmitted noise, providing more powerful technical support for applications such as defogging of foggy images and enhancement of underwater images.

[0070] As an alternative, the method further includes: obtaining the above-mentioned first sample image; inputting the above-mentioned first sample image into the above-mentioned initial restoration model to obtain the above-mentioned background light data, the above-mentioned transmitted light data, and the above-mentioned initial image data, where the above-mentioned transmitted light data represents image data obtained by adding noise to the depth image data of the above-mentioned first sample image, the above-mentioned transmitted light data is used to indicate the light transmission state of the above-mentioned first sample image, and the above-mentioned background light data is used to indicate the image brightness of the above-mentioned first sample image; generating a second sample image based on the above-mentioned background light data, the above-mentioned transmitted light data, and the above-mentioned initial image data; using the above-mentioned target output branch to remove the transmitted noise of the above-mentioned second sample image to obtain the above-mentioned clear sample image, where the above-mentioned target output branch includes multiple convolutional layers; calculating a training loss value based on the above-mentioned clear sample image, and in the case where the above-mentioned training loss value is less than or equal to a loss threshold, determining the above-mentioned initial restoration model as the above-mentioned target restoration model.

[0071] It should be noted that there are various methods for extracting the above-mentioned background light data, which can be based on statistical analysis of local maximum brightness values, or can be achieved through frequency domain transformation, edge detection, or specific image preprocessing algorithms. This application does not limit this.

[0072] Similarly, the generation of the transmitted light data is not limited to adding Gaussian noise to the depth image data, and various types of noise can also be used, or even simulation data based on physical models, to better reflect the light transmission characteristics of the real atmosphere or underwater environment. The generation strategy of the above-mentioned initial image data is also diverse. In addition to convolutional operations, residual learning, skip connections, or specific denoising algorithms can also be used. This application also does not limit this.

[0073] In an exemplary embodiment, a first sample image (which can be multiple sample images) is obtained from a foggy traffic surveillance camera, and the first sample image is blurred due to fog scattering. Subsequently, the first sample image is input into an initial restoration model based on a deep learning architecture. First, the background light data, transmitted light data, and initial image data of the first sample image are obtained. Based on these data, the initial restoration model generates a second sample image, that is, an image that contains transmitted noise but has been preliminarily processed. Next, the target output branch in the initial restoration model processes the second sample image. This branch is composed of multiple convolutional layers and is specifically used to remove the transmitted noise. Through this processing, the initial restoration model generates a clear sample image, which is visually closer to the original fog-free state, and the clarity and details have been significantly improved.

[0074] Based on the clear sample images, the training loss value of the initial restoration model can be calculated, which reflects the gap between the denoising effect of the model and the ideal effect. If the training loss value is less than or equal to a pre-set loss threshold, that is, the defogging effect of the model has reached or exceeded the set performance standard, then the current initial restoration model can be determined as the final target restoration model, indicating that the initial restoration model has passed the optimization of the unsupervised training process and can achieve high-quality defogging on unseen foggy images.

[0075] Through the embodiments of the present application, an unsupervised training method is adopted to achieve the technical effect of effectively removing transmission noise from degraded images, and the purpose of improving image clarity and information extraction ability is achieved. This technical solution not only simplifies the preparation process of training data, but also improves the generalization ability of the target restoration model in complex scattering environments, and has a significant promotion effect on various applications in the field of image processing, such as traffic monitoring, underwater exploration, medical image analysis, etc.

[0076] As an alternative solution, generating the second sample image according to the above background light data, the above transmission light data, and the above initial image data includes: determining the product of the above initial image data and the above transmission light data as the first data; determining the product of the difference between the above transmission light data and the constant 1 and the above background light data as the second data; and determining the above second sample image according to the above first data and the above second data.

[0077] Optionally, in the embodiments of the present application, the above first data refers to the result data obtained by performing a per-pixel multiplication operation on the initial image data and the transmission light data, and this data reflects the degradation situation of each pixel point in the image under the influence of the scattering medium. It includes, but is not limited to, performing a dot product operation on the preliminarily processed image and the transmission light data generated by simulating the scattering environment to obtain a preliminary estimate of the degraded image.

[0078] It should be noted that the constant 1 in the embodiments of the present application can refer to a standardized all-one matrix, which is used to calculate the difference with the transmission light data, and the generated data reflects the transmittance information. This operation is multiplied by the background light data when generating the second data to simulate the fog-free image information in the atmospheric scattering equation. The present application does not limit this, and the constant 1 can be any form of standardized matrix, or even a specific numerical matrix matching the image size, to adapt to the scattering models and calculation requirements of different scenarios.

[0079] In an exemplary embodiment, first, a depth extraction module is used to estimate the depth image data of the first sample image. Noise is added to the depth image data through a data augmentation operation to obtain transmitted light data. Then, the initial image data is multiplied pointwise with the transmitted light data to obtain first data. At the same time, the difference between the transmitted light data and the constant 1 is calculated, and the result is multiplied with the background light data to generate second data. Finally, the sum of the first data and the second data is determined as the second sample image. This image, while retaining the original image structure, adds a simulated degradation effect based on depth information and a scattering model, providing rich transmitted noise examples for unsupervised training.

[0080] Through the embodiments of the present application, by performing a multiplication operation on the image data and the transmitted light data, the technical effect of reliably separating the scattered light and the transmitted light information from the degraded image is achieved, and the calculation of the loss function in the unsupervised learning framework is optimized, thereby achieving the purpose of improving the accuracy and robustness of the model in image dehazing and underwater image enhancement tasks. By introducing prior knowledge of depth information and a scattering model, the problem of the lack of clean image pairs as supervision signals in unsupervised training is effectively solved, providing a more reasonable training objective and constraint for the model, and thus improving the performance and quality of image restoration in a complex scattering environment.

[0081] As an alternative solution, the above method further includes: using the depth extraction module in the above transmitted light extraction branch to perform a depth information extraction operation on the above first sample image to obtain depth image data, and determining initial transmitted light data according to the above depth image data and the scattering coefficient; using the adapter module in the above transmitted light extraction branch to add Gaussian noise to the above initial transmitted light data to obtain the above transmitted light data.

[0082] Optionally, in the embodiments of the present application, the above depth information extraction operation refers to using a depth estimation model to extract depth image data related to the scene depth from the first sample image, including but not limited to using a pre-trained depth estimation network, a three-dimensional reconstruction algorithm, or a depth perception technology based on structured light. These methods can provide key depth prior information for subsequent transmitted noise removal.

[0083] Further, the following formula is used to determine the initial transmitted light data: ;

[0084] where d(x) is the depth image data and beta is the scattering coefficient (i.e., the scattering intensity).

[0085] After obtaining the initial transmitted light data, it can be input into the above adapter module to perform a data augmentation operation on the initial transmitted light data to add Gaussian noise to the initial transmitted light data to obtain the transmitted light data.

[0086] It should be noted that the data augmentation operation is not limited to power operations, but may also include scale transformation, rotation, shearing, color adjustment, or the introduction of other types of noise to increase the diversity and complexity of the training data. This application does not make any restrictions in this regard.

[0087] At the same time, the number and type of convolutional layers in the transmitted light extraction branch can also be adjusted according to specific application scenarios and model design requirements, including but not limited to using deeper network structures, introducing dilated convolutions, or combining attention mechanisms to improve the feature extraction ability and generalization performance of the model.

[0088] In an exemplary embodiment, first, a pre-trained depth estimation model is used to perform depth information extraction on the acquired first sample image to obtain initial depth image data. Subsequently, Gaussian noise is added to the initial depth image data to generate depth image data, aiming to simulate the depth uncertainty in the real scattering environment. Next, the transmitted light extraction branch performs a data augmentation operation on the depth image data, and transmitted light data is obtained through a power operation. This data augmentation operation can simulate the influence of the scattering medium on the propagation of light and provide a basis for generating the second sample image later. Based on the transmitted light data, background light data, and initial image data, the second sample image is generated. The degradation degree of the second sample image is controllable, and the transmitted noise characteristics are clear, providing high-quality training samples for model training.

[0089] Through the embodiments of this application, by adopting depth information extraction and data augmentation operations, the technical effect of accurately estimating transmitted light data from degraded images is achieved, and the loss calculation of the model in the unsupervised training process is optimized, thereby achieving the purpose of improving the defogging and underwater image enhancement performance of the model. By using the depth information of the sample image and image augmentation means, the robustness and accuracy of the target restoration model in dealing with complex scattering scenarios are effectively enhanced, opening up a new path for improving the image processing effect.

[0090] As an alternative solution, the method further includes: performing a Fourier transform operation on the first sample image using the above-mentioned background light extraction branch to obtain the background light data, where the background light extraction branch includes multiple convolutional layers; generating a second sample image based on the background light data.

[0091] Optionally, in the embodiments of this application, the background light extraction branch refers to a specific component in the model, which is used to process the first sample image and extract the background light data therefrom, that is, the luminance information in the image that is not affected by the scattering medium. It includes but is not limited to using a convolutional neural network to perform a frequency domain conversion operation, analyzing the luminance distribution of the image based on the Fourier transform, so as to obtain the background light data for the subsequent generation of the second sample image.

[0092] It should be noted that the method for obtaining background light data in this application is not limited to Fourier transform operations. It can also be through other frequency-domain or spatial-domain analysis techniques, such as wavelet transform, Laplace transform, or direct color space analysis. This application does not make any limitations in this regard, aiming to cover all calculation methods that can effectively evaluate the overall brightness of an image.

[0093] In addition, the number and type of convolutional layers in the background light extraction branch can also be adjusted according to actual needs to optimize the estimation accuracy of the model for background light data.

[0094] In an exemplary embodiment, first, the background light extraction branch in the initial restoration model is used to process the first sample image. This branch includes multiple convolutional layers and can analyze the frequency-domain features of the image based on Fourier transform operations to obtain the image brightness information. Through this processing, background light data is obtained, which reflects the brightness distribution of the image without the influence of the scattering medium.

[0095] Next, after obtaining this background light data, a second sample image is generated by combining the transmitted light data and the depth information background to simulate the atmospheric scattering or underwater scattering effect, so as to provide degraded image samples for model training, enabling the model to learn the skill of removing transmission noise under unsupervised conditions.

[0096] Through the embodiments of this application, using the background light extraction branch to process the image realizes the technical effect of accurately estimating the background light data from images in foggy or underwater environments, achieving the purpose of providing key prior information for the unsupervised training process, improving the training efficiency and effect of the image dehazing and underwater image enhancement models, and effectively enhancing the robustness and accuracy of the target restoration model.

[0097] As an alternative solution, the method further includes: establishing an objective loss function based on the first sample image, the second sample image, and the clear sample image; using the objective loss function to determine the training loss value; and when the training loss value is less than or equal to the loss threshold, determining the initial restoration model as the target restoration model.

[0098] Optionally, in the embodiments of this application, the above objective loss function refers to a quantization index used to evaluate the performance of the initial restoration model when processing scattering effect images. It is not limited to a single loss metric, including but not limited to mean square error (MSE), structural similarity index (SSIM), perceptual loss, or adversarial loss, etc. By comprehensively analyzing the differences between the clear sample image output by the model and the first sample image and the second sample image, it guides the optimization process of the model.

[0099] It should be noted that the calculation method of the training loss value can be flexibly changed in this application. It is not limited to the combination of loss functions mentioned above, but can also be other forms of error metrics or performance indicators, such as peak signal-to-noise ratio (PSNR), Jaccard similarity coefficient, or loss functions customized for specific tasks. The setting of the loss threshold also depends on the specific application scenario and the expected image restoration quality, which can be an empirical value or the optimal value determined through cross-validation. This application does not make any limitations in this regard.

[0100] In an exemplary embodiment, first, based on the first sample image collected in a foggy or underwater environment, a second sample image is generated by simulating the scattering effect. Subsequently, the second sample image is processed by the initial restoration model to obtain a preliminary clear sample image. Immediately afterwards, the system constructs an objective loss function containing multiple loss components according to the differences among the first sample image, the second sample image, and the clear sample image, for example, simultaneously considering MSE and SSIM, to comprehensively evaluate the restoration effect of the model. During the model training process, the model parameters are continuously optimized, and the value of the objective loss function, that is, the training loss value, is calculated to measure the current performance of the model. When the training loss value is less than or equal to the pre-set loss threshold, it indicates that the model has reached the expected performance standard. At this time, the initial restoration model is officially determined as the target restoration model that can effectively process scattering effect images.

[0101] Through the embodiments of this application, by adopting the loss function evaluation method based on multiple images, the technical effect of accurately training the initial restoration model under unsupervised conditions is achieved, and the purposes of optimizing the model performance, effectively removing or reducing the scattering effect, and improving the image clarity and information extraction ability are achieved.

[0102] As an alternative solution, an objective loss function is established based on the background light data, the transmitted light data, and the initial image data, including at least one of the following: a first loss function is established based on the mean square error of the image data of the first sample image and the image data of the second sample image, where the objective loss function includes the first loss function; a second loss function is established based on the mean square error of the image data of the clear sample image and the image data obtained after performing a gradient stopping operation on the initial image data, where the objective loss function includes the second loss function; a third loss function is established using the partial derivative of the initial transmitted light data, where the above initial transmitted light data is determined by the depth image data of the first sample image and the scattering coefficient, and the above depth image data is obtained by performing a depth information extraction operation on the first sample image through the depth extraction module in the transmitted light extraction branch, and the objective loss function includes the third loss function.

[0103] Optionally, in the embodiments of the present application, the above-mentioned target loss function refers to a set of functions used to quantify and guide the accuracy of image restoration during model training, including but not limited to the first loss function, the second loss function, and the third loss function. These functions work together to optimize the image de-scattering effect.

[0104] Specifically, obtain the first image vector corresponding to the above-mentioned first sample image and the second image vector corresponding to the above-mentioned second sample image; establish a first loss function using the square of the L2 norm of the above-mentioned first image vector and the above-mentioned second image vector, where the above-mentioned target loss function includes the above-mentioned first loss function; obtain the third image vector corresponding to the above-mentioned clear sample image, and perform a gradient stopping operation on the above-mentioned initial image data to obtain a fourth image vector; establish a second loss function using the square of the L2 norm of the above-mentioned third image vector and the above-mentioned fourth image vector, where the above-mentioned target loss function includes the above-mentioned second loss function.

[0105] It should be noted that the square of the L2 norm is used as part of the loss function to calculate the distance between image vectors, thereby measuring the consistency between the model prediction result and the target image. In the present application, this consistency evaluation is not limited to pixel-level comparison, but can also be extended to the feature vector level. For example, feature vectors generated by a pre-trained feature extraction network can be used to capture higher-level image structure information. In addition, the use of gradient stopping operations and partial derivative constraints aims to avoid information leakage during the training process and ensure the rationality of the relationship between transmitted light data and depth information. These operation methods can have multiple implementation paths in practical applications, and the present application does not limit this.

[0106] In an exemplary embodiment, when processing the first sample image, first use a depth information extraction operation to obtain initial depth image data, and then use the transmitted light extraction branch in the initial restoration model to perform a data augmentation operation to obtain initial transmitted light data. This process involves performing a power operation on the initial depth image data to adapt to the mathematical characteristics of the scattering model. Then, convert the first sample image into a first image vector, and the second sample image into a second image vector. By calculating the square of the L2 norm of these two vectors, a first loss function is constructed to evaluate the pixel-level difference between the second sample image generated by the model and the original image. Similarly, obtain the feature representation of the clear sample image, that is, the third image vector, and perform a gradient stopping operation on the initial image data to generate a fourth image vector. Based on the square of the L2 norm of the third image vector and the fourth image vector, a second loss function is constructed to supervise the consistency between the clear image output by the model and the image estimated through depth information. Finally, the system uses the partial derivative of the initial transmitted light data and the depth image data to establish a third loss function to ensure that the estimation of the transmitted light data is consistent with the change trend of the depth information and avoid estimation bias during the training process.

[0107] Through the embodiments of the present application, a set of loss functions including the square of the L2 norm, gradient stopping operation, and partial derivative loss are adopted to effectively train and optimize the image de-scattering model under unsupervised conditions, achieving the purpose of improving the effects of image de-hazing and underwater image enhancement, while ensuring the generalization ability and robustness of the model. By reasonably constructing multi-dimensional loss constraints, the problems of unclear objectives and insufficient constraints in unsupervised learning are effectively solved, providing strong technical support for image processing in complex scattering environments.

[0108] In an exemplary embodiment, the method for processing image data proposed in the present application can be applied to the field of video image processing technology. For example, it can be applied to scenarios such as haze image enhancement related to scattering effects and underwater image enhancement. Among them, image scattering effect removal generally refers to image de-hazing and underwater image enhancement, which have a wide range of application requirements in the industrial field, such as traffic monitoring information traceability, endoscopic surgery image processing, marine science exploration, underwater search and rescue, and other scenarios. The mainstream method uses supervised learning to construct synthetic clean images and degraded images to train neural networks. However, due to the domain difference between the synthetic images and the real images, the neural networks obtained by supervised training have poor generalization ability and poor effects in actual applications.

[0109] In other words, for the existing technologies that directly adopt the supervised learning method with a large amount of paired data, it is often difficult to collect the enhanced images corresponding to underwater or haze images in real scenes. Moreover, there is a domain difference between the existing publicly available paired data sets and real scenes, and the performance of the models trained based on the publicly available data will decline when directly migrated and used.

[0110] Based on this, the embodiments of the present application use the deep prior information of the pre-trained large model. To make more reasonable use of the deep prior information, an Adapter (adapter, the above-mentioned transmitted light extraction branch) is designed to estimate the TransmissionMap (the above-mentioned transmitted light data). Using the power operation relationship between the deep prior and the Transmission Map, a scattering coefficient loss function is designed. Using the characteristics of the degradation equation, reasonable quadratic augmented degraded images are designed to construct prior constraints and consistency losses to train the initial restoration model and obtain the target restoration model. Specifically, the training process of the initial restoration model includes but is not limited to:

[0111] S1, sending the sample degraded image I(x) (the above-mentioned first sample image) into three networks to respectively obtain the global background light A(x) (the above-mentioned background light data), the transmission T(x) (the above-mentioned transmitted light data), and the clean image J(x) (the above-mentioned initial image data).

[0112] It should be noted that after the training is completed, during the actual usage phase of the target restoration model, the re-degraded image obtained based on these three components (i.e., the degraded image I''(x) generated according to A(x), T(x), and J(x)) should be consistent with the original degraded image I(x). The atmospheric scattering equations for image dehazing and underwater image enhancement are as follows:

[0113] I''(x)=J(x)T'(x)+(1 - T'(x))A(x);

[0114] d(x)=DepthModel(x);

[0115] T(x)=Adapter(e -d(x) );

[0116] Among them, DepthModel can be a pre-trained large model (used to obtain depth image information and can be a component module in the above target restoration model), which can be implemented using open-source models such as DINO and DepthAnything. Its output is an estimated depth map. According to the principle of T(x), there is a power operation process here, and an Adapter can be designed to complete the fine-tuned output of T(x).

[0117] S2. Image secondary augmentation. Based on the T(x) component estimated by the initial restoration model, secondary augmentation is performed, mainly by modifying T(x):

[0118] ;

[0119] That is, a tiny random noise noise is added to the power exponent to achieve the purpose of augmenting the transmission map while still retaining the original information. This noise follows a Gaussian distribution with a mean of 0 and a variance of 0.01. At the same time, a new degraded image is recombined according to A(x) and J(x) (the above second sample image), and the formula is as follows:

[0120] I''(x)=J(x)T'(x)+(1 - T'(x))A(x);

[0121] Next, is sent to J-Net (the above target output branch) to obtain a new estimated value of the clean image (the above clear sample image).

[0122] Specifically, the loss function design of the initial restoration model can include but is not limited to:

[0123] S1, the re-degradation loss, that is, according to the atmospheric degradation scattering equation, the reconstructed image of the component should be consistent with the original degraded image, and the above first loss function L is designed. reson To achieve this constraint, it is expressed by the formula: ;

[0124] S2, image quadratic augmentation. Using the image J(x) processed for the first time (the above initial image data) and the regenerated clean image J'(x) (the above clear sample image), this component should be consistent, and the above second loss function L is designed. sim To achieve this constraint, in addition, considering the need to avoid falling into a degenerate solution, the sg, that is, gradient stop operation, can be adopted. It is expressed by the formula as follows: ;

[0125] S3, according to the power operation law of the Transmission-Map, considering the gradient consistency loss, the above third loss function L is designed. tmap To achieve this constraint, it is expressed by the formula as follows:

[0126] L tmap = ;

[0127] That is, the scattering coefficient is a constant, and the partial derivative of the transmitted light data with respect to the depth image data is a constant. After the initial reduction model training is completed, its second derivative is 0.

[0128] Furthermore, according to the first loss function, the second loss function, and the third loss function, the target loss function is determined as follows:

[0129] L total = L sim + L reson + L tmap ;

[0130] It should also be noted that in the embodiments of the present application, considering the principle of the atmospheric scattering equation, its equation is as follows:

[0131] ;

[0132] Among them, I(x) is the degraded image, that is, the foggy image or the underwater image, J(x) is the clear image, that is, the variable to be finally solved by the defogging or underwater enhancement algorithm, T(x) is the Transmission-Map (the above transmitted light data), and A(x) is the global background light (the above background light data). Specifically, T(x) is defined as follows:

[0133] ;

[0134] Among them, d(x) is the depth image data, and beta is the scattering coefficient (i.e., the scattering intensity). Specifically, Figure 3 is a schematic diagram of an optional image data processing method according to an embodiment of the present application. Figure 3 Both J-Net (the above-mentioned target output branch) and A-Net (the above-mentioned background light extraction branch) in it can be composed of simple convolutional layers, Norm, and the activation function Sigmoid. The number of layers can be flexibly set. For example, it can be set to any layer from 3 to 5 layers, and is used to solve the clean image J(x). Image A-Net is used to solve the global background light A(x), and its composition is a frequency domain adaptation network. Adapter is used to output the transmission light data (Transmission-Map), and its composition can include multiple convolutional layers with a convolution kernel of 1x1. The present application does not limit the size of the convolution kernel, and this is only an example here.

[0135] It can be understood that in the specific implementation manner of the present application, data related to user information and the like are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.

[0136] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0137] According to another aspect of the embodiments of the present application, there is also provided an image data processing device for implementing the above-mentioned image data processing method. As Figure 4 shown, the device includes:

[0138] An acquisition module 402, configured to acquire a degraded image, where the degraded image represents an image with transmission noise;

[0139] A processing module 404 is configured to input a degraded image into a target restoration model to obtain a clear image. The target restoration model is obtained by training an initial restoration model using first sample images. The initial restoration model includes a background light extraction branch, a transmitted light extraction branch, and a target output branch. The background light extraction branch is configured to extract background light data of the first sample images. The transmitted light extraction branch is configured to extract transmitted light data of the first sample images. The target output branch is configured to remove transmitted noise of the first sample images, generate initial image data, and remove transmitted noise of second sample images determined by the background light data, the transmitted light data, and the initial image data to obtain clear sample images, and the clear sample images are used to adjust model parameters of the initial restoration model to obtain the target restoration model.

[0140] As an alternative solution, the above device is further configured to: obtain the above first sample images; input the above first sample images into the above initial restoration model to obtain the above background light data, the above transmitted light data, and the above initial image data, where the above transmitted light data represents image data with noise added to the depth image data of the above first sample images, the above transmitted light data is used to indicate the light transmission state of the above first sample images, and the above background light data is used to indicate the image brightness of the above first sample images; generate second sample images according to the above background light data, the above transmitted light data, and the above initial image data; use the above target output branch to remove transmitted noise of the above second sample images to obtain the above clear sample images, where the above target output branch includes a plurality of convolutional layers; calculate a training loss value based on the above clear sample images, and determine the above initial restoration model as the above target restoration model when the above training loss value is less than or equal to a loss threshold.

[0141] As an alternative solution, the above device is configured to generate second sample images according to the above background light data, the above transmitted light data, and the above initial image data in the following manner: determine a product of the above initial image data and the above transmitted light data as first data; determine a product of a difference between the above transmitted light data and a constant 1 and the above background light data as second data; and determine the above second sample images according to the above first data and the above second data.

[0142] As an alternative solution, the above device is further configured to: perform a depth information extraction operation on the above first sample images using a depth extraction module in the above transmitted light extraction branch to obtain depth image data, and determine initial transmitted light data according to the depth image data and a scattering coefficient; add Gaussian noise to the above initial transmitted light data using an adapter module in the above transmitted light extraction branch to obtain the above transmitted light data; and generate the above second sample images based on the above transmitted light data.

[0143] As an alternative, the above-mentioned device is further configured to: perform a Fourier transform operation on the first sample image using the above-mentioned background light extraction branch to obtain the above-mentioned background light data, where the above-mentioned background light extraction branch includes a plurality of convolutional layers; generate the above-mentioned second sample image based on the above-mentioned background light data.

[0144] As an alternative, the above-mentioned device is further configured to: establish an objective loss function based on the first sample image, the second sample image, and the clear sample image; determine the above-mentioned training loss value using the above-mentioned objective loss function; and when the above-mentioned training loss value is less than or equal to a loss threshold, determine the above-mentioned initial restoration model as the above-mentioned target restoration model.

[0145] As an alternative, the above-mentioned device is configured to establish an objective loss function based on the background light data, the transmitted light data, and the initial image data in at least one of the following manners: establish a first loss function based on the mean square error of the image data of the first sample image and the image data of the second sample image, where the objective loss function includes the first loss function; establish a second loss function based on the mean square error of the image data of the clear sample image and the image data obtained after performing a gradient stopping operation on the initial image data, where the objective loss function includes the second loss function; establish a third loss function using the partial derivative of the initial transmitted light data, where the above-mentioned initial transmitted light data is determined by the depth image data of the first sample image and the scattering coefficient, and the above-mentioned depth image data is obtained by performing a depth information extraction operation on the first sample image through a depth extraction module in the above-mentioned transmitted light extraction branch, and the above-mentioned objective loss function includes the above-mentioned third loss function.

[0146] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0147] Regarding the device in the above-mentioned embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0148] According to one aspect of the present application, there is provided a computer program product, which includes a computer program.

[0149] The serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0150] Figure 5 A computer system block diagram of an electronic device for implementing the embodiments of the present application is schematically shown.

[0151] It should be noted that Figure 5 The computer system 500 of the shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0152] As Figure 5 shown, the computer system 500 includes a central processing unit 501 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 502 (Read-Only Memory, ROM) or the program loaded from the storage section 508 into the random access memory 503 (Random Access Memory, RAM). In the random access memory 503, various programs and data required for system operation are also stored. The central processing unit 501, the read-only memory 502, and the random access memory 503 are connected to each other via a bus 504. The input / output interface 505 (Input / Output interface, i.e., I / O interface) is also connected to the bus 504.

[0153] The following components are connected to the input / output interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including such as a cathode ray tube (Cathode Ray Tube, CRT), a liquid crystal display (Liquid Crystal Display, LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a local area network card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that the computer program read from it can be installed into the storage section 508 as needed.

[0154] In particular, according to an embodiment of the present application, the processes described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, various functions defined in the system of the present application are executed.

[0155] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, various functions provided by the embodiments of the present application are executed.

[0156] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above-mentioned method for processing image data. The electronic device may be Figure 1 the terminal device or server shown. This embodiment takes the electronic device as the terminal device as an example for illustration. As Figure 6 shown, the electronic device includes a memory 602 and a processor 604. A computer program is stored in the memory 602, and the processor 604 is configured to execute the steps in any one of the above method embodiments through the computer program.

[0157] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.

[0158] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the methods in the embodiments of the present application through a computer program.

[0159] Optionally, those of ordinary skill in the art can understand that Figure 6 the structure shown is only schematic, Figure 6 and it does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 6 , or have a different configuration from that shown in Figure 6 .

[0160] Among them, the memory 602 can be used to store software programs and modules, such as the program instructions / modules corresponding to the image data processing method and apparatus in the embodiments of the present application. The processor 604 executes various functional applications and data processing by running the software programs and modules stored in the memory 602, that is, implements the above-mentioned image data processing method. The memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 602 may further include a memory remotely disposed relative to the processor 604, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof. Among them, the memory 602 can specifically but not limitedly be used to store information such as degraded images, clear images, background light data, transmitted light data, and initial image data. As an example, as Figure 6 shown, the above-mentioned memory 602 may but not limitedly include the acquisition module 402 and the processing module 404 in the above-mentioned image data processing apparatus. In addition, it may also include but not limited to other module units in the above-mentioned image data processing apparatus, which will not be elaborated in this example.

[0161] Optionally, the above-mentioned transmission device 606 is used to receive or send data via a network. Specific examples of the above network may include wired networks and wireless networks. In one instance, the transmission device 606 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, thereby enabling communication with the Internet or local area network. In one instance, the transmission device 606 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.

[0162] In addition, the above-mentioned electronic device further includes: a display 608, which is used to display the above-mentioned degraded images and clear images; and a connection bus 610, which is used to connect each module component in the above-mentioned electronic device.

[0163] In other embodiments, the above-mentioned terminal device or server can be a node in a distributed system. Among them, the distributed system can be a blockchain system, and the blockchain system can be a distributed system formed by connecting the multiple nodes in the form of network communication. Among them, the nodes can form a peer-to-peer network, and any form of computing device, such as servers, terminals, and other electronic devices, can become a node in the blockchain system by joining the peer-to-peer network.

[0164] According to one aspect of the present application, a computer-readable storage medium is provided. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the image data processing methods provided in various optional implementations of the above-mentioned image data processing aspect.

[0165] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be configured to store instructions for executing the methods in the various embodiments of the present application.

[0166] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the relevant hardware of the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0167] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0168] If the integrated unit in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in the storage medium and includes several instructions for causing one or more electronic devices to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0169] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0170] In the several embodiments provided by the present application, it should be understood that the disclosed application program can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.

[0171] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional unit in various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0173] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for processing image data, characterized in that, Including: Obtain a degraded image, where the degraded image represents an image with transmission noise; Input the degraded image into a target restoration model to obtain a clear image, where the target restoration model is obtained by training an initial restoration model using a first sample image. The initial restoration model includes a background light extraction branch, a transmission light extraction branch, and a target output branch. The background light extraction branch is used to extract the background light data of the first sample image. The transmission light extraction branch is used to extract the transmission light data of the first sample image. The target output branch is used to remove the transmission noise of the first sample image, generate initial image data, and remove the transmission noise of a second sample image determined by the background light data, the transmission light data, and the initial image data to obtain a clear sample image. The clear sample image is used to adjust the model parameters of the initial restoration model to obtain the target restoration model; Establish a target loss function based on the first sample image, the second sample image, and the clear sample image, including: Establish a first loss function based on the mean square error of the image data of the first sample image and the image data of the second sample image, where the target loss function includes the first loss function; Establish a second loss function based on the mean square error of the image data of the clear sample image and the image data obtained after performing a gradient stop operation on the initial image data, where the target loss function includes the second loss function; Establish a third loss function using the partial derivative of the initial transmission light data, where the initial transmission light data is determined by the depth image data and the scattering coefficient of the first sample image, and the depth image data is obtained by performing a depth information extraction operation on the first sample image through a depth extraction module in the transmission light extraction branch. The target loss function includes the third loss function.

2. The method according to claim 1, wherein The method further includes: Obtain the first sample image; Input the first sample image into the initial restoration model to obtain the background light data, the transmission light data, and the initial image data, where the transmission light data represents image data with noise added to the depth image data of the first sample image, the transmission light data is used to indicate the light transmission state of the first sample image, and the background light data is used to indicate the image brightness of the first sample image; Generate a second sample image according to the background light data, the transmission light data, and the initial image data; Use the target output branch to remove the transmission noise of the second sample image to obtain the clear sample image, where the target output branch includes a plurality of convolutional layers; Calculate a training loss value based on the clear sample image. When the training loss value is less than or equal to a loss threshold, determine the initial restoration model as the target restoration model.

3. The method according to claim 2, characterized in that, The generating a second sample image according to the background light data, the transmission light data, and the initial image data includes: Determine the product of the initial image data and the transmission light data as the first data; Determine the product of the difference between the transmitted light data and the constant 1 and the background light data as the second data; Determine the second sample image according to the first data and the second data.

4. The method according to claim 1, characterized in that, The method further includes: Use the depth extraction module in the transmitted light extraction branch to perform a depth information extraction operation on the first sample image to obtain depth image data, and determine the initial transmitted light data according to the depth image data and the scattering coefficient; Use the adapter module in the transmitted light extraction branch to add Gaussian noise to the initial transmitted light data to obtain the transmitted light data; Generate the second sample image based on the transmitted light data.

5. The method according to claim 1, wherein The method further includes: Use the background light extraction branch to perform a Fourier transform operation on the first sample image to obtain the background light data, wherein the background light extraction branch includes a plurality of convolutional layers; Generate the second sample image based on the background light data.

6. The method according to claim 1, wherein The method further includes: Use the target loss function to determine the training loss value; In the case where the training loss value is less than or equal to the loss threshold, determine the initial restoration model as the target restoration model.

7. An apparatus for processing image data, characterized in that, Includes: An acquisition module for acquiring a degraded image, wherein the degraded image represents an image with transmitted noise; A processing module for inputting the degraded image into a target restoration model to obtain a clear image, wherein the target restoration model is obtained by training an initial restoration model using a first sample image, the initial restoration model includes a background light extraction branch, a transmitted light extraction branch, and a target output branch, the background light extraction branch is used to extract the background light data of the first sample image, the transmitted light extraction branch is used to extract the transmitted light data of the first sample image, the target output branch is used to remove the transmitted noise of the first sample image, generate initial image data, and remove the transmitted noise of the second sample image determined by the background light data, the transmitted light data, and the initial image data to obtain a clear sample image, and the clear sample image is used to adjust the model parameters of the initial restoration model to obtain the target restoration model; establish a target loss function based on the first sample image, the second sample image, and the clear sample image, including: establish a first loss function based on the mean square error of the image data of the first sample image and the image data of the second sample image, wherein the target loss function includes the first loss function; establish a second loss function based on the mean square error of the image data of the clear sample image and the image data obtained after performing a gradient stop operation on the initial image data, wherein the target loss function includes the second loss function; establish a third loss function using the partial derivative of the initial transmitted light data, wherein the initial transmitted light data is determined by the depth image data and the scattering coefficient of the first sample image, and the depth image data is obtained by performing a depth information extraction operation on the first sample image by the depth extraction module in the transmitted light extraction branch, and the target loss function includes the third loss function.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein the computer program, when being run by an electronic device, executes the method described in any one of claims 1 to 6.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 6 through the computer program.

Citation Information

Patent Citations

  • Image defogging method for extremely complex environment, medium and program product

    CN119579443A