Image denoising methods, apparatus, vehicles and storage media

By using a generative adversarial network model to train a dataset, a denoising network model is generated, which solves the problem of noise affecting vehicle cameras during nighttime driving, improving image clarity and driving safety.

CN115115531BActive Publication Date: 2025-10-28GREAT WALL MOTOR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210044611.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-10-28
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

During nighttime driving, reduced noise from onboard cameras can degrade the performance of driver monitoring systems, impacting driving safety.

Method used

By acquiring a training dataset, including noisy video images under low light conditions and noiseless video images under high light conditions, the generative adversarial network model is iteratively trained to generate a denoising network model for processing vehicle images to remove noise.

Benefits of technology

It improves image quality and enhances driving safety, especially in low-light conditions, by increasing image clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115531B_ABST
    Figure CN115115531B_ABST
Patent Text Reader

Abstract

This application relates to the field of image processing technology, and provides an image denoising method, apparatus, vehicle, and storage medium. The method includes: acquiring a training dataset, wherein the training dataset includes multiple data pairs, each data pair including N first video images and N second video images, the light intensity value inside the vehicle corresponding to the first video image is less than the light intensity value inside the vehicle corresponding to the second video image, the first video image and the second video image each correspond to different shooting objects inside the same vehicle, and N is a positive integer; iteratively training a denoising network model to be trained based on the training dataset to obtain a trained denoising network model; and denoising a third video image based on the trained denoising network model to obtain a denoised fourth video image. The above method can remove noise from in-vehicle images, improve image quality, and thus improve driving safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, and in particular relates to an image denoising method, apparatus, vehicle and storage medium. Background Technology

[0002] With the development of smart cockpit technology, in-vehicle cameras equipped with Driver Monitor Systems (DMS) have become a hallmark of cockpit intelligence. These cameras can not only identify the driver and determine their level of attention, but also infer their emotional state through facial expressions, thus aiding in safer driving. However, noise during nighttime driving degrades the image quality of the cameras, significantly reducing the performance of the DMS system and causing considerable inconvenience for drivers at night. Summary of the Invention

[0003] This application provides an image denoising method, apparatus, vehicle, and storage medium, which can remove noise from vehicle images, improve image quality, and thus enhance driving safety.

[0004] In a first aspect, embodiments of this application provide an image denoising method, including:

[0005] Obtain a training dataset, wherein the training dataset includes multiple data pairs, each data pair including N first video images and N second video images, the light brightness value inside the car corresponding to the first video image is less than the light brightness value inside the car corresponding to the second video image, the first video image and the second video image each correspond to different shooting objects inside the same car, and N is a positive integer;

[0006] Based on the training dataset, the denoising network model to be trained is iteratively trained to obtain the trained denoising network model.

[0007] The third video image is denoised based on the trained denoising network model to obtain the denoised fourth video image.

[0008] Secondly, embodiments of this application provide an image denoising apparatus, including:

[0009] The acquisition module is used to acquire a training dataset, wherein the training dataset includes multiple data pairs, each data pair includes N first video images and N second video images, the light brightness value inside the vehicle corresponding to the first video image is less than the light brightness value inside the vehicle corresponding to the second video image, the first video image and the second video image each correspond to different shooting objects inside the same vehicle, and N is a positive integer;

[0010] The training module is used to iteratively train the denoising network model to be trained based on the training dataset to obtain the trained denoising network model.

[0011] The denoising module is used to denoise the third video image based on the trained denoising network model to obtain the denoised fourth video image.

[0012] Thirdly, embodiments of this application provide a vehicle including a terminal device, the terminal device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the image denoising method described in any one of the first aspects above.

[0013] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, characterized in that the computer program, when executed by a processor, implements the image denoising method described in any one of the first aspects above.

[0014] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the image denoising method described in any one of the first aspects.

[0015] The beneficial effects of the first aspect of this application compared with the prior art are as follows: By acquiring a training dataset, wherein the training dataset includes multiple data pairs, each data pair including N first video images and N second video images, the light intensity value inside the vehicle corresponding to the first video image is less than the light intensity value inside the vehicle corresponding to the second video image, the first video image and the second video image each correspond to different shooting objects inside the same vehicle, and N is a positive integer; the denoising network model to be trained is iteratively trained based on the training dataset to obtain the trained denoising network model; the third video image is denoised based on the trained denoising network model to obtain the denoised fourth video image. In this way, the denoising network model trained using the unpaired training dataset with low acquisition difficulty can remove noise from the vehicle images, improve the image quality, and thus improve driving safety.

[0016] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic flowchart of an image denoising method provided in an embodiment of this application;

[0019] Figure 2 yes Figure 1 A schematic diagram illustrating the specific implementation steps of step 102;

[0020] Figure 3 This is a schematic diagram of the structure of an image denoising device provided in an embodiment of this application;

[0021] Figure 4 This is a schematic diagram of the structure of a terminal device in a vehicle provided in an embodiment of this application. Detailed Implementation

[0022] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0023] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0024] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0025] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0026] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0027] Figure 1 A schematic flowchart of the image denoising method provided in this application is shown, with reference to... Figure 1 The image denoising method is described in detail below:

[0028] S101, Obtain the training dataset, wherein the training dataset includes multiple data pairs, each data pair including N first video images and N second video images, the light brightness value inside the car corresponding to the first video image is less than the light brightness value inside the car corresponding to the second video image, the first video image and the second video image each correspond to different shooting objects inside the same car, and N is a positive integer.

[0029] Here, the training dataset can be obtained through acquisition devices. These acquisition devices include DMS cameras and Occupancy Monitoring System (OMS) cameras, both of which are used to acquire video images inside the vehicle.

[0030] In one possible implementation, the process of step S101 may include:

[0031] S1011, Collect video images inside the vehicle when the first condition is met, and use the video images inside the vehicle when the first condition is met as the first video image. The first condition is that the light brightness value is less than the second preset threshold.

[0032] Here, the first condition is that the light intensity value is less than the second preset threshold, indicating that the lighting conditions in the in-vehicle shooting environment are poor. For example, the vehicle is driving at night, in a tunnel, or in an underground parking garage. The first video image is a noisy image.

[0033] S1012, Collect video images inside the vehicle when the second condition is met, and use the video images inside the vehicle when the second condition is met as the second video images. The second condition is that the light brightness value is greater than or equal to the second preset threshold.

[0034] Here, the second condition is that the light intensity value is greater than or equal to the second preset threshold, indicating that the lighting conditions in the in-vehicle shooting environment are relatively good. For example, the vehicle is driving during the day. The second video image is a noise-free image.

[0035] S1013, construct a training dataset based on the first video image and the second video image.

[0036] It should be noted that each first video image captured by the camera has a corresponding second video image. That is, the first and second video images each correspond to different subjects within the same vehicle. In other words, while the image capture environments of the first video image and its corresponding second video image are related, they are not paired. That is, the content captured in the two video images does not need to be exactly the same.

[0037] As can be seen from the above description, the training dataset in this embodiment is an unpaired dataset. Unpaired datasets are easier to collect and obtain, which reduces the difficulty of training the denoising network model to a certain extent.

[0038] This step may specifically include: dividing the acquired first and second video images into multiple data pairs, each data pair comprising N first video images and N second video images. It should be noted that each of the N first video images corresponds to one second video image in that data pair. Different first video images can be video images from the same vehicle or video images from different vehicles.

[0039] To enrich the diversity of the collected data, the camera can be placed in multiple locations, such as the A-pillar, center console, steering wheel, and rearview mirror. In addition, the age and gender of the people being filmed inside the vehicle, such as the driver and passengers, also need to be diverse.

[0040] S102, iteratively train the denoising network model to be trained based on the training dataset to obtain the trained denoising network model.

[0041] Optionally, the denoising network model to be trained is a Generative Adversarial Network (GAN) model. A GAN model consists of a generator network and a discriminator network.

[0042] The general process of iteratively training the denoising network model to be trained based on the training dataset is as follows: the untrained generator network and the discriminator network are trained adversarially together. The generator network is used to generate a random video image to deceive the discriminator network by inputting the first video image from the data pair. Then, the discriminator network judges whether the random video image is real or fake. Finally, during the training process of these two types of networks, the capabilities of the two types of networks become stronger and stronger, eventually reaching a steady state. The generator network that has reached a steady state is determined as the trained denoising network model. For details, please refer to Example 1.

[0043] S103, the third video image is denoised based on the trained denoising network model to obtain the denoised fourth video image.

[0044] In one possible implementation, step S103 may include:

[0045] S1031, acquire a third video image of the vehicle interior.

[0046] Here, third-party video images of the vehicle's interior can be obtained using cameras installed on the vehicle, such as DMS cameras and OMS cameras.

[0047] S1032, based on the brightness value of each pixel in the third video image, calculate the average light brightness value corresponding to the third video image.

[0048] Specifically, the average luminance value corresponding to the third video image is calculated using the following formula (1):

[0049]

[0050] Where (x,y) represents the position coordinates of a pixel in the third video image, Lum(x,y) represents the brightness value of a pixel in the third video image, N represents the total number of pixels in the third video image, and δ is a preset parameter value.

[0051] S1033, if the average light intensity value is less than the first preset threshold, then the third video image is denoised based on the trained denoising network model to obtain the denoised fourth video image.

[0052] Here, if the average light intensity value is less than the first preset threshold, it means that the lighting conditions of the in-vehicle shooting environment corresponding to the third video image are poor. The third video image is a noisy video image and needs to be denoised. That is, the third video image is input into the trained denoising network model for denoising processing, and the denoised fourth video image is output. The fourth video image is a clear image after removing noise, thereby improving the image quality during driving and thus improving driving safety.

[0053] Example 1

[0054] like Figure 2 As shown, in one possible implementation, the process of step S102 may include:

[0055] S1021, input the first video image from the data pair into the denoising network model to obtain the fifth video image.

[0056] This step may specifically include: extracting the Y component of the YUV spatial image corresponding to the first video image in the data pair; inputting the Y component corresponding to the first video image into the denoising network model to obtain the fifth video image.

[0057] Optionally, the denoising network model can be a generative network.

[0058] It should be noted that the main structural information of an image is in the Y component. Using the corresponding Y component of the image for model training can reduce the amount of computation and improve the robustness of the denoising network model, making it unaffected by the color space of the image.

[0059] It should be noted that the first video image is a noisy image. The network parameters of the generator network are randomized at the start of training. The N first video images from the data pair are input into the generator network to obtain N fifth video images, which are N random video images.

[0060] Optionally, the network structure of the generator adopts the U-Net network structure, which includes M convolutional modules. Each convolutional module contains convolutional layers, activation functions, and normalization layers, with multiple convolutional layers. Each convolutional module is a residual structure, divided into downsampling and upsampling stages. In the upsampling stage, bilinear upsampling is combined with convolutional layers to replace the traditional deconvolution operation.

[0061] For example, the U-Net network structure may include 8 convolutional modules, each of which contains 2 3×3 convolutional layers, a LeakyReLu activation function, and a normalization layer.

[0062] S1022, the fifth video image is input into the global discrimination network to obtain the first discrimination result.

[0063] In this step, the fifth video image is input into the global discriminant network to determine whether the fifth video image is real or fake. The output result is generally represented by 0 or 1. 0 indicates a false result, and 1 indicates a true result.

[0064] S1023, input the second video image from the data pair into the global discrimination network to obtain the second discrimination result.

[0065] This step may specifically include: extracting the Y component of the YUV spatial image corresponding to the second video image in the data pair; inputting the Y component corresponding to the second video image into the global discrimination network to obtain the second discrimination result.

[0066] Here, the second video image, i.e. the real noiseless image, is input into the global discriminant network in order to determine whether the second video image is real or fake.

[0067] S1024, input the local image of the fifth video image into the local discrimination network to obtain the third discrimination result.

[0068] Here, each fifth video image is divided into multiple local images and input into a local discriminant network to determine the authenticity of local images in the fifth video image, thereby guiding the training of the subsequent generative network in a more refined manner.

[0069] It should be noted that, in this embodiment, the discriminator network specifically refers to the global discriminator network and the local discriminator network. For the second video image in the data pair, i.e., the real noise-free image, the discriminator network should correctly output the value 1 as much as possible. However, for the fifth video image, which is the fifth video image (fake denoised image) obtained by the generator network after the first video image (noisy image) in the data pair, the discriminator network should incorrectly output the value 0 as much as possible. Meanwhile, the generator network tries to deceive the discriminator network as much as possible. The two networks work against each other, constantly adjusting their parameters, with the ultimate goal of making it impossible for the discriminator network (both the global and local discriminator networks) to determine whether the output of the generator network is real.

[0070] In this embodiment, the two networks compete against each other and continuously adjust their parameters, specifically referring to the continuous adjustment of the network parameters of the global discriminant network, the local discriminant network, and the generator network.

[0071] First, keep the network parameters of the generator network unchanged, and update the network parameters of the global discriminant network and the local discriminant network multiple times to improve their discrimination capabilities. Then, keep the network parameters of the global discriminant network and the local discriminant network unchanged, and update the network parameters of the generator network multiple times.

[0072] In this implementation, the network parameters of the global discriminant network and the local discriminant network have been updated multiple times by traversing some data in the training dataset. This implementation specifically describes the process of keeping the network parameters of the global discriminant network and the local discriminant network unchanged, and updating the network parameters of the generated network multiple times based on some unused data.

[0073] The following section uses a global discriminant network as an example to illustrate the process of updating the network parameters of the global discriminant network in the previous implementation.

[0074] a1: Input the N first video images from the first data pair into the generator network G to obtain N random video images x. f .

[0075] a2: Take N random video images x f Inputting a global discriminant network D1 yields N discriminant results D1(x) f ).

[0076] a3: Take N second video images from the first data pair x r Inputting a global discriminant network D1 yields N discriminant results D1(x) r ).

[0077] a4: Based on N discrimination results D1(x f ) and N discrimination results D1(x r The loss value of the global discriminant network D1 is calculated.

[0078] Specifically, the loss value of the global discriminant network D1 is calculated using the following formulas (2), (3), and (4).

[0079]

[0080]

[0081]

[0082] Among them, D c (x r ,x f D is used to represent the global discriminator's ability to distinguish between the second video image. c (x r ,x f The larger the value of ), the greater the probability that the second video image is identified as a real image; D c (x f ,x r D is used to represent the ability of a global discriminator to distinguish random video images. c (x f ,x r The larger the value of ), the greater the probability that a random video image is identified as a real image.

[0083] D1(x) represents N discrimination results. f The expected value of ) σ(D1(x) r )) represents N discrimination results D1(x r The activation function value; D1(x) represents N discrimination results. r The expected value of ) σ(D1(x) f )) represents N discrimination results D1(x f The activation function value of ).

[0084] a5: Loss value based on global discriminant network D1 Update the network parameters of the global discriminant network D1 to obtain the updated global discriminant network D1.

[0085] The above process is one network parameter update process for the global discriminant network D1. Afterwards, the first data pair in step a1 is replaced with the second data pair, and the above steps are repeated to achieve another network parameter update for the global discriminant network D1 after the parameter update, until the first preset number of network parameter updates has been performed.

[0086] It should be noted that the network parameter update process of the local discriminant network D2 is roughly the same as that of the global discriminant network described above. The difference lies in the different data pairs used for parameter updates, the fact that the input to the local discriminant network D2 is a local image of the video image, and the different loss function used to calculate the loss value of the local discriminant network D2. Specifically, the loss function used to calculate the loss value of the local discriminant network D2 is as follows: Formula (5):

[0087]

[0088] Wherein, D2(x r ) represents the discrimination result obtained after inputting a local image of the second video image in the data pair into the local discriminant network D2, where D2(x) f ) represents the discrimination result obtained after a local image of a random video image is input into the local discriminant network D2. The random video image is the video image obtained after the first video image in the data pair is processed by the generator network.

[0089] S1025, calculate the target loss value of the denoising network model based on the first discrimination result, the second discrimination result and the third discrimination result.

[0090] Here, the training of the denoising network model, i.e. the generative network, is guided by the discrimination results of two discriminant networks (i.e., global discriminant network D1 and local discriminant network D2). For details of the implementation process, please refer to Example 2.

[0091] S1026, Update the network parameters of the denoising network model according to the target loss value to obtain the trained denoising network model.

[0092] This step may specifically include: after updating the network parameters of the denoising network model according to the target loss value, determining whether the updated denoising network model meets the preset conditions. If it meets the preset conditions, the updated denoising network model is determined as the trained denoising network model; otherwise, the updated denoising network model is updated again until the updated denoising network model meets the preset conditions.

[0093] Specifically, updating the denoised network model again involves updating the denoised network model to a denoised network model, then returning to step S1021 and subsequent steps. At this point, the data pair in step S1021 is replaced with the next data pair from the training dataset.

[0094] Optionally, the preset condition is that the target loss value of the denoised network model after this update is equal to the target loss value of the denoised network model after the last update, or the preset condition is that the cumulative number of network parameter updates corresponding to the updated denoised network model reaches a second preset number.

[0095] In one example, if the target loss value of the denoised network model after this update is not equal to the target loss value of the denoised network model after the last update, and the number of times the network parameters of the denoised network model are updated in this round has not reached the third preset number, then after updating the denoised network model to the denoised network model, the process returns to step S1021 and subsequent steps until the target loss value of the denoised network model after this update is equal to the target loss value of the denoised network model after the last update, or the number of times the network parameters of the denoised network model are updated in this round reaches the third preset number.

[0096] Here, completing one round of network parameter updates includes multiple updates to the global discriminator network, multiple updates to the local discriminator network, and multiple updates to the denoising network model. The number of updates to the network parameters of the two discriminator networks is greater than the number of updates to the network parameters of the denoising network model. Different data pairs are used when updating the network parameters of different networks. It should be noted that during the update of one network's parameters, the network parameters of the other two networks remain unchanged.

[0097] When the number of network updates for the current denoising network model reaches the third preset number, and the target loss value of the updated denoising network model is no longer equal to the target loss value of the previous updated denoising network model, then the next round of network parameter update process will be executed until the target loss value of the current updated denoising network model is equal to the target loss value of the previous updated denoising network model. It should be noted that at this point, the loss values ​​of the global discriminant network and the local discriminant network will no longer decrease.

[0098] Example 2

[0099] In one possible implementation, step S1025 may specifically include:

[0100] b1: Calculate the first loss value, which represents the feature distance between the first video image and the fifth video image.

[0101] It should be noted that, since the data pairs in the training data are unpaired and there is no real image (second video image) that truly corresponds to the first video image, in order to ensure the consistency of the video image content before and after the network output, this embodiment uses the perceptual loss function to calculate a first loss value representing the feature distance between the first video image and the fifth video image. The specific implementation process includes:

[0102] b11: Extract the image features of the first video image and the image features of the fifth video image respectively.

[0103] Specifically, the image features of the first video image and the image features of the fifth video image are extracted using a pre-trained feature extraction model.

[0104] Optionally, the feature extraction model is the VGG-16 model.

[0105] b12: Calculate the feature distance between the first video image and the fifth video image based on the image features of the first video image and the fifth video image.

[0106] Specifically, the feature distance between the first video image and the fifth video image is calculated using the following formula (6):

[0107]

[0108] in, Let I represent the image features extracted by the j-th convolutional layer after the i-th max pooling layer of the feature extraction model from the first video image; G(I) represents the image features extracted by the j-th convolutional layer after the i-th max pooling layer of the feature extraction model, G(I) represents the fifth video image obtained after the first video image is processed by the generative network, and y represents the height of the image, with a value ranging from 1 to H. i,j x represents the width of the image, with a value ranging from 1 to w. i,j , w i,j H i,j L is used to represent the area of ​​an image. p This represents the first loss value, which is the feature distance between the first video image and the fifth video image.

[0109] b13: The feature distance between the first video image and the fifth video image is determined as the first loss value.

[0110] b2: Calculate the second loss value based on the first discrimination result and the second discrimination result.

[0111] In one possible implementation, step b2 may include:

[0112] b21: Calculate the first expected value of the first discrimination result and the second expected value of the second discrimination result respectively.

[0113] b22: Based on the preset activation function, calculate the first function value of the first discrimination result and the second function value of the second discrimination result respectively.

[0114] b23: Calculate the second loss value based on the first expected value, the second expected value, the first function value, and the second function value.

[0115] It should be noted that steps b21 to b23 above can be expressed by formula (7), that is, the second loss value is calculated using the following formula (7):

[0116]

[0117] Here, D c (x f ,x r D can be expressed by the above formula (4), c (x r ,x f It can be expressed by the above formula (3), that is:

[0118]

[0119]

[0120] At this time, D1(x) f D1(x) represents the first discrimination result. r () indicates the second discrimination result. This represents the second loss value.

[0121] b3: Calculate the third loss value based on the third discrimination result.

[0122] Specifically, the third loss value is calculated using the following formula (8):

[0123]

[0124] Wherein, D2(x f () indicates the third discrimination result. This represents the third loss value.

[0125] b4: Sum the first loss value, the second loss value, and the third loss value to obtain the target loss value of the denoising network model.

[0126] Here, the target loss value

[0127] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0128] Corresponding to the method described in the above embodiments, Figure 3 A structural block diagram of an image denoising apparatus provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.

[0129] Reference Figure 3 The device 300 may include: an acquisition module 310, a training module 320, and a noise reduction processing module 330.

[0130] The acquisition module 310 is used to acquire a training dataset, which includes multiple data pairs. Each data pair includes N first video images and N second video images. The light intensity value inside the vehicle corresponding to the first video image is less than the light intensity value inside the vehicle corresponding to the second video image. The first video image and the second video image each correspond to different shooting objects inside the same vehicle, and N is a positive integer.

[0131] Training module 320 is used to iteratively train the denoising network model to be trained based on the training dataset to obtain the trained denoising network model.

[0132] The denoising module 330 is used to denoise the third video image based on the trained denoising network model to obtain the denoised fourth video image.

[0133] In one possible implementation, the training module 320 may specifically include:

[0134] The first processing submodule is used to input the first video image from the data pair into the denoising network model to obtain the fifth video image.

[0135] The second processing submodule is used to input the fifth video image into the global discrimination network to obtain the first discrimination result.

[0136] The third processing submodule is used to input the second video image from the data pair into the global discrimination network to obtain the second discrimination result.

[0137] The fourth processing submodule is used to input a local image of the fifth video image into the local discrimination network to obtain the third discrimination result.

[0138] The first calculation submodule is used to calculate the target loss value of the denoising network model based on the first discrimination result, the second discrimination result, and the third discrimination result.

[0139] The parameter update submodule is used to update the network parameters of the denoising network model based on the target loss value, so as to obtain the trained denoising network model.

[0140] In one possible implementation, the first computation submodule may specifically include:

[0141] The first calculation unit is used to calculate the first loss value, which represents the feature distance between the first video image and the fifth video image.

[0142] The second calculation unit is used to calculate the second loss value based on the first discrimination result and the second discrimination result.

[0143] The third calculation unit is used to calculate the third loss value based on the third discrimination result.

[0144] The fourth calculation unit is used to sum the first loss value, the second loss value, and the third loss value to obtain the target loss value of the denoising network model.

[0145] In one possible implementation, the first computing unit is specifically used for:

[0146] Image features of the first video image and image features of the fifth video image are extracted respectively.

[0147] The feature distance between the first video image and the fifth video image is calculated based on the image features of the first video image and the fifth video image.

[0148] The feature distance between the first video image and the fifth video image is determined as the first loss value.

[0149] In one possible implementation, the second computational unit is specifically used for:

[0150] Calculate the first expected value of the first discrimination result and the second expected value of the second discrimination result respectively.

[0151] Based on the preset activation function, the first function value of the first discrimination result and the second function value of the second discrimination result are calculated respectively.

[0152] Calculate the second loss value based on the first expected value, the second expected value, the first function value, and the second function value.

[0153] In one possible implementation, the noise reduction module 330 can specifically be used for:

[0154] Acquire third-party video images of the vehicle's interior.

[0155] The average light intensity value corresponding to the third video image is calculated based on the brightness value of each pixel in the third video image.

[0156] If the average light intensity value is less than the first preset threshold, the third video image is denoised based on the trained denoising network model to obtain the denoised fourth video image.

[0157] In one possible implementation, the acquisition module 310 can specifically be used for:

[0158] The video image inside the vehicle is acquired when the first condition is met, and the video image inside the vehicle when the first condition is met is used as the first video image. The first condition is that the light brightness value is less than the second preset threshold.

[0159] The video image inside the vehicle is acquired when the second condition is met, and the video image inside the vehicle when the second condition is met is used as the second video image. The second condition is that the light brightness value is greater than or equal to the second preset threshold.

[0160] A training dataset is constructed based on the first and second video images.

[0161] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0162] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0163] This application also provides a vehicle, see [link to example]. Figure 4The vehicle includes a terminal device 400, which may include at least one processor 410, a memory 420, and a computer program stored in the memory 420 and executable on the at least one processor 410. When the processor 410 executes the computer program, it implements the steps in any of the above-described method embodiments, for example... Figure 1 Steps S101 to S103 in the illustrated embodiment. Alternatively, when the processor 410 executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 3 The functions of modules 310 to 330 are shown.

[0164] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 420 and executed by processor 410 to complete this application. The one or more modules / units may be a series of computer program segments capable of performing a specific function, which are used to describe the execution process of the computer program in terminal device 400.

[0165] Those skilled in the art will understand that Figure 4 This is merely an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, buses, etc.

[0166] The processor 410 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0167] The memory 420 can be an internal storage unit of the terminal device or an external storage device, such as a plug-in hard drive, a smart media card (SMC), a secure digital card (SD), or a flash card. The memory 420 is used to store the computer program and other programs and data required by the terminal device. The memory 420 can also be used to temporarily store data that has been output or will be output.

[0168] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0169] The image denoising method provided in this application embodiment can be applied to terminal devices such as computers, tablets, laptops, netbooks, and personal digital assistants (PDAs). This application embodiment does not impose any restrictions on the specific type of terminal device.

[0170] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0171] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0172] In the embodiments provided in this application, it should be understood that the disclosed terminal devices, apparatuses, and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, apparatuses, or units, and may be electrical, mechanical, or other forms.

[0173] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0174] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0175] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by one or more processors, it can implement the steps of the various method embodiments described above.

[0176] Similarly, as a computer program product, when the computer program product is run on a terminal device, it enables the terminal device to implement the steps in the above-described method embodiments.

[0177] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0178] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. An image denoising method, characterized in that, include: A training dataset is obtained, comprising multiple data pairs. Each data pair includes N first video images and N second video images. The light intensity value inside the vehicle corresponding to the first video image is less than the light intensity value inside the vehicle corresponding to the second video image. The first video image and the second video image each correspond to different subjects captured inside the same vehicle, and N is a positive integer. The second video images are noise-free images. Each first video image has a corresponding second video image. The image acquisition environments of the first video image and the second video image are associated but not paired. Iterative training is performed on the denoising network model to be trained based on the training dataset to obtain the trained denoising network model. This includes: inputting the first video image from the data pair into the denoising network model to obtain a fifth video image; inputting the fifth video image into a global discriminant network to obtain a first discrimination result; inputting the second video image from the data pair into the global discriminant network to obtain a second discrimination result; inputting a local image of the fifth video image into a local discriminant network to obtain a third discrimination result; calculating the target loss value of the denoising network model based on the first discrimination result, the second discrimination result, and the third discrimination result; and updating the network parameters of the denoising network model based on the target loss value to obtain the trained denoising network model. The denoising network model is a generative network. The third video image is denoised based on the trained denoising network model to obtain the denoised fourth video image. The step of calculating the target loss value of the denoising network model based on the first discrimination result, the second discrimination result, and the third discrimination result includes: Extract the image features of the first video image and the image features of the fifth video image respectively; Calculate the feature distance between the first video image and the fifth video image based on the image features of the first video image and the image features of the fifth video image; The feature distance between the first video image and the fifth video image is determined as the first loss value; the first loss value represents the feature distance between the first video image and the fifth video image. Calculate the second loss value based on the first discrimination result and the second discrimination result; Calculate the third loss value based on the third discrimination result; The target loss value of the denoising network model is obtained by summing the first loss value, the second loss value, and the third loss value.

2. The image denoising method as described in claim 1, characterized in that, The step of calculating the second loss value based on the first discrimination result and the second discrimination result includes: Calculate the first expected value of the first discrimination result and the second expected value of the second discrimination result respectively; Based on a preset activation function, the first function value of the first discrimination result and the second function value of the second discrimination result are calculated respectively; Calculate the second loss value based on the first expected value, the second expected value, the first function value, and the second function value.

3. The image denoising method as described in claim 1, characterized in that, The step of denoising the third video image based on the trained denoising network model to obtain a denoised fourth video image includes: Acquire the third video image of the vehicle's interior; Based on the brightness value of each pixel in the third video image, the average light brightness value corresponding to the third video image is calculated. If the average light intensity value is less than the first preset threshold, the third video image is denoised based on the trained denoising network model to obtain the denoised fourth video image.

4. The image denoising method as described in claim 1, characterized in that, The acquisition of the training dataset includes: Collect video images of the vehicle interior when a first condition is met, and use the video images of the vehicle interior when the first condition is met as the first video image. The first condition is that the light brightness value is less than a second preset threshold. Collect video images of the vehicle interior when the second condition is met, and use the video images of the vehicle interior when the second condition is met as the second video image. The second condition is that the light brightness value is greater than or equal to the second preset threshold. The training dataset is constructed based on the first video image and the second video image.

5. An image denoising device, characterized in that, include: An acquisition module is used to acquire a training dataset, wherein the training dataset includes multiple data pairs, each data pair including N first video images and N second video images, the light intensity value inside the vehicle corresponding to the first video image is less than the light intensity value inside the vehicle corresponding to the second video image, the first video image and the second video image each correspond to different shooting objects inside the same vehicle, and N is a positive integer; wherein the second video image is a noise-free image; each first video image has one corresponding second video image; the image acquisition environments of the first video image and the second video image are associated, but not paired; The training module is used to iteratively train the denoising network model to be trained based on the training dataset to obtain the trained denoising network model. Specifically, the training module includes: a first processing submodule, used to input the first video image from the data pair into the denoising network model to obtain the fifth video image; a second processing submodule, used to input the fifth video image into a global discriminant network to obtain the first discrimination result; a third processing submodule, used to input the second video image from the data pair into the global discriminant network to obtain the second discrimination result; a fourth processing submodule, used to input a local image of the fifth video image into a local discriminant network to obtain the third discrimination result; a first calculation submodule, used to calculate the target loss value of the denoising network model based on the first, second, and third discrimination results; and a parameter update submodule, used to update the network parameters of the denoising network model based on the target loss value to obtain the trained denoising network model. The denoising network model is a generative network. The denoising module is used to denoise the third video image based on the trained denoising network model to obtain the denoised fourth video image. The first calculation submodule specifically includes: The first calculation unit is used to calculate the first loss value, which represents the feature distance between the first video image and the fifth video image. The second calculation unit is used to calculate the second loss value based on the first discrimination result and the second discrimination result; The third calculation unit is used to calculate the third loss value based on the third discrimination result; The fourth calculation unit is used to sum the first loss value, the second loss value, and the third loss value to obtain the target loss value of the denoising network model; The first computing unit is specifically used for: Extract image features from the first video image and the fifth video image respectively; Calculate the feature distance between the first video image and the fifth video image based on the image features of the first video image and the fifth video image; The feature distance between the first video image and the fifth video image is determined as the first loss value.

6. A vehicle, comprising a terminal device, said terminal device including a memory, a processor, and a computer program stored in said memory and executable on said processor, characterized in that, When the processor executes the computer program, it implements the image denoising method as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image denoising method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Image denoising method and device and storage medium

    CN113643189A

  • Method and device for constructing image processing model, equipment and storage medium

    CN113887390A