Image processing method and image processing apparatus

By employing image processing methods such as scratch restoration, bad pixel recovery, noise removal, and color cast correction to video images, the problem of poor video image display quality has been solved, and image quality has been improved.

CN115398469BActive Publication Date: 2026-04-24BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BOE TECHNOLOGY GROUP CO LTD
Filing Date
2021-03-22
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Video images may suffer from poor display quality due to scratches, dead pixels, noise, or color cast, which are difficult to repair effectively with existing technologies.

Method used

Image processing methods are used to restore scratches, recover bad pixels, remove noise, and correct color cast in video images. Image restoration is performed through techniques such as difference operations, multi-frame filtering, denoising networks, and color balance adjustment.

Benefits of technology

It improves the display effect of video images, effectively removes scratches, dead pixels and noise, corrects color cast, and enhances image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115398469B_ABST
    Figure CN115398469B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method and an image processing device, the method comprising: performing at least one of the following steps on a to-be-processed video image: a scratch repairing step, a bad point repairing step, a noise removing step and a color cast correcting step. In the present disclosure, the video image can be restored for scratches, bad points, noise and / or color cast, so as to improve the display effect of the video image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of display technology, and more particularly to an image processing method and an image processing apparatus. Background Technology

[0002] Video images, such as film reels, may develop scratches, dead pixels, noise, or color casts due to the passage of time or improper storage. How to repair these problems in video images to improve display quality is an urgent issue to be addressed. Summary of the Invention

[0003] This disclosure provides an image processing method and an image processing apparatus for performing scratch restoration, dead pixel restoration, noise removal, and / or color cast correction on video images to improve display performance.

[0004] To solve the above-mentioned technical problems, this disclosure is implemented as follows:

[0005] In a first aspect, embodiments of this disclosure provide an image processing method, including:

[0006] Perform at least one of the following steps on the video image to be processed:

[0007] Scratch repair steps: Perform scratch removal processing on the video image to be processed to obtain a first image; perform difference operation on the video image to be processed and the first image to obtain a difference image; process the difference image to obtain a scratch image that retains only the scratches; obtain a scratch repair image based on the video image to be processed and the scratch image;

[0008] Defective pixel repair steps: Acquire N1 consecutive video images, wherein each N1 video image includes a video image to be processed, at least one video image preceding the video image to be processed, and at least one video image following the video image to be processed; Filter the video image to be processed based on the N1 video images to obtain a defective pixel repair image; Perform artifact removal processing on the defective pixel repair image based on the at least one video image preceding the video image to be processed and at least one video image following the video image to be processed to obtain an artifact repair image;

[0009] Denoising step: A denoising network is used to denoise the video image to be processed. The denoising network is obtained by the following training method: a target non-motion mask is obtained based on N2 consecutive video images, wherein the N2 video images include the video image to be denoised; the denoising network to be trained is trained based on the N2 video images and the target non-motion mask to obtain the denoising network.

[0010] Color cast correction steps: Determine the target color cast values ​​for each RGB channel of the video image to be processed; perform color balance adjustment on the video image to be processed according to the target color cast values ​​to obtain a first corrected image; perform color migration on the first corrected image according to a reference image to obtain a second corrected image.

[0011] In a second aspect, embodiments of this disclosure provide an image processing apparatus, comprising:

[0012] The processing module includes at least one of the following modules:

[0013] Scratch Repair Submodule: Performs scratch removal processing on the video image to be processed to obtain a first image; performs difference calculation on the video image to be processed and the first image to obtain a difference map; processes the difference map to obtain a scratch image that retains only the scratches; and obtains a scratch repair image based on the video image to be processed and the scratch image.

[0014] The dead pixel repair submodule acquires N1 consecutive video images, each N1 frame including a video image to be processed, at least one video image preceding the video image to be processed, and at least one video image following the video image to be processed; it then filters the video image to be processed based on the N1 video images to obtain a dead pixel repair image; and finally, it performs artifact removal processing on the dead pixel repair image based on the at least one video image preceding the video image to be processed and at least one video image following the video image to be processed to obtain an artifact repair image.

[0015] Denoising submodule: A denoising network is used to denoise the video image to be processed. The denoising network is obtained by the following training method: a target non-motion mask is obtained based on N2 consecutive video images, wherein the N2 video images include the video image to be denoised; the denoising network to be trained is trained based on the N2 video images and the target non-motion mask to obtain the denoising network.

[0016] Color cast correction submodule: Determines the target color cast values ​​for each RGB channel of the video image to be processed; performs color balance adjustment on the video image to be processed based on the target color cast values ​​to obtain a first corrected image; performs color migration on the first corrected image based on a reference image to obtain a second corrected image.

[0017] Thirdly, embodiments of this disclosure provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the above-described image processing method.

[0018] Fourthly, embodiments of this disclosure provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the above-described image processing method.

[0019] In the embodiments of this disclosure, scratch recovery, dead pixel recovery, noise removal and / or color cast correction can be performed on video images to improve the display effect of video images. Attached Figure Description

[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this disclosure. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0021] Figure 1 This is a flowchart illustrating the scratch repair steps according to an embodiment of the present disclosure;

[0022] Figure 2 This is a schematic diagram illustrating the specific process of scratch repair steps according to an embodiment of the present disclosure;

[0023] Figure 3 This is a schematic diagram of the video image to be processed and the first image according to an embodiment of this disclosure;

[0024] Figure 4 This is a schematic diagram of the first and second difference images according to an embodiment of this disclosure;

[0025] Figure 5 This is a schematic diagram of a first scratch image and a second scratch image according to an embodiment of this disclosure;

[0026] Figure 6 This is a comparative schematic diagram of the video image to be processed and the scratch repair image according to an embodiment of this disclosure;

[0027] Figure 7 This is a flowchart illustrating the dead pixel repair steps according to an embodiment of the present disclosure;

[0028] Figure 8 This is a comparative diagram of the video image to be processed and the image with damaged pixels repaired, according to an embodiment of this disclosure.

[0029] Figure 9 This is a schematic diagram of the structure of a multi-scale cascaded network according to an embodiment of the present disclosure;

[0030] Figure 10 This is a schematic diagram of the input image of a multi-scale cascaded network according to an embodiment of the present disclosure;

[0031] Figure 11This is a schematic diagram comparing the video image to be processed and the output image of the multi-scale cascaded network in an embodiment of this disclosure;

[0032] Figure 12 This is a schematic diagram of the sub-network structure in an embodiment of this disclosure;

[0033] Figure 13 This is a schematic diagram comparing the output image and the post-processed image of the multi-scale cascaded network according to an embodiment of this disclosure;

[0034] Figure 14 This is a flowchart illustrating the noise reduction steps in an embodiment of the present disclosure.

[0035] Figure 15 This is a schematic diagram of the process of obtaining a motion mask using an optical flow network according to an embodiment of the present disclosure;

[0036] Figure 16 This is a schematic diagram illustrating the specific process of obtaining a motion mask using an optical flow network according to an embodiment of the present disclosure.

[0037] Figure 17 This is a schematic diagram of a training method for a denoising network according to an embodiment of the present disclosure;

[0038] Figure 18 This is a schematic diagram illustrating an implementation method of a denoising network according to an embodiment of this disclosure;

[0039] Figure 19 This is a schematic diagram of the video image to be processed according to an embodiment of this disclosure;

[0040] Figure 20 for Figure 19 A magnified view of a portion of the video image to be processed;

[0041] Figure 21 for Figure 19 The target non-motion mask M corresponding to the video image to be processed in the image;

[0042] Figure 22 The image is the result of denoising the video image to be processed using the denoising network of this embodiment.

[0043] Figure 23 This is a flowchart illustrating the color cast correction steps in an embodiment of the present disclosure.

[0044] Figure 24 This is a schematic diagram of the structure of an image processing apparatus according to an embodiment of the present disclosure. Detailed Implementation

[0045] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0046] This disclosure provides an image processing method, including:

[0047] Perform at least one of the following steps on the video image to be processed:

[0048] Scratch repair steps: Perform scratch removal processing on the video image to be processed to obtain a first image; perform difference operation on the video image to be processed and the first image to obtain a difference image; process the difference image to obtain a scratch image that retains only the scratches; obtain a scratch repair image based on the video image to be processed and the scratch image;

[0049] Defective pixel repair steps: Acquire N1 consecutive video images, wherein each N1 video image includes a video image to be processed, at least one video image preceding the video image to be processed, and at least one video image following the video image to be processed; Filter the video image to be processed based on the N1 video images to obtain a defective pixel repair image; Perform artifact removal processing on the defective pixel repair image based on the at least one video image preceding the video image to be processed and at least one video image following the video image to be processed to obtain an artifact repair image;

[0050] Denoising step: A denoising network is used to denoise the video image to be processed; wherein, the denoising network is obtained by the following training method: a target non-motion mask M is obtained based on N2 consecutive video images, wherein the N2 video images include the video image to be denoised; the denoising network to be trained is trained based on the N2 video images and the target non-motion mask M to obtain the denoising network.

[0051] Color cast correction steps: Determine the target color cast values ​​for each RGB channel of the video image to be processed; perform color balance adjustment on the video image to be processed according to the target color cast values ​​to obtain a first corrected image; perform color migration on the first corrected image according to a reference image to obtain a second corrected image.

[0052] It should be noted that at least one of the above four steps can be performed on the video image. If multiple steps need to be performed, the order of execution of these steps is not limited. For example, if scratch repair and dead pixel repair need to be performed, scratch repair can be performed on the video image first, followed by dead pixel repair, or dead pixel repair can be performed first, followed by scratch repair.

[0053] In this embodiment of the disclosure, at least one step of scratch repair, dead pixel repair, noise reduction, or color cast correction is performed on the video image to improve the display effect of the video image.

[0054] The four steps described above will be explained below.

[0055] I. Scratch Repair

[0056] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the scratch repair steps in an embodiment of the present disclosure, including:

[0057] Step 11: Perform scratch removal processing on the video image to be processed to obtain the first image;

[0058] In this embodiment of the disclosure, the video image to be processed can be filtered to remove scratches, such as median filtering.

[0059] Step 12: Perform a difference operation on the video image to be processed and the first image to obtain a difference image;

[0060] Step 13: Process the difference image to obtain a scratch image that retains only the scratches;

[0061] Step 14: Obtain the scratch repair image based on the video image to be processed and the scratch image.

[0062] In this embodiment, scratch removal processing is performed on a single frame of the video image to be processed to obtain an image with scratches removed. Difference operation is performed on the video image to be processed and the scratch-removed image to obtain a difference image containing scratches and image details. Then, the difference image is processed again to filter out image details and retain scratches to obtain a scratch image. Finally, a scratch-repaired image with scratches removed is obtained based on the video image to be processed and the scratch image. This process removes scratches without affecting image clarity.

[0063] The following sections will provide a detailed explanation of each of the above steps.

[0064] (1) Step 11:

[0065] In this embodiment of the disclosure, optionally, the scratch removal processing of the video image to be processed includes: performing median filtering processing on the video image to be processed according to at least one of the filter type and the scratch type in the video image to be processed, to obtain a scratch-removed image.

[0066] Optionally, performing median filtering on the video image to be processed based on at least one of the filter type and the scratch type in the video image to be processed includes: selecting a corresponding filter type based on the scratch type in the video image to be processed, and performing median filtering on the video image to be processed, wherein:

[0067] When the scratches in the video image to be processed are vertical scratches, a median filter in the horizontal direction is used to perform median filtering on the video image to be processed.

[0068] When the scratches in the video image to be processed are horizontal scratches, median filtering is performed on the video image to be processed using a median filter in the vertical direction.

[0069] In other words, in this embodiment of the disclosure, the median filter to be used can be determined based on the direction of the scratches in the video image to be processed.

[0070] Of course, in some other embodiments of this disclosure, the median filter may not be changed, but the video image to be processed may be rotated so that the scratches in the video image to be processed match the median filter.

[0071] That is, performing median filtering on the video image to be processed based on at least one of the filter type and the scratch type in the video image to be processed includes: performing different preprocessing on the video image to be processed based on the filter type and the scratch type in the video image to be processed, and performing median filtering on the preprocessed video image to be processed, wherein:

[0072] When a horizontal median filter is used, and the scratches in the video image to be processed are non-vertical scratches, the video image to be processed is rotated to convert the scratches into vertical scratches; non-vertical scratches include horizontal scratches and diagonal scratches. Of course, if the scratches in the video image to be processed are vertical scratches, then there is no need to rotate the image to be processed.

[0073] If a vertical median filter is used, and the scratches in the video image are not horizontal, the video image to be processed is rotated to convert the scratches into horizontal scratches. Non-horizontal scratches include vertical scratches and diagonal scratches. Of course, if the scratches in the video image to be processed are horizontal, there is no need to rotate the image.

[0074] Furthermore, if the video image to be processed contains both horizontal and vertical scratches, median filters in both the horizontal and vertical directions can be applied simultaneously for median filtering. For example, horizontal median filtering can be performed first, followed by vertical median filtering, or vice versa.

[0075] In this embodiment of the disclosure, scratch removal processing of the video image to be processed includes: applying a median filter of size 1×k and / or k×1 to the video image to be processed for median filtering; wherein, the 1×k median filter is a horizontal median filter, and the k×1 median filter is a vertical median filter.

[0076] For example, a 1×k median filter is used to filter the video image I to be processed, resulting in a first image I. median .

[0077] I median =M 1×k (I)

[0078] Among them, M 1×k (x) represents filtering x using a median filter of size 1×k.

[0079] The following explains how to determine the size of the filter.

[0080] In this embodiment of the present disclosure, optionally, before performing scratch removal processing on the video image to be processed, the method further includes: sequentially increasing the k of the median filter from a preset value to perform median filtering on the video image to be processed to obtain a second image; and determining the final value of k based on the filtering effect of the second image.

[0081] For example, assuming the median filter used is a 1×k horizontal median filter, you can first set k to 3 (preset value), that is, use a 1×3 median filter to filter the video image to be processed, obtain the second image, and observe the filtering effect of the second image. If the scratch removal effect is not obvious, you can set the value of k to 4 (or other values ​​greater than 3), that is, use a 1×4 median filter to filter the video image to be processed, obtain the second image, and observe the filtering effect of the second image. If the scratch removal effect is not obvious, increase the value of l = k again until there are no obvious scratches in the second image.

[0082] Of course, in some other embodiments of this disclosure, the value of k can also be determined directly according to the thickness of the scratch. For example, for an image with a resolution of 2560×1440, k can be set to a value below 11.

[0083] The following example illustrates this. Please refer to it. Figure 3 , Figure 3 Image (a) shows the video image to be processed and a magnified view of a portion of it. As can be seen from the image, the video image to be processed contains vertical scratches. Please refer to [reference needed]. Figure 2 A horizontal median filter can be used to remove scratches from the video image to be processed, resulting in the first image. See also... Figure 3 (b) Figure 3 (b) shows the first image and a magnified view of a portion of the first image. It can be seen from the image that the vertical scratches have been removed from the first image, but the first image is more blurry compared to the image to be processed.

[0084] (2) Step 12:

[0085] In this embodiment of the disclosure, performing a difference operation on the video image to be processed and the first image to obtain a difference image includes: performing a difference operation on the video image to be processed and the first image to obtain a first difference image and / or a second difference image, wherein the first difference image is obtained by subtracting the first image from the video image to be processed, and the second difference image is obtained by subtracting the video image to be processed from the first image.

[0086] In this embodiment of the disclosure, subtracting two images indicates that the pixels at corresponding positions in the two images are subtracted.

[0087] The first difference image is a white texture map containing image details and scratches, while the second difference image is a black texture map containing image details and scratches.

[0088] In this embodiment of the application, the first difference image can also be called a positive residual image, and the second difference image can also be called a negative residual image.

[0089] For example, the video image I to be processed is compared with the first image I. median Subtracting each other, we obtain positive residuals Err. white With negative residual Errb lack The calculation formula is as follows, where both positive and negative residuals are positive values.

[0090] Err white =II median

[0091] Err black =I median -I

[0092] Still with Figure 2For example, subtract the video image to be processed from the first image. The first image is obtained by multiplying the video image to be processed by 1 (×1) and the first image by -1 (×-1), then adding them together. This results in the first difference image (the video image to be processed minus the first image). The second difference image is obtained by multiplying the video image to be processed by -1 (×-1) and the first image by 1 (×1), then adding them together. This results in the second difference image (the video image to be processed minus the first image). Please refer to [reference needed]. Figure 4 , Figure 4 Image (a) shows the first difference image and a magnified view of a portion of the first difference image. Figure 4 (b) in the image is a magnified view of the second difference image and a local part of the second difference image. It can be seen that the first difference image is a white texture image containing image details and scratches, and the second difference image is a black texture image containing image details and scratches.

[0093] In the above example, the first difference image and the second difference image are calculated simultaneously. Of course, in other embodiments of this disclosure, it is not excluded that only the first difference image or only the second difference image is calculated.

[0094] (3) Step 13:

[0095] In this embodiment of the disclosure, processing the difference image to obtain a scratch image that retains only the scratches includes: processing the first difference image to obtain a first scratch image that retains only the scratches; and / or, processing the second difference image to obtain a second scratch image that retains only the scratches;

[0096] A scratch image is an image that has had its details filtered out, leaving only the scratches.

[0097] In this embodiment of the application, the first scratch image can also be called a positive scratch image, and the second scratch image can also be called a negative scratch image.

[0098] In this embodiment of the disclosure, optionally, processing the first difference image to obtain a first scratch image that retains only the scratches includes:

[0099] The first difference image is filtered by median filtering in the vertical direction and median filtering in the horizontal direction, respectively, to obtain the first vertically filtered image and the first horizontally filtered image.

[0100] If the scratch in the video image to be processed is a vertical scratch, the first scratch image is obtained by subtracting the first horizontally filtered image from the first vertically filtered image.

[0101] If the scratch in the video image to be processed is a horizontal scratch, the first scratch image is obtained by subtracting the first vertically filtered image from the first horizontally filtered image.

[0102] For example, suppose the scratches in the video image to be processed are vertical scratches, for Err white (First difference image) Perform median filtering in the vertical direction and median filtering in the horizontal direction to obtain the first vertically filtered image M. n×1 (Err white ) and the first level filtered image M 1×n (Err white Then, the first vertically filtered image M n×1 (Err white Image M after subtracting the first level of filtering 1×n (Err white The filtered first scratch image L was obtained respectively. white At this point, the first scratch image is represented by a positive number:

[0103] L white =M n×1 (Err white )-M 1×n (Err white )

[0104] Among them, M n×1 (Err white ) indicates that median filtering is performed on the first difference image in the vertical direction, M 1×n (Err white This indicates that median filtering is applied to the first difference image in the horizontal direction.

[0105] Processing the second difference image to obtain a second scratch image that retains only the scratches includes:

[0106] The second difference image is filtered by median filtering in the vertical direction and median filtering in the horizontal direction, respectively, to obtain the second vertically filtered image and the second horizontally filtered image.

[0107] If the scratch in the video image to be processed is a vertical scratch, the second scratch image is obtained by subtracting the second horizontally filtered image from the second vertically filtered image;

[0108] If the scratch in the video image to be processed is a horizontal scratch, the second scratch image is obtained by subtracting the second vertically filtered image from the second horizontally filtered image.

[0109] For example, suppose the scratches in the video image to be processed are vertical scratches, for Err black (Second difference image) Perform median filtering in the vertical direction and median filtering in the horizontal direction to obtain the second vertically filtered image M. n×1 (Err black The second level filtered image M1×n (Err black Then, the second vertically filtered image M n×1 (Err black Image M after subtracting the second level filter 1×n (Err black The filtered second scratch image L was obtained respectively. black At this point, the second scratch image is represented by a positive number:

[0110] L black =M n×1 (Err black )-M 1×n (Err black )

[0111] Among them, M n×1 (Err black M represents median filtering of the second difference image in the vertical direction. 1×n (Err black This indicates that median filtering is applied to the second difference image in the horizontal direction.

[0112] In this embodiment of the disclosure, since the length of the scratch is usually greater than the length of the lines in the image details, in order to filter out the image details and retain only the scratch, the values ​​of n in the vertical median filter and the horizontal median filter can be set to be large. For example, they can be set to half the average length of the scratch. For example, if the maximum length of the scratch is 180, the value of n can be set to between 80 and 100.

[0113] Still with Figure 2 For example, the first and second difference images are subjected to median filtering in the vertical and horizontal directions, respectively, to obtain the vertically filtered image and the horizontally filtered image. Since... Figure 2 If the scratches in the video image to be processed are vertical scratches, then the horizontally filtered image needs to be subtracted from the vertically filtered image to obtain the first scratch image and the second scratch image. Please refer to... Figure 5 , Figure 5 Image (a) is a magnified view of the first scratch image and a portion thereof. Figure 5 Image (b) is a magnified view of the second scratch image and the second scratch image. It can be seen that the first scratch image contains white scratches, and the second scratch image contains black scratches.

[0114] (4) Step 14:

[0115] Obtaining a scratch repair image based on the video image to be processed and the scratch image includes: performing calculations on the video image to be processed, the first scratch image, and / or the second scratch image to obtain the scratch repair image.

[0116] In this embodiment of the disclosure, the scratch repair image can be calculated using the following formula:

[0117] I deline =IL white -(L black ×-1)=IL white +L black

[0118] Among them, I deline I is the scratch repair image, L is the video image to be processed. white For the first scratch image, L black Let L be the second scratch image. In the formula, since the second scratch image L... black Since it is a positive value, it needs to be multiplied by -1 to restore it to a negative value.

[0119] In this embodiment of the disclosure, when calculating the scratch repair image, only the first scratch image or only the second scratch image may be used.

[0120] Subtracting the scratched image from the original image ensures scratch removal while preserving image sharpness. Please refer to this method. Figure 6 (a) shows the video image to be processed and a magnified view of its portion; (b) shows the scratch-repaired image and a magnified view of its portion. Figure 6 As can be seen, the scratches in the scratch-repaired image have been removed, and the image clarity has not changed compared to the video image to be processed.

[0121] II. Dead Pixel Repair

[0122] Please refer to Figure 7 , Figure 7 This is a flowchart illustrating the dead pixel repair steps in an embodiment of the present disclosure, including:

[0123] Step 71: Acquire N1 consecutive video images, wherein the N1 video images include the video image to be processed, at least one video image before the video image to be processed, and at least one video image after the video image to be processed;

[0124] In this embodiment of the disclosure, N1 is a positive integer greater than or equal to 3, and can be set as needed, for example, it can be 3.

[0125] Step 72: Based on at least one frame of video image before and at least one frame of video image after the video image to be processed, perform filtering processing on the video image to be processed to obtain a bad pixel repair image;

[0126] Step 73: Perform artifact removal processing on the bad pixel repair image based on at least one frame of video image before and at least one frame of video image after the video image to be processed to obtain an artifact repair image.

[0127] Since multiple frames of video images are used to process the image to remove bad pixels, the resulting image with repaired bad pixels will introduce motion artifacts. At this point, the problem of removing bad pixels is transformed into the problem of artifact removal, that is, the filtering for removing bad pixels and the multi-scale cascaded network for artifact removal are combined to achieve the repair of bad pixels in the video image.

[0128] In other words, bad pixel repair can include two processes: removing bad pixels and removing artifacts, which will be introduced separately below.

[0129] (1) Remove dead pixels

[0130] The defective pixel repair method of this disclosure can be applied to the repair of defects in film, and of course, it is not excluded from the repair of defects in other types of video images.

[0131] Defective pixels are a common type of damage to film. They are white or black spots that form on the film surface due to the loss of gel or the presence of stains during storage. Defective pixels on film generally have the following three characteristics:

[0132] 1) The standard deviation of pixel grayscale within a bad pixel is very small, and the grayscale within the block is basically consistent;

[0133] 2) The grayscale of defective pixels is discontinuous in both the temporal and spatial domains. Because this damage is randomly distributed within a frame, defective pixels are unlikely to reappear at the same location in two adjacent frames, thus exhibiting a kind of impulsive damage in the temporal domain. Within a single image frame, the grayscale of the defective pixel area generally differs significantly from the surrounding background grayscale, making it perceptible to the human eye.

[0134] 3) Spatial adjacency. That is, if a pixel is within a bad pixel area, then the pixels around it are also very likely to be in the bad pixel area.

[0135] This disclosure mainly focuses on the second characteristic of bad pixels to achieve their repair. Since bad pixels are discontinuous in the time domain, and the pixel values ​​at the same position in adjacent frames are often similar, this disclosure uses the content of adjacent frames to repair bad pixels in the current frame.

[0136] In this embodiment of the disclosure, optionally, median filtering is performed on the video image to be processed based on at least one frame of video image preceding and at least one frame of video image following the video image to be processed to obtain a bad pixel repair image. For example, assuming N1 equals 3, for the current video image to be processed I... t and adjacent frames I t-1 I t+1 The median is calculated pixel by pixel. Since the pixel values ​​at the same position in adjacent frames within the same scene generally do not differ too much, during the median calculation process, bad pixel areas with large differences in grayscale values ​​from the surrounding images will be replaced by pixels from the previous or next frame, thereby eliminating bad pixels in the intermediate frame image.

[0137] Of course, in other embodiments of this disclosure, other filtering methods, such as mean filtering, are also possible to remove bad pixels from the image to be processed.

[0138] like Figure 8 As shown, Figure 8 In the image, (a) is the video image to be processed with bad pixels, and (b) is the image with bad pixels repaired. Figure 8 As can be seen, median filtering can fill in bad pixels in the intermediate frame using the content of the preceding and following frames, but it also introduces error information from the preceding and following frames into the intermediate frame, resulting in motion artifacts. Figure 8 In the image, (a) the small image in the upper right corner is a magnified view of the bad pixel area, (b) the small image in the upper right corner is a magnified view of the bad pixel repair area, and the small image in the lower right corner is a magnified view of the motion artifact portion. Therefore, in this embodiment of the present disclosure, it is also necessary to perform artifact removal processing on the bad pixel repair image to eliminate the artifacts generated by filtering.

[0139] (2) Remove artifacts

[0140] In this embodiment of the disclosure, optionally, performing artifact removal processing on the bad pixel repair image based on at least one frame of video image preceding and at least one frame of video image following the video image to be processed to obtain the artifact repair image includes:

[0141] The defective pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the video image to be processed are downsampled N3-1 times to obtain N3-1 downsampled images at different resolutions. The N3 images at different resolutions are then input into a multi-scale cascaded network for artifact removal processing to obtain an artifact-repaired image. The N3 images at different resolutions include: the defective pixel repair image, at least one frame of video image before the video image to be processed, at least one frame of video image after the video image to be processed, and the N3-1 downsampled images at different resolutions. The multi-scale cascaded network includes N3 cascaded sub-networks, and the images processed by the N cascaded sub-networks are generated based on the N3 images at different resolutions.

[0142] Where N3 is a positive integer greater than or equal to 2.

[0143] Further, optionally, images at N3 resolutions can be input into a multi-scale cascaded network for artifact removal processing to obtain artifact-repaired images, including:

[0144] For the first sub-network in the N3 cascaded sub-networks: the bad pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the bad pixel repair image are downsampled by A times to obtain N1 frame first downsampled image. The N1 frame first downsampled image is then stitched together with itself to obtain a first stitched image. The first stitched image is then input into the first sub-network to obtain a first output image.

[0145] For the intermediate sub-network between the first and last sub-networks: the output image of the previous sub-network is upsampled to obtain a first upsampled image; the bad pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the video image are downsampled by a factor of B to obtain N1 frames of second downsampled images. The second downsampled image and the first upsampled image have the same scale. The two sets of images are stitched together to obtain a second stitched image. The second stitched image is input into the intermediate sub-network to obtain a second output image. Among the two sets of images, one set is the N1 frame of second downsampled images, and the other set includes: other downsampled images in the N1 frame of second downsampled images except for the downsampled image corresponding to the bad pixel repair image, and the first upsampled image.

[0146] For the last sub-network: the output image of the previous sub-network is upsampled to obtain a second upsampled image. The second upsampled image has the same scale as the video image to be processed. The two sets of images are stitched together to obtain a third stitched image. The third stitched image is input into the last sub-network to obtain the artifact restoration image. Among the two sets of images, one set is the N1 frame video image, and the other set includes: other images in the N1 frame video image except for the bad pixel restoration image, and the second upsampled image.

[0147] In this embodiment of the disclosure, optionally, the aforementioned sub-network can be the encoder-decoder resblock network structure proposed in SRN. Of course, in other embodiments of this disclosure, other networks may also be used, and this disclosure does not limit the scope of the network.

[0148] To improve network performance, in this embodiment of the disclosure, optionally, the N3 cascaded sub-networks have the same structure but different parameters.

[0149] In this embodiment of the disclosure, optionally, N3 equals 3, A equals 4, and B equals 2.

[0150] The following example illustrates this.

[0151] like Figure 9 As shown, the multi-scale cascaded network consists of three sub-networks with the same structure cascaded together, and the three sub-networks process inputs at different scales (i.e., resolutions).

[0152] The input to the multi-scale cascaded network is three consecutive frames of image I. t-1 、I′ t and I t+1 Please refer to Figure 10 , where (b) is I′ t , is the image with artifacts after the above processing, (a) is I t-1 It is the video image I to be processed. t The previous frame image, (c) is I t+1 It is the video image I to be processed. t The next frame of the image, the output of the multi-scale cascaded network and I′ t Corresponding image with artifacts removed

[0153] The computational steps for multi-scale cascaded networks are as follows:

[0154] 1) Input the three frames of image I t-1 、I′ t and I t+1 Each image was downsampled by a factor of 4, resulting in a first downsampled image with a resolution reduced by a factor of 4. and The three downsampled images are each stitched together with the image itself to obtain the first stitched image, which is then input into Network 1. Network 1 outputs a first output image.

[0155] The above stitching refers to stitching in the fourth dimension. Each image is a three-dimensional array with an array structure of H×W×C, i.e., height, width, and channel. The fourth dimension is the channel dimension.

[0156] In this embodiment of the disclosure, optionally, a bicubic interpolation method is used to perform I. t-1 、I′ t and I t+1 Downsample by a factor of 4. Of course, other downsampling methods can also be used.

[0157] After image downsampling, its artifacts will also be reduced, which is more conducive to the network's elimination and repair of artifacts. Therefore, the input image sizes of the three sub-networks are 1 / 4, 1 / 2, and 1 of the original input image size, respectively. Here, the image is to be input into the first sub-network, so a 4x downsampling is performed.

[0158] 2) Output of Network 1 Upsampled by 2 times, resulting in the first upsampled image. and image I t-1 、I′ t and I t+1 The second downsampled image was obtained by downsampling by a factor of 2. and Take these three frames of the second downsampled image as a set of inputs, and in addition, and As one set of inputs, the two sets of inputs are concatenated to obtain a second concatenated image, which is then input into Network 2. Network 2 outputs a frame of the second output image.

[0159] In this embodiment of the disclosure, optionally, the output of network 1 is obtained by bicubic interpolation. Upsample by a factor of 2. Of course, other methods can also be used for upsampling.

[0160] In this embodiment of the disclosure, optionally, image I is obtained by bicubic interpolation. t-1 、I′ t and I t+1 Downsample by a factor of 2. Of course, other downsampling methods can also be used.

[0161] 3) Output of Network 2 Upsampled by 2 times to obtain the second upsampled image Image I t-1 、I′ t and I t +1 As a set of inputs, in addition, I t-1 , and I t+1 As one set of inputs, the two sets of inputs are concatenated to obtain a third concatenated image, which is then input into network 3. Network 3 outputs one frame of image. This is the final result of the network.

[0162] In this embodiment of the disclosure, optionally, the output of network 2 is obtained by bicubic interpolation. Upsample by a factor of 2. Of course, other methods can also be used for upsampling.

[0163] Please refer to Figure 11 , Figure 11 In the image, (a) represents a portion of the original video image to be processed, (b) represents a portion of the image after filtering for bad pixel restoration, which contains motion artifacts caused by filtering, and (c) represents a portion of the restored image output by the multi-scale cascaded network. Figure 11 As can be seen, motion artifacts have been eliminated.

[0164] In this embodiment of the disclosure, optionally, network 1, network 2 and network 3 can be the encoder-decoder resblock network structure proposed in SRN. In order to improve the network performance, the parameters of these three sub-networks are not shared in this embodiment of the disclosure.

[0165] In this embodiment of the disclosure, optionally, each of the sub-networks includes a plurality of 3D convolutional layers, a plurality of 3D deconvolutional layers, and a plurality of 3D average pooling layers.

[0166] Please refer to Figure 12 , Figure 12 This is a schematic diagram of the sub-network structure in an embodiment of this disclosure. The sub-network is an encoder-decoder resblock network, where Conv3d(n32f5s1) represents a 3D convolutional layer with 32 filters (n), a filter size (f) of (1×5×5), and a filter stride (s) of 1; DeConv3d(n32f5s2) represents a 3D deconvolutional layer with 32 filters (n), a filter size (f) of (1×5×5), and a filter stride (s) of 2; AvgPool3D(f2s1) represents a 3D average pooling layer with a kernel size (f) of (2×1×1) and a stride (s) of 1; and (B,3,H,W,32) above the arrow represents the size of the feature map output by the current layer. Here, (B,3,H,W,32) refers to the size of the intermediate output of each layer of the network. Each layer outputs a 5-dimensional array, and (B,3,H,W,32) refers to the structure of the array, namely B×3×H×W×32, where B is the batch size.

[0167] like Figure 12 As shown, the sub-network contains 16 3D convolutional layers and 2 3D deconvolutional layers. Information between adjacent frames is fused through 3D average pooling layers. The output feature maps of the 3rd convolutional layer and the 1st deconvolutional layer are fused through 3D average pooling layers, then summed pixel by pixel, and the result is used as the input of the 14th convolutional layer. The output feature maps of the 6th convolutional layer and the 2nd deconvolutional layer are fused through 3D average pooling layers, then summed pixel by pixel, and their sum is used as the input of the 12th convolutional layer. Finally, the network outputs the final image after removing artifacts through a convolutional layer with one filter.

[0168] In this embodiment of the disclosure, the multi-scale cascaded network is optionally obtained using the following training method:

[0169] Step 1: Obtain N1 consecutive training images, wherein the N1 training images include the training image to be processed, at least one training image before the training image to be processed, and at least one training image after the training image.

[0170] Step 2: Filter the training image to be processed based on the N1 frames of training images to obtain the first training image;

[0171] Step 3: Based on the first training image, at least one training image before the training image to be processed, and at least one training image after the training image to be processed, train the multi-scale cascaded network to be trained to obtain the trained multi-scale cascaded network.

[0172] In this embodiment of the disclosure, optionally, when training the multi-scale cascaded network to be trained, the total loss used includes at least one of the following: image content loss, color loss, edge loss, and perceptual loss.

[0173] Image content loss is primarily used to improve the fidelity of the output image. Optionally, in this embodiment, the image content loss can be calculated using an L1 loss function or a mean squared error loss function.

[0174] Optionally, the L1 loss is calculated using the following formula:

[0175]

[0176] Among them, l content For L1 loss, For the artifact removal training image, y i Let n be the first training image, and n be the number of images in a batch.

[0177] The color loss function corrects image color by applying Gaussian blur to the texture and content of both the artifact-removed training image and the target image, preserving only the color information. Optionally, in this embodiment, the color loss is calculated using the following formula:

[0178]

[0179] Among them, l color For the color loss, For the artifact removal training image, y i Let be the first training image, n be the number of images in a batch, and Blur(x) be the Gaussian blur function.

[0180] The edge loss function primarily improves the accuracy of the contour information in the artifact-removed training image by calculating the difference in edge information between the artifact-removed training image and the target image. In this embodiment, a Holistically-Nested Network (HED) can be used to extract the edge information of the image. Optionally, in this embodiment, the edge loss is calculated using the following formula:

[0181]

[0182] Among them, l edge For the edge loss, For the artifact removal training image, y i The first training image is n, where n is the number of images in a batch, and H is the number of images in a batch. j (x) represents the image edge map extracted by the j-th layer of the HED network.

[0183] In this embodiment of the disclosure, high-level features extracted by the VGG network are used to calculate the perceptual loss function to measure the semantic difference between the output image and the target image. Optionally, the perceptual loss is calculated using the following formula:

[0184]

[0185] Among them, l feature For the perceived loss, For the artifact removal training image, y i The first training image is denoted as n, where n is the number of images in a batch. This represents the image feature map extracted from the j-th layer of the VGG network.

[0186] In this embodiment of the disclosure, optionally, the total loss is equal to the weighted sum of image content loss, color loss, edge loss, and perceptual loss.

[0187] Optionally, the total loss is calculated using the following formula:

[0188] L = l content +λ1l color +λ2l edge +λ3l feature

[0189] Where λ1 = 0.5, λ2 = 10 -2 λ3=10 -4 Of course, in other embodiments of this disclosure, the weights of each loss may not be limited to this.

[0190] In this embodiment of the disclosure, training data provided by the video temporal super-resolution track of the 2020-AIM competition can be used for training. This training set contains 240 sets of frame sequences, each set containing 181 clear images of 1280×720 resolution. The reasons for using this training dataset are as follows:

[0191] a) Each group of 181 images in this dataset was taken in the same scene. Each training session uses images from the same scene, which can avoid the large differences in image content between different scenes and cause interference.

[0192] b) This training dataset is used for training video temporal super-resolution tracks. Objects in the same scene have appropriate motion between adjacent frames, which meets the requirement of generating artifacts in the simulation training data disclosed herein.

[0193] c) The images in this training dataset are relatively clean, free of noise, and have a high resolution, which is beneficial for the network to generate clearer images.

[0194] Since the video images to be processed in this embodiment have had their bad pixels repaired before being input into the multi-scale cascaded network, and the main purpose of network training is to remove artifacts generated by filtering, the same filtering operation is only performed on the training dataset when generating simulation data, and there is no need to simulate the generation of bad pixels.

[0195] The network model of this disclosure can be trained on the Ubuntu 16.04 system, compiled using the Python language, and based on the TensorFlow deep learning framework and open-source image and video processing tools such as OpenCV and FFmpeg.

[0196] In this embodiment of the disclosure, optionally, training the multi-scale cascaded network to be trained based on the first training image, at least one training image preceding the training image to be processed, and at least one training image following the training image to be processed includes:

[0197] A random image block is cropped from the first training image. An image block is cropped from the same position in at least one training image before the training image to be processed and at least one training image after the training image to be processed, respectively, to obtain N1 frame image blocks.

[0198] The N1 frame image blocks are input into the multi-scale cascaded network to be trained for training.

[0199] In this embodiment, the Adam optimization algorithm can be used to optimize the network parameters. The learning rate of the Adam algorithm is set to 10⁻⁴. During network training, three consecutive training images are taken from the training dataset for median filtering preprocessing. Then, a 512×512 image block is randomly cropped from the middle frame, and corresponding image blocks are cropped from the same positions in the preceding and following frames as input for one network iteration. When all images in the training dataset have been read once, one epoch is completed. Every 10 epochs (one epoch is the process of training all training samples once), the learning rate of Adam is reduced to 0.8 times the original rate.

[0200] In this embodiment, downsampling is performed on the cropped image patches, which is a method for augmenting the dataset. That is, multiple image patches can be randomly cropped multiple times from the same image for network training, thereby increasing the number of images used for network training. Random cropping allows different locations within the same image to be used. Furthermore, cropping into image patches also reduces the image resolution, decreases the amount of data processed by the network, and improves processing speed.

[0201] (3) Post-processing

[0202] In this embodiment of the disclosure, the image after the bad pixel removal and artifact removal processes not only repairs the bad pixels in the image but also removes the artifacts caused by the motion of objects. However, the overall clarity of the image output by the network still differs from that of the original video image to be processed. Therefore, by performing filtering operations on the video image to be processed, the bad pixel repair image, and the image repaired by the multi-scale cascaded network, the details in the original video image to be processed are added to the repaired image to improve the clarity of the repaired image.

[0203] Therefore, in this embodiment of the present disclosure, optionally, after obtaining the artifact-repaired image, the method further includes: filtering the artifact-repaired image based on the video image to be processed and the bad pixel repair image to obtain an output image artifact-repaired image.

[0204] Optionally, the artifact restoration image can be subjected to median filtering based on the video image to be processed and the bad pixel restoration image to obtain an output image.

[0205] Please refer to Figure 13 , Figure 13 In the image, (a) is the output image of the multi-scale cascaded network, and (b) is the post-processed image. Figure 13 As can be seen, the clarity of the post-processed image is significantly higher than that of the output image of the multi-scale cascaded network.

[0206] III. Noise Reduction

[0207] Please refer to Figure 14 , Figure 14 This is a flowchart illustrating the denoising step in an embodiment of the present disclosure. The denoising step includes:

[0208] Step 141: Use a denoising network to denoise the video image to be processed; wherein, the denoising network is obtained by the following training method: obtain a target non-motion mask based on N2 consecutive video images, wherein the N2 video images include the video image to be denoised; train the denoising network to be trained based on the N2 video images and the target non-motion mask to obtain the denoising network.

[0209] In this embodiment, a blind denoising technique is used when training the denoising network. This technique does not require paired training datasets; only the sequence of video frames to be denoised needs to be input. Using a non-motion mask, temporal denoising is performed only on the non-motion data as a reference image. This is suitable for training denoising networks without a clear reference image. Furthermore, it is applicable to various types of video noise removal, without needing to consider the noise type; only a subset of video frames is needed to learn the denoising network.

[0210] In this embodiment of the disclosure, optionally, the denoising network to be trained is trained based on the N2 frames of video images and the target non-motion mask to obtain the denoising network, including:

[0211] Step 151: Obtain a reference image based on the N2 frame video images and the target non-motion mask;

[0212] The reference image is equivalent to the ground truth of the video image to be denoised, i.e., the image without noise.

[0213] Step 152: Input the video image to be denoised into the denoising network to be trained to obtain the first denoised image;

[0214] Step 153: Obtain a second denoised image based on the first denoised image and the target non-motion mask;

[0215] Step 154: Determine the loss function of the denoising network to be trained based on the reference image and the second denoised image, and adjust the parameters of the denoising network to be trained based on the loss function to obtain the denoising network.

[0216] In this embodiment of the disclosure, an optical flow method can be used to obtain a non-motion mask for a video image, which will be described in detail below.

[0217] Optical flow is the instantaneous velocity of pixels moving on the imaging plane of a spatially moving object. The optical flow method utilizes the temporal changes of pixels in an image sequence and the correlation between adjacent frames to find the correspondence between the previous and current frames, thereby calculating the motion information of objects between adjacent frames.

[0218] Before obtaining the non-motion mask, the motion mask must first be obtained. The motion mask refers to the motion information in the image. The non-motion mask refers to the information in the image other than the motion mask, i.e., the non-motion information.

[0219] Please refer to Figure 15 , Figure 15 This is a schematic diagram of the process of obtaining a motion mask using an optical flow network according to an embodiment of the present disclosure. Two video images F and Fr are input into the optical flow network to obtain an optical flow graph, and then the motion mask is obtained through the optical flow graph.

[0220] In this embodiment, the choice of optical flow network is not limited. It can be any currently open-source optical flow network, such as flownet or flownet2, or a traditional optical flow algorithm (not deep learning), such as TV-L1flow. It is only necessary to use the optical flow algorithm to obtain the optical flow graph.

[0221] Please refer to Figure 16 , Figure 16 This is a schematic diagram illustrating the specific process of obtaining a motion mask using an optical flow network according to an embodiment of this disclosure. Assume the size of the two frames of images input to the optical flow network is (720, 576). The optical flow network outputs two frames of optical flow maps, with the size of the optical flow maps also being (720, 576). These two frames represent the first optical flow map (representing vertical motion information in two consecutive frames) and the second optical flow map (representing horizontal motion information in two consecutive frames). The last X-X1 rows of the first optical flow map are subtracted from the first X-X1 rows (in the example, the last three rows are subtracted from the first three rows) to obtain the first difference map. The last Y-Y1 columns of the second optical flow map are subtracted from the first Y-Y1 columns (in the example, the last three columns are subtracted from the first three columns) to obtain the second difference map. The last X1 rows of the first difference map are padded with 0, and the last Y1 column of the second difference map is padded with 0, resulting in a map of the same size as the optical flow map. Finally, the two difference maps are added together. Wherein, || represents absolute value, >T means that after the absolute value operation, pixels with values ​​greater than T are assigned a value of 1, and all other values ​​less than or equal to T are assigned a value of 0, where T is a preset threshold. A binarized image is obtained after thresholding. Optionally, dilation can also be performed on the binarized image. The specific method of dilation is to find the pixel values ​​of 1 in the binarized image, and according to the dilation kernel, set the pixel positions in the binarized image corresponding to the positions where 1 is found in the kernel to 1. This is in the embodiment of this disclosure. Figure 16 The example given is a 3x3 dilation kernel. If a pixel in a row and column of a binarized image has a value of 1, then all pixels above, below, to the left, and to the right of that pixel are also set to 1. Other dilation kernels can be used, as long as they achieve the dilation effect. The purpose of this dilation operation is to expand the range of the motion mask, marking as many motion positions as possible and reducing errors.

[0222] The above process yields the motion mask Mask_move in the video image. This mask value is binary, meaning it's either 0 or 1. A value of 1 represents a moving area, and a value of 0 represents a relatively still area. The non-motion mask is calculated as follows: Mask_static = 1 - Mask_move.

[0223] The training process of the denoising network will be explained below in conjunction with the above method for determining the non-motion mask of the target.

[0224] In this embodiment of the disclosure, the optional method for determining the target non-moving mask includes:

[0225] Step 181: Each first video image in the N2 frames of video images is paired with the video image to be denoised and input into the optical flow network to obtain a first optical flow map representing vertical motion information and a second optical flow map representing horizontal motion information. The first video image is any video image in the N2 frames of video images other than the video image to be denoised. The resolution of the first optical flow map and the second optical flow map is X*Y.

[0226] Step 182: Based on the first optical flow map and the second optical flow map, calculate the motion mask of each frame of the first video image and the video image to be denoised, to obtain N2-1 motion masks;

[0227] Assuming the N2 video frames are F1, F2, F3, F4, and F5, with F3 being the current video image to be denoised, then sample pairs of F1 and F3 are input into the optical flow network to obtain motion mask Mask_move1, sample pairs of F2 and F3 are input into the optical flow network to obtain motion mask Mask_move2, sample pairs of F4 and F3 are input into the optical flow network to obtain motion mask Mask_move4, and sample pairs of F5 and F3 are input into the optical flow network to obtain motion mask Mask_mov5.

[0228] Step 183: Obtain the target non-motion mask based on the N2-1 motion masks.

[0229] In this embodiment of the disclosure, optionally, obtaining the target non-motion mask based on the N2-1 motion masks includes:

[0230] Based on the N2-1 moving masks, N2-1 non-moving masks are obtained; non-moving mask = 1 - moving mask;

[0231] Multiply the N2-1 non-moving target masks together to obtain the non-moving target mask.

[0232] In this embodiment of the disclosure, optionally, calculating the motion mask of each frame of the first video image and the video image to be denoised based on the first optical flow map and the second optical flow map includes:

[0233] Step 191: Subtract the last X-X1 rows of the first optical flow map from the first X-X1 rows to obtain the first difference map, and fill the last X1 rows of the first difference map with 0 to obtain the processed first difference map;

[0234] X1 is a positive integer less than X. For example, X1 can be 1, which means subtracting the previous X-1 rows from the next X-1 rows to obtain the first difference graph.

[0235] Step 192: Subtract the last Y-Y1 column from the first Y-Y1 column of the second optical flow map to obtain the second difference map, and fill the last Y1 column of the second difference map with 0 to obtain the processed second difference map;

[0236] Y1 is a positive integer less than Y. For example, Y1 can be 1, which means subtracting the previous Y-1 column from the next Y-1 column to obtain the second difference graph.

[0237] Step 193: Add the processed first difference map and the processed second difference map to obtain the third difference map;

[0238] Step 194: Assign a value of 1 to pixels whose absolute value is greater than a preset threshold in the third difference image, and assign a value of 0 to pixels whose absolute value is less than the preset threshold, to obtain a binary image;

[0239] Step 195: Obtain the motion mask based on the binary image.

[0240] In this embodiment of the present disclosure, optionally, obtaining the motion mask from the binary image includes: performing a dilation operation on the binary image to obtain the motion mask.

[0241] In this embodiment of the disclosure, optionally, obtaining the reference image based on the N2 frame video image and the target non-motion mask includes:

[0242] Multiply the N2 frames of video images by the target non-motion mask to obtain N2 products;

[0243] The reference diagram is obtained by adding the N2 products together and taking the average value.

[0244] Alternatively, the reference graph can be obtained by weighted summation of the N2 products and taking the average. The weights can be set as needed.

[0245] In this embodiment of the present disclosure, optionally, obtaining a second denoised image based on the first denoised image and the target non-motion mask includes: multiplying the first denoised image by the target non-motion mask to obtain the second denoised image.

[0246] In this embodiment of the disclosure, N2 may optionally be 5 to 9.

[0247] The training method of the above denoising network will be illustrated below using N2 as an example of 5.

[0248] Please refer to Figure 17 , Figure 17 In the diagram, F1, F2, F3, F4, and F5 are five consecutive video frames, F3 is the current video image to be denoised, DN3 is the denoised image output by the denoising network, corresponding to F3, M is the target non-motion mask, * represents multiplication of the corresponding pixel positions, and Ref3 represents the reference image, which can be considered as the ground truth corresponding to DN3.

[0249] First, according to Figure 16 The process involves inputting F1 and F3 into the optical flow network to obtain the motion mask Mask_move1, F2 and F3 to obtain the motion mask Mask_move2, F4 and F3 to obtain the motion mask Mask_move4, and F5 and F3 to obtain the motion mask Mask_mov5. From these, the non-motion masks Mask_static1, Mask_static2, Mask_static4, and Mask_static5 are calculated. The final value M is the product of the four non-motion masks, meaning that the non-motion parts of all non-motion masks are retained, while the motion parts are removed.

[0250] The method for obtaining the reference image is as follows: multiply F1, F2, F3, F4, and F5 by M respectively, then sum them and take the average. This is a temporal denoising principle, utilizing the principle that the effective information distribution is the same between consecutive frames, but noise is randomly distributed. Adding multiple frames and taking the average preserves the effective information while canceling out the influence of random noise. The purpose of calculating the non-motion mask is to ensure that the effective information of corresponding pixels in the denoised image and the reference image is the same. Without the non-motion mask step, directly adding multiple frames and taking the average preserves the effective information of non-motion positions, but motion positions will produce severe artifacts, destroying the original effective information, and thus cannot be used as a reference image for training. Using the non-motion mask, the generated reference image retains only the non-motion positions, while the denoised image also retains the corresponding pixels, forming training data pairs for training.

[0251] The method to obtain the denoised image is as follows: input F3, or F3 and its adjacent video images into the denoising network to obtain the first denoised image, and multiply the first denoised image by M to obtain the second denoised image (i.e., DN3).

[0252] In this embodiment of the disclosure, the denoising network can be any denoising network. Please refer to... Figure 18 , Figure 18 This is a schematic diagram illustrating an implementation method of a denoising network according to an embodiment of the present disclosure. The input to the denoising network is five consecutive frames of video images. The denoising network includes: multiple cascaded filters, each filter including multiple cascaded convolutional kernels. Figure 18 (The vertical bars in the middle). Figure 18 In the illustrated embodiment, each filter includes four cascaded convolutional kernels. However, in other embodiments of this disclosure, the number of convolutional kernels in a filter is not limited to four. In embodiments of this disclosure, among multiple cascaded filters, every two filters have the same resolution, and the output of each filter, except the last one, serves as the input to the next filter and the filter with the same resolution. Figure 18 In the illustrated embodiment, the denoising network includes six filters connected in series, wherein the first filter has the same resolution as the sixth filter, the second filter has the same resolution as the fifth filter, the third filter has the same resolution as the fourth filter, the output of the first filter serves as the input to the second and sixth filters (which have the same resolution as the first filter), the output of the second filter serves as the input to the third and fifth filters (which have the same resolution as the second filter), and the output of the third filter serves as the input to the fourth filter (which has the same resolution as the third filter).

[0253] In this embodiment of the disclosure, the parameters saved after the denoising network is trained can be used as the initialization parameters for the next video denoising, so a new training can be completed in about 100 frames of a new video.

[0254] Please refer to Figures 19-22 , Figure 19 The video image to be processed. Figure 20 for Figure 19 A magnified view of a portion of the video image to be processed. Figure 21 for Figure 19 The target non-motion mask M corresponding to the video image to be processed in the image. Figure 22 The image shown is the result of denoising the video image after it has been processed using the denoising network of this embodiment. The comparison results show that the denoising effect is significant.

[0255] IV. Color Correction

[0256] Color digital images captured by digital imaging devices such as digital cameras are synthesized from three channels: red (R), green (G), and blue (B). However, during the imaging process, digital imaging devices often produce images with color deviations from the original scene due to factors such as lighting and image sensor characteristics. This is known as color cast. Typically, a color-cast image is characterized by a significantly higher average pixel value in one or more of the R, G, and B channels. The color distortion caused by color cast severely affects the visual quality of the image; therefore, color cast correction in digital images is an important issue in the field of digital image processing. When processing old photographs and video materials, color cast issues are frequently addressed due to their age and preservation requirements.

[0257] To resolve color cast issues in video images, please refer to... Figure 23 The color cast correction steps in this embodiment of the disclosure include:

[0258] Step 191: Determine the target color cast values ​​for each RGB channel of the video image to be processed;

[0259] Step 192: Based on the target color cast value, perform color balance adjustment on the video image to be processed to obtain the first corrected image;

[0260] Step 193: Based on the reference image, perform color transfer on the first corrected image to obtain the second corrected image.

[0261] The reference image is an input image with virtually no color cast.

[0262] In this embodiment of the disclosure, the degree of color cast in the image is first automatically estimated, and the color balance adjustment is performed on the video image to be processed to initially correct the color cast. Then, based on the reference image, the color transfer processing is performed on the color balance adjusted image to further adjust the color cast, so that the color cast correction result is more in line with the ideal expectation.

[0263] In this embodiment of the disclosure, optionally, determining the target color cast values ​​of each RGB channel of the video image to be processed includes:

[0264] Step 201: Obtain the mean value of each RGB channel of the video image to be processed;

[0265] The calculation method for the average values ​​(avgR, avgG, avgB) of the RGB three channels is as follows: sum the gray values ​​of all R sub-pixels in the video image to be processed, and then calculate the average value to obtain avgR; sum the gray values ​​of all G sub-pixels in the video image to be processed, and then calculate the average value to obtain avgG; sum the gray values ​​of all B sub-pixels in the video image to be processed, and then calculate the average value to obtain avgB.

[0266] Step 202: Convert the mean values ​​of each RGB channel to the Lab color space to obtain the Lab color components (l, a, b) corresponding to the mean values ​​of each RGB channel.

[0267] Lab is a device-independent color system and also a color system based on physiological characteristics. This means that it uses a digital method to describe human visual perception. In the Lab color space, the L component is used to represent the brightness of a pixel, with a value range of [0,100], representing pure black to pure white; a represents the range from red to green, with a value range of [127,-128]; and b represents the range from yellow to blue, with a value range of [127,-128].

[0268] Generally speaking, for a normal, unbiased image, the mean values ​​of a and b should be close to 0. If a > 0, the image is reddish; otherwise, it is greenish. If b > 0, the image is yellowish; otherwise, it is bluish.

[0269] Step 203: Determine the degree of color cast (l, 0-a, 0-b) corresponding to the mean value of each RGB channel based on the color components (l, a, b) of the Lab space.

[0270] According to the grayscale world hypothesis, for an image without color cast, the mean color components a and b should be close to 0. Therefore, the color cast corresponding to the mean of each RGB channel is (l, 0-a, 0-b).

[0271] Step 204: Convert the color cast degree (l, 0-a, 0-b) to the RGB color space to obtain the target color cast value for each RGB channel.

[0272] The RGB color space cannot be directly converted to the Lab color space. In this embodiment, the XYZ color space is used to convert the RGB color space to the XYZ color space, and then the XYZ color space is converted to the Lab color space.

[0273] That is, converting the mean value of each RGB channel to the Lab color space includes: converting the mean value of each RGB channel to the XYZ color space to obtain the mean value of the XYZ color space; and converting the mean value of the XYZ color space to the Lab color space.

[0274] Similarly, converting the color cast degree (l, 0-a, 0-b) to the RGB color space includes: converting the color cast degree to the XYZ color space to obtain the color cast degree in the XYZ color space; and converting the color cast degree in the XYZ color space to the RGB color space.

[0275] In this embodiment of the disclosure, the mutual conversion relationship between RGB and XYZ can be as follows:

[0276]

[0277]

[0278] The conversion relationship between XYZ and Lab can be seen as follows:

[0279]

[0280]

[0281]

[0282]

[0283] Among them, X n Y n Z n The default values ​​are usually 0.95047, 1.0, and 1.08883.

[0284] The following explains how to adjust color balance.

[0285] The concept of white balance defines a region as the standard, considered white (more precisely, 18 degrees gray). The colors of other regions are then shifted based on this standard. The principle of color balance adjustment is to increase or decrease the contrasting colors to eliminate color cast in the image.

[0286] In this embodiment of the disclosure, optionally, performing color balance adjustment on the video image to be processed based on the target color cast value to obtain a first corrected image includes:

[0287] Based on the target color cast values ​​of each RGB channel, perform at least one color balance adjustment process among highlight function processing, shadow function processing, and midtone function processing on the video image to be processed;

[0288] The highlight function and shadow function are linear functions, and the midtone function is an exponential function.

[0289] In this embodiment of the disclosure, optionally, the specular function is: y = a(v)*x + b(v);

[0290] The shadow function is: y = c(v)*x + d(v);

[0291] The intermediate modulation function is: y = x f(v) ;

[0292] y is the first corrected image, x is the video image to be processed, v is determined according to the target color cast value of each RGB channel, and f(v), a(v), b(v), c(v), d(v) are functions of v.

[0293] In this embodiment of the disclosure, when adjusting the midtones, changing any one parameter will cause the current channel to change in one direction, while the other two channels will change in the opposite direction. For example, increasing the R channel parameter by 50 will increase the pixel value of the R channel, while decreasing the pixel values ​​of the G and B channels (G-50, B-50), with the two sides changing in completely opposite directions.

[0294] When adjusting highlights, for positive adjustments, such as +50 to the R channel, the algorithm only increases the value of the R channel while keeping the other two channels unchanged; for negative adjustments, such as -50 to the R channel, the algorithm keeps the R channel unchanged while increasing the values ​​of the other two channels.

[0295] When adjusting shadows, for positive adjustments, such as +50 to the R channel, the algorithm results in the R channel value remaining unchanged while the values ​​of the other two channels decrease; for negative adjustments, such as -50 to the R channel, the algorithm results in the R channel value decreasing while the values ​​of the other two channels remain unchanged.

[0296] Alternatively, f(v) = e -v .

[0297] Further optional, b(v) = 0.

[0298] Further optional,

[0299] Mixing equal amounts of RGB colors yields grays of varying brightness. If we change ΔR, ΔGd, and ΔB by the same value, theoretically, the original image will remain unchanged (adding or subtracting gray shouldn't change the color, and we need to maintain brightness, so it shouldn't change either). For example, ΔR, ΔGd, and ΔB of (+20, +35, +15) are equivalent to (+5, +20, 0) and (0, +15, -5). Therefore, to reduce the total change, we need to satisfy the condition min... d The (ΔR-d, ΔG-d, ΔB-d) values ​​when |ΔR-d|+|ΔG-d|+|ΔB-d| are used as the final target color cast value. Then, the three target color cast values ​​are combined to obtain v.

[0300] Alternatively, for the R channel, v = (ΔR - d) - (ΔG - d) - (ΔB - d);

[0301] For channel G, v = (ΔG - d) - (ΔR - d) - (ΔB - d);

[0302] For channel B, v = (ΔB-d) - (ΔR-d) - (ΔG-d);

[0303] Where ΔR, ΔG, and ΔB are the target color cast values ​​for each RGB channel, and d is the median value after sorting ΔR, ΔG, and ΔB in order of magnitude. For example, if ΔR is 10, ΔG is 15, and ΔB is 5, then d = 10.

[0304] The color migration method in the embodiments of this disclosure will be described below.

[0305] In this embodiment of the disclosure, optionally, performing color transfer on the first corrected image based on the reference image to obtain the second corrected image includes:

[0306] Step 211: Convert the reference image and the first corrected image to the Lab color space;

[0307] For conversion methods, please refer to the above-mentioned RGB to Lab conversion process.

[0308] Step 212: In the Lab color space, determine the mean and standard deviation of the reference image and the first corrected image;

[0309] Step 213: Determine the color transfer result of the k-th channel in the Lab color space based on the mean and standard deviation of the reference image and the first corrected image;

[0310] Step 214: Convert the color migration result to the RGB color space to obtain the second corrected image.

[0311] Optionally, the color migration result is calculated as follows:

[0312]

[0313] Among them, I k Let be the color transfer result of the k-th channel in the Lab color space, t be the reference image, and S be the first corrected image. This represents the mean value of the k-th channel of the first corrected image. This represents the standard deviation of the k-th channel of the first corrected image. This represents the mean value of the k-th channel of the reference image. This represents the standard deviation of the k-th channel of the reference image.

[0314] Experiments have shown that during color migration, the migration of the luminance channel causes changes in image brightness, especially in images with large areas of the same color, where changes in the luminance channel result in visual alterations. Therefore, in this embodiment, only the ab channels are migrated, meaning the k-th channel is at least one of channels a and b, thus correcting color cast while maintaining image brightness.

[0315] Please refer to Figure 24 This disclosure also provides an image processing apparatus 200, comprising:

[0316] Processing module 201, the processing module includes at least one of the following modules:

[0317] Scratch Repair Submodule 2011: Performs scratch removal processing on the video image to be processed to obtain a first image; performs difference operation on the video image to be processed and the first image to obtain a difference image; processes the difference image to obtain a scratch image that retains only the scratches; and obtains a scratch repair image based on the video image to be processed and the scratch image.

[0318] The dead pixel repair submodule 2012: acquires N1 consecutive video images, wherein each N1 video image includes a video image to be processed, at least one video image preceding the video image to be processed, and at least one video image following the video image to be processed, wherein N1 is a positive integer greater than or equal to 3; filters the video image to be processed based on the at least one video image preceding the video image to be processed and at least one video image following the video image to be processed to obtain a dead pixel repair image; and performs artifact removal processing on the dead pixel repair image based on the at least one video image preceding the video image to be processed and at least one video image following the video image to be processed to obtain an artifact repair image.

[0319] Denoising Submodule 2013: A denoising network is used to denoise the video image to be processed. The denoising network is obtained by the following training method: a target non-motion mask is obtained based on N2 consecutive video images, the N2 video images including the video image to be denoised; the denoising network to be trained is trained based on the N2 video images and the target non-motion mask to obtain the denoising network.

[0320] Color cast correction submodule 2014: Determines the target color cast values ​​for each RGB channel of the video image to be processed; performs color balance adjustment on the video image to be processed according to the target color cast values ​​to obtain a first corrected image; performs color migration on the first corrected image according to a reference image to obtain a second corrected image.

[0321] In this embodiment of the present disclosure, optionally, the scratch repair submodule performs scratch removal processing on the video image to be processed by: performing median filtering processing on the video image to be processed according to at least one of the filter type and the scratch type in the video image to be processed, to obtain a scratch-removed image.

[0322] Optionally, performing median filtering on the video image to be processed based on at least one of the filter type and the scratch type in the video image to be processed includes:

[0323] Based on the scratch type in the video image to be processed, a corresponding filter type is selected, and median filtering is performed on the video image to be processed, wherein:

[0324] When the scratches in the video image to be processed are vertical scratches, a median filter in the horizontal direction is used to perform median filtering on the video image to be processed.

[0325] When the scratches in the video image to be processed are horizontal scratches, median filtering is performed on the video image to be processed using a median filter in the vertical direction.

[0326] In this embodiment of the disclosure, optionally, performing median filtering on the video image to be processed based on at least one of the filter type and the scratch type in the video image to be processed includes:

[0327] Based on the filter type and the scratch type in the video image to be processed, different preprocessing methods are applied to the video image to be processed. Median filtering is then performed on the preprocessed video image to be processed, wherein:

[0328] When a horizontal median filter is used, and the scratches in the video image to be processed are non-vertical scratches, the video image to be processed is rotated to convert the scratches into vertical scratches.

[0329] When a vertical median filter is used, and the scratches in the video image are not horizontal, the video image to be processed is rotated so that the scratches are converted into horizontal scratches.

[0330] In this embodiment of the disclosure, optionally, scratch removal processing of the video image to be processed includes: applying a median filter of size 1×k and / or k×1 to the video image to be processed for median filtering;

[0331] The scratch repair submodule is also used to: increase the value of k of the median filter sequentially from a preset value to perform median filtering on the video image to be processed to obtain a second image; and determine the final value of k based on the filtering effect of the second image.

[0332] In this embodiment of the disclosure, optionally, performing a difference operation on the video image to be processed and the first image to obtain a difference image includes: performing a difference operation on the video image to be processed and the first image to obtain a first difference image and / or a second difference image, wherein the first difference image is obtained by subtracting the first image from the video image to be processed, and the second difference image is obtained by subtracting the video image to be processed from the first image;

[0333] Processing the difference image to obtain a scratch image that retains only the scratches includes: processing the first difference image to obtain a first scratch image that retains only the scratches; and / or, processing the second difference image to obtain a second scratch image that retains only the scratches;

[0334] Obtaining a scratch repair image based on the video image to be processed and the scratch image includes: performing calculations on the video image to be processed, the first scratch image, and / or the second scratch image to obtain the scratch repair image.

[0335] In this embodiment of the disclosure, optionally, processing the first difference image to obtain a first scratch image that retains only the scratches includes:

[0336] The first difference image is filtered by median filtering in the vertical direction and median filtering in the horizontal direction, respectively, to obtain the first vertically filtered image and the first horizontally filtered image.

[0337] If the scratch in the video image to be processed is a vertical scratch, the first scratch image is obtained by subtracting the first horizontally filtered image from the first vertically filtered image.

[0338] If the scratch in the video image to be processed is a horizontal scratch, the first scratch image is obtained by subtracting the first vertically filtered image from the first horizontally filtered image.

[0339] Processing the second difference image to obtain a second scratch image that retains only the scratches includes:

[0340] The second difference image is filtered by median filtering in the vertical direction and median filtering in the horizontal direction, respectively, to obtain the second vertically filtered image and the second horizontally filtered image.

[0341] If the scratch in the video image to be processed is a vertical scratch, the second scratch image is obtained by subtracting the second horizontally filtered image from the second vertically filtered image;

[0342] If the scratch in the video image to be processed is a horizontal scratch, the second scratch image is obtained by subtracting the second vertically filtered image from the second horizontally filtered image.

[0343] In this embodiment of the disclosure, optionally, the scratch repair image is calculated using the following formula:

[0344] I deline =IL white +Lb lack

[0345] Among them, I deline I is the scratch repair image, L is the video image to be processed. white For the first scratch image, L black This is the second scratch image.

[0346] In this embodiment of the disclosure, optionally, a bad pixel repair submodule is used to perform median filtering on the video image to be processed based on at least one video image before and at least one video image after the video image to be processed to obtain a bad pixel repair image.

[0347] In this embodiment of the disclosure, optionally, performing artifact removal processing on the bad pixel repair image based on at least one frame of video image preceding and at least one frame of video image following the video image to be processed to obtain the artifact repair image includes:

[0348] The defective pixel repair image, at least one video frame before the video image to be processed, and at least one video frame after the video image to be processed are downsampled N3-1 times to obtain N3-1 downsampled images at different resolutions. Each downsampled image at different resolutions includes N1 downsampled images corresponding to the defective pixel repair image, at least one video frame before the video image to be processed, and at least one video frame after the video image to be processed. The N3 images at different resolutions are input into a multi-scale cascaded network for artifact removal processing to obtain an artifact-repaired image. The N3 images at different resolutions include the defective pixel repair image, at least one video frame before the video image to be processed, at least one video frame after the video image to be processed, and the N3-1 downsampled images at different resolutions. The multi-scale cascaded network includes N3 cascaded sub-networks, and the images processed by the N3 cascaded sub-networks are generated based on the N3 images at different resolutions.

[0349] Where N3 is a positive integer greater than or equal to 2.

[0350] In this embodiment of the disclosure, optionally, inputting images of N3 different resolutions into a multi-scale cascaded network for artifact removal processing to obtain artifact-repaired images includes:

[0351] For the first sub-network in the N3 cascaded sub-networks: the bad pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the bad pixel repair image are downsampled by A times to obtain N1 frame first downsampled image. The N1 frame first downsampled image is then stitched together with itself to obtain a first stitched image. The first stitched image is then input into the first sub-network to obtain a first output image.

[0352] For the intermediate sub-network between the first and last sub-networks: the output image of the previous sub-network is upsampled to obtain a first upsampled image; the bad pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the video image are downsampled by a factor of B to obtain N1 frames of second downsampled images. The second downsampled image and the first upsampled image have the same scale. The two sets of images are stitched together to obtain a second stitched image. The second stitched image is input into the intermediate sub-network to obtain a second output image. Among the two sets of images, one set is the N1 frame of second downsampled images, and the other set includes: other downsampled images in the N1 frame of second downsampled images except for the downsampled image corresponding to the bad pixel repair image, and the first upsampled image.

[0353] For the last sub-network: the output image of the previous sub-network is upsampled to obtain a second upsampled image. The second upsampled image has the same scale as the video image to be processed. The two sets of images are stitched together to obtain a third stitched image. The third stitched image is input into the last sub-network to obtain the artifact restoration image. Among the two sets of images, one set is the N1 frame video image, and the other set includes: other images in the N1 frame video image except for the bad pixel restoration image, and the second upsampled image.

[0354] In this embodiment of the disclosure, optionally, the N3 cascaded sub-networks have the same structure but different parameters.

[0355] In this embodiment of the disclosure, optionally, each of the sub-networks includes a plurality of 3D convolutional layers, a plurality of 3D deconvolutional layers, and a plurality of 3D average pooling layers.

[0356] In this embodiment of the disclosure, optionally, N3 equals 3, A equals 4, and B equals 2.

[0357] In this embodiment of the disclosure, the multi-scale cascaded network is optionally obtained using the following training method:

[0358] Obtain N1 consecutive training images, wherein the N1 training images include the training image to be processed, at least one training image before the training image to be processed, and at least one training image after the training image.

[0359] The first training image is obtained by filtering the training image to be processed based on the N1 frames of training images;

[0360] The multi-scale cascaded network to be trained is trained based on the first training image, at least one training image before the training image to be processed, and at least one training image after the training image to be processed, to obtain the trained multi-scale cascaded network.

[0361] In this embodiment of the disclosure, optionally, when training the multi-scale cascaded network to be trained, the total loss used includes at least one of the following: image content loss, color loss, edge loss, and perceptual loss.

[0362] In this embodiment of the disclosure, optionally, the total loss is equal to the weighted sum of image content loss, color loss, edge loss, and perceptual loss. The image content L1 loss is calculated using the following formula:

[0363]

[0364] Among them, l content For L1 loss, For the artifact removal training image, y iLet n be the first training image, and n be the number of images in a batch.

[0365] In this embodiment of the disclosure, optionally, the color loss is calculated using the following formula:

[0366]

[0367] Among them, l color For the color loss, For the artifact removal training image, y i Let be the first training image, n be the number of images in a batch, and Blur(x) be the Gaussian blur function.

[0368] In this embodiment of the disclosure, optionally, the edge loss is calculated using the following formula:

[0369]

[0370] Among them, l edge For the edge loss, For the artifact removal training image, y i The first training image is n, where n is the number of images in a batch, and H is the number of images in a batch. j (x) represents the image edge map extracted by the j-th layer of the HED network.

[0371] In this embodiment of the disclosure, optionally, the perception loss is calculated using the following formula:

[0372]

[0373] Among them, l feature For the perceived loss, For the artifact removal training image, y i The first training image is denoted as n, where n is the number of images in a batch. This represents the image feature map extracted from the j-th layer of the VGG network.

[0374] In this embodiment of the disclosure, optionally, training the multi-scale cascaded network to be trained based on the first training image, at least one training image preceding the training image to be processed, and at least one training image following the training image to be processed includes:

[0375] A random image block is cropped from the first training image. An image block is cropped from the same position in at least one training image before the training image to be processed and at least one training image after the training image to be processed, respectively, to obtain N1 frame image blocks.

[0376] The N1 frame image blocks are input into the multi-scale cascaded network to be trained for training.

[0377] In this embodiment of the disclosure, optionally, the bad pixel repair submodule is used to: filter the artifact repair image based on the video image to be processed and the bad pixel repair image to obtain an output image.

[0378] In this embodiment of the disclosure, optionally, the bad pixel repair submodule is used to: perform median filtering on the artifact repair image based on the video image to be processed and the bad pixel repair image to obtain an output image.

[0379] In this embodiment of the disclosure, N1 may optionally be equal to 3.

[0380] In this embodiment of the disclosure, optionally, the denoising submodule, used for obtaining the target non-motion mask based on N2 consecutive video images, includes:

[0381] A reference image is obtained based on the N2 frames of video images and the target non-motion mask;

[0382] The video image to be denoised is input into the denoising network to be trained to obtain the first denoised image;

[0383] Based on the first denoised image and the target non-motion mask, a second denoised image is obtained;

[0384] The loss function of the denoising network to be trained is determined based on the reference image and the second denoised image, and the parameters of the denoising network to be trained are adjusted according to the loss function to obtain the denoising network.

[0385] In this embodiment of the disclosure, optionally, obtaining the reference image based on the N2 frame video image and the target non-motion mask includes:

[0386] Each of the first video images in the N2 frames of video images is paired with the video image to be denoised and input into the optical flow network to obtain a first optical flow map representing vertical motion information and a second optical flow map representing horizontal motion information. The first video image is any video image in the N2 frames of video images other than the video image to be denoised. The resolution of the first optical flow map and the second optical flow map is X*Y.

[0387] Based on the first optical flow map and the second optical flow map, calculate the motion mask of each frame of the first video image and the video image to be denoised, and obtain N2-1 target motion masks;

[0388] The target non-motion mask is obtained based on the N2-1 motion masks.

[0389] In this embodiment of the disclosure, optionally, calculating the motion mask of each frame of the first video image and the video image to be denoised based on the first optical flow map and the second optical flow map includes:

[0390] Subtract the last X-X1 rows of the first optical flow map from the first X-X1 rows to obtain the first difference map, and fill the last X1 rows of the first difference map with 0 to obtain the processed first difference map.

[0391] Subtract the last Y-Y1 column from the first Y-Y1 column of the second optical flow map to obtain the second difference map, and fill the last Y1 column of the second difference map with 0 to obtain the processed second difference map;

[0392] The processed first difference map and the processed second difference map are added together to obtain the third difference map;

[0393] In the third difference image, pixels with absolute values ​​greater than a preset threshold are assigned a value of 1, and pixels with absolute values ​​less than the preset threshold are assigned a value of 0, thus obtaining a binary image.

[0394] The motion mask is obtained from the binary image.

[0395] In this embodiment of the disclosure, optionally, obtaining the motion mask based on the binary image includes:

[0396] The motion mask is obtained by performing a dilation operation on the binary image.

[0397] In this embodiment of the disclosure, optionally, obtaining the target non-motion mask based on the N2-1 motion masks includes:

[0398] N2-1 non-motion masks are obtained from the N2-1 motion masks; where, non-motion mask = 1 - motion mask;

[0399] Multiply the N2-1 non-motion masks together to obtain the target non-motion mask.

[0400] In this embodiment of the disclosure, optionally, obtaining the reference image based on the N2 frame video image and the target non-motion mask M includes:

[0401] Multiply the N2 video frames by the target non-motion mask M to obtain N2 products;

[0402] The reference diagram is obtained by adding the N2 products together and taking the average value.

[0403] In this embodiment of the disclosure, optionally, obtaining the second denoised image based on the first denoised image and the target non-motion mask M includes:

[0404] The first denoised image is multiplied by the target non-motion mask M to obtain the second denoised image.

[0405] In this embodiment of the disclosure, N2 may optionally be 5 to 9.

[0406] In this embodiment of the disclosure, optionally, the color cast correction submodule is used to determine the target color cast values ​​of each RGB channel of the video image to be processed, including:

[0407] Obtain the average value of each RGB channel of the video image to be processed;

[0408] The mean values ​​of each RGB channel are converted to the Lab color space to obtain the Lab color components (l, a, b) corresponding to the mean values ​​of each RGB channel.

[0409] Based on the color components (l, a, b) of the Lab space, determine the degree of color cast (l, 0-a, 0-b) corresponding to the mean value of each RGB channel;

[0410] The color cast degree (l, 0-a, 0-b) is converted to the RGB color space to obtain the target color cast value for each RGB channel.

[0411] In this embodiment of the disclosure, optionally, converting the mean value of each RGB channel to the Lab color space includes: converting the mean value of each RGB channel to the XYZ color space to obtain the mean value of the XYZ color space; and converting the mean value of the XYZ color space to the Lab color space.

[0412] Converting the color cast degree (l, 0-a, 0-b) to the RGB color space includes: converting the color cast degree to the XYZ color space to obtain the color cast degree in the XYZ color space; and converting the color cast degree in the XYZ color space to the RGB color space.

[0413] In this embodiment of the disclosure, optionally, performing color balance adjustment on the video image to be processed based on the target color cast value to obtain a first corrected image includes:

[0414] Based on the target color cast values ​​of each RGB channel, the video image to be processed is subjected to at least one of the following color balance adjustment processes: highlight function processing, shadow function processing, and midtone function processing. The highlight function and shadow function are linear functions, and the midtone function is an exponential function.

[0415] In this embodiment of the disclosure, optionally, the specular function is: y = a(v)*x + b(v);

[0416] The shadow function is: y = c(v)*x + d(v);

[0417] The intermediate modulation function is: y = x f(v) ;

[0418] y is the first corrected image, x is the video image to be processed, v is determined according to the target color cast value of each RGB channel, and f(v), a(v), b(v), c(v), d(v) are functions of v.

[0419] In this embodiment of the disclosure, f(v) = e -v .

[0420] In this embodiment of the disclosure, optionally, b(v) = 0.

[0421] In this embodiment of the disclosure, optionally,

[0422] In this embodiment of the disclosure, optionally, for channel R, v = (ΔR-d) - (ΔG-d) - (ΔB-d);

[0423] For channel G, v = (ΔG - d) - (ΔR - d) - (ΔB - d);

[0424] For channel B, v = (ΔB-d) - (ΔR-d) - (ΔG-d);

[0425] Where ΔR, ΔG, and ΔB are the target color cast values ​​for each RGB channel, and d is the median value after sorting ΔR, ΔG, and ΔB in order of magnitude.

[0426] In this embodiment of the disclosure, optionally, performing color transfer on the first corrected image based on the reference image to obtain the second corrected image includes:

[0427] Convert the reference image and the first corrected image to the Lab color space;

[0428] In the Lab color space, determine the mean and standard deviation of the reference image and the first corrected image;

[0429] Based on the mean and standard deviation of the reference image and the first corrected image, the color transfer result of the k-th channel in the Lab color space is determined;

[0430] The color migration result is converted to the RGB color space to obtain the second corrected image.

[0431] In this embodiment of the disclosure, the calculation method for the color migration result is optionally as follows:

[0432]

[0433] Among them, Ik Let be the color transfer result of the k-th channel in the Lab color space, t be the reference image, and S be the first corrected image. This represents the mean value of the k-th channel of the first corrected image. This represents the standard deviation of the k-th channel of the first corrected image. This represents the mean value of the k-th channel of the reference image. This represents the standard deviation of the k-th channel of the reference image.

[0434] In this embodiment of the disclosure, optionally, the k-th channel is at least one of channels a and b.

[0435] This application also provides an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects.

[0436] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0437] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0438] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0439] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0440] The embodiments of this disclosure have been described above with reference to the accompanying drawings. However, this disclosure is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this disclosure without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this disclosure.

Claims

1. An image processing method, characterized in that, include: Dead pixel repair steps: Obtain continuous Frame video image, the The video frame includes a video image to be processed, at least one video frame preceding the video image to be processed, and at least one video frame following the video image to be processed. The value is a positive integer greater than or equal to 3; the video image to be processed is filtered based on at least one video image before and at least one video image after the video image to be processed to obtain a bad pixel repair image; the bad pixel repair image is subjected to artifact removal processing based on at least one video image before and at least one video image after the video image to be processed to obtain an artifact repair image. The process of performing artifact removal processing on the bad pixel restoration image based on at least one frame of video image preceding and at least one frame of video image following the video image to be processed to obtain the artifact restoration image includes: The bad pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the video image to be processed are respectively processed. -1 downsampling, resulting in -1 downsampled images at different resolutions, wherein each downsampled image at different resolutions includes at least one video image preceding and at least one video image following the bad pixel repair image and the video image to be processed. Frame downsampled image; Images of various resolutions are input into a multi-scale cascaded network for artifact removal processing to obtain artifact-repaired images. The image of a certain resolution includes: the bad pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the video image to be processed, and the... -1 resolution downsampled images, the multi-scale cascaded network includes A cascaded subnetwork, the The images processed by the cascaded sub-networks are respectively based on the Image generation at various resolutions; in, It is a positive integer greater than or equal to 2.

2. The method according to claim 1, characterized in that, The process of filtering the video image to be processed to obtain a defect-repaired image, based on at least one frame of video image preceding and at least one frame of video image following the video image to be processed, includes: Based on at least one frame of video image preceding and at least one frame of video image following the video image to be processed, a median filter is applied to the video image to be processed to obtain a bad pixel repair image.

3. The method according to claim 1, characterized in that, Will Images of various resolutions are input into a multi-scale cascaded network for artifact removal processing to obtain artifact-restored images, including: Regarding the above The first subnetwork in a cascaded subnetwork: The bad pixel repair image, at least one frame of video image preceding the video image to be processed, and at least one frame of video image following the video image are all downsampled by a factor of A to obtain... The first downsampled image of the frame, the The first downsampled image of each frame is stitched together with itself to obtain the first stitched image. The first stitched image is then input into the first sub-network to obtain the first output image. For the intermediate subnetwork between the first and last subnetworks: the output image of the previous subnetwork is upsampled to obtain a first upsampled image; the bad pixel repair image, at least one frame of video image before the video image to be processed, and at least one frame of video image after the image are downsampled by a factor of B to obtain... The second downsampled image is taken as a frame, and the scale of the second downsampled image is the same as that of the first upsampled image. The two sets of images are stitched together to obtain a second stitched image. The second stitched image is then input into the intermediate sub-network to obtain a second output image. One of the two sets of images is... The second downsampled image of the frame, another group includes: the The other downsampled images in the second downsampled image frame, excluding the downsampled image corresponding to the bad pixel repair image, and the first upsampled image; For the last sub-network: the output image of the previous sub-network is upsampled to obtain a second upsampled image. The second upsampled image has the same scale as the video image to be processed. The two sets of images are stitched together to obtain a third stitched image. The third stitched image is input into the last sub-network to obtain the artifact-repaired image. One of the two sets of images is the... Another set of frame video images includes: The other images in the frame video image besides the bad pixel repair image, and the second upsampled image.

4. The method according to claim 1, characterized in that, The Each cascaded subnetwork has the same structure but different parameters.

5. The method according to claim 1, characterized in that, Each of the sub-networks includes multiple 3D convolutional layers, multiple 3D deconvolutional layers, and multiple 3D average pooling layers.

6. The method according to claim 3, characterized in that, A equals 3, B equals 4, and C equals 2.

7. The method according to any one of claims 1-6, characterized in that, The multi-scale cascaded network was obtained using the following training method: Get continuous The training image of the frame, The training image frame includes the training image to be processed, at least one training image preceding the training image to be processed, and at least one training image following the training image; According to the above The first training image is obtained by filtering the training image to be processed using the training image frame; The multi-scale cascaded network to be trained is trained based on the first training image, at least one training image before the training image to be processed, and at least one training image after the training image to be processed, to obtain the trained multi-scale cascaded network.

8. The method according to claim 7, characterized in that, When training a multi-scale cascaded network to be trained, the total loss used includes at least one of the following: image content loss, color loss, edge loss, and perceptual loss.

9. The method according to claim 8, characterized in that, The total loss is equal to the weighted sum of image content loss, color loss, edge loss, and perceptual loss.

10. The method according to claim 9, characterized in that, The image content loss is calculated using the following formula: ; in, For L1 loss, For training images to remove artifacts, Let n be the first training image, and n be the number of images in a batch.

11. The method according to claim 8, characterized in that, The color loss is calculated using the following formula: ; in, For the color loss, For training images to remove artifacts, Let be the first training image, n be the number of images in a batch, and Blur(x) be the Gaussian blur function.

12. The method according to claim 8, characterized in that, The edge loss is calculated using the following formula: ; in, For the edge loss, For training images to remove artifacts, The first training image is denoted as n, where n is the number of images in a batch. This represents the image edge map extracted from the j-th layer of the HED network.

13. The method according to claim 8, characterized in that, The perceived loss is calculated using the following formula: ; in, For the perceived loss, For training images to remove artifacts, The first training image is denoted as n, where n is the number of images in a batch. This represents the image feature map extracted from the j-th layer of the VGG network.

14. The method according to claim 7, characterized in that, Training the multi-scale cascaded network to be trained based on the first training image, at least one training image preceding the training image to be processed, and at least one training image following the training image to be processed includes: A random image patch is cropped from the first training image. Image patches are also cropped from the same positions in at least one training image preceding and at least one training image following the training image to be processed, resulting in... Frame image blocks; The Frame image blocks are input into the multi-scale cascaded network to be trained.

15. The method according to claim 1, characterized in that, After obtaining the artifact-repaired image, the following is also included: Based on the video image to be processed and the image with damaged pixels, the image with damaged pixels is filtered to obtain the output image.

16. The method according to claim 15, characterized in that, Based on the video image to be processed and the image with damaged pixels, the image with damaged pixels is subjected to median filtering to obtain the output image.

17. The method according to claim 1, characterized in that, It equals 3.

18. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the image processing method as described in any one of claims 1 to 17.

19. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image processing method as described in any one of claims 1 to 17.