Video processing method and device

The video processing is decoupled into three branches of denoising, fusion and demosaic through the residual branch model, which solves the poor universality problem caused by the independent settings of different tasks in the prior art video processing model, and realizes the comprehensive processing of HDR and NonHDR videos.

CN120374473APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410674541.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, video processing models cannot simultaneously handle the three tasks of noise reduction, demosaic and high dynamic fusion, especially lacking universality for processing HDR and NonHDR type videos.

Method used

The residual branch model is adopted, and the video processing is decoupled into three independent processing branches, denoising, fusion and de-mosaic, and connecting these branches using the residual structure to establish their logical relationships, forming a model that can handle these three tasks simultaneously.

Benefits of technology

The comprehensive processing of HDR and NonHDR videos is achieved, improving the versatility of the model, allowing it to handle noise reduction, demosaic and high dynamic fusion tasks simultaneously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374473A_ABST
    Figure CN120374473A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method and device, and relates to the technical field of video processing. The method comprises the steps of obtaining a to-be-processed video; dividing an image in the video to be processed according to a preset moment node to obtain a frame image at each moment; and inputting the frame image at each moment into the trained residual branch model for image processing to obtain a video processing result. Compared with the prior art, the method has the advantages that the corresponding processing branches are set for decoupling of video processing, noise reduction, fusion and demosaicing, and connection is performed through the residual structure, so that the processing logic relationship among the processing branches is defined, the functions of the trained residual branch model are more comprehensive, and the robustness of the residual branch model is improved. According to the method, noise reduction, demosaicing and fusion tasks in the video can be processed, processing of HDR type and NonHDR type videos is considered at the same time, and the problem that in a related method, processing models are independently set for different tasks, and universality is poor is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and in particular, to a video processing method and apparatus. Background Art

[0002] Video processing is an important function of ISP (Image Singal Process), mainly including three tasks: noise reduction, demosaicing, and high dynamic range fusion. Among them, the video can include HDR (High-Dynamic Range) video and non-HDR video (also known as NonHDR or SDR). For NonHDR type videos, it includes noise reduction and demosaicing. HDR type videos also include high dynamic range fusion on the basis of NonHDR. The existing methods usually decouple them and independently design processing modules for different processing tasks or video types. There is no processing model that can handle the above three tasks simultaneously, so the generality is poor. Summary of the Invention

[0003] In view of this, this application provides a video processing method and apparatus, and the main purpose is to propose a new processing model that can simultaneously handle three tasks: noise reduction, demosaicing, and high dynamic range fusion, taking into account the processing of HDR type and NonHDR type videos, and solving the problem of poor generality in related methods.

[0004] In a first aspect, this application provides a video processing method, including:

[0005] Obtain a video to be processed;

[0006] Divide the images in the video to be processed according to preset time nodes to obtain frame images at each moment;

[0007] Input the frame images at each moment into a trained residual branch model for image processing to obtain a video processing result;

[0008] Among them, each branch in the residual branch model adopts a residual structure, and at least includes a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame images according to the processing result of the second processing branch, and the fourth processing branch is used for demosaicing processing in the denoising processing and integration processing.

[0009] Optionally, when the residual branch model includes four processing branches, the step of obtaining the trained residual branch model includes: obtaining a training video; the training video includes training frame images divided into each moment; setting four processing branches for the initial model and allocating computing power resources for each processing branch; training the initial model based on the training frame images to obtain the trained residual branch model.

[0010] Optionally, the training frame images include short frame images and long frame images; the step of training the initial model based on the training frame images to obtain the trained residual branch model includes: using the short frame images to train the first processing branch and the fourth processing branch to obtain a first processing result output by the first processing branch; the first processing result includes a first image and a first branch weight; fixing the first processing branch and the fourth processing branch to the first weight, and using the first image and the long frame images to train the second processing branch and the fourth processing branch to obtain a second processing result output by the second processing branch; the second processing result includes a second image and a second branch weight; fixing the second processing branch and the fourth processing branch to the second weight, and using the second image and historical frame images to train the third processing branch and the fourth processing branch to obtain a third processing result output by the third processing branch; the third processing result includes a third image and a third branch weight; respectively allocating the first branch weight, the second branch weight, and the third branch weight to the corresponding processing branches to obtain the trained residual branch model; wherein, the fourth processing branch is used for demosaicing during the training process; the historical frame images are the third images of the frame images at the previous moment with the current moment as the reference; the third image is a processing result obtained by fusing the first image and the second image with the short frame image as the reference frame.

[0011] Optionally, allocating computing power resources for each processing branch includes: setting the computing power resource allocation ratios of the first processing branch, the second processing branch, and the third processing branch to 1:1:1.

[0012] Optionally, after obtaining the video to be processed, the method further includes: judging the type of the video to be processed; when the type of the video to be processed is NonHDR type, masking the second processing branch.

[0013] Optionally, when the video type to be processed is the HDR type, the frame image at each moment further includes a short-frame image and a long-frame image; before inputting the frame image at each moment into the trained residual branch model for image processing, the method further includes: performing brightness alignment processing on the short-frame image by using the exposure ratio of the short-frame image and the long-frame image; and / or, performing denoising on the short-frame image by a preset non-blind denoising method.

[0014] In a second aspect, the present application provides a video processing device, including:

[0015] An acquisition unit, configured to acquire a video to be processed;

[0016] A partitioning unit, configured to partition the images in the video to be processed according to preset time nodes to obtain frame images at each moment;

[0017] A processing unit, configured to input the frame images at each moment into a trained residual branch model for image processing to obtain a video processing result;

[0018] Wherein, each branch in the residual branch model adopts a residual structure and at least includes a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame image according to the processing result of the second processing branch, and the fourth processing branch is used for demosaicing processing in the denoising processing and integration processing.

[0019] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the video processing method described in the first aspect is implemented.

[0020] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the video processing method described in the first aspect is implemented.

[0021] In a fifth aspect, the present application provides a chip, including one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from the memory of the electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is enabled to execute the video processing method described in the first aspect.

[0022] With the above technical solutions, a video processing method and apparatus provided by the present application first obtain a video to be processed; divide the images in the video to be processed at preset time nodes to obtain frame images at each moment; input the frame images at each moment into a trained residual branch model for image processing to obtain a video processing result. Among them, each branch in the residual branch model adopts a residual structure and at least includes a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame images according to the processing result of the second processing branch, and the fourth processing branch is used for demosaicing processing in the denoising processing and the integration processing. Compared with the related technologies, the present application decouples the video processing, sets corresponding processing branches for noise reduction, fusion, and demosaicing respectively and connects them through a residual structure, clarifying the processing logic relationship between the processing branches, so that the trained residual branch model has a more comprehensive function, can handle the noise reduction, demosaicing, and fusion tasks in the video, and at the same time takes into account the processing of HDR type and NonHDR type videos, solving the problem of poor generality in the related methods where processing models are independently set for different tasks.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0025] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0026] Figure 1 It shows a schematic flowchart of a video processing method provided by an embodiment of the present application;

[0027] Figure 2 It shows a schematic flowchart of the training steps of a residual branch model provided by an embodiment of the present application;

[0028] Figure 3 It shows a schematic diagram of the processing logic of a residual branch model provided by an embodiment of the present application;

[0029] Figure 4The figure shows a schematic structural diagram of a video processing device provided by an embodiment of the present application. Detailed implementation manners

[0030] Here, some embodiments of the present disclosure will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, deformations, and equivalents of the methods, apparatuses, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those set forth herein. Instead, it can be changed as will be apparent after understanding the present disclosure, except for operations that must be performed in a specific order. Additionally, descriptions of features known in the art may be omitted for the sake of clarity and conciseness.

[0031] The implementation manners described in some embodiments of the present disclosure below do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0032] In order to be able to simultaneously process three tasks of noise reduction, demosaicing, and high dynamic range fusion, and take into account the processing of HDR type and NonHDR type videos, and solve the problem of poor generality in related methods. This embodiment provides a video processing method, as Figure 1 shown, the method includes:

[0033] S101. Obtain a video to be processed.

[0034] There are two types of videos to be processed, namely HDR videos and NonHDR videos. Among them, an HDR video refers to a high dynamic range video. An HDR video uses a wider color range, making the colors in the picture more vivid and saturated, bringing a more realistic visual experience to the user. It can not only show deeper blacks, making the black parts in the picture more profound and layered, but also has a higher dynamic range, being able to capture more detailed information, making the details of the items in the picture clearer and more distinct. An HDR video has a wider brightness range, so there will be no overexposure phenomenon in bright scenes, and the user can see every detail in the picture. A NonHDR video refers to a non-high dynamic range video. Compared with HDR videos, the color performance of NonHDR videos is relatively dull. Due to the limited dynamic range, NonHDR videos may not be able to accurately restore the color details in high-brightness and low-brightness areas, resulting in a relatively weak overall color performance. In addition, it will also cause overexposure in high-light areas or loss of details in shadow areas. However, the NonHDR video format is more popular in the market and has better compatibility.

[0035] In this embodiment, since the trained residual branch model is provided with processing branches with different functions and adopts a residual structure, it can take into account the processing of both HDR and NonHDR types of videos. Therefore, regardless of the type of input video, it can be processed without inputting different models for different video types or tasks.

[0036] S102. Divide the images in the video to be processed at preset time nodes to obtain frame images at each moment;

[0037] Dividing at preset time nodes means that when processing a video, according to a certain specific preset time point or time period, the images in the video are extracted or processed. For example, the video is divided into images at multiple moments in seconds, and then further refined into frame images. A frame image refers to that in video processing, a video is usually decomposed into a series of frame images. Each frame is a static picture in the video, similar to a photo. After division, the images corresponding to each preset time point are extracted to form "frame images". These frame images represent the static state of the video at a specific time point.

[0038] For HDR type videos, each frame of video image in the video to be processed includes long frame images and short frame images with two different exposure parameters. The main differences between short frame images and long frame images are in the exposure time and exposure gain. The main differences between long and short frames are in the exposure time and exposure gain. Long frames include more dark part information, while short frames include more bright part details. For NonHDR type videos, long frame images are not included.

[0039] S103. Input the frame images at each moment into the trained residual branch model for image processing to obtain a video processing result;

[0040] Among them, each branch in the residual branch model adopts a residual structure and includes at least a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame images according to the processing result of the second processing branch, and the fourth processing branch is used for demosaicing processing in the denoising processing and integration processing.

[0041] Residual Network (ResNet) is a convolutional neural network (CNN) structure in deep learning. It solves the problems of gradient vanishing and representation bottleneck in deep networks by introducing "residual blocks". Residual blocks allow the network to learn the residuals between the input and output, which helps the network better capture details and features in images. The residual branch model refers to one or more residual blocks in the ResNet. In the ResNet, data is usually passed and processed between multiple residual blocks. Each residual block receives the input data and outputs the processed data. This data may be directly passed to the next residual block, or combined with the input data in some form (such as addition), and then passed to the next residual block. Each frame image at each moment is input into the trained residual branch model, and the model will perform a series of processing and analysis on these images. After being processed by the residual branch model, each frame image will obtain a corresponding result.

[0042] Among them, the first processing branch is mainly used for denoising processing. More specifically, it performs denoising processing on short-frame images. Denoising is a basic task in image processing, aiming to remove unnecessary noise from images to improve image quality. The second processing branch further improves denoising based on the first processing branch. Specifically, the second processing branch performs denoising processing on long-frame images in the HDR video type based on the output result of the first processing branch; if it is of the NonHDR type, the second processing branch can be blocked, and the first processing result directly passes through the second processing branch without any processing and is input into the third processing branch. The third processing branch integrates the frame images according to the denoising processing result of the second processing branch. The fourth processing branch is used to perform demosaicing processing on the image while the first three processing branches are processing. Demosaicing means removing or modifying the mosaic in an image or video, which refers to restoring the occluded or modified content through some technical means.

[0043] In this embodiment, first, a video to be processed is obtained; at a preset time node, the images in the video to be processed are divided to obtain frame images at each moment; the frame images at each moment are input into a trained residual branch model for image processing to obtain a video processing result. Among them, each branch in the residual branch model adopts a residual structure and at least includes a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame images according to the processing result of the second processing branch, and the fourth processing branch is used for demosaicing processing in the denoising processing and the integration processing. Compared with the related technology, in this embodiment, by decoupling the video processing, corresponding processing branches are set for noise reduction, fusion, and demosaicing respectively and connected through a residual structure, clarifying the processing logic relationship between the processing branches, so that the trained residual branch model has a more comprehensive function, can handle the noise reduction, demosaicing, and fusion tasks in the video, and at the same time takes into account the processing of HDR type and NonHDR type videos, solving the problem of poor generality in the related methods where processing models are independently set for different tasks.

[0044] Optionally, in the case where the residual branch model includes four processing branches, the steps of obtaining the trained residual branch model include: obtaining a training video; the training video includes training frame images that have been divided into each moment; setting four processing branches for the initial model and allocating computing power resources for each processing branch; training the initial model based on the training frame images to obtain the trained residual branch model.

[0045] In this embodiment, obtaining the training video is the basis of the training process, and video data for training needs to be prepared. Specifically, for example, the training video is of the HDR (High Dynamic Range) type, which means that the video contains a wide range of brightness, with rich details from dark to bright. The video is composed of consecutive frame images. During the training process, the video will be decomposed into these individual frame images to facilitate the model's learning and processing. Before starting the training, an initial residual branch model needs to be set up. This model has four processing branches respectively applied to denoise short frame images, denoise long frame images, integrate images according to time series, and perform demosaicing processing. Each branch is responsible for performing different image processing tasks. Computing power resources refer to the hardware resources of a computer used to execute computing tasks, such as CPUs, GPUs, etc. In deep learning, the training of a model requires a large amount of computing power support. Therefore, these resources need to be reasonably allocated to different processing branches to ensure that they can execute tasks in parallel and efficiently. During the training process, the model will continuously try to predict the output of the training frame images and adjust the parameters according to the error between the prediction result and the true result. After a certain number of iterative trainings, the model will converge to a relatively stable state, and at this time, the training can be stopped to obtain a trained residual branch model.

[0046] Optionally, the training frame images include short frame images and long frame images; based on the training frame images, training the initial model to obtain a trained residual branch model includes: using the short frame images to train the first processing branch and the fourth processing branch to obtain a first processing result output by the first processing branch; the first processing result includes a first image and a first branch weight; fixing the first processing branch and the fourth processing branch as the first weight, and using the first image and the long frame images to train the second processing branch and the fourth processing branch to obtain a second processing result output by the second processing branch; the second processing result includes a second image and a second branch weight; fixing the second processing branch and the fourth processing branch as the second weight, and using the second image and the historical frame images to train the third processing branch and the fourth processing branch to obtain a third processing result output by the third processing branch; the third processing result includes a third image and a third branch weight; respectively allocating the first branch weight, the second branch weight, and the third branch weight to the corresponding processing branches to obtain a trained residual branch model; wherein, the fourth processing branch is used to perform demosaicing processing during the training process; the historical frame image is the third image of the frame image at the previous moment based on the current moment; the third image is the processing result obtained by fusing the first image and the second image with the short frame image as the reference frame.

[0047] In this embodiment, the training frame images include short-frame images and long-frame images. The main differences between short-frame images and long-frame images lie in the exposure time and exposure gain. The main differences between long and short frames are in the exposure time (the exposure time of the long frame is t_l, and the exposure time of the short frame is t_s) and the exposure gain (the gain of the long frame is gain_l, and the gain of the short frame is gain_s). The long frame includes more dark part information, while the short frame includes more bright part details.

[0048] First stage: Use the short-frame images to train the first processing branch and the fourth processing branch. The process of training the first processing branch and the fourth processing branch is also to process the short-frame images through the first processing branch and the fourth processing branch. The purpose of training the first processing branch is to enable the first processing branch to learn to denoise the long-frame images. The fourth processing branch is used to perform demosaicing while the first processing branch is performing noise reduction. The reason for training the fourth processing branch is that the fourth processing branch is needed to perform demosaicing on the images during the training process. After training, the first processing branch will output the first processing result, including the first image (specifically, it can include the denoised image and corresponding image features such as feature map1) and the first branch weight (the parameters learned by this branch during the training process).

[0049] Second stage: After fixing the parameters of the first processing branch and the fourth processing branch to the weights learned in the first stage, use the first image and the long-frame images to train the second processing branch and the fourth processing branch. The purpose of training the fourth processing branch and the role of the fourth processing branch are the same as those in the first stage. The purpose of training the second processing branch is to enable the second processing branch to learn to denoise the long-frame images. After training, the second processing branch will output the second processing result, including the second image (specifically, it can include the denoised image and corresponding image features such as feature map2) and the second branch weight.

[0050] Third stage: After fixing the parameters of the second processing branch and the fourth processing branch to the weights learned in the second stage, use the second image and the historical frame images to train the third processing branch and the fourth processing branch. The historical frame image is the third image of the frame image at the previous moment based on the current moment (i.e., the output result of the third processing branch at the previous moment). The main role of the third processing branch is to integrate the images according to the temporal features. Therefore, it is necessary to determine the current processing result in combination with the output result of the previous moment. After training, the third processing branch will output the third processing result, including the third image and the third branch weight.

[0051] The third image is the final result obtained by using the short-frame image as the reference frame and fusing the first image (the denoising result of the short-frame image) and the second image (the denoising result of the long-frame image) with the original short-frame image. Of course, it also includes the demosaicing result of the fourth processing branch. This kind of fusion processing includes frame-by-frame fusion in time or image fusion in space to improve the quality or continuity of the image.

[0052] Finally, the first-branch weight, second-branch weight, and third-branch weight learned in the first stage, second stage, and third stage are respectively assigned to the corresponding processing branches, so as to obtain a complete trained residual branch model. For the fourth processing branch, there is no fixed weight, and usually, it inherits the weight of the previous stage (for example, when training the third branch, it inherits the weight of the second processing branch), so as to train the subsequent branch based on the weight of the previous stage.

[0053] Optionally, allocate computing power resources for each processing branch, including: setting the computing power resource allocation ratios of the first processing branch, the second processing branch, and the third processing branch to 1:1:1.

[0054] In this embodiment, when allocating the computing power resources of each processing branch in the residual branch model, the computing power resource allocation ratios of the first processing branch, the second processing branch, and the third processing branch are set to 1:1:1, which means that these three branches will obtain equal computing resources during training. Such an allocation method can ensure that each branch has sufficient ability to learn and optimize its specific image processing tasks.

[0055] Optionally, after obtaining the video to be processed, the method further includes: judging the type of the video to be processed; in the case where the type of the video to be processed is NonHDR type, shielding the second processing branch.

[0056] In this embodiment, before video processing, the type of the video will be judged first. If the video to be processed is judged to be of NonHDR type, then the second processing branch will be shielded. This is because the second processing branch mainly denoises long-frame images during training. Shielding this branch can avoid unnecessary waste of computing resources and may improve the processing efficiency and quality of NonHDR videos. By adjusting the processing flow according to the video type or characteristics, different video types may require different processing strategies or algorithms to achieve the best processing effect.

[0057] Optionally, when the video type to be processed is of the HDR type, the frame image at each moment further includes a short-frame image and a long-frame image; before inputting the frame image at each moment into the trained residual branch model for image processing, the method further includes: performing brightness alignment processing on the short-frame image by using the exposure ratio between the short-frame image and the long-frame image; and / or denoising the short-frame image by a preset non-blind denoising method.

[0058] In this embodiment, before video processing, the type of the video is first determined. If the video to be processed is determined to be of the HDR type, before inputting the frame image at each moment into the trained residual branch model for image processing, the short-frame is aligned in exposure before input, that is, the brightness is matched with the long-frame, and non-blind denoising is adopted in data processing. The exposure ratio is a parameter describing the brightness difference between images under different exposure times. The long-frame image has higher brightness, while the short-frame image has lower brightness. Due to the brightness difference between the long-frame and short-frame images, directly using them for training the model may lead to a decline in model performance because the model needs to process this unnecessary brightness change. Therefore, performing brightness alignment processing is to ensure that the brightness of the short-frame image is similar to or matches that of the long-frame image, ensuring the consistency of image brightness under different exposure conditions, so that the model can focus more on learning other features of the image rather than the brightness difference.

[0059] Non-blind denoising can also be called post-correction denoising, which is a common step in image processing for reducing or eliminating noise in an image. Non-blind denoising means that the denoising process does not depend on the specific type or characteristics of the noise, but is based on some preset algorithms or methods for processing. In this embodiment, instead of using the standard non-blind denoising conversion (where the standard non-blind denoising refers to converting the noise distribution into a Gaussian distribution with a variance of 1), modified non-blind denoising is used. The modified non-blind denoising not only considers the original value of the image (multiplied by the exposure time and shot noise), but also takes into account the square term of the shot noise and the read noise, and is adjusted by using the square root and the proportional coefficient, making it more robust to the full scene. For example, the formula for the modified non-blind denoising conversion is shown as follows:

[0060] img_vst = vst(img * expo, read * expo, shot * sqrt(expo))

[0061] In the formula, img represents the pixel value of the original image, which can be understood as the brightness or color information of any point in the image. expo represents the exposure parameter, which is used to adjust the gain or attenuation of the pixel value, similar to adjusting the brightness of the image. Multiplying the original pixel value img means adjusting the overall exposure of the image. read represents the readout noise, which usually refers to the inherent noise level of the image sensor during the readout process. shot represents the shot noise, which originates from the photodiode noise and is related to the photon arrival amount, representing the random fluctuations in the image. sqrt(expo) represents the square root, which is used to perform operations on the parameter exp here, and may be used to adjust the weight of the noise or its influence. It can be understood that the input image pixel values are adjusted, taking into account adjusting the exposure (expo multiplied to increase or decrease the brightness), the readout noise (read multiplied by expo for adjustment), and the shot noise (shot multiplied by the square root of exp(o)` adjusted weight) to achieve the purpose of optimization or special visual effects processing.

[0062] Further, as Figure 2 shown, a schematic flowchart of the training steps of a residual branch model provided in this embodiment is shown. Designing a unified model architecture mainly relies on two design philosophies: one is decoupling, which is an efficient model design paradigm, and composite tasks are disassembled into subtasks through decoupling. The other is that for composite tasks, the most important thing is to determine the computing power allocation, connection order, and training relationship between subtasks. The task subtasks include noise reduction, demosaicing, and high dynamic range fusion.

[0063] The specific details are as follows:

[0064] S201, Obtain the training frame images that have been divided into each moment in the training HDR video.

[0065] Divide the images in the to-be-processed HDR video according to the preset moment nodes to obtain the frame images at each moment. Among them, the frame images include short-frame images and long-frame images. The main differences between the short-frame and long-frame images lie in the exposure time and exposure gain.

[0066] S202, Use the exposure ratio of the short-frame image and the long-frame image to perform brightness alignment processing on the short-frame image.

[0067] Due to the differences in exposure time and exposure gain, the long-frame includes more dark information, while the short-frame includes more bright details. The brightness of the long-frame and the short-frame can be aligned through the exposure ratio. The short-frame is aligned in terms of exposure before input, that is, the brightness is made consistent with the long-frame.

[0068] S203, Denoise the short-frame image through a preset non-blind denoising method.

[0069] Non-blind denoising is adopted for data processing, but the standard non-blind denoising transformation (here, the standard non-blind denoising refers to converting the noise distribution into a Gaussian distribution with a variance of 1) is not used. Instead, a modified non-blind denoising transformation formula is used, and the denoising formula is not elaborated here. The actual exposure ratio is adopted when processing HDR videos, and the exposure ratio can be set to 1 when processing NonHDR videos, that is, it degrades to the standard vst formula for NonHDR.

[0070] S204. Use the short-frame image to train the first processing branch and the fourth processing branch to obtain the first image and the first branch weight.

[0071] S205. Fix the first processing branch and the fourth processing branch to the first weight, and use the first image and the long-frame image to train the second processing branch and the fourth processing branch to obtain the second image and the second branch weight.

[0072] S206. Fix the second processing branch and the fourth processing branch to the second weight, and use the second image and the historical-frame image to train the third processing branch and the fourth processing branch to obtain the third image and the third branch weight.

[0073] S207. Assign the first branch weight, the second branch weight, and the third branch weight to the corresponding processing branches respectively to obtain the trained residual branch model.

[0074] S204 to S207 are to set four processing branches for the initial model under the residual branch model, and based on the training frame images, train the initial model to obtain the trained residual branch model. The specific details during the training process are as Figure 3 shown in the schematic diagram of the processing logic of a residual branch model provided in this embodiment. The specific details are as follows:

[0075] The model includes 4 parts: Short_unet, Long_unet, Time_unet, and Demosaic, and the connection order is as Figure 3 shown, and the specific training includes 3 stages:

[0076] (1) Train Short_unet and Demosaic until convergence. The input is the original short frame.

[0077] (2) Fix the Short_unet weight, train Long_unet and Demosaic until convergence, and Demosaic inherits the weight of the first stage. The input is the original long frame and the abstract feature map1 output by Short_unet.

[0078] (3) Fix the weights of Short_unet and Long_unet, train Time_unet and Demosaic until convergence, and Demosaic inherits the weights of the second stage. The input is the backpropagated historical frames and the abstract feature map2 output by Long_unet.

[0079] Each unet can also use other network structures, such as transform, etc., as long as it has the ability of feature extraction and non-linear mapping. Each part adopts a residual structure, that is, the original short frame plus 4 learned residuals to get the final output. Without considering Demosaic, the computing power of the other 3 network structures is allocated 1:1:1.

[0080] During inference, the entire model participates in the calculation when processing HDR videos, and the inputs include short frames, long frames, and historical frames; when processing NonHDR videos, the long frame input is all 0, and the role of Long_unet is masked, that is, the inputs include the current frame and historical frames. NonHDR can be regarded as a special case of HDR.

[0081] Among them, the design of Demosaic only requires two layers. The following gives an example:

[0082]

[0083] First, convert the four-channel 4*H / 2*W / 2 RAW image to 4*H*W through unsample. The shortcut branch takes the average of the green channel and converts it to 3*H*W. The residual branch includes two layers of convolution.

[0084] In this embodiment, by decoupling video processing, corresponding processing branches are set for noise reduction, fusion, and demosaicing respectively and connected through a residual structure, clarifying the processing logic relationship between each processing branch. As a result, the function of the trained residual branch model is more comprehensive, capable of processing noise reduction, demosaicing, and fusion tasks in videos, while taking into account the processing of both HDR type and NonHDR type videos, and solving the problem of poor generality in related methods where processing models are independently set for different tasks.

[0085] Further, as Figures 1 to 3 a specific implementation of the method shown, this embodiment provides a video processing device, as Figure 4 shown, the device includes: an acquisition unit 41, a division unit 42, and a processing unit 43.

[0086] The acquisition unit 41 is configured to acquire the video to be processed;

[0087] The division unit 42 is configured to divide the images in the video to be processed according to a preset time node to obtain frame images at each moment;

[0088] A processing unit 43, configured to input the frame image at each moment into a trained residual branch model for image processing to obtain a video processing result;

[0089] Wherein, each branch in the residual branch model adopts a residual structure, and at least includes a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame image according to the processing result of the second processing branch, and the fourth processing branch is used for demosaicing processing in the denoising processing and integration processing.

[0090] In a specific application scenario, the acquisition unit 41 is specifically configured to acquire a training video; the training video includes training frame images divided into each moment; set four processing branches for the initial model and allocate computing power resources for each processing branch; based on the training frame images, train the initial model to obtain the trained residual branch model.

[0091] In a specific application scenario, the processing unit 43 is further specifically configured to use the short frame image to train the first processing branch and the fourth processing branch to obtain a first processing result output by the first processing branch; the first processing result includes a first image and a first branch weight; fix the first processing branch and the fourth processing branch to the first weight, and use the first image and the long frame image to train the second processing branch and the fourth processing branch to obtain a second processing result output by the second processing branch; the second processing result includes a second image and a second branch weight; fix the second processing branch and the fourth processing branch to the second weight, and use the second image and the historical frame image to train the third processing branch and the fourth processing branch to obtain a third processing result output by the third processing branch; the third processing result includes a third image and a third branch weight; allocate the first branch weight, the second branch weight, and the third branch weight to the corresponding processing branches respectively to obtain the trained residual branch model; wherein, the fourth processing branch is used for demosaicing processing during the training process; the historical frame image is the third image of the frame image at the previous moment based on the current moment; the third image is a processing result obtained by fusing the first image and the second image with the short frame image as the reference frame.

[0092] In a specific application scenario, the processing unit 43 is further specifically configured to set the computing power resource allocation ratios of the first processing branch, the second processing branch, and the third processing branch to 1:1:1.

[0093] In a specific application scenario, the partitioning unit 42 is further specifically configured to determine the type of the video to be processed; in the case where the type of the video to be processed is the NonHDR type, the second processing branch is blocked.

[0094] In a specific application scenario, the processing unit 43 is further specifically configured to perform brightness alignment processing on the short-frame image by using the exposure ratio of the short-frame image and the long-frame image; and / or perform denoising on the short-frame image by a preset non-blind denoising method.

[0095] It should be noted that for other corresponding descriptions of the various functional units involved in the video processing method provided in this embodiment, reference can be made to Figures 1 to 3 the corresponding descriptions therein, which will not be elaborated here.

[0096] Based on the method as shown in Figures 1 to 3 above, correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as shown in Figures 1 to 3 above is implemented.

[0097] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in the various implementation scenarios of this application.

[0098] Based on the method as shown in Figures 1 to 3 above, and Figure 4 the virtual device embodiment as shown in Figures 1 to 3 above, for the purpose of achieving the above object, this embodiment of the application further provides an electronic device, such as an intelligent terminal such as a smart phone, a tablet computer, a drone, a smart robot, etc., and the device includes a storage medium and a processor; the storage medium is used for storing a computer program; the processor is used for executing the computer program to implement the method as shown in

[0099] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a Radio Frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.

[0100] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and it may include more or fewer components, or combine some components, or have different component arrangements.

[0101] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication between other hardware and software in the information processing physical device.

[0102] Based on the above method as Figures 1 to 3 shown, and Figure 4 the virtual device embodiment shown, this embodiment further provides a chip, including one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from the memory of the electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to execute the above method as Figures 1 to 3 shown.

[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or by hardware. By applying the solution of this embodiment, compared with the related art, in this embodiment, by decoupling video processing, corresponding processing branches are respectively set for noise reduction, fusion, and demosaicing and connected through a residual structure, clarifying the processing logic relationship between the processing branches, so that the trained residual branch model has a more comprehensive function and can handle noise reduction, demosaicing, and fusion tasks in videos, while taking into account the processing of HDR type and NonHDR type videos, and solving the problem of poor generality in the related methods where processing models are independently set for different tasks.

[0104] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0105] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments described herein, but rather will conform to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A video processing method, characterized in that, Including: Obtain the video to be processed; Divide the images in the video to be processed according to preset time nodes to obtain frame images for each moment; Input the frame images for each moment into the trained residual branch model for image processing to obtain the video processing result; Among them, each branch in the residual branch model adopts a residual structure, and at least includes a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame image according to the processing result of the second processing branch, and the fourth processing branch is used for demosaicking processing in the denoising processing and integration processing.

2. The method according to claim 1, characterized in that When the residual branch model includes four processing branches, the steps of obtaining the trained residual branch model include: Obtain a training video; the training video includes training frame images that have been divided into each moment; Set four processing branches for the initial model and allocate computing power resources for each processing branch; Based on the training frame images, train the initial model to obtain the trained residual branch model.

3. The method according to claim 2, wherein The training frame images include short frame images and long frame images; The training the initial model based on the training frame images to obtain the trained residual branch model includes: Use the short frame images to train the first processing branch and the fourth processing branch to obtain the first processing result output by the first processing branch; the first processing result includes a first image and a first branch weight; Fix the first processing branch and the fourth processing branch to the first weight, and use the first image and the long frame images to train the second processing branch and the fourth processing branch to obtain the second processing result output by the second processing branch; the second processing result includes a second image and a second branch weight; Fix the second processing branch and the fourth processing branch to the second weight, and use the second image and the historical frame images to train the third processing branch and the fourth processing branch to obtain the third processing result output by the third processing branch; the third processing result includes a third image and a third branch weight; Respectively allocate the first branch weight, the second branch weight, and the third branch weight to the corresponding processing branches to obtain the trained residual branch model; Among them, the fourth processing branch is used for demosaicking processing during the training process; the historical frame image is the third image of the frame image at the previous moment based on the current moment; the third image is the processing result obtained by fusing the first image and the second image with the short frame image as the reference frame.

4. The method according to claim 3, wherein Allocating computing power resources for each processing branch includes: Set the computing power resource allocation ratios of the first processing branch, the second processing branch, and the third processing branch to 1:1:

1.

5. The method according to claim 1, wherein After obtaining the video to be processed, the method further includes: Judge the type of the video to be processed; When the video type to be processed is of the NonHDR type, the second processing branch is blocked.

6. The method according to claim 5, characterized in that, When the video type to be processed is of the HDR type, the frame image at each moment further includes a short-frame image and a long-frame image; Before inputting the frame image at each moment into a trained residual branch model for image processing, the method further includes: Performing brightness alignment processing on the short-frame image by using the exposure ratio between the short-frame image and the long-frame image; And / or Performing denoising on the short-frame image by a preset non-blind denoising method.

7. A video processing device, characterized in that, Including: An acquisition unit configured to acquire a video to be processed; A division unit configured to divide the images in the video to be processed according to preset time nodes to obtain a frame image at each moment; A processing unit configured to input the frame image at each moment into a trained residual branch model for image processing to obtain a video processing result; Wherein, each branch in the residual branch model adopts a residual structure and at least includes a first processing branch, a second processing branch, a third processing branch, and a fourth processing branch; the first processing branch is used for denoising processing, the second processing branch is used for denoising processing according to the processing result of the first processing branch, the third processing branch is used for integrating the frame image according to the processing result of the second processing branch, and the fourth processing branch is used for performing demosaicing processing in the denoising processing and the integration processing.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

10. A chip, characterized in that, Including one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from the memory of the electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to execute the method according to any one of claims 1 to 6.