A video enhancement method, model training method and device

By performing feature extraction and model training on the reference video frames and adjacent video frames of motion video, an offset matrix and weight matrix are generated, which solves the problem of low enhancement quality for motion video in the prior art, and achieves high-quality enhancement of RAW image videos.

CN119090755BActive Publication Date: 2025-05-30INTELLINDUST INFORMATION TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310650683.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-02
Publication Date
2025-05-30
Estimated Expiration
2043-06-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively enhance videos for moving videos, especially when maintaining video clarity and reducing noise, blurred details and color distortion.

Method used

By determining the reference video frame in the original video and its adjacent video frames, and inputting it to pre-trained enhancement parameters to determine the model, an offset matrix and a weight matrix are generated to calculate the enhanced pixel value of each video frame, and then perform video enhancement.

Benefits of technology

High-quality enhancement of RAW image videos is achieved, improving the clarity and sensory quality of the video, reducing noise and blurring of details, and enhancing the overall quality of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119090755B_ABST
    Figure CN119090755B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a video enhancement method, a model training method, and an apparatus. The video enhancement method includes: determining a specified reference video frame from each video frame of an original video; inputting the reference video frame and an adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model to obtain an offset matrix and a weight matrix; for each first image region in the reference video frame, determining a corresponding second image region in the adjacent video frame; for each first image region in the reference video frame, calculating an enhanced pixel value of the first image region according to the weight of the first image region and the weight of the corresponding second image region in the adjacent video frame to obtain a target video frame corresponding to the reference video frame; and obtaining a target video corresponding to the original video based on the target video frame, which can enhance RAW images; and can improve the quality of the obtained target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a video enhancement method, a model training method, and an apparatus. Background Art

[0002] When a video is captured by a capture device, it is affected by the capture environment. For example, when the capture environment is a low-illumination environment at night, due to the limited number of incident photons of the sensor of the capture device, the signal-to-noise ratio of each video frame in the captured video may be low, and thus the clarity of the video is also low. Video enhancement utilizes the redundant information of each video frame in the video in the time domain and the correlation between the video frames in the video to process each video frame in the video, so as to improve the signal-to-noise ratio of each video frame in the video, and further improve the clarity of the video, that is, improve the sensory quality of the video. Correspondingly, the experience of a user watching the video can be improved. For a video in a static scene captured by a static capture device (which can be referred to as a static video), the 3D DNR (3-Dimension Digital Noise Reducer) technology can be used to perform video enhancement on the static video to obtain an enhanced video with higher quality. However, when the capture device is moving, or when a video in a moving scene is captured by the capture device (which can be referred to as a moving video), if the 3D DNR technology is used to perform video enhancement on the moving video, the quality of the enhanced video is low, and problems such as serious noise, blurred details, color distortion, and motion ghosting often occur in the video frames in the enhanced video. In order to perform video enhancement on a moving video, for each video frame in the moving video, multiple video frames in the moving video need to be used to enhance the video frame to enhance the information of the moving area in the video frame.

[0003] In the prior art, when performing video enhancement on a video, the video frames in the video can be registered first through a preset registration method (for example, a registration method based on feature point matching, or a registration method based on block matching technology). Registration is to determine the mapping relationship between video frames, such as determining the corresponding feature points of a feature point in one video frame in another video frame, or determining the corresponding image area of an image area in one video frame in another video frame. Subsequently, video enhancement is performed according to the mapping relationship between the video frames.

[0004] However, the preset registration method is applicable to RGB (Red, Green, Blue) images. The pixel values of RGB images have a non-linear relationship with the intensity of incident light, making it difficult to analyze RGB images. The pixel values of RAW (raw) images have a good linear relationship with the incident light intensity and retain the most original information of the image. Therefore, based on the physical imaging process, it is less difficult to accurately model the pixel values of RAW images, that is, it is less difficult to analyze RAW images. However, the preset registration method may not be applicable to RAW images. When each video frame in a video is a RAW image, the video frames in the video cannot be registered, and thus the video cannot be enhanced. Moreover, the accuracy of using the preset registration method to register video frames in a video is relatively low. For example, when using a registration method based on feature point matching to register video frames, since the feature points in each video frame are not evenly distributed in the video frame, for an area in a video frame where no feature points are detected, it is impossible to determine the corresponding feature points in another video frame. Or, when using a registration method based on block matching technology to register video frames, multiple image regions in one video frame may correspond to the same image region in another video frame, and thus ghosting may occur in the enhanced video frame. As a result, the quality of the enhanced video is not high. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a video enhancement method, a model training method, and a device to enhance RAW images and improve the quality of the obtained target video. The specific technical solutions are as follows:

[0006] In the first aspect of the embodiments of the present invention, a video enhancement method is first provided. The method includes:

[0007] Determine a specified reference video frame from each video frame of the original video; wherein each video frame in the original video is a RAW image;

[0008] Input the reference video frame and the adjacent video frames of the reference video frame into a pre-trained enhancement parameter determination model to obtain an offset matrix and a weight matrix output by the enhancement parameter determination model;

[0009] Wherein, the elements in the offset matrix represent the offset amount of the first image region in the reference video frame relative to the adjacent video frame; the elements in the weight matrix are in one-to-one correspondence with the first image region in the reference video frame and the second image region in the adjacent video frame. If an element corresponds to the first image region in the reference video frame, the element represents the weight of the corresponding first image region. If an element corresponds to the second image region in the adjacent video frame, the element represents the weight of the corresponding second image region;

[0010] For each first image region in the reference video frame, based on the position of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame, determine the corresponding second image region of the first image region in the adjacent video frame;

[0011] For each first image region in the reference video frame, calculate the enhanced pixel value of the first image region according to the weight of the first image region and the weight of the corresponding second image region of the first image region in the adjacent video frame, to obtain the target video frame corresponding to the reference video frame;

[0012] Based on the target video frame, obtain the target video corresponding to the original video.

[0013] Optionally, the enhancement parameter determination model includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network;

[0014] The step of inputting the reference video frame and the adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model to obtain the offset matrix and the weight matrix output by the enhancement parameter determination model includes:

[0015] Input the reference video frame and the adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model, and perform feature extraction on the reference video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the reference video frame and the image features of the adjacent video frame, as the first feature vectors;

[0016] Perform feature extraction on the first feature vectors through the second feature extraction network to obtain second feature vectors;

[0017] Perform normalization processing on the second feature vectors through the offset prediction network to obtain the corresponding offset matrix;

[0018] Perform feature extraction on the first feature vectors through the third feature extraction network to obtain third feature vectors;

[0019] Perform normalization processing on the third feature vectors through the weight prediction network to obtain the corresponding weight matrix.

[0020] Optionally, the step of for each first image region in the reference video frame, based on the position of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame, determining the corresponding second image region of the first image region in the adjacent video frame includes:

[0021] For each first image region in the reference video frame, obtain the coordinates of the first image region in the reference video frame;

[0022] Calculate the sum of the coordinates of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame as the reference coordinates;

[0023] In the adjacent video frame, determine the image region corresponding to the reference coordinates to obtain the second image region corresponding to the first image region in the adjacent video frame.

[0024] Optionally, the calculating the enhanced pixel value of the first image region according to the weight of the first image region and the weight of the second image region corresponding to the first image region in the adjacent video frame to obtain the target video frame corresponding to the reference video frame includes:

[0025] For each pixel point in the first image region of the reference video frame, if the reference coordinates corresponding to the pixel point in the adjacent video frame are integers, calculate the weighted sum of the pixel value of the pixel point and the pixel value of the pixel point at the reference coordinates according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinates corresponding to the pixel point in the adjacent video frame belong, to obtain the target video frame corresponding to the reference video frame;

[0026] If the reference coordinates corresponding to the pixel point in the adjacent video frame are not integers, determine the pixel value corresponding to the reference coordinates based on the pixel values of the pixel points adjacent to the reference coordinates, and calculate the weighted sum of the pixel value of the pixel point and the pixel value corresponding to the reference coordinates according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinates corresponding to the pixel point in the adjacent video frame belong, to obtain the target video frame corresponding to the reference video frame.

[0027] Optionally, the determining the pixel value corresponding to the reference coordinates based on the pixel values of the pixel points adjacent to the reference coordinates includes:

[0028] Based on the nearest neighbor interpolation algorithm, determine the pixel value of the pixel point with the smallest distance from the reference coordinates as the pixel value corresponding to the reference coordinates;

[0029] Or,

[0030] Based on the bilinear interpolation algorithm, calculate the pixel value corresponding to the reference coordinates based on the pixel values of a preset number of pixel points within a preset neighborhood range of the reference coordinates.

[0031] Optionally, the adjacent video frames of the reference video frame include: the video frames in the original video that are before the reference video frame, and / or, the video frames that are after the reference video frame.

[0032] In a second aspect of the embodiments of the present invention, a model training method is provided for generating the enhanced parameter determination model according to any one of the first aspects above. The method includes:

[0033] Obtain a first sample video frame and a second sample video frame; wherein, the first sample video frame and the second sample video frame are RAW images; the first sample video frame is obtained by adding noise to the second sample video frame;

[0034] Input the first sample video frame and the adjacent video frames of the first sample video frame in the sample video into the enhanced parameter determination model with an initial structure, and obtain the offset matrix and weight matrix output by the enhanced parameter determination model with the initial structure;

[0035] Wherein, the elements in the offset matrix represent the offset amount of the first image area in the first sample video frame relative to the adjacent video frame; the elements in the weight matrix are in one-to-one correspondence with the first image area in the first sample video frame and the second image area in the adjacent video frame; if an element corresponds to the first image area in the first sample video frame, the element represents the weight of the corresponding first image area; if an element corresponds to the second image area in the adjacent video frame, the element represents the weight of the corresponding second image area;

[0036] For each first image area in the first sample video frame, based on the position of the first image area in the first sample video frame and the offset amount of the first image area relative to the adjacent video frame, determine the corresponding second image area of the first image area in the adjacent video frame;

[0037] For each first image area in the first sample video frame, calculate the enhanced pixel value of the first image area according to the weight of the first image area and the weight of the corresponding second image area of the first image area in the adjacent video frame, and obtain the predicted video frame corresponding to the first sample video frame;

[0038] Calculate the loss function value representing the difference between the predicted video frame and the second sample video frame;

[0039] Based on the calculated loss function value, adjust the model parameters of the enhanced parameter determination model with the initial structure until a preset convergence condition is reached, and obtain the trained enhanced parameter determination model.

[0040] Optionally, the enhancement parameter determination model of the initial structure includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network;

[0041] The step of inputting the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model of the initial structure to obtain the offset matrix and the weight matrix output by the enhancement parameter determination model of the initial structure includes:

[0042] Input the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model of the initial structure, and perform feature extraction on the first sample video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the first sample video frame and the image features of the adjacent video frame, as the first feature vector;

[0043] Perform feature extraction on the first feature vector through the second feature extraction network to obtain a second feature vector;

[0044] Perform normalization processing on the second feature vector through the offset prediction network to obtain the corresponding offset matrix;

[0045] Perform feature extraction on the first feature vector through the third feature extraction network to obtain a third feature vector;

[0046] Perform normalization processing on the third feature vector through the weight prediction network to obtain the corresponding weight matrix.

[0047] In the third aspect of the embodiments of the present invention, a video enhancement device is provided, and the device includes:

[0048] A reference video frame determination module, configured to determine a specified reference video frame from each video frame of the original video; wherein, each video frame in the original video is a RAW image;

[0049] A matrix acquisition module, configured to input the reference video frame and the adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model to obtain the offset matrix and the weight matrix output by the enhancement parameter determination model;

[0050] Among them, the elements in the offset matrix represent the offset of the first image region in the reference video frame relative to the adjacent video frame; the elements in the weight matrix correspond one-to-one to the first image region in the reference video frame and the second image region in the adjacent video frame; if an element corresponds to the first image region in the reference video frame, this element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, this element represents the weight of the corresponding second image region.

[0051] An image region determination module, configured to, for each first image region in the reference video frame, determine a second image region corresponding to the first image region in the adjacent video frame based on the position of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame.

[0052] A target video frame acquisition module, configured to, for each first image region in the reference video frame, calculate an enhanced pixel value of the first image region according to the weight of the first image region and the weight of the second image region corresponding to the first image region in the adjacent video frame, and obtain a target video frame corresponding to the reference video frame.

[0053] A target video acquisition module, configured to obtain a target video corresponding to the original video based on the target video frame.

[0054] Optionally, the enhancement parameter determination model includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network.

[0055] The matrix acquisition module is specifically configured to:

[0056] Input the reference video frame and the adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model, perform feature extraction on the reference video frame and the adjacent video frame through the first feature extraction network, and obtain a feature vector representing the image feature of the reference video frame and the image feature of the adjacent video frame as a first feature vector.

[0057] Perform feature extraction on the first feature vector through the second feature extraction network to obtain a second feature vector.

[0058] Perform normalization processing on the second feature vector through the offset prediction network to obtain a corresponding offset matrix.

[0059] Perform feature extraction on the first feature vector through the third feature extraction network to obtain a third feature vector.

[0060] The third feature vector is normalized by the weight prediction network to obtain a corresponding weight matrix.

[0061] Optionally, the image region determination module is specifically configured to:

[0062] For each first image region in the reference video frame, obtain the coordinates of the first image region in the reference video frame;

[0063] Calculate the sum of the coordinates of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame as the reference coordinates;

[0064] In the adjacent video frame, determine the image region corresponding to the reference coordinates to obtain the second image region corresponding to the first image region in the adjacent video frame.

[0065] Optionally, the target video frame acquisition module is specifically configured to:

[0066] For each pixel point in the first image region of the reference video frame, if the reference coordinates corresponding to the pixel point in the adjacent video frame are integers, calculate the weighted sum of the pixel value of the pixel point and the pixel value of the pixel point at the reference coordinates according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinates corresponding to the pixel point in the adjacent video frame belong, to obtain the target video frame corresponding to the reference video frame;

[0067] If the reference coordinates corresponding to the pixel point in the adjacent video frame are not integers, determine the pixel value corresponding to the reference coordinates based on the pixel values of the pixel points adjacent to the reference coordinates, and calculate the weighted sum of the pixel value of the pixel point and the pixel value corresponding to the reference coordinates according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinates corresponding to the pixel point in the adjacent video frame belong, to obtain the target video frame corresponding to the reference video frame.

[0068] Optionally, the target video frame acquisition module is specifically configured to:

[0069] Based on the nearest neighbor interpolation algorithm, determine the pixel value of the pixel point with the smallest distance from the reference coordinates as the pixel value corresponding to the reference coordinates;

[0070] Or,

[0071] Based on the bilinear interpolation algorithm, calculate the pixel value corresponding to the reference coordinates based on the pixel values of a preset number of pixel points within a preset neighborhood range of the reference coordinates.

[0072] Optionally, the adjacent video frames of the reference video frame include: video frames in the original video that are before the reference video frame, and / or video frames that are after the reference video frame.

[0073] In a fourth aspect of the embodiments of the present invention, there is provided a model training device for generating the enhancement parameter determination model according to any one of the first aspects above. The device includes:

[0074] A sample video frame acquisition module, configured to acquire a first sample video frame and a second sample video frame; wherein, the first sample video frame and the second sample video frame are RAW images; the first sample video frame is obtained by adding noise to the second sample video frame;

[0075] A matrix acquisition module, configured to input the first sample video frame and the adjacent video frames of the first sample video frame in the sample video into the enhancement parameter determination model with an initial structure, and obtain an offset matrix and a weight matrix output by the enhancement parameter determination model with the initial structure;

[0076] Wherein, the elements in the offset matrix represent the offset amount of the first image area in the first sample video frame relative to the adjacent video frame; the elements in the weight matrix are in one-to-one correspondence with the first image area in the first sample video frame and the second image area in the adjacent video frame; if an element corresponds to the first image area in the first sample video frame, the element represents the weight of the corresponding first image area; if an element corresponds to the second image area in the adjacent video frame, the element represents the weight of the corresponding second image area;

[0077] An image area determination module, configured to, for each first image area in the first sample video frame, determine the second image area corresponding to the first image area in the adjacent video frame based on the position of the first image area in the first sample video frame and the offset amount of the first image area relative to the adjacent video frame;

[0078] A predicted video frame acquisition module, configured to, for each first image area in the first sample video frame, calculate the enhanced pixel value of the first image area according to the weight of the first image area and the weight of the second image area corresponding to the first image area in the adjacent video frame, and obtain the predicted video frame corresponding to the first sample video frame;

[0079] A difference calculation module, configured to calculate a loss function value representing the difference between the predicted video frame and the second sample video frame;

[0080] An adjustment module is used to adjust the model parameters of the enhancement parameter determination model of the initial structure based on the calculated loss function value until a preset convergence condition is reached, and a trained enhancement parameter determination model is obtained.

[0081] Optionally, the enhancement parameter determination model of the initial structure includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network;

[0082] The matrix acquisition module is specifically used for:

[0083] Input the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model of the initial structure, and extract features from the first sample video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the first sample video frame and the image features of the adjacent video frame, as the first feature vector;

[0084] Extract features from the first feature vector through the second feature extraction network to obtain a second feature vector;

[0085] Normalize the second feature vector through the offset prediction network to obtain a corresponding offset matrix;

[0086] Extract features from the first feature vector through the third feature extraction network to obtain a third feature vector;

[0087] Normalize the third feature vector through the weight prediction network to obtain a corresponding weight matrix.

[0088] In the fifth aspect of the embodiments of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0089] The memory is used to store a computer program;

[0090] When the processor is used to execute the program stored in the memory, it implements the video enhancement method described in any one of the first aspects above.

[0091] In the sixth aspect of the embodiments of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0092] The memory is used to store a computer program;

[0093] A processor, when executing a program stored in a memory, implements the model training method according to any one of the above second aspects.

[0094] In a seventh aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the video enhancement method according to any one of the above first aspects is implemented.

[0095] In an eighth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the model training method according to any one of the above second aspects is implemented.

[0096] The embodiments of the present invention further provide a computer program product containing instructions, which when running on a computer, causes the computer to execute the steps of the video enhancement method according to any one of the above first aspects.

[0097] The embodiments of the present invention further provide a computer program product containing instructions, which when running on a computer, causes the computer to execute the steps of the model training method according to any one of the above second aspects.

[0098] A video enhancement method provided by the embodiments of the present invention includes: determining a specified reference video frame from each video frame of an original video; where each video frame in the original video is a RAW image; inputting the reference video frame and adjacent video frames of the reference video frame into a pre-trained enhancement parameter determination model to obtain an offset matrix and a weight matrix output by the enhancement parameter determination model; where elements in the offset matrix represent the offset amount of a first image region in the reference video frame relative to an adjacent video frame; elements in the weight matrix are in one-to-one correspondence with the first image region in the reference video frame and a second image region in the adjacent video frame; if an element corresponds to the first image region in the reference video frame, the element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, the element represents the weight of the corresponding second image region; for each first image region in the reference video frame, determining a corresponding second image region of the first image region in the adjacent video frame based on the position of the first image region in the reference video frame and the offset amount of the first image region relative to the adjacent video frame; for each first image region in the reference video frame, calculating an enhanced pixel value of the first image region according to the weight of the first image region and the weight of the corresponding second image region of the first image region in the adjacent video frame to obtain a target video frame corresponding to the reference video frame; and obtaining a target video corresponding to the original video based on the target video frame.

[0099] Based on the above processing, the original video with each video frame being a RAW image can be enhanced based on the offset matrix and the weight matrix to obtain the target video corresponding to the original video, that is, the enhancement of the RAW image can be realized. Moreover, since the offset matrix and the weight matrix are obtained based on the pre-trained enhancement parameter determination model, for each first image region in the reference video frame, the second image region corresponding to the first image region in the adjacent video frame of the reference video frame can be determined based on the offset matrix, which can improve the accuracy of the determined second image region. Furthermore, based on the weights of the first image regions in the reference video frame and the weights of the second image regions corresponding to the first image regions, the reference video frame can be enhanced, which can improve the quality of the obtained target video.

[0100] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0102] Figure 1 The first flowchart of the video enhancement method provided by the embodiment of the present invention;

[0103] Figure 2 The second flowchart of the video enhancement method provided by the embodiment of the present invention;

[0104] Figure 3 The third flowchart of the video enhancement method provided by the embodiment of the present invention;

[0105] Figure 4 The fourth flowchart of the video enhancement method provided by the embodiment of the present invention;

[0106] Figure 5 The fifth flowchart of the video enhancement method provided by the embodiment of the present invention;

[0107] Figure 6 The first flowchart of the model training method provided by the embodiment of the present invention;

[0108] Figure 7 The second flowchart of the model training method provided by the embodiment of the present invention;

[0109] Figure 8 A working principle diagram of the enhancement parameter determination model provided by the embodiment of the present invention;

[0110] Figure 9 A comparative example diagram of the video enhancement effect provided by the embodiment of the present invention;

[0111] Figure 10 A structural diagram of the video enhancement device provided by the embodiment of the present invention;

[0112] Figure 11 A structural diagram of the model training device provided by the embodiment of the present invention;

[0113] Figure 12 The first structural diagram of the electronic device provided by the embodiment of the present invention;

[0114] Figure 13 The second structural diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0115] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on the present invention belong to the protection scope of the present invention.

[0116] In the related art, when performing video enhancement on a video, it is necessary to first determine the mapping relationship between video frames through a preset registration method, and then perform video enhancement according to the mapping relationship between video frames. However, the preset registration method is not applicable to RAW images. When each video frame in the video is a RAW image, it is impossible to register each video frame in the video, and thus it is impossible to perform video enhancement on the video; moreover, the accuracy of registering video frames in the video using the preset registration method is relatively low, and as a result, the quality of the enhanced video is not high.

[0117] To solve the above problems, an embodiment of the present invention provides a video enhancement method, which is applied to an electronic device. The electronic device can obtain the original video to be enhanced and determine the specified reference video frame in the original video. Each video frame in the original video is a RAW image. Further, the electronic device can, according to the video enhancement method provided by the embodiment of the present invention, determine a model through pre-trained enhancement parameters to obtain an offset matrix and a weight matrix. For each first image region in the reference video frame, based on the position of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame, determine the corresponding second image region of the first image region in the adjacent video frame; for each first image region in the reference video frame, calculate the enhanced pixel value of the first image region according to the weight of the first image region and the weight of the corresponding second image region of the first image region in the adjacent video frame to obtain the target video frame corresponding to the reference video frame; based on the target video frame, obtain the target video corresponding to the original video, which can realize the enhancement of the RAW image; and can improve the quality of the obtained target video.

[0118] See Figure 1 , Figure 1 is the first flowchart of the video enhancement method provided by the embodiment of the present invention, and the method may include the following steps:

[0119] S101: Determine the specified reference video frame from each video frame of the original video.

[0120] Wherein, each video frame in the original video is a RAW image.

[0121] S102: Input the reference video frame and the adjacent video frame of the reference video frame into the pre-trained enhancement parameter determination model to obtain the offset matrix and the weight matrix output by the enhancement parameter determination model.

[0122] Wherein, the elements in the offset matrix represent the offset of the first image region in the reference video frame relative to the adjacent video frame; the elements in the weight matrix correspond one by one to the first image region in the reference video frame and the second image region in the adjacent video frame; if an element corresponds to the first image region in the reference video frame, the element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, the element represents the weight of the corresponding second image region.

[0123] S103: For each first image region in the reference video frame, based on the position of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame, determine the corresponding second image region of the first image region in the adjacent video frame.

[0124] S104: For each first image region in the reference video frame, calculate the enhanced pixel value of the first image region according to the weight of the first image region and the weight of the corresponding second image region of the first image region in the adjacent video frame, so as to obtain the target video frame corresponding to the reference video frame.

[0125] S105: Based on the target video frame, obtain the target video corresponding to the original video.

[0126] Based on the video enhancement method provided in the embodiments of the present invention, the original video including each video frame in RAW image format can be enhanced based on the offset matrix and the weight matrix to obtain the target video corresponding to the original video, that is, the enhancement of the RAW image can be realized; and since the offset matrix and the weight matrix are obtained based on the pre-trained enhancement parameter determination model, for each first image region in the reference video frame, the corresponding second image region of the first image region in the adjacent video frame of the reference video frame can be determined based on the offset matrix, which can improve the accuracy of the determined second image region. Furthermore, based on the weights of the first image regions in the reference video frame and the weights of the corresponding second image regions of the first image regions, the reference video frame can be enhanced, which can improve the quality of the obtained target video.

[0127] Regarding step S101, the original video is a RAW format video that has not undergone ISP (Image Signal Processing), and each video frame in the original video is a RAW image, and the pixel value of each pixel point in each RAW image is RAW domain data. The RAW domain data is the original data after the optical signal recorded by the photosensitive element of the acquisition device is converted into a digital signal. The pixel value of a pixel point can include single-channel RAW domain data. For example, the pixel value of a pixel point can be a G (Green) value recorded by a photosensitive element.

[0128] The electronic device can determine the video frames that need to be enhanced from each video frame of the original video to obtain the specified reference video frame. For example, each video frame in the original video can be a reference video frame, or some video frames in the original video can be reference video frames.

[0129] Regarding step S102, the adjacent video frames of the reference video frame include: the video frames in the original video that are before the reference video frame, and / or the video frames that are after the reference video frame.

[0130] The adjacent video frames of the reference video frame include the video frames in the original video that are before the reference video frame. When the reference video frame is the first video frame in the original video, since there are no video frames in the original video that are before the reference video frame, the video frames in the original video that are after the reference video frame can be selected as the adjacent video frames of the reference video frame.

[0131] The adjacent video frames of the reference video frame include the video frames in the original video that are after the reference video frame. When the reference video frame is the last video frame in the original video, since there are no video frames in the original video that are after the reference video frame, the video frames in the original video that are before the reference video frame can be selected as the adjacent video frames of the reference video frame.

[0132] The adjacent video frames of the reference video frame include: the first number of video frames in the original video that are before the reference video frame, and the second number of video frames that are after the reference video frame. The first number and the second number can be set based on requirements. When it is necessary to improve the quality of the obtained target video, the first number and the second number can be set to larger values; when it is necessary to improve the efficiency of obtaining the target video, the first number and the second number can be set to smaller values. Moreover, the first number and the second number can be the same, for example, the first number and the second number can be 2. Or, the first number and the second number can be different, for example, the first number can be 1 and the second number can be 3.

[0133] The enhancement parameter determination model can be a CNN (Convolutional Neural Network) model. The enhancement parameter determination model is trained based on the first sample video frame and the second sample video frame, and the first sample video frame is obtained by adding noise to the second sample video frame. The training method of the enhancement parameter determination model refers to the relevant introduction in the subsequent embodiments.

[0134] After the electronic device obtains the reference video frame, it can also obtain the pixel matrix composed of the pixel values of each pixel point of the reference video frame (which can be called the first pixel matrix). After the electronic device obtains the adjacent video frames of the reference video frame, it can also obtain the pixel matrix of the pixel values of each pixel point of the adjacent video frame (which can be called the second pixel matrix). Then, the electronic device can concatenate the first pixel matrix and the second pixel matrix to obtain the concatenated pixel matrix (which can be called the third pixel matrix).

[0135] Subsequently, the electronic device inputs the third pixel matrix into a pre-trained enhancement parameter determination model, and obtains an offset matrix and a weight matrix output by the enhancement parameter determination model. The elements in the offset matrix represent the offsets of the first image region in the reference video frame relative to the adjacent video frames of the reference video frame; the elements in the weight matrix correspond one-to-one to the first image region in the reference video frame and the second image region in the adjacent video frames of the reference video frame; if an element corresponds to the first image region in the reference video frame, this element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frames of the reference video frame, this element represents the weight of the corresponding second image region.

[0136] Exemplarily, the height and width dimensions of the third pixel matrix are h*w. The height and width dimensions of the offset matrix output by the enhancement parameter determination model are . An element in the offset matrix corresponds to a first image region in the reference video frame, then the size of a first image region in the reference video frame is 4*4, and the element corresponding to this first image region in the offset matrix represents the offsets of each pixel point in this first image region relative to the adjacent video frames of the reference video frame. The offset of a first image region relative to the adjacent video frames of the reference video frame includes: the offset of this first image region in the horizontal direction relative to the adjacent video frames of the reference video frame, and the offset of this first image region in the vertical direction relative to the adjacent video frames of the reference video frame.

[0137] The height and width dimensions of the weight matrix output by the enhancement parameter determination model are . If an element in the weight matrix corresponds to a first image region in the reference video frame, then the size of a first image region in the reference video frame is 4*4, and the element corresponding to this first image region in the weight matrix represents the weight of this first image region; if an element in the weight matrix corresponds to a second image region in an adjacent video frame of the reference video frame, then the size of a second image region in the adjacent video frame of the reference video frame is 4*4, and the element corresponding to this second image region in the weight matrix represents the weight of this second image region.

[0138] In some embodiments, the enhancement parameter determination model includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network.

[0139] Correspondingly, on the basis of Figure 1 , referring to Figure 2 , step S102 includes the following steps:

[0140] S1021: Input the reference video frame and the adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model, extract features from the reference video frame and the adjacent video frame through the first feature extraction network, and obtain feature vectors representing the image features of the reference video frame and the image features of the adjacent video frame, which are used as the first feature vectors.

[0141] S1022: Extract features from the first feature vectors through the second feature extraction network to obtain second feature vectors.

[0142] S1023: Normalize the second feature vectors through the offset prediction network to obtain the corresponding offset matrix.

[0143] S1024: Extract features from the first feature vectors through the third feature extraction network to obtain third feature vectors.

[0144] S1025: Normalize the third feature vectors through the weight prediction network to obtain the corresponding weight matrix.

[0145] The first feature extraction network may include multiple convolutional layers, and the multiple convolutional layers include: a convolutional layer with a 3×3 convolutional kernel and a stride of 1, and a convolutional layer with a 3×3 convolutional kernel and a stride of 2. Moreover, an activation function is connected after each convolutional layer included in the first feature extraction network. For example, the activation function can be a ReLu (Rectified Linear Units) function.

[0146] Through the multiple convolutional layers included in the first feature extraction network, extract features from the third pixel matrix, implement downsampling of the third pixel matrix, and obtain feature vectors (i.e., the first feature vectors) representing the image features of the reference video frame and the image features of the adjacent video frame of the reference video frame. If the height and width dimensions of the third pixel matrix are h*w, and the electronic device inputs the third pixel matrix into the pre-trained enhancement parameter determination model, then extract features from the third pixel matrix through the first feature network, and obtain a feature extraction result with a height and width dimension of , which is used as the first feature vector. And one element in the first feature vector can correspond to a 4*4 region in the reference video frame or the adjacent video frame of the reference video frame.

[0147] The second feature extraction network may include two convolutional layers, which are convolutional layers with a convolution kernel of 3×3 and a stride of 1. An activation function is connected after the first convolutional layer in the second feature extraction network. For example, the activation function can be a ReLu function. Since the output data of the second feature extraction network is the offset of each first image region in the reference video frame relative to the adjacent video frame of the reference video frame, and the offset includes the horizontal offset and the vertical offset, that is, there are two offsets for each first image region in the reference video frame relative to each adjacent video frame of the reference video frame. Therefore, each adjacent video frame of the reference video frame corresponds to two output channels of the second feature extraction network. Thus, the number of output channels of the second feature extraction network is twice the number of adjacent video frames of the reference video frame. If the reference video frame corresponds to 4 adjacent video frames, the number of output channels of the second feature extraction network is 8.

[0148] Through the two convolutional layers included in the second feature extraction network, feature extraction is performed on the first feature vector to obtain a feature vector (i.e., the second feature vector) representing the offset of the reference video frame relative to the adjacent video frame of the reference video frame. Subsequently, the offset prediction network normalizes the second feature vector to obtain a corresponding offset matrix. For example, the offset prediction network can use a normalization function to normalize the second feature vector. For example, the normalization function can be the tanh (hyperbolic tangent) function, which normalizes the second feature vector to the interval [-1, 1] to obtain a corresponding offset matrix.

[0149] The third feature extraction network may include two convolutional layers, which are convolutional layers with a convolution kernel of 3×3 and a stride of 1. An activation function is connected after the first convolutional layer in the third feature extraction network. For example, the activation function can be a ReLu function. Since the output data of the third feature extraction network is the weight of each first image region in the reference video frame and the weight of each second image region in the adjacent video frame of the reference video frame, that is, the reference video frame corresponds to one output channel of the third feature extraction network, and each adjacent video frame of the reference video frame also corresponds to one output channel of the third feature extraction network. Therefore, the number of output channels of the third feature extraction network is the same as the sum of the number of reference video frames and the number of adjacent video frames of the reference video frame. If one reference video frame corresponds to 4 adjacent video frames, the number of output channels of the third feature extraction network is 5.

[0150] Through two convolutional layers included in the third feature extraction network, feature extraction is performed on the first feature vector to obtain a feature vector (i.e., the third feature vector) representing the weights of the reference video frame and the adjacent video frames of the reference video frame. Subsequently, through the weight prediction network, normalization processing is performed on the third feature vector to obtain a corresponding weight matrix. For example, the weight prediction network can use a normalization function to perform normalization processing on the third feature vector. For example, the normalization function can be softmax (normalized exponential function), which normalizes the third feature vector to the interval [0, 1] to obtain a corresponding weight matrix, and the sum of the elements included in the weight matrix is 1.

[0151] Based on the above processing, inputting the third pixel matrix into the pre-trained enhancement parameter determination model can obtain the offset matrix and the weight matrix output by the enhancement parameter determination model, and the accuracy of the obtained offset matrix and weight matrix is higher; performing downsampling on the third pixel matrix can increase the receptive field of the first feature extraction network and reduce the computational complexity of subsequent processing, which can improve the efficiency of the enhancement parameter determination model in outputting the offset matrix and the weight matrix. Subsequently, based on the offset matrix, the weight matrix, and the adjacent video frames of the reference video frame, enhancing the reference video frame can improve the quality of the obtained target video and the efficiency of generating the target video.

[0152] Regarding step S103, for each first image region in the reference video frame, the element corresponding to the first image region in the offset matrix represents the offset amount of the first image region relative to the adjacent video frame of the reference video frame. Furthermore, according to the position of the first image region in the reference video frame and the offset amount of the first image region relative to the adjacent video frame of the reference video frame, the second image region corresponding to the first image region in the adjacent video frame of the reference video frame is obtained.

[0153] In some embodiments, on the basis of Figure 1 referring to Figure 3 , step S103 may include the following steps:

[0154] S1031: For each first image region in the reference video frame, obtain the coordinates of the first image region in the reference video frame.

[0155] S1032: Calculate the sum of the coordinates of the first image region in the reference video frame and the offset amount of the first image region relative to the adjacent video frame as the reference coordinates.

[0156] S1033: In the adjacent video frame, determine the image region corresponding to the reference coordinates to obtain the second image region corresponding to the first image region in the adjacent video frame.

[0157] For each first image region in the reference video frame, the position of the first image region in the reference video frame can be represented using coordinates. For example, the electronic device can establish a coordinate system with the vertex at the upper left corner of the reference video frame as the origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis. The first image region can be a rectangular region, and the electronic device can determine the coordinates of the four vertices of the first image region in the coordinate system, and the coordinates of the four vertices in the coordinate system can represent the position of the first image region in the reference video frame. Alternatively, the electronic device can determine the coordinates of two vertices of the first image region in the coordinate system, as well as the length and width of the first image region, and the coordinates of the two vertices in the coordinate system, as well as the length and width of the first image region, can represent the position of the first image region in the reference video frame.

[0158] In some embodiments, since the offset matrix output by the pre-trained enhancement parameter model is obtained through normalization processing, that is, each element in the offset matrix does not represent how many pixel points the first image region is offset relative to the adjacent video frame of the reference video frame. Therefore, the electronic device can first restore the elements in the offset matrix so that the restored elements can represent how many pixel points the first image region is offset relative to the adjacent video frame of the reference video frame, that is, obtain the actual offset of the first image region relative to the adjacent video frame of the reference video frame.

[0159] Exemplarily, for each first image region in the reference video frame, the electronic device can calculate the product of the offset of the first image region relative to the adjacent video frame of the reference video frame and a third number to obtain the actual offset of the first image region relative to the adjacent video frame of the reference video frame, including the actual offset of the first image region relative to the adjacent video frame of the reference video frame in the horizontal direction and the actual offset of the first image region relative to the adjacent video frame of the reference video frame in the vertical direction.

[0160] For example, if the motion range of a first image region is a circle with a radius of 64 pixel points, then the third number can be 64. If the offset of a first image region relative to the adjacent video frame of the reference video frame is 1, it means that the first image region has moved a distance of 64 pixel points relative to the adjacent video frame of the reference video frame, that is, the actual offset of the first image region relative to the adjacent video frame of the reference video frame is 64.

[0161] Then, the electronic device can calculate the sum of the abscissa of the first image region in the reference video frame and the actual horizontal offset of the first image region relative to the adjacent video frame of the reference video frame to obtain the abscissa of the reference coordinate corresponding to the first image region; calculate the sum of the ordinate of the first image region in the reference video frame and the actual vertical offset of the first image region relative to the adjacent video frame of the reference video frame to obtain the ordinate of the reference coordinate corresponding to the first image region.

[0162] For example, for each first image region in the reference video frame, the position of the first image region in the reference video frame can be expressed as: (1, 1), (1, 2), (2, 1), (2, 2); if the actual horizontal offset of the first image region relative to the adjacent video frame of the reference video frame is 1 and the actual vertical offset is 2, then calculate the sum of the abscissa of the coordinate of the first image region in the reference video frame and the actual horizontal offset of the first image region relative to the adjacent video frame of the reference video frame, and calculate the sum of the ordinate of the coordinate of the first image region in the reference video frame and the actual vertical offset of the first image region relative to the adjacent video frame of the reference video frame to obtain the reference coordinates (2, 3), (2, 4), (3, 3), (3, 4) corresponding to the first image region.

[0163] Subsequently, the electronic device determines the image region corresponding to the reference coordinate corresponding to the first image region in the adjacent video frame of the reference video frame to obtain the second image region corresponding to the first image region in the adjacent video frame of the reference video frame.

[0164] Based on the above processing, compared with the prior art registration method based on feature point matching, first extract the feature points in two video frames, such as extracting SIFT (Scale-Invariant Feature Transform) feature points, and then perform matching calculations on the feature points in the two video frames to obtain the corresponding feature points of the feature points in one video frame in the other video frame. According to the corresponding relationship between the feature points in the two video frames, obtain the transformation matrix, and then use the transformation matrix to transform one video frame to achieve the registration of the two video frames. Or, a registration method based on the block-matching technology. For each image region in one video frame, in the preset search region in the other video frame, by sliding the window, calculate the SAD (Sum of Absolute Difference) between the pixel values of each pixel point in the image region and the pixel values of each pixel point in each image region in the preset search region in the other video frame, and determine the image region with the smallest SAD in the other video frame as the corresponding image region of the image region in one video frame. Then, according to the corresponding relationship between the image regions in the two video frames, obtain the transformation matrix, and then use the transformation matrix to transform one video frame to achieve the registration of the two video frames.

[0165] In the embodiment of the present invention, the offset matrix and the weight matrix can be obtained through the pre-trained enhanced parameter determination model, which can improve the accuracy of the obtained offset matrix and weight matrix. And by restoring each element in the offset matrix, the restored element can represent the actual offset of the first image region relative to the adjacent video frame of the reference video frame. The actual offset can represent the actual moving distance of the first image region relative to the adjacent video frame of the reference video frame, that is, the elements in the offset matrix can adapt to the reference video frame and the adjacent video frame of the reference video frame. Then, according to the restored elements, determine the reference coordinates corresponding to the first image region, which can improve the accuracy of the determined reference coordinates corresponding to the first image region.

[0166] For step S104, for each first image region in the reference video frame, the content included in the first image region is the same as that included in the corresponding second image region in the adjacent video frame of the reference video frame. If the reference video frame is before the adjacent video frame of the reference video frame, then the first image region is the image region before the movement of the second image region; if the reference video frame is after the adjacent video frame of the reference video frame, then the first image region is the image region after the movement of the second image region.

[0167] After determining the second image area corresponding to the first image area in the adjacent video frame of the reference video frame, the electronic device may enhance the first image area according to the weight of the first image area and the weight of the second image area to obtain the target video frame corresponding to the reference video frame.

[0168] In some embodiments, based on Figure 3 , referring to Figure 4 , step S104 may include the following steps:

[0169] S1041: For each pixel point in the first image area of the reference video frame, if the reference coordinate corresponding to the pixel point in the adjacent video frame is an integer, calculate the weighted sum of the pixel value of the pixel point and the pixel value of the pixel point at the reference coordinate according to the weight of the first image area to which the pixel point belongs and the weight of the second image area to which the reference coordinate corresponding to the pixel point in the adjacent video frame belongs, so as to obtain the target video frame corresponding to the reference video frame.

[0170] S1042: If the reference coordinate corresponding to the pixel point in the adjacent video frame is not an integer, determine the pixel value corresponding to the reference coordinate based on the pixel values of the pixel points adjacent to the reference coordinate, and calculate the weighted sum of the pixel value of the pixel point and the pixel value corresponding to the reference coordinate according to the weight of the first image area to which the pixel point belongs and the weight of the second image area to which the reference coordinate corresponding to the pixel point in the adjacent video frame belongs, so as to obtain the target video frame corresponding to the reference video frame.

[0171] For each first image area in the reference video frame, the first image area includes a plurality of pixel points (which may be referred to as first pixel points), and the weights of the plurality of first pixel points are the same, and are all the weight of the first image area; for each second image area in the adjacent video frame of the reference video frame, the second image area includes a plurality of pixel points (which may be referred to as second pixel points), and the weights of the plurality of second pixel points are the same, and are all the weight of the second image area.

[0172] For each first pixel point in the first image region, after the electronic device calculates the corresponding reference coordinates of the first pixel point in the adjacent video frames of the reference video frame, if the reference coordinates corresponding to the first pixel point in the adjacent video frames of the reference video frame are integers, that is, the electronic device can determine the corresponding second pixel point of the first pixel point in the adjacent video frames of the reference video frame, the electronic device can directly calculate the weighted sum of the pixel value of the first pixel point and the pixel value of the second pixel point corresponding to the first pixel point according to the weight of the first image region to which the first pixel point belongs and the weight of the second image region to which the second pixel point corresponding to the first pixel point belongs, to obtain the enhanced pixel value of the first image region. The second image region to which the second pixel point corresponding to the first pixel point belongs, that is, the second image region corresponding to the first image region to which the first pixel point belongs in the adjacent video frames of the reference video frame. The electronic device calculates the enhanced pixel values of each first image region in the reference video frame, and can thus obtain the enhanced reference video frame, that is, the target video frame.

[0173] If the reference coordinates corresponding to the first pixel point in the adjacent video frames of the reference video frame are not integers, that is, the electronic device fails to determine the corresponding second pixel point of the first pixel point in the adjacent video frames of the reference video frame, the electronic device cannot obtain the pixel value of the second pixel point corresponding to the first pixel point, and thus cannot calculate the enhanced pixel value of the first pixel point based on the pixel value of the second pixel point corresponding to the first pixel point, and further cannot enhance the reference video frame. Correspondingly, the electronic device can determine the pixel value corresponding to the reference coordinates of the first pixel point based on the pixel values of the pixel points adjacent to the reference coordinates corresponding to the first pixel point in the adjacent video frames of the reference video frame.

[0174] In some embodiments, on the basis of Figure 4 referring to Figure 5 , step S1042 may include the following steps:

[0175] S10421: If the reference coordinates corresponding to the pixel point in the adjacent video frame are not integers, based on the nearest neighbor interpolation algorithm, determine the pixel value of the pixel point with the smallest distance from the reference coordinates as the pixel value corresponding to the reference coordinates; or, based on the bilinear interpolation algorithm, calculate the pixel value corresponding to the reference coordinates based on the pixel values of a preset number of pixel points within a preset neighborhood range of the reference coordinates, and calculate the weighted sum of the pixel value of the pixel point and the pixel value corresponding to the reference coordinates according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinates corresponding to the pixel point in the adjacent video frame belong, to obtain the target video frame corresponding to the reference video frame.

[0176] The pixel points adjacent to the reference coordinate corresponding to the first pixel point in the adjacent video frame of the reference video frame may be: the pixel points with the smallest distance from the reference coordinate corresponding to the first pixel point.

[0177] For example, if the reference coordinate corresponding to the first pixel point is (2.1, 2.2), then the integer coordinate with the smallest distance from the reference coordinate corresponding to the first pixel point is (2, 2). Then the electronic device may use the pixel value of the pixel point with the coordinate (2, 2) as the pixel value corresponding to the reference coordinate corresponding to the first pixel point.

[0178] The pixel points adjacent to the reference coordinate corresponding to the first pixel point in the adjacent video frame of the reference video frame may be: a preset number of pixel points within the preset neighborhood range of the reference coordinate corresponding to the first pixel point and having the closest distance from the reference coordinate corresponding to the first pixel point. For example, the preset number may be 4.

[0179] For example, if the reference coordinate corresponding to the first pixel point is (12.5, 13.5), and the preset neighborhood range of the reference coordinate corresponding to the first pixel point is: a circular area with the reference coordinate corresponding to the first pixel point as the center and a radius of 1. Then the pixel points adjacent to the reference coordinate corresponding to the first pixel point are: the pixel point with the coordinate (12, 13), the pixel point with the coordinate (12, 14), the pixel point with the coordinate (13, 13), and the pixel point with the coordinate (13, 14).

[0180] Furthermore, the electronic device uses the bilinear interpolation algorithm to perform bilinear interpolation on the pixel values of a preset number of pixel points adjacent to the reference coordinate corresponding to the first pixel point to obtain the pixel value corresponding to the reference coordinate corresponding to the first pixel point.

[0181] Subsequently, the electronic device may calculate the weighted sum of the pixel value of the first pixel point and the pixel value corresponding to the reference coordinate corresponding to the first pixel point according to the weight of the first image region to which the first pixel point belongs and the weight of the second image region to which the reference coordinate corresponding to the first pixel point belongs, to obtain the enhanced pixel value of the first image region. By calculating the enhanced pixel values of each first image region in the reference video frame, the electronic device may obtain the enhanced reference video frame, that is, the target video frame.

[0182] Based on the above processing, if the reference coordinates corresponding to the first pixel point are integers, that is, the electronic device can determine the second pixel point corresponding to the first pixel point in the adjacent video frame of the reference video frame, the electronic device directly calculates the enhanced pixel value of the first pixel point according to the pixel value of the second pixel point corresponding to the first pixel point; if the reference coordinates corresponding to the first pixel point are not integers, that is, the electronic device fails to determine the second pixel point corresponding to the first pixel point in the adjacent video frame of the reference video frame, the electronic device determines the pixel value corresponding to the reference coordinates of the first pixel point based on the nearest neighbor interpolation algorithm or the bilinear interpolation algorithm according to the pixel values of the pixel points adjacent to the reference coordinates corresponding to the first pixel point, and then calculates the enhanced pixel value of the first pixel point according to the pixel value corresponding to the reference coordinates of the first pixel point. The accuracy of the calculated enhanced pixel value can be improved, and further the accuracy of the target video frame corresponding to the obtained reference video frame can be improved.

[0183] Regarding step S105, after the electronic device obtains the target video frames corresponding to the specified reference video frames in the original video, the enhancement of each reference video frame in the original video is achieved. Subsequently, the electronic device can combine the target video frames corresponding to each reference video frame to obtain the target video corresponding to the original video.

[0184] Exemplarily, there may be a preset order among the specified reference video frames in the original video. After the electronic device obtains the target video frames corresponding to each reference video frame, it sorts the target video frames corresponding to each reference video frame according to the preset order among the reference video frames, and then splices the sorted target video frames corresponding to each reference video frame into a video to obtain the target video.

[0185] See Figure 6 , Figure 6 which is the first flowchart of the model training method provided by the embodiment of the present invention. This method is used to generate the enhanced parameter determination model in any of the foregoing embodiments. This method may include the following steps:

[0186] S601: Obtain a first sample video frame and a second sample video frame.

[0187] Among them, the first sample video frame and the second sample video frame are RAW images; the first sample video frame is obtained by adding noise to the second sample video frame.

[0188] S602: Input the first sample video frame and the adjacent video frames of the first sample video frame in the sample video to the enhanced parameter determination model with an initial structure, and obtain the offset matrix and weight matrix output by the enhanced parameter determination model with the initial structure.

[0189] Among them, the elements in the offset matrix represent the offset of the first image region in the first sample video frame relative to the adjacent video frame; the elements in the weight matrix correspond one by one to the first image region in the first sample video frame and the second image region in the adjacent video frame; if an element corresponds to the first image region in the first sample video frame, the element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, the element represents the weight of the corresponding second image region.

[0190] S603: For each first image region in the first sample video frame, based on the position of the first image region in the first sample video frame and the offset of the first image region relative to the adjacent video frame, determine the corresponding second image region of the first image region in the adjacent video frame.

[0191] S604: For each first image region in the first sample video frame, calculate the enhanced pixel value of the first image region according to the weight of the first image region and the weight of the corresponding second image region of the first image region in the adjacent video frame, and obtain the predicted video frame corresponding to the first sample video frame.

[0192] S605: Calculate the loss function value representing the difference between the predicted video frame and the second sample video frame.

[0193] S606: Based on the calculated loss function value, adjust the model parameters of the enhanced parameter determination model with the initial structure until the preset convergence condition is reached, and obtain the trained enhanced parameter determination model.

[0194] Based on the model training method provided by the embodiments of the present invention, an enhanced parameter determination model can be obtained. Subsequently, based on the pre-trained enhanced parameter determination model, an offset matrix and a weight matrix corresponding to the reference video frame input to the enhanced parameter determination model and the adjacent video frame of the reference video frame can be obtained. Subsequently, the reference video frame can be enhanced according to the offset matrix and the weight matrix. The reference video frame is a RAW image, that is, the RAW image can be enhanced; based on the pre-trained enhanced parameter determination model to obtain the offset matrix and the weight matrix, for each first image region in the reference video frame, based on the offset matrix, determine the corresponding second image region of the first image region in the adjacent video frame of the reference video frame, which can improve the accuracy of the determined second image region. Furthermore, based on the weights of the first image regions in the reference video frame and the weights of the corresponding second image regions of the first image regions, enhancing the reference video frame can improve the quality of the obtained target video.

[0195] Regarding step S601, the first sample video frame is obtained by adding noise to the second sample video frame. For example, under normal illumination, using the minimum ISO (sensitivity) of the camera, a technician holds the camera and shoots the video to be processed in a way that simulates the movement of the camera lens. Subsequently, the electronic device can determine the first sample video frame, the second sample video frame, and the sample video based on the video to be processed.

[0196] In one implementation, each video frame in the video to be processed is a second sample video frame. The electronic device directly adds noise to each video frame in the video to be processed to obtain the sample video. The video frame in the sample video corresponding to the second sample video frame in the video to be processed is the first sample video frame corresponding to the second sample video frame.

[0197] In another implementation, the electronic device performs a reduction in illumination on each video frame in the video to be processed to obtain the video frame (which can be referred to as the low-illumination video frame) of each video frame in the video to be processed in a low-illumination environment, thereby obtaining the video (which can be referred to as the low-illumination video) in the low-illumination environment corresponding to the video to be processed, and uses the low-illumination video frame as the second sample video frame.

[0198] Exemplarily, for each video frame in the video to be processed, the electronic device can calculate the quotient of the pixel value of each pixel point of the video frame and a preset ratio to obtain the low-illumination video frame corresponding to the video frame, and can obtain the low-illumination video corresponding to the video to be processed. The preset ratio can be 10, 20, 30, etc. The pixel value of each pixel point of the second sample video frame can be recorded as dark RAW domain data, or dark gt data.

[0199] Then, the electronic device can add noise to each video frame in the low-illumination video to obtain the sample video. The video frame in the sample video corresponding to the second sample video frame in the low-illumination video is the first sample video frame corresponding to the second sample video frame.

[0200] Regarding step S602, the adjacent video frames of the first sample video frame in the sample video to which it belongs include: the first number of video frames in the sample video that are before the first sample video frame, and / or the second number of video frames that are after the first sample video frame. The manner in which the electronic device determines the adjacent video frames of the first sample video frame in the sample video to which it belongs is similar to the manner in which the electronic device determines the adjacent video frames of the reference video frame in the foregoing embodiments, and the relevant introduction in the foregoing embodiments can be referred to.

[0201] The enhancement parameter determination model of the initial structure can be a CNN model. After the electronic device obtains the first sample video frame, it can also obtain the pixel matrix composed of the pixel values of each pixel point of the first sample video frame (which can be called the fourth pixel matrix). After the electronic device obtains the adjacent video frame of the first sample video frame, it can also obtain the pixel matrix of the pixel values of each pixel point of the adjacent video frame (which can be called the fifth pixel matrix). Then, the electronic device can splice the fourth pixel matrix and the fifth pixel matrix to obtain the spliced pixel matrix (which can be called the sixth pixel matrix).

[0202] Exemplarily, the pixel matrix composed of the pixel values of each pixel point of each video frame in the video to be processed can be the RAW domain data of the video frame. After the electronic device obtains the video to be processed, every 5 consecutive video frames in the video to be processed are used as a minimum processing unit. For each minimum processing unit in the video to be processed, the 5 consecutive video frames included in the minimum processing unit are respectively denoted as: alt_img0 video frame, alt_img1 video frame, ref_img video frame, alt_img2 video frame, alt_img3 video frame.

[0203] For each video frame in the minimum processing unit, perform a reduction in illuminance processing on the video frame, that is, divide the RAW domain data of the video frame by a preset magnification factor to obtain the low-light RAW domain data corresponding to the video frame, that is, obtain the low-illuminance video frame corresponding to the video frame, and further obtain the low-illuminance processing unit corresponding to the minimum processing unit. For example, perform a reduction in illuminance processing on the alt_img0 video frame to obtain the low-illuminance video frame corresponding to the alt_img0 video frame, denoted as the gt_alt_img0 video frame; perform a reduction in illuminance processing on the ref_img video frame to obtain the low-illuminance video frame corresponding to the ref_img video frame, denoted as the gt_ref_img video frame. The electronic device can select the low-illuminance video frame corresponding to the ref_img video frame, that is, the gt_ref_img video frame, as the second sample video frame.

[0204] Then, the electronic device can add a preset noise to each low-light video frame in the low-light processing unit to obtain the low-light video frames with added noise corresponding to the low-light video frames in the low-light processing unit, and then obtain a sample video. Each video frame in the sample video is the low-light video frame with added noise corresponding to each video frame in the video to be processed. For example, add a preset noise to the low-light video frame gt_alt_img0 corresponding to the alt_img0 video frame to obtain the corresponding low-light video frame with added noise, denoted as the noise_alt_img0 video frame; add a preset noise to the low-light video frame gt_ref_img corresponding to the ref_img video frame to obtain the corresponding low-light video frame with added noise, denoted as the noise_ref_img video frame. Further, the electronic device can select the low-light video frame with added noise corresponding to the ref_img video frame as the first sample video frame, that is, use the noise_ref_img video frame as the first sample video frame corresponding to the second sample video frame. Correspondingly, the adjacent video frames of the first sample video frame include: the noise_alt_img0 video frame, the noise_alt_img1 video frame, the noise_alt_img2 video frame, and the noise_alt_img3 video frame.

[0205] Then, the electronic device splices the low-light video frames with added noise corresponding to 5 consecutive video frames included in the minimum processing unit, and the splicing result can be denoted as {noise_alt_img0; noise_alt_img1; noise_ref_img; noise_alt_img2, noise_alt_img3}, which is used as the input data of the enhanced parameter determination model with the initial structure.

[0206] Subsequently, the electronic device can input the sixth pixel matrix into the enhanced parameter determination model with the initial structure to obtain the offset matrix and weight matrix output by the enhanced parameter determination model with the initial structure. The elements in the offset matrix represent the offset of the first image area in the first sample video frame relative to the adjacent video frames of the first sample video frame; the elements in the weight matrix are in one-to-one correspondence with the first image area in the first sample video frame and the second image area in the adjacent video frames of the first sample video frame; if an element corresponds to the first image area in the first sample video frame, the element represents the weight of the corresponding first image area; if an element corresponds to the second image area in the adjacent video frames of the first sample video frame, the element represents the weight of the corresponding second image area.

[0207] In some embodiments, the enhanced parameter determination model with the initial structure includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network.

[0208] Correspondingly, based on Figure 6 , referring to Figure 7 , step S602 includes the following steps:

[0209] S6021: Input the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model of the initial structure. Extract features from the first sample video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the first sample video frame and the image features of the adjacent video frame, which are used as the first feature vectors.

[0210] S6022: Extract features from the first feature vectors through the second feature extraction network to obtain second feature vectors.

[0211] S6023: Normalize the second feature vectors through the offset prediction network to obtain the corresponding offset matrix.

[0212] S6024: Extract features from the first feature vectors through the third feature extraction network to obtain third feature vectors.

[0213] S6025: Normalize the third feature vectors through the weight prediction network to obtain the corresponding weight matrix.

[0214] The first feature extraction network may include multiple convolutional layers, and the multiple convolutional layers include: a convolutional layer with a convolution kernel of 3×3 and a stride of 1, and a convolutional layer with a convolution kernel of 3×3 and a stride of 2. Moreover, an activation function is connected after each convolutional layer included in the first feature extraction network. For example, the activation function can be a ReLu function.

[0215] Extract features from the sixth pixel matrix through the multiple convolutional layers included in the first feature extraction network to downsample the sixth pixel matrix and obtain feature vectors representing the image features of the first sample video frame and the image features of the adjacent video frame of the first sample video frame (i.e., the first feature vectors).

[0216] The second feature extraction network may include two convolutional layers, and the two convolutional layers are convolutional layers with a convolution kernel of 3×3 and a stride of 1. Moreover, an activation function is connected after the first convolutional layer in the second feature extraction network. For example, the activation function can be a ReLu function; since there are two offsets of each first image region in the first sample video frame relative to each adjacent video frame of the first sample video frame, that is, each adjacent video frame of the first sample video frame corresponds to two output channels of the second feature extraction network. Therefore, the number of output channels of the second feature extraction network is twice the number of adjacent video frames of the first sample video frame.

[0217] Through two convolutional layers included in the second feature extraction network, feature extraction is performed on the first feature vector to obtain a feature vector (i.e., the second feature vector) representing the offset of the first sample video frame relative to the adjacent video frame of the first sample video frame. Subsequently, the offset prediction network performs normalization processing on the second feature vector to obtain a corresponding offset matrix. For example, the offset prediction network can use a normalization function to perform normalization processing on the second feature vector. For example, the normalization function can be the tanh function, which normalizes the second feature vector to the interval [-1, 1] to obtain a corresponding offset matrix.

[0218] The third feature extraction network may include two convolutional layers, which are convolutional layers with a convolutional kernel of 3×3 and a stride of 1, and an activation function is connected after the first convolutional layer in the third feature extraction network. For example, the activation function can be the ReLu function; since the first sample video frame corresponds to one output channel of the third feature extraction network, and each adjacent video frame of the first sample video frame also corresponds to one output channel of the third feature extraction network, therefore, the number of output channels of the third feature extraction network is the same as the sum of the number of the first sample video frames and the number of adjacent video frames of the first sample video frames.

[0219] Through two convolutional layers included in the third feature extraction network, feature extraction is performed on the first feature vector to obtain a feature vector (i.e., the third feature vector) representing the weights of the first sample video frame and the adjacent video frames of the first sample video frame. Subsequently, the weight prediction network performs normalization processing on the third feature vector to obtain a corresponding weight matrix. For example, the weight prediction network can use a normalization function to perform normalization processing on the third feature vector. For example, the normalization function can be softmax, which normalizes the third feature vector to the interval [0, 1] to obtain a corresponding weight matrix, and the sum of the elements included in the weight matrix is 1.

[0220] Based on the above processing, inputting the sixth pixel matrix into the enhanced parameter determination model with the initial structure can obtain the offset matrix and the weight matrix output by the enhanced parameter determination model. Subsequently, enhancing the first sample video frame according to the offset matrix and the weight matrix, and adjusting the parameters of the enhanced parameter determination model according to the enhancement result can improve the accuracy of the offset matrix and the weight matrix output by the enhanced parameter determination model. Downsampling the sixth pixel matrix can increase the receptive field of the first feature extraction network and reduce the computational complexity of subsequent processing, which can improve the efficiency of the offset matrix and the weight matrix output by the enhanced parameter determination model.

[0221] Regarding step S603, for each first image region in the first sample video frame, the position of the first image region in the first sample video frame can be represented by coordinates.

[0222] In some embodiments, since the offset matrix output by the enhancement parameter model of the initial structure is obtained through normalization, that is, each element in the offset matrix cannot directly represent how many pixel points the first image region is offset from the adjacent video frame of the first sample video frame. Therefore, the electronic device can first restore the elements in the offset matrix so that the restored elements can represent how many pixel points the first image region is offset from the adjacent video frame of the first sample video frame, that is, obtain the actual offset of the first image region relative to the adjacent video frame of the first sample video frame. The actual offset can represent the actual moving distance of the first image region relative to the adjacent video frame of the first sample video frame. The actual offset of the first image region relative to the adjacent video frame of the first sample video frame includes: the actual offset of the first image region relative to the adjacent video frame of the first sample video frame in the horizontal direction, and the actual offset of the first image region relative to the adjacent video frame of the first sample video frame in the vertical direction.

[0223] Then, the electronic device can calculate the sum of the abscissa of the first image region in the first sample video frame and the actual offset of the first image region relative to the adjacent video frame of the first sample video frame in the horizontal direction to obtain the abscissa of the reference coordinate corresponding to the first image region; calculate the sum of the ordinate of the first image region in the first sample video frame and the actual offset of the first image region relative to the adjacent video frame of the first sample video frame in the vertical direction to obtain the ordinate of the reference coordinate corresponding to the first image region. That is, the second image region corresponding to the first image region in the adjacent video frame of the first sample video frame can be determined.

[0224] Regarding step S604, the way the electronic device calculates the enhanced pixel value of each first image region in the first sample video frame to obtain the predicted video frame corresponding to the first sample video frame is similar to the way the electronic device calculates the enhanced pixel value of each first image region in the reference video frame to obtain the target video frame corresponding to the reference video frame in the foregoing embodiments, and the relevant introduction of the foregoing embodiments can be referred to.

[0225] Exemplarily, the electronic device can enhance the first sample video frame based on the following formula (1) based on the weights of the first image regions in the first sample video frame and the weights of the second image regions corresponding to each first image region in the adjacent video frame of the first sample video frame to obtain the predicted video frame corresponding to the first sample video frame.

[0226]

[0227] where out represents the predicted video frame, noise_ref m represents the m-th first image region in the first sample video frame, wts_refm Represents the weight of the m-th first image region, warped_noise_alt m [i] represents the second image region corresponding to the m-th first image region in the i-th adjacent video frame, wts_alt m [i] represents the weight of the second image region corresponding to the m-th first image region in the i-th adjacent video frame, and j represents that the first sample video frame has j + 1 adjacent video frames. For example, if the first sample video frame has 4 adjacent video frames, then j in the above formula (1) is 3.

[0228] For steps S605 and S606, after the electronic device obtains the predicted video frame corresponding to the first sample video frame, it can calculate the difference between the predicted video frame and the second sample video frame.

[0229] The loss function can be an absolute difference loss function, and the electronic device can calculate the absolute difference loss function value representing the difference between the predicted video frame and the second sample video frame. Exemplarily, the electronic device can calculate the absolute difference loss function value representing the difference between the predicted video frame and the second sample video frame based on the following formula (2):

[0230] L1_loss = |out - ref_img| (2)

[0231] Where, L1_loss represents the absolute difference loss function value, out represents the predicted video frame, and ref_img represents the second sample video frame.

[0232] Furthermore, the electronic device adjusts the enhancement parameters of the initial structure based on the calculated loss function value to determine the model parameters of the model. For example, the model parameters of the enhancement parameter determination model of the initial structure are adjusted in the manner of gradient descent until the preset convergence condition is reached, and the trained enhancement parameter determination model is obtained. The process of training the enhancement parameter determination model of the initial structure can also be called a supervised learning process.

[0233] See Figure 8 , Figure 8 which is a working principle diagram of the enhancement parameter determination model provided by the embodiment of the present invention.

[0234] Figure 8 The input in [] is the input data of the enhancement parameter determination model, and the input data is the first sample video frame and the adjacent video frames of the first sample video frame.

[0235] Figure 8The first type of convolutional layer is represented by a black rectangle, denoted as conv3×3_stride = 1 + relu. The convolutional kernel of the first type of convolutional layer is 3×3, the stride is 1, and it is connected to an activation function. The second type of convolutional layer is represented by a white rectangle, denoted as conv3×3_stride = 2 + relu. The convolutional kernel of the second type of convolutional layer is 3×3, the stride is 2, and it is connected to an activation function.

[0236] The first feature extraction network includes nine convolutional layers. The first three convolutional layers, the fifth convolutional layer, the sixth convolutional layer, the eighth convolutional layer, and the ninth convolutional layer among these nine convolutional layers are all of the first type; the fourth convolutional layer and the seventh convolutional layer among these nine convolutional layers are of the second type. Through the nine convolutional layers included in the first feature extraction network, feature extraction is performed on the input data to achieve downsampling of the input data, and a first feature vector representing the image features of the first sample video frame and the image features of the adjacent video frames of the first sample video frame is obtained.

[0237] Figure 8 The third type of convolutional layer is represented by a black rectangle filled with white horizontal stripes, denoted as conv3×3_stride = 1, c = 8. The convolutional kernel of the third type of convolutional layer is 3×3, and the stride is 1.

[0238] The second feature extraction network includes two convolutional layers. The first convolutional layer among these two convolutional layers is of the first type, and the second convolutional layer among these two convolutional layers is of the third type. Through the two convolutional layers included in the second feature extraction network, feature extraction is performed on the first feature vector to obtain a second feature vector representing the offset of the first sample video frame relative to the adjacent video frames of the first sample video frame.

[0239] Figure 8 The offset prediction network is represented by a black rectangle filled with white diagonal stripes and can be denoted as tanh. The offset prediction network can use a normalization function to normalize the second feature vector. For example, the normalization function can be the tanh function to normalize the second feature vector to the interval [-1, 1] to obtain a corresponding offset matrix. The output of the offset prediction network can be denoted as out_flow, c = 8 (output_offset, the number of output channels is 8). The output of the offset prediction network can also be referred to as motion vectors.

[0240] Figure 8 The fourth type of convolutional layer is represented by a white rectangle filled with black horizontal stripes, denoted as conv3×3_stride = 1, c = 5. The convolutional kernel of the fourth type of convolutional layer is 3×3, and the stride is 1.

[0241] The third feature extraction network includes two convolutional layers. The first convolutional layer of the two convolutional layers is a first type of convolutional layer, and the second convolutional layer of the two convolutional layers is a fourth type of convolutional layer. Through the two convolutional layers included in the third feature extraction network, feature extraction is performed on the first feature vector to obtain a third feature vector representing the weights of the first sample video frame and the adjacent video frames of the first sample video frame.

[0242] Figure 8 In the figure, a rectangle filled with black diagonal stripes in white represents the weight prediction network, and the weight prediction network can be denoted as softmax. The weight prediction network can use a normalization function to normalize the third feature vector. For example, the normalization function can be softmax to normalize the third feature vector to the interval [0, 1] to obtain a corresponding weight matrix, and the sum of the elements included in the weight matrix is 1. The output of the weight prediction network can be denoted as out_wts, and c = 5 (output_weight, the number of output channels is 5). The output of the weight prediction network can also be referred to as the fusion weight.

[0243] Furthermore, the electronic device can, based on formula (1) in the foregoing embodiment, enhance the first sample video frame based on the weights of the first image regions in the first sample video frame and the weights of the corresponding second image regions of each first image region in the adjacent video frames of the first sample video frame to obtain a predicted video frame corresponding to the first sample video frame. Then, the electronic device can, based on formula (2) in the foregoing embodiment, calculate the absolute difference loss function value representing the difference between the predicted video frame and the second sample video frame, and based on the calculated absolute difference loss function value, adjust the model parameters of the enhancement parameter determination model with the initial structure until a preset convergence condition is reached to obtain a trained enhancement parameter determination model. Subsequently, the electronic device can input 5 consecutive video frames in the sample video into the trained enhancement parameter determination model, and the enhancement parameter determination model can perform a forward inference process to obtain an enhanced video image output.

[0244] Based on the above processing, an enhancement parameter determination model can be obtained. Subsequently, based on the pre-trained enhancement parameter determination model, an offset matrix and a weight matrix corresponding to the reference video frame input to the enhancement parameter determination model and the adjacent video frames of the reference video frame can be obtained. Subsequently, the reference video frame can be enhanced according to the offset matrix and the weight matrix. The reference video frame is a RAW image, that is, the enhancement of the RAW image can be realized; based on the pre-trained enhancement parameter determination model to obtain the offset matrix and the weight matrix, for each first image region in the reference video frame, the second image region corresponding to the first image region in the adjacent video frame of the reference video frame can be determined based on the offset matrix, which can improve the accuracy of the determined second image region. Furthermore, based on the weights of the first image regions in the reference video frame and the weights of the second image regions corresponding to the first image regions, the reference video frame can be enhanced, which can improve the quality of the obtained target video.

[0245] Moreover, the enhancement parameter determination model can output relatively accurate offsets and weights, simplifying the subsequent processing steps, enabling real-time registration and enhancement of RAW images, and achieving good time-domain enhancement effects on the original video.

[0246] In addition, directly outputting the offsets and weights between RAW images based on the enhancement parameter determination model for time-domain enhancement has good effects, and the subsequent processing is simple, and real-time operation can be achieved on a general NPU (Neural-network Process Units, embedded neural network processor) platform.

[0247] See Figure 9 , Figure 9 which is a comparative example diagram of the video enhancement effect provided by the embodiment of the present invention. Figure 9 The left image in Figure 9 is the reference video frame, and there are multiple noise points in the reference video frame, that is, the clarity of the reference video frame is low; Figure 9 The middle image in

[0248] Based on the same inventive concept as the above video enhancement method, an embodiment of the present invention further provides a video enhancement device. Refer to Figure 10 , Figure 10 which is a structural diagram of the video enhancement device provided by an embodiment of the present invention. The device includes:

[0249] A reference video frame determination module 1001, configured to determine a specified reference video frame from each video frame of the original video; wherein, each video frame in the original video is a RAW image;

[0250] A matrix acquisition module 1002, configured to input the reference video frame and an adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model, and obtain an offset matrix and a weight matrix output by the enhancement parameter determination model;

[0251] Wherein, an element in the offset matrix represents an offset amount of a first image area in the reference video frame relative to the adjacent video frame; an element in the weight matrix corresponds one-to-one with the first image area in the reference video frame and a second image area in the adjacent video frame; if an element corresponds to the first image area in the reference video frame, the element represents the weight of the corresponding first image area; if an element corresponds to the second image area in the adjacent video frame, the element represents the weight of the corresponding second image area;

[0252] An image area determination module 1003, configured to, for each first image area in the reference video frame, determine a second image area corresponding to the first image area in the adjacent video frame based on the position of the first image area in the reference video frame and the offset amount of the first image area relative to the adjacent video frame;

[0253] A target video frame acquisition module 1004, configured to, for each first image area in the reference video frame, calculate an enhanced pixel value of the first image area according to the weight of the first image area and the weight of the second image area corresponding to the first image area in the adjacent video frame, and obtain a target video frame corresponding to the reference video frame;

[0254] A target video acquisition module 1005, configured to obtain a target video corresponding to the original video based on the target video frame.

[0255] Optionally, the enhancement parameter determination model includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network;

[0256] The matrix acquisition module 1002 is specifically configured to:

[0257] Input the reference video frame and the adjacent video frames of the reference video frame into a pre-trained enhancement parameter determination model, and perform feature extraction on the reference video frame and the adjacent video frames through the first feature extraction network to obtain feature vectors representing the image features of the reference video frame and the image features of the adjacent video frames, which are used as the first feature vectors;

[0258] Perform feature extraction on the first feature vectors through the second feature extraction network to obtain second feature vectors;

[0259] Perform normalization processing on the second feature vectors through the offset prediction network to obtain corresponding offset matrices;

[0260] Perform feature extraction on the first feature vectors through the third feature extraction network to obtain third feature vectors;

[0261] Perform normalization processing on the third feature vectors through the weight prediction network to obtain corresponding weight matrices.

[0262] Optionally, the image region determination module 1003 is specifically configured to:

[0263] For each first image region in the reference video frame, obtain the coordinates of the first image region in the reference video frame;

[0264] Calculate the sum of the coordinates of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame as the reference coordinates;

[0265] In the adjacent video frame, determine the image region corresponding to the reference coordinates to obtain the second image region corresponding to the first image region in the adjacent video frame.

[0266] Optionally, the target video frame acquisition module 1004 is specifically configured to:

[0267] For each pixel point in the first image region of the reference video frame, if the reference coordinates corresponding to the pixel point in the adjacent video frame are integers, calculate the weighted sum of the pixel value of the pixel point and the pixel value of the pixel point at the reference coordinates according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinates corresponding to the pixel point in the adjacent video frame belong, to obtain the target video frame corresponding to the reference video frame;

[0268] If the reference coordinates corresponding to the pixel point in the adjacent video frame are not integers, determine the pixel value corresponding to the reference coordinates based on the pixel values of the pixel points adjacent to the reference coordinates, and calculate the weighted sum of the pixel value of the pixel point and the pixel value corresponding to the reference coordinates according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinates corresponding to the pixel point in the adjacent video frame belong, so as to obtain the target video frame corresponding to the reference video frame.

[0269] Optionally, the target video frame acquisition module 1004 is specifically configured to:

[0270] Based on the nearest neighbor interpolation algorithm, determine the pixel value of the pixel point with the smallest distance from the reference coordinates as the pixel value corresponding to the reference coordinates;

[0271] Or,

[0272] Based on the bilinear interpolation algorithm, calculate the pixel value corresponding to the reference coordinates based on the pixel values of a preset number of pixel points within a preset neighborhood range of the reference coordinates.

[0273] Optionally, the adjacent video frames of the reference video frame include: the video frames in the original video that are before the reference video frame, and / or, the video frames that are after the reference video frame.

[0274] Based on the video enhancement device provided in the embodiments of the present invention, the original video including each video frame as a RAW image can be enhanced based on the offset matrix and the weight matrix to obtain the target video corresponding to the original video, that is, the RAW image can be enhanced; and, since the offset matrix and the weight matrix are determined based on the pre-trained enhancement parameter determination model, for each first image region in the reference video frame, the second image region corresponding to the first image region in the adjacent video frame of the reference video frame can be determined based on the offset matrix, which can improve the accuracy of the determined second image region. Furthermore, based on the weights of the first image regions in the reference video frame and the weights of the second image regions corresponding to the first image regions, the reference video frame can be enhanced, which can improve the quality of the obtained target video.

[0275] Based on the same inventive concept as the above model training method, the embodiments of the present invention further provide a model training device for generating the enhancement parameter determination model in any one of the above embodiments. See Figure 11 , Figure 11 is a structural diagram of the model training device provided in the embodiments of the present invention. The device includes:

[0276] A sample video frame acquisition module 1101 is configured to acquire a first sample video frame and a second sample video frame; wherein, the first sample video frame and the second sample video frame are RAW images; the first sample video frame is obtained by adding noise to the second sample video frame;

[0277] A matrix acquisition module 1102 is configured to input the first sample video frame and an adjacent video frame of the first sample video frame in the sample video into an enhancement parameter determination model with an initial structure, and obtain an offset matrix and a weight matrix output by the enhancement parameter determination model with the initial structure;

[0278] Wherein, an element in the offset matrix represents an offset amount of a first image region in the first sample video frame relative to the adjacent video frame; elements in the weight matrix are in one-to-one correspondence with the first image region in the first sample video frame and a second image region in the adjacent video frame; if an element corresponds to the first image region in the first sample video frame, the element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, the element represents the weight of the corresponding second image region;

[0279] An image region determination module 1103 is configured to, for each first image region in the first sample video frame, determine a second image region corresponding to the first image region in the adjacent video frame based on the position of the first image region in the first sample video frame and the offset amount of the first image region relative to the adjacent video frame;

[0280] A predicted video frame acquisition module 1104 is configured to, for each first image region in the first sample video frame, calculate an enhanced pixel value of the first image region according to the weight of the first image region and the weight of the second image region corresponding to the first image region in the adjacent video frame, and obtain a predicted video frame corresponding to the first sample video frame;

[0281] A difference calculation module 1105 is configured to calculate a loss function value representing the difference between the predicted video frame and the second sample video frame;

[0282] An adjustment module 1106 is configured to adjust model parameters of the enhancement parameter determination model with the initial structure based on the calculated loss function value until a preset convergence condition is reached, and obtain a trained enhancement parameter determination model.

[0283] Optionally, the enhancement parameter determination model with the initial structure includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network;

[0284] The matrix acquisition module 1102 is specifically configured to:

[0285] Input the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model with an initial structure, and extract features from the first sample video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the first sample video frame and the image features of the adjacent video frame, as the first feature vectors;

[0286] Extract features from the first feature vectors through the second feature extraction network to obtain second feature vectors;

[0287] Perform normalization processing on the second feature vectors through the offset prediction network to obtain corresponding offset matrices;

[0288] Extract features from the first feature vectors through the third feature extraction network to obtain third feature vectors;

[0289] Perform normalization processing on the third feature vectors through the weight prediction network to obtain corresponding weight matrices.

[0290] Based on the model training device provided in the embodiments of the present invention, an enhancement parameter determination model can be obtained. Subsequently, based on the pre-trained enhancement parameter determination model, the offset matrix and weight matrix corresponding to the reference video frame input to the enhancement parameter determination model and the adjacent video frame of the reference video frame can be obtained. Subsequently, the reference video frame can be enhanced according to the offset matrix and weight matrix. The reference video frame is a RAW image, that is, the enhancement of the RAW image can be realized; the offset matrix and weight matrix are obtained based on the pre-trained enhancement parameter determination model, then for each first image region in the reference video frame, the second image region corresponding to the first image region in the adjacent video frame of the reference video frame can be determined based on the offset matrix, which can improve the accuracy of the determined second image region. Furthermore, based on the weights of the first image regions in the reference video frame and the weights of the second image regions corresponding to the first image regions, the reference video frame can be enhanced, which can improve the quality of the obtained target video.

[0291] The embodiments of the present invention also provide an electronic device, as Figure 12 shown, including a processor 1201, a communication interface 1202, a memory 1203, and a communication bus 1204, wherein the processor 1201, the communication interface 1202, and the memory 1203 communicate with each other through the communication bus 1204,

[0292] The memory 1203 is used to store a computer program;

[0293] The processor 1201 is configured to implement the steps of any of the video enhancement methods in the above embodiments when executing the program stored in the memory 1203.

[0294] An embodiment of the present invention further provides an electronic device, such as Figure 13 as shown, including a processor 1301, a communication interface 1302, a memory 1303, and a communication bus 1304. Among them, the processor 1301, the communication interface 1302, and the memory 1303 communicate with each other through the communication bus 1304.

[0295] The memory 1303 is used to store a computer program.

[0296] The processor 1301 is configured to implement the steps of any of the model training methods in the above embodiments when executing the program stored in the memory 1303.

[0297] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0298] The communication interface is used for communication between the above electronic device and other devices.

[0299] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0300] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0301] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of any of the above video enhancement methods are implemented.

[0302] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of any of the above model training methods are implemented.

[0303] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, and when it runs on a computer, the computer is made to execute any of the video enhancement methods in the above embodiments.

[0304] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, and when it runs on a computer, the computer is made to execute any of the model training methods in the above embodiments.

[0305] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media integrated. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0306] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0307] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, electronic device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0308] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A video enhancement method, characterized in that, the method includes: determining a specified reference video frame from each video frame of the original video; wherein, each video frame in the original video is a RAW image; inputting the reference video frame and an adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model to obtain an offset matrix and a weight matrix output by the enhancement parameter determination model; wherein, an element in the offset matrix represents an offset amount of a first image area in the reference video frame relative to the adjacent video frame; elements in the weight matrix are in one-to-one correspondence with the first image area in the reference video frame and a second image area in the adjacent video frame; if an element corresponds to the first image area in the reference video frame, the element represents the weight of the corresponding first image area; if an element corresponds to the second image area in the adjacent video frame, the element represents the weight of the corresponding second image area; for each first image area in the reference video frame, determining a corresponding second image area of the first image area in the adjacent video frame based on the position of the first image area in the reference video frame and the offset amount of the first image area relative to the adjacent video frame; for each first image area in the reference video frame, calculating an enhanced pixel value of the first image area according to the weight of the first image area and the weight of the corresponding second image area of the first image area in the adjacent video frame to obtain a target video frame corresponding to the reference video frame; obtaining a target video corresponding to the original video based on the target video frame.

2. The method according to claim 1, characterized in that, the enhancement parameter determination model includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network; the inputting the reference video frame and an adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model to obtain an offset matrix and a weight matrix output by the enhancement parameter determination model includes: inputting the reference video frame and an adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model, and performing feature extraction on the reference video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the reference video frame and the image features of the adjacent video frame as a first feature vector; performing feature extraction on the first feature vector through the second feature extraction network to obtain a second feature vector; performing normalization processing on the second feature vector through the offset prediction network to obtain a corresponding offset matrix; performing feature extraction on the first feature vector through the third feature extraction network to obtain a third feature vector; performing normalization processing on the third feature vector through the weight prediction network to obtain a corresponding weight matrix.

3. The method according to claim 1, characterized in that, For each first image region in the reference video frame, determining a corresponding second image region of the first image region in the adjacent video frame based on the position of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame includes: For each first image region in the reference video frame, obtaining the coordinates of the first image region in the reference video frame; Calculating the sum of the coordinates of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame as a reference coordinate; In the adjacent video frame, determining the image region corresponding to the reference coordinate to obtain the corresponding second image region of the first image region in the adjacent video frame.

4. The method according to claim 3, wherein, For each first image region in the reference video frame, calculating an enhanced pixel value of the first image region based on the weight of the first image region and the weight of the corresponding second image region of the first image region in the adjacent video frame to obtain a target video frame corresponding to the reference video frame includes: For each pixel point in the first image region of the reference video frame, if the reference coordinate corresponding to the pixel point in the adjacent video frame is an integer, calculating the weighted sum of the pixel value of the pixel point and the pixel value of the pixel point at the reference coordinate according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinate corresponding to the pixel point in the adjacent video frame belongs, to obtain a target video frame corresponding to the reference video frame; If the reference coordinate corresponding to the pixel point in the adjacent video frame is not an integer, determining the pixel value corresponding to the reference coordinate based on the pixel values of the pixel points adjacent to the reference coordinate, and calculating the weighted sum of the pixel value of the pixel point and the pixel value corresponding to the reference coordinate according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinate corresponding to the pixel point in the adjacent video frame belongs, to obtain a target video frame corresponding to the reference video frame.

5. The method according to claim 4, wherein, Determining the pixel value corresponding to the reference coordinate based on the pixel values of the pixel points adjacent to the reference coordinate includes: Based on the nearest neighbor interpolation algorithm, determining the pixel value of the pixel point with the smallest distance to the reference coordinate as the pixel value corresponding to the reference coordinate; Or, Based on the bilinear interpolation algorithm, calculating the pixel value corresponding to the reference coordinate based on the pixel values of a preset number of pixel points within a preset neighborhood range of the reference coordinate.

6. The method according to claim 1, wherein, The adjacent video frames of the reference video frame include: video frames in the original video that are before the reference video frame, and / or, video frames that are after the reference video frame.

7. A model training method, wherein, For generating the enhanced parameter determination model according to any one of claims 1-6, the method includes: Obtain a first sample video frame and a second sample video frame; wherein, the first sample video frame and the second sample video frame are RAW images; the first sample video frame is obtained by adding noise to the second sample video frame; Input the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model with the initial structure, and obtain the offset matrix and weight matrix output by the enhancement parameter determination model with the initial structure; Wherein, the elements in the offset matrix represent the offset amount of the first image region in the first sample video frame relative to the adjacent video frame; the elements in the weight matrix correspond one-to-one with the first image region in the first sample video frame and the second image region in the adjacent video frame; if an element corresponds to the first image region in the first sample video frame, this element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, this element represents the weight of the corresponding second image region; For each first image region in the first sample video frame, based on the position of this first image region in the first sample video frame and the offset amount of this first image region relative to the adjacent video frame, determine the corresponding second image region of this first image region in the adjacent video frame; For each first image region in the first sample video frame, calculate the enhanced pixel value of this first image region according to the weight of this first image region and the weight of the corresponding second image region of this first image region in the adjacent video frame, and obtain the predicted video frame corresponding to the first sample video frame; Calculate the loss function value representing the difference between the predicted video frame and the second sample video frame; Based on the calculated loss function value, adjust the model parameters of the enhancement parameter determination model with the initial structure until the preset convergence condition is reached, and obtain the trained enhancement parameter determination model.

8. The method according to claim 7, characterized in that, The enhancement parameter determination model with the initial structure includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network and a weight prediction network; The step of inputting the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model with the initial structure, and obtaining the offset matrix and weight matrix output by the enhancement parameter determination model with the initial structure includes: Input the first sample video frame and the adjacent video frame of the first sample video frame in the sample video into the enhancement parameter determination model with the initial structure, and perform feature extraction on the first sample video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the first sample video frame and the image features of the adjacent video frame, as the first feature vector; Perform feature extraction on the first feature vector through the second feature extraction network to obtain a second feature vector; Normalize the second feature vector through the offset prediction network to obtain a corresponding offset matrix; Extract features from the first feature vector through the third feature extraction network to obtain a third feature vector; Normalize the third feature vector through the weight prediction network to obtain a corresponding weight matrix.

9. A video enhancement device, characterized in that, the device includes: A reference video frame determination module, configured to determine a specified reference video frame from each video frame of the original video; wherein, each video frame in the original video is a RAW image; A matrix acquisition module, configured to input the reference video frame and an adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model, and obtain an offset matrix and a weight matrix output by the enhancement parameter determination model; wherein, an element in the offset matrix represents an offset amount of a first image region in the reference video frame relative to the adjacent video frame; an element in the weight matrix corresponds one-to-one to the first image region in the reference video frame and a second image region in the adjacent video frame; if an element corresponds to the first image region in the reference video frame, the element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, the element represents the weight of the corresponding second image region; An image region determination module, configured to, for each first image region in the reference video frame, determine a second image region corresponding to the first image region in the adjacent video frame based on the position of the first image region in the reference video frame and the offset amount of the first image region relative to the adjacent video frame; A target video frame acquisition module, configured to, for each first image region in the reference video frame, calculate an enhanced pixel value of the first image region based on the weight of the first image region and the weight of the second image region corresponding to the first image region in the adjacent video frame, and obtain a target video frame corresponding to the reference video frame; A target video acquisition module, configured to obtain a target video corresponding to the original video based on the target video frame.

10. The device according to claim 9, characterized in that, the enhancement parameter determination model includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network; The matrix acquisition module is specifically configured to: Input the reference video frame and an adjacent video frame of the reference video frame into a pre-trained enhancement parameter determination model, and extract features from the reference video frame and the adjacent video frame through the first feature extraction network to obtain a feature vector representing the image feature of the reference video frame and the image feature of the adjacent video frame, as the first feature vector; Extract features from the first feature vector through the second feature extraction network to obtain a second feature vector; Normalize the second feature vector through the offset prediction network to obtain a corresponding offset matrix; Feature extraction is performed on the first feature vector through the third feature extraction network to obtain a third feature vector; Normalization processing is performed on the third feature vector through the weight prediction network to obtain a corresponding weight matrix.

11. The apparatus according to claim 9, wherein, the image region determination module is specifically configured to: For each first image region in the reference video frame, obtain the coordinates of the first image region in the reference video frame; Calculate the sum value of the coordinates of the first image region in the reference video frame and the offset of the first image region relative to the adjacent video frame as the reference coordinate; In the adjacent video frame, determine the image region corresponding to the reference coordinate to obtain the second image region corresponding to the first image region in the adjacent video frame.

12. The apparatus according to claim 11, wherein, the target video frame acquisition module is specifically configured to: For each pixel point in the first image region of the reference video frame, if the reference coordinate corresponding to the pixel point in the adjacent video frame is an integer, calculate the weighted sum of the pixel value of the pixel point and the pixel value of the pixel point at the reference coordinate according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinate corresponding to the pixel point in the adjacent video frame belongs, to obtain the target video frame corresponding to the reference video frame; If the reference coordinate corresponding to the pixel point in the adjacent video frame is not an integer, determine the pixel value corresponding to the reference coordinate based on the pixel values of the pixel points adjacent to the reference coordinate, and calculate the weighted sum of the pixel value of the pixel point and the pixel value corresponding to the reference coordinate according to the weight of the first image region to which the pixel point belongs and the weight of the second image region to which the reference coordinate corresponding to the pixel point in the adjacent video frame belongs, to obtain the target video frame corresponding to the reference video frame.

13. The apparatus according to claim 12, wherein, the target video frame acquisition module is specifically configured to: Based on the nearest neighbor interpolation algorithm, determine the pixel value of the pixel point with the smallest distance from the reference coordinate as the pixel value corresponding to the reference coordinate; Or, Based on the bilinear interpolation algorithm, calculate the pixel value corresponding to the reference coordinate based on the pixel values of a preset number of pixel points within a preset neighborhood range of the reference coordinate.

14. The apparatus according to claim 9, wherein, The adjacent video frames of the reference video frame include: video frames in the original video that are before the reference video frame, and / or, video frames that are after the reference video frame.

15. A model training apparatus, wherein, for generating the enhanced parameter determination model according to any one of claims 1-6, the apparatus includes: A sample video frame acquisition module for acquiring a first sample video frame and a second sample video frame; wherein, the first sample video frame and the second sample video frame are RAW images; the first sample video frame is obtained by adding noise to the second sample video frame; A matrix acquisition module, configured to input the first sample video frame and an adjacent video frame of the first sample video frame in the sample video into an enhancement parameter determination model with an initial structure, and obtain an offset matrix and a weight matrix output by the enhancement parameter determination model with the initial structure; Wherein, an element in the offset matrix represents an offset amount of a first image region in the first sample video frame relative to the adjacent video frame; elements in the weight matrix correspond one by one to the first image region in the first sample video frame and a second image region in the adjacent video frame; if an element corresponds to the first image region in the first sample video frame, the element represents the weight of the corresponding first image region; if an element corresponds to the second image region in the adjacent video frame, the element represents the weight of the corresponding second image region; An image region determination module, configured to, for each first image region in the first sample video frame, determine a second image region corresponding to the first image region in the adjacent video frame based on the position of the first image region in the first sample video frame and the offset amount of the first image region relative to the adjacent video frame; A predicted video frame acquisition module, configured to, for each first image region in the first sample video frame, calculate an enhanced pixel value of the first image region according to the weight of the first image region and the weight of the second image region corresponding to the first image region in the adjacent video frame, and obtain a predicted video frame corresponding to the first sample video frame; A difference calculation module, configured to calculate a loss function value representing the difference between the predicted video frame and the second sample video frame; An adjustment module, configured to adjust model parameters of the enhancement parameter determination model with the initial structure based on the calculated loss function value until a preset convergence condition is reached, and obtain a trained enhancement parameter determination model.

16. The apparatus according to claim 15, wherein, the enhancement parameter determination model with the initial structure includes: a first feature extraction network, a second feature extraction network, a third feature extraction network, an offset prediction network, and a weight prediction network; The matrix acquisition module is specifically configured to: input the first sample video frame and an adjacent video frame of the first sample video frame in the sample video into an enhancement parameter determination model with an initial structure, and perform feature extraction on the first sample video frame and the adjacent video frame through the first feature extraction network to obtain feature vectors representing the image features of the first sample video frame and the image features of the adjacent video frame, as a first feature vector; perform feature extraction on the first feature vector through the second feature extraction network to obtain a second feature vector; perform normalization processing on the second feature vector through the offset prediction network to obtain a corresponding offset matrix; perform feature extraction on the first feature vector through the third feature extraction network to obtain a third feature vector; perform normalization processing on the third feature vector through the weight prediction network to obtain a corresponding weight matrix.

17. An electronic device, characterized in that, it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used for storing a computer program; the processor is used for implementing the method steps described in any one of claims 1-6 when executing the program stored on the memory.

18. An electronic device, characterized in that, it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used for storing a computer program; the processor is used for implementing the method steps described in any one of claims 7-8 when executing the program stored on the memory.

19. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and the computer program implements the method steps described in any one of claims 1-6 when executed by a processor.

20. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and the computer program implements the method steps described in any one of claims 7-8 when executed by a processor.

Citation Information

Patent Citations

  • Method for improving video compression coding efficiency based on deep learning

    CN107820085A

  • Video frame reconstruction method

    CN111147804A