Video Processing Method, Device, Computer Equipment and Storage Medium
The Mobile-Former structure enhances video resolution by fusing spatial, temporal, and scale features, addressing the limitations of existing methods in local processing to improve video quality in mobile devices.
Patent Information
- Application Number
- CN202210929241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-08-03
AI Technical Summary
The existing video super-resolution technology has poor performance in local processing, making it difficult to effectively improve the resolution of mobile phone videos.
Using a combination of spatial model, temporal model and codec model, the feature parameters of the video frame are obtained, spatial features, temporal features and scale features are fusion, and lightweight feature extraction is used to reconstruct super-resolution video.
It improves the resolution and quality of video, reduces the amount of computing, and is suitable for real-time processing of mobile devices such as mobile phones.
Smart Images

Figure CN115147284B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a video processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of intelligent terminal technology, mobile phone videos have become an important means of communication and entertainment for people. Due to the limitations of hardware conditions and data compression during communication, the details of video images are prone to be missing, resulting in a decrease in resolution. Therefore, video super-resolution technology has emerged, which can effectively improve video resolution, increase video details, and improve video quality.
[0003] In the prior art, Vision Transformer (a vision model based on the Transformer architecture) is usually used for video super-resolution processing. However, due to the use of a large number of network layers and the extensive use of attention mechanisms, although Vision Transformer can establish a perfect global dependence relationship, its performance in local processing is poor, and it is difficult to effectively improve video resolution.
[0004] Therefore, the current mobile phone video processing technology has the problem of limited improvement in video resolution. Summary of the Invention
[0005] Based on this, it is necessary to provide a video processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can effectively improve resolution for the above technical problems.
[0006] In a first aspect, this application provides a video processing method. The method includes:
[0007] Obtain the feature parameters of video frames in the video to be processed;
[0008] Input the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frames;
[0009] Input the spatial features into a pre-trained temporal model, and fuse the spatial features and the temporal features of the video frames through the temporal model to obtain the first fused features of the video frames; the temporal features are obtained through the temporal model;
[0010] Input the first fused features into a pre-trained encoding and decoding model, and fuse the first fused features and the scale features of the video frames through the encoding and decoding model to obtain the second fused features of the video frames; the scale features are obtained through the encoding and decoding model;
[0011] Based on the second fusion feature, a super-resolution video corresponding to the video to be processed is obtained.
[0012] In one embodiment, the feature parameters include image features and tags; the obtaining of the feature parameters of the video frames in the video to be processed includes:
[0013] Obtain the original video to be processed;
[0014] Perform data cleaning on the original video to be processed to obtain a cleaned video;
[0015] Group the cleaned video to obtain the video to be processed;
[0016] Perform feature mapping processing on each of the video frames in the video to be processed to obtain the image features of each of the video frames, and perform embedding processing on each of the video frames in the video to be processed to obtain the tags of each of the video frames.
[0017] In one embodiment, the video to be processed includes at least one group of video frames, and each of the video frames in each group of video frames corresponds to a spatial model; the inputting of the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frames includes:
[0018] Input the feature parameters of each of the video frames in each group of video frames into the spatial model corresponding to the video frame respectively to obtain the spatial features of each of the video frames.
[0019] In one embodiment, the encoding and decoding model includes two downsampling sub-models, a scale-invariant sub-model, and two upsampling sub-models; the inputting of the first fusion feature into a pre-trained encoding and decoding model to fuse the first fusion feature and the scale feature of the video frame through the encoding and decoding model to obtain the second fusion feature of the video frame includes:
[0020] Input the first fusion feature into the two downsampling sub-models, a scale-invariant sub-model, and two upsampling sub-models in sequence to obtain the second fusion feature of the video frame.
[0021] In one embodiment, the obtaining of the super-resolution video corresponding to the video to be processed according to the second fusion feature includes:
[0022] Fuse the second fusion feature of the video frame with the image feature of the video frame to obtain the third fusion feature of the video frame;
[0023] Perform deconvolution layer reconstruction processing on the third fusion feature to obtain a reconstructed video frame;
[0024] Superimpose the reconstructed video frame and the video frame of the video to be processed to obtain the super-resolution video.
[0025] In one embodiment, the superimposing the reconstructed video frame and the video frame of the video to be processed to obtain the super-resolution video includes:
[0026] Superimpose the reconstructed video frame and the video frame of the video to be processed to obtain a superimposed video frame;
[0027] Connect at least one of the superimposed video frames to obtain a superimposed video;
[0028] Adjust the parameters of the superimposed video according to preset video display parameters to obtain the super-resolution video.
[0029] In one embodiment, before obtaining the feature parameters of the video frame in the video to be processed, it further includes:
[0030] Obtain model training data and the data identifier corresponding to the model training data;
[0031] Input the model training data into the super-resolution model to be trained to obtain the recognition result of the model training data;
[0032] Train the super-resolution model to be trained according to the difference between the recognition result of the model training data and the data identifier to obtain a pre-trained super-resolution model; the pre-trained super-resolution model includes the pre-trained spatial model, the pre-trained temporal model, and the pre-trained codec model.
[0033] In a second aspect, the present application further provides a video processing device. The device includes:
[0034] A parameter acquisition module, configured to acquire feature parameters of video frames in a video to be processed;
[0035] A first processing module, configured to input the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frame;
[0036] A second processing module, configured to input the spatial features into a pre-trained temporal model, and fuse the spatial features and the temporal features of the video frame through the temporal model to obtain the first fused features of the video frame; the temporal features are obtained through the temporal model;
[0037] A third processing module, configured to input the first fusion feature into a pre-trained encoding and decoding model, and fuse the first fusion feature and the scale feature of the video frame through the encoding and decoding model to obtain a second fusion feature of the video frame; the scale feature is obtained through the encoding and decoding model;
[0038] A super-resolution module, configured to obtain a super-resolution video corresponding to the video to be processed according to the second fusion feature.
[0039] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0040] Obtain the feature parameters of the video frames in the video to be processed;
[0041] Input the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frame;
[0042] Input the spatial features into a pre-trained temporal model, and fuse the spatial features and the temporal features of the video frame through the temporal model to obtain a first fusion feature of the video frame; the temporal features are obtained through the temporal model;
[0043] Input the first fusion feature into a pre-trained encoding and decoding model, and fuse the first fusion feature and the scale feature of the video frame through the encoding and decoding model to obtain a second fusion feature of the video frame; the scale feature is obtained through the encoding and decoding model;
[0044] Obtain a super-resolution video corresponding to the video to be processed according to the second fusion feature.
[0045] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the following steps are implemented:
[0046] Obtain the feature parameters of the video frames in the video to be processed;
[0047] Input the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frame;
[0048] Input the spatial features into a pre-trained temporal model, and fuse the spatial features and the temporal features of the video frame through the temporal model to obtain a first fusion feature of the video frame; the temporal features are obtained through the temporal model;
[0049] Input the first fusion feature into a pre-trained encoding and decoding model, and fuse the first fusion feature and the scale feature of the video frame through the encoding and decoding model to obtain the second fusion feature of the video frame; the scale feature is obtained through the encoding and decoding model;
[0050] Obtain a super-resolution video corresponding to the video to be processed according to the second fusion feature of the video frame.
[0051] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0052] Obtain the feature parameters of the video frames in the video to be processed;
[0053] Input the feature parameters into a pre-trained spatial model to obtain the spatial feature of the video frame;
[0054] Input the spatial feature into a pre-trained temporal model, and fuse the spatial feature and the temporal feature of the video frame through the temporal model to obtain the first fusion feature of the video frame; the temporal feature is obtained through the temporal model;
[0055] Input the first fusion feature into a pre-trained encoding and decoding model, and fuse the first fusion feature and the scale feature of the video frame through the encoding and decoding model to obtain the second fusion feature of the video frame; the scale feature is obtained through the encoding and decoding model;
[0056] Obtain a super-resolution video corresponding to the video to be processed according to the second fusion feature of the video frame.
[0057] The above video processing method, device, computer device, storage medium and computer program product first obtain the feature parameters of the video frames in the video to be processed, then input the feature parameters into a pre-trained spatial model to obtain the spatial feature of the video frame, input the spatial feature into a pre-trained temporal model, fuse the spatial feature and the temporal feature of the video frame through the temporal model to obtain the first fusion feature of the video frame, input the first fusion feature into a pre-trained encoding and decoding model, fuse the first fusion feature and the scale feature of the video frame through the encoding and decoding model to obtain the second fusion feature of the video frame, and finally obtain a super-resolution video corresponding to the video to be processed according to the second fusion feature; by fusing spatial features, temporal features and scale features, video information can be fully utilized for super-resolution processing of videos, effectively improving the quality of mobile phone videos.
[0058] Moreover, by using the Mobile-Former structure with both global attention mechanism and efficient local processing ability to implement the spatial model, temporal model, and encoding-decoding model, the computational complexity can be reduced, lightweight feature extraction can be achieved, and the implementation of super-resolution video processing on mobile phones is ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic flowchart of a video processing method in an embodiment;
[0060] Figure 2 It is a structural block diagram of a mobile phone video quality improvement system based on the Mobile-Former block in an embodiment;
[0061] Figure 3 It is a schematic flowchart of the processing flow of a data preprocessing module in an embodiment;
[0062] Figure 4 It is a schematic flowchart of the processing flow of a super-resolution module in an embodiment;
[0063] Figure 5 It is a schematic flowchart of the processing flow of a result processing module in an embodiment;
[0064] Figure 6 It is a schematic flowchart of the steps for generating a super-resolution model in an embodiment;
[0065] Figure 7 It is a structural block diagram of a super-resolution network in an embodiment;
[0066] Figure 8 It is a schematic flowchart of a mobile phone video quality improvement method based on the Mobile-Former block in an embodiment;
[0067] Figure 9 It is a structural block diagram of a video processing device in an embodiment;
[0068] Figure 10 It is an internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0070] The video processing method provided by the embodiments of this application can be applied to a terminal or a server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0071] In one embodiment, as Figure 1 shown, a video processing method is provided. Taking the application of this method to a terminal as an example, the method includes the following steps:
[0072] Step S110, obtain the feature parameters of the video frames in the video to be processed.
[0073] Among them, the feature parameters can be the image features and tokens of the video frames.
[0074] In a specific implementation, feature mapping can be performed on each video frame in the video to be processed to obtain the image features of each video frame, and embedding processing can also be performed on each video frame in the video to be processed to obtain the tokens of each video frame.
[0075] Among them, feature mapping can be a method for establishing a mapping relationship between a frame image and a feature matrix.
[0076] Among them, embedding can be a method for embedding tokens and frame images.
[0077] In practical applications, the original data of the mobile phone video can be cleaned through a data preprocessing module to obtain the cleaned video. The data preprocessing module can also group the cleaned video and, taking the group as a unit, input the video frames into a super-resolution module. The super-resolution module can perform feature mapping and embedding on each video frame in each group respectively to obtain the image features and tokens of each video frame.
[0078] Step S120, input the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frames.
[0079] Among them, the spatial model can be a spatial Mobile-Former block. Among them, the Mobile-Former block can be a feature extraction module based on MobileNet (a lightweight deep neural network) and transformer (an attention mechanism network).
[0080] In a specific implementation, a spatial model with a parallel structure can be pre-trained, and the image features and tokens of the video frames are input into the pre-trained spatial model with a parallel structure to obtain the spatial features of the video frames output by each spatial model in the parallel structure.
[0081] In practical applications, spatial Mobile-Former blocks matching the number of video frames in each group can be designed in the super-resolution module. After obtaining the image features and tokens of each video frame in each group, the image features and tokens of each video frame can be respectively input into the corresponding spatial Mobile-Former blocks to obtain the spatial features and tokens output by each spatial Mobile-Former block.
[0082] Step S130: Input the spatial features into a pre-trained temporal model, and the temporal model fuses the spatial features and the temporal features of the video frames to obtain the first fused feature of the video frames; the temporal features are obtained through the temporal model.
[0083] Among them, the temporal module can be a temporal Mobile-Former block.
[0084] Among them, the first fused feature can be the fusion of the spatial features and the temporal features.
[0085] In a specific implementation, a temporal model can be pre-trained to concatenate the spatial features output by each spatial model to obtain the concatenated spatial features, and the concatenated spatial features are input into the pre-trained temporal model for temporal feature extraction to obtain the temporal features of the video frames. The temporal model can also fuse the extracted temporal features with the concatenated spatial features to obtain the first fused feature.
[0086] In practical applications, a temporal Mobile-Former block can be designed in the super-resolution module. The temporal Mobile-Former block is used to concatenate the spatial features output by the spatial Mobile-Former blocks to construct the temporal features of each video frame, simulate the time step, perform fusion processing on different frames, obtain the tokens extracted at different times, fuse the global temporal interaction information and the local temporal features, optimize all the video frame data, and improve the quality of the features that fuse space and time.
[0087] Step S140: Input the first fused feature into a pre-trained encoding-decoding model, and the encoding-decoding model fuses the first fused feature and the scale features of the video frames to obtain the second fused feature of the video frames; the scale features are obtained through the encoding-decoding model.
[0088] Among them, the encoding and decoding model can be composed of 2 downsampled Mobile-Former blocks, 1 Mobile-Former block with an unchanged scale, and 2 upsampled Mobile-Former blocks.
[0089] In a specific implementation, the encoding and decoding model can be pre-trained, and the first fused feature is input into the trained encoding and decoding model for scale feature extraction to obtain the scale feature of the video frame. The encoding and decoding model can also fuse the extracted scale feature with the first fused feature to obtain the second fused feature.
[0090] In practical applications, an encoder-decoder block can be designed in the super-resolution module. The encoder-decoder block includes 2 downsampled Mobile-Former blocks, 1 Mobile-Former block with an unchanged scale, and 2 upsampled Mobile-Former blocks. The features and tokens output by the temporal Mobile-Former block are refined through 2 downsampled Mobile-Former blocks, 1 Mobile-Former block with an unchanged scale, and 2 upsampled Mobile-Former blocks to obtain video frame features of different sizes, facilitating the acquisition of more scale-related information during reconstruction, thereby reconstructing high-resolution video frames with rich details.
[0091] Step S150, obtain a super-resolution video corresponding to the video to be processed according to the second fused feature.
[0092] In a specific implementation, the image features of the video frame can be fused with the second fused feature to obtain a third fused feature, and three-dimensional transposed convolution reconstruction is performed on the third fused feature to obtain a reconstructed video frame. The video frames in the video to be processed can also be superimposed with the reconstructed video frames to obtain super-resolution video frames, and multiple super-resolution video frames are connected to obtain the super-resolution video corresponding to the video to be processed.
[0093] The above video processing method first obtains the feature parameters of video frames in the video to be processed, then inputs the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frames, inputs the spatial features into a pre-trained temporal model, and the temporal model fuses the spatial features and the temporal features of the video frames to obtain the first fused features of the video frames. The first fused features are input into a pre-trained codec model, and the codec model fuses the first fused features and the scale features of the video frames to obtain the second fused features of the video frames. Finally, according to the second fused features, a super-resolution video corresponding to the video to be processed is obtained. By fusing spatial features, temporal features, and scale features, video information can be fully utilized for super-resolution processing of videos, effectively improving the quality of mobile phone videos.
[0094] Moreover, by using the Mobile-Former structure with both global attention mechanism and efficient local processing ability to implement the spatial model, temporal model, and codec model, the amount of computation can be reduced, lightweight feature extraction can be achieved, and the implementation of super-resolution video processing on mobile phones is ensured.
[0095] In one embodiment, the feature parameters include image features and tokens. The above step S110 may specifically include: obtaining the original video to be processed; performing data cleaning on the original video to be processed to obtain the cleaned video; grouping the cleaned video to obtain the video to be processed; performing feature mapping processing on each video frame in the video to be processed to obtain the image features of each video frame, and performing embedding processing on each video frame in the video to be processed to obtain the tokens of each video frame.
[0096] In specific implementation, the original data of the mobile phone video can be obtained, the original data of the mobile phone video can be cleaned to remove interference elements such as abnormal frequencies, pulse burrs, and background noise to obtain the cleaned video, the video frames in the cleaned video can be grouped to obtain the video to be processed including one or more groups of video frames, and feature mapping and embedding are performed on each video frame in the video to be processed to obtain the image features and tokens of each video frame respectively.
[0097] For example, the data preprocessing module can divide every 7 video frames in the cleaned video into a group to obtain the video to be processed. When the last group has less than 7 video frames, it can be padded forward. In the order of the video frames, taking the group as a unit, the 7 video frames in each group are input into the super-resolution module, and the super-resolution module can perform feature mapping and embedding on each video frame to obtain the image features and tokens.
[0098] In this embodiment, by obtaining the original video to be processed, cleaning the data of the original video to be processed to obtain the cleaned video, grouping the cleaned video to obtain the video to be processed, performing feature mapping processing on each video frame in the video to be processed to obtain the image features of each video frame, and performing embedding processing on each video frame in the video to be processed to obtain the tokens of each video frame, it is possible to remove the interference elements in the original data of the mobile phone video through data cleaning, improve the reliability of video processing, and improve the efficiency of video processing by grouping and parallel processing multiple video frames.
[0099] In one embodiment, the video to be processed includes at least one group of video frames, and each video frame in each group of video frames corresponds to a spatial model respectively; the above step S120 may specifically include: inputting the feature parameters of each video frame in each group of video frames into the spatial model corresponding to the video frame respectively to obtain the spatial features of each video frame.
[0100] In specific implementation, multiple parallel spatial models matching the number of video frames in each group of video frames can be designed. After obtaining the image features and tokens of each video frame in each group of video frames, the image features and tokens of each video frame can be input into the corresponding spatial model respectively to obtain the spatial features and tokens output by the spatial model.
[0101] For example, 7 spatial Mobile-Former blocks can be designed in parallel in the super-resolution module, and the image features and tokens of 7 video frames in each group are input into the 7 spatial Mobile-Former blocks. The spatial Mobile-Former blocks are used to model the global interaction between the tokens extracted from 7 video frames at the same time, and perform local processing on the data features of a single picture to optimize the data features of a single frame and improve the quality of spatial features.
[0102] In this embodiment, by inputting the feature parameters of each video frame in each group of video frames into the spatial model corresponding to the video frame respectively to obtain the spatial features of each video frame, the spatial features of multiple video frames can be obtained in parallel, improving the efficiency of video processing and being beneficial to the real-time performance of the system.
[0103] In one embodiment, the encoding and decoding model includes two downsampling sub-models, one scale-invariant sub-model, and two upsampling sub-models; the above step S140 may specifically include: inputting the first fusion feature into the two downsampling sub-models, one scale-invariant sub-model, and two upsampling sub-models in sequence to obtain the second fusion feature of the video frame.
[0104] Among them, the downsampling sub-model can be a downsampled Mobile-Former block. The scale-invariant sub-model can be a Mobile-Former block with an unchanged scale. The upsampling sub-model can be an upsampled Mobile-Former block.
[0105] In a specific implementation, the first fusion feature output by the temporal Mobile-Former block can be sequentially input into 2 downsampled Mobile-Former blocks, 1 Mobile-Former block with an unchanged scale, and 2 upsampled Mobile-Former blocks to obtain the scale features of the video frame, and the first fusion feature and the scale features are fused to obtain the second fusion feature of the video frame.
[0106] In this embodiment, by sequentially inputting the first fusion feature into two downsampling sub-models, one scale-invariant sub-model, and two upsampling sub-models to obtain the second fusion feature of the video frame, the spatial information of a single video frame, the spatial information of consecutive video frames, and the size information of video frames at different scales can be fused, making full use of video data, which is beneficial to improving the quality of the video.
[0107] In one embodiment, the above step S150 may specifically include: fusing the second fusion feature of the video frame with the image feature of the video frame to obtain the third fusion feature of the video frame; performing deconvolution layer reconstruction processing on the third fusion feature to obtain the reconstructed video frame; and superimposing the reconstructed video frame with the video frame of the video to be processed to obtain the super-resolution video.
[0108] In a specific implementation, the image feature of the video frame can be fused with the second fusion feature to obtain the third fusion feature, and three-dimensional deconvolution reconstruction is performed on the third fusion feature to obtain the reconstructed video frame. It is also possible to perform upsampling on the video frames in the video to be processed to obtain upsampled low-resolution video frames, and superimpose the upsampled low-resolution video frames with the reconstructed video frame to obtain super-resolution video frames. Connecting multiple super-resolution video frames can obtain the super-resolution video corresponding to the video to be processed.
[0109] In this embodiment, by fusing the second fusion feature of the video frame with the image feature of the video frame to obtain the third fusion feature of the video frame, performing deconvolution layer reconstruction processing on the third fusion feature to obtain the reconstructed video frame, and superimposing the reconstructed video frame with the video frame of the video to be processed to obtain the super-resolution video, spatial information, temporal information, and scale information can be superimposed on the original low-resolution video to reconstruct a high-resolution video with rich details and improve the video quality.
[0110] In one embodiment, the step of superimposing the reconstructed video frame and the video frame of the video to be processed to obtain a super-resolution video may specifically include: superimposing the reconstructed video frame and the video frame of the video to be processed to obtain a superimposed video frame; connecting at least one superimposed video frame to obtain a superimposed video; according to preset video display parameters, adjusting the parameters of the superimposed video to obtain a super-resolution video.
[0111] Among them, the video display parameters may include a video display size and a video display format.
[0112] In specific implementation, the reconstructed video frame obtained by three-dimensional deconvolution reconstruction may be superimposed on the video frame of the video to be processed to obtain a superimposed video frame, and a plurality of consecutive superimposed video frames may be obtained. According to the order of the video frames in the video to be processed, the plurality of superimposed video frames may be connected to obtain a superimposed video. According to the preset video display size and video display format, the superimposed video may be adjusted to obtain a super-resolution video.
[0113] In this embodiment, by superimposing the reconstructed video frame and the video frame of the video to be processed to obtain a superimposed video frame, connecting at least one superimposed video frame to obtain a superimposed video, and adjusting the parameters of the superimposed video according to the preset video display parameters to obtain a super-resolution video, a super-resolution video that meets the screen display requirements can be output to meet the video display needs.
[0114] In one embodiment, before the above step S110, it may specifically further include: obtaining model training data and a data identifier corresponding to the model training data; inputting the model training data into a super-resolution model to be trained to obtain an identification result of the model training data; training the super-resolution model to be trained according to the difference between the identification result of the model training data and the data identifier to obtain a pre-trained super-resolution model; the pre-trained super-resolution model includes a pre-trained spatial model, a pre-trained temporal model, and a pre-trained encoding and decoding model.
[0115] Among them, the super-resolution model may be composed of 7 parallel spatial Mobile-Former blocks, 1 temporal Mobile-Former block, 2 downsampling Mobile-Former blocks, 1 Mobile-Former block with an invariant scale, and 2 upsampling Mobile-Former blocks.
[0116] In a specific implementation, the Vimeo90K dataset can be used as the training dataset. This dataset contains model training data. Obtain the data identifier corresponding to the model training data, input the model training data into the super-resolution model to be trained, obtain the recognition result of the super-resolution model to be trained for the model training data, compare the recognition result of the model training data with the data identifier, and adjust the parameters of the super-resolution model to be trained according to the difference between the two. Repeat the above process. After multiple adjustments, a pre-trained super-resolution model can be obtained.
[0117] In this embodiment, by obtaining the model training data and the data identifier corresponding to the model training data, inputting the model training data into the super-resolution model to be trained, obtaining the recognition result of the model training data, and training the super-resolution model to be trained according to the difference between the recognition result of the model training data and the data identifier, a pre-trained super-resolution model can be obtained, and a trained super-resolution model can be obtained, which is convenient for improving the video quality through the super-resolution model and increasing the efficiency of video processing.
[0118] To facilitate the in-depth understanding of the embodiments of the present application by those skilled in the art, the following will be described with a specific example.
[0119] Current super-resolution technologies usually use convolutional neural networks to process video sequences, generally by processing the reconstructed frames from the support frames or optical flow estimation. Since the number of frames in a video is generally large, processing frame by frame results in poor parallel efficiency and wastes certain resources. Vision transformer can also be used for video super-resolution. Vision transformer often adopts more network layers and uses the attention mechanism extensively to establish a perfect global dependence relationship. However, its performance in local processing is poor, the stacked transformers have a large depth, and the computational amount for processing each video frame is also large, making it difficult to be applied to mobile devices such as mobile phones and tablets. The recently proposed Mobile-Former network combines the advantages of MobileNet and transformer, can build a global dependence relationship while maintaining lightweight, and perform efficient image classification. However, this network can only process single images, does not consider reconstruction, only has an encoder function, and cannot perform super-resolution.
[0120] Figure 2 A structural block diagram of a mobile phone video quality improvement system based on the Mobile-Former block is provided. According to Figure 2 , the mobile phone video quality improvement system 201 may include a data preprocessing module 202, a super-resolution module 203, and a result processing module 204.
[0121] Among them, the data preprocessing module 202 is responsible for collecting the original video data and performing preprocessing to obtain the data features available for the super-resolution module 203, mainly including: obtaining videos, data cleaning, and extracting video frames. The data preprocessing module 202 regards 7 video frames to be processed as a group each time and inputs them into the super-resolution module 203. Due to the grouped processing of video frames, under the condition that resources permit, the video frame data of different groups can be processed in parallel, improving the overall processing efficiency.
[0122] Among them, the super-resolution module 203 uses a deep learning model composed of Mobile-Former blocks to model the joint features of audio-visual data, extract features, and reconstruct high-resolution video data. First, the Mobile-Former blocks are used in parallel for the 7 input video frames to extract features and construct the internal spatial features of each video frame. The Former structure in this part models the global interaction between the tokens extracted at the same time for each frame image. Mobile performs local processing on individual image frames, and through the interaction between Mobile and Former, the global and local information is organically fused, and finally the feature maps of 7 frames of pictures are obtained. Secondly, the features processed from the 7 frames of pictures are concatenated to construct the internal temporal features of the 7 video frames, simulate the time step, perform fusion processing on different frames, obtain the tokens extracted at different times, and use the Mobile-Former blocks to extract features, fusing the global temporal interaction information and the local temporal features. Then, the feature maps and tokens that have undergone spatial and temporal feature extraction processing are refined through 5 symmetric Mobile-Former blocks, including two downsampling and two upsampling Mobile-Former blocks, for obtaining video information of different sizes. Finally, through an anti-convolution reconstruction module that fuses global information once, combined with the 7 initially low-resolution frames after upsampling, the final super-resolution video is reconstructed.
[0123] Among them, the result processing module 204 processes the data output by the super-resolution, assembles the super-resolved video frames in the order of the video frames extracted by the data preprocessing module 202. Since the video frames become larger after video super-resolution, after assembly, the video size is adjusted as needed to meet the screen requirements for playing the video, and the complete video data is output.
[0124] Figure 3 A schematic diagram of the processing flow of the data preprocessing module is provided. According to Figure 3 , Figure 2 the data preprocessing module 202 in
[0125] Step S301: Obtain the original video data.
[0126] Step S302: Perform data cleaning on the original video data to remove interference elements such as abnormal frequencies, pulse glitches, and background noise.
[0127] Step S303: Split the cleaned video, group all video frames, with 7 frames in a group. Pad the last group with less than 7 frames forward. Then, in order, input the video frame data into the super-resolution module 2 in groups for super-resolution.
[0128] Figure 4 A schematic diagram of the processing flow of the super-resolution module is provided. According to Figure 4 , Figure 2 the super-resolution module 203 in, call the deep learning model composed of Mobile-Former blocks to enhance the super-resolution of consecutive video frames, thereby improving the video quality and outputting clear high-quality video data. The specific processing steps include:
[0129] Step S401: Obtain the preprocessed data, perform feature mapping and embedding on each video frame to obtain image features and tokens.
[0130] Step S402: Use the spatial Mobile-Former block in the deep learning model to model the global interaction between the tokens extracted from 7 video frames at the same time, perform local processing on the single-image data features, optimize the single-frame data features, and improve the spatial feature quality.
[0131] Step S403: Use the temporal Mobile-Former block in the deep learning model to concatenate the optimized spatial features of 7 video frames, construct the internal temporal features of 7 video frames, simulate the time step, perform fusion processing on different frames, obtain the tokens extracted at different times, fuse the global temporal interaction information and the local temporal features, optimize all video frame data, and improve the quality of the fused spatial and temporal features.
[0132] Step S404: Use the encoder - decoder block in the deep - learning model to refine the features and tokens output by the Mobile - Former block in step S403 through 2 downsampled Mobile - Former blocks, 1 Mobile - Former block with unchanged scale, and 2 upsampled Mobile - Former blocks, forming an overall encoder - decoder structure. This is used to obtain video - frame features of different sizes, facilitating the acquisition of more scale - related information during reconstruction, thereby reconstructing high - resolution video frames with rich details.
[0133] Step S405: Use the reconstruction module in the deep - learning model to process the data features and generate high - quality speech and video outputs. The reconstruction module first fuses the features of the 7 video frames extracted in step S401 and the features output by the Mobile - Former block in step S404. Secondly, it reconstructs the fused features through a transposed convolutional layer. Finally, it superimposes the reconstructed video frames and the initial low - resolution video frames obtained through upsampling respectively to improve the accuracy of the overall structure.
[0134] Among them, the Mobile - Former block is a parallel design of MobileNet and transformer, with a bidirectional bridging structure, which can combine the advantages of local processing of MobileNet and global interaction of transformer. It is an efficient and lightweight feature extraction module.
[0135] Among them, the decoder / encoder is a common model framework in deep learning. The model can adopt CNN, RNN, LSTM, etc. The encoder converts the input sequence into a vector of a fixed dimension, and the decoder generates the target translation from the activation state.
[0136] Figure 5 A schematic diagram of the processing flow of a result - processing module is provided. According to Figure 5 , Figure 2 the result - processing module 204 in
[0137] is responsible for assembling the reconstructed video frames output by the super - resolution module 203 in sequence and outputting appropriate video data. The specific processing steps include:
[0138] Step S501: Assemble the super - resolved video frames in the order of video frames extracted by the data pre - processing module to form video data.
[0139] Step S502: Adjust the video size and format according to the screen requirements and output requirements.
[0140] Figure 6 A flowchart showing the steps of generating a super-resolution model is provided. Figure 7 A structural block diagram of a super-resolution network is provided. According to Figure 6 and Figure 7 , the specific steps of generating the super-resolution model are as follows:
[0141] Step S601: Use the widely used vimeo90K dataset in the industry as the training dataset, which has already been divided into a training set and a test set.
[0142] Step S602: Use the training data to train a deep learning model based on the Mobile-Former block until the model accuracy reaches the threshold.
[0143] Step S603: Output the trained quality improvement model.
[0144] Figure 8 A flowchart showing the method for improving the quality of mobile phone videos based on the Mobile-Former block is provided. According to Figure 8 , the method for improving the quality of mobile phone videos based on the Mobile-Former block is as follows:
[0145] Step S801: The data preprocessing module collects the original video data, performs data cleaning, and divides the video into different groups of video frames, and inputs them into the super-resolution module.
[0146] Step S802: The super-resolution module improves the quality of each group of 7 input video frames to obtain high-quality video frame data.
[0147] Step S803: The result processing module adjusts the video format and size according to the playback requirements and outputs the video data.
[0148] The above-mentioned system and method for improving the quality of mobile phone videos based on the Mobile-Former block use the Mobile-Former structure with global attention mechanism and efficient local processing ability, which can combine the advantages of both the transformer and MobileNet. At the same time, the size and computational complexity of this module are relatively small, making it more suitable for processing and improving mobile phone videos.
[0149] Moreover, when processing the video super-resolution task, the spatial features of multiple video frames are processed in parallel, and multiple video frames are reconstructed simultaneously, improving the parallelism of the network and being beneficial to the real-time performance of the system.
[0150] Furthermore, due to the fusion of the spatial information of individual video frames, the spatial information of consecutive video frames, and the size information of video frames at different scales, the video data is fully utilized, which is beneficial to improving the quality of the video.
[0151] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0152] Based on the same inventive concept, an embodiment of the present application also provides a video processing device for implementing the above-mentioned video processing method. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following video processing device can refer to the limitations on the video processing method in the above text, and will not be repeated here.
[0153] In one embodiment, as Figure 9 shown, a video processing device is provided, including: a parameter acquisition module 910, a first processing module 920, a second processing module 930, a third processing module 940, and a super-resolution module 950, where:
[0154] The parameter acquisition module 910 is used to acquire the feature parameters of the video frames in the video to be processed;
[0155] The first processing module 920 is used to input the feature parameters into a pre-trained spatial model to obtain the spatial features of the video frames;
[0156] The second processing module 930 is used to input the spatial features into a pre-trained temporal model, and fuse the spatial features and the temporal features of the video frames through the temporal model to obtain the first fused features of the video frames; the temporal features are obtained through the temporal model;
[0157] The third processing module 940 is used to input the first fused features into a pre-trained codec model, and fuse the first fused features and the scale features of the video frames through the codec model to obtain the second fused features of the video frames; the scale features are obtained through the codec model;
[0158] The super-resolution module 950 is used to obtain a super-resolution video corresponding to the video to be processed according to the second fused features.
[0159] In one embodiment, the above-mentioned parameter acquisition module 910 is further configured to obtain the original video to be processed; perform data cleaning on the original video to be processed to obtain a cleaned video; group the cleaned video to obtain the video to be processed; perform feature mapping processing on each video frame in the video to be processed to obtain the image features of each video frame, and perform embedding processing on each video frame in the video to be processed to obtain the tags of each video frame.
[0160] In one embodiment, the above-mentioned first processing module 920 is further configured to input the feature parameters of each video frame in each group of video frames into the spatial model corresponding to the video frame respectively to obtain the spatial features of each video frame.
[0161] In one embodiment, the above-mentioned third processing module 940 is further configured to input the first fusion feature into the two downsampling sub-models, a scale-invariant sub-model, and two upsampling sub-models in sequence to obtain the second fusion feature of the video frame.
[0162] In one embodiment, the above-mentioned super-resolution module 950 further includes:
[0163] A feature fusion module, configured to fuse the second fusion feature of the video frame with the image feature of the video frame to obtain the third fusion feature of the video frame;
[0164] A video reconstruction module, configured to perform deconvolution layer reconstruction processing on the third fusion feature to obtain a reconstructed video frame;
[0165] A video superposition module, configured to superpose the reconstructed video frame with the video frame of the video to be processed to obtain the super-resolution video.
[0166] In one embodiment, the above-mentioned video superposition module is further configured to superpose the reconstructed video frame with the video frame of the video to be processed to obtain a superposed video frame; connect at least one of the superposed video frames to obtain a superposed video; adjust the parameters of the superposed video according to preset video display parameters to obtain the super-resolution video.
[0167] In one embodiment, the above-mentioned video processing device further includes:
[0168] A sample acquisition module, configured to obtain model training data and the data identifier corresponding to the model training data;
[0169] A sample recognition module, configured to input the model training data into a super-resolution model to be trained to obtain the recognition result of the model training data;
[0170] A model training module, configured to train the super-resolution model to be trained according to the difference between the recognition result of the model training data and the data identifier, so as to obtain a pre-trained super-resolution model; the pre-trained super-resolution model includes the pre-trained spatial model, the pre-trained temporal model, and the pre-trained encoding and decoding model.
[0171] Each module in the above video processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0172] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a video processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0173] Those skilled in the art can understand that Figure 10 the structure shown in
[0174] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0175] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0176] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0177] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.
[0178] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0179] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0180] The above-described embodiments merely represent several implementation manners of the present application, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A video processing method, characterized in that, The method includes: Obtaining feature parameters of video frames in a video to be processed, where the feature parameters include image features; Inputting the feature parameters into a pre-trained spatial model to obtain spatial features of the video frames; Inputting the spatial features into a pre-trained temporal model, and fusing the spatial features and the temporal features of the video frames through the temporal model to obtain first fusion features of the video frames; the temporal features are obtained through the temporal model; Inputting the first fusion features into a pre-trained codec model, and fusing the first fusion features and the scale features of the video frames through the codec model to obtain second fusion features of the video frames; the scale features are obtained through the codec model; Fusing the second fusion features with the image features of the video frames to obtain third fusion features of the video frames; Performing deconvolution layer reconstruction processing on the third fusion features to obtain reconstructed video frames; Overlaying the reconstructed video frames with the video frames of the video to be processed to obtain overlaid video frames; Connecting at least one of the overlaid video frames to obtain an overlaid video; Adjusting parameters of the overlaid video according to preset video display parameters to obtain a super-resolution video corresponding to the video to be processed.
2. The method according to claim 1, wherein The feature parameters further include tags; the obtaining of the feature parameters of video frames in the video to be processed includes: Obtaining an original video to be processed; Performing data cleaning on the original video to be processed to obtain a cleaned video; Grouping the cleaned video to obtain the video to be processed; Performing feature mapping processing on each of the video frames in the video to be processed to obtain image features of each of the video frames, and performing embedding processing on each of the video frames in the video to be processed to obtain tags of each of the video frames.
3. The method according to claim 2, wherein, Before obtaining the feature parameters of video frames in the video to be processed, it further includes: Inputting the feature parameters of each of the video frames in each group of video frames into the spatial model corresponding to the video frame respectively to obtain spatial features of each of the video frames.
4. The method according to claim 2, wherein The codec model includes two downsampling sub-models, a scale-invariant sub-model, and two upsampling sub-models; the inputting of the first fusion features into the pre-trained codec model, and the fusing of the first fusion features and the scale features of the video frames through the codec model to obtain the second fusion features of the video frames includes: Inputting the first fusion features into the two downsampling sub-models, the scale-invariant sub-model, and the two upsampling sub-models in sequence to obtain the second fusion features of the video frames.
5. The method according to claim 1, wherein Before obtaining the feature parameters of video frames in the video to be processed, it further includes: Obtaining model training data and data identifiers corresponding to the model training data; Inputting the model training data into a super-resolution model to be trained to obtain an identification result of the model training data; Training the super-resolution model to be trained based on the difference between the recognition result of the model training data and the data identifier, to obtain a pre-trained super-resolution model; the pre-trained super-resolution model includes the pre-trained spatial model, the pre-trained temporal model, and the pre-trained encoding and decoding model.
6. A video processing device, characterized in that, The device includes: A parameter acquisition module, configured to acquire feature parameters of video frames in a video to be processed, where the feature parameters include image features; A first processing module, configured to input the feature parameters into the pre-trained spatial model to obtain spatial features of the video frames; A second processing module, configured to input the spatial features into the pre-trained temporal model, and fuse the spatial features and the temporal features of the video frames through the temporal model to obtain first fusion features of the video frames; the temporal features are obtained through the temporal model; A third processing module, configured to input the first fusion features into the pre-trained encoding and decoding model, and fuse the first fusion features and the scale features of the video frames through the encoding and decoding model to obtain second fusion features of the video frames; the scale features are obtained through the encoding and decoding model; A super-resolution module, configured to obtain a super-resolution video corresponding to the video to be processed according to the second fusion features; The super-resolution module is specifically configured to: Fuse the second fusion features with the image features of the video frames to obtain third fusion features of the video frames; Perform deconvolution layer reconstruction processing on the third fusion features to obtain reconstructed video frames; Overlay the reconstructed video frames with the video frames of the video to be processed to obtain overlaid video frames; Connect at least one of the overlaid video frames to obtain an overlaid video; Adjust the parameters of the overlaid video according to preset video display parameters to obtain a super-resolution video corresponding to the video to be processed.
7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Progressive feature flow deep fusion network for monitoring video enhancement
CN112348766A
Video super-resolution processing method and device, super-resolution reconstruction model and medium
CN112950471A