Video encoding method, video decoding method, electronic device, and storage medium
By adjusting video frames to a fixed resolution, the method simplifies video encoding and decoding, allowing a single set of neural networks to handle frames with various resolutions, thus enhancing operational convenience and versatility.
Patent Information
- Application Number
- JP2024575719
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-23
- Filing Date
- 2023-06-21
- Publication Date
- 2025-06-26
AI Technical Summary
Existing video encoding and decoding methods based on deep video generation require separate neural network models for each video frame resolution, making them complex and inconvenient for handling multiple resolutions.
A method that adjusts the resolution of original video frames to a fixed preset resolution, allowing a single feature extraction network, motion estimation network, and generation network to handle video frames of various resolutions.
This approach simplifies the encoding and decoding processes by eliminating the need for multiple neural networks, enabling efficient encoding and decoding of video frames with different resolutions using a single set of neural network models.
Smart Images

Figure 2025519941000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the priority of a Chinese patent application titled "Video Encoding Method, Video Decoding Method, Electronic device and Storage Medium" with application number 202210716251.4, filed with the Chinese Patent Office on June 23, 2022, the entire content of which is incorporated herein by reference.
[0002] Embodiments of this application relate to the field of computer technology, and in particular, to video encoding methods, video decoding methods, electronic devices, and storage media.
Background Art
[0003] With the continuous development of computer technology, people's lifestyles have also changed significantly. For example, in daily work and life, people's demand for video conferencing and video live broadcasts has been continuously increasing.
[0004] Video encoding and decoding are the keys to realizing video conferencing and video live broadcasts. With the continuous development of machine learning, codec methods based on deep video generation can be used to encode and decode videos (especially face videos). This method mainly uses a neural network model to deform a reference frame based on the motion of the frame to be encoded, so as to obtain a reconstructed frame corresponding to the frame to be encoded. The above method executes the encoding operation and decoding operation on the video frame end-to-end to realize the reconstruction of the video frame.
[0005] The above codec method based on deep video generation by a set of fully trained neural network models can usually be used only to reconstruct video frames having a fixed resolution to be encoded and cannot support multiple different resolutions. However, in practical applications, due to factors such as network bandwidth, there may be multiple resolutions for the video frames to be encoded instead of a fixed resolution. At this time, a corresponding set of neural network models must be trained for each resolution, and at the application stage, the corresponding network model is loaded according to the actual resolution of the video frames to be encoded. Such an operation is complex and very inconvenient.
Summary of the Invention
Problems to be Solved by the Invention
[0006] In view of this, embodiments of the present application provide a video encoding method, a decoding method, an electronic device, and a storage medium for at least partially solving the above problems.
Means for Solving the Problems
[0007] According to a first aspect of the embodiments of the present application, a video encoding method is provided. obtaining an original reference video frame and an original target video frame to be encoded; adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and performing feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution; Encoding the original reference video frame and the target feature respectively to obtain a video bitstream, and performing reconstruction of the video frame based on the video bitstream to generate a reconstructed video frame having the same resolution as the original target video frame.
[0008] According to a second aspect of the embodiments of the present application, a video decoding method is provided. Obtaining and decoding a video bitstream to obtain an original reference video frame and a target feature; Adjusting the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, and extracting features from the adjusted reference video frame through a feature extraction network to obtain reference features; Performing motion estimation based on the reference feature and the target feature through a motion estimation network to obtain a motion estimation result; Generating a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through a generation network.
[0009] According to a third aspect of the embodiments of the present application, a video encoding method is provided. Obtaining an original reference video frame and an original target video frame to be encoded; Adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and performing feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution. The decoding device decodes the video bitstream to obtain the original reference video frame and the target feature, adjusts the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, extracts the feature of the adjusted reference video frame to obtain the reference feature through the feature extraction network, executes motion estimation based on the reference feature and the target feature through the motion estimation network to obtain the motion estimation result, and generates a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through the generation network, and includes the step of encoding the original reference video frame and the target feature respectively to obtain the video bitstream.
[0010] According to a fourth aspect of the embodiment of the present application, a video encoding method is provided. The step of acquiring the original video clip captured by the video acquisition device. The step of determining the original reference video frame and the original target video frame to be encoded from the original video clip. The step of adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and extracting the feature of the adjusted target video frame through the feature extraction network corresponding to the first preset resolution to obtain the target feature. The step of encoding the original reference video frame and the target feature respectively to obtain the video bitstream. The step of transmitting the video bitstream to the conference terminal device to cause the conference terminal device to execute the reconstruction of the video frame based on the video bitstream, and generate and display a reconstructed video frame having the same resolution as the original target video frame.
[0011] According to a fifth aspect of an embodiment of the present application, an electronic device is provided, including a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory is configured to store at least one executable instruction, and the executable instruction enables the processor to execute operations corresponding to the video encoding method described in the first aspect, the third aspect, or the fourth aspect, or operations corresponding to the video decoding method described in the second aspect.
[0012] According to a sixth aspect of an embodiment of the present application, a computer storage medium storing a computer program is provided. When the program is executed by a processor, the video encoding method described in the first aspect, the third aspect, or the fourth aspect is implemented, or the video decoding method described in the second aspect is implemented.
[0013] According to a seventh aspect of an embodiment of the present application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions command a computing device to execute operations corresponding to the video encoding method described in the first aspect, the third aspect, or the fourth aspect, or operations corresponding to the video decoding method described in the second aspect.
[0014] According to the video encoding method and decoding method provided by the embodiments of the present application, in the encoding stage, after the original target video frame to be encoded is obtained, the original target video frame is universally resolved by a resolution adjustment operation, and the original target video frame is converted into an adjusted target video frame having a fixed resolution (a first preset resolution). Thus, even if the original target video frame has various resolutions, a video frame having a fixed resolution is finally input to the feature extraction network. In this way, it is not necessary to train multiple feature extraction networks for different resolutions, and only one feature extraction network corresponding to the first preset resolution (a feature extraction network for extracting features from a video frame having the first preset resolution) is required to implement the encoding of the original target video frames having multiple different resolutions, which has a wider range of applications and higher versatility. At the same time, the operation is simpler and more convenient. In addition, correspondingly, in the decoding stage, the original reference video frame is also universally resolved, and the original reference video frame is converted into an adjusted reference video frame having a fixed resolution, thereby obtaining reference features and target features of a fixed resolution. Therefore, the features finally input to the motion estimation network and the generation network are features of a fixed resolution, whereby only one motion estimation network and one generation network are required to implement decoding in various scenes of target video frames having different resolutions. In summary, the embodiments of the present application only require one set of neural network models to perform the encoding and decoding operations for the original target video frames of various resolutions, which has a wider range of applications along with a simpler and more convenient process of operation.
[0015] To more clearly illustrate the embodiments of the present application or the technical solutions of the existing technology, the accompanying drawings used in the embodiments or the existing technology are briefly described below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. Those skilled in the art can obtain other drawings based on these drawings.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying out the Invention
[0017] To enable those skilled in the art to more deeply understand the technical solutions of the embodiments of the present application, the technical solutions of the embodiments of the present application are clearly and completely described in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in the embodiments of the present application should belong to the protection scope of the embodiments of the present application.
[0018] FIG. 1 is a schematic diagram of the framework of an encoding and decoding method based on deep video generation. The main principle of this method is to deform the reference frame based on the motion of the frame to be encoded to obtain a reconstructed frame corresponding to the frame to be encoded. The basic framework of the encoding and decoding method based on deep video generation is described in combination with FIG. 1.
[0019] The first step is the encoding stage. The encoder uses a feature extractor to extract the target keypoint information of the target face video frame to be encoded and encodes the target keypoint information. At the same time, the reference face video frame is encoded using a conventional image encoding method (such as VVC, HEVC, etc.).
[0020] The second step is the decoding stage. The motion estimation module in the decoder extracts the reference keypoint information of the reference face video frame through the keypoint extractor, and performs dense motion estimation based on the reference keypoint information and the target keypoint information to obtain a dense motion estimation map and an occlusion map. The dense motion estimation map represents the relative motion relationship between the target face video frame and the reference face video frame in the feature domain represented by the keypoint information, and the occlusion map represents the degree to which each pixel in the target face video frame is occluded.
[0021] The third step is the decoding stage. The generation module in the decoder performs a deformation process on the reference face video frame based on the dense motion estimation map to obtain a deformation process result, and multiplies the deformation process result by the occlusion map to output the reconstructed face video frame.
[0022] The method shown in FIG. 1 is based on a neural network model composed of a feature extractor (feature extraction module), a motion estimation module, and a generation module to perform encoding and decoding operations on video frames. After the training of each neural network module of the above model is completed, its internal parameters and the resolution sizes of the input and output data remain fixed. Therefore, in the inference stage, the set of trained neural network models can only be used to reconstruct video frames with a specific resolution to be encoded, and cannot be compatible with multiple different resolutions.
[0023] However, in practical applications, due to factors such as network bandwidth, there may be multiple resolutions instead of a fixed resolution for the video frames to be encoded. At this time, when performing the above-described method of encoding and decoding based on deep video generation, one set of corresponding neural network models must be trained for each resolution, and at the inference stage, according to the actual resolution of the video frames to be encoded, the corresponding model is loaded. Such an operation is complex and very inconvenient.
[0024] In an embodiment of the present application, the resolution of the original target video frame is unified by a resolution adjustment operation, and the original target video frame is converted into an adjusted target video frame having a specific resolution. Then, subsequent feature extraction and other operations are performed to encode and decode the downsampled target video frame, and a reconstructed video frame having the same resolution as the original target video frame of each original target video frame is finally output. In this way, even if the original target video frames have various different resolutions, the data finally input to the feature extraction network, the motion estimation network, and the generation network still has a fixed resolution. Therefore, it is not necessary to train multiple neural networks for different resolutions, and only one neural network is required to realize the encoding of multiple original target video frames having different resolutions, and thus has a wider range of applications and higher versatility. At the same time, the operation is simpler and more convenient.
[0025] The details of the implementation of the embodiment of the present application are further described in connection with the accompanying drawings of the embodiment of the present application.
[0026] First Embodiment Referring to FIG. 2, FIG. 2 is a flowchart of a video encoding method according to the first embodiment of the present application. In particular, the video encoding method provided in this embodiment includes the following steps.
[0027] Step 202: Obtain the original reference video frame and the original target video frame to be encoded.
[0028] In particular, the original reference video frame and the original target video frame in the present application are video frames having the same resolution, and both the original reference video frame and the original target video frame can be face video frames. Further, in the embodiments of the present application, the resolution sizes of the original reference video frame and the original target video frame are not limited.
[0029] Furthermore, in order to obtain a relatively high-quality reconstructed video frame when the video frame is subsequently reconstructed, the original reference frame and the original target video frame to be encoded can be selected from the same video clip, that is, in this step, the original reference video frame and the original target video frame to be encoded from the same video clip can be obtained.
[0030] Step 204: Adjust the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution.
[0031] In particular, in the embodiments of the present application, the specific value of the first preset resolution is not limited and can be customized according to the encoding and decoding computing power resources of the encoding device and the decoding device, the network status, and the timeliness requirements.
[0032] In particular, in order to reduce the bit rate, the first preset resolution can be set to a lower value. Correspondingly, in this step, the original target video frame can be downsampled to obtain an adjusted target video frame having the first preset resolution.
[0033] Optionally, in some embodiments, the adjusted target video frame can be obtained by determining a first target scaling factor based on the resolution of the original target video frame, i.e., and using the first target scaling factor to scale the original target video frame to obtain an adjusted target video frame having a first preset resolution. The first target scaling factor can be determined based on the size relationship between the resolution of the original target video frame and the first preset resolution. In particular, the ratio of the resolution of the original target video frame to the first preset resolution can be determined as the first target scaling factor.
[0034] Furthermore, for some possible resolutions of the original target video frame, the respective scaling factors corresponding to each resolution can be pre-calculated and put into a sequence of the first scaling factors. Then, after the original target video frame is obtained, the first target scaling factor corresponding to the resolution of the original target video frame can be determined from the preset sequence of the first scaling factors according to the preset corresponding relationship between the resolution and the scaling factor.
[0035] Step 206: Perform feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution.
[0036] In the embodiments of the present application, in order to obtain target features, feature extraction can be performed on the adjusted target video frame with the help of a machine learning model (feature extraction network). In particular, the adjusted target video frame can be input into a pre-trained feature extraction network to output the target features of the adjusted target video frame of the feature extraction network.
[0037] The feature extraction network of the embodiment of the present application is a feature extraction network corresponding to a first preset resolution, that is, a network model for extracting features from a video frame having the first preset resolution.
[0038] The target feature can be information used to characterize the features of the adjusted target video frame. For a face video frame, the target feature can be, in particular, a target key point feature or a target compact feature. The target key point feature characterizes the feature information of preset key points in the adjusted target video frame, and the target compact feature characterizes important information such as the position information, pose information, and expression information of the face features in the adjusted target video frame.
[0039] In the embodiment of the present application, the structure and parameters of the feature extraction network are not limited and can be set according to actual needs. For example, the feature extraction network can be a U-Net network composed of a convolutional layer and a generalized division normalization layer.
[0040] Step 208: Separate the original reference video frame and the target feature for encoding to obtain a video bitstream, and perform reconstruction of the video frame based on the video bitstream to generate a reconstructed video frame having the same resolution as the original target video frame.
[0041] In particular, for the original reference video frame, a relatively small quantization distortion can be used for encoding, and the encoding process retains the complete data of the original reference video frame. For example, the original reference video frame can be encoded by the method of Versatile Video Coding (VVC). For the target feature, the encoding can be performed by quantization and entropy encoding.
[0042] Furthermore, in some embodiments of the present application, in order to further reduce the bitrate of video encoding, the original reference video frame can also be adjusted in resolution to obtain an adjusted reference video frame, and feature extraction can be performed on the adjusted reference video frame to obtain reference features. Then, a differential operation can be performed on the target features and the reference features, and the difference obtained by the differential operation is encoded to form a video bitstream.
[0043] Compared with the encoding method directly based on the target features, the above method is based on the difference between the target features and the reference features to obtain a video bitstream. Obviously, the amount of data of the difference between the target features and the reference features is less than the amount of data of the target features themselves. Therefore, encoding based on the difference between the target features and the reference features can effectively reduce the bitrate of video encoding.
[0044] Referring to FIG. 3, FIG. 3 is a schematic diagram of a scenario corresponding to the first embodiment of the present application. Hereinafter, the embodiments of the present application will be described with reference to the schematic diagram shown in FIG. 3, and a specific scenario will be used as an example for explanation.
[0045] The original reference video frame and the original target video frame to be encoded are obtained respectively. The resolutions of the original reference video frame and the original target video frame are both W×H (the number of pixels included in the unit size in the width direction is W, and the number of pixels included in the unit size in the height direction is H). The sequence of the first scaling factor s = {s1, s2, s3, ..., s nFrom {}, the first target scaling factor is determined, and the original target video frame is downsampled based on the first target scaling factor to obtain an adjusted target video frame of W1×H1 (the first preset resolution). Feature extraction is performed on the adjusted target video frame through a feature extraction network to obtain target features. The original reference video frame and the target features are separately encoded to obtain a video bitstream. The target features are encoded by entropy encoding, and the original reference video frame is encoded by VVC.
[0046] In the embodiments of the present application, in the encoding stage, after the original target video frame to be encoded is obtained, the resolution of the original target video frame is unified by a resolution adjustment operation, and the original target video frame is converted into an adjusted target video frame having a fixed resolution (the first preset resolution). Therefore, even if the original target video frame has various resolutions, a video frame having a fixed resolution is finally input to the feature extraction network. In this way, it is not necessary to train multiple feature extraction networks for different resolutions, and only one feature extraction network corresponding to the first preset resolution (a feature extraction network for extracting features from a video frame having the first preset resolution) is required to realize the encoding of multiple original target video frames having different resolutions, which has broader applications and higher versatility. At the same time, the operation is simpler and more convenient.
[0047] The video encoding method provided in the first embodiment of the present application can be executed by a video encoding terminal (encoder) to encode video files with different resolutions, particularly face video files, so as to compress the digital bandwidth of the video file. The video encoding method provided in the first embodiment of the present application can be applied to various different scenarios such as normal storage and streaming of video games at various resolutions including faces. In particular, the video encoding method provided in the embodiments of the present application can be used to encode game video frames to form corresponding video bitstreams for storage and transmission in video streaming services or other similar applications. Another example includes low-latency scenarios such as video conferencing and live video broadcasts. In particular, the video encoding method provided in the embodiments of the present application can be used to encode face video data with various resolutions collected by a video acquisition device to form corresponding video bitstreams, and the corresponding video bitstreams are transmitted to a conference terminal, and the conference terminal decodes the video bitstreams to obtain corresponding face video images. Another example includes virtual reality scenarios. The face video encoding method provided in the embodiments of the present application can be used to encode face video data with various resolutions collected by a video acquisition device to form corresponding video bitstreams, and the corresponding video bitstreams are transmitted to devices related to virtual reality (such as VR virtual glasses), and the video bitstreams can be decoded by the VR device to obtain corresponding face video images. Based on face video images and the like, corresponding VR functions are realized.
[0048] Second Embodiment Referring to FIG. 4, FIG. 4 is a flowchart of a video decoding method according to the second embodiment of the present application. In particular, the video decoding method provided in this embodiment includes the following steps.
[0049] Step 402: Obtain and decode a video bitstream to obtain the original reference video frame and the target feature.
[0050] The target feature is obtained by extracting features from the adjusted target video frame, and the adjusted target video frame is a video frame with a first preset resolution obtained by adjusting the resolution of the original target video frame.
[0051] Step 404: Adjust the resolution of the original reference video frame to obtain an adjusted reference video frame with the first preset resolution.
[0052] In this step, the specific method of adjusting the resolution of the original reference video frame is the same as the specific method of adjusting the resolution of the original target video frame in step 204 of the first embodiment. In particular, the original reference video frame can be downsampled to obtain an adjusted reference video frame with the first preset resolution.
[0053] Step 406: Extract features from the adjusted reference video frame through a feature extraction network to obtain reference features.
[0054] In this step, the specific method of obtaining the reference features can refer to the specific method of obtaining the target features in step 206 of the first embodiment, which will not be repeated herein.
[0055] Step 408: Perform motion estimation based on the reference features and the target features through a motion estimation network to obtain a motion estimation result.
[0056] In particular, in one method, sparse motion estimation can be performed based on reference features and target features to obtain a sparse motion estimation map, and the obtained sparse motion estimation map is directly used as the motion estimation result. The sparse motion estimation map represents the relative motion relationship between the original reference video frame corresponding to the reference features and the original target video frame corresponding to the target features in a preset sparse feature region.
[0057] In another method, after obtaining the sparse motion estimation map, dense motion estimation is performed again to obtain a dense motion estimation map and a occlusion map as the final motion estimation result based on the sparse motion estimation map and the first reconstructed video frame generated by the first reference video frame. The dense motion estimation map represents the relative motion relationship between the original target video frame and the original reference video frame in a preset dense feature region. The occlusion map represents the degree to which each pixel of the original target video frame is occluded.
[0058] Compared with the above two methods, the former method has a simple calculation process, and thus has high calculation efficiency and can quickly obtain the motion estimation result. The latter method obtains the relative motion relationship between the original target video frame and the original reference video frame in a denser feature region, and the relative motion relationship is more accurate than the relative motion relationship represented by the sparse motion estimation map.
[0059] Step 410: Generate a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through the generation network.
[0060] Specifically, in order to obtain the deformation processing result, the generation network deforms the original reference video frame based on the motion estimation result obtained in step 408, and outputs a video frame reconstructed based on the deformation processing result.
[0061] In the embodiment of the present application, in the decoding stage, the resolution of the original reference video frame is also unified, and the original reference video frame is converted into an adjusted reference video frame having a fixed resolution, thereby obtaining reference features and target features with a fixed resolution. Therefore, the features finally input to the motion estimation network and the generation network are also features with a fixed resolution, whereby only one motion estimation network and one generation network are required to realize decoding in various scenarios having target video frames with different resolutions. In the embodiment of the present application, only one set of neural network models is required to perform the encoding operation and the decoding operation for the original target video frames with various resolutions, which has a wider range of applications and a simpler and more convenient operation process.
[0062] The video decoding method of the present embodiment can be executed by any suitable electronic device having a data capability, including but not limited to a server, a PC, etc.
[0063] The Third Embodiment Referring to FIG. 5, FIG. 5 is a flowchart of a video decoding method according to the third embodiment of the present application. Specifically, the video decoding method provided in the present embodiment includes the following steps.
[0064] Step 502: Obtain and decode a video bitstream to obtain an original reference video frame and target features.
[0065] Step 504: Adjust the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, and extract features from the adjusted reference video frame through a feature extraction network to obtain reference features.
[0066] In particular, for example, the original reference video frame can be downsampled to obtain an adjusted reference video frame having a first preset resolution, and features are extracted from the adjusted reference video frame through a feature extraction network to obtain reference features.
[0067] Step 506: Input the reference features and the target features into a motion estimation network, and perform motion estimation through the motion estimation network to obtain a first motion estimation result.
[0068] Step 508: Adjust the resolution of the original reference video frame to obtain an adjusted reference video frame having a second preset resolution.
[0069] In this step, the original reference video frame can be downsampled to obtain an adjusted reference video frame having a second preset resolution.
[0070] Optionally, in some embodiments, the adjusted reference video frame having a second preset resolution can be obtained by the following method, that is, determining a second target scaling coefficient based on the resolution of the original reference video frame, and scaling the original reference video frame using the second target scaling coefficient to obtain an adjusted reference video frame having a second preset resolution, and the second target scaling coefficient can be determined based on the size relationship between the resolution of the original reference video frame and the second preset resolution. In particular, the ratio of the resolution of the original reference video frame to the second preset resolution can be determined as the second target scaling coefficient.
[0071] Furthermore, regarding a situation where there may be multiple different resolutions of the original reference video frame, respective scaling factors corresponding to each resolution may be pre-calculated and placed within a sequence of second scaling factors. Thereafter, after the original reference video frame is obtained, a second target scaling factor corresponding to the resolution of the original reference video frame may be determined from a pre-set sequence of second scaling factors according to a pre-set correspondence between the resolution and the scaling factor.
[0072] Step 510: Input an adjusted reference video frame having a first motion estimation result and a second pre-set resolution into a generation network, and perform a deformation process on the adjusted reference video frame having the second pre-set resolution through the generation network to generate an intermediate reconstructed video frame having the second pre-set resolution.
[0073] The second pre-set resolution in step 508 above is configured according to the first pre-set resolution and the structural parameters of the generation network.
[0074] In particular, the first motion estimation result obtained in step 506 has a first preset resolution. Further, the generation network generally includes a downsampling sub-network, a deformation sub-network, and an upsampling sub-network. In this step, the specific operations performed through the generation network are, first, to downsample the downsampled reference video frame by the internal downsampling sub-network to obtain a second downsampled reference frame, then, to deform the second downsampled reference frame through the deformation sub-network with reference to the first motion estimation result to obtain a deformed reference frame, and finally, to upsample the deformed reference frame by the upsampling sub-network to output the result. In order to perform the deformation process smoothly, the resolutions of the second downsampled reference frame and the first motion estimation result need to match each other, that is, the second downsampled reference frame and the first motion estimation result have the same resolution (in the embodiments of the present application, both have the first preset resolution). Therefore, in the embodiments of the present application, when setting the second preset resolution, the second downsampled reference frame obtained after the downsampled reference video frame with the second preset resolution is downsampled by the downsampling sub-network can match the resolution of the first motion estimation result, and both of them are the first preset resolution.
[0075] Step 512: Adjust the resolution of the intermediate reconstructed video frame to obtain a reconstructed video frame having the same resolution as the original reference video frame.
[0076] The intermediate reconstructed video frame generated through the generation network has a second preset resolution. If the second preset resolution is obtained by downsampling the original reference video frame, in this step, in order to obtain a reconstructed video frame having the same resolution as the original reference video frame, contrary to step 508, it is necessary to perform an upsampling operation on the intermediate reconstructed video frame.
[0077] In particular, the upsampling in this step can be performed in the following manner, that is, determining the reciprocal of the second target scaling factor in step 508 above as the third target scaling factor, and using the third target scaling factor to downsample the intermediate reconstructed video frame in order to obtain a reconstructed video frame having the same resolution as the original reference video frame.
[0078] Referring to FIG. 6, FIG. 6 is a schematic diagram of a scenario corresponding to the third embodiment of the present application. In connection with the schematic diagram shown in FIG. 6, an example of a specific scenario for explaining the embodiment of the present application is given below.
[0079] Decoding a video bitstream to obtain an original reference video frame having a resolution of W×H and target features, downsampling the original reference video frame to obtain an adjusted reference video frame having W1×H1 (the first preset resolution), obtaining corresponding reference features through a feature extraction network, inputting the reference features and the target features into a motion estimation network to obtain a first motion estimation result, and at the same time, a sequence of second scaling factors x = {x1, x2, x3, ..., x for downsampling the original reference video frame to obtain an adjusted reference video frame having a second preset resolution (not shown) 1nDetermine a second target scaling factor from}, and based on the adjusted reference video frame having a second preset resolution and the first motion estimation result through the generation network, obtain an intermediate reconstructed video frame having the second preset resolution, and a sequence of third scaling factors for upsampling the intermediate reconstructed video frame to obtain a reconstructed video frame having a resolution of W×H [Number] Determine a third target scaling factor from.
[0080] In the embodiments of the present application, in the decoding stage, the resolution of the original reference video frame is also unified, and the original reference video frame is converted into an adjusted reference video frame having a fixed resolution, thereby obtaining reference features and target features of the fixed resolution. Therefore, the features finally input to the motion estimation network and the generation network are also features of the fixed resolution, and only one motion estimation network and one generation network are required to realize decoding in a scenario having a plurality of target video frames with different resolutions. In the embodiments of the present application, for original target video frames having various different resolutions, only one set of neural network models is required to perform the encoding operation and the decoding operation, which has a simpler and more convenient process for a wider range of applications and operations.
[0081] In addition, in the embodiments of the present application, the resolution adjustment process (upsampling process and downsampling process) is performed on the video frame, that is, it is performed in the image domain rather than the feature domain. Therefore, it is beneficial for each network in the neural network model to learn correct motion information, etc., thereby improving the quality of video frame reconstruction.
[0082] The video decoding method of this embodiment can be executed by any suitable electronic device having data capabilities, including but not limited to servers, PCs, etc.
[0083] The Fourth Embodiment Referring to FIG. 7, FIG. 7 is a flowchart of a video decoding method according to the fourth embodiment of the present application. In particular, the video decoding method provided in this embodiment includes the following steps.
[0084] Step 702: Obtain and decode a video bitstream to obtain an original reference video frame and target features.
[0085] Step 704: Adjust the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, and perform feature extraction on the adjusted reference video frame through a feature extraction network to obtain reference features.
[0086] In particular, in this step, the original reference video frame can be downsampled to obtain an adjusted reference video frame having a first preset resolution.
[0087] Step 706: Perform resolution adjustment on the reference features and target features to obtain adjusted reference features and adjusted target features.
[0088] If the original reference video frame is downsampled in step 704, in this step, the reference features and target features can be upsample accordingly to obtain adjusted reference features and adjusted target features.
[0089] Step 708: Input the adjusted reference features and adjusted target features into a motion estimation network, and perform motion estimation through the motion estimation network to obtain a second motion estimation result.
[0090] Step 710: Input the second motion estimation result and the original reference video frame into the generation network, and deform the original reference video frame through the generation network to generate a reconstructed video frame having the same resolution as the original reference video frame.
[0091] The scaling factor (sampling factor) used in the resolution adjustment in step 706 above is set according to the resolution of the original reference video frame, the structural parameters of the motion estimation network, and the structural parameters of the generation network. In particular, similar to step 510, in order to smoothly execute the deformation process in the generation network, after motion estimation is performed on the adjusted reference features and the adjusted target features obtained according to the sampling factor through the motion estimation network, the obtained second motion estimation result is downsampled by the downsampling subnetwork and may have the same resolution as the resolution of the original reference video frame.
[0092] Referring to FIG. 8, FIG. 8 is a schematic diagram of a scenario corresponding to the fourth embodiment of the present application. Examples of specific scenarios will be described below with reference to the schematic diagram shown in FIG. 8 according to the embodiments of the present application.
[0093] Decoding the video bitstream to obtain the original reference video frame having a resolution of W×H and the target features, downsampling the original reference video frame to obtain an adjusted reference video frame having W1×H1 (the first preset resolution), obtaining the corresponding reference features through the feature extraction network, and a sequence of scaling factors x = {x1, x2, x3, ..., x for upsampling the reference features and the target features to obtain the adjusted reference features and the adjusted target features 1nDetermine the target scaling factor from {}, obtain a second motion estimation result through the motion estimation network, and through the generation network, generate a reconstructed video frame having the same resolution as the original reference video frame based on the second motion estimation result and the original reference video frame.
[0094] In the embodiment of the present application, in the decoding stage, the resolution of the original reference video frame is also unified, and the original reference video frame is converted into an adjusted reference video frame having a fixed resolution, thereby obtaining reference features and target features of the fixed resolution. Therefore, the features finally input into the motion estimation network and the generation network are also features of the fixed resolution, and only one motion estimation network and one generation network are required to realize decoding in a scenario having a plurality of target video frames with different resolutions. In the embodiment of the present application, for original target video frames having various different resolutions, only one set of neural network models is required to perform the encoding operation and the decoding operation, which has a simpler and more convenient process for a wider range of applications and operations.
[0095] In addition, in the embodiment of the present application, when the generation network outputs a result, the result is not upsampled or downsampled. Therefore, visual artifacts in the final reconstructed video frame can be effectively prevented.
[0096] The video decoding method of this embodiment can be executed by any suitable electronic device having data capabilities, including but not limited to servers, PCs, etc.
[0097] The Fifth Embodiment Referring to FIG. 9, FIG. 9 is a flowchart of a video decoding method according to the fifth embodiment of the present application. In particular, the video decoding method provided in this embodiment includes the following steps.
[0098] Step 902: Obtain and decode a video bitstream to obtain an original reference video frame and target features.
[0099] Step 904: Adjust the resolution of the original reference video frame to obtain an adjusted reference video frame with a first preset resolution, and perform feature extraction on the adjusted reference video frame through a feature extraction network to obtain reference features.
[0100] In particular, in this step, the original reference video frame can be downsampled to obtain an adjusted reference video frame with a first preset resolution.
[0101] Step 906: Input the reference features and target features into a motion estimation network, and perform motion estimation through the motion estimation network to obtain a first motion estimation result.
[0102] Step 908: Adjust the resolution of the first motion estimation result to obtain a third motion estimation result.
[0103] If the original reference video frame is downsampled in step 904, correspondingly, in this step, the first motion estimation result can be upsampled to obtain a third motion estimation result.
[0104] Step 910: Input the third motion estimation result and the original reference video frame into a generation network, and deform the original reference video frame through the generation network to generate a reconstructed video frame with the same resolution as the original reference video frame.
[0105] The sampling coefficient used in the resolution adjustment in step 908 above is set according to the resolution of the original reference video frame and the structural parameters of the generation network. In particular, in order to smoothly execute the deformation process in the generation network, the resolution of the third motion estimation result can (may) coincide with (be equal to) the resolution of the original reference video frame downsampled by the downsampling subnetwork.
[0106] Referring to FIG. 10, FIG. 10 is a schematic diagram of a scenario corresponding to the fifth embodiment of the present application. Examples of specific scenarios will be described below with reference to the schematic diagram shown in FIG. 10 according to the embodiments of the present application.
[0107] Decoding the video bitstream to obtain the original reference video frame with a resolution of W×H and the target feature, downsampling the original reference video frame to obtain an adjusted reference video frame with W1×H1 (the first preset resolution), obtaining the corresponding reference feature through the feature extraction network, performing motion estimation on the reference feature and the target feature to obtain the first motion estimation result with the first preset resolution, and obtaining a sequence of scaling coefficients x = {x1, x2, x3, ..., x 1n} to determine the target scaling coefficient, and generating a reconstructed video frame having the same resolution as the original reference video frame based on the third motion estimation result and the original reference video frame through the generation network.
[0108] In the embodiments of the present application, in the decoding stage, the resolution of the original reference video frame is also unified, and the original reference video frame is converted into an adjusted reference video frame with a fixed resolution, thereby obtaining reference features and target features with a fixed resolution. Therefore, the features finally input to the motion estimation network and the generation network are also features with a fixed resolution, and only one motion estimation network and one generation network are required to realize decoding in a scenario having a plurality of target video frames with different resolutions. In the embodiments of the present application, for original target video frames having various different resolutions, only one set of neural network models is required to perform the encoding operation and the decoding operation, which has a simpler and more convenient process for a wider range of applications and operations.
[0109] In addition, in the embodiments of the present application, when the generation network outputs a result, resolution adjustment (upsampling or downsampling) is not performed on the result. Therefore, visual artifacts in the final reconstructed video frame can be effectively prevented.
[0110] The video decoding method of this embodiment can be executed by any suitable electronic device having data capabilities, including but not limited to servers, PCs, etc.
[0111] Sixth Embodiment Referring to FIG. 11, FIG. 11 is a flowchart of a video decoding method according to the sixth embodiment of the present application. In particular, the video decoding method provided in this embodiment includes the following steps.
[0112] Step 1102: Obtain and decode a video bitstream to obtain an original reference video frame and target features.
[0113] Step 1104: Adjust the resolution of the original reference video frame to obtain an adjusted reference video frame with a first preset resolution, and perform feature extraction on the downsampled reference video frame through a feature extraction network to obtain reference features.
[0114] In particular, in this step, the original reference video frame can be downsampled to obtain an adjusted reference video frame with a first preset resolution.
[0115] Step 1106: Input the reference features and target features into a motion estimation network, and perform motion estimation through the motion estimation network to obtain a first motion estimation result.
[0116] Step 1108: Input the original reference video frame and the first motion estimation result into a generation network, downsample the original reference video frame by a downsampling subnetwork to obtain a first downsampled reference frame, downsample the first downsampled reference frame by a downsampling layer to obtain a second downsampled reference frame, deform the second downsampled reference frame by a deformation subnetwork to obtain a deformed reference frame, upsample the deformed reference frame by an upsampling layer to obtain a first upsampled deformed frame, and upsample the first upsampled deformed frame by an upsampling subnetwork to obtain a reconstructed video frame with the same resolution as the original reference video frame.
[0117] In particular, the sampling coefficient used by the downsampling layer when downsampling the first downsampled reference frame is the reciprocal of the sampling coefficient used by the upsampling layer when upsampling the first upsampled and deformed frame. In other words, if the sampling coefficient used by the downsampling layer when downsampling the first downsampled reference frame is x1, the sampling coefficient used by the upsampling layer when upsampling the first upsampled and deformed frame is 1 / x1.
[0118] The sampling coefficient used by the downsampling layer to downsample the first downsampled reference frame is set according to the resolution of the original reference video frame, the first preset resolution, and the structural parameters of the generation network. In particular, after the original reference video frame is input into the generation network, the resolution of the second downsampled reference frame finally output by the downsampling layer is the first preset resolution.
[0119] Referring to FIG. 12, FIG. 12 is a schematic diagram of a scenario corresponding to the sixth embodiment of the present application. Examples of specific scenarios will be described below with reference to the schematic diagram shown in FIG. 12 for the embodiments of the present application.
[0120] Decode the video bitstream to obtain the original reference video frame with a resolution of W×H and the target features, downsample the original reference video frame to obtain an adjusted reference video frame with W1×H1 (the first preset resolution), obtain the corresponding reference features through the feature extraction network, perform motion estimation on the reference features and the target features to obtain the first motion estimation result at the first preset resolution, input the original reference video frame and the first motion estimation result into the generation network, and after passing through the downsampling subnetwork, use the target scaling factor determined from the sequence
Number
[0121] In the embodiments of the present application, in the decoding stage, the resolution of the original reference video frame is also unified, and the original reference video frame is converted into an adjusted reference video frame having a fixed resolution, thereby obtaining reference features and target features of the fixed resolution. Therefore, the features finally input to the motion estimation network and the generation network are also features of the fixed resolution, and only one motion estimation network and one generation network are required to realize decoding in a scenario having a plurality of target video frames with different resolutions. In the embodiments of the present application, for the original target video frames having various different resolutions, only one set of neural network models is required to perform the encoding operation and the decoding operation, which has a simpler and more convenient process for a wider range of applications and operations.
[0122] The video decoding method of this embodiment can be executed by any suitable electronic device having data capabilities, including but not limited to servers, PCs, etc.
[0123] The Seventh Embodiment Referring to FIG. 13, FIG. 13 is a flowchart of a video encoding method according to the seventh embodiment of the present application. In particular, the video encoding method provided in this embodiment includes the following steps.
[0124] Step 1302: Obtain the original reference video frame and the original target video frame to be encoded.
[0125] Step 1304: Adjust the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and perform feature extraction on the adjusted target video frame through a feature extraction network corresponding to the first preset resolution to obtain target features.
[0126] Step 1306: The decoding terminal device decodes the video bitstream to obtain the original reference video frame and the target feature, adjusts the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, performs feature extraction on the adjusted reference video frame through a feature extraction network to obtain a reference feature, performs motion estimation based on the reference feature and the target feature through a motion estimation network to obtain a motion estimation result, and generates a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through a generation network. To enable this, the original reference video frame and the target feature are each encoded to obtain the video bitstream.
[0127] In the video encoding method of this embodiment, the details of the implementation of each step can be referred to the corresponding steps of any one of the second to sixth embodiments described above, and will not be repeated herein.
[0128] According to the video encoding method provided by the embodiments of the present application, in the encoding stage, after the original target video frame to be encoded is obtained, the resolution of the original target video frame is unified by a resolution adjustment operation, and the original target video frame is converted into an adjusted target video frame having a fixed resolution (a first preset resolution). Therefore, even if the original target video frame has various different resolutions, a video frame having a fixed resolution is finally input to the feature extraction network. In this way, it is not necessary to train a plurality of feature extraction networks for different resolutions, but only one feature extraction network corresponding to the first preset resolution (a feature extraction network for extracting features from a video frame having the first preset resolution) is required to realize the encoding of a plurality of original target video frames having different resolutions, which has a wider range of applications and higher versatility, and at the same time has a simpler and more convenient operation. In addition, correspondingly, in the decoding stage, the resolution of the original reference video frame is also unified, and the original reference video frame is converted into an adjusted reference video frame having a fixed resolution, thereby obtaining reference features and target features of a fixed resolution. Therefore, the features finally input to the motion estimation network and the generation network are also features of a fixed resolution, and only one motion estimation network and one generation network are required to realize decoding in a scenario having target video frames with various different resolutions. In summary, in the embodiments of the present application, only one set of neural network models is required for the encoding operation and the decoding operation for the original target video frames having various resolutions, which has a wider range of applications and a simpler and more convenient operation process.
[0129] The Eighth Embodiment Referring to FIG. 14, FIG. 14 is a flowchart of a video encoding method according to the eighth embodiment of the present application. The application scenario of the video encoding method is that a video acquisition device obtains a conference video and performs video encoding using the video encoding method provided in this embodiment to form a corresponding video bitstream, and the video bitstream is transmitted to a conference terminal, and the conference terminal decodes the video bitstream to obtain a corresponding conference video screen for display.
[0130] In particular, the video encoding method provided in this embodiment includes the following steps.
[0131] Step 1402: Obtain the original video clip captured by the video acquisition device.
[0132] Step 1404: Determine the original reference video frame and the original target video frame to be encoded from the original video clip.
[0133] Step 1406: Adjust the resolution of the original target video frame to obtain an adjusted target video frame with a first preset resolution, and perform feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution.
[0134] Step 1408: Encode the original reference video frame and the target features respectively to obtain a video bitstream.
[0135] Step 1410: The conference terminal device performs reconstruction of the video frame based on the video bitstream, generates a reconstructed video frame having the same resolution as the original target video frame, and transmits the video bitstream to the conference terminal device to display the reconstructed video frame.
[0136] The Ninth Embodiment Referring to FIG. 15, FIG. 15 is a structural block diagram of a video encoding device according to the ninth embodiment of the present application. The video encoding device provided in the embodiment of the present application includes an original video frame acquisition module 1502 configured to obtain an original reference video frame and an original target video frame to be encoded, a target feature acquisition module 1504 configured to adjust the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and perform feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution, a bitstream acquisition module 1506 configured to encode the original reference video frame and the target features respectively to obtain a video bitstream, and perform video frame reconstruction based on the video bitstream to generate a reconstructed video frame having the same resolution as the original target video frame.
[0137] Optionally, in some embodiments, when adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, the target feature acquisition module 1504 determines a first target scaling factor based on the resolution of the original target video frame, and is specifically configured to scale the original target video frame using the first target scaling factor to obtain an adjusted target video frame having a first preset resolution.
[0138] Optionally, in some embodiments of the implementation, when determining the first target scaling factor based on the resolution of the original target video frame, the target feature acquisition module 1504 It is specifically configured to determine a first target scaling factor corresponding to the resolution of the original target video frame from a sequence of preset first scaling factors according to a preset correspondence between the resolution and the scaling factor.
[0139] The video encoding device of this embodiment is used to implement the corresponding video encoding method of the above-mentioned embodiments of the plurality of methods, and has the beneficial effects of the corresponding method embodiments, which will not be repeated herein. In addition, for the implementation of the functions of each module of the video encoding device of this embodiment, reference can be made to the description of the corresponding parts of the above-mentioned method embodiments, which will not be repeated herein.
[0140] The 10th embodiment Referring to FIG. 16, FIG. 16 is a structural block diagram of a video decoding device according to the 10th embodiment of the present application. The video decoding device provided in the embodiments of the present application includes a decoding module 1602 configured to obtain and decode a video bitstream to obtain an original reference video frame and target features; a reference feature acquisition module 1604 configured to adjust the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, and perform feature extraction on the adjusted reference video frame to obtain reference features through a feature extraction network; a motion estimation module 1606 configured to perform motion estimation based on the reference features and the target features to obtain a motion estimation result through a motion estimation network; and a generation module 1608 configured to generate a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through a generation network.
[0141] Optionally, in some embodiments, the motion estimation module 1606 The reference features and the target features are input into a motion estimation network, and are specifically configured to perform motion estimation through the motion estimation network in order to obtain a first motion estimation result.
[0142] The generation module 1608 adjusts the resolution of the original reference video frame in order to obtain an adjusted reference video frame having a second preset resolution. The first motion estimation result and the adjusted reference video frame having the second preset resolution are input into a generation network, and a deformation process is performed on the adjusted reference video frame having the second preset resolution through the generation network in order to generate an intermediate reconstructed video frame having the second preset resolution. It is specifically configured to adjust the resolution of the intermediate reconstructed video frame in order to obtain a reconstructed video frame having the same resolution as the target video frame.
[0143] Optionally, in some embodiments, the motion estimation module 1606 adjusts the resolution of the reference features and the target features in order to obtain adjusted reference features and adjusted target features. The adjusted reference features and the adjusted target features are input into a motion estimation network, and are specifically configured to perform motion estimation through the motion estimation network in order to obtain a second motion estimation result.
[0144] The generation module 1608 is specifically configured to input the second motion estimation result and the original reference video frame into a generation network, and perform a deformation process on the original reference video frame through the generation network in order to generate a reconstructed video frame having the same resolution as the target video frame.
[0145] Optionally, in some embodiments, the motion estimation module 1606 The reference features and the target features are input into a motion estimation network, and are specifically configured to perform motion estimation through the motion estimation network in order to obtain a first motion estimation result.
[0146] The generation module 1608 In order to obtain a third motion estimation result, adjust the resolution of the first motion estimation result, Input the third motion estimation result and the original reference video frame into a generation network, and is specifically configured to perform a deformation process on the original reference video frame through the generation network in order to generate a reconstructed video frame having the same resolution as the target video frame.
[0147] Optionally, in some embodiments, the generation module includes a downsampling sub-network, a downsampling layer, a deformation sub-network, an upsampling layer, and an upsampling sub-network.
[0148] The motion estimation module 1606 is specifically configured to input the reference features and the target features into a motion estimation network and perform motion estimation through the motion estimation network in order to obtain a first motion estimation result.
[0149] The generation module 1608 inputs the original reference video frame and the first motion estimation result into the generation network, downsamples the original reference video frame by the downsampling sub-network to obtain the first downsampled reference frame, downsamples the first downsampled reference frame by the downsampling layer to obtain the second downsampled reference frame, deforms the second downsampled reference frame by the deformation sub-network to obtain the deformed reference frame, upsamples the deformed reference frame by the upsampling layer to obtain the first upsampled deformed frame, and is specifically configured to upsample the first upsampled deformed frame by the upsampling sub-network to obtain a reconstructed video frame having the same resolution as the original reference video frame.
[0150] The video decoding apparatus of this embodiment is used to implement the corresponding video decoding method of the above-described multiple method embodiments, has the beneficial effects of the corresponding method embodiments, and they will not be repeated herein. In addition, the implementation of the functions of each module of the video decoding apparatus of this embodiment can refer to the description of the corresponding parts of the above-described method embodiments, which will not be repeated herein.
[0151] The 11th embodiment FIG. 17 shows a schematic structural diagram of an electronic device according to the 11th embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the electronic device.
[0152] As shown in FIG. 17, the conference terminal may include a processor 1702, a communication interface 1704, a memory 1706, and a communication bus 1708.
[0153] The processor 1702, the communication interface 1704, and the memory 1706 communicate with each other through the communication bus 1708.
[0154] The communication interface 1704 is configured to communicate with another electronic device or server.
[0155] The processor 1702 is configured to execute a program 1710 that can particularly execute the related steps of the video encoding method or the video decoding method embodiments described above.
[0156] In particular, the program 1710 may include program code, and the program code includes computer operation instructions.
[0157] The processor 1702 may be a CPU configured to implement the embodiments of the present application, or an application specific integrated circuit (ASIC), or one or more integrated circuits. One or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0158] The memory 1706 is configured to store the program 1710. The memory 1706 may include high-speed RAM memory and may also include non-volatile memory such as at least one disk memory.
[0159] Program 1710 can be specifically configured to enable the processor 1702 to perform the following operations, namely, obtaining the original reference video frame and the original target video frame to be encoded, adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, extracting features from the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution, encoding the original reference video frame and the target features respectively to obtain a video bitstream, and performing video frame reconstruction based on the video bitstream to generate a reconstructed video frame having the same resolution as the original target video frame.
[0160] Alternatively, Program 1710 can be specifically configured to enable the processor 1702 to perform the following operations, namely, obtaining and decoding a video bitstream to obtain the original reference video frame and target features, adjusting the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, extracting features from the adjusted reference video frame through a feature extraction network to obtain reference features, performing motion estimation based on the reference features and the target features through a motion estimation network to obtain a motion estimation result, and generating a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through a generation network.
[0161] Alternatively, program 1710 enables the processor 1702 to perform the following operations: obtaining the original reference video frame and the original target video frame to be encoded; adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution; performing feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution; and enabling the decoding terminal device to decode the video bitstream to obtain the original reference video frame and the target features, adjusting the resolution of the original reference video frame to obtain an adjusted reference video frame having the first preset resolution, extracting the features of the adjusted reference video frame to obtain reference features through the feature extraction network, performing motion estimation based on the reference features and the target features through a motion estimation network to obtain a motion estimation result, and generating a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through a generation network. In particular, it can be configured to enable encoding the original reference video frame and the target features respectively to obtain the video bitstream.
[0162] Alternatively, the program 1710 enables the processor 1702 to perform the following operations: obtaining the original video clip captured by the video acquisition device; determining the original reference video frame and the original target video frame to be encoded from the original video clip; adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution; extracting the features of the adjusted target video frame through a feature extraction network corresponding to the first preset resolution to obtain target features; encoding the original reference video frame and the target features respectively to obtain a video bitstream; and transmitting the video bitstream to the conference terminal device to enable the conference terminal device to perform reconstruction of the video frame based on the video bitstream, and generate and display a reconstructed video frame having the same resolution as the original target video frame.
[0163] Details of the implementation of each step of the program 1710 can be referred to the corresponding descriptions of the corresponding steps and units of the video encoding method embodiment or the video decoding method embodiment described above, which will not be repeated herein. Those skilled in the art can clearly understand that for the convenience and simplification of the description, the specific working processes of the above-mentioned devices and modules can refer to the descriptions of the corresponding processes in the embodiments of the above-mentioned methods, and they will not be repeated herein.
[0164] By the electronic device of this embodiment, In the symbolization stage, after obtaining the original target video frame to be symbolized, the resolution of the original target video frame is unified by a resolution adjustment operation, and the original target video frame is converted into an adjusted target video frame having a fixed resolution (a first preset resolution). Therefore, even if the original target video frame has various different resolutions, a video frame having a fixed resolution is still input to the feature extraction network. Thus, it is not necessary to train a plurality of feature extraction networks for different resolutions, and only one feature extraction network corresponding to the first preset resolution (a feature extraction network for extracting features from a video frame having the first preset resolution) is required to realize the symbolization of a plurality of original target video frames having different resolutions, which has a wider range of applications and higher versatility. At the same time, the operation is simpler and more convenient. In addition, correspondingly, in the decoding stage, the resolution of the original reference video frame is also unified, and the original reference video frame is converted into an adjusted reference video frame having a fixed resolution, thereby obtaining reference features and target features of a fixed resolution. Therefore, features of a fixed resolution are finally input to the motion estimation network and the generation network, and only one motion estimation network and one generation network are required to realize decoding in various scenarios having target video frames with different resolutions. In summary, in the embodiments of the present application, only one set of neural network models is required to perform the symbolization operation and the decoding operation for original target video frames of various resolutions, which has a wider range of applications and a simpler and more convenient operation process.
[0165] Embodiments of the present application also provide a computer program product including computer instructions instructing a computing device to perform operations corresponding to any of the methods in the embodiments of the above-mentioned plurality of methods.
[0166] Depending on the implementation requirements, in order to achieve the objectives of the embodiments of the present application, it should be noted that various components / steps described in the embodiments of the present application can be divided into more components / steps, and two or more components / steps or parts of the operations of the components / steps can be combined into new components / steps.
[0167] The above-described method according to the embodiments of the present application can be implemented in hardware, firmware, or as software or computer code stored in a recording medium (such as CD ROM, RAM, floppy disk, hard disk, magneto-optical disk, etc.), or downloaded through a remote recording medium or network and implemented as computer code originally stored in a non-transitory machine-readable medium stored in a local recording medium. Therefore, the method described herein can be stored on a recording medium when such software processing is performed using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the video encoding method or video decoding method described herein is implemented. Further, when a general-purpose computer accesses the code for implementing the video encoding method or video decoding method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the video encoding method or video decoding method shown herein.
[0168] Those skilled in the art may recognize that the units and method steps of each example described in the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed by hardware or by software depends on the specific application of the technical solution and the design constraints. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered outside the scope of the embodiments of this application.
[0169] The above method of implementation is only used to explain the embodiments of this application and is not a limitation to the embodiments of this application. Also, those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of this application, and the scope of patent protection of the embodiments of this application should be defined by the scope of the claims.
Explanation of Reference Signs
[0170] 1502 Original Video Frame Acquisition Module 1504 Target Feature Acquisition Module, Target Feature Obtaining Module 1506 Bitstream Acquisition Module 1602 Decoding Module 1604 Reference Feature Acquisition Module 1606 Motion Estimation Module 1608 Generation Module 1702 Processor 1704 Communication Interface 1706 Memory 1708 Communication Bus 1710 Program
Claims
1. obtaining an original reference video frame and an original target video frame to be encoded; adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and performing feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution; encoding the original reference video frame and the target features respectively to obtain a video bitstream, and performing video frame reconstruction based on the video bitstream to generate a reconstructed video frame having the same resolution as the original target video frame A video encoding method comprising the steps of:
2. The step of adjusting the resolution of the original target video frame to obtain the adjusted target video frame having the first preset resolution comprises: determining a first target scaling factor based on the resolution of the original target video frame; scaling the original target video frame using the first target scaling factor to obtain the adjusted target video frame having the first preset resolution The video encoding method according to claim 1, comprising the steps of:
3. The step of determining the first target scaling factor based on the resolution of the original target video frame comprises: determining the first target scaling factor corresponding to the resolution of the original target video frame from a sequence of preset first scaling factors according to a preset correspondence between the resolution and the scaling factor The video encoding method according to claim 2, comprising the steps of:
4. obtaining and decoding a video bitstream to obtain an original reference video frame and target features; Adjusting the resolution of the original reference video frame to obtain an adjusted reference video frame having a first preset resolution, and extracting features from the adjusted reference video frame through a feature extraction network to obtain reference features; Performing motion estimation based on the reference features and the target features to obtain a motion estimation result through a motion estimation network; Generating a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through a generation network A video decoding method including.
5. The step of performing motion estimation based on the reference features and the target features to obtain the motion estimation result through the motion estimation network includes Inputting the reference features and the target features into the motion estimation network and performing motion estimation through the motion estimation network to obtain a first motion estimation result; The step of generating the reconstructed video frame having the same resolution as the target video frame based on the motion estimation result and the original reference video frame through a generation network includes Adjusting the resolution of the original reference video frame to obtain an adjusted reference video frame having a second preset resolution; Inputting the first motion estimation result and the adjusted reference video frame having the second preset resolution into the generation network, performing a deformation process on the adjusted reference video frame having the second preset resolution through the generation network, and generating an intermediate reconstructed video frame having the second preset resolution; Adjusting the resolution of the intermediate reconstructed video frame to obtain a reconstructed video frame having the same resolution as the target video frame The video decoding method according to claim 4, including.
6. The step of performing motion estimation based on the reference features and the target features to obtain the motion estimation result through the motion estimation network includes To obtain the adjusted reference feature and the adjusted target feature, the step of adjusting the resolution of the reference feature and the target feature; Inputting the adjusted reference feature and the adjusted target feature into the motion estimation network, executing the motion estimation through the motion estimation network, and obtaining a second motion estimation result; comprising; The step of generating the reconstructed video frame having the same resolution as the target video frame based on the motion estimation result and the original reference video frame through the generation network; Inputting the second motion estimation result and the original reference video frame into the generation network, executing a deformation process on the original reference video frame through the generation network, and generating the reconstructed video frame having the same resolution as the target video frame; The video decoding method according to claim 4, comprising.
7. The step of executing the motion estimation based on the reference feature and the target feature to obtain the motion estimation result through the motion estimation network; Inputting the reference feature and the target feature into the motion estimation network, and executing the motion estimation through the motion estimation network to obtain the first motion estimation result; The step of generating the reconstructed video frame having the same resolution as the target video frame based on the motion estimation result and the original reference video frame through the generation network; The step of adjusting the resolution of the first motion estimation result to obtain a third motion estimation result; Inputting the third motion estimation result and the original reference video frame into the generation network, executing a deformation process on the original reference video frame through the generation network, and generating the reconstructed video frame having the same resolution as the target video frame; The video decoding method according to claim 4, comprising.
8. The generation network includes a downsampling subnet, a downsampling layer, a deformation subnet, an upsampling layer, and an upsampling subnet; To obtain the motion estimation result through the motion estimation network, the step of performing motion estimation based on the reference feature and the target feature includes: inputting the reference feature and the target feature into the motion estimation network, and performing motion estimation through the motion estimation network to obtain a first motion estimation result; generating, through a generation network, the reconstructed video frame having the same resolution as the target video frame based on the motion estimation result and the original reference video frame includes: inputting the original reference video frame and the first motion estimation result into the generation network, downsampling the original reference video frame by the downsampling sub-network to obtain a first downsampled reference frame, downsampling the first downsampled reference frame by the downsampling layer to obtain a second downsampled reference frame, deforming the second downsampled reference frame by the deformation sub-network to obtain a deformed reference frame, upsampling the deformed reference frame by the upsampling layer to obtain a first upsampled deformed frame, and upsampling the first upsampled deformed frame by the upsampling sub-network to obtain the reconstructed video frame having the same resolution as the target video frame; The video decoding method according to claim 4, comprising: **Claim 9** obtaining an original reference video frame and an original target video frame to be encoded; adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and performing feature extraction on the adjusted target video frame to obtain target features through a feature extraction network corresponding to the first preset resolution; The decoding device decodes the video bitstream to obtain the original reference video frame and the target feature, adjusts the resolution of the original reference video frame to obtain an adjusted reference video frame having the first preset resolution, extracts the features of the adjusted reference video frame to obtain reference features through the feature extraction network, performs motion estimation based on the reference features and the target feature through a motion estimation network to obtain a motion estimation result, and generates a reconstructed video frame having the same resolution as the original reference video frame based on the motion estimation result and the original reference video frame through a generation network. To enable this, the step of encoding the original reference video frame and the target feature respectively to obtain the video bitstream A video encoding method, comprising. Claim 10 The step of obtaining an original video clip captured by a video acquisition device, The step of determining, from the original video clip, an original reference video frame and an original target video frame to be encoded, The step of adjusting the resolution of the original target video frame to obtain an adjusted target video frame having a first preset resolution, and the step of extracting the features of the adjusted target video frame through a feature extraction network corresponding to the first preset resolution to obtain target features, The step of encoding the original reference video frame and the target feature respectively to obtain a video bitstream, The step of causing a conference terminal device to perform reconstruction of a video frame based on the video bitstream, and transmitting the video bitstream to the conference terminal device to generate and display a reconstructed video frame having the same resolution as the original target video frame A video encoding method, comprising. Claim 11 An electronic device comprising a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus An electronic device, wherein the memory is configured to store at least one executable instruction, and the executable instruction enables the processor to execute an operation corresponding to the video encoding method according to any one of claims 1 to 3, 9, or 10, or an operation corresponding to the video decoding method according to any one of claims 4 to 8.
12. A computer storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the video encoding method according to any one of claims 1 to 3, 9, or 10, or implements the video decoding method according to any one of claims 4 to 8.
13. A computer program product including computer instructions, wherein the computer instructions direct a computing device to execute an operation corresponding to the video encoding method according to any one of claims 1 to 3, 9, or 10, or to execute an operation corresponding to the video decoding method according to any one of claims 4 to 8.