Video encoding method and system

Through a learning-based spatial scalable video coding method, a neural network is used to extract spatial features and predict motion information, and inter-layer information is combined to generate multi-scale mixed context, which solves the problem of low video coding performance in existing technologies and achieves high-quality and high-resolution video coding in low-bandwidth environments.

CN117939146BActive Publication Date: 2025-09-19UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites -1 Cited by

Patent Information

Application Number
CN202410105013.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-09-19
Estimated Expiration
2044-01-25

Smart Images

  • Figure CN117939146B_ABST
    Figure CN117939146B_ABST
Patent Text Reader

Abstract

The present application provides a learning-based spatial scalable video coding method and system. In the above method, the base layer information of the base layer video coding frame with lower resolution in the target coded video frame is obtained to obtain the first inter-layer information. Finally, based on the target frame to be coded and the reconstructed video frame of the previous frame, the encoding of the video frame can be completed to obtain the first code stream. Since the encoding resolution of the first inter-layer information is the same as the resolution of the enhancement layer code stream, the encoding of the high-resolution video frame in the enhancement layer can be combined with the first inter-layer information in the base layer video coding frame for encoding, thereby improving the performance of video frame encoding. At the same time, the encoding of the video frame will be further combined with the reconstructed video frame in the previous video frame for encoding, so that the video frame is encoded by the mixed use of inter-frame information and inter-layer information, and the video encoding performance is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video coding technology, and in particular to a video coding method and system. Background Art

[0002] Video coding is the process of compressing video signals into digital data for storage, transmission, and processing. During video coding, it is often necessary to encode data at different resolutions for the same video into a more compact bitstream, thereby reducing the cost of data transmission and storage. Existing video coding schemes for different resolutions are typically layer-based scalable video coding schemes. In these layer-based coding schemes, the bitstream is defined as a base layer and multiple enhancement layers. The base layer provides basic global video information, while the enhancement layers provide additional coding information to achieve higher video quality. In coding schemes that compress video data into bitstreams, the scalability of video data is primarily reflected in three dimensions: temporal, quality, and spatial. Spatial scalability lies in the ability to encode video data into corresponding bitstream data at different resolution levels.

[0003] In current spatial domain-based video coding schemes, video coding is usually performed based on traditional video coding standards. Video coding standards are a set of specifications and algorithms used to digitally compress and encode video signals so that they occupy less bandwidth and storage space during storage and transmission while maintaining high video quality. Existing spatial domain-based video coding schemes usually perform video encoding based on traditional video coding standards such as H.264 / AVC and H.265 / HEVC. Although this method has good interoperability, traditional video coding standards need to ensure a high bit rate when processing complex scenes with multiple resolutions and high frame rates. Therefore, they cannot meet the processing requirements of high quality and high resolution in low-bandwidth environments, and the encoding performance of video data is relatively low.

[0004] Therefore, how to solve the problem of low video data encoding performance in the prior art has become a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0005] Based on the above problems, in order to solve the problem of low video data encoding performance in the prior art, the present application provides a video encoding method and system.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] In a first aspect, the present application discloses a learning-based spatial scalable video coding method, which is applied to a preset neural network. The method comprises:

[0008] Obtaining a target coded video frame; the target coded video frame includes a base layer video coded frame and an enhancement layer video coded frame; the coding resolution of the base layer video coded frame is smaller than the coding resolution of the enhancement layer video coded frame; the base layer video coded frame and the enhancement layer video coded frame are in the same time domain;

[0009] Acquiring base layer information from the base layer video frame to obtain first inter-layer information; the coding resolution of the first inter-layer information is the same as the coding resolution of the enhancement layer coded video frame; the first inter-layer information includes: spatial features, predicted motion information, and layer prior information;

[0010] Get the previous frame to reconstruct the video frame;

[0011] The target coded video frame is encoded according to the previous reconstructed video frame and the first inter-layer information to obtain a first code stream.

[0012] Optionally, encoding the target coded video frame according to the previous reconstructed video frame and the first inter-layer information to obtain a first bitstream specifically includes:

[0013] Performing encoding and decoding reconstruction according to the target coded video frame, the previous reconstructed video frame, and the predicted motion information to obtain reconstructed high-resolution motion information;

[0014] Performing context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal domain features in the previous reconstructed video frame to generate a multi-scale mixed context;

[0015] The target coded video frame is encoded using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first code stream.

[0016] Optionally, performing encoding and decoding reconstruction according to the target coded video frame, the previous reconstructed video frame, and the predicted motion information to obtain reconstructed high-resolution motion information specifically includes:

[0017] Inputting the target coded video frame and the previous reconstructed video frame into a preset optical flow network to obtain high-resolution motion information; the coding resolution of the high-resolution motion information is the same as the coding resolution of the enhancement layer coded video frame;

[0018] Encoding is performed according to the predicted motion information and the high-resolution motion information to obtain a bit stream of the high-resolution motion information;

[0019] The code stream of the high-resolution motion information is decoded and reconstructed based on the predicted motion information to obtain the reconstructed high-resolution motion information.

[0020] Optionally, performing context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal features in the previous reconstructed video frame to generate a multi-scale mixed context specifically includes:

[0021] Determining multi-scale spatial domain features and multi-scale temporal domain features based on the spatial domain features and the temporal features in the last reconstructed video frame;

[0022] downsampling the reconstructed high-resolution motion information to obtain multi-scale motion information;

[0023] Performing motion compensation on the multi-scale temporal features based on the multi-scale motion information to obtain aligned multi-scale temporal features;

[0024] The multi-scale mixed context is generated according to the multi-scale spatial domain features and the aligned multi-scale temporal domain features.

[0025] Optionally, encoding the target coded video frame by using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first bitstream specifically includes:

[0026] Determining probability distribution parameters of the first bitstream according to the layer prior information and a preset inter-layer prior entropy model;

[0027] The target coded video frame is encoded based on the probability distribution parameter of the first code stream to obtain the first code stream.

[0028] Optionally, performing base layer information acquisition on the base layer video frame to obtain first inter-layer information specifically includes:

[0029] Performing encoding and decoding reconstruction on the base layer video coding frame to obtain second inter-layer information; the coding resolution of the second inter-layer information is lower than the coding resolution of the enhancement layer video coding frame;

[0030] performing domain transformation processing on the second inter-layer information to obtain transformed second inter-layer information;

[0031] The transformed second inter-layer information is up-sampled according to the coding resolution of the enhancement layer video coding frame to obtain the first inter-layer information.

[0032] Optionally, encoding according to the predicted motion information and the high-resolution motion information to obtain a bitstream of the high-resolution motion information specifically includes:

[0033] Determining probability distribution parameters of a bitstream of the high-resolution motion information based on the predicted motion information and a preset motion entropy model;

[0034] The code stream of the high-resolution motion information is determined according to the probability distribution parameter of the code stream of the high-resolution motion information.

[0035] Optionally, generating the multi-scale mixed context according to the spatial domain features and the aligned multi-scale temporal domain features specifically includes:

[0036] Constructing a feature weight map between the aligned multi-scale temporal features and the spatial features;

[0037] Based on the feature weight map, the aligned multi-scale time domain features and the spatial domain features are subjected to feature fusion to obtain multi-scale hybrid features;

[0038] The multi-scale mixed context is generated according to the multi-scale mixed feature.

[0039] In a second aspect, the present application discloses a learning-based spatial scalable video coding system, which is applied to a preset neural network. The system includes:

[0040] A first acquisition module is configured to acquire a target coded video frame; the target coded video frame includes a base layer video coded frame and an enhancement layer video coded frame; the coding resolution of the base layer video coded frame is lower than the coding resolution of the enhancement layer video coded frame; the base layer video coded frame and the enhancement layer video coded frame are in the same time domain;

[0041] an inter-layer information acquisition module, configured to acquire base layer information from the base layer video frame to obtain first inter-layer information; the coding resolution of the first inter-layer information is the same as the coding resolution of the enhancement layer coded video frame; the first inter-layer information includes: spatial features, predicted motion information, and layer prior information;

[0042] The second acquisition module is used to acquire the previous reconstructed video frame;

[0043] The encoding module is used to encode the target coded video frame according to the reconstructed video frame of the previous frame and the first inter-layer information to obtain a first code stream

[0044] Optionally, the encoding module is specifically used to:

[0045] Performing encoding and decoding reconstruction according to the target coded video frame, the previous reconstructed video frame, and the predicted motion information to obtain reconstructed high-resolution motion information;

[0046] Performing context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal domain features in the previous reconstructed video frame to generate a multi-scale mixed context;

[0047] The target coded video frame is encoded using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first code stream.

[0048] Compared to the prior art, the present application has the following advantageous effects: The present application provides a learning-based spatial scalable video coding method and system. First, a target coded video frame is obtained; the target coded video frame includes a base layer video coding frame and an enhancement layer video coding frame; the coding resolution of the base layer video coding frame is lower than the coding resolution of the enhancement layer video coding frame; the base layer video coding frame and the enhancement layer video coding frame are in the same time domain; base layer information is acquired from the base layer video frame to obtain first inter-layer information; the coding resolution of the first inter-layer information is the same as the coding resolution of the enhancement layer video coding frame; the first inter-layer information includes: spatial features, predicted motion information, and layer prior information; a previously reconstructed video frame is obtained; and the target coded video frame is encoded based on the previously reconstructed video frame and the first inter-layer information to obtain a first bitstream. In the above method, base layer information is acquired from the lower-resolution base layer video coding frame in the target coded video frame to obtain first inter-layer information. The first inter-layer information includes spatial features, predicted motion information, and layer prior information at the same coding resolution as the enhancement layer. Finally, encoding is performed based on the target frame to be encoded and the reconstructed video frame from the previous frame to complete the encoding of the video frame and obtain the first bitstream. Since the encoding resolution of the first inter-layer information is the same as the resolution of the enhancement layer bitstream, the encoding of the high-resolution video frame in the enhancement layer can be combined with the first inter-layer information in the base layer video coding frame, thereby improving the video frame encoding performance. At the same time, the encoding of the video frame is further combined with the reconstructed video frame from the previous video frame. This hybrid use of inter-frame information and inter-layer information to encode the video frame significantly improves video encoding performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0050] Figure 1A flowchart of a learning-based spatial scalable video coding method provided in an embodiment of the present application;

[0051] Figure 2 A flowchart of another learning-based spatial scalable video coding method provided in an embodiment of the present application;

[0052] Figure 3 A performance parameter indicator diagram of a learning-based spatial scalable video coding method provided in an embodiment of the present application;

[0053] Figure 4 A structural diagram of a learning-based spatial scalable video coding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] As described above, in current spatial domain-based video coding schemes, video coding is usually completed based on traditional video coding standards. Video coding standards are a set of specifications and algorithms used to digitally compress and encode video signals so that they occupy less bandwidth and storage space during storage and transmission while maintaining high video quality. Existing spatial domain-based video coding schemes usually perform video encoding based on traditional video coding standards such as H.264 / AVC and H.265 / HEVC. Although this method has good interoperability, traditional video coding standards need to ensure a high bit rate when processing complex scenes with multiple resolutions and high frame rates. Therefore, they cannot meet the processing requirements of high quality and high resolution in a low-bandwidth environment, and the encoding performance of video data is low.

[0055] Therefore, how to solve the problem of low video data encoding performance in the prior art has become a technical problem that those skilled in the art urgently need to solve.

[0056] To address the above-mentioned issues, the present application provides a learning-based spatial scalable video coding method and system. The method first obtains a target coded video frame, comprising a base layer video coding frame and an enhancement layer video coding frame. The base layer video coding frame has a lower coding resolution than the enhancement layer video coding frame. The base layer video coding frame and the enhancement layer video coding frame are in the same time domain. Base layer information is acquired from the base layer video frame to obtain first inter-layer information. The coding resolution of the first inter-layer information is the same as that of the enhancement layer video coding frame. The first inter-layer information includes spatial features, predicted motion information, and layer prior information. A previously reconstructed video frame is acquired. The target coded video frame is encoded based on the previously reconstructed video frame and the first inter-layer information to obtain a first bitstream. In the method, base layer information is acquired from the lower-resolution base layer video coding frame in the target coded video frame to obtain the first inter-layer information. The first inter-layer information includes spatial features, predicted motion information, and layer prior information at the same coding resolution as the enhancement layer. Finally, encoding is performed based on the target frame to be encoded and the reconstructed video frame from the previous frame to complete the encoding of the video frame and obtain a first bitstream. Since the encoding resolution of the first inter-layer information is the same as the resolution of the enhancement layer bitstream, the encoding of the high-resolution video frame in the enhancement layer can be combined with the first inter-layer information in the base layer video coding frame for encoding, thereby improving the performance of video frame encoding. At the same time, the encoding of the video frame is further combined with the reconstructed video frame from the previous video frame for encoding. Thus, by combining the inter-frame information and the inter-layer information to encode the video frame, the video encoding performance is greatly improved. To help those skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all of them. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0057] Figure 2This is a flow chart of another video encoding method provided by the present application. As shown in the figure, the encoding of video frames is divided into base layer encoding and enhancement layer encoding. In actual application scenarios, the encoding resolution of the enhancement layer is often higher than the encoding resolution of the base layer. In the figure, I1 is a frame encoded by the base layer intra-frame encoder, and P1 is a frame encoded by the enhancement layer intra-frame encoder, which refers to I1 of the base layer. P2 is a frame encoded by the base layer inter-frame encoder, which refers to the previous frame of the base layer. B1 is a frame encoded by the enhancement layer inter-frame encoder, which refers to the current frame of the base layer and the previous frame of the enhancement layer. It can be seen that in the video encoding method provided by the present application, the encoding of the enhancement layer will refer to the frame encoded in the base layer and the video frame in the previous enhancement layer. The encoding of the base layer will refer to the previous frame in the base layer.

[0058] See also Figure 1 , which is a flow chart of a video encoding method provided in an embodiment of the present application, specifically comprising the following steps:

[0059] S101: Obtain a target coded video frame; the target coded video frame includes a base layer video coded frame and an enhancement layer video coded frame; the coding resolution of the base layer video coded frame is smaller than the coding resolution of the enhancement layer video coded frame; the base layer video coded frame and the enhancement layer video coded frame are in the same time domain

[0060] In actual application scenarios, video encoding often involves encoding of the base layer and encoding of the enhancement layer. The encoding of the base layer is used to create a basic video quality level, which usually includes a lower resolution and bit rate. The video encoding of the enhancement layer will further encode the video of the base layer to improve the quality of the video or increase the resolution and other aspects. The encoding of the enhancement layer usually relies on the information of the base layer to achieve higher quality video transmission. Therefore, the target encoded video frame obtained in the present application includes a base layer video encoding frame and an enhancement layer video encoding frame. The two are in the same time domain, and the encoding resolution of the enhancement layer video encoding frame is greater than the encoding resolution of the base layer video encoding frame.

[0061] In the learning-based spatial scalable video coding method of this application, the overall method is applied within a pre-set neural network framework. The inclusion of the pre-set neural network enables the extraction of more key features, such as spatial features and residual information, from the target coded video data during the video coding process. Furthermore, the pre-set neural network framework reduces the storage and transmission costs of video data, thereby improving overall video coding performance.

[0062] In traditional video coding schemes, the reference relationship between video coding performance at different resolutions is single and cannot be globally optimized. In recent years, development has been slow, and performance improvement has tended to be saturated. In this application, in order to further improve the performance and performance of spatially scalable video coding, spatially scalable video coding based on neural network learning is adopted. When encoding high-resolution video data, low-resolution coded data is used as an inter-layer reference, so that the high-resolution enhancement layer video coding can refer to the information in the low-resolution base layer video coding frame. The encoding of the higher-resolution enhancement layer video coding frame uses a learnable neural network to fully combine inter-layer information and time domain information, thereby improving the performance of video coding.

[0063] Performance S102: Obtain the basic layer information of the basic layer video frame to obtain the first inter-layer information; the coding resolution of the first inter-layer information is the same as the coding resolution of the enhanced layer code video coding frame; the first inter-layer information includes: spatial domain features, predicted motion information and layer prior information.

[0064] As can be seen from the above, in the video coding scheme of this application, the encoding of high-resolution video frames in the enhancement layer can complete its own video coding based on the codec information in the low-resolution base layer video frames, thereby improving video coding performance. Therefore, after obtaining the target coded video frame, the base layer information is obtained from the base layer video frame to obtain the first inter-layer information. The coding resolution of the high-resolution first inter-layer information is the same as the coding resolution of the enhancement layer code stream. Therefore, when the enhancement layer video coding frame in the target coded video frame is subsequently encoded, the video encoding can be performed based on the first inter-layer information in the base layer video frame.

[0065] The first inter-layer information includes spatial features, predicted motion information, and prior information. Spatial features of a video refer to characteristics such as the spatial distribution and color distribution of each frame. In video coding, spatial features can be used for inter-frame and intra-frame compression. Predicted motion information refers to the motion characteristics between adjacent frames in a video sequence. It is used to assist in encoding motion information within frames in the enhancement layer video coding. Layer prior information refers to prior knowledge of the statistical and structural characteristics of the video content. In video coding, fully leveraging prior information to model and compress the bitstream can improve coding performance, reduce distortion, and reduce bitrate.

[0066] Specifically, the above process of obtaining the first inter-layer information through the base layer video frame is implemented through the following three steps:

[0067] Step 1: perform encoding and decoding reconstruction on the base layer video coding frame to obtain second inter-layer information; the coding resolution of the second inter-layer information is lower than the coding resolution of the enhancement layer video coding frame.

[0068] In the process of obtaining the first inter-layer information based on the base layer video frame, the base layer video frame is first reconstructed by encoding and decoding to obtain the second inter-layer information. The coding resolution of the second inter-layer information is lower than that of the enhancement layer bitstream, and includes low-resolution motion information, implicit residual expression, and base layer reconstruction features.

[0069] During the encoding and decoding process of the base layer video frame, the base layer video frame is encoded and decoded once. During the encoding process, the motion information and residual information in the base layer stream corresponding to the base layer video frame are encoded. That is, the low-resolution information of the current frame is encoded and written into the stream. This information is then decoded to obtain the second layer information.

[0070] Step 2: Perform domain transformation processing on the second inter-layer information to obtain transformed second inter-layer information.

[0071] In the obtained second-layer inter-layer information, since the second-layer inter-layer information essentially belongs to the information features of the base layer, and the video coding for the high-resolution portion often involves the video coding of the enhancement layer, the second-layer inter-layer information and the information in the enhancement layer often have different semantics and information content. Therefore, it is necessary to perform domain transformation processing on the second-layer inter-layer information to convert the second-layer inter-layer information into a feature domain suitable for feature fusion in the enhancement layer. Among them, the low-resolution motion information in the second-layer inter-layer information will be transformed from two-channel features to multi-channel features after domain transformation processing, while the residual implicit expression and the base layer reconstruction features only undergo nonlinear changes through domain transformation processing, and their number of channels does not change.

[0072] Step three: up-sample the transformed second inter-layer information according to the coding resolution of the enhancement layer video coding frame to obtain the first inter-layer information.

[0073] Finally, based on the coding resolution of the enhancement layer bitstream, the magnification required to map the transformed second-layer inter-layer information to the enhancement layer features can be determined. Data upsampling is then performed on each type of feature information in the second-layer inter-layer information to obtain a high-resolution representation of the second-layer inter-layer information, namely, the first-layer inter-layer information. The upsampling algorithm can employ neighbor interpolation, bilinear interpolation, bicubic interpolation, and other algorithms. The specific upsampling algorithm can vary depending on the actual application scenario and is not specifically limited in this embodiment.

[0074] In actual application scenarios, in the process of obtaining the first inter-layer information by upsampling the second inter-layer information, the obtained first inter-layer information can be finely adjusted through refined residuals, thereby improving the data accuracy of the first inter-layer information.

[0075] S103: Obtain the previous reconstructed video frame;

[0076] S104: Encode the target coded video frame according to the previous reconstructed video frame and the first inter-layer information to obtain a first code stream reconstructed video frame.

[0077] After obtaining the first inter-layer information, the encoding of the target coded video frame will be performed based on the inter-frame information between the previous video frame and the target coded video frame and the first inter-layer information. This is intended to improve the video encoding performance of the target coded video frame by combining the inter-frame information and the inter-layer information. Specifically, the process of encoding the target coded video frame based on the previous video frame and the first inter-layer information is implemented through the following three steps:

[0078] Step 1: performing encoding and decoding reconstruction according to the target coded video frame, the previous reconstructed video frame, and the predicted motion information to obtain reconstructed high-resolution motion information;

[0079] In the process of encoding the target coded video frame based on the previous video frame and the first inter-layer information, the encoding and decoding reconstruction is first performed based on the combined predicted motion information, the target coded video frame and the previous video frame to obtain reconstructed high-resolution motion information.

[0080] Specifically, the target coded video frame and the previous video frame are first input into a preset optical flow network to generate high-resolution motion information with the same encoding resolution as the enhancement layer. The optical flow network is a computer vision algorithm used to estimate the motion information of pixels in an image. When the target coded video frame and the previous video frame are input into the preset optical flow network, the corresponding optical flow, i.e., high-resolution motion information, is output. Optical flow refers to the change in brightness pattern on the surface of an object in an image caused by the movement of a camera or object. The goal of the optical flow network is to estimate the direction and speed of an object's movement by analyzing the brightness changes between pixels in an image.

[0081] After obtaining high-resolution motion information, encoding is performed based on the predicted motion information in the first inter-layer information and the high-resolution motion information to obtain a corresponding high-resolution motion information bitstream. Finally, the high-resolution motion information bitstream is decoded to obtain reconstructed high-resolution motion information. Specifically, during the encoding process based on the predicted motion information and high-resolution motion, the predicted motion information is input into a preset motion entropy model to determine the probability distribution parameters of the high-resolution motion information bitstream, further improving the performance of high-resolution motion information encoding.

[0082] When determining the probability distribution parameters of a bitstream, a preset motion entropy model and predicted motion information can provide important auxiliary information. Motion entropy models are typically based on modeling the temporal or spatial variations of data to provide statistical characteristics of the data. Predicted motion information can help infer the future temporal or spatial trends of the data, thereby more accurately establishing a probability distribution model. Specifically, the preset motion entropy model can be based on motion entropy theory, including modeling factors such as the frequency and amplitude of data changes to determine the data's probability distribution characteristics. For example, if we know that the data exhibits a certain periodicity or trend in temporal variations, we can use this information to establish a probability distribution model and, in turn, determine the probability distribution parameters of the bitstream. Predicted motion information, on the other hand, can predict the future development trends of the data through methods such as motion prediction algorithms, providing more accurate information for determining the probability distribution parameters of the bitstream. Based on predicted motion information, we can better understand the evolution of the data and build a more accurate probability distribution model accordingly. Therefore, by presetting the motion entropy model and predicting the motion information to determine the probability distribution parameters of the bitstream, we can help more accurately understand the statistical characteristics and development trends of the data, thereby selecting a suitable probability distribution model and determining the probability distribution parameters of the bitstream, providing more effective support for data compression encoding.

[0083] Step 2: performing context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal domain features in the previous reconstructed video frame to generate a multi-scale mixed context;

[0084] After completing the reconstruction of high-resolution motion information, context mining is further performed on the temporal features in the previous reconstructed video frame based on the reconstructed high-resolution motion information and spatial features to generate multi-scale mixed context, thereby improving the performance and efficiency of video encoding through the mixed use of inter-frame information and inter-layer information.

[0085] In the process of generating a multi-scale mixed context, multi-scale spatial and temporal features are first determined based on the spatial features in the first inter-layer information and the temporal features in the previous video frame. In determining the multi-scale temporal features, traditional image processing methods such as convolutional neural networks and multi-scale filters can be used to obtain spatial features of different scales in the target coded video frame. Then, methods such as optical flow estimation are used to extract temporal features between the previous video frame and the target coded video frame. The features of the two are fused and extracted to obtain multi-scale spatial and temporal features.

[0086] After this, the reconstructed high-resolution motion information needs to be further downsampled to obtain multi-scale motion information, and the multi-scale temporal features are motion compensated using this multi-scale motion information to generate aligned multi-scale temporal features. Multi-scale motion information corresponds to video motion features at different spatial scales. These multi-scale motion information can more comprehensively describe the dynamic changes in video sequences, which is of great significance for some visual tasks that require consideration of multi-scale information, such as target tracking and action recognition. In the process of motion compensation of multi-scale temporal features, the temporal features in the current frame are aligned to the position of the reference frame by performing operations such as translation and interpolation according to the corresponding motion vector. This can eliminate temporal deformation or blurring caused by object motion.

[0087] Finally, a multiscale mixed context is generated using the multiscale spatial features and the aligned multiscale temporal features. During this process, a feature weight map is constructed between spatial and temporal features at the same scale. By performing feature fusion on the feature weight map, multiscale mixed features are further generated. Finally, context modeling is performed on the multiscale mixed features to obtain the final multiscale mixed context.

[0088] Specifically, in the process of feature fusion based on the feature weight map, the aligned multi-scale time domain features and multi-scale spatial domain features can be weighted fused, and the weight maps corresponding to the features can be multiplied, or the two can be fused in a weighted summation manner according to the weights to obtain a more comprehensive multi-scale mixed feature.

[0089] Step three: Encode the target coded video frame using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first bitstream.

[0090] Finally, the multi-scale context obtained by combining the inter-layer information and inter-frame information of the base layer in the above steps is combined with the target coded video frame and the prior information in the first inter-layer information and encoded to complete the encoding process of the target coded video frame and obtain the first code stream.

[0091] In the process of encoding the target coded video frame, the probability distribution parameters of the first code stream are determined according to the layer prior information and the preset inter-layer prior entropy model, and the target coded video frame is encoded based on the probability distribution parameters of the first code stream to obtain the first code stream.

[0092] Layer prior information includes the probability of each layer occurring, inter-layer correlation properties, and more. The inter-layer prior entropy model can estimate the relationship between different layers and provide an estimate of inter-layer entropy. Inter-layer entropy measures the correlation and redundancy between one layer and other layers during the encoding process. Various statistical methods or models, such as Gaussian models and conditional entropy models, can be used to establish a preset inter-layer prior entropy model. This embodiment does not specifically limit the method or type of pre-setting the preset inter-layer entropy model.

[0093] See also Figure 3 , this figure is a performance parameter indicator diagram of a learning-based spatial scalable video coding method provided in an embodiment of the present application.

[0094] As shown in the figure, the present invention achieves better coding performance than the existing SHVC (Scalable High Efficiency Video Coding) method. Specifically, when using BD-Rate to measure coding gain, the solution described in Example 1 surpasses the H.265 / SHVC standard reference software SHM-12.4, whether in RGB or YUV420 color space, and whether using PSNR or MS-SSIM as the distortion metric.

[0095] This embodiment provides a learning-based spatial scalable video coding method. The method first obtains a target coded video frame, comprising a base layer video coded frame and an enhancement layer video coded frame. The base layer video coded frame has a lower coding resolution than the enhancement layer video coded frame. The base layer video coded frame and the enhancement layer video coded frame are in the same temporal domain. Base layer information is acquired from the base layer video frame to obtain first inter-layer information. The coding resolution of the first inter-layer information is the same as that of the enhancement layer video coded frame. The first inter-layer information includes spatial features, predicted motion information, and layer prior information. A previously reconstructed video frame is acquired. The target coded video frame is encoded based on the previously reconstructed video frame and the first inter-layer information to obtain a first bitstream. In the method, base layer information is acquired from the lower-resolution base layer video coded frame within the target coded video frame to obtain the first inter-layer information. The first inter-layer information includes spatial features, predicted motion information, and layer prior information at the same coding resolution as the enhancement layer. Finally, encoding is performed based on the target frame to be encoded and the reconstructed video frame from the previous frame to complete the encoding of the video frame and obtain the first bitstream. Since the encoding resolution of the first inter-layer information is the same as the resolution of the enhancement layer bitstream, the encoding of the high-resolution video frame in the enhancement layer can be combined with the first inter-layer information in the base layer video coding frame, thereby improving the video frame encoding performance. At the same time, the encoding of the video frame is further combined with the reconstructed video frame from the previous video frame. This hybrid use of inter-frame information and inter-layer information to encode the video frame significantly improves video encoding performance.

[0096] The following introduces a learning-based spatial scalable video coding system provided in an embodiment of the present application. The learning-based spatial scalable video coding system described below and the learning-based spatial scalable video coding method described above can refer to each other.

[0097] Reference Figure 4 , which is a schematic diagram of the structure of a learning-based spatial scalable video coding system provided in an embodiment of the present application, specifically including the following modules:

[0098] The first acquisition module 100 is configured to acquire a target coded video frame; the target coded video frame includes a base layer coded video frame and an enhancement layer coded video frame; the coding resolution of the base layer coded video frame is lower than the coding resolution of the enhancement layer coded video frame; the base layer coded video frame and the enhancement layer coded video frame are in the same time domain;

[0099] The inter-layer information acquisition module 200 is configured to acquire base layer information from the base layer video frame to obtain first inter-layer information; the coding resolution of the first inter-layer information is the same as the coding resolution of the enhancement layer coded video frame; the first inter-layer information includes: spatial features, predicted motion information, and layer prior information;

[0100] The second acquisition module 300 is used to acquire the previous reconstructed video frame;

[0101] The encoding module 400 is used to encode the target coded video frame according to the reconstructed video frame of the previous frame and the first inter-layer information to obtain a first code stream.

[0102] Optionally, the encoding module 400 is specifically configured to:

[0103] Performing encoding and decoding reconstruction according to the target coded video frame, the previous reconstructed video frame, and the predicted motion information to obtain reconstructed high-resolution motion information;

[0104] Performing context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal domain features in the previous reconstructed video frame to generate a multi-scale mixed context;

[0105] The target coded video frame is encoded using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first code stream.

[0106] Optionally, the encoding module 500 is specifically configured to:

[0107] Inputting the target coded video frame and the previous reconstructed video frame into a preset optical flow network to obtain high-resolution motion information; the coding resolution of the high-resolution motion information is the same as the coding resolution of the enhancement layer coded video frame;

[0108] Encoding is performed according to the predicted motion information and the high-resolution motion information to obtain a bit stream of the high-resolution motion information;

[0109] The code stream of the high-resolution motion information is decoded and reconstructed based on the predicted motion information to obtain the reconstructed high-resolution motion information.

[0110] Optionally, the encoding module 500 is specifically configured to:

[0111] Determining multi-scale spatial domain features and multi-scale temporal domain features based on the spatial domain features and the temporal features in the last reconstructed video frame;

[0112] downsampling the reconstructed high-resolution motion information to obtain multi-scale motion information;

[0113] Performing motion compensation on the multi-scale temporal features based on the multi-scale motion information to obtain aligned multi-scale temporal features;

[0114] The multi-scale mixed context is generated according to the multi-scale spatial domain features and the aligned multi-scale temporal domain features.

[0115] Optionally, the encoding module is specifically used to:

[0116] Determining probability distribution parameters of the first bitstream according to the layer prior information and a preset inter-layer prior entropy model;

[0117] The target coded video frame is encoded based on the probability distribution parameter of the first code stream to obtain the first code stream.

[0118] Optionally, the inter-layer information acquisition module 300 is specifically configured to:

[0119] Performing encoding and decoding reconstruction on the base layer video coding frame to obtain second inter-layer information; the coding resolution of the second inter-layer information is lower than the coding resolution of the enhancement layer video coding frame;

[0120] performing domain transformation processing on the second inter-layer information to obtain transformed second inter-layer information;

[0121] The transformed second inter-layer information is up-sampled according to the coding resolution of the enhancement layer video coding frame to obtain the first inter-layer information.

[0122] Optionally, the encoding module 500 is specifically configured to:

[0123] Determining probability distribution parameters of a bitstream of the high-resolution motion information based on the predicted motion information and a preset motion entropy model;

[0124] The code stream of the high-resolution motion information is determined according to the probability distribution parameter of the code stream of the high-resolution motion information.

[0125] Optionally, the encoding module 500 is specifically configured to:

[0126] Constructing a feature weight map between the aligned multi-scale temporal features and the spatial features;

[0127] Based on the feature weight map, the aligned multi-scale time domain features and the spatial domain features are subjected to feature fusion to obtain multi-scale hybrid features;

[0128] The multi-scale mixed context is generated according to the multi-scale mixed feature.

[0129] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the method device, electronic device and vehicle, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The method device, electronic device and vehicle described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0130] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A learning-based spatial scalable video coding method, characterized in that: Applied to a preset neural network, the method includes: Obtaining a target coded video frame; the target coded video frame includes a base layer video coded frame and an enhancement layer video coded frame; the coding resolution of the base layer video coded frame is lower than the coding resolution of the enhancement layer video coded frame; the base layer video coded frame and the enhancement layer video coded frame are in the same time domain; Acquiring base layer information from the base layer video frame to obtain first inter-layer information; the coding resolution of the first inter-layer information is the same as the coding resolution of the enhancement layer video coding frame; the first inter-layer information includes: spatial features, predicted motion information, and layer prior information; Get the previous frame to reconstruct the video frame; Encoding the target coded video frame according to the previous reconstructed video frame and the first inter-layer information to obtain a first bitstream; The encoding of the target coded video frame according to the previous reconstructed video frame and the first inter-layer information to obtain a first bitstream specifically includes: Performing encoding and decoding reconstruction according to the target coded video frame, the previous reconstructed video frame, and the predicted motion information to obtain reconstructed high-resolution motion information; Performing context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal domain features in the previous reconstructed video frame to generate a multi-scale mixed context; Encoding the target coded video frame using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first bitstream; The performing of context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal domain features in the previous reconstructed video frame to generate a multi-scale mixed context specifically includes: Determining multi-scale spatial domain features and multi-scale temporal domain features based on the spatial domain features and the temporal features in the last reconstructed video frame; downsampling the reconstructed high-resolution motion information to obtain multi-scale motion information; Performing motion compensation on the multi-scale temporal features based on the multi-scale motion information to obtain aligned multi-scale temporal features; The multi-scale mixed context is generated according to the multi-scale spatial domain features and the aligned multi-scale temporal domain features.

2. The method according to claim 1, characterized in that The encoding, decoding and reconstructing according to the target coded video frame, the previous reconstructed video frame and the predicted motion information to obtain reconstructed high-resolution motion information specifically includes: Inputting the target coded video frame and the previous reconstructed video frame into a preset optical flow network to obtain high-resolution motion information; the coding resolution of the high-resolution motion information is the same as the coding resolution of the enhancement layer video coding frame; Encoding is performed according to the predicted motion information and the high-resolution motion information to obtain a bit stream of the high-resolution motion information; The code stream of the high-resolution motion information is decoded and reconstructed based on the predicted motion information to obtain the reconstructed high-resolution motion information.

3. The method according to claim 1, characterized in that The encoding of the target coded video frame by using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first bitstream specifically includes: Determining probability distribution parameters of the first bitstream according to the layer prior information and a preset inter-layer prior entropy model; The target coded video frame is encoded based on the probability distribution parameter of the first code stream to obtain the first code stream.

4. The method according to claim 1, wherein The acquiring of the base layer information of the base layer video frame to obtain the first inter-layer information specifically includes: Performing encoding and decoding reconstruction on the base layer video coding frame to obtain second inter-layer information; the coding resolution of the second inter-layer information is lower than the coding resolution of the enhancement layer video coding frame; performing domain transformation processing on the second inter-layer information to obtain transformed second inter-layer information; The transformed second inter-layer information is up-sampled according to the coding resolution of the enhancement layer video coding frame to obtain the first inter-layer information.

5. The method according to claim 2, characterized in that The encoding according to the predicted motion information and the high-resolution motion information to obtain a code stream of the high-resolution motion information specifically includes: Determining probability distribution parameters of a bitstream of the high-resolution motion information based on the predicted motion information and a preset motion entropy model; The code stream of the high-resolution motion information is determined according to the probability distribution parameter of the code stream of the high-resolution motion information.

6. The method according to claim 1, characterized in that Generating the multi-scale mixed context according to the spatial domain features and the aligned multi-scale temporal domain features specifically includes: Constructing a feature weight map between the aligned multi-scale temporal features and the spatial features; Based on the feature weight map, the aligned multi-scale time domain features and the spatial domain features are subjected to feature fusion to obtain multi-scale hybrid features; The multi-scale mixed context is generated according to the multi-scale mixed feature.

7. A learning-based spatial scalable video coding system, characterized in that Applied to a preset neural network, the system includes: A first acquisition module is configured to acquire a target coded video frame; the target coded video frame includes a base layer video coded frame and an enhancement layer video coded frame; the coding resolution of the base layer video coded frame is lower than the coding resolution of the enhancement layer video coded frame; the base layer video coded frame and the enhancement layer video coded frame are in the same time domain; an inter-layer information acquisition module, configured to acquire base layer information from the base layer video frame to obtain first inter-layer information; the coding resolution of the first inter-layer information is the same as the coding resolution of the enhancement layer video coding frame; the first inter-layer information includes: spatial features, predicted motion information, and layer prior information; The second acquisition module is used to acquire the previous reconstructed video frame; an encoding module, configured to encode the target coded video frame according to the previous frame reconstructed video frame and the first inter-layer information to obtain a first code stream; The encoding module is specifically used to: Performing encoding and decoding reconstruction according to the target coded video frame, the previous reconstructed video frame, and the predicted motion information to obtain reconstructed high-resolution motion information; Performing context mining based on the reconstructed high-resolution motion information, the spatial domain features, and the temporal domain features in the previous reconstructed video frame to generate a multi-scale mixed context; Encoding the target coded video frame using the multi-scale mixed context, the target coded video frame, and the layer prior information to obtain a first bitstream; The encoding module is further used for: Determining multi-scale spatial domain features and multi-scale temporal domain features based on the spatial domain features and the temporal features in the last reconstructed video frame; downsampling the reconstructed high-resolution motion information to obtain multi-scale motion information; Performing motion compensation on the multi-scale temporal features based on the multi-scale motion information to obtain aligned multi-scale temporal features; The multi-scale mixed context is generated according to the multi-scale spatial domain features and the aligned multi-scale temporal domain features.

Citation Information

Patent Citations

  • Encoding method and device and decoding method and device for hierarchical video

    CN112702604A

  • Video encoding and decoding methods and systems for video streaming service

    CN1926873A