Video encoding and decoding method, device, equipment, medium and program product

By generating pixel difference features between the target video frame and the first decoded video frame, and using a pre-trained model for encoding, the problem of low compression efficiency in traditional video coding standards is solved, achieving more efficient video compression and image quality improvement.

CN121967698APending Publication Date: 2026-05-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional video coding standards, which use techniques such as motion estimation, motion compensation, and discrete cosine transform, have low compression efficiency and are difficult to effectively reduce redundant content.

Method used

By employing pre-trained encoding and decoding models, second encoded data is generated by analyzing the pixel difference features between the target video frame and the first decoded video frame, thereby reducing the encoding of the entire target video frame and improving video compression efficiency.

Benefits of technology

By identifying the correlation between the target video frame and the first decoded video frame, the amount of encoded data is reduced, video compression efficiency is improved, and image quality and user experience are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967698A_ABST
    Figure CN121967698A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video encoding and decoding method and device, equipment, a medium and a program product. The method comprises the following steps: acquiring at least one video frame to be coded, and determining a target video frame according to the at least one video frame; generating first coded data and a first decoded video frame of the target video frame according to the target video frame; determining pixel difference characteristics of the target video frame and the first decoded video frame through a pre-trained coding model, and generating second coded data according to the first decoded video frame and the pixel difference characteristics; and sending the first coded data and the second coded data to a decoding end, so that the decoding end determines a pixel difference feature according to the first decoded video frame and the second coded data, and generates a second decoded video frame according to the pixel difference feature and the first decoded video frame. According to the embodiment of the invention, the pixel difference characteristics of the target video frame and the first decoded video frame are coded to obtain the second coded data, so that the coded data volume can be reduced, and the video compression efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer technology, and more particularly to a video encoding / decoding method, apparatus, device, medium, and program product. Background Technology

[0002] Video encoding uses specific compression techniques to remove redundant content from a video, making it easier for storage, transmission, and other processing.

[0003] Currently, video encoding is typically performed based on traditional video coding standards. For example, H.264, H.265, or H.266 video coding standards are used to compress video content, thereby reducing transmission bandwidth and storage costs. However, traditional video coding standards, which reduce redundant content through techniques such as motion estimation, motion compensation, and discrete cosine transform, have relatively low compression efficiency. Summary of the Invention

[0004] This disclosure provides a video encoding / decoding method, apparatus, device, medium, and program product that can improve video compression efficiency.

[0005] In a first aspect, embodiments of this disclosure provide a video encoding method, the method being applied at an encoding end, comprising:

[0006] Obtain at least one video frame to be encoded, and determine the target video frame based on the at least one video frame;

[0007] First encoded data and a first decoded video frame of the target video frame are generated based on the target video frame, wherein the first encoded data represents the encoded data of the target video frame;

[0008] The pixel difference features between the target video frame and the first decoded video frame are determined by a pre-trained encoding model. Second encoded data is generated based on the first decoded video frame and the pixel difference features, wherein the second encoded data represents the encoded data of the pixel difference features.

[0009] The first encoded data and the second encoded data are sent to the decoding end so that the decoding end can decode the first encoded data to obtain the first decoded video frame. The pixel difference features are determined by the pre-trained decoding model based on the first decoded video frame and the second encoded data. The second decoded video frame of the target video frame is generated based on the pixel difference features and the first decoded video frame.

[0010] Secondly, this disclosure also provides a video decoding method, which is applied at a decoding end and includes:

[0011] The encoding end sends first encoded data and second encoded data. The second encoded data is obtained by the encoding end encoding the pixel difference features of the target video frame and the first decoded video frame of the target video frame. The first encoded data represents the encoded data of the target video frame.

[0012] Decoding the first encoded data yields the first decoded video frame of the target video frame;

[0013] The pixel difference features are determined by a pre-trained decoding model based on the first decoded video frame and the second encoded data.

[0014] A second decoded video frame of the target video frame is generated based on the pixel difference features and the first decoded video frame.

[0015] Thirdly, embodiments of this disclosure also provide a video encoding apparatus, the apparatus being applied at an encoding end, comprising:

[0016] The first acquisition module is used to acquire at least one video frame to be encoded, and to determine a target video frame based on the at least one video frame.

[0017] A first decoding module is configured to generate first encoded data and a first decoded video frame of the target video frame based on the target video frame, wherein the first encoded data represents the encoded data of the target video frame;

[0018] The encoded data generation module is used to determine the pixel difference features between the target video frame and the first decoded video frame through a pre-trained encoding model, and generate second encoded data based on the first decoded video frame and the pixel difference features, wherein the second encoded data represents the encoded data of the pixel difference features;

[0019] The encoded data sending module is used to send the first encoded data and the second encoded data to the decoding end, so that the decoding end can decode the first encoded data to obtain the first decoded video frame, determine the pixel difference features based on the first decoded video frame and the second encoded data through a pre-trained decoding model, and generate the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

[0020] Fourthly, embodiments of this disclosure also provide a video decoding apparatus, the apparatus being applied at a decoding end, comprising:

[0021] The second acquisition module is used to acquire first encoded data and second encoded data sent by the encoding end. The second encoded data is obtained by the encoding end encoding the pixel difference features of the target video frame and the first decoded video frame of the target video frame. The first encoded data represents the encoded data of the target video frame.

[0022] The second decoding module is used to decode the first encoded data to obtain the first decoded video frame of the target video frame;

[0023] A feature determination module is used to determine the pixel difference features based on the first decoded video frame and the second encoded data using a pre-trained decoding model;

[0024] The video frame determination module is used to generate a second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

[0025] Fifthly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0026] One or more processors, wherein the processors include a central processing unit and a neural network processor;

[0027] Storage device for storing one or more programs.

[0028] When the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in any embodiment of this disclosure.

[0029] Sixthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the methods described in any embodiment of this disclosure.

[0030] In a seventh aspect, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in any embodiment of this disclosure.

[0031] This disclosure provides a video encoding method. It involves generating first encoded data of a target video frame and a first decoded video frame of the target video frame. Then, a pre-trained encoding model is used to determine the pixel difference features between the target video frame and the first decoded video frame. Based on the first decoded video frame and the pixel difference features, second encoded data of the target video frame is generated, and both the first and second encoded data are sent to the decoding end. Since the first decoded video frame is the decoded data of the target video frame, they are correlated. By identifying the correlation between the target video frame and the first decoded video frame through the encoding model, determining the pixel difference features between them based on the correlation, and encoding the pixel difference features to obtain the second encoded data corresponding to the target video frame, the encoding of the entire target video frame can be avoided, reducing the amount of encoded data and improving video compression efficiency. Attached Figure Description

[0032] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0033] Figure 1 This is a flowchart illustrating a video encoding method provided in an embodiment of the present disclosure;

[0034] Figure 2 This is a schematic diagram of the structure of a video coding framework in a video coding method provided in this embodiment of the disclosure;

[0035] Figure 3 This is a schematic flowchart of a video decoding method provided in an embodiment of the present disclosure;

[0036] Figure 4 This is a schematic diagram of the structure of a video decoding framework in a video decoding method provided in an embodiment of this disclosure;

[0037] Figure 5 This is a schematic diagram of the structure of a video encoding device provided in an embodiment of the present disclosure;

[0038] Figure 6 This is a schematic diagram of the structure of a video decoding device provided in an embodiment of the present disclosure;

[0039] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0040] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0041] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0042] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0043] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0044] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0045] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0046] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0047] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0048] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0049] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0050] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0051] Figure 1 This is a flowchart illustrating a video encoding method provided in an embodiment of this disclosure, applicable to video transmission scenarios, such as video uploading, video sharing, or cloud gaming. The method is applied at the encoding end. This method can be executed by a video encoding device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.

[0052] like Figure 1 As shown, the method includes:

[0053] S110. Obtain at least one video frame to be encoded, and determine the target video frame based on the at least one video frame.

[0054] At least one video frame is used to represent a video or a single frame of an image. The target video frame is the video frame to be encoded. For example, the target video frame can be a video frame from a video to be transmitted. The video to be encoded can be video captured by a camera, or video generated by recording the client's screen, etc. The video to be transmitted can be transmitted from the client device to the server or other client devices. Optionally, the target video frame can also be game footage from a cloud game.

[0055] For example, a target video to be encoded is obtained, and the target video frame is determined based on the timestamps of the video frames in the target video. In some embodiments where the encoding end is a client, the client obtains the target video to be encoded according to a video upload operation. The client determines the target video frame based on the timestamps of each video frame in the target video. The video upload operation is used by the client to upload the video to the server. Optionally, the target video frames are determined according to the timestamp order and then encoded. In other embodiments where the encoding end is a client, the client obtains the target video to be encoded according to a video sharing operation. The client determines the target video frame based on the timestamps of each video frame in the target video. The video sharing operation is used by the client to send the video to other clients. It should be noted that in video upload or video sharing scenarios, the client can obtain the original video based on user interaction operations, and use the complete original video or video segments from the original video as the target video to be encoded. The original video can be a video captured by the client device's camera, or a video obtained by recording the screen content of the client device, etc., and this disclosure does not specifically limit the scope. In some embodiments where the encoding end is the server, the server obtains the target video frame according to the target operation instruction of the client, and determines the target video frame according to the timestamp of the video frame in the target video.

[0056] Optionally, each video frame in the video to be encoded is obtained as a target video frame based on its timestamp. Alternatively, at least two video frames are obtained as at least two target video frames based on their timestamps, and these at least two target video frames are processed in parallel.

[0057] Optionally, a target operation instruction sent by the decoding end is obtained, and a target video frame is generated based on the target operation instruction, wherein the target operation instruction is generated based on user interaction operations in the decoding end. In other embodiments where the encoding end is a server, the server generates the target video frame based on the target operation instruction sent by the client. For example, in cloud gaming, the server generates the target video frame based on the target operation instruction sent by the user terminal.

[0058] S120. Generate first encoded data and a first decoded video frame of the target video frame based on the target video frame.

[0059] The first encoded data represents the encoded data of the target video frame. The first decoded video frame is obtained by decoding the first encoded data.

[0060] For example, the target video frame is downsampled to obtain a downsampled video frame. The downsampled video frame is encoded to obtain the first encoded data. The first encoded data is decoded to obtain the decoded data of the target video frame. The decoded data is then upsampled according to the resolution of the target video frame to obtain the first decoded video frame.

[0061] The downsampled video frame is a low-resolution version of the target video frame. The resolution of the downsampled video frame is lower than that of the target video frame. For example, the downsampled video frame is obtained by downsampling the target video frame.

[0062] This disclosure provides a hybrid scalable video coding framework, comprising a base layer and an enhancement layer. The base layer implements basic video coding, while the enhancement layer implements enhanced video coding. The base layer coding is an encoding operation implemented using a CPU and hardware codec, based on traditional video coding standards. The enhancement layer coding is an encoding operation implemented using a neural network processor. Specifically, the base layer provides basic video quality and resolution, ensuring that the decoder can receive a watchable video stream even under poor network conditions or limited decoding capabilities. The base layer can implement video coding based on traditional video coding standards. The enhancement layer provides additional video details under better network conditions to improve video quality and resolution. In this disclosure, the decoder can select between the encoded stream generated by the base layer or the bitstream generated by the enhancement layer for decoding based on the hardware information of the decoding device, achieving compatibility with different hardware devices using the same video coding and decoding scheme.

[0063] Figure 2 This is a schematic diagram of the video coding framework in a video coding method provided in an embodiment of this disclosure. Figure 2 As shown, the video coding framework includes a base layer 210 and an enhancement layer 220. The base layer 210 includes a downsampling module 230, an encoder 240, a decoder 250, and an upsampling module 260. The enhancement layer 220 includes a coding model 280. The coding model 280 includes a feature extraction network 270, an encoder network 281, and an entropy model 282. The downsampling module 230 is used to downsample the target video frame. The upsampling module 260 is used to upsample the decoded data corresponding to the first encoded data. Downsampling is used to reduce the resolution of the video frame. Upsampling is used to increase the resolution of the video frame.

[0064] In some embodiments, a downsampled video frame corresponding to the target video frame is encoded to obtain first encoded data. For example, the downsampled video frame is encoded using a traditional video coding standard to obtain the first encoded data. The first encoded data is then decoded using the same decoding algorithm as the encoding of the downsampled video frame to obtain decoded data corresponding to the downsampled video frame. The decoded data is then upsampled according to the resolution of the target video frame to obtain a first decoded video frame. The resolution of the first decoded video frame is the same as the resolution of the target video frame.

[0065] Reference Figure 2 As shown, a downsampled video frame is input to encoder 240, which encodes the downsampled video frame and outputs the first encoded data corresponding to the downsampled video frame. The first encoded data is input to decoder 250, which decodes the first encoded data to obtain decoded data. The decoded data is input to upsampling module 260, which outputs a first decoded video frame with the same resolution as the target video frame.

[0066] S130. Determine the pixel difference features between the target video frame and the first decoded video frame using a pre-trained encoding model, and generate second encoded data based on the first decoded video frame and the pixel difference features.

[0067] The second encoded data represents the encoded data of the pixel difference features. The encoding model is a neural network model. For example, the encoding model includes a feature extraction network, an encoder network, and an entropy model. Optionally, the feature extraction network includes a network formed by stacking multiple convolutional layers. The encoder network includes a network formed by stacking multiple convolutional layers, residual blocks, and generalized divisive normalization (GDN) layers. The entropy model is used to entropy encode the output of the encoder network.

[0068] Pixel difference features characterize the differences in spatial and / or temporal features between the target video frame and the first decoded video frame. Spatial features represent the spatial and / or color distribution of pixels, while temporal features represent the motion information of pixels. Spatial and temporal features of the first decoded video frame can be extracted using a feature extraction network. Semantic understanding of these features yields the contextual features corresponding to the first decoded video frame. These contextual features are then used as conditions to assist the encoder network in encoding the target video frame.

[0069] For example, the number of channels in the target video frame is the same as the number of channels in the context features. The target video frame and context features are input into an encoder network, which concatenates them channel-by-channel to obtain the concatenated features. These concatenated features are then processed through convolutional layers, normalization layers, and residual blocks to automatically perceive the correlation between the target video frame and the first decoded video frame, and based on this correlation, outputs the pixel difference features between the target video frame and the first decoded video frame.

[0070] For example, determining the pixel difference features between the target video frame and the first decoded video frame using a pre-trained encoding model, and generating second encoded data based on the first decoded video frame and the pixel difference features, includes: determining the pixel difference features using the encoding model based on the pixel features of the target video frame and the first decoded video frame; determining the codeword length associated with the pixel difference features using the encoding model based on the pixel difference features and the first decoded video frame; and encoding the pixel difference features according to the codeword length to obtain the second encoded data.

[0071] Pixel features can include spatial and color information. For example, spatial information includes pixel location and depth. Color information includes pixel color, brightness, and saturation. Codeword length represents the length of the codeword corresponding to the pixel difference feature. A codeword is a binary code used to represent the pixel difference feature. For example, strings with a high probability of occurrence in the pixel difference feature can be assigned shorter codewords, while strings with a low probability of occurrence can be assigned longer codewords, thus achieving higher encoding efficiency.

[0072] Optionally, determining the pixel difference features based on the pixel features of the target video frame and the first decoded video frame using the encoding model includes:

[0073] The first decoded video frame is input into the feature extraction network, which obtains the pixel features of the first decoded video frame. The pixel features of the first decoded video frame are then output to the encoder network and the entropy model. The target video frame is input into the encoder network, which obtains the pixel features of the target video frame. Based on the pixel features of the target video frame and the pixel features of the first decoded video frame, the pixel difference features are determined, and the pixel difference features are output to the entropy model.

[0074] like Figure 2As shown, the first decoded video frame is input into the feature extraction network 270, which outputs the context features of the first decoded video frame. These context features are then input into the encoder network 281 and the entropy model 282 included in the encoding model 280. The encoder network 281 determines pixel difference features based on the context features and the target video frame. The encoder network 281 outputs the pixel difference features to the entropy model 282, which encodes the pixel difference features and outputs the second encoded data 290.

[0075] Specifically, the context features corresponding to the first decoded video frame are concatenated along the channel dimension to the pixel features of the target video frame using an encoder network. Then, multi-scale feature transformation is performed on the pixel features of each channel to obtain the spatial correlation between the target video frame and the first decoded video frame. The pixel difference features between the target video frame and the first decoded video frame are determined based on the spatial correlation.

[0076] For example, the method of determining the codeword length associated with the pixel difference feature based on the pixel difference feature and the first decoded video frame using the encoding model, and encoding the pixel difference feature according to the codeword length to obtain the second encoded data includes: inputting the pixel difference feature into the entropy model in the encoding model, determining the codeword length associated with the pixel difference feature based on the pixel difference feature and the pixel features of the first decoded video frame using the entropy model, encoding the pixel difference feature according to the codeword length, and outputting the second encoded data.

[0077] In this embodiment, pixel difference features and context features corresponding to the first decoded video frame are input into an entropy model. A super-prior model within the entropy model determines super-prior parameters based on the pixel difference features. Specifically, the super-prior model includes a prior encoder, an arithmetic encoder, an arithmetic decoder, and a prior decoder. Pixel difference features are input into the super-prior model, and the prior encoder processes the pixel difference features to obtain a super-prior processing result. The super-prior processing result is quantized to obtain a quantization result. The arithmetic encoder and decoder process the quantization result, and the processing result is input into the prior decoder. The prior decoder outputs the super-prior parameters. The entropy model determines the codeword length associated with the pixel difference features based on the super-prior parameters and the context features corresponding to the first decoded video frame. Optionally, the entropy model further includes a context encoder. The context encoder includes convolutional layers and activation layers. The context features corresponding to the first decoded video frame are input into the context encoder, and the context encoder outputs the context prior parameters corresponding to the context features. Gaussian distribution parameters are determined based on the super-prior parameters and the context prior parameters, and a probability distribution model is determined based on the Gaussian distribution parameters. The codeword length associated with pixel difference features is determined based on a probability distribution model.

[0078] Optionally, the first decoded video frame of the preceding video frame is obtained. The first decoded video frame corresponding to the target video frame and the first decoded video frame of the preceding video frame are input into a feature extraction network. The feature extraction network extracts spatial and temporal features, and generates context features based on these features. An encoder network concatenates these context features along the channel dimension to the pixel features of the target video frame. Then, multi-scale feature transformations are performed on the pixel features of each channel to obtain the spatial and temporal correlations between the target video frame and the first decoded video frame. Based on these spatial and temporal correlations, the pixel difference features between the target video frame and its corresponding first decoded video frame are determined.

[0079] S140. Send the first encoded data and the second encoded data to the decoding end so that the decoding end can decode the first encoded data to obtain the first decoded video frame. Determine the pixel difference features based on the first decoded video frame and the second encoded data using a pre-trained decoding model. Generate the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

[0080] In this embodiment, a priori parameters are added to the second encoded data, and the first and second encoded data are sent to the decoding end. The decoding end decodes the first encoded data to obtain decoded data, upsamples the decoded data to obtain a first decoded video frame, determines the pixel difference features based on the second encoded data and the first decoded video frame using a pre-trained decoding model, and generates a second decoded video frame based on the pixel difference features and the first decoded video frame. The decoding end can be a client or a server.

[0081] The encoding model and the decoding model are neural network models trained based on sample pairs consisting of sample frames in the training sample set and their corresponding decoded sample frames. In this embodiment, the training process of the encoding and decoding models is as follows: A training sample set is constructed based on video samples. Optionally, the video samples can be videos with a resolution higher than a set resolution threshold. Any video frame in the video samples is used as a sample frame. The sample frame is downsampled to obtain a sample downsampled frame. The sample downsampled frame is encoded and decoded using the codec of the base layer to obtain sample decoded data. The sample decoded data is upsampled according to the resolution of the sample frame to obtain a decoded sample frame. Sample pairs are formed based on the sample frames and their corresponding decoded sample frames. The encoding and decoding models are trained using the sample pairs corresponding to each sample frame in the video samples.

[0082] The technical solution of this disclosure involves generating first encoded data and a first decoded video frame of the target video frame. Then, a pre-trained encoding model is used to determine the pixel difference features between the target video frame and the first decoded video frame. Based on the first decoded video frame and the pixel difference features, second encoded data of the target video frame is generated, and both are sent to the decoding end. Since the first decoded video frame is the decoded data of the target video frame, they are correlated. By identifying the correlation between the target video frame and the first decoded video frame through the encoding model, determining the pixel difference features between them based on the correlation, and encoding the pixel difference features to obtain the second encoded data corresponding to the target video frame, encoding the entire target video frame can be avoided, reducing the amount of encoded data and improving video compression efficiency.

[0083] Figure 3 This is a flowchart illustrating a video decoding method provided in an embodiment of this disclosure, applicable to video transmission scenarios, such as video uploading, video sharing, or cloud gaming. The method is applied at the decoding end. This method can be executed by a video decoding device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.

[0084] like Figure 3 As shown, the method includes:

[0085] S310. Obtain the first encoded data and the second encoded data sent by the encoding end. The second encoded data is obtained by the encoding end encoding the pixel difference features of the target video frame and the first decoded video frame of the target video frame.

[0086] The first encoded data represents the encoded data of the target video frame. The target video frame is downsampled to obtain a downsampled video frame, and the downsampled video frame is encoded to obtain the first encoded data. The downsampled video frame represents the video frame obtained by downsampling the target video frame. The first decoded video frame is a video frame obtained by upsampling the decoded data corresponding to the first encoded data, and the resolution of the first decoded video frame is the same as the resolution of the target video frame.

[0087] In some cases where the decoding end is a client, it retrieves the first and second encoded data sent by the server. Alternatively, it retrieves the first and second encoded data from other clients in the communication connection. In some cases where the decoding end is a server, it retrieves the first and second encoded data uploaded by the client.

[0088] S320. Decode the first encoded data to obtain the first decoded video frame of the target video frame.

[0089] For example, the first encoded data is decoded to obtain the decoded data of the target video frame, and the decoded data is upsampled according to the resolution of the target video frame to obtain the first decoded video frame.

[0090] Specifically, the decoder decodes the first encoded data using the same video encoding standard as the encoding operation to obtain decoded data. The decoded data is then input to an upsampling module, which upsamples the decoded data according to the resolution of the target video frame, outputting the first decoded video frame.

[0091] Optionally, if the client device does not have a neural network processor, the first decoded video frame can be used as the decoding result of the target video frame to avoid excessive CPU resource consumption. If the current CPU utilization is less than a set utilization threshold, the CPU can be used to decode the second encoded data. If the user device has a neural network processor, the neural network processor is used to decode the second encoded data.

[0092] S330. The pixel difference features are determined by a pre-trained decoding model based on the first decoded video frame and the second encoded data.

[0093] The encoding model and the decoding model are neural network models trained on sample pairs consisting of sample frames in the training sample set and their corresponding decoded sample frames. The decoding model includes a feature extraction network, an entropy model, and a decoder network.

[0094] For example, the first decoded video frame is input into the feature extraction network, the feature extraction network obtains the pixel features of the first decoded video frame, and the pixel features of the first decoded video frame are output to the entropy model. The second encoded data is input into the entropy model, the entropy model determines the pixel difference features based on the second encoded data and the pixel features of the first decoded video frame, and the pixel difference features are output to the decoder network.

[0095] In some embodiments, a first decoded video frame is input into a feature extraction network, which outputs the contextual features of the first decoded video frame to a decoder network and an entropy model. Second encoded data is input into the entropy model, which determines the codeword length associated with the pixel difference features based on the prior parameters in the second encoded data and the contextual features of the first decoded video frame. The second encoded data is then decoded according to the codeword length to obtain the pixel difference features corresponding to the second encoded data.

[0096] Specifically, a priori parameters are determined based on the second encoded data, and these parameters are input into the entropy model. The entropy model then determines the codeword length associated with the pixel difference features based on the priori parameters and contextual features. The second encoded data is then decoded according to the codeword length to obtain the pixel difference features corresponding to the second encoded data.

[0097] S340. Generate a second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

[0098] For example, a pixel difference image is generated by the decoder network based on the pixel difference features. The pixel difference image is then superimposed onto the first decoded video frame to obtain the second decoded video frame.

[0099] Since pixel difference features represent differences within the feature domain, multi-scale feature transformations are required to obtain a pixel difference image. For example, the pixel difference features and the context features of the first decoded video frame are input into the decoder network. The decoder network includes convolutional layers, inverse generalized divisive normalization (IGDN) layers, and residual blocks. Using the context features of the first decoded video frame as a condition, the decoder network performs multi-scale feature transformations on the pixel difference features through convolutional layers, IGDN layers, and residual blocks, outputting the transformed pixel difference image. By superimposing the pixel difference image onto the first decoded video frame, a second decoded video frame of the target video frame is obtained. Because the pixel difference image retains more detailed information about the target video frame, superimposing it onto the first decoded video frame yields a second decoded video frame with better image quality.

[0100] Figure 4 This is a schematic diagram of the video decoding framework in a video decoding method provided in an embodiment of this disclosure. Figure 4As shown, the video decoding framework includes a base layer 410 and an enhancement layer 420. The base layer 410 includes a decoder 430 and an upsampling module 440. The enhancement layer includes a decoding model 460. The decoding model 460 includes a feature extraction network 450, a decoder network 461, and an entropy model 462. The decoder 430 decodes the first encoded data 470 to obtain decoded data 480. The upsampling module 440 upsamples the decoded data 480 based on the resolution of the target video frame to obtain a first decoded video frame 490. The feature extraction network 450 extracts features from the first decoded video frame 490 to obtain the context features of the first decoded video frame 490. The context features are input to the decoder network 461 and the entropy model 462, respectively. The second encoded data 400 undergoes entropy decoding processing through the entropy model 462 to obtain pixel difference features. Pixel difference features and context features are input into decoder network 461. Decoder network 461 performs feature transformation on pixel difference features based on context features to transform pixel difference features from the feature domain to the pixel domain, resulting in pixel difference image 4100. Pixel difference image 4100 and first decoded video frame 490 can be superimposed using addition operator 4110 to obtain high-resolution second decoded video frame 4120. Specifically, for any pixel in the first decoded video frame, the pixel information corresponding to the pixel difference image 4100 is superimposed onto the pixel to obtain the second decoded video frame.

[0101] Optionally, after generating the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame, the method further includes: generating a decoded video based on the second decoded video frame and a timestamp of each target video frame. For example, in a scenario where a client shares a video with other clients, the decoding end can save the second decoded video frame after decoding it, and then generate a decoded video based on the second decoded video frame and timestamp of each target video frame for subsequent display and other processing. In a scenario where a client uploads a video to a server, the decoding end can save the second decoded video frame after decoding it, and then generate a decoded video based on the second decoded video frame and timestamp of each target video frame for subsequent download and other processing. In a scenario where a client downloads a video from a server, the decoding end can save the second decoded video frame after decoding it, and then generate a decoded video based on the second decoded video frame and timestamp of each target video frame for subsequent display and other processing.

[0102] Optionally, the second decoded video frame can be displayed. For example, in a cloud gaming scenario, after the decoding end decodes the second decoded video frame, it displays the interface video frame on the screen.

[0103] This embodiment of the disclosure obtains first encoded data and second encoded data of a target video frame, decodes and upsamples the first encoded data to obtain a first decoded video frame. Then, a decoding model determines pixel difference features based on the second encoded data and the first decoded video frame. Since the pixel difference features retain more detailed information of the target video frame, a second decoded video frame is generated based on the pixel difference features and the first decoded video frame, resulting in a second decoded video frame with better image quality. This improves the image quality of the decoded video frame and enhances the user experience.

[0104] Figure 5 This is a schematic diagram of a video encoding device provided in an embodiment of the present disclosure. The device can be implemented in software and / or hardware, and optionally, it can be implemented using an electronic device, such as a mobile terminal, a PC, or a server. Figure 5 As shown, the device includes:

[0105] The first acquisition module 510 is used to acquire at least one video frame to be encoded and to determine a target video frame based on the at least one video frame.

[0106] The first decoding module 520 is configured to generate first encoded data and a first decoded video frame of the target video frame based on the target video frame, wherein the first encoded data represents the encoded data of the target video frame;

[0107] The encoding data generation module 530 is used to determine the pixel difference features between the target video frame and the first decoded video frame through a pre-trained encoding model, and generate second encoding data based on the first decoded video frame and the pixel difference features, wherein the second encoding data represents the encoding data of the pixel difference features;

[0108] The encoding data sending module 540 is used to send the first encoded data and the second encoded data to the decoding end, so that the decoding end can decode the first encoded data to obtain the first decoded video frame, determine the pixel difference features based on the first decoded video frame and the second encoded data through a pre-trained decoding model, and generate the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

[0109] Optionally, the first acquisition module 510 is specifically used for:

[0110] Obtain the target video to be encoded, and determine the target video frame based on the timestamp of the video frame in the target video;

[0111] Alternatively, the target operation instruction sent by the decoding end can be obtained, and a target video frame can be generated according to the target operation instruction, wherein the target operation instruction is generated based on the user interaction operation in the decoding end.

[0112] Optionally, the first decoding module 520 is specifically used for:

[0113] The target video frame is downsampled to obtain a downsampled video frame;

[0114] The downsampled video frame is encoded to obtain the first encoded data;

[0115] The first encoded data is decoded to obtain the decoded data of the target video frame, and the decoded data is upsampled according to the resolution of the target video frame to obtain the first decoded video frame.

[0116] Optionally, the encoded data generation module 530 is specifically used for:

[0117] The pixel difference features are determined by the encoding model based on the pixel features of the target video frame and the first decoded video frame;

[0118] The encoding model determines the codeword length associated with the pixel difference features based on the pixel difference features and the first decoded video frame, and encodes the pixel difference features according to the codeword length to obtain the second encoded data.

[0119] Optionally, the encoding model and the decoding model are neural network models trained on sample pairs consisting of sample frames in the training sample set and the corresponding decoded sample frames. The encoding model includes a feature extraction network, an encoder network, and an entropy model.

[0120] Further, determining the pixel difference features based on the pixel features of the target video frame and the first decoded video frame using the encoding model includes:

[0121] The first decoded video frame is input into the feature extraction network, the pixel features of the first decoded video frame are obtained through the feature extraction network, and the pixel features of the first decoded video frame are output to the encoder network and the entropy model.

[0122] The target video frame is input into the encoder network, and the pixel features of the target video frame are obtained through the encoder network. Based on the pixel features of the target video frame and the pixel features of the first decoded video frame, the pixel difference features are determined, and the pixel difference features are output to the entropy model.

[0123] Optionally, the step of determining the codeword length associated with the pixel difference features based on the pixel difference features and the first decoded video frame using the encoding model, and encoding the pixel difference features according to the codeword length to obtain the second encoded data, includes:

[0124] The pixel difference features are input into the entropy model in the encoding model. The entropy model determines the codeword length associated with the pixel difference features based on the pixel difference features and the pixel features of the first decoded video frame. The pixel difference features are then encoded according to the codeword length, and the second encoded data is output.

[0125] The video encoding apparatus provided in this disclosure can execute the video encoding method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0126] Figure 6 This is a schematic diagram of a video decoding device provided in an embodiment of the present disclosure. The device can be implemented in software and / or hardware, and optionally, it can be implemented using an electronic device, such as a mobile terminal, a PC, or a server. Figure 6 As shown, the device includes:

[0127] The second acquisition module 610 is used to acquire first encoded data and second encoded data sent by the encoding end. The second encoded data is obtained by the encoding end encoding the pixel difference features of the target video frame and the first decoded video frame of the target video frame. The first encoded data represents the encoded data of the target video frame.

[0128] The second decoding module 620 is used to decode the first encoded data to obtain the first decoded video frame of the target video frame;

[0129] The feature determination module 630 is used to determine the pixel difference features based on the first decoded video frame and the second encoded data using a pre-trained decoding model;

[0130] The video frame determination module 640 is used to generate a second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

[0131] Optionally, it also includes:

[0132] After generating the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame, a decoded video is generated based on the second decoded video frame and the timestamp of each target video frame.

[0133] Alternatively, display the second decoded video frame.

[0134] Optionally, the second decoding module 620 is specifically used for:

[0135] The first encoded data is decoded to obtain the decoded data of the target video frame, and the decoded data is upsampled according to the resolution of the target video frame to obtain the first decoded video frame.

[0136] Optionally, the decoding model includes a feature extraction network, an entropy model, and a decoder network.

[0137] Optionally, the feature determination module 630 is specifically used for:

[0138] The first decoded video frame is input into the feature extraction network, the pixel features of the first decoded video frame are obtained through the feature extraction network, and the pixel features of the first decoded video frame are output to the entropy model.

[0139] The second encoded data is input into the entropy model, and the entropy model determines the pixel difference features based on the pixel features of the second encoded data and the first decoded video frame, and outputs the pixel difference features to the decoder network.

[0140] Optionally, the decoded video frame determination module 640 is specifically used for:

[0141] The decoder network generates a pixel difference image based on the pixel difference features.

[0142] The pixel difference image is superimposed onto the first decoded video frame to obtain the second decoded video frame.

[0143] The video decoding apparatus provided in this disclosure can execute the video decoding method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0144] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0145] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 700. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0146] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.

[0147] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0148] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.

[0149] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0150] The electronic device provided in this embodiment belongs to the same inventive concept as the video encoding method or video decoding method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0151] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video encoding or video decoding method provided in the above embodiments.

[0152] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0153] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0154] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0155] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire at least one video frame to be encoded, and determine a target video frame based on the at least one video frame;

[0156] First encoded data and a first decoded video frame of the target video frame are generated based on the target video frame, wherein the first encoded data represents the encoded data of the target video frame;

[0157] The pixel difference features between the target video frame and the first decoded video frame are determined by a pre-trained encoding model. Second encoded data is generated based on the first decoded video frame and the pixel difference features, wherein the second encoded data represents the encoded data of the pixel difference features.

[0158] The first encoded data and the second encoded data are sent to the decoding end so that the decoding end can decode the first encoded data to obtain the first decoded video frame. The pixel difference features are determined by the pre-trained decoding model based on the first decoded video frame and the second encoded data. The second decoded video frame of the target video frame is generated based on the pixel difference features and the first decoded video frame.

[0159] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire first encoded data and second encoded data sent by the encoding end, wherein the second encoded data is obtained by the encoding end encoding pixel difference features of a target video frame and a first decoded video frame of the target video frame, wherein the first encoded data characterizes the encoded data of the target video frame;

[0160] Decoding the first encoded data yields the first decoded video frame of the target video frame;

[0161] The pixel difference features are determined by a pre-trained decoding model based on the first decoded video frame and the second encoded data.

[0162] A second decoded video frame of the target video frame is generated based on the pixel difference features and the first decoded video frame.

[0163] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0165] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0166] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0167] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0168] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0169] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0170] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video encoding method, characterized in that, The method is applied to the encoding end and includes: Obtain at least one video frame to be encoded, and determine the target video frame based on the at least one video frame; First encoded data and a first decoded video frame of the target video frame are generated based on the target video frame, wherein the first encoded data represents the encoded data of the target video frame; The pixel difference features between the target video frame and the first decoded video frame are determined by a pre-trained encoding model. Second encoded data is generated based on the first decoded video frame and the pixel difference features, wherein the second encoded data represents the encoded data of the pixel difference features. The first encoded data and the second encoded data are sent to the decoding end so that the decoding end can decode the first encoded data to obtain the first decoded video frame. The pixel difference features are determined by the pre-trained decoding model based on the first decoded video frame and the second encoded data. The second decoded video frame of the target video frame is generated based on the pixel difference features and the first decoded video frame.

2. The method according to claim 1, characterized in that, The step of acquiring at least one video frame to be encoded and determining a target video frame based on the at least one video frame includes: Obtain the target video to be encoded, and determine the target video frame based on the timestamp of the video frame in the target video; Alternatively, the target operation instruction sent by the decoding end can be obtained, and a target video frame can be generated according to the target operation instruction, wherein the target operation instruction is generated based on the user interaction operation in the decoding end.

3. The method according to claim 1, characterized in that, The step of generating first encoded data and a first decoded video frame based on the target video frame includes: The target video frame is downsampled to obtain a downsampled video frame; The downsampled video frame is encoded to obtain the first encoded data; The first encoded data is decoded to obtain the decoded data of the target video frame, and the decoded data is upsampled according to the resolution of the target video frame to obtain the first decoded video frame.

4. The method according to claim 1, characterized in that, The step of determining the pixel difference features between the target video frame and the first decoded video frame using a pre-trained encoding model, and generating second encoded data based on the first decoded video frame and the pixel difference features, includes: The pixel difference features are determined by the encoding model based on the pixel features of the target video frame and the first decoded video frame; The encoding model determines the codeword length associated with the pixel difference features based on the pixel difference features and the first decoded video frame, and encodes the pixel difference features according to the codeword length to obtain the second encoded data.

5. The method according to claim 4, characterized in that, The encoding model and the decoding model are neural network models trained on sample pairs consisting of sample frames in the training sample set and the corresponding decoded sample frames. The encoding model includes a feature extraction network, an encoder network, and an entropy model. The step of determining the pixel difference features based on the pixel features of the target video frame and the first decoded video frame using the encoding model includes: The first decoded video frame is input into the feature extraction network, the pixel features of the first decoded video frame are obtained through the feature extraction network, and the pixel features of the first decoded video frame are output to the encoder network and the entropy model. The target video frame is input into the encoder network, and the pixel features of the target video frame are obtained through the encoder network. Based on the pixel features of the target video frame and the pixel features of the first decoded video frame, the pixel difference features are determined, and the pixel difference features are output to the entropy model.

6. The method according to claim 4, characterized in that, The step of determining the codeword length associated with the pixel difference features based on the pixel difference features and the first decoded video frame using the encoding model, and encoding the pixel difference features according to the codeword length to obtain the second encoded data, includes: The pixel difference features are input into the entropy model in the encoding model. The entropy model determines the codeword length associated with the pixel difference features based on the pixel difference features and the pixel features of the first decoded video frame. The pixel difference features are then encoded according to the codeword length, and the second encoded data is output.

7. A video decoding method, characterized in that, The method is applied at the decoding end and includes: The encoding end sends first encoded data and second encoded data. The second encoded data is obtained by the encoding end encoding the pixel difference features of the target video frame and the first decoded video frame of the target video frame. The first encoded data represents the encoded data of the target video frame. Decoding the first encoded data yields the first decoded video frame of the target video frame; The pixel difference features are determined by a pre-trained decoding model based on the first decoded video frame and the second encoded data. A second decoded video frame of the target video frame is generated based on the pixel difference features and the first decoded video frame.

8. The method according to claim 7, characterized in that, After generating the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame, the method further includes: A decoded video is generated based on the second decoded video frame and timestamp of each of the target video frames; Alternatively, display the second decoded video frame.

9. The method according to claim 7, characterized in that, The first decoded video frame obtained by decoding the first encoded data to obtain the target video frame includes: The first encoded data is decoded to obtain the decoded data of the target video frame, and the decoded data is upsampled according to the resolution of the target video frame to obtain the first decoded video frame.

10. The method according to claim 7, characterized in that, The decoding model includes a feature extraction network, an entropy model, and a decoder network; The step of determining the pixel difference features using a pre-trained decoding model based on the first decoded video frame and the second encoded data includes: The first decoded video frame is input into the feature extraction network, the pixel features of the first decoded video frame are obtained through the feature extraction network, and the pixel features of the first decoded video frame are output to the entropy model. The second encoded data is input into the entropy model, and the entropy model determines the pixel difference features based on the pixel features of the second encoded data and the first decoded video frame, and outputs the pixel difference features to the decoder network.

11. The method according to claim 10, characterized in that, The step of generating the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame includes: The decoder network generates a pixel difference image based on the pixel difference features. The pixel difference image is superimposed onto the first decoded video frame to obtain the second decoded video frame.

12. A video encoding device, characterized in that, The device is applied to the encoding end and includes: The first acquisition module is used to acquire at least one video frame to be encoded, and to determine a target video frame based on the at least one video frame. A first decoding module is configured to generate first encoded data and a first decoded video frame of the target video frame based on the target video frame, wherein the first encoded data represents the encoded data of the target video frame; The encoded data generation module is used to determine the pixel difference features between the target video frame and the first decoded video frame through a pre-trained encoding model, and generate second encoded data based on the first decoded video frame and the pixel difference features, wherein the second encoded data represents the encoded data of the pixel difference features; The encoded data sending module is used to send the first encoded data and the second encoded data to the decoding end, so that the decoding end can decode the first encoded data to obtain the first decoded video frame, determine the pixel difference features based on the first decoded video frame and the second encoded data through a pre-trained decoding model, and generate the second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

13. A video decoding device, characterized in that, The device is used at the decoding end and includes: The second acquisition module is used to acquire first encoded data and second encoded data sent by the encoding end. The second encoded data is obtained by the encoding end encoding the pixel difference features of the target video frame and the first decoded video frame of the target video frame. The first encoded data represents the encoded data of the target video frame. The second decoding module is used to decode the first encoded data to obtain the first decoded video frame of the target video frame; A feature determination module is used to determine the pixel difference features based on the first decoded video frame and the second encoded data using a pre-trained decoding model; The video frame determination module is used to generate a second decoded video frame of the target video frame based on the pixel difference features and the first decoded video frame.

14. An electronic device, characterized in that, The electronic device includes: One or more processors, wherein the processors include a central processing unit and a neural network processor; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-11.

15. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the method as described in any one of claims 1-11.

16. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1-11.