Coding method and apparatus, electronic device, and computer-readable storage medium

By updating the model parameters to adapt to the differences in multimedia data, efficient encoding and quality improvement of multimedia data is achieved, and the problem of insufficient quality of multimedia data compression and decoding in the prior art is solved.

CN117915106BActive Publication Date: 2025-05-27SHUXING TECH (BEIJING) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410070532.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-05-27
Estimated Expiration
2044-01-17

AI Technical Summary

Technical Problem

In the multimedia data distribution scenario, it is difficult for the prior art to effectively compress multimedia data and improve the quality of multimedia data decoded by client.

Method used

By acquiring the target multimedia data and the first encoding result, decoding the first encoding result using the second model, calculating the difference between the target multimedia data and the first decoding result, and updating the parameters of the first model to obtain the third model. Then, the target multimedia data is encoded using the third model to obtain the second encoding result.

Benefits of technology

This method not only reduces the code rate of the second encoding result, but also improves the quality of the multimedia data when the client decoder decodes the second encoding result.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117915106B_ABST
    Figure CN117915106B_ABST
Patent Text Reader

Abstract

The present application discloses an encoding method and apparatus, an electronic device, and a computer-readable storage medium. The method is used to encode target multimedia data before sending it to a client, and the method includes: obtaining the target multimedia data and a first encoding result, where the first encoding result is obtained by encoding the target multimedia data using a first model, the first model has trainable parameters, and the first model is used to encode multimedia data; decoding the first encoding result using a second model to obtain a first decoding result, where the second model is the same as the decoding model deployed on the client, and the decoding model is used to decode the encoding result of the multimedia data; updating the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain a third model; and encoding the target multimedia data using the third model to obtain a second encoding result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of encoding technologies, and in particular, to an encoding method and apparatus, an electronic device, and a computer-readable storage medium. Background Art

[0002] In the scenario of multimedia data distribution, due to the large amount of multimedia data, it is usually necessary to encode the multimedia data to compress it, and then send the encoding result to the client. Therefore, how to encode multimedia data and send it to the client is of great significance. Summary of the Invention

[0003] This application provides an encoding method and apparatus, an electronic device, and a computer-readable storage medium to send multimedia data to a client.

[0004] In a first aspect, an encoding method is provided. The method is used to encode target multimedia data before sending it to a client. The method includes:

[0005] Obtain target multimedia data and a first encoding result, where the first encoding result is obtained by encoding the target multimedia data using a first model. The first model has trainable parameters and is used to encode multimedia data;

[0006] Decode the first encoding result using a second model to obtain a first decoding result. The second model is the same as the decoding model deployed on the client, and the decoding model is used to decode the encoding result of multimedia data;

[0007] Based on the difference between the target multimedia data and the first decoding result, update the parameters of the first model to obtain a third model;

[0008] Encode the target multimedia data using the third model to obtain a second encoding result.

[0009] In combination with any implementation manner of this application, the first model and the second model are trained using at least two reference multimedia data in a dataset. Among them, when encoding the at least two reference multimedia data using the first model and decoding the encoding results of the at least two reference multimedia data using the second model to obtain at least two second decoding results, the difference between the at least two reference multimedia data and the at least two second decoding results is less than or equal to a preset value.

[0010] Combined with any implementation manner of the present application, updating the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain a third model includes:

[0011] Determine the first mean square error between the target multimedia data and the first decoding result;

[0012] Based on the first mean square error, update the parameters of the first model to obtain a fourth model;

[0013] Train the fourth model for n rounds based on the target multimedia data and the second model to obtain n second mean square errors;

[0014] Determine the minimum value between the first mean square error and the n second mean square errors to obtain a target mean square error;

[0015] Use the model that obtains the target mean square error as the third model.

[0016] Combined with any implementation manner of the present application, before training the fourth model for n rounds based on the target multimedia data and the second model to obtain n second mean square errors, the method further includes:

[0017] Obtain the network speed of the client;

[0018] Based on the network speed, determine n, and n is positively correlated with the network speed.

[0019] Combined with any implementation manner of the present application, after obtaining the second encoding result, the method further includes:

[0020] Send the second encoding result to the client.

[0021] Combined with any implementation manner of the present application, before using the second model to decode the first encoding result to obtain a first decoding result, the method further includes:

[0022] Obtain the decoding performance index of the client, where the decoding performance index characterizes the performance of the decoding resources in the client, and the decoding resources are the performance of the resources used for decoding;

[0023] Based on the decoding performance index, change the parameters of the second model and / or change the structure of the second model to obtain a fifth model, and the requirements of running the fifth model for the performance of the decoding resources match the decoding performance index;

[0024] The step of using the second model to decode the first encoding result to obtain a first decoding result includes:

[0025] Decode the first encoding result by using the fifth model to obtain the first decoding result.

[0026] Combined with any implementation manner of the present application, the target multimedia data includes a first image, the first model is used to encode the image, and the second model is used to decode the encoding result of the image;

[0027] Obtaining the first encoding result includes:

[0028] Encode the first foreground image by using the first model to obtain a third encoding result, where the first foreground image is a pixel region covered by the foreground in the first image;

[0029] The step of using the second model to decode the first encoding result to obtain the first decoding result includes:

[0030] Decode the third encoding result by using the second model to obtain the first decoding result;

[0031] The step of updating the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain a third model includes:

[0032] Update the parameters of the first model based on the difference between the first foreground image and the first decoding result to obtain a first foreground encoder;

[0033] The step of using the third model to encode the target multimedia data to obtain a second encoding result includes:

[0034] Encode the background image by using the first model to obtain a fourth encoding result, where the background image is a pixel region in the first image other than the first foreground image;

[0035] Encode the first foreground image by using the first foreground encoder to obtain a fifth encoding result;

[0036] Based on the fourth encoding result and the fifth encoding result, obtain the second encoding result.

[0037] Combined with any implementation manner of the present application, the first image is a frame image in a live video stream, the live video stream is a video stream obtained by video-live streaming a target scene, and the foreground is a pixel region in the first image other than the target scene;

[0038] The step of encoding the background image by using the first model to obtain a fourth encoding result includes:

[0039] The background encoder is used to encode the background image to obtain the fourth encoding result. The background encoder is obtained by training the first model using a scene image and the second model, and the scene image is an image captured of the target scene.

[0040] Combined with any implementation manner of the present application, the live video stream further includes a second image, the timestamp of the second image is greater than the timestamp of the first image, and the difference between the timestamp of the first image and the timestamp of the second image is less than or equal to a threshold value;

[0041] The method further includes:

[0042] The first model is used to encode a second foreground image to obtain a sixth encoding result, and the second foreground image is a pixel region covered by the foreground in the second image;

[0043] The second model is used to decode the sixth encoding result to obtain a third decoding result;

[0044] Based on the difference between the second foreground image and the third decoding result, the parameters of the first model are updated to obtain a second foreground encoder;

[0045] The second foreground encoder is used to encode the second foreground image to obtain a seventh encoding result;

[0046] Based on the fourth encoding result and the seventh encoding result, an eighth encoding result of the second image is obtained.

[0047] In a second aspect, an encoding device is provided. The encoding device is used to encode target multimedia data before sending it to a client. The encoding device includes:

[0048] An acquisition unit, configured to acquire target multimedia data and a first encoding result. The first encoding result is obtained by encoding the target multimedia data using a first model. The first model has trainable parameters, and the first model is used to encode multimedia data;

[0049] A decoding unit, configured to decode the first encoding result using a second model to obtain a first decoding result. The second model is the same as the decoding model deployed on the client, and the decoding model is used to decode the encoding result of multimedia data;

[0050] An update unit, configured to update the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain a third model;

[0051] An encoding unit, configured to encode the target multimedia data by using the third model to obtain a second encoding result.

[0052] Combined with any embodiment of the present application, the updating unit is configured to:

[0053] Determine a first mean square error between the target multimedia data and the first decoding result;

[0054] Update the parameters of the first model based on the first mean square error to obtain a fourth model;

[0055] Train the fourth model for n rounds based on the target multimedia data and the second model to obtain n second mean square errors;

[0056] Determine the minimum value between the first mean square error and the n second mean square errors to obtain a target mean square error;

[0057] Use the model that obtains the target mean square error as the third model.

[0058] Combined with any embodiment of the present application, the obtaining unit is further configured to obtain the network speed of the client;

[0059] The encoding device further includes: a determining unit, configured to determine the n based on the network speed, and the n is positively correlated with the network speed.

[0060] Combined with any embodiment of the present application, the encoding device further includes: a sending unit, configured to send the second encoding result to the client.

[0061] Combined with any embodiment of the present application, the obtaining unit is further configured to obtain a decoding performance index of the client, where the decoding performance index characterizes the performance of decoding resources in the client, and the decoding resources are the performance of resources for decoding;

[0062] The encoding device further includes: a changing unit, further configured to change the parameters of the second model and / or change the structure of the second model based on the decoding performance index to obtain a fifth model, and the requirements for the performance of the decoding resources when running the fifth model match the decoding performance index;

[0063] The use of the decoding unit is specifically configured to:

[0064] Decode the first encoding result by using the fifth model to obtain the first decoding result.

[0065] Combined with any embodiment of the present application, the target multimedia data includes a first image, the first model is used to encode the image, and the second model is used to decode the encoding result of the image;

[0066] The obtaining unit is specifically configured to:

[0067] Encode the first foreground image by using the first model to obtain a third encoding result, where the first foreground image is a pixel region covered by the foreground in the first image;

[0068] The decoding unit is specifically configured to:

[0069] Decode the third encoding result by using the second model to obtain the first decoding result;

[0070] The updating unit is specifically configured to:

[0071] Update the parameters of the first model based on the difference between the first foreground image and the first decoding result to obtain a first foreground encoder;

[0072] The encoding unit is specifically configured to:

[0073] Encode the background image by using the first model to obtain a fourth encoding result, where the background image is a pixel region in the first image other than the first foreground image;

[0074] Encode the first foreground image by using the first foreground encoder to obtain a fifth encoding result;

[0075] Obtain the second encoding result based on the fourth encoding result and the fifth encoding result.

[0076] Combined with any implementation manner of the present application, the first image is a frame image in a live video stream, the live video stream is a video stream obtained by video-livecasting a target scene, and the foreground is a pixel region in the first image other than the target scene;

[0077] The encoding unit is specifically configured to:

[0078] Encode the background image by using a background encoder to obtain the fourth encoding result, where the background encoder is obtained by training the first model by using a scene image and the second model, and the scene image is an image taken of the target scene.

[0079] Combined with any implementation manner of the present application, the live video stream further includes a second image, the timestamp of the second image is greater than the timestamp of the first image, and the difference between the timestamp of the first image and the timestamp of the second image is less than or equal to a threshold;

[0080] The encoding unit is further configured to encode a second foreground image by using the first model to obtain a sixth encoding result, where the second foreground image is a pixel region covered by the foreground in the second image;

[0081] The decoding unit is further configured to decode the sixth encoding result by using the second model to obtain a third decoding result;

[0082] The updating unit is further configured to update parameters of the first model based on a difference between the second foreground image and the third decoding result to obtain a second foreground encoder;

[0083] The encoding unit is further configured to encode the second foreground image by using the second foreground encoder to obtain a seventh encoding result;

[0084] The encoding unit is further configured to obtain an eighth encoding result of the second image based on the fourth encoding result and the seventh encoding result.

[0085] In a third aspect, an electronic device is provided, including: a processor and a memory, where the memory is configured to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method according to the first aspect and any one of its possible implementation manners as described above.

[0086] In a fourth aspect, another electronic device is provided, including: a processor, a sending device, an input device, an output device, and a memory, where the memory is configured to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the first aspect and any one of its embodiments as described above.

[0087] In a fifth aspect, a computer-readable storage medium is provided, where a computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the first aspect and any one of its embodiments as described above.

[0088] In a sixth aspect, a computer program product is provided, where the computer program product includes a computer program or instructions. When the computer program or instructions run on a computer, the computer is caused to execute the first aspect and any one of its embodiments as described above.

[0089] It should be understood that the above general description and subsequent detailed description are merely exemplary and explanatory, and do not limit this application.

[0090] In this application, before sending target multimedia data to a client, an encoding device obtains the target multimedia data and a first encoding result, where the first encoding result is obtained by encoding the target multimedia data using a first model. The first encoding result is decoded using a second model to obtain a first decoding result, where the second model is the same as the decoding model deployed on the client. Then, based on the difference between the target multimedia data and the first decoding result, the parameters of the first model are updated to obtain a third model. In this way, the first model can be specifically trained using the target multimedia data to obtain the third model. The target multimedia data is then encoded using the third model to obtain a second encoding result. This can not only reduce the bit rate of the second encoding result, but also improve the quality of the decoded multimedia data when the decoder on the client decodes the second encoding result. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the following will describe the drawings required for use in the embodiments of this application or the background art.

[0092] The drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with this application and are used together with the specification to illustrate the technical solutions of this application.

[0093] Figure 1 It is a schematic diagram of the architecture of a multimedia data distribution system provided by an embodiment of this application;

[0094] Figure 2 It is a schematic flowchart of an encoding method provided by an embodiment of this application;

[0095] Figure 3 It is a schematic diagram of the architecture for training a first model and a second model provided by an embodiment of this application;

[0096] Figure 4 It is a schematic flowchart of encoding an image provided by an embodiment of this application;

[0097] Figure 5 It is a schematic diagram of the structure of an encoding device provided by an embodiment of this application;

[0098] Figure 6 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0099] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0100] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0101] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0102] In the scenario of multimedia data distribution, since the amount of target multimedia data to be distributed is relatively large, it is usually necessary to encode the target multimedia data to compress it, and then distribute the encoding result to the client. Please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of a multimedia data distribution system provided by an embodiment of this application. As Figure 1 shown, the multimedia data distribution system 1 includes a client 11, a client 12, and a server 13. Among them, there is a communication connection between both the client 11 and the client 12 and the server 13. Both the client 11 and the client 12 can upload multimedia data to the server 13 through this communication connection, and the server 13 can also distribute multimedia data to the client 11 and the client 12 through this communication connection.

[0103] Optionally, both the client 11 and the client 12 can be one of the following: a mobile phone, a computer, a tablet computer, a wearable intelligent device. For example, the client 11 is a mobile phone and the client 12 is a computer. Another example is that both the client 11 and the client 12 are tablet computers. Optionally, the server 13 is a server.

[0104] It should be understood thatFigure 1 The shown clients 11 and 12 are only examples and should not be understood that there can only be 2 clients having communication connections with the server 13. In actual applications, the number of clients having communication connections with the server 13 can be m, where m is a positive integer.

[0105] The embodiment of the present application provides an encoding method, which is used to encode target multimedia data before sending it to a client. The execution subject of the encoding method is an encoding device, where the encoding device can be the server 13 in the above multimedia data distribution system 1. The encoding device can be any electronic device that can execute the technical solutions disclosed in the method embodiments of the present application. Optionally, the encoding device can be one of the following: a computer, a server.

[0106] It should be understood that the method embodiments of the present application can also be implemented by a processor executing computer program code. The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an encoding method provided by the embodiment of the present application.

[0107] 201. Obtain the above target multimedia data and the first encoding result.

[0108] In the embodiments of the present application, the target multimedia data may be multimedia data, where the multimedia data includes: images, videos, and audios. The first encoding result is obtained by encoding the target multimedia data using a first model, where the first model has trainable parameters, and the first model is used to encode multimedia data. The trainable parameters are parameters that can be updated through training. Since the first model has trainable parameters, the trainable parameters in the first model can be updated by training the first model. In a possible implementation manner, the multimedia data is an image, and the first model is a deep learning model. By training the first model, the trainable parameters in the first model can be updated, so that the first model has the ability to encode images. In another possible implementation manner, the multimedia data is a video, and the first model is a deep learning model. By training the first model, the trainable parameters in the first model can be updated, so that the first model has the ability to encode videos. Optionally, the first model is an encoder in a variational auto encoder (VAE). At this time, the trainable parameters include the parameters in the encoder of the VAE. By updating the trainable parameters, the ability of the encoder of the VAE to encode multimedia data can be changed. Specifically, when the multimedia data input to the encoder of the VAE remains unchanged, if the parameters in the encoder of the VAE are changed by training the encoder of the VAE, the encoding result of the multimedia data output by the encoder of the VAE will change. Optionally, the first model is a convolutional neural network (CNN). For example, the first model includes a convolutional layer for performing convolutional processing, a pooling layer for performing pooling processing, and a fully connected layer (FCL). The trainable parameters of the CNN include the parameters in the convolutional layer, the parameters in the pooling layer, and the parameters in the FCL. When the multimedia data input to the CNN remains unchanged, if the parameters in the convolutional layer, the parameters in the pooling layer, and the parameters in the FCL are changed by training the CNN, the encoding result of the multimedia data output by the CNN will change.

[0109] In a possible implementation manner of obtaining the target multimedia data, the encoding device receives the target multimedia data input by the user through the input component, where the input component includes: a keyboard, a mouse, a touch screen, a touch pad, and an audio input device.

[0110] In another possible implementation manner of obtaining the target multimedia data, the encoding device receives the target multimedia data sent by the terminal. Optionally, the terminal may be any one of the following: a mobile phone, a computer, a tablet computer, and a server.

[0111] In an implementation of obtaining a first encoding result, an encoding device receives the first encoding result input by a user through an input component.

[0112] In another implementation of obtaining a first encoding result, the encoding device receives the first encoding result sent by a terminal.

[0113] In yet another implementation of obtaining a first encoding result, after obtaining target multimedia data, the encoding device encodes the target multimedia data using a first model to obtain the first encoding result.

[0114] It should be understood that in the embodiments of the present application, the steps of the encoding device obtaining the target multimedia data and obtaining the first encoding result may be executed separately or simultaneously, and the present application does not make any limitations thereto.

[0115] 202. Use a second model to decode the above first encoding result to obtain a first decoding result.

[0116] In the embodiments of the present application, the second model is the same as the decoding model deployed on the client side, where the decoding model is used to decode the encoding result of multimedia data, and the second model is also used to decode the encoding result of multimedia data. For example, when the multimedia data is an image, the second model is used to decode the encoding result of the image. Optionally, the second model has trainable parameters, and the trainable parameters in the second model can be updated by training the second model. In a possible implementation manner, the multimedia data is an image, and the second model is a deep learning model. By training the second model, the trainable parameters in the second model can be updated, so that the first model has the ability to encode images.

[0117] Optionally, the second model is the decoder in VAE. At this time, the trainable parameters include the parameters in the decoder of VAE. By updating the trainable parameters, the ability of the decoder of VAE to encode multimedia data can be changed. Specifically, when the encoding result input to the decoder of VAE remains unchanged, if the decoder of VAE is trained to change the parameters in the decoder of VAE, then the multimedia data output by the decoder of VAE will change. Optionally, the second model is a CNN. For example, the second model includes a convolutional layer for performing convolutional processing and an upsampling layer for performing upsampling. The trainable parameters of the CNN include the parameters in the convolutional layer and the parameters in the upsampling layer. When the encoding result input to the CNN remains unchanged, if the CNN is trained to change the parameters in the convolutional layer and the parameters in the upsampling layer, then the multimedia data output by the CNN will change.

[0118] For the client, after receiving the encoded result, the multimedia data can be obtained by decoding the encoded result using the second model. For example, the multimedia data is an image. The encoding device encodes image a to obtain the encoded result of image a, and then sends the encoded result of image a to the client. The client then decodes the encoded result of image a using the second model to obtain image b, where the content of image b is the same as the content of image a.

[0119] In a possible implementation, the first model and the second model are trained using at least two reference multimedia data in the dataset. Specifically, the reference image is encoded using the first model to obtain the encoded result of the reference image. Then, the second model decodes the encoded result of the reference multimedia data to obtain the second decoded result. Next, the mean square error (MSE) between the second decoded result and the reference multimedia data is calculated as the loss, and the parameters of the first model and the second model are updated according to this loss until the loss converges, completing the training of the first model and the second model.

[0120] Optionally, when the first model is a model for encoding an image and the second model is a model for decoding the encoded result of the image, Figure 3 This is a schematic diagram of an architecture for training the first model and the second model provided by an embodiment of the present application. As Figure 3 shown, the input of the first model is the reference image. After the first model processes the reference image, the feature map of the reference image can be obtained, where the feature map carries the feature information of the reference image. Then, the feature map of the reference image is quantized and arithmetically encoded in sequence to obtain the bitstream (bits) of the reference image. Among them, arithmetic coding is a type of entropy coding, which is a way of encoding different characters with different codes based on the probability of character occurrence in the data. Through arithmetic coding, the statistical redundancy in the data can be compressed and lossless compression of the data can be achieved. Then, the bitstream of the reference image is arithmetically decoded, and the result of the arithmetic decoding is input into the second model. After the second model decodes the result of the arithmetic decoding, a reconstructed image is obtained, where arithmetic decoding is the inverse process of arithmetic coding. Finally, the MSE between the reference image and the reconstructed image is calculated, and this MSE is used as the loss. Through backpropagation, the parameters of the second model and the first model are updated until the loss converges, completing the training of the first model and the training of the second model.

[0121] It should be understood that Figure 3 the specific structures of the first model and the second model are not shown. Any model with trainable parameters and used for encoding an image can be used as Figure 3 the first model in. Any model with trainable parameters and used for decoding the encoded result of an image can be used asFigure 3 The second model in

[0122] As an alternative implementation, when at least two reference multimedia data are encoded and decoded using the first model and the second model to obtain at least two second decoding results, the difference between the at least two reference multimedia data and the at least two second decoding results is less than or equal to a preset value. In this implementation, the fact that the difference between the at least two reference multimedia data and the at least two second decoding results is less than or equal to the preset value indicates that the difference between the at least two reference multimedia data and the at least two second decoding results is small, which also means that through training, the first model and the second model can achieve good encoding and decoding effects on the at least two reference multimedia data.

[0123] In a possible implementation, the fact that the difference between the at least two reference multimedia data and the at least two second decoding results is less than or equal to the preset value means that the MSE between any one of the reference multimedia data and the corresponding second decoded image is less than or equal to the preset value. In another possible implementation, the fact that the difference between the at least two reference multimedia data and the at least two second decoding results is less than or equal to the preset value means that the average value of the MSE between the reference multimedia data and the corresponding second decoded images is less than or equal to the preset value.

[0124] In step 202, the encoding device decodes the first encoding result using the second model to obtain a first decoding result.

[0125] 203. Update the parameters of the first model based on the difference between the above target multimedia data and the above first decoding result to obtain a third model.

[0126] The difference between the target multimedia data and the first decoding result represents the difference between the multimedia data obtained by decoding and the target multimedia data when the target multimedia data is encoded using the first model to obtain a first encoding result and the first encoding result is decoded using the second model. That is to say, if the encoding device encodes the target multimedia data using the first model to obtain a first encoding result and sends the first encoding result to the client, then the difference between the multimedia data obtained by the client decoding the first encoding result and the target multimedia data is the difference between the target multimedia data and the first decoding result. Therefore, the difference between the target multimedia data and the first decoding result can represent the quality of the multimedia data obtained by the client decoding the first encoding result when the encoding device encodes the target multimedia data using the first model.

[0127] Therefore, the encoding device updates the parameters of the first model based on the difference between the target multimedia data and the first decoding result, obtains a third model, and can use the third model to encode the target multimedia data to obtain a second encoding result. When the client decodes the second encoding result to obtain the first decoding result, the difference between the multimedia data decoded by the client and the target multimedia data can be reduced, and thus the quality of the multimedia data decoded by the client can be improved.

[0128] Based on this, the encoding device updates the parameters of the first model based on the difference between the target multimedia data and the first decoding result, and obtains a third model. In a possible implementation manner, the encoding device determines the loss of the first model based on the difference between the target multimedia data and the first decoding result, where the loss of the first model is positively correlated with the difference. Based on the loss of the first model, backpropagation is performed on the parameters of the first model.

[0129] Specifically, after determining the loss of the first model, the following formula can be used to calculate the gradient of the output layer in the first model, where the result output by the output layer is the result output by the first model:

[0130]

[0131] where L is the loss of the first model, y is the first decoding result, t is the target multimedia data, is the gradient of the output layer.

[0132] After determining the gradient of the output layer, the gradient of the layers in the first model other than the output layer can be calculated layer by layer forward according to the gradient of the output layer. After obtaining the gradient of each layer in the first model, the parameters of each layer can be updated according to the gradients of each layer and the gradient descent algorithm to obtain a third model.

[0133] 204. Use the above third model to encode the above target multimedia data to obtain a second encoding result.

[0134] Since the third model is trained using the target multimedia data, using the third model to encode the target multimedia data to obtain a second encoding result can improve the encoding effect. Specifically, using the second model to decode the second encoding result can improve the quality of the multimedia data decoded by the second model.

[0135] In the embodiments of the present application, before the encoding device sends the target multimedia data to the client, it obtains the target multimedia data and the first encoding result, where the first encoding result is obtained by encoding the target multimedia data using the first model. The first encoding result is decoded using the second model to obtain the first decoding result, where the second model is the same as the decoding model deployed on the client. Then, based on the difference between the target multimedia data and the first decoding result, the parameters of the first model are updated to obtain the third model. Thus, the first model can be specifically trained using the target multimedia data to obtain the third model. The third model is then used to encode the target multimedia data to obtain the second encoding result. This can not only reduce the bit rate of the second encoding result, but also improve the quality of the decoded multimedia data when the decoder of the client is used to decode the second encoding result.

[0136] As an optional implementation manner, after the encoding device obtains the second encoding result, it further performs the following steps: sending the second encoding result to the client. Thus, it can be realized that the target multimedia data is sent to the client by sending the second encoding result to the client.

[0137] As an optional implementation manner, in the process of the encoding device executing step 203, it performs the following steps:

[0138] 301. Determine the first MSE between the above-mentioned target multimedia data and the above-mentioned first decoding result.

[0139] 302. Update the parameters of the above-mentioned first model based on the above-mentioned first MSE to obtain the fourth model.

[0140] 303. Train the fourth model for n rounds based on the above-mentioned target multimedia data and the above-mentioned second model to obtain n second MSEs.

[0141] Specifically, the fourth model is used to encode the target multimedia data to obtain the ninth encoding result, and the second model is used to decode the ninth encoding result to obtain the third decoding result. The parameters of the fourth model are updated based on the MSE between the target multimedia data and the third decoding result to obtain the sixth model. Thus, one round of training of the fourth model is completed based on the target multimedia data and the second model to obtain one second MSE, where the obtained second MSE is the MSE between the target multimedia data and the third decoding result.

[0142] 304. Determine the minimum value between the above-mentioned first MSE and the above-mentioned n second MSEs to obtain the target MSE.

[0143] 305. Use the encoder that obtains the above-mentioned target MSE as the above-mentioned third model.

[0144] For example, when n is 1, the parameters of the fourth model are updated based on the second MSE to obtain the sixth model. If the first MSE is smaller than the second MSE, then the target MSE is the first MSE, the encoder that obtains the first MSE is the first model, and at this time the third model is the first model. If the second MSE is smaller than the first MSE, then the target MSE is the second MSE, the encoder that obtains the second MSE is the fourth model, and at this time the third model is the fourth model.

[0145] In this embodiment, the encoding device first determines the first MSE between the target multimedia data and the first decoding result, and then updates the parameters of the first model based on the first MSE to obtain the fourth model. Then, based on the target multimedia data and the second model, the fourth model is trained for n rounds to obtain n second MSEs, thereby enabling the fourth model to be fully trained using the target multimedia data. Then, the minimum value between the first MSE and the n second MSEs is determined to obtain the target MSE, and the encoder that obtains the MSE is used as the third model, which can be used to select an encoder with the best encoding effect on the target multimedia data from the encoders obtained through n rounds of training as the third model.

[0146] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for encoding an image provided by an embodiment of the present application. As Figure 4 shown, after the process starts, the first model is optimized jointly with the second model on the dataset. Specifically, the first model is trained using at least two reference images in the dataset and the second model. The specific implementation process can refer to the description of training the first model using the target multimedia data and the second model in steps 201 to 203, where each of the at least two reference images is the target multimedia data. Specifically, as Figure 4 shown, the second model is fixed, and the first model is trained using the first image. The implementation of this process can refer to the implementation of training the first model using the target multimedia data to obtain the third model in steps 201 to 203, where fixing the second model means not updating the parameters of the second model, and the first image is the target multimedia data. Then, it is determined whether the number of training rounds of the first model has reached n + 1. If the determination result is no, the second model is continued to be fixed, and the first model is trained using the first image. If the determination result is yes, the optimal encoder is found as the third model, and the first image is encoded using the third model. The implementation process of finding the optimal encoder as the third model can refer to steps 304 and 305. After encoding the first image using the third model, the process of encoding the target image ends.

[0147] In the case where the third model is obtained by training the first model through the encoding method described above, encoding the image using the third model to obtain an encoding result can save the bit rate of the encoding result, and in the case where the second model is used to decode the encoding result, the quality of the decoded image can be improved. Table 1 below shows the bit rate of the encoding result obtained by encoding the test image using the first model, and the peak signal-to-noise ratio (PSNR) of the image obtained by decoding the encoding result using the second model, where the unit of the bit rate is the encoding length required for encoding per pixel (bit per pixel, BPP). Table 1 below also shows the bit rate of the encoding result obtained by encoding the test image using the third model (i.e., the optimized BPP), and the PSNR of the image obtained by decoding the encoding result using the second model (i.e., the optimized PSNR), where the unit of the bit rate is BPP.

[0148]

[0149] Table 1

[0150] As shown in Table 1 above, there are 10 test images as follows: Image 1, Image 2, Image 3, Image 4, Image 5, Image 6, Image 7, Image 8, Image 9, Image 10. After encoding these 10 images respectively using the first model, 20 encoding results can be obtained. Among them, the bit rates of the encoding results obtained by encoding these 20 images are 0.590 BPP, 0.883 BPP, 0.186 BPP, 1.048 BPP, 0.937 BPP, 1.429 BPP, 0.401 BPP, 0.508 BPP, 0.571 BPP, 0.269 BPP respectively, and the average value of these 10 bit rates is 0.6822 BPP. Decoding the 10 encoding results obtained by the first model using the second model can obtain 10 images. The PSNRs of these 10 images are: 38.053, 35.543, 40.648, 34.855, 36.014, 35.437, 39.985, 38.787, 38.216, 39.417 respectively, and the average value of the PSNRs of these 10 images is 37.6955.

[0151] Similarly, after encoding these 10 images using the third model, 10 encoding results can also be obtained. The optimized BPP column represents the bitrates of these 10 encoding results. Specifically, the bitrates of these 10 encoding results are: 0.580, 0.882, 0.185, 1.047, 0.936, 1.408, 0.396, 0.500, 0.570, 0.269, and the average value of these 10 bitrates is 0.6773. Decoding the 10 encoding results obtained by the third model using the second model can obtain 10 images. The PSNR values of these 10 images are: 40.744, 38.042, 42.523, 37.179, 38.037, 38.271, 41.911, 41.122, 40.153, 41.356, and the average value of the PSNR of these 10 images is 39.9338.

[0152] By comparison, it can be seen that the average value of the bitrates of the encoding results obtained by the first model is larger than the average value of the bitrates of the encoding results obtained by the third model. This indicates that encoding the images using the third model can reduce the bitrate of the encoding results compared to using the first model. The average value of the PSNR of the images obtained by decoding the encoding results obtained by the third model is larger than the average value of the PSNR of the images obtained by decoding the encoding results obtained by the first model. This indicates that encoding the images using the third model can improve the quality of the images obtained by decoding the encoding results compared to using the first model.

[0153] As an alternative implementation, before the encoding device executes step 303, the following steps are also executed:

[0154] 401. Obtain the network speed of the above-mentioned client.

[0155] In this implementation, the encoding device communicates with the client via the Internet. Then, the network speed of the client will affect the time for the client to receive the encoding results sent by the encoding device. Specifically, the faster the network speed of the client, the shorter the time consumed for the client to receive the encoding results, and the slower the network speed of the client, the longer the time consumed for the client to receive the encoding results.

[0156] In an implementation of obtaining the network speed of the client, the encoding device receives the network speed sent by the client.

[0157] 402. Determine the above-mentioned n based on the above-mentioned network speed.

[0158] The encoding device needs to encode the target multimedia data only when it receives a request to send the target multimedia data to the client. That is to say, the encoding device executes step 303 when it receives a request to send the target multimedia data to the client. Therefore, the execution duration of step 303 will affect the response duration of the request to send an image to the client. Specifically, the longer the execution duration of step 303, the longer this response duration. Since the larger n is, the more rounds of training for the first model, correspondingly, the longer the time consumed for training the first model, that is, the longer the execution duration of step 303. Therefore, the larger n is, the longer this response duration.

[0159] Also, because the slower the network speed of the client, the longer this response duration, so if the response duration needs to be controlled within a reasonable range, the network speed of the client should be considered comprehensively. Therefore, the encoding device can determine the value of n according to the network speed of the client. Specifically, when the network speed and n are positively correlated, the encoding device determines n according to the network speed of the client.

[0160] In this implementation manner, before the encoding device executes step 303, it first obtains the network speed of the client, and then when the network speed of the client and n are positively correlated, it determines n according to the network speed of the client. Then, training the first model according to n to obtain the third model can fully train the first model while the response duration of the request to send the target multimedia data to the client is controllable, thereby improving the experience of the client in displaying the target multimedia data.

[0161] As an alternative implementation manner, after obtaining the second encoding result, the encoding device sends the second encoding result to the client, thereby realizing sending the target multimedia data to the client.

[0162] As an alternative implementation manner, before the encoding device executes step 202, it also executes the following steps:

[0163] 501. Obtain the decoding performance index of the above client.

[0164] In the embodiments of the present application, the decoding performance index characterizes the performance of the decoding resources in the client, where the decoding resources are the performance of the resources for decoding. For example, if the decoding resources in the client are a graphics processing unit (GPU), then the decoding performance index characterizes the performance of the GPU of the client.

[0165] In an implementation manner of obtaining the decoding performance index of the client, the encoding device obtains the decoding performance index sent by the client.

[0166] Optionally, the decoding performance metric characterizes the performance of the decoding resources in the client at the target time, where the target time is the time when the client sends a request to the encoding device to download the target multimedia data. For example, the client sends a request to the encoding device to download the target multimedia data at time t1, and the request includes a decoding performance metric, where the decoding performance metric characterizes the performance of the decoding resources in the client at time t1.

[0167] Since the performance of the decoding resources in the client changes in real time, and the time when the client receives the encoding result sent by the encoding device is close to the target time, after the client receives the encoding result sent by the encoding device, the performance of the decoding resources in the client is close to the performance of the decoding resources in the client at the target time. Therefore, in the case where the decoding performance metric characterizes the performance of the decoding resources in the client at the target time, the encoding device can more accurately determine the performance of the target resources in the client, where the target resources are the resources for decoding the encoding result sent by the encoding device. It should be understood that the target resources and the decoding resources may be the same or different. Specifically, the decoding resources are all the resources available for decoding in the client, while the target resources are the resources available for decoding the encoding result in the client. For example, the decoding resources are the GPU. If the utilization rate of the GPU is 30% when the client needs to decode the encoding result, then the resources available for decoding the encoding result are 70% of the GPU. Obviously, in the case where the decoding performance metric characterizes the performance of the decoding resources in the client at the target time, based on the decoding performance metric, the performance of the target resources in the client can be more accurately determined.

[0168] 502. Based on the above decoding performance metric, change the parameters of the above second model and / or change the structure of the above second model to obtain a fifth model.

[0169] Since the client uses a decoder for decoding, it consumes decoding resources, and the performance of the decoding resources will directly affect the use effect of the decoder. Specifically, the more complex the structure of the decoder and the more parameters in the decoder, the higher the requirements for the performance of the decoding resources for running the decoder.

[0170] Therefore, when the performance of the decoding resources in the client is sufficient to support running the decoder, the decoding resources can perform decoding using all the structures and all the parameters in the decoder. When the performance of the decoding resources in the client is insufficient to support running the decoder, the decoding resources can perform decoding using some of the structures in the decoder. When the performance of the decoding resources in the client is insufficient to support running the decoder, the decoding resources may perform decoding using some of the parameters in the decoder. When the performance of the decoding resources in the client is insufficient to support running the decoder, the decoding resources may perform decoding using some of the structures and some of the parameters in the decoder. That is to say, when the performance of the decoding resources in the client is insufficient to support running the decoder, the decoder actually used for decoding in the client is adjusted based on the deployed decoder. In other words, the decoder actually used for decoding in the client is adjusted based on the second model.

[0171] Thus, in order to better simulate the decoding process of the client, when the encoding device trains the first model using the second model, it should adjust the second model based on the performance of the decoding resources in the client. Therefore, the encoding device changes the parameters of the second model and / or changes the structure of the second model based on the decoding performance metric to obtain the fifth model. Specifically, the encoding device changes the parameters of the second model based on the decoding performance metric to obtain the fifth model. Among them, the weaker the performance indicated by the decoding performance metric, the more the parameters of the second model are reduced by changing the parameters of the second model. The encoding device changes the structure of the second model based on the decoding performance metric to obtain the fifth model. Among them, the weaker the performance indicated by the decoding performance metric, the more the structure of the second model is reduced by changing the structure of the second model.

[0172] In the embodiments of the present application, the requirements for the performance of the resources used for decoding when running the fifth model match the decoding performance metric. That is to say, the resources used for decoding in the client can support running the fifth model.

[0173] After obtaining the fifth model, the encoding device performs the following steps during the execution of step 202:

[0174] 503. Decode the first encoding result using the above fifth model to obtain the first decoding result.

[0175] In this embodiment, the encoding device obtains the decoding performance metrics of the client, where the decoding performance metrics characterize the performance of the decoding resources in the client. Then, based on the decoding performance metrics, the parameters and / or the structure of the second model are changed to obtain a fifth model, where the requirements for the performance of the decoding resources when running the fifth model match the decoding performance metrics. In this way, a fifth model that matches the actual decoder running during the client's decoding of the encoding result sent by the encoding device can be obtained. Then, the first encoding result is decoded using the fifth model to obtain a first decoding result, which can better simulate the effect of the client decoding the first encoding result. Thus, when the third model is obtained by training the first model based on the fifth model and the target multimedia data, the target multimedia data is encoded using the third model to obtain a second encoding result, and the second encoding result is sent to the client, which can improve the quality of the multimedia data obtained by the client decoding the second encoding result.

[0176] As an alternative embodiment, the target multimedia data includes a first image, the first model is used to encode the image, and the second model is used to decode the encoding result of the image.

[0177] The encoding device obtains the first encoding result by performing the following steps: encoding the first foreground image using the above-mentioned first model to obtain a third encoding result.

[0178] In the embodiment of the present application, the first foreground image is the pixel area covered by the foreground in the first image. For example, the foreground is a person or the foreground is a vehicle.

[0179] In this embodiment, during the execution of step 202, the encoding device performs the following steps: decoding the above-mentioned third encoding result using the above-mentioned second model to obtain a first decoding result.

[0180] In this embodiment, during the execution of step 203, the encoding device performs the following steps: updating the parameters of the above-mentioned first model based on the difference between the above-mentioned first foreground image and the above-mentioned first decoding result to obtain a first foreground encoder.

[0181] In this embodiment, the first foreground encoder obtained by the encoding device has a better encoding effect on the foreground in the first image, and thus can improve the quality of the foreground in the image decoded by the client.

[0182] In this embodiment, during the execution of step 204, the encoding device performs the following steps: encoding the background image using the above-mentioned first model to obtain a fourth encoding result, where the background image is the pixel region in the first image except for the first foreground image; encoding the first foreground image using the above-mentioned first foreground encoder to obtain a fifth encoding result; and obtaining the second encoding result based on the fourth encoding result and the fifth encoding result. At this time, the second encoding result is the encoding result of the first image. In a possible implementation manner, the encoding device obtains the second encoding result by combining the fourth encoding result and the fifth encoding result.

[0183] In this embodiment, the encoding device uses the foreground in the first image to train the first model to obtain a first foreground encoder. Then, it encodes the background image using the first model to obtain a fourth encoding result, and encodes the foreground in the first image using the first foreground encoder to obtain a fifth encoding result, which can reduce the bit rate of the encoding result of the foreground in the first image and improve the quality of the encoding result of the foreground in the first image. Finally, the second encoding result is obtained based on the fourth encoding result and the fifth encoding result. Since the first foreground image has a smaller data volume than the first image, training the first model using the first foreground image consumes less time than training the first model using the first image. Moreover, through this embodiment, the encoding effect of the foreground in the target multimedia data can be improved, and correspondingly, the quality of the foreground in the image decoded by the client can be improved.

[0184] As an optional embodiment, the first image is a frame in a live video stream, where the live video stream is a video stream obtained by video-livecasting a target scene. For example, the target scene is a room for livecasting, and the live video stream is a video stream generated by livecasting the room for livecasting. That is to say, the scene in the live video stream is the target scene, that is, the scene in the live video stream is fixed.

[0185] It should be understood that there may be objects outside the target scene in the target scene. For example, there is at least one person in the target scene, or there are goods in the target scene. Moreover, the objects in the target scene can change. For example, from time t1 to time t2, the anchor in the target scene is A, and from time t2 to time t3, the anchor in the target scene is B. That is to say, the people in the target scene from time t1 to time t2 are different from the people in the target scene from time t2 to time t3. Another example is that from time t1 to time t2, the goods in the target scene are C, and from time t2 to time t3, the goods in the target scene are D. That is to say, the goods in the target scene from time t1 to time t2 are different from the goods in the target scene from time t2 to time t3.

[0186] In this embodiment, the foreground is the pixel region in the first image except for the target scene. For example, if there is a person in the target scene, then the live video stream includes the person, and correspondingly, the first image includes the person. At this time, the foreground is the person in the first image. Another example is that if there is a commodity in the target scene, then the live video stream includes the commodity, and correspondingly, the first image includes the commodity. At this time, the foreground is the commodity in the target multimedia data.

[0187] In this embodiment, the encoding device obtains the fourth encoding result by performing the following steps: encoding the background image using the background encoder to obtain the fourth encoding result.

[0188] In the embodiments of the present application, the background encoder is obtained by training the first model using the scene image and the second model, where the scene image is an image captured of the target scene. Specifically, encoding the scene image using the first model to obtain the tenth encoding result, and decoding the tenth encoding result using the second model to obtain the fourth decoded image. Determine the third MSE between the fourth image and the scene image. Based on the third MSE, determine the loss of the first model, and according to the loss of the first model, update the parameters of the first model until the loss of the first model converges to obtain the background encoder.

[0189] Since the background encoder is trained using the scene image, and the content in the background image has a high degree of matching with the content in the scene image, therefore, encoding the background image using the background encoder to obtain the fourth encoding result can improve the quality of the fourth encoding result and reduce the bit rate of the fourth encoding result.

[0190] In this embodiment, since the scene in the live video stream is fixed, after training the first model using the scene image to obtain the background encoder, the background encoder can be used to encode the background image in any frame of the live video stream. This eliminates the need to train the first model separately for the background image of each frame, thereby improving the encoding efficiency of the live video stream.

[0191] As an optional embodiment, the live video stream further includes a second image, where the timestamp of the second image is greater than the timestamp of the first image, and the difference between the timestamp of the first image and the timestamp of the second image is less than or equal to a threshold. In other words, during the playback of the live video stream, the second image is played after the first image is played, and the time difference between playing the first image and playing the second image is less than the threshold. In the embodiments of the present application, the time difference between playing the first image and playing the second image is less than the threshold, indicating that the time difference between playing the first image and playing the second image is small. Optionally, the threshold is 0.5 seconds.

[0192] In this embodiment, the encoding device further performs the following steps:

[0193] 601. Encode the second foreground image using the above first model to obtain a sixth encoding result.

[0194] In the embodiments of the present application, the second foreground image is the pixel region covered by the foreground in the second image.

[0195] 602. Decode the above sixth encoding result using the above second model to obtain a third decoding result.

[0196] 603. Update the parameters of the above first model based on the difference between the above second foreground image and the above third decoding result to obtain a second foreground encoder.

[0197] 604. Encode the above second foreground image using the above second foreground encoder to obtain a seventh encoding result.

[0198] 605. Obtain an eighth encoding result of the above second image based on the above fourth encoding result and the above seventh encoding result.

[0199] Since the scene in the live video stream is fixed and the background in the live video stream is fixed. Also, because the time difference between playing the first image and playing the second image is small, the change in the background in the second image compared to the background in the second image is small. Therefore, the similarity between the background in the first image and the background in the second image is high. Thus, in the case where the encoding result of the background image in the first image has been obtained, the encoding result of this background image can be directly used without encoding the background in the second image again. Therefore, in this embodiment, when the encoding device encodes the second image after obtaining the fourth encoding result by encoding the background image in the first image, it uses the second foreground image in the second image to train the first model to obtain a second foreground encoder, then uses the second foreground encoder to encode the second foreground image to obtain a seventh encoding result, and finally obtains the eighth encoding result of the second image based on the obtained fourth encoding result and seventh encoding result, thereby improving the encoding efficiency of the live video stream.

[0200] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0201] The method of the embodiments of the present application is elaborated in detail above, and the device of the embodiments of the present application is provided below.

[0202] Please refer to Figure 5 , Figure 5Schematic structural diagram of an encoding device provided by an embodiment of the present application. The encoding device 2 includes: an acquisition unit 21, a decoding unit 22, an update unit 23, and an encoding unit 24. Optionally, the encoding device 2 further includes: a sending unit 25, a determination unit 26, and a change unit 27. Specifically:

[0203] The acquisition unit 21 is configured to acquire target multimedia data and a first encoding result, where the first encoding result is obtained by encoding the target multimedia data using a first model. The first model has trainable parameters and is used for encoding multimedia data.

[0204] The decoding unit 22 is configured to decode the first encoding result using a second model to obtain a first decoding result. The second model is the same as the decoding model deployed on the client, and the decoding model is used for decoding the encoding result of multimedia data.

[0205] The update unit 23 is configured to update the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain a third model.

[0206] The encoding unit 24 is configured to encode the target multimedia data using the third model to obtain a second encoding result.

[0207] Combined with any implementation manner of the present application, the update unit 23 is configured to:

[0208] Determine a first mean square error between the target multimedia data and the first decoding result;

[0209] Update the parameters of the first model based on the first mean square error to obtain a fourth model;

[0210] Train the fourth model for n rounds based on the target multimedia data and the second model to obtain n second mean square errors;

[0211] Determine the minimum value between the first mean square error and the n second mean square errors to obtain a target mean square error;

[0212] Use the model that obtains the target mean square error as the third model.

[0213] Combined with any implementation manner of the present application, the acquisition unit 21 is further configured to acquire the network speed of the client;

[0214] The encoding device 2 further includes: a determination unit 26, configured to determine the n based on the network speed, and the n is positively correlated with the network speed.

[0215] In combination with any embodiment of the present application, the encoding device 2 further includes: a sending unit 25, configured to send the second encoding result to the client.

[0216] In combination with any embodiment of the present application, the obtaining unit 21 is further configured to obtain a decoding performance metric of the client, where the decoding performance metric characterizes the performance of decoding resources in the client, and the decoding resources are the performance of resources for decoding;

[0217] The encoding device 2 further includes: a changing unit 27, further configured to change the parameters of the second model and / or change the structure of the second model based on the decoding performance metric to obtain a fifth model, and the requirements of running the fifth model for the performance of the decoding resources match the decoding performance metric;

[0218] The decoding unit 22 is specifically configured to:

[0219] Decode the first encoding result by using the fifth model to obtain the first decoding result.

[0220] In combination with any embodiment of the present application, the target multimedia data includes a first image, the first model is used to encode the image, and the second model is used to decode the encoding result of the image;

[0221] The obtaining unit 21 is specifically configured to:

[0222] Encode a first foreground image by using the first model to obtain a third encoding result, where the first foreground image is a pixel region covered by the foreground in the first image;

[0223] The decoding unit 22 is specifically configured to:

[0224] Decode the third encoding result by using the second model to obtain the first decoding result;

[0225] The updating unit 23 is specifically configured to:

[0226] Update the parameters of the first model based on the difference between the first foreground image and the first decoding result to obtain a first foreground encoder;

[0227] The encoding unit 24 is specifically configured to:

[0228] Encode a background image by using the first model to obtain a fourth encoding result, where the background image is a pixel region in the first image other than the first foreground image;

[0229] Encode the first foreground image by using the first foreground encoder to obtain a fifth encoding result;

[0230] Based on the fourth encoding result and the fifth encoding result, the second encoding result is obtained.

[0231] Combined with any embodiment of the present application, the first image is a frame image in a live video stream, the live video stream is a video stream obtained by video-livecasting a target scene, and the foreground is the pixel region in the first image other than the target scene;

[0232] The encoding unit 24 is specifically configured to:

[0233] Encode the background image by using a background encoder to obtain the fourth encoding result, where the background encoder is obtained by training the first model by using a scene image and the second model, and the scene image is an image obtained by photographing the target scene.

[0234] Combined with any embodiment of the present application, the live video stream further includes a second image, the timestamp of the second image is greater than the timestamp of the first image, and the difference between the timestamp of the first image and the timestamp of the second image is less than or equal to a threshold;

[0235] The encoding unit 24 is further configured to encode a second foreground image by using the first model to obtain a sixth encoding result, where the second foreground image is the pixel region covered by the foreground in the second image;

[0236] The decoding unit 22 is further configured to decode the sixth encoding result by using the second model to obtain a third decoding result;

[0237] The updating unit 23 is further configured to update the parameters of the first model based on the difference between the second foreground image and the third decoding result to obtain a second foreground encoder;

[0238] The encoding unit 24 is further configured to encode the second foreground image by using the second foreground encoder to obtain a seventh encoding result;

[0239] The encoding unit 24 is further configured to obtain an eighth encoding result of the second image based on the fourth encoding result and the seventh encoding result.

[0240] In the embodiments of the present application, before the encoding device sends target multimedia data to the client, it obtains the target multimedia data and the first encoding result, where the first encoding result is obtained by encoding the target multimedia data using the first model. The first encoding result is decoded using the second model to obtain the first decoding result, where the second model is the same as the decoding model deployed on the client. Then, based on the difference between the target multimedia data and the first decoding result, the parameters of the first model are updated to obtain the third model. Thus, it is possible to perform targeted training on the first model using the target multimedia data to obtain the third model. Then, the target multimedia data is encoded using the third model to obtain the second encoding result. In this way, not only can the bit rate of the second encoding result be reduced, but also when the decoder of the client is used to decode the second encoding result, the quality of the decoded multimedia data can be improved.

[0241] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.

[0242] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in the embodiments of the present application. The electronic device 3 includes a processor 31 and a memory 32. Optionally, the electronic device 3 further includes an input device 33 and an output device 34. The processor 31, the memory 32, the input device 33, and the output device 34 are coupled through a connector, which includes various interfaces, transmission lines, or buses, etc. The embodiments of the present application do not limit this. It should be understood that in each embodiment of the present application, coupling means being interconnected in a specific manner, including being directly connected or indirectly connected through other devices. For example, they can be connected through various interfaces, transmission lines, buses, etc.

[0243] The processor 31 may include one or more processors, for example, including one or more central processing units (CPUs). In the case where the processor is a single CPU, the CPU can be a single-core CPU or a multi-core CPU. Optionally, the processor 31 can be a processor group composed of multiple CPUs, and the multiple processors are coupled to each other through one or more buses. Optionally, the processor can also be other types of processors, etc. The embodiments of the present application do not limit this.

[0244] The memory 32 can be used to store computer program instructions and various computer program codes including the program codes for executing the solution of this application. Optionally, the memory includes but is not limited to random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), and this memory is used for relevant instructions and data.

[0245] The input device 33 is used to input data and / or signals, and the output device 34 is used to output data and / or signals. The input device 33 and the output device 34 can be independent devices or an integrated device.

[0246] It can be understood that in the embodiments of this application, the memory 32 can not only be used to store relevant instructions, but also be used to store relevant data. For example, the memory 32 can be used to store the to-be-distributed image and the first encoding result obtained through the input device 33, or the memory 32 can also be used to store the second encoding result obtained through the processor 31, etc. The embodiments of this application do not limit the specific data stored in this memory.

[0247] It can be understood that Figure 6 Only a simplified design of an electronic device is shown. In practical applications, the electronic device may also separately include necessary other components, including but not limited to any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of this application are within the protection scope of this application.

[0248] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraint conditions of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.

[0249] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. Those skilled in the art can also clearly understand that each embodiment of the present application has different focuses in description. For the convenience and brevity of description, the same or similar parts may not be elaborated in different embodiments. Therefore, parts not described or not described in detail in a certain embodiment can be referred to the descriptions of other embodiments.

[0250] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0251] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0252] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0253] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital versatile disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0254] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The foregoing storage medium includes: various media that can store program codes such as read-only memory (ROM) or random access memory (RAM), magnetic disks, or optical discs.

Claims

1. A coding method, characterized in that: The method is used to encode the target multimedia data before sending the target multimedia data to the client, and the method includes: Acquire target multimedia data and a first encoding result, where the first encoding result is obtained by encoding the target multimedia data using a first model, where the first model has trainable parameters, and the first model is used to encode the multimedia data; The first encoding result is decoded using a fifth model to obtain a first decoding result; wherein the fifth model is obtained by changing the parameters of the second model and / or changing the structure of the second model according to the decoding performance indicator obtained from the client; the decoding performance indicator represents the performance of the decoding resources in the client, the second model is the same as the decoding model deployed on the client, and the decoding model is used to decode the encoding result of the multimedia data; Based on the difference between the target multimedia data and the first decoding result, updating the parameters of the first model to obtain a third model; The target multimedia data is encoded using the third model to obtain a second encoding result.

2. The method according to claim 1, characterized in that The first model and the second model are trained using at least two reference multimedia data in a data set, wherein, when the at least two reference multimedia data are encoded using the first model and the encoding results of the at least two reference multimedia data are decoded using the second model to obtain at least two second decoding results, the difference between the at least two reference multimedia data and the at least two second decoding results is less than or equal to a preset value.

3. The method according to claim 1, characterized in that The updating of the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain the third model comprises: Determining a first mean square error between the target multimedia data and the first decoding result; Based on the first mean square error, updating the parameters of the first model to obtain a fourth model; Based on the target multimedia data and the second model, training the fourth model for n rounds to obtain n second mean square errors; Determine the minimum value of the first mean square error and the n second mean square errors to obtain a target mean square error; The model of the target mean square error is obtained as the third model.

4. The method according to claim 3, characterized in that Before training the fourth model for n rounds based on the target multimedia data and the second model to obtain n second mean square errors, the method further includes: Obtaining the network speed of the client; Based on the network speed, n is determined, and n is positively correlated with the network speed.

5. The method according to any one of claims 1 to 4, characterized in that: After obtaining the second encoding result, the method further includes: Send the second encoding result to the client.

6. The method according to any one of claims 1 to 4, characterized in that The target multimedia data includes a first image, the first model is used to encode the image, and the second model is used to decode the encoding result of the image; Obtain the first encoding result, including: Encoding a first foreground image using the first model to obtain a third encoding result, wherein the first foreground image is a pixel area covered by the foreground in the first image; The step of decoding the first encoding result by using the second model to obtain a first decoding result includes: Decoding the third encoding result using the second model to obtain the first decoding result; The updating of the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain the third model comprises: Based on the difference between the first foreground image and the first decoding result, updating the parameters of the first model to obtain a first foreground encoder; The step of encoding the target multimedia data by using the third model to obtain a second encoding result includes: Encoding a background image using the first model to obtain a fourth encoding result, wherein the background image is a pixel area in the first image excluding the first foreground image; encoding the first foreground image using the first foreground encoder to obtain a fifth encoding result; The second encoding result is obtained based on the fourth encoding result and the fifth encoding result.

7. The method according to claim 6, characterized in that The first image is a frame of image in a live video stream, the live video stream is a video stream obtained by live broadcasting of a target scene, and the foreground is a pixel area in the first image excluding the target scene; The step of encoding the background image by using the first model to obtain a fourth encoding result includes: The background image is encoded using a background encoder to obtain the fourth encoding result, wherein the background encoder is obtained by training the first model using a scene image and the second model, and the scene image is an image obtained by photographing the target scene.

8. The method according to claim 7, characterized in that The live video stream further includes a second image, a timestamp of the second image is greater than a timestamp of the first image, and a difference between the timestamp of the first image and the timestamp of the second image is less than or equal to a threshold; The method further comprises: encoding a second foreground image using the first model to obtain a sixth encoding result, wherein the second foreground image is a pixel area covered by the foreground in the second image; Decoding the sixth encoding result using the second model to obtain a third decoding result; Based on the difference between the second foreground image and the third decoding result, updating the parameters of the first model to obtain a second foreground encoder; encoding the second foreground image using the second foreground encoder to obtain a seventh encoding result; Based on the fourth encoding result and the seventh encoding result, an eighth encoding result of the second image is obtained.

9. An encoding device, characterized in that: The encoding device is used to encode the target multimedia data before sending the target multimedia data to the client, and the encoding device includes: an acquisition unit, configured to acquire target multimedia data and a first encoding result, wherein the first encoding result is obtained by encoding the target multimedia data using a first model, wherein the first model has trainable parameters, and the first model is used to encode the multimedia data; A decoding unit, configured to decode the first encoding result using a fifth model to obtain a first decoding result; wherein the fifth model is obtained by changing the parameters of the second model and / or changing the structure of the second model by using a decoding performance indicator obtained from the client; the decoding performance indicator represents the performance of a decoding resource in the client, the second model is the same as a decoding model deployed on the client, and the decoding model is used to decode the encoding result of multimedia data; an updating unit, configured to update the parameters of the first model based on the difference between the target multimedia data and the first decoding result to obtain a third model; The encoding unit is used to encode the target multimedia data using the third model to obtain a second encoding result.

10. An electronic device, characterized in that: include: A processor and a memory, wherein the memory is used to store computer program codes, wherein the computer program codes include computer instructions, and when the processor executes the computer instructions, the electronic device executes the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is enabled to execute the method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The computer program product comprises a computer program or instructions; when the computer program or instructions are run on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Time delay data adjustment method and device, electronic equipment and storage medium

    CN112994981A

  • Video coding method, model training method, equipment and storage medium

    CN116095328A

  • Hierarchical video coding method, device and product based on semantic information

    CN116723333A

  • Image generation and model training method and device, equipment and storage medium

    CN117351299A