Video coding method and device, video decoding method and device, electronic equipment and storage medium

By converting high-quality images into lower-quality images and encoding, the problem of improving video encoding and codec quality while the device configuration remains unchanged, and the encoding and reconstruction of high-quality images are realized.

CN120238652APending Publication Date: 2025-07-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311837893.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, it is difficult to improve the video encoding and decoding quality without changing the device configuration, especially when encoding and decoding images in high-quality image formats, the device coverage is low and cannot meet the high requirements of users for video quality.

Method used

The image of the high-quality first image format is converted into a relatively low-quality second image format, and the second image is obtained based on the image conversion loss. By encoding the first image and the second image respectively, encoding the high-quality image is realized.

Benefits of technology

With the unchanged device configuration, the video encoding quality is improved, the dependence on the device is reduced, and the encoding and reconstruction of high-quality images is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238652A_ABST
    Figure CN120238652A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video coding method and device, a video decoding method and device, electronic equipment and a storage medium, and relates to the fields of video coding and decoding, video live broadcast, games, cloud technologies and the like. The method comprises the steps that a video to be coded is acquired, for each frame of image to be coded in the video to be coded, the frame of image to be coded is converted into a first image in a second image format, a second image is obtained based on image conversion loss between the frame of image to be coded and the first image, the image format of the image to be coded is the first image format, and the image format of the image to be coded is the second image format. The image quality of the to-be-coded image is higher than that of the first image, and the code stream of each frame of image is obtained by respectively coding the first image and the second image corresponding to each frame of to-be-coded image. Based on the method, dependence on equipment is reduced, high-quality images can be coded while equipment configuration is not changed, video coding quality is improved, and video coding requirements are better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology and may be involved in fields such as games, video encoding and decoding, video live streaming, and cloud technology. Specifically, this application relates to a video encoding method, a decoding method, a device, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of technology, people have higher and higher requirements for video quality, and also put forward higher requirements for video resolution and color display effects.

[0003] Currently, in related application scenarios, video encoding and decoding are usually based on low-quality images. For example, images in the YUV420 format (a common color encoding format for media streams) are used for encoding and decoding. However, the video display effects reconstructed by encoding and decoding based on low-quality images are poor and cannot meet the needs of users. Therefore, it is necessary to improve the quality of image encoding and decoding to achieve more accurate color display.

[0004] However, encoding and decoding images in a higher-quality image format often require higher requirements for devices, resulting in a lower coverage rate of devices. Therefore, how to improve the quality of video encoding and decoding without changing the device configuration has become an urgent problem to be solved. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a video encoding method, a decoding method, a device, an electronic device, and a storage medium that can effectively improve the quality of video encoding and decoding.

[0006] On the one hand, the embodiments of this application provide a video encoding method, which includes:

[0007] Obtain a video to be encoded, where the video to be encoded includes at least one frame of image to be encoded, and the image format of the image to be encoded is a first image format;

[0008] For each frame of image to be encoded, convert the frame of image to be encoded into a first image in a second image format, and obtain a second image based on the image conversion loss between the frame of image to be encoded and the first image, where the image quality of the image to be encoded is higher than the image quality of the first image;

[0009] Encode the first image and the second image corresponding to each frame of image to be encoded respectively to obtain the bitstreams of each frame of image.

[0010] On the other hand, the embodiments of this application also provide a video encoding device, which includes:

[0011] A video acquisition module for acquiring a video to be encoded, where the video to be encoded includes at least one frame of image to be encoded, and the image format of the image to be encoded is a first image format;

[0012] A conversion module for converting each frame of the image to be encoded into a first image in a second image format and obtaining a second image based on the image conversion loss between the frame of the image to be encoded and the first image, where the image quality of the image to be encoded is higher than that of the first image;

[0013] An encoding module for encoding the first image and the second image corresponding to each frame of the image to be encoded respectively to obtain the bitstreams of each frame of image.

[0014] Optionally, the first image format and the second image format are two image formats in the YUV color mode;

[0015] The conversion module can be used for:

[0016] Determining the luminance component values of each pixel point in the frame of the image to be encoded as the luminance component values of each pixel point in the first image corresponding to the frame of the image to be encoded;

[0017] Sampling the first chrominance component value and the second chrominance component value of the frame of the image to be encoded according to the difference between the first image format and the second image format to obtain the first chrominance component value and / or the second chrominance component value of some pixel points in the first image corresponding to the frame of the image to be encoded;

[0018] Constructing a second image based on the first chrominance component value and the second chrominance component value of each pixel point in the frame of the image to be encoded that are not sampled.

[0019] Optionally, the image sizes of the first image and the second image are the same;

[0020] The conversion module can be used for:

[0021] Taking the first partial component values as the luminance component values of each pixel point in the second image to be constructed respectively; where the first partial component values are some of the first chrominance component values and the second chrominance component values that are not sampled;

[0022] Determining the first chrominance component value and / or the second chrominance component value of the pixel points in the second image corresponding to the some pixel points in the first image based on the second partial component values, where the second partial component values are the component values of the first chrominance component values and the second chrominance component values that are not sampled except the first partial component values.

[0023] Optionally, the first image format is YUV444 format, and the second image format is NV12 format;

[0024] The conversion module can be used for:

[0025] For each 2×2 image block in the frame of the image to be encoded, sample the first chrominance component value of the first pixel point in the image block to obtain the first chrominance component value of the corresponding pixel point in the first image, and sample the second chrominance component value of the second pixel point in the image block to obtain the second chrominance component value of the corresponding pixel point in the first image; the first pixel point and the second pixel point are any pixel points in the image block;

[0026] For each 2×2 image block in the frame of the image to be encoded, construct an image block corresponding to the image block in the second image based on the first chrominance component values and the second chrominance component values of the pixel points not sampled in the image block.

[0027] Optionally, the conversion module can be used for:

[0028] Take any 4 component values from the first chrominance component values and the second chrominance component values not sampled in the image block, and respectively use them as the luminance component values of the pixel points in the image block corresponding to the image block in the second image;

[0029] Take the 2 component values not sampled except the 4 component values and respectively use them as the first chrominance component value or the second chrominance component value of the pixel points corresponding to the first pixel point and the second pixel point in the second image.

[0030] Optionally, the conversion module is further used for:

[0031] Create a first view and a second view corresponding to the image to be processed;

[0032] Store the luminance component values of the pixel points in the image to be encoded corresponding to the image to be processed into the first view;

[0033] Store the first chrominance component value of the first pixel point and the second chrominance component value of the second pixel point in the frame of the image to be encoded into the second view;

[0034] The encoding module can be used for:

[0035] Encode the first view and the second view of the image to be processed to obtain a bitstream corresponding to the image to be processed.

[0036] On the other hand, an embodiment of the present application further provides a video decoding method, and the method includes:

[0037] Obtain the data to be decoded; wherein, the data to be decoded includes the bitstreams corresponding to at least one frame of images, the bitstream corresponding to one frame of images includes the first bitstream of the first image corresponding to this frame of images and the second bitstream of the second image corresponding to this frame of images, this frame of images is an image in the first image format, the first image is an image in the second image format, and the image quality of the image in the first image format is higher than that of the image in the second image format; the second image is obtained based on the image conversion loss between this frame of images and the first image;

[0038] Analyze the bitstreams corresponding to the at least one frame of images to obtain the first bitstream and the second bitstream corresponding to the at least one frame of images;

[0039] Decode the first bitstream and the second bitstream corresponding to each frame of images respectively, and obtain the first image and the second image corresponding to each frame of images based on the decoding results;

[0040] By fusing the first image and the second image corresponding to each frame of images, obtain the reconstructed image corresponding to each frame of images.

[0041] On the other hand, an embodiment of the present application further provides a video decoding device, and this device includes:

[0042] An obtaining module, configured to obtain the data to be decoded; wherein, the data to be decoded includes the bitstreams corresponding to at least one frame of images, the bitstream corresponding to one frame of images includes the first bitstream of the first image corresponding to this frame of images and the second bitstream of the second image corresponding to this frame of images, this frame of images is an image in the first image format, the first image is an image in the second image format, and the image quality of the image in the first image format is higher than that of the image in the second image format; the second image is obtained based on the image conversion loss between this frame of images and the first image;

[0043] An analysis module, configured to analyze the bitstreams corresponding to the at least one frame of images to obtain the first bitstream and the second bitstream corresponding to the at least one frame of images;

[0044] A decoding module, configured to decode the first bitstream and the second bitstream corresponding to each frame of images respectively, and obtain the first image and the second image corresponding to each frame of images based on the decoding results;

[0045] A fusion and reconstruction module, configured to obtain the reconstructed image corresponding to each frame of images by fusing the first image and the second image corresponding to each frame of images.

[0046] Optionally, for each frame of images, the first bitstream corresponding to this frame of images carries the first identifier of this frame of images, and the second bitstream corresponding to this frame of images carries the second identifier of this frame of images;

[0047] The decoding module can be used for:

[0048] Store the first image corresponding to each decoded frame image into the first queue, and store the second image corresponding to each decoded frame image into the second queue;

[0049] For each first image in the first queue, based on the first identifier of the first image and the second identifiers of the second images stored in the second queue, determine the second image corresponding to the first image from the second queue.

[0050] Optionally, the obtaining module can be used to:

[0051] Receive the bitstream of the video in real time;

[0052] The decoding module can be used to:

[0053] At a preset time interval, match and retrieve the first identifier of the first image with the second identifiers of the second images stored in the second queue. If a matching second image is retrieved before reaching the first threshold, use the retrieved second image as the second image corresponding to the first image;

[0054] The decoding module is further used to:

[0055] If no matching second image is retrieved when reaching the first threshold, determine the reconstructed image corresponding to this frame image based on the first image.

[0056] Optionally, the decoding module is further used to:

[0057] Create a third view and a fourth view corresponding to the image to be processed;

[0058] Store the luminance component value corresponding to the image to be processed into the third view;

[0059] Store the first chrominance component value and the second chrominance component value corresponding to the image to be processed into the fourth view;

[0060] The fusion and reconstruction module can be used to:

[0061] For each frame image, obtain the reconstructed image corresponding to this frame image by fusing the third view and the fourth view corresponding to this frame image.

[0062] Optionally, the decoding module can be used to:

[0063] Decode the first bitstream and the second bitstream corresponding to each frame image respectively by means of hardware decoding.

[0064] The embodiments of the present application further provide an electronic device, which includes a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement the method provided in any optional embodiment of the present application.

[0065] On the other hand, the embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method provided in any optional embodiment of the present application is implemented.

[0066] On the other hand, the embodiments of the present application further provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the method provided in any optional embodiment of the present application is implemented.

[0067] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are as follows:

[0068] In the video encoding method provided by the embodiments of the present application, during the video encoding process, an image in the first image format with high quality in the video is converted into a second image in the second image format with relatively low quality, and the second image is obtained based on the image conversion loss. By encoding the first image and the second image respectively, the encoding of high-quality images is realized without changing the device configuration, reducing the dependence on the device, improving the video encoding quality, and better meeting the video encoding requirements.

[0069] During the video decoding process, by decoding the first bitstream and the second bitstream in the data to be decoded respectively, the first image and the second image are obtained, and the first image and the second image are fused to obtain a high-quality reconstructed image. The reconstruction of high-quality images is realized without changing the device configuration, reducing the dependence on the device, and improving the video decoding quality. Description of the Drawings

[0070] Figure 1 It is a schematic flowchart of a video encoding method provided by an embodiment of the present application;

[0071] Figure 2 It is a schematic structural diagram of the YUV444 and YUV420 image formats provided by an embodiment of the present application;

[0072] Figure 3 It is a schematic structural diagram of the YUV444 and YUV422 image formats provided by an embodiment of the present application;

[0073] Figure 4 It is a schematic structural diagram of image splitting provided by an embodiment of the present application;

[0074] Figure 5 It is a schematic flowchart of splitting an image to be encoded provided by an embodiment of the present application;

[0075] Figure 6 It is a schematic flowchart of a video decoding method provided by an embodiment of the present application;

[0076] Figure 7 It is a schematic diagram of fusing a first image and a second image provided by an embodiment of the present application;

[0077] Figure 8 It is a schematic flowchart of video encoding and decoding provided by the present application;

[0078] Figure 9 It is a schematic structural diagram of a video conferencing system in a video conferencing scenario provided by the present application;

[0079] Figure 10 It is a schematic structural diagram of a video transmission system in a cloud gaming scenario provided by the present application;

[0080] Figure 11 It is a schematic structural diagram of a call system in a video call scenario provided by the present application;

[0081] Figure 12 It is a schematic structural diagram of a video encoding device provided by an embodiment of the present application;

[0082] Figure 13 It is a schematic structural diagram of a video decoding device provided by an embodiment of the present application;

[0083] Figure 14 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0084] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0085] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present invention. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items can refer to one, more or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it can be implemented that parameter A includes A1 or A2 or A3, and it can also be implemented that parameter A includes at least two of the three items of parameter A1, A2, A3.

[0086] The embodiments of the present application provide a video encoding method, a decoding method, a device, an electronic device and a storage medium. During the video encoding process, by converting an image in a first image format with high quality in the video into a second image in a second image format with relatively low quality, and obtaining the second image based on the image conversion loss, by encoding the first image and the second image respectively, the encoding of high-quality images is achieved without changing the device configuration, reducing the dependence on the device, improving the video encoding quality, and better meeting the video encoding requirements.

[0087] During the video decoding process, by decoding the first bitstream and the second bitstream in the data to be decoded respectively, obtaining the first image and the second image, and fusing the first image and the second image, a high-quality reconstructed image is obtained. The reconstruction of high-quality images is achieved without changing the device configuration, reducing the dependence on the device, and improving the video decoding quality.

[0088] Optionally, the solution provided by the embodiments of the present application may involve cloud technology. For example, the solution of the embodiments of the present application may be executed by a server or a user terminal. Among them, the server may be a cloud server. The data processing involved in the implementation process of this solution may be realized based on cloud technology, and the data storage involved in the implementation process may adopt cloud storage. For example, the data calculation involved in the video encoding and decoding process may be realized by using cloud computing technology, and the storage of the video may adopt cloud storage.

[0089] Among them, cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services according to needs. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed file systems, and works together through application software or application interfaces to jointly provide data storage and business access functions.

[0090] It should be noted that in the optional embodiments of the present application, for the relevant data of the object information (such as the uploaded video), when the embodiments in the present application are applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments in the present application involve data related to an object, it needs to be obtained under the condition of object authorization and consent, relevant department authorization and consent, and compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information needs to obtain the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained, and the embodiments also need to be implemented under the condition of object authorization and consent.

[0091] For the convenience of understanding, the following explanations are made for the nouns involved in the present application.

[0092] 1. RGBA: It is used for computer image encoding and represents the color space of Red, Green, Blue, and transparent Alpha. It is a variant of RGB color encoding.

[0093] 2. YUV: A color encoding method where Y represents luminance, and UV represents chrominance. UV is used to describe color and saturation. YUV is usually divided into YUV444, YUV422, YUV420, YUV411, etc. according to the sampling method.

[0094] Among them, YUV444 represents full sampling, that is, the sampling ratios of the Y, U, and V components are the same. Each pixel in the corresponding image contains a Y, U, and V component value.

[0095] YUV422 represents 2:1 horizontal sampling and full vertical sampling. That is, the UV components are half of the Y component sampling. The Y component and the UV components are sampled in a 2:1 ratio, and every two pixel points share a set of UV components.

[0096] YUV420 represents 2:1 horizontal sampling and 2:1 vertical sampling. Every four pixel points share a set of UV components.

[0097] For the storage method of YUV, there are generally two methods: the packed format, the planar format, and the semi-planar format. Among them, in the packed format, Y, U, and V are stored together; in the planar format, the Y, U, and V components are stored separately; in the semi-planar format, the Y component and the UV components are stored separately.

[0098] 3. NV12: It belongs to the YUV420SP (semi-planar) format and uses two planes to store the Y component and the UV components respectively. Among them, the UV components share one plane and are interleaved in the order of U, V, U, V.

[0099] NV21: It belongs to the YUV420SP (semi-planar) format and uses two planes to store the Y component and the UV components respectively. Among them, the UV components share one plane and are interleaved in the order of V, U, V, U.

[0100] Next, through the description of several exemplary embodiments, the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0101] Figure 1The figure shows a schematic flowchart of a video encoding method provided by an embodiment of the present application. This method can be executed by any electronic device, such as a user terminal (which can also be referred to as a terminal, terminal device, or user equipment, etc.) or a server, or can be implemented by multiple electronic devices in cooperation.

[0102] Among them, the above-mentioned server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server (which can be referred to as the cloud) that provides cloud computing services. The terminal (which can also be referred to as a user terminal or user equipment) can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle-mounted terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0103] As Figure 1 shown, the video encoding method provided by an embodiment of the present application can include the following S110 to S130.

[0104] S110: Obtain the video to be encoded.

[0105] The solution provided by an embodiment of the present application can be applied to any application scenario with encoding and decoding requirements, that is, the video to be encoded can be any video that needs to be encoded in any application scenario, such as short videos, movie videos, game videos, etc. The video can be a real-time video, such as a live video, an online meeting video, an online call video, etc., or a non-real-time video, such as a video that has been recorded, a video on a video website (such as a video in the video database of an application program that provides video viewing services), etc. Among them, the video to be encoded includes at least one frame of image to be encoded. That is to say, the method provided by an embodiment of the present application is also applicable to encoding a single image.

[0106] Among them, the image format of the image to be encoded is the first image format, and here the image format refers to the color encoding format of the image. Optionally, the first image format is an image format in the YUV color mode, such as YUV444, YUV422, etc.

[0107] To facilitate the processing of the video (such as transmission, storage, etc.), the images in the video to be encoded are usually set to be in the YUV format when performing video encoding. However, the videos collected by image acquisition devices are often in the RGBA format (a color space composed of red (Red), green (Green), blue (Blue), and transparency Alpha). Therefore, before performing video encoding, it is also necessary to convert the images to be encoded in the video to be encoded into images in the YUV format.

[0108] S120: For each frame of the image to be encoded, convert the image to be encoded of this frame into a first image in a second image format, and obtain a second image based on the image conversion loss between the image to be encoded of this frame and the first image.

[0109] As people's requirements for video quality increase, continuing to use images of lower quality for video encoding can no longer meet the needs of users. It is necessary to improve the image quality and achieve more accurate color display. The image quality is directly related to the image format. Encoding images in a higher-quality image format often requires higher requirements for devices, and the coverage rate of devices is low.

[0110] In the embodiments of the present application, in order to reduce the dependence on devices, for each frame of the image to be encoded in the first image format in the video to be encoded, the image to be encoded of this frame can be converted into a first image in a second image format, and a second image can be obtained based on the image conversion loss between the image to be encoded of this frame and the first image. Among them, the image quality of the image to be encoded is higher than the image quality of the first image, that is, the image quality corresponding to the first image format is higher than the image quality corresponding to the second image format. The image to be encoded can be obtained by combining the first image and the second image.

[0111] Optionally, the first image format and the second image format are two different image formats in the YUV color mode. When converting the image to be encoded into the first image, the luminance component value (Y component value) of each pixel point in the image to be encoded of this frame is determined as the luminance component value of each pixel point in the first image corresponding to the image to be encoded of this frame; according to the difference between the first image format and the second image format, sample the first chrominance component (U component) and the second chrominance component (V component) of the image to be encoded of this frame to obtain the first chrominance component value and / or the second chrominance component value of some pixel points in the first image corresponding to the image to be encoded of this frame.

[0112] Exemplarily, when the first image format is YUV444 format and the second image format is YUV420 format, as Figure 2 shown, on the left is the image to be encoded in YUV444 format, and each pixel point includes a Y component value, a U component value, and a V component value; on the right is the first image in YUV420 format, and each pixel point includes a Y component value, and every four pixel points share a U component value and a V component value.

[0113] When the first image format is YUV444 format and the second image format is YUV422 format, as Figure 3As shown in the figure, on the left is the image to be encoded in YUV444 format, where each pixel includes a Y component value, a U component value, and a V component value; on the right is the first image in YUV422 format, where each pixel includes a Y component value, and every two pixels share a U component value and a V component value.

[0114] The image conversion loss between the image to be encoded and the first image is the first chrominance component value (U component value) and the second chrominance component value (V component value) of each pixel that is not sampled during the image conversion process. Based on the first chrominance component value and the second chrominance component value that are not sampled in this frame of the image to be encoded, a second image is constructed.

[0115] Continuing with the example where the first image format is YUV444 format and the second image format is YUV420 format, assuming the size of the image to be encoded is W×H×3, after conversion to an image in YUV420 format, the size becomes W×H×3 / 2, resulting in a loss in the UV components on the screen. In the embodiments of the present application, this part of the loss components is combined into the second image for video encoding, and the reconstructed image obtained by combining the first image and the second image is closer to the original image.

[0116] Optionally, the image sizes of the first image and the second image are the same. When constructing the second image, the first part of the component values are respectively used as the luminance component values of each pixel in the second image to be constructed, where the first part of the component values are part of the U component values and V component values that are not sampled. Based on the second part of the component values, the U component value and / or V component value of the pixel in the second image corresponding to some pixels in the first image are determined, and the second part of the component values are the component values of the U component values and V component values that are not sampled except for the first part of the component values. That is, the structural distributions of the corresponding YUV components in the first image and the second image are the same.

[0117] It can be understood that the structural distributions of the corresponding YUV components in the first image and the second image are the same, and the second image can be regarded as having the same image format as the first image, that is, the second image is an image in the second image format.

[0118] Optionally, when the first image format is YUV444 format and the second image format is NV12 format, when converting the image to be encoded into the first image, for each 2*2 image block (also known as pixel block, macroblock) in the frame of the image to be encoded, the Y component value of each pixel point in the image block is used as the Y component value of the corresponding pixel point in the image block in the first image. The U component value of the first pixel point in the image block is sampled to obtain the U component value of the corresponding pixel point in the first image, and the V component value of the second pixel point in the image block is sampled to obtain the V component value of the corresponding pixel point in the first image. Among them, the first pixel point and the second pixel point are any pixel points in the image block, and the first pixel point and the second pixel point can be the same pixel point or different pixel points.

[0119] For each 2*2 image block in the frame of the image to be encoded, based on the U component values and V component values of the pixel points in the image block that have not been sampled, an image block corresponding to the image block in the second image is constructed. Optionally, any 4 component values among the first chrominance component value and the second chrominance component value that have not been sampled in the image block are used as the Y component values of the pixel points in the image block corresponding to the image block in the second image, and the 2 component values that have not been sampled except the 4 component values are used as the U component values or V component values of the pixel points corresponding to the first pixel point and the second pixel point in the second image.

[0120] Exemplarily, as Figure 4 shown, assuming that the first image format is YUV444 format and the second image format is NV12 format, (a) in the figure is a 2*2 image block in the image to be encoded, which includes 4 pixel points, and each pixel point includes a Y component value, a U component value and a V component value, which are (Y0, U0, V0), (Y1, U1, V1), (Y2, U2, V2), (Y3, U3, V3) respectively. (b) in the figure is the corresponding image block in the first image, and each pixel point includes a Y component value. The four pixel points correspond to a U0 component value and a V0 component value. Among them, the Y component value of each pixel point in (b) corresponds to the Y component value of the corresponding pixel point in (a). By sampling the U component values of the pixel points in (a), the U component value of pixel point 0 in (a) is used as the U component value of the corresponding pixel point 0 in (b); by sampling the V component values of the pixel points in (a), the V component value of pixel point 0 in (a) is used as the V component value of the corresponding pixel point 0 in (b).

[0121] For each image block, the image conversion loss generated by converting the image to be encoded into the first image includes U1, V1, U2, V2, U3, and V3. Arbitrarily select any four component values U1, V1, U3, and V3 as the first part of the component values. Use U1, V1, U3, and V3 as the Y component values of each pixel point in the corresponding image block of the second image, use U2 as the U component value of the corresponding pixel point of the image block, and use V2 as the V component value of the corresponding pixel point of the image block. It can be seen that the second image also belongs to an image in NV12 format.

[0122] Further, for each 2*2 image block in the image to be encoded, according to the sampling rule of the Y component, sample the Y component values of each pixel point in the image block corresponding to the second image from the remaining UV component values (image conversion loss), and according to the sampling rule of the UV component, sample the U component value or V component value of the pixel point corresponding to the first pixel point and the second pixel point in the second image from the component values other than the Y component value in the remaining UV component values.

[0123] Continue with Figure 4 as an example for illustration. The image conversion loss generated by converting the image to be encoded into the first image includes U1, V1, U2, V2, U3, and V3. According to the sampling rule of the Y component, U1, V1, U3, and V3 can be used as the Y component values of each pixel point in the corresponding image block of the second image, U2 can be used as the U component value of the corresponding pixel point of the image block, and V2 can be used as the V component value of the corresponding pixel point of the image block. Or, use U1, V1, U2, and V2 as the Y component values of each pixel point in the corresponding image block of the second image, use U3 as the U component value of the corresponding pixel point of the image block, and use V3 as the V component value of the corresponding pixel point of the image block.

[0124] In the embodiments of the present application, the first image and the second image corresponding to each frame of the image to be encoded in the video to be encoded are respectively used as the images to be processed. For each image to be processed, it can be stored in the following manner:

[0125] Create the first view and the second view corresponding to the image to be encoded. Among them, the first view and the second view can be created through CreateRenderTargetView, and the first view and the second view belong to render target views.

[0126] Store the Y component values of each pixel point in the image to be encoded corresponding to the image to be processed into the first view, and store the U component value of the first pixel point and the V component value of the second pixel point in the frame of the image to be encoded into the second view. Among them, the image to be processed adopts a planar (planer) storage mode, with Y stored as one plane and UV stored as another plane.

[0127] For the first image, store the Y component values of each pixel point in the first image into the first view, that is, draw the Y component values in the first view through a shader, and store the U component value of the first pixel point and the V component value of the second pixel point sampled from the first image into the second view. Similarly, draw the UV component values in the second view through the shader.

[0128] For the second image, store the corresponding Y component values in the second image into the first view, and store the corresponding U component value and V component value in the second image into the second view. For example, Figure 4 the corresponding Y component values in the second image in [example] are U1, V1, U3, V3, and the corresponding UV component values in the second image include U2 and V2.

[0129] Figure 5 is a schematic flow chart for splitting the image to be encoded. For each frame of the image to be encoded in the video to be encoded, the image format of the image to be encoded is the YUV444 format. Split the YUV components of the YUV444 format image to obtain the Y component values and UV component values corresponding to the first image, and the Y' component values and UV' component values corresponding to the second image. Taking Figure 4 the component values of each pixel point in [example] as an example, among them, the Y component values corresponding to the first image include Y0 to Y4, the UV component values corresponding to the first image include U0 and V0, the Y' component values corresponding to the second image include U1, V1, U3, V3, and the UV' component values corresponding to the second image include U2 and V2.

[0130] Store Y0 to Y4 into the first view of the first image, store U0 and V0 into the second view of the first image, store U1, V1, U3, V3 into the first view of the second image, and store U2 and V2 into the second view of the second image.

[0131] Based on the first view and the second view of the first image, obtain the first image. Based on the first view and the second view of the second image, obtain the second image. Among them, both the first image and the second image are images in the NV12 format.

[0132] It can be understood that the first image format and the second image format are any formats in the YUV color mode, as long as the image quality corresponding to the first image format is higher than the image quality corresponding to the second image format. For example, when the first image format is the YUV444 format, the second image format can be the YUV422 format, the YUV420 format (NV12 format, NV21 format, YV12, etc.). When the first image format is the YUV422 format, the second image format can be the YUV420 format (NV12 format, NV21 format, YV12, etc.).

[0133] S130: Obtain the bitstreams of each frame of images by encoding the first image and the second image corresponding to each frame of the image to be encoded respectively.

[0134] For each frame of the image to be encoded, encode the first view and the second view of the first image corresponding to this frame of the image to be encoded to obtain the first bitstream corresponding to the first image; encode the first view and the second view of the second image corresponding to this frame of the image to be encoded to obtain the second bitstream corresponding to the second image. Obtain the bitstreams of each frame of images according to the first bitstream and the second bitstream corresponding to each frame of the image to be encoded.

[0135] Optionally, when encoding the first image and the second image, any encoding method can be adopted, such as H.264 encoding, H.265 encoding, etc. The embodiments of the present application do not limit this and can be set according to needs.

[0136] Based on Figure 1 The video encoding method shown above, in the video encoding process, by converting the image in the first image format with high quality in the video into the second image in the second image format with relatively low quality, and obtaining the second image based on the image conversion loss, and encoding the first image and the second image respectively, the encoding of high-quality images is realized without changing the device configuration, reducing the dependence on the device, improving the video encoding quality, and better meeting the video encoding requirements.

[0137] The video encoding method provided by the embodiments of the present application can be applied to the video transmission process. After compressing the video to be transmitted by the above encoding method, then transmit the compressed video. It can also be applied to the compression process of local videos, and use the above encoding method to compress the locally stored videos. It can also pre-encode and compress the local videos or the videos collected in real time and then perform video transmission. The embodiments of the present application do not limit this and can be set according to needs.

[0138] The embodiments of the present application also provide a video decoding method. Figure 6 As the schematic flowchart of the video decoding method provided by the embodiments of the present application, this method can be executed by any electronic device, such as a user terminal (which can also be referred to as a terminal, terminal device or user device, etc.) or a server, or can also be realized by the cooperation of multiple electronic devices.

[0139] Such as Figure 6 As shown, the video decoding method provided by the embodiments of the present application may include the following S210 to S240.

[0140] S210: Obtain the data to be decoded.

[0141] Among them, the data to be decoded includes the bitstreams corresponding to at least one frame of image. The bitstream corresponding to one frame of image includes the first bitstream of the first image corresponding to this frame of image and the second bitstream of the second image corresponding to this frame of image. This frame of image is an image in the first image format, the first image is an image in the second image format, the image quality of the image in the first image format is higher than that of the image in the second image format, and the second image is obtained based on the image conversion loss between this frame of image and the first image.

[0142] It should be noted that for the specific generation method of the data to be decoded, reference can be made to the above video encoding process. The embodiments of the present application will not elaborate here and can refer to the above content.

[0143] Optionally, the data to be decoded can be transmitted in real time, and the decoding end can receive the bitstream of the video in real time for decoding. The data to be decoded can also be stored locally, and the locally stored video is decoded. The embodiments of the present application do not limit this.

[0144] Optionally, the first image format and the second image format are different image formats in the YUV color mode, and the image quality corresponding to the first image format is higher than that corresponding to the second image format. For example, the first image format is YUV444 format, and the second image format is NV12 format.

[0145] S220: Parse the bitstreams corresponding to at least one frame of image to obtain the first bitstream and the second bitstream corresponding to at least one frame of image.

[0146] S230: Decode the first bitstream and the second bitstream corresponding to each frame of image respectively, and obtain the first image and the second image corresponding to each frame of image based on the decoding results.

[0147] Among them, the parsed first bitstream and second bitstream are respectively placed in different buffer queues, and different decoding threads read and call the decoder for decoding. The first image corresponding to each frame of decoded image is stored in the first queue, and the second image corresponding to each frame of decoded image is stored in the second queue. Among them, for each frame of image, the first bitstream corresponding to this frame of image carries the first identifier of this frame of image, the second bitstream corresponding to this frame of image carries the second identifier of this frame of image, and the first identifier and the second identifier represent the corresponding relationship between the first bitstream and the second bitstream corresponding to the same frame of image.

[0148] For each first image in the first queue, based on the first identifier of this first image and the second identifiers of the second images stored in the second queue, determine the second image corresponding to this first image from the second queue.

[0149] Optionally, if there is a second image that matches the first image, fuse the first image and the corresponding second image to obtain a reconstructed image, and render and display the reconstructed image. If there is no second image that matches the first image, render and display the first image with lower quality.

[0150] Optionally, when the data to be decoded is the bitstream of a real-time transmitted video, considering that there may be performance differences in the decoder, the first identifier of the first image can be matched and retrieved with the second identifiers of the second images stored in the second queue at a preset time interval. If a matching second image is retrieved before reaching the first threshold, the retrieved second image is used as the second image corresponding to the first image. Otherwise, it is determined that there is no second image that matches the first image, and the first image with lower quality is directly rendered and displayed to avoid video frame freezing caused by excessive waiting time.

[0151] Optionally, when decoding the video bitstream, the corresponding decoding algorithm can be determined, and the first bitstream and the second bitstream are decoded using the corresponding decoding algorithm. For example, when the first bitstream and the second bitstream use H.264 encoding, the first bitstream and the second bitstream are decoded through the H.264 decoding algorithm.

[0152] Video decoding can generally be divided into two categories. One is software decoding with the Central Processing Unit (CPU) as the core, and the other is hardware decoding with the Graphics Processing Unit (GPU) as the core. Among them, hardware decoding usually has higher decoding efficiency and is more suitable for scenarios with low latency and high smoothness.

[0153] Optionally, the first image in the first bitstream is an image in NV12 format, and the second image in the second bitstream is also an image in NV12 format. The hardware decoding method can be used to decode the first bitstream and the second bitstream through a hardware decoder, which improves the video decoding rate and makes the video playback smoother.

[0154] S240: By fusing the first image and the second image corresponding to each frame of the image, a reconstructed image corresponding to each frame of the image is obtained.

[0155] Since the second image format is an image format in the YUV color mode, the reconstructed image obtained by fusion is also an image format in the YUV color mode. Therefore, it is also necessary to convert the reconstructed image into an image in RGBA format and display it.

[0156] Optionally, for each first image, a third view and a fourth view corresponding to the first image are created. Among them, the third view and the fourth view can be created by CreateShaderResourceView (create shader resource view), and the third view and the fourth view are shader resource views. The Y component value corresponding to the first image is stored in the third view corresponding to the first image, and the UV component value corresponding to the first image is stored in the fourth view corresponding to the first image. Among them, the Y component value corresponding to the first image is the Y component value of each pixel in each frame of the video to be decoded, and the UV component value corresponding to the first image is the UV component value corresponding to the first pixel and the second pixel obtained by sampling.

[0157] For each second image, a third view and a fourth view corresponding to the second image are created. The luminance component value corresponding to the second image is stored in the third view corresponding to the second image, and the first chrominance component value and the second chrominance component value corresponding to the second image are stored in the fourth view corresponding to the second image. Among them, the Y component value corresponding to the second image is the first part of the component values that are not sampled, and the UV component value corresponding to the second image is the second part of the component values that are not sampled.

[0158] For each frame of image, the reconstructed image corresponding to the frame of image is obtained by fusing the third view and the fourth view corresponding to the frame of image.

[0159] Figure 7 is a schematic diagram of fusing the first image and the second image. The first image and the second image corresponding to each frame of image in the data to be decoded are both images in NV12 format. Taking the component values of each pixel in Figure 4 as an example, the Y component values Y0 to Y4 corresponding to the first image are stored in the third view, the UV component values U0 and V0 corresponding to the first image are stored in the fourth view, the Y component values U1, V1, U3, and V3 corresponding to the second image are stored in the third view, and the UV component values U2 and V2 corresponding to the second image are stored in the fourth view. Based on the third view and the fourth view of the first image, and the third view and the fourth view of the second image, the reconstructed image is obtained by fusion.

[0160] In the embodiments of the present application, the splitting rule of the YUV components of each pixel during video encoding corresponds to the merging (fusion) rule of the YUV components of each pixel during video decoding. When performing image fusion, for each pair of matching first image and second image, the corresponding UV component value in the second image is supplemented to the corresponding pixel in the first image. For example, Figure 4When fusing the first image (b) and the second image (c) obtained by splitting, the U1V1 component in (c) is supplemented to pixel point 1 in (b), the U2V2 component in (c) is supplemented to pixel point 2 in (b), and the U3V3 component in (c) is supplemented to pixel point 3 in (b).

[0161] The video decoding method provided by the embodiments of this application can be applied during the video transmission process. After receiving the video bitstream in real time, the above decoding method is used to decode the real-time transmitted video bitstream. It can also be applied during the decompression process of the local compressed video, and the above decoding method is used to decode the compressed video stored locally. The embodiments of this application do not limit this and can be set according to needs.

[0162] Based on Figure 6 The video decoding method shown, during the video decoding process, by separately decoding the first bitstream and the second bitstream in the data to be decoded, a first image and a second image are obtained, and the first image and the second image are fused to obtain a high-quality reconstructed image. It realizes the reconstruction of high-quality images without changing the device configuration, reduces the dependence on the device, and improves the video decoding quality.

[0163] In summary, the video encoding and decoding method provided by the embodiments of this application adopts a dual-channel encoding and decoding method. During the video encoding process, the image to be encoded is converted into a first image, and based on the image conversion loss, a second image is obtained. The first image and the second image are respectively encoded to obtain a first bitstream and a second bitstream. During the video decoding process, by parsing the data to be decoded, the first bitstream and the second bitstream are obtained. Among them, the second bitstream is supplementary content of the first bitstream. When an abnormal situation occurs, such as the second bitstream is lost, the rendering display can be directly based on the decoding result of the first bitstream.

[0164] Figure 8 It is a schematic flow diagram of the video encoding and decoding provided by the embodiments of this application. For each frame of the image to be encoded in the video to be encoded, the encoding end can convert the image to be encoded in the first image format into a first image in the second image format, where the image quality corresponding to the first image format is higher than the image corresponding to the second image format. The first image and the second image are respectively encoded to obtain a first bitstream corresponding to the first image and a second bitstream corresponding to the second image. The corresponding bitstream identifiers are respectively added to the first bitstream and the second bitstream, and the two bitstreams are merged into one bitstream and sent to the decoding end through the network.

[0165] After receiving the video bitstream, the decoding end parses and splits the video bitstream to obtain the first bitstream and the second bitstream, and respectively decodes the first bitstream and the second bitstream to obtain the corresponding first image and second image. The first image and the second image corresponding to the same frame of image are fused to obtain a reconstructed image.

[0166] To facilitate a better understanding and description of the method provided in the embodiments of the present application, the following introduces the optional implementation manners of the video encoding and decoding method provided in the present application in combination with several specific scenario embodiments. Taking the video conferencing scenario as an example, video conferencing requires sharing the screen simultaneously on multiple devices and has relatively high requirements for transmission latency, smoothness, etc.

[0167] Figure 9 FIG. is a schematic structural diagram of a video conferencing system in a video conferencing scenario. Among them, the video conferencing system includes a first terminal, multiple second terminals, and a conference server. The first terminal is the terminal of the presenter, and the second terminal is the terminal of the audience.

[0168] When the video conference is started, the first terminal collects the current conference video in real time, converts each frame of RGB image in the collected conference video into a YUV image to obtain a YUV444 image. For each frame of YUV444 image in the conference video, convert the frame image into a first NV12 image, and construct a second image based on the image conversion loss generated by converting the YUV444 image to the NV12 image. Among them, the constructed second image also belongs to the NV12 format image. Respectively perform H.264 encoding on the first image and the second image of each frame of image in the video to obtain a first bitstream and a second bitstream, and merge the first bitstream and the second bitstream into one bitstream and send it to the conference server.

[0169] Based on the bitstream of the video to be transmitted received, the conference server sends the bitstream of the video to each second terminal. For each second terminal, after receiving the video bitstream, the video bitstream can be parsed to obtain the first bitstream and the second bitstream corresponding to each frame of image, and based on the hardware decoding method, the first bitstream and the second bitstream are respectively decoded through the H.264 decoding algorithm to obtain the first image and the second image corresponding to each frame of image. By fusing the first image and the second image of each frame of image, the reconstructed image corresponding to each frame of image is obtained. Based on each frame of reconstructed image, the video content is displayed.

[0170] The video encoding and decoding method provided in the present application can also be applied to the scenario of video transmission, such as the transmission of long / short videos in a video player, the distribution of game scene videos in a game application, etc. Taking the cloud game scenario as an example for illustration, as Figure 10 shown, the video transmission system includes a game server and multiple user terminals. Among them, the game server is a cloud server.

[0171] When any user terminal receives a user's operation to obtain a game scene video, it sends a game scene video acquisition request to the game server. After receiving the acquisition request, the game server determines the game scene video to be sent and encodes the game scene video. Specifically, it converts the color mode of each frame image in the game scene video to the YUV444 format, and converts each frame of YUV444 to a first image in the lower-quality NV12 format. Based on the loss generated by the image conversion, a second image is constructed. The second image also belongs to the NV12 format image. The first image and the second image of each frame image in the video are respectively encoded by H.265 to obtain a first bitstream and a second bitstream, and the first bitstream and the second bitstream are merged into one bitstream and sent to the user terminal.

[0172] After receiving the bitstream of the game scene video, the user terminal can parse the bitstream of the game scene video to obtain the first bitstream and the second bitstream corresponding to each frame image in the game scene video, and use the hardware decoding method to decode the first bitstream and the second bitstream respectively through the H.265 decoding algorithm to obtain the first image and the second image corresponding to each frame image. The first image corresponding to each decoded frame image is stored in the first queue, and the second image corresponding to each decoded frame image is stored in the second queue. For each first image in the first queue, based on the first identifier of the first image and the second identifiers of the second images stored in the second queue, the second image corresponding to the first image is determined from the second queue. If the second image corresponding to the first image is matched, the first image and the corresponding second image are fused to obtain a reconstructed image and displayed. If the second image matching the first image is not found, the first image of lower quality is displayed.

[0173] The video encoding and decoding method provided by this application can also be applied to the video call scenario, such as Figure 11 shown, the structural schematic diagram of the call system corresponding to this video call scenario is as Figure 11 shown, including a call server and multiple user terminals.

[0174] For each user terminal participating in the video call, the user terminal real-time collects the current user video and encodes the collected user video. Specifically, the image in the first image format of each frame in the user video is converted to the first image in the second image format, and a second image is constructed based on the loss generated by the image conversion. The image quality corresponding to the first image format is higher than the image quality corresponding to the second image format. The first image and the second image of each frame image are respectively encoded to obtain a first bitstream and a second bitstream, and the first bitstream and the second bitstream are merged into one bitstream and sent to the call server.

[0175] After receiving the video bitstreams sent by each user terminal, the call server decodes the video bitstreams sent by each user terminal respectively. Specifically, for each video bitstream, the first bitstream and the second bitstream corresponding to each frame of image are parsed, the first bitstream and the second bitstream are decoded respectively to obtain the first image and the second image corresponding to each frame of image, and the first reconstructed image corresponding to each frame of image is obtained by fusing the first image and the second image of each frame of image.

[0176] Based on each frame of the first reconstructed image in the decoded video streams of each user terminal, the call server performs multi-terminal mixing and encodes the mixed video. The mixed video bitstream is sent to each user terminal. For each user terminal, according to the received mixed video bitstream, video decoding is performed to obtain the second reconstructed image corresponding to each frame of the mixed image and display it.

[0177] Based on the same principle as the video encoding method provided in the embodiments of the present application, the embodiments of the present application provide a video encoding device, as Figure 12 shown, the video encoding device 300 may include a video acquisition module 310, a conversion module 320, and an encoding module 330.

[0178] The video acquisition module 310 is used to acquire the video to be encoded, and the video to be encoded includes at least one frame of image to be encoded, and the image format of the image to be encoded is the first image format;

[0179] The conversion module 320 is used to convert each frame of the image to be encoded into a first image in the second image format for each frame of the image to be encoded, and obtain a second image based on the image conversion loss between the frame of the image to be encoded and the first image, wherein the image quality of the image to be encoded is higher than the image quality of the first image;

[0180] The encoding module 330 is used to obtain the bitstream of each frame of image by encoding the first image and the second image corresponding to each frame of the image to be encoded respectively.

[0181] Optionally, the first image format and the second image format are two image formats in the YUV color mode;

[0182] The conversion module 320 may be used for:

[0183] Determine the luminance component value of each pixel point in the frame of the image to be encoded as the luminance component value of each pixel point in the first image corresponding to the frame of the image to be encoded;

[0184] According to the differences between the first image format and the second image format, sample the first chrominance component value and the second chrominance component value of the image frame to be encoded, so as to obtain the first chrominance component value and / or the second chrominance component value of some pixel points in the first image corresponding to the image frame to be encoded;

[0185] Based on the first chrominance component value and the second chrominance component value of each pixel point in the image frame to be encoded that are not sampled, construct a second image.

[0186] Optionally, the image sizes of the first image and the second image are the same;

[0187] The conversion module 320 can be used to:

[0188] Use the first partial component values as the luminance component values of each pixel point in the second image to be constructed; wherein, the first partial component values are partial component values among the first chrominance component values and the second chrominance component values that are not sampled;

[0189] Based on the second partial component values, determine the first chrominance component value and / or the second chrominance component value of the pixel points in the second image corresponding to the partial pixel points in the first image, and the second partial component values are the component values among the first chrominance component values and the second chrominance component values that are not sampled except the first partial component values.

[0190] Optionally, the first image format is YUV444 format and the second image format is NV12 format;

[0191] The conversion module 320 can be used to:

[0192] For each 2×2 image block in the image frame to be encoded, sample the first chrominance component value of the first pixel point in the image block to obtain the first chrominance component value of the corresponding pixel point in the first image, and sample the second chrominance component value of the second pixel point in the image block to obtain the second chrominance component value of the corresponding pixel point in the first image; the first pixel point and the second pixel point are any pixel points in the image block;

[0193] For each 2×2 image block in the image frame to be encoded, based on the first chrominance component value and the second chrominance component value of each pixel point in the image block that are not sampled, construct the image block in the second image corresponding to the image block.

[0194] Optionally, the conversion module 320 can be used to:

[0195] Any 4 component values among the first chrominance component values and the second chrominance component values that are not sampled in the image block are respectively used as the luminance component values of each pixel point in the image block corresponding to the image block in the second image;

[0196] The 2 component values that are not sampled except for the 4 component values are respectively used as the first chrominance component value or the second chrominance component value of the pixel points corresponding to the first pixel point and the second pixel point in the second image.

[0197] Optionally, the conversion module 320 is further configured to:

[0198] Create a first view and a second view corresponding to the image to be processed;

[0199] Store the luminance component values of each pixel point in the image to be processed corresponding to the image to be encoded into the first view;

[0200] Store the first chrominance component value of the first pixel point and the second chrominance component value of the second pixel point in the frame to be encoded image into the second view;

[0201] The encoding module 330 can be used to:

[0202] Encode the first view and the second view of the image to be processed to obtain a bitstream corresponding to the image to be processed.

[0203] Based on Figure 12 The video encoding device shown. During the video encoding process, by converting an image in a first image format with high quality in the video into a second image in a second image format with relatively low quality, and obtaining the second image based on the image conversion loss, by encoding the first image and the second image respectively, the encoding of high-quality images is achieved without changing the device configuration, reducing the dependence on the device, improving the video encoding quality, and better meeting the video encoding requirements.

[0204] Based on the same principle as the video decoding method provided in the embodiments of the present application, the embodiments of the present application also provide a video decoding device, as Figure 13 shown. The video decoding device 400 may include an acquisition module 410, an analysis module 420, a decoding module 430, and a fusion reconstruction module 440. Among them:

[0205] An acquisition module 410 is configured to acquire data to be decoded. The data to be decoded includes a bitstream corresponding to at least one frame of image. The bitstream corresponding to one frame of image includes a first bitstream of a first image corresponding to this frame of image and a second bitstream of a second image corresponding to this frame of image. This frame of image is an image in a first image format, the first image is an image in a second image format, and the image quality of the image in the first image format is higher than that of the image in the second image format. The second image is obtained based on the image conversion loss between this frame of image and the first image.

[0206] An analysis module 420 is configured to analyze the bitstream corresponding to the at least one frame of image to obtain the first bitstream and the second bitstream corresponding to the at least one frame of image.

[0207] A decoding module 430 is configured to separately decode the first bitstream and the second bitstream corresponding to each frame of image, and obtain the first image and the second image corresponding to each frame of image based on the decoding result.

[0208] A fusion and reconstruction module 440 is configured to obtain a reconstructed image corresponding to each frame of image by fusing the first image and the second image corresponding to each frame of image.

[0209] Optionally, for each frame of image, a first identifier of this frame of image is carried in the first bitstream corresponding to this frame of image, and a second identifier of this frame of image is carried in the second bitstream corresponding to this frame of image.

[0210] The decoding module 430 may be configured to:

[0211] Store the first image corresponding to each decoded frame of image into a first queue, and store the second image corresponding to each decoded frame of image into a second queue.

[0212] For each first image in the first queue, based on the first identifier of this first image and the second identifiers of the second images stored in the second queue, determine the second image corresponding to this first image from the second queue.

[0213] Optionally, the acquisition module 410 may be configured to:

[0214] Receive the bitstream of the video in real time.

[0215] The decoding module may be configured to:

[0216] At a preset time interval, match and retrieve the first identifier of this first image with the second identifiers of the second images stored in the second queue. If a matching second image is retrieved before reaching a first threshold, use the retrieved second image as the second image corresponding to this first image.

[0217] The decoding module 430 is further configured to:

[0218] If no matching second image is retrieved when reaching the first threshold, a reconstructed image corresponding to the frame image is determined based on the first image.

[0219] Optionally, the decoding module 430 is further configured to:

[0220] Create a third view and a fourth view corresponding to the image to be processed;

[0221] Store the luminance component value corresponding to the image to be processed into the third view;

[0222] Store the first chrominance component value and the second chrominance component value corresponding to the image to be processed into the fourth view;

[0223] The fusion and reconstruction module 440 may be configured to:

[0224] For each frame of image, a reconstructed image corresponding to the frame of image is obtained by fusing the third view and the fourth view corresponding to the frame of image.

[0225] Optionally, the decoding module 430 may be configured to:

[0226] Decode the first bitstream and the second bitstream corresponding to each frame of image respectively by using a hardware decoding method.

[0227] Based on Figure 13 The video decoding device shown, during video decoding, by decoding the first bitstream and the second bitstream in the data to be decoded respectively, a first image and a second image are obtained, and the first image and the second image are fused to obtain a high-quality reconstructed image. The reconstruction of high-quality images is realized without changing the device configuration, the dependence on the device is reduced, and the video decoding quality is improved.

[0228] The device of the embodiments of the present application can execute the methods provided by the embodiments of the present application, and their implementation principles are similar. The actions performed by each module in the device of the embodiments of the present application correspond to the steps in the methods of the embodiments of the present application. For the detailed function descriptions of each module of the device, reference can be specifically made to the descriptions in the corresponding methods shown above, and details are not described herein again.

[0229] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit including the function of the module or unit.

[0230] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program stored in the memory, the method in any optional embodiment of the present application can be implemented.

[0231] Figure 14 The structural schematic diagram of an electronic device applicable to the embodiment of the present invention is shown, as Figure 14 shown. The electronic device can be a server or a user terminal, and the electronic device can be used to implement the method provided in any embodiment of the present invention.

[0232] As Figure 14 shown in Figure 14 it, the electronic device 2000 mainly includes at least one processor 2001 ( Figure 13 one is shown in

[0233] ), a memory 2002, a communication module 2003, an input / output interface 2004 and other components. Optionally, the components can be connected and communicate through a bus 2005. It should be noted that the structure of the electronic device 2000 shown in

[0234] The processor 2001 is connected to the memory 2002 via the bus 2005, and realizes corresponding functions by invoking the application programs stored in the memory 2002. Among them, the processor 2001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 2001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0235] The electronic device 2000 can be connected to a network through the communication module 2003 (which can include but is not limited to components such as a network interface) to communicate with other devices (such as user terminals or servers, etc.) through the network to achieve data interaction, such as sending data to other devices or receiving data from other devices. Among them, the communication module 2003 can include a wired network interface and / or a wireless network interface, etc., that is, the communication module can include at least one of a wired communication module or a wireless communication module.

[0236] The electronic device 2000 can be connected to the required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 2004. The electronic device 2000 itself can have a display device, and can also externally connect other display devices through the interface 2004. Optionally, a storage device, such as a hard disk, etc., can also be connected through the interface 2004 to store the data in the electronic device 2000 into the storage device, or read the data in the storage device, and can also store the data in the storage device into the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 2004 can be components of the electronic device 2000 or external devices connected to the electronic device 2000 when needed.

[0237] The bus 2005 for connecting each component may include a path for transmitting information between the above components. The bus 2005 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. According to different functions, the bus 2005 may be divided into an address bus, a data bus, a control bus, etc.

[0238] Optionally, for the solution provided in the embodiments of the present invention, the memory 2002 may be used to store a computer program for executing the solution of the present invention, and the processor 2001 runs the computer program to implement the actions of the method or device provided in the embodiments of the present invention.

[0239] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the corresponding content of the foregoing method embodiments can be implemented.

[0240] The embodiments of the present application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the corresponding content of the foregoing method embodiments can be implemented.

[0241] It should be noted that the terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than the illustrated or textually described order.

[0242] It should be understood that although the flowchart in the embodiments of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0243] The above are only alternative implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also fall within the protection scope of the embodiments of the present application.

Claims

1. A video encoding method, characterized in that, Including: Obtain a video to be encoded, where the video to be encoded includes at least one image to be encoded, and the image format of the image to be encoded is a first image format; For each image to be encoded, convert the image to be encoded into a first image in a second image format, and obtain a second image based on the image conversion loss between the image to be encoded and the first image, where the image quality of the image to be encoded is higher than the image quality of the first image; By encoding the first image and the second image corresponding to each image to be encoded respectively, obtain the bitstreams of each image.

2. The method according to claim 1, wherein The first image format and the second image format are two image formats in the YUV color mode; The converting the image to be encoded into a first image in a second image format and obtaining a second image based on the image conversion loss between the image to be encoded and the first image includes: Determine the luminance component values of each pixel point in the image to be encoded as the luminance component values of each pixel point in the first image corresponding to the image to be encoded; According to the difference between the first image format and the second image format, sample the first chrominance component value and the second chrominance component value of the image to be encoded to obtain the first chrominance component value and / or the second chrominance component value of some pixel points in the first image corresponding to the image to be encoded; Based on the first chrominance component value and the second chrominance component value of each pixel point in the image to be encoded that are not sampled, construct a second image.

3. The method according to claim 2, wherein The image sizes of the first image and the second image are the same; The constructing a second image based on the first chrominance component value and the second chrominance component value of each pixel point in the image to be encoded that are not sampled includes: Use a first part of the component values as the luminance component values of each pixel point in the second image to be constructed; where the first part of the component values is part of the first chrominance component value and the second chrominance component value that are not sampled; Based on a second part of the component values, determine the first chrominance component value and / or the second chrominance component value of the pixel points in the second image corresponding to the some pixel points in the first image, and the second part of the component values is the component values of the first chrominance component value and the second chrominance component value that are not sampled except the first part of the component values.

4. The method according to claim 2, characterized in that The first image format is YUV444 format, and the second image format is NV12 format; The sampling the first chrominance component value and the second chrominance component value of the image to be encoded according to the difference between the first image format and the second image format to obtain the first chrominance component value and / or the second chrominance component value of some pixel points in the first image corresponding to the image to be encoded includes: For each 2×2 image block in the frame of the image to be encoded, sample the first chrominance component value of the first pixel point in the image block to obtain the first chrominance component value of the corresponding pixel point in the first image, and sample the second chrominance component value of the second pixel point in the image block to obtain the second chrominance component value of the corresponding pixel point in the first image; the first pixel point and the second pixel point are any pixel points in the image block. Constructing a second image based on the first chrominance component values and second chrominance component values of the pixel points in the frame of the image to be encoded that have not been sampled includes: For each 2×2 image block in the frame of the image to be encoded, construct an image block corresponding to the image block in the second image based on the first chrominance component values and second chrominance component values of the pixel points in the image block that have not been sampled.

5. The method according to claim 4, characterized in that, For each 2×2 image block in the frame of the image to be encoded, constructing an image block corresponding to the image block in the second image based on the first chrominance component values and second chrominance component values of the pixel points in the image block that have not been sampled includes: Take any 4 component values from the first chrominance component values and second chrominance component values of the pixel points in the image block that have not been sampled, and respectively use them as the luminance component values of the pixel points in the image block corresponding to the image block in the second image; Take the 2 component values that have not been sampled except the above 4 component values, and respectively use them as the first chrominance component value or the second chrominance component value of the pixel points corresponding to the first pixel point and the second pixel point in the second image.

6. The method according to any one of claims 3 to 5, characterized in that, Taking the first image and the second image corresponding to each frame of the image to be encoded in the video to be encoded as the images to be processed respectively, for each image to be processed, the method further includes: Create a first view and a second view corresponding to the image to be processed; Store the luminance component values of the pixel points in the image to be processed corresponding to the image to be encoded into the first view; Store the first chrominance component value of the first pixel point and the second chrominance component value of the second pixel point in the frame of the image to be encoded into the second view; For each image to be processed, encode the image to be processed respectively, including: Encode the first view and the second view of the image to be processed to obtain a bitstream corresponding to the image to be processed.

7. A video decoding method, characterized in that, The method includes: Obtain data to be decoded; wherein, the data to be decoded includes bitstreams corresponding to at least one frame of image, the bitstream corresponding to one frame of image includes the first bitstream of the first image corresponding to the frame of image and the second bitstream of the second image corresponding to the frame of image, the frame of image is an image in the first image format, the first image is an image in the second image format, and the image quality of the image in the first image format is higher than the image quality of the image in the second image format; the second image is obtained based on the image conversion loss between the frame of image and the first image; Parse the bitstreams corresponding to the at least one frame of image to obtain the first bitstream and the second bitstream corresponding to the at least one frame of image; Decode the first bitstream and the second bitstream corresponding to each frame of image respectively, and obtain the first image and the second image corresponding to each frame of image based on the decoding results. By fusing the first image and the second image corresponding to each frame of image, a reconstructed image corresponding to each frame of image is obtained.

8. The method according to claim 7, wherein For each frame of image, a first identifier of the frame of image is carried in the first bitstream corresponding to the frame of image, and a second identifier of the frame of image is carried in the second bitstream corresponding to the frame of image; Said obtaining the first image and the second image corresponding to each frame of image based on the decoding result includes: Storing the first image corresponding to each decoded frame of image into a first queue, and storing the second image corresponding to each decoded frame of image into a second queue; For each first image in the first queue, based on the first identifier of the first image and the second identifiers of the second images stored in the second queue, determining the second image corresponding to the first image from the second queue.

9. The method according to claim 8, wherein Said obtaining the data to be decoded includes: Receiving the bitstream of the video in real time; Said determining the second image corresponding to the first image from the second queue based on the first identifier of the first image and the second identifiers of the second images stored in the second queue includes: At a preset time interval, matching and retrieving the first identifier of the first image with the second identifiers of the second images stored in the second queue. If a matching second image is retrieved before reaching a first threshold, the retrieved second image is used as the second image corresponding to the first image; The method further includes: If no matching second image is retrieved when reaching the first threshold, determining the reconstructed image corresponding to the frame of image based on the first image.

10. The method according to claim 7, characterized in that Taking the first image and the second image corresponding to each frame of image in the data to be decoded as images to be processed respectively. For each image to be processed, the method further includes: Creating a third view and a fourth view corresponding to the image to be processed; Storing the luminance component value corresponding to the image to be processed into the third view; Storing the first chrominance component value and the second chrominance component value corresponding to the image to be processed into the fourth view; Said obtaining the reconstructed image corresponding to each frame of image by fusing the first image and the second image corresponding to each frame of image includes: For each frame of image, by fusing the third view and the fourth view corresponding to the frame of image, obtaining the reconstructed image corresponding to the frame of image.

11. The method according to claim 7, wherein Said separately decoding the first bitstream and the second bitstream corresponding to each frame of image includes: Using a hardware decoding method to separately decode the first bitstream and the second bitstream corresponding to each frame of image.

12. A video encoding device, characterized in that, Includes: A video acquisition module, configured to acquire a video to be encoded, the video to be encoded includes at least one frame of image to be encoded, and the image format of the image to be encoded is a first image format; A conversion module, configured to, for each frame of image to be encoded, convert the frame of image to be encoded into a first image in a second image format, and based on the image conversion loss between the frame of image to be encoded and the first image, obtain a second image, wherein the image quality of the image to be encoded is higher than the image quality of the first image; An encoding module, configured to obtain the bitstreams of each frame of image by separately encoding the first image and the second image corresponding to each frame of image to be encoded.

13. A video decoding device, characterized in that, Includes: An acquisition module, configured to acquire data to be decoded; wherein the data to be decoded includes a bitstream corresponding to at least one frame of image, and the bitstream corresponding to one frame of image includes a first bitstream of a first image corresponding to this frame of image and a second bitstream of a second image corresponding to this frame of image, this frame of image is an image in a first image format, the first image is an image in a second image format, and the image quality of the image in the first image format is higher than that of the image in the second image format; the second image is obtained based on the image conversion loss between this frame of image and the first image; An analysis module, configured to analyze the bitstream corresponding to the at least one frame of image to obtain the first bitstream and the second bitstream corresponding to the at least one frame of image; A decoding module, configured to separately decode the first bitstream and the second bitstream corresponding to each frame of image, and obtain the first image and the second image corresponding to each frame of image based on the decoding result; A fusion and reconstruction module, configured to obtain a reconstructed image corresponding to each frame of image by fusing the first image and the second image corresponding to each frame of image.

14. An electronic device, characterized in that, The electronic device includes a memory and a processor, and a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 6 or claims 7 to 11.

15. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 or claims 7 to 11 is implemented.