Coding and decoding method and related device
By generating global reference tensors and motion information across time domains, the problem of poor encoding and decoding in existing video compression schemes is solved, and more efficient encoding and decoding is achieved, reducing cache pressure and improving image reconstruction quality.
Patent Information
- Application Number
- CN202410445360.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2024-04-12
- Publication Date
- 2025-07-25
AI Technical Summary
The existing video compression scheme has poor encoding and decoding performance and low efficiency, so it is impossible to effectively use the redundant information in the video for efficient encoding and decoding.
By determining the global reference tensor and motion information corresponding to the current image, the target prediction information across time domain is generated, the amount of data bits of the residual information is reduced, the buffering pressure at the encoding end is reduced, and the global reference tensor is generated using multiple coded images during the encoding process to improve the encoding and decoding efficiency.
Improve the efficiency and performance of encoding and decoding, ensure the accuracy of target prediction information, reduce the cache pressure on the encoding side, and improve the image reconstruction quality.
Smart Images

Figure CN120378612A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202410095535.5 and the invention title "Coding and Decoding Method, Device, Equipment, Storage Medium and Computer Program" filed on January 23, 2024, the entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the field of data compression, and particularly to a coding and decoding method and related devices. Background Art
[0003] Video compression refers to a technology that utilizes redundant information in a video to represent the original video with less data (video bitstream). Video compression can relieve the pressure on video storage and network bandwidth occupied by video transmission. This video compression technology includes encoding and decoding, and the coding and decoding performance (reflecting video quality) and coding and decoding efficiency (reflecting time consumption) are elements that need to be considered in video compression technology. However, the coding and decoding performance of some current video compression schemes is poor and the efficiency is low. Summary of the Invention
[0004] This application provides a coding and decoding method, device, equipment, storage medium and computer program, which can solve the problems of poor coding and decoding performance and low efficiency in related technologies. The technical solutions are as follows:
[0005] In a first aspect, a coding method is provided. The method includes: determining a global reference tensor corresponding to a current image and motion information of the current image, where the global reference tensor is determined based on a plurality of previously encoded images before the current image; determining target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image; determining residual information of the current image based on the target prediction information of the current image and the current image; and encoding the residual information and the motion information of the current image into a bitstream.
[0006] This application can determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image. Since the global reference tensor is determined based on multiple previously encoded images before the current image, that is, the embodiments of this application can generate a global reference tensor according to multiple previously encoded images before the current image, so that the global reference tensor can represent the image information available for reference in the current image among multiple previously encoded images before the current image, and has cross-temporal globality. In this way, it can ensure that the target prediction information corresponding to the current image obtained based on the global reference tensor corresponding to the current image and the motion information of the current image is more accurate, so that the data bit amount of the residual information corresponding to the current image determined subsequently is smaller, thereby effectively improving the encoding and decoding efficiency. In addition, generating a global reference tensor through multiple previously encoded images before the current image can enable the encoding end to ensure that there is more information available for the current image to reference when encoding the current image without caching multiple previously encoded images, thereby greatly reducing the caching pressure of the encoding end while improving the encoding and decoding efficiency.
[0007] Optionally, if the current image is in a video, when the current image is the first frame image in the video, the encoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to the first frame image is a feature of all 0s. When the current image is the Mth frame in the video, the encoding end can determine the global reference tensor corresponding to the current image and the motion information of the current image, where M is an integer greater than 1.
[0008] Optionally, if the current image is in a video and the video includes at least one GOP, the encoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to the I frame is a feature of all 0s.
[0009] When the current image is the I frame in the Nth GOP in the video, or is a P frame in the video, the encoding end can determine the global reference tensor corresponding to the current image and the motion information of the current image, where N is an integer greater than 1.
[0010] It should be noted that when the current image is the I frame in the GOP, if the encoding end determines that the global reference tensor corresponding to the current image is 0, that is, no matter which GOP in the video the current image is the I frame of, the encoding end determines that the global reference tensor corresponding to the current image is 0. In this case, the global reference tensor is determined based on multiple previously encoded images before the current image, and the multiple previously encoded images are the previously encoded images in the GOP where the current image is located.
[0011] Optionally, if the encoding end determines the global reference tensor corresponding to the current image to be 0 only when the current image is an I-frame in the first GOP of the video, and when the current image is an I-frame in the Nth GOP of the video, the encoding end determines the global reference tensor corresponding to the current image and the motion information of the current image. In this case, the global reference tensor corresponding to the current image is determined based on all the previously encoded images before the current image. At this time, the global reference tensor is not limited to one or several previously encoded images, but can represent all the available reference information in multiple previously encoded images before the current image, thereby increasing the information available for the current image to reference and further improving the encoding and decoding efficiency.
[0012] Optionally, the encoding end determines the reconstruction information of the first encoded image, where the first encoded image is an encoded image adjacent to the current image. Based on the reconstruction information of the first encoded image, the global reference tensor corresponding to the first encoded image is updated to obtain the global reference tensor corresponding to the current image.
[0013] It should be noted that the first encoded image is usually the previous encoded image adjacent to the current image. For example, if the current image is the t-th frame of the video, then the first encoded image is the (t - 1)-th frame.
[0014] Since in the subsequent steps, the decoding end needs to update the global reference tensor corresponding to the first decoded image, where the first decoded image is a decoded image adjacent to the current image, and the first decoded image at the decoding end and the first encoded image at the encoding end are the same image. For the sake of description, the first decoded image and the first encoded image will be referred to as the first image hereinafter. When the decoding end decodes, it cannot obtain the original first image at the encoding end, but can only obtain the reconstructed first image. That is to say, there may be a certain deviation between the first image reconstructed by the decoding end and the original first image. Therefore, the encoding end updates the global reference tensor corresponding to the first image based on the reconstruction information of the first image to ensure that the global reference tensors corresponding to the current image used by the encoding end and the decoding end subsequently are the same, so that the target prediction information of the current image obtained by the encoding end and the decoding end is more accurate, thereby effectively improving the encoding and decoding efficiency and performance.
[0015] If the current image is located in a video that includes at least one GOP, when the current image is the first P-frame in the GOP, the first encoded image can be the I-frame in the GOP. When the current image is the second P-frame in the GOP, the first encoded image is the first P-frame in the GOP.
[0016] If the current image is in the video, and when the encoding end determines that the global reference tensor corresponding to the current image is 0 when the current image is the first frame in the video, and when the current image is the Mth frame in the video, the global reference tensor corresponding to the current image and the motion information of the current image are determined. In this case, the global reference tensor corresponding to the Nth frame image in the video bitstream is obtained based on the global reference tensor corresponding to the encoded image of the (N - 1)th frame and the reconstruction information of the encoded image of the (N - 1)th frame. The global reference tensor corresponding to the second frame image is the reconstruction information of the encoded image of the first frame.
[0017] Optionally, the reconstruction information of the first encoded image includes the reconstructed first encoded image or the image features of the reconstructed first encoded image. In different cases, the implementation methods for determining the reconstruction information of the first encoded image are different, which will be introduced separately below.
[0018] If the reconstruction information of the first encoded image includes the reconstructed first encoded image, the encoding end can reconstruct the first encoded image to obtain the reconstructed image of the first encoded image. That is, based on the bitstream, the reconstructed motion information of the first encoded image is obtained, and then based on the global reference tensor corresponding to the first encoded image and the reconstructed motion information of the first encoded image, the target prediction information of the first encoded image is determined; the reconstructed residual information of the first encoded image is obtained based on the bitstream; based on the target prediction information of the first encoded image and the reconstructed residual information of the first encoded image, the first encoded image is reconstructed to obtain the reconstructed image of the first encoded image.
[0019] If the reconstruction information of the first encoded image includes the reconstructed first encoded image, the encoding end can reconstruct the first encoded image to obtain the reconstructed image of the first encoded image, and based on the reconstructed image of the first encoded image, the image features of the reconstructed first encoded image are determined.
[0020] Optionally, the encoding end inputs the reconstructed image of the first encoded image into the image feature extraction network to obtain the image features of the reconstructed first encoded image output by the image feature extraction network.
[0021] Optionally, the implementation process of updating the global reference tensor corresponding to the first encoded image based on the reconstruction information of the first encoded image to obtain the global reference tensor corresponding to the current image includes: inputting the reconstruction information of the first encoded image and the global reference tensor corresponding to the first encoded image into the global reference tensor update network to obtain the global reference tensor corresponding to the current image output by the global reference tensor update network.
[0022] Optionally, the motion information of the current image includes a motion vector, which is used to indicate the offset of the target block in the current image from its position in the reference image of the current image, where the reference image is the reconstructed image of a coded image adjacent to the current image. In this case, the reference image of the current image is the reconstructed image of the first coded image described above.
[0023] Before determining the motion information of the current image based on the image features of the current image and the image features of the reconstructed first coded image, the encoding end can extract the image features from the current image to obtain the image features of the current image.
[0024] Optionally, the encoding end inputs the current image into an image feature extraction network to obtain the image features of the current image output by the image feature extraction network.
[0025] Optionally, based on the global reference tensor corresponding to the current image and the motion information of the current image, the global prediction information corresponding to the current image is determined, and based on the global prediction information corresponding to the current image, the target prediction information of the current image is determined.
[0026] Optionally, the encoding end can determine the reconstructed motion information of the current image based on the motion information of the current image, and determine the global motion information corresponding to the current image based on the reconstructed motion information of the current image. The global motion information includes a global motion vector, which indicates the offset of the target block in the current image from its position in the global reference tensor corresponding to the current image. Based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image, the global prediction information corresponding to the current image is determined.
[0027] Since the decoding end also needs to determine the global motion information corresponding to the current image in subsequent steps, but the decoding end cannot obtain the original motion information of the current image during decoding and can only obtain the reconstructed motion information of the current image. That is to say, there may be a certain deviation between the motion information reconstructed by the decoding end and the motion information corresponding to the original current image. Therefore, the encoding end determines the global motion information corresponding to the current image through the reconstructed motion information of the current image to ensure that the global motion information corresponding to the current image adopted by the encoding end and the decoding end in the subsequent steps is the same, so that the encoding end and the decoding end can obtain the target prediction information of the current image more accurately, thereby effectively improving the efficiency and performance of encoding and decoding.
[0028] Optionally, the encoding end updates the global motion information corresponding to the first coded image based on the reconstructed motion information of the current image to obtain the global motion information corresponding to the current image.
[0029] Optionally, the encoding end inputs the reconstructed motion information of the current image and the global motion information corresponding to the first encoded image into the global motion information update network to obtain the global motion information corresponding to the current image output by the global motion information update network.
[0030] Optionally, the encoding end determines the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
[0031] Optionally, the encoding end can also directly determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the motion information of the current image.
[0032] Since the motion information of the current image can, to a certain extent, characterize the motion of the target block in the current image relative to the encoded image, the encoding end can directly determine the global prediction information corresponding to the current image based on the motion information of the current image and the global reference tensor corresponding to the current image. Compared with determining the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image, in this way, while ensuring the accuracy of the subsequent determined target prediction information, the computational amount of video compression can be effectively reduced, and the encoding and decoding efficiency can be improved.
[0033] Optionally, the global prediction information corresponding to the current image can be directly determined as the target prediction information of the current image.
[0034] Optionally, the encoding end determines the local prediction information corresponding to the current image based on the reconstructed motion information of the current image and the reconstructed information of the first encoded image, and determines the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image.
[0035] The implementation process of determining the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image includes: fusing the local prediction information corresponding to the current image and the global prediction information corresponding to the current image to obtain the target prediction information of the current image.
[0036] It should be noted that when the reconstructed information of the first encoded image includes the reconstructed first encoded image, the target prediction information of the current image is the same as the channel dimension and spatial dimension of the current image. When the reconstructed information of the first encoded image includes the image features of the reconstructed first encoded image, the target prediction information of the current image is the same as the channel dimension and spatial dimension of the image features of the current image.
[0037] Optionally, if the reconstruction information of the first encoded image includes the image features of the reconstructed first encoded image, in this case, the encoding end can determine the residual information corresponding to the current image based on the target prediction information of the current image and the image features of the current image.
[0038] Optionally, if the reconstruction information of the first encoded image includes the reconstructed first encoded image, in this case, the encoding end can determine the residual information of the current image based on the target prediction information of the current image and the current image.
[0039] It should be noted that the residual information of the current image and the motion information of the current image can be encoded into different bitstreams respectively, or the residual information of the current image and the motion information of the current image can be encoded into the same bitstream.
[0040] In a second aspect, a decoding method is provided. The method includes: obtaining the reconstructed motion information of the current image based on the bitstream; determining the global reference tensor corresponding to the current image, where the global reference tensor is determined based on a plurality of decoded images before the current image; determining the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image; obtaining the reconstructed residual information of the current image based on the bitstream; and obtaining the reconstructed image of the current image based on the target prediction information of the current image and the reconstructed residual information of the current image.
[0041] In the decoding process of the present application, accurate target prediction information can be obtained through the reconstructed motion information of the current image and the global reference tensor corresponding to the current image. In this way, it can be ensured that the image quality of the reconstructed image of the current image obtained subsequently based on the target prediction information and the reconstructed residual information corresponding to the current image is better, thereby greatly improving the performance of encoding and decoding.
[0042] Optionally, the reconstructed motion information of the current image includes a motion vector, and the motion vector is used to indicate the offset of the target block in the current image from its position in the reference image of the current image, where the reference image is the reconstructed image of a decoded image adjacent to the current image. In this case, the reference image of the current image is the reconstructed image of the first decoded image in the following text, or rather, the reference image of the current image is the reconstructed first decoded image.
[0043] If the current image is in a video, when the current image is the first frame image in the video, the decoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to the first frame image is a feature of all 0s. When the current image is the Mth frame in the video, the decoding end can determine the global reference tensor corresponding to the current image, where M is an integer greater than 1.
[0044] If the current image is in a video that includes at least one GOP, the decoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to this I-frame is a feature of all 0s. When the current image is an I-frame in the Nth GOP in the video, or is a P-frame in the video, the decoding end can determine the global reference tensor corresponding to the current image, where N is an integer greater than 1.
[0045] If the decoding end determines the global reference tensor corresponding to the current image as 0 only when the current image is an I-frame in the first GOP in the video, and determines the global reference tensor corresponding to the current image when the current image is an I-frame in the Nth GOP in the video, in this case, the global reference tensor corresponding to the current image is determined based on all the decoded images before the current image.
[0046] Optionally, the reconstruction information of the first decoded image, where the first decoded image is a decoded image adjacent to the current image, and based on the reconstruction information of the first decoded image, update the global reference tensor corresponding to the first decoded image to obtain the global reference tensor corresponding to the current image.
[0047] It should be noted that the first decoded image is usually the previous decoded image adjacent to the current image.
[0048] When the current image is in a video that includes at least one GOP, if the current image is the first P-frame in the GOP, the first decoded image can be the I-frame in the GOP. If the current image is the second P-frame in the GOP, the first decoded image is the first P-frame in the GOP.
[0049] If the current image is in a video, and the decoding end determines that the global reference tensor corresponding to the current image is 0 when the current image is the first frame image in the video, and determines the global reference tensor corresponding to the current image when the current image is the Mth frame in the video, in this case, the global reference tensor corresponding to the Nth frame image in the video bitstream is obtained based on the global reference tensor corresponding to the decoded image of the N-1th frame and the reconstruction information of the decoded image of the N-1th frame, and the global reference tensor corresponding to the second frame image is the reconstruction information of the decoded image of the first frame.
[0050] Optionally, the reconstruction information of the first decoded image includes the reconstructed first decoded image, or the image features of the reconstructed first decoded image.
[0051] In the case where the reconstruction information of the first decoded image includes the reconstructed first decoded image, the decoding end stores the reconstructed first decoded image, so that the decoding end can directly determine the reconstruction information of the first decoded image.
[0052] Optionally, the decoding end can also reconstruct the first decoded image, that is, obtain the reconstructed motion information of the first decoded image based on the bitstream, and then determine the target prediction information of the first decoded image based on the global reference tensor corresponding to the first decoded image and the reconstructed motion information of the first decoded image; obtain the reconstructed residual information of the first decoded image based on the bitstream; reconstruct the first decoded image based on the target prediction information and the reconstructed residual information of the first decoded image to obtain the reconstructed image of the first decoded image.
[0053] If the reconstruction information of the first decoded image includes the reconstructed first decoded image, the decoding end can reconstruct the first decoded image to obtain the reconstructed first decoded image, and determine the image features of the reconstructed first decoded image based on the reconstructed first decoded image.
[0054] Optionally, the process of updating the global reference tensor corresponding to the first decoded image based on the reconstruction information of the first decoded image to obtain the global reference tensor corresponding to the current image includes: inputting the reconstruction information of the first decoded image and the global reference tensor corresponding to the first decoded image into the global reference tensor update network to obtain the global reference tensor corresponding to the current image output by the global reference tensor update network.
[0055] Optionally, determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image, and determine the target prediction information of the current image based on the global prediction information corresponding to the current image.
[0056] Optionally, the decoding end determines the reconstructed motion information of the current image based on the reconstructed motion information of the current image, determines the global motion information corresponding to the current image based on the reconstructed motion information of the current image, where the global motion information is used to describe the position offset of the current image relative to the global reference tensor corresponding to the current image, and determines the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
[0057] Optionally, the decoding end can update the global motion information corresponding to the first decoded image based on the reconstructed motion information of the current image to obtain the global motion information corresponding to the current image.
[0058] Optionally, the decoding end can directly determine the global prediction information corresponding to the current image as the target prediction information of the current image.
[0059] Optionally, the decoding end determines the local prediction information corresponding to the current image based on the reconstructed motion information of the current image and the reconstruction information of the first decoded image, and determines the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image.
[0060] Optionally, if the reconstruction information of the first decoded image includes the image features of the reconstructed first decoded image, in this case, the decoding end determines the image features of the reconstructed image of the current image according to the relevant algorithm based on the target prediction information of the current image and the reconstruction residual information of the current image; based on the image features of the reconstructed image of the current image, the current image is reconstructed according to the relevant algorithm to obtain the reconstructed image of the current image.
[0061] Optionally, the decoding end inputs the target prediction information of the current image and the reconstruction residual information of the current image into the reconstruction network to obtain the image features of the reconstructed image of the current image output by the reconstruction network.
[0062] Optionally, if the reconstruction information of the first decoded image includes the reconstructed first decoded image, in this case, the decoding end can reconstruct the current image based on the target prediction information of the current image and the reconstruction residual information of the current image to obtain the reconstructed image of the current image.
[0063] In a third aspect, an encoding device is provided, and the encoding device has a function of implementing the behavior of the encoding method in the first aspect above. The encoding device includes at least one module, and the at least one module is used to implement the encoding method provided in the first aspect above.
[0064] In a fourth aspect, a decoding device is provided, and the decoding device has a function of implementing the behavior of the decoding method in the first aspect above. The decoding device includes at least one module, and the at least one module is used to implement the decoding method provided in the first aspect above.
[0065] In a fifth aspect, an encoding device is provided, and the encoding device includes: a processor, the processor is coupled with a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the encoding device executes the encoding method described in the first aspect above.
[0066] In a sixth aspect, a decoding device is provided, and the decoding device includes: a processor, the processor is coupled with a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the decoding device executes the decoding method described in the second aspect above.
[0067] In a seventh aspect, a coding and decoding system is provided. The coding and decoding system includes the coding device described in the fifth aspect and / or the decoding device described in the sixth aspect.
[0068] In an eighth aspect, a computer-readable storage medium is provided, including program code which, when running on a computer, causes the computer to execute the method described in the first aspect above.
[0069] In a ninth aspect, a computer-readable storage medium is provided, including program code which, when running on a computer, causes the computer to execute the method described in the second aspect above.
[0070] In a tenth aspect, a computer program product is provided, including instructions which, when running on a computer, cause the computer to execute the method described in the first aspect above.
[0071] In an eleventh aspect, a computer program product is provided, including instructions which, when running on a computer, cause the computer to execute the method described in the second aspect above.
[0072] In a twelfth aspect, a computer-readable storage medium is provided, on which a bitstream obtained by the method described in the first aspect above and executed by one or more processors is stored.
[0073] In a thirteenth aspect, a device for storing a bitstream is provided, which is characterized by including at least one storage medium and a communication interface; the communication interface is used for receiving or sending the bitstream; the at least one storage medium is used for storing the bitstream; the bitstream is encoded by an encoder according to the encoding method described in the first aspect above.
[0074] In a fourteenth aspect, a method for storing a bitstream is provided, including: receiving the bitstream through the communication interface; storing the bitstream in one or more storage media, where the bitstream is encoded by an encoder according to the encoding method described in the first aspect above.
[0075] In a fifteenth aspect, a system for distributing a bitstream is provided, including at least one storage medium and a video stream device; the at least one storage medium is used for storing the bitstream, where the bitstream is encoded by an encoder according to the encoding method described in the first aspect above;
[0076] the video stream device is used for, in response to a request from a decoder, causing the target bitstream in the at least one storage medium to be sent to the decoder.
[0077] Sixteenth aspect, a method for distributing a bitstream is provided, including: receiving a first request; in response to the first request, selecting a target bitstream from at least one storage medium; sending the target bitstream to a destination device; the at least one storage medium is used for storing bitstreams, and the bitstreams are encoded by an encoder according to the encoding method described in the first aspect above.
[0078] Seventeenth aspect, a system for processing a bitstream is provided, characterized by including an image source device, an encoder device, one or more storage media, and a destination device; the image source device is used for providing image data; the encoder device is used for obtaining the image data of the image source device through an interface and encoding the image data to obtain one or more bitstreams, and the bitstreams are encoded by the encoder according to the encoding method described in the first aspect above; the encoder device is used for storing the one or more bitstreams into one or more storage media; alternatively, the encoder device is used for encapsulating the one or more bitstreams to obtain a transport bitstream; the encoder device is used for transmitting the transport bitstream to the destination device through a communication link or a communication network; the destination device is used for de-encapsulating the transport bitstream to obtain the one or more bitstreams; the destination device is used for decoding the one or more bitstreams to obtain decoded data.
[0079] The technical effects obtained in the second aspect to the seventeenth aspect above are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be elaborated here. Description of the Drawings
[0080] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0081] Figure 2 is a schematic diagram of another implementation environment provided by an embodiment of the present application;
[0082] Figure 3 is a flowchart of an encoding method provided by an embodiment of the present application;
[0083] Figure 4 is a schematic structural diagram of an image feature extraction network provided by an embodiment of the present application;
[0084] Figure 5 is a schematic structural diagram of a ResBlock sub-network provided by an embodiment of the present application;
[0085] Figure 6 is a schematic structural diagram of a global reference tensor update network provided by an embodiment of the present application;
[0086] Figure 7It is a schematic structural diagram of a global motion information update network provided by an embodiment of the present application;
[0087] Figure 8 It is a flowchart of a decoding method provided by an embodiment of the present application;
[0088] Figure 9 It is a schematic structural diagram of a reconstruction network provided by an embodiment of the present application;
[0089] Figure 10 It is a flowchart of an encoding and decoding method provided by an embodiment of the present application;
[0090] Figure 11 It is a flowchart of another encoding and decoding method provided by an embodiment of the present application;
[0091] Figure 12 It is a schematic diagram of a test result provided by an embodiment of the present application;
[0092] Figure 13 It is a flowchart of another encoding and decoding method provided by an embodiment of the present application;
[0093] Figure 14 It is a schematic structural diagram of an encoding device provided by an embodiment of the present application;
[0094] Figure 15 It is a schematic structural diagram of a decoding device provided by an embodiment of the present application. Detailed implementation manners
[0095] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0096] For ease of understanding, before explaining the encoding and decoding method provided by the embodiments of the present application in detail, the nouns, application scenarios, and implementation environments involved in the embodiments of the present application will be introduced first.
[0097] First, the nouns involved in the embodiments of the present application will be introduced.
[0098] Video compression: including intra-frame prediction coding and inter-frame prediction coding. Intra-frame prediction coding does not require the use of reference images, and inter-frame prediction coding requires using the current image and reference images to determine the inter-frame motion information, and using the inter-frame motion information to compress the video.
[0099] Group of pictures (GOP): The video bitstream includes multiple GOPs. A GOP is a group of consecutive pictures and is the basic unit accessed by a video image encoder and decoder. Each GOP includes an I frame and a P frame.
[0100] Reference Image: In video compression, a reference image is an encoded or decoded image that is used for the encoding or decoding of the current image.
[0101] I-Frame: An I-frame is usually an image encoded using intra-frame prediction and is also called a key frame. An I-frame is compressed without referring to other pictures. An I-frame describes the details of the image background and moving objects, and a complete image can be reconstructed using only the data of the I-frame during decoding. An I-frame is usually the first frame of each GOP.
[0102] P-Frame: A P-frame is usually an image encoded using inter-frame prediction and is also called a forward prediction frame (forward reference frame). A P-frame represents the difference between this frame and a previous I-frame (or P-frame). A P-frame uses motion compensation to transmit the prediction residual and motion vector between it and the previous reference image (i.e., an I or P-frame).
[0103] Inter-Frame Prediction Coding: It mainly includes two parts, one is the inter-frame prediction part, and the other is the residual compression part. The inter-frame prediction part includes a prediction and compression module for motion information and a transformation module. In some related technologies, motion information is embodied as optical flow. During the encoding process, the images of the reference frame and the current frame are input into an optical flow estimation network to obtain the predicted optical flow and compress the optical flow. In some other related technologies, motion information is embodied as motion features. During the encoding process, the image features of the current frame and the reference frame are extracted, and the image features of the current frame and the reference frame are input into a convolutional neural network to obtain the predicted motion features and compress the motion features. The transformation module usually uses a wrap operation. During the encoding process, using inter-frame side information, the reference frame is transformed into the prediction result of the current frame.
[0104] In the embodiments of the present application, the above-mentioned motion information (also called motion vector) can be optical flow or motion features, and the embodiments of the present application do not limit this.
[0105] IPPP Coding Mode: A coding mode in video compression. The first frame is an I-frame, and the subsequent frames are P-frames, and a P-frame only uses one forward reference frame. For example, for the current image (i.e., the t-th frame), the reference image of the t-th frame is usually the (t - 1)-th frame.
[0106] Bitrate: In image compression, it refers to the encoding length required for encoding per pixel. The higher the bitrate, the larger the size of the transmitted file and the better the image reconstruction quality.
[0107] R-D-λ: R refers to the bitrate; D refers to the reconstruction quality; λ is a network parameter used to adjust the bitrate. Usually, there is a corresponding relationship between λ and R.
[0108] Rate - distortion curve (RD - Curve): The abscissa is the bit rate, and the ordinate is the PSNR. Generally, the higher the curve, the better the encoding and decoding performance.
[0109] Next, the application scenarios related to the embodiments of this application will be introduced.
[0110] Video compression refers to a technology that uses redundant information in a video to represent the original video with less data (video bitstream). Video compression can relieve the pressure on video storage and network bandwidth occupied by video transmission. This video compression technology includes encoding and decoding. Encoding and decoding performance (reflecting video quality) and encoding and decoding efficiency (reflecting time consumption) are elements that need to be considered in video compression technology. With the development of society and the improvement of technology level, in recent years, video content has accounted for more than 80% of the entire consumer Internet traffic, and the file size of the original video is getting larger and larger, which poses great challenges to both encoding and decoding performance and encoding and decoding efficiency.
[0111] In the end - to - end video compression method, a neural network model is trained by using a single loss function to balance the distortion degree and the bit rate, and the compression performance is improved through a multi - reference video frame mechanism. However, the number of reference video frames in this multi - reference video frame mechanism is still relatively fixed, that is, the number of reference video frames is fixed after the neural network model is trained, and the number and position of reference video frames cannot be freely adjusted and controlled. Moreover, as the number of reference video frames increases, the requirements for the computing power and cache size of the device will also be higher.
[0112] Based on this, an embodiment of the present application provides an encoding and decoding method, which can determine the target prediction information of the current image through the global reference tensor corresponding to the current image and the motion information of the current image. Since the global reference tensor is determined based on multiple previously encoded images before the current image, that is, the embodiment of the present application can generate a global reference tensor according to multiple previously encoded images before the current image, so that the global reference tensor can represent the image information in multiple previously encoded images before the current image that can be used as a reference for the current image, and has cross-temporal globality. In this way, it can be ensured that the target prediction information corresponding to the current image obtained based on the global reference tensor corresponding to the current image and the motion information of the current image is more accurate, so that the data bit amount of the residual information corresponding to the current image determined subsequently is smaller, thereby effectively improving the encoding and decoding efficiency. In addition, by generating a global reference tensor from multiple previously encoded images before the current image, it can be ensured that the encoding end has more information for the current image to refer to when encoding without caching multiple previously encoded images, thereby greatly reducing the caching pressure of the encoding end while improving the encoding and decoding efficiency. During the decoding process, accurate target prediction information can be obtained through the reconstructed motion information of the current image and the global reference tensor corresponding to the current image. In this way, it can be ensured that the image quality of the reconstructed image of the current image obtained subsequently based on the target prediction information and the reconstructed residual information corresponding to the current image is better, thereby greatly improving the performance of encoding and decoding.
[0113] Next, the implementation environment involved in the embodiment of the present application will be introduced.
[0114] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a source device 10, a destination device 20, a link 30, and a storage device 40. Among them, the source device 10 can generate an encoded image, that is, a bitstream. Therefore, the source device 10 can also be referred to as an encoding device. The destination device 20 can decode the bitstream generated by the source device 10. Therefore, the destination device 20 can also be referred to as a decoding device. The link 30 can receive the encoded image generated by the source device 10 and can transmit the encoded image to the destination device 20. The storage device 40 can receive the encoded image generated by the source device 10 and can store the encoded image. Under such conditions, the destination device 20 can directly obtain the encoded image from the storage device 40. Alternatively, the storage device 40 can correspond to a file server or another intermediate storage device that can store the encoded image generated by the source device 10. Under such conditions, the destination device 20 can obtain the encoded image stored in the storage device 40 via streaming or downloading.
[0115] Both the source device 10 and the destination device 20 may include one or more processors and a memory coupled to the one or more processors. The memory may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium that can be used to store the desired program code in the form of computer-accessible instructions or data structures, etc. For example, both the source device 10 and the destination device 20 may include a mobile phone, a smart phone, a personal digital assistant (PDA), a wearable device, a pocket PC (PPC), a tablet computer, a smart in-vehicle unit, a smart TV, a smart speaker, a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video game console, an in-vehicle computer, or the like.
[0116] The link 30 may include one or more media or devices capable of transmitting the encoded image from the source device 10 to the destination device 20. In one possible implementation, the link 30 may include one or more communication media capable of enabling the source device 10 to directly send the encoded image to the destination device 20 in real time. In the embodiments of the present application, the source device 10 may modulate the encoded image based on a communication standard, which may be a wireless communication protocol, etc., and may send the modulated image to the destination device 20. The one or more communication media may include wireless and / or wired communication media. For example, the one or more communication media may include the radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, and the packet-based network may be a local area network, a wide area network, or a global network (e.g., the Internet), etc. The one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from the source device 10 to the destination device 20. The embodiments of the present application do not make specific limitations thereto.
[0117] In a possible implementation, the storage device 40 can store the received encoded image sent by the source device 10, and the destination device 20 can directly obtain the encoded image from the storage device 40. Under such conditions, the storage device 40 can include any one of a variety of distributed or locally accessible data storage media. For example, any one of the variety of distributed or locally accessible data storage media can be a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing a bitstream, etc.
[0118] In a possible implementation, the storage device 40 can correspond to a file server or another intermediate storage device that can store the bitstream generated by the source device 10, and the destination device 20 can obtain the image stored in the storage device 40 via streaming or downloading. The file server can be any type of server capable of storing the encoded image and sending the encoded image to the destination device 20. In a possible implementation, the file server can include a web server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive, etc. The destination device 20 can obtain the encoded image through any standard data connection (including an Internet connection). Any standard data connection can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for obtaining the encoded image stored on the file server. The transmission of the encoded image from the storage device 40 can be a streaming transmission, a download transmission, or a combination of both.
[0119] Figure 1 The shown implementation environment is only one possible implementation, and the technology of the embodiments of the present application can not only be applicable to Figure 1 the shown source device 10 capable of encoding an image and the destination device 20 capable of decoding the encoded image, but also applicable to other devices capable of encoding an image and decoding a bitstream. The embodiments of the present application do not make specific limitations thereto.
[0120] In Figure 1In the illustrated implementation environment, source device 10 includes data source 120, encoder 100, and output interface 140. In some embodiments, output interface 140 may include a regulator / demodulator (modem) and / or a transmitter, where the transmitter may also be referred to as a transceiver. Data source 120 may include an image capture device (e.g., a camera, etc.), an archive containing previously captured images, a feed interface for receiving images from an image content provider, and / or a computer graphics system for generating images, or a combination of these sources of images.
[0121] Data source 120 may send an image to encoder 100, and encoder 100 may encode the received image sent by data source 120 to obtain an encoded image. The encoder may send the encoded image to the output interface. In some embodiments, source device 10 directly sends the encoded image to destination device 20 via output interface 140. In other embodiments, the encoded image may also be stored on storage device 40 for later retrieval by destination device 20 and used for decoding and / or display.
[0122] In Figure 1 the illustrated implementation environment, destination device 20 includes input interface 240, decoder 200, and display device 220. In some embodiments, input interface 240 includes a receiver and / or a modem. Input interface 240 may receive the encoded image via link 30 and / or from storage device 40, and then send it to decoder 200, which may decode the received encoded image to obtain a decoded image. The decoder may send the decoded image to display device 220. Display device 220 may be integrated with destination device 20 or may be external to destination device 20. Generally, display device 220 displays the decoded image. Display device 220 may be any of a variety of types of display devices. For example, display device 220 may be a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0123] Although Figure 1Although not shown in the figure, in some aspects, the encoder 100 and the decoder 200 may each be integrated with an encoder and a decoder, and may include appropriate multiplexer - demultiplexer (MUX - DEMUX) units or other hardware and software for encoding both audio and images in a common data stream or a separate data stream. In some embodiments, if applicable, the MUX - DEMUX unit may conform to the ITU H.223 multiplexer protocol or other protocols such as the user datagram protocol (UDP).
[0124] Each of the encoder 100 and the decoder 200 may be any one of the following circuits: one or more microprocessors, digital signal processors (DSPs), application - specific integrated circuits (ASICs), field - programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the techniques of the embodiments of the present application are implemented partially in software, the device may store instructions for the software in a suitable non - volatile computer - readable storage medium and may execute the instructions in hardware using one or more processors to implement the techniques of the embodiments of the present application. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) may be regarded as one or more processors. Each of the encoder 100 and the decoder 200 may be included in one or more encoders or decoders, and any one of the encoders or decoders may be integrated as part of a combined encoder / decoder (codec) in the corresponding device.
[0125] The embodiments of the present application may generally refer to the encoder 100 as "signaling" or "transmitting" certain information to another device such as the decoder 200. The terms "signaling" or "transmitting" may generally refer to the transfer of syntax elements and / or other data for decoding a compressed image. This transfer may occur in real - time or almost real - time. Alternatively, this communication may occur after a period of time, for example, when storing syntax elements in a computer - readable storage medium in an encoded bitstream during encoding, and the decoding device may then retrieve the syntax elements at any time after storing the syntax elements in this medium.
[0126] Figure 2It is a schematic diagram of another implementation environment provided by an embodiment of the present application. This implementation environment includes an encoding end and a decoding end. The encoding end includes an AI encoding module, an entropy encoding module, and a file saving module. The decoding end includes a file loading module, an entropy decoding module, and an AI decoding module.
[0127] During the compression process, after the encoding end obtains the image to be compressed (for example, the image in the video collected by a camera, etc.), the AI encoding module obtains the residual information and motion information of the current image, and then the entropy encoding module performs entropy encoding on the residual information and motion information of the current image to obtain a bitstream file, and the file saving module saves the bitstream file to obtain a compressed file. The compressed file is input to the decoding end. The decoding end loads the compressed file through the file loading module, and obtains the reconstructed image through the entropy decoding module and the AI decoding module.
[0128] Optionally, the process of the AI encoding module and the AI decoding module for processing data is implemented on an embedded neural network processing unit (NPU) to improve the data processing efficiency, and the processes such as entropy encoding, saving files, and loading files are implemented on a central processing unit (CPU).
[0129] Optionally, the encoding end and the decoding end are one device, or the encoding end and the decoding end are two independent devices. If the encoding end and the decoding end are one device, this device can encode an image through the encoding method provided by the embodiment of the present application, and can also decode the image through the decoding method provided by the embodiment of the present application. If the encoding end and the decoding end are two independent devices, the encoding method provided by the embodiment of the present application can be applied to the encoding end of these two devices, and the decoding method provided by the embodiment of the present application can be applied to the decoding end of these two devices. That is to say, for one device, this device has both the function of image compression and the function of image decompression, or this device has the function of image compression or the function of image decompression.
[0130] The encoding and decoding methods provided by the embodiments of the present application can be applied to various scenarios, such as business scenarios such as cloud storage, video surveillance, live broadcast, and transmission, and can be specifically applied to terminal video recording, video albums, cloud storage, etc. The images encoded and decoded in various scenarios can be the images included in a video file. It should be noted that in combination with Figure 1 the shown implementation environment, any of the encoding methods hereinafter can be executed by the encoder 100 in the source device 10. Any of the decoding methods hereinafter can be executed by the decoder 200 in the destination device 20. In combination with Figure 2In the implementation environment shown, any of the encoding methods described below can be executed by the encoding end. Any of the decoding methods described below can be executed by the decoding end.
[0131] Figure 3 is a flowchart of an encoding method provided by an embodiment of the present application. This method is applied to the encoding end. Please refer to Figure 3 and the method includes the following steps.
[0132] Step 301: Determine the global reference tensor corresponding to the current image and the motion information of the current image. This global reference tensor is determined based on multiple previously encoded images before the current image.
[0133] In some embodiments, if the current image is in a video that includes at least one GOP, the encoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to this I-frame is a feature of all 0s.
[0134] When the current image is an I-frame in the Nth GOP in the video or is a P-frame in the video, the encoding end can determine the global reference tensor corresponding to the current image and the motion information of the current image, where N is an integer greater than 1.
[0135] That is to say, when the current image is an I-frame in the first GOP in the video, the encoding end determines that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to this I-frame is a feature of all 0s. When the current image is an I-frame in the Nth GOP in the video, the encoding end determines the global reference tensor corresponding to the current image and the motion information of the current image, or directly determines that the global reference tensor corresponding to the current image is 0.
[0136] It should be noted that when the current image is an I-frame in the GOP, if the encoding end determines that the global reference tensor corresponding to the current image is 0, that is, no matter which GOP in the video the current image is in, the encoding end determines that the global reference tensor corresponding to the current image is 0. In this case, the global reference tensor is determined based on multiple previously encoded images before the current image, and these multiple previously encoded images are the encoded images in the GOP where the current image is located. In other words, the global reference tensor is updated in units of GOP, and the global reference tensor is updated within each GOP and reset to 0 when crossing GOPs.
[0137] If the encoding end determines that the current image is an I-frame in the first GOP of the video, the global reference tensor corresponding to the current image is determined to be 0. When the current image is an I-frame in the Nth GOP of the video, the encoding end determines the global reference tensor corresponding to the current image and the motion information of the current image. In this case, the global reference tensor corresponding to the current image is determined based on all the previously encoded images before the current image. At this time, the global reference tensor is not limited to one or several previously encoded images, but can represent all the available reference information in multiple previously encoded images before the current image, thereby increasing the information available for the current image to reference and further improving the encoding and decoding efficiency.
[0138] Optionally, when the current image is an I-frame in the GOP, the encoding end can encode the current image according to relevant I-frame encoding techniques. As an example, the encoding end can encode the current image according to the intra-frame prediction encoding technique.
[0139] In some other embodiments, if the current image is in the video, when the current image is the first frame image in the video, the encoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to the first frame image has all-zero characteristics. When the current image is the Mth frame in the video, the encoding end can determine the global reference tensor corresponding to the current image and the motion information of the current image, where M is an integer greater than 1.
[0140] Optionally, when the current image is the first frame image in the GOP, the encoding end can encode the current image according to relevant I-frame encoding techniques. As an example, the encoding end can encode the current image according to the intra-frame prediction encoding technique.
[0141] Next, the implementation method for determining the global reference tensor corresponding to the current image and the implementation method for determining the motion information of the current image will be introduced separately through (1)-(2).
[0142] (1) Determine the global reference tensor corresponding to the current image.
[0143] In some embodiments, the encoding end determines the reconstruction information of the first encoded image, where the first encoded image is an encoded image adjacent to the current image. Based on the reconstruction information of the first encoded image, the global reference tensor corresponding to the first encoded image is updated to obtain the global reference tensor corresponding to the current image.
[0144] It should be noted that the first encoded image is usually the previous encoded image adjacent to the current image. For example, if the current image is the tth frame of the video, the first encoded image is the (t - 1)th frame.
[0145] Since the decoding end needs to update the global reference tensor corresponding to the first decoded image in subsequent steps, the first decoded image is a decoded image adjacent to the current image, and the first decoded image at the decoding end and the first encoded image at the encoding end are the same image. For ease of description, the first decoded image and the first encoded image will be referred to as the first image hereinafter. When the decoding end decodes, it cannot obtain the original first image at the encoding end, but only the reconstructed first image. That is to say, there may be a certain deviation between the first image reconstructed by the decoding end and the original first image. Therefore, the encoding end updates the global reference tensor corresponding to the first image based on the reconstruction information of the first image to ensure that the global reference tensors corresponding to the current image adopted by the encoding end and the decoding end are the same, so that the target prediction information of the current image obtained by the encoding end and the decoding end is more accurate, thereby effectively improving the encoding and decoding efficiency and performance.
[0146] Exemplarily, if the current image is in a video and the video includes at least one GOP, if the current image is the first P frame in the GOP, the first encoded image can be the I frame in the GOP. If the current image is the second P frame in the GOP, the first encoded image is the first P frame in the GOP.
[0147] Exemplarily, if the current image is in a video and the encoding end determines that the global reference tensor corresponding to the current image is 0 when the current image is the first frame image in the video, and determines the global reference tensor corresponding to the current image and the motion information of the current image when the current image is the Mth frame in the video. In this case, the global reference tensor corresponding to the Nth frame image in the video stream is obtained based on the global reference tensor corresponding to the (N - 1)th encoded image and the reconstruction information of the (N - 1)th encoded image, and the global reference tensor corresponding to the second frame image is the reconstruction information of the first encoded image of the first frame.
[0148] In some embodiments, the reconstruction information of the first encoded image includes the reconstructed first encoded image or the image features of the reconstructed first encoded image. In different cases, the implementation methods for determining the reconstruction information of the first encoded image are different, which will be introduced separately hereinafter.
[0149] If the reconstruction information of the first encoded image includes the reconstructed first encoded image, the encoding end can reconstruct the first encoded image to obtain the reconstructed image of the first encoded image. That is, the reconstructed motion information of the first encoded image is obtained based on the bitstream, and then the target prediction information of the first encoded image is determined based on the global reference tensor corresponding to the first encoded image and the reconstructed motion information of the first encoded image; the reconstructed residual information of the first encoded image is obtained based on the bitstream; the first encoded image is reconstructed based on the target prediction information of the first encoded image and the reconstructed residual information of the first encoded image to obtain the reconstructed image of the first encoded image.
[0150] The implementation manner of determining the target prediction information of the first encoded image based on the global reference tensor corresponding to the first encoded image and the reconstructed motion information of the first encoded image is similar to the implementation manner of determining the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image in the subsequent encoding method; the implementation manner of obtaining the reconstructed image of the first encoded image based on the target prediction information of the first encoded image and the reconstructed residual information of the first encoded image is similar to the implementation manner of obtaining the reconstructed image of the current image based on the target prediction information of the current image and the reconstructed residual information of the current image in the subsequent decoding method. The detailed content will be introduced later and will not be elaborated here.
[0151] If the reconstruction information of the first encoded image includes the reconstructed first encoded image, the encoding end can reconstruct the first encoded image to obtain the reconstructed image of the first encoded image, and determine the image features of the reconstructed first encoded image based on the reconstructed image of the first encoded image.
[0152] In some embodiments, the encoding end can extract image features from the reconstructed image of the first encoded image according to relevant feature extraction techniques to obtain the image features of the reconstructed first encoded image.
[0153] Optionally, the encoding end inputs the reconstructed image of the first encoded image into an image feature extraction network to obtain the image features of the reconstructed first encoded image output by the image feature extraction network.
[0154] Exemplarily, Figure 4 is a schematic structural diagram of an image feature extraction network provided by an embodiment of the present application. The image feature extraction network includes a convolutional layer (Conv) and three consecutive ResBlock sub-networks with the same configuration. Among them, [c in , h in , w in indicates that the number of channels of the input tensor of the image feature extraction network is c in, the size of the input tensor of the image feature extraction network in the x-axis direction is h in , the size of the input tensor of the image feature extraction network in the y-axis direction is w in ; Conv(ks, c in , c out indicates that the convolution kernel size of the convolutional layer is ks, the number of channels of the input tensor of the convolutional layer is c in , the number of channels of the output tensor of the convolutional layer is c out , and the stride of the convolution kernel movement is 2. ResBlock(c out ) indicates that the number of channels of the output tensor of the ResBlock sub-network is c out . [c out , h in / 2, w in / 2] indicates that the number of channels of the output tensor of the image feature extraction network is c out , the size of the output tensor of the image feature extraction network in the x-axis direction is h in / 2, and the size of the output tensor of the image feature extraction network in the y-axis direction is w in / 2. In this case, the input tensor of the image feature extraction network is the reconstructed image of the first encoded image above, and the output tensor of the image feature extraction network is the image feature of the reconstructed first encoded image.
[0155] Exemplarily, Figure 5 is a schematic structural diagram of a ResBlock sub-network provided by an embodiment of the present application. The ResBlock sub-network includes two convolutional layers (Conv) and an activation layer (such as an activation layer constructed based on Relu, LRelu, or other activation functions). Among them, in Conv(3, c, c, 1), the convolution kernel size of the convolutional layer is 3×3, the number of channels of the input tensor and the output tensor of the convolutional layer are both c, and the stride of the convolution kernel movement is 1.
[0156] It should be noted that the schematic structural diagrams of all the networks shown in the embodiments of the present application are only for example, and in actual applications, they can also be other structures. The embodiments of the present application do not limit the structure and dimensions of the networks.
[0157] In some embodiments, based on the reconstruction information of the first encoded image, the implementation process of updating the global reference tensor corresponding to the first encoded image to obtain the global reference tensor corresponding to the current image includes: inputting the reconstruction information of the first encoded image and the global reference tensor corresponding to the first encoded image into the global reference tensor update network to obtain the global reference tensor corresponding to the current image output by the global reference tensor update network.
[0158] Exemplarily,Figure 6 It is a schematic structural diagram of a global reference tensor update network provided by an embodiment of the present application. The global reference tensor update network includes three convolutional layers (Conv) and three consecutive ResBlock sub-networks with the same configuration. Among them, It is indicated that the number of channels of the reconstruction information of the first encoded image is 3, the size of the reconstruction information of the first encoded image in the x-axis direction is h, and the size of the reconstruction information of the first encoded image in the y-axis direction is w. Conv(ks, 3, c, 2) indicates that the size of the convolutional kernel of the convolutional layer is ks, the number of channels of the input tensor of the convolutional layer is 3, the number of channels of the output tensor of the convolutional layer is c, and the step size of the convolutional kernel movement is 2. It is indicated that the number of channels of the global reference tensor corresponding to the first encoded image is c, and the size of the global reference tensor corresponding to the first encoded image in the x-axis direction is h ref and the size of the global reference tensor corresponding to the first encoded image in the y-axis direction is w ref . Conv(ks, c×2, c, 2) indicates that the size of the convolutional kernel of the convolutional layer is ks, the number of channels of the input tensor of the convolutional layer is c×2, the number of channels of the output tensor of the convolutional layer is c, and the step size of the convolutional kernel movement is 2. ResBlock(c) indicates that the number of channels of the output tensor of the ResBlock sub-network is c, and the structure of the ResBlock sub-network is as Figure 5 shown. Conv(ks, c, c, 2) indicates that the size of the convolutional kernel of the convolutional layer is ks, the number of channels of the input tensor of the convolutional layer and the number of channels of the output tensor of the convolutional layer are both c, and the step size of the convolutional kernel movement is 2. It is indicated that the number of channels of the global reference tensor corresponding to the current image is c, and the size of the global reference tensor corresponding to the current image in the x-axis direction is h ref and the size of the global reference tensor corresponding to the current image in the y-axis direction is w ref .
[0159] (2) Determine the motion information of the current image.
[0160] Among them, the motion information of the current image includes a motion vector, and the motion vector is used to indicate the offset of the target block in the current image from that in the reference image of the current image. The reference image is the reconstructed image of the encoded image adjacent to the current image. In this case, the reference image of the current image is the reconstructed image of the above-mentioned first encoded image, or rather, the reference image of the current image is the reconstructed first encoded image.
[0161] It should be noted that in the embodiment of the present application, the reconstructed first encoded image and the reconstructed image of the first encoded image have the same meaning.
[0162] In some embodiments, the current image includes at least one target block, which may be a coding unit (CU) in the HEVC standard. In other embodiments, the target block may also be a pixel point in the current image or a set composed of multiple pixel points. The specific composition of the target block is not limited in the embodiments of the present application.
[0163] It should also be noted that if the motion information of the current image is determined through a network model, the current image needs to be input into the network model. In this case, the encoder can convert the current image into a tensor corresponding to the current image, and then input the tensor corresponding to the current image into the network model. At this time, the target block may also be an element in the tensor corresponding to the current image or a set composed of multiple elements.
[0164] Based on the above description, the reconstruction information of the first encoded image includes the reconstructed first encoded image or the image features of the reconstructed first encoded image. In different cases, the implementation methods for determining the motion information of the current image are different, and will be introduced separately below.
[0165] In the case where the reconstruction information of the first encoded image includes the reconstructed first encoded image, the encoding end can determine the motion information of the current image based on the current image and the reconstructed first encoded image according to relevant motion estimation algorithms such as optical flow estimation.
[0166] In the case where the reconstruction information of the first encoded image includes the image features of the reconstructed first encoded image, the encoding end can determine the motion information of the current image based on the image features of the current image and the image features of the reconstructed first encoded image according to relevant motion estimation algorithms such as optical flow estimation.
[0167] In some embodiments, before determining the motion information of the current image based on the image features of the current image and the image features of the reconstructed first encoded image, the encoding end can extract the image features from the current image according to relevant feature extraction techniques to obtain the image features of the current image.
[0168] Optionally, the encoding end inputs the current image into an image feature extraction network to obtain the image features of the current image output by the image feature extraction network.
[0169] Exemplarily, the image feature extraction network may be as Figure 4 shown. In this case, the input tensor of the image feature extraction network is the above-mentioned current image, and the output tensor of the image feature extraction network is the image features of the current image.
[0170] In summary, when the reconstruction information of the first encoded image includes the reconstructed first encoded image, the motion vector is used to indicate the offset of the target block in the current image from that in the reconstructed first encoded image; when the reconstruction information of the first encoded image includes the image features of the reconstructed first encoded image, the motion vector is used to indicate the offset of the target block in the image features of the current image from those in the image features of the reconstructed first encoded image.
[0171] Step 302: Determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image.
[0172] Next, the implementation manner of determining the target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image will be introduced in detail through (1)-(2).
[0173] (1) Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the motion information of the current image.
[0174] In some embodiments, the encoding end can determine the reconstructed motion information of the current image based on the motion information of the current image, determine the global motion information corresponding to the current image based on the reconstructed motion information of the current image, the global motion information includes a global motion vector, and the global motion vector indicates the offset of the target block in the current image from that in the global reference tensor corresponding to the current image, and determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
[0175] Optionally, the encoding end can encode the motion information of the current image into the bitstream, and then parse the bitstream to obtain the reconstructed motion information of the current image.
[0176] Similarly to the above, since the decoding end also needs to determine the global motion information corresponding to the current image in subsequent steps, but the decoding end cannot obtain the original motion information of the current image during decoding and can only obtain the reconstructed motion information of the current image. That is to say, there may be a certain deviation between the motion information reconstructed by the decoding end and the motion information corresponding to the original current image. Therefore, the encoding end determines the global motion information corresponding to the current image through the reconstructed motion information of the current image to ensure that the global motion information corresponding to the current image adopted by the encoding end and the decoding end subsequently is the same, so that the encoding end and the decoding end can obtain the target prediction information of the current image more accurately, thereby effectively improving the efficiency and performance of encoding and decoding.
[0177] In some embodiments, the encoding end updates the global motion information corresponding to the first encoded image based on the reconstructed motion information of the current image to obtain the global motion information corresponding to the current image.
[0178] Optionally, the encoding end inputs the reconstructed motion information of the current image and the global motion information corresponding to the first encoded image into the global motion information update network to obtain the global motion information corresponding to the current image output by the global motion information update network.
[0179] Exemplarily, Figure 7 is a schematic structural diagram of a global motion information update network provided by an embodiment of the present application. The global motion information update network includes three convolutional layers (Conv) and three consecutive ResBlock sub-networks with the same configuration. Among them, represents that the number of channels of the reconstructed motion information of the current image is c m , the size of the reconstructed motion information of the current image in the x-axis direction is h ref , the size of the reconstructed information of the first encoded image in the y-axis direction is w ref . Conv(ks, 3, c m , 1) indicates that the convolution kernel size of the convolutional layer is ks, the number of channels of the input tensor of the convolutional layer is 3, the number of channels of the output tensor of the convolutional layer is c m , and the step size of the convolution kernel movement is 1. represents that the number of channels of the global motion information corresponding to the first encoded image is c m , the size of the global motion information corresponding to the first encoded image in the x-axis direction is h ref , the size of the global motion information corresponding to the first encoded image in the y-axis direction is w ref . Conv(ks, c m ×2, c m ), 2) indicates that the convolution kernel size of the convolutional layer is ks, the number of channels of the input tensor of the convolutional layer is c m ×2, the number of channels of the output tensor of the convolutional layer is c m , and the step size of the convolution kernel movement is 2. ResBlock(c m ) indicates that the number of channels of the output tensor of the ResBlock sub-network is c m , and the structure of the ResBlock sub-network is as Figure 5 shown. Conv(ks, c m , c m , 2) indicates that the convolution kernel size of the convolutional layer is ks, and the number of channels of the input tensor and the output tensor of the convolutional layer are both c m , and the step size of the convolution kernel movement is 2. The number of channels representing the global motion information corresponding to the current image is c m , the size of the global motion information corresponding to the current image in the x-axis direction is h ref , the size of the global motion information corresponding to the current image in the y-axis direction is w ref .
[0180] In some embodiments, the encoding end determines the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image, according to relevant prediction techniques such as optical flow mapping (warping) and deformable convolution. The embodiments of the present application do not limit this
[0181] In some other embodiments, the encoding end can also directly determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the motion information of the current image, according to relevant prediction techniques such as optical flow mapping (warping) and deformable convolution. The embodiments of the present application do not limit this
[0182] Since the motion information of the current image can, to a certain extent, characterize the motion of the target block in the current image relative to the encoded image, the encoding end can directly determine the global prediction information corresponding to the current image based on the motion information of the current image and the global reference tensor corresponding to the current image. Compared with determining the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image, in this way, while ensuring the accuracy of the subsequent determined target prediction information, the computational amount of image compression can be effectively reduced, and the encoding and decoding efficiency can be improved
[0183] (2) Determine the target prediction information of the current image based on the global prediction information corresponding to the current image
[0184] In some embodiments, the encoding end can directly determine the global prediction information corresponding to the current image as the target prediction information of the current image
[0185] In some other embodiments, the encoding end determines the local prediction information corresponding to the current image based on the reconstructed motion information of the current image and the reconstructed information of the first encoded image, and determines the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image
[0186] Optionally, based on the reconstructed motion information of the current image and the reconstruction information of the first encoded image, the encoding end determines the local prediction information corresponding to the current image according to relevant prediction techniques such as optical flow warping and deformable convolution.
[0187] The implementation process of determining the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image includes: fusing the local prediction information corresponding to the current image and the global prediction information corresponding to the current image to obtain the target prediction information of the current image.
[0188] It should be noted that when the reconstruction information of the first encoded image includes the reconstructed first encoded image, the target prediction information of the current image is the same as the channel dimension and spatial dimension of the current image. When the reconstruction information of the first encoded image includes the image features of the reconstructed first encoded image, the target prediction information of the current image is the same as the channel dimension and spatial dimension of the image features of the current image.
[0189] In some embodiments, the channel dimensions and spatial dimensions of the target prediction information of the current image, the local prediction information corresponding to the current image, and the global prediction information corresponding to the current image are all the same, where the spatial dimension includes the dimension of the first tensor in the x-axis direction and the dimension in the y direction, and the channel dimension includes the dimension of the first tensor in the z direction.
[0190] For any element at any position in the target prediction information of the current image for any channel, based on the first data corresponding to the local prediction information and the global prediction information respectively, the target first data is determined, and the target first data is used as the element value at that position in the target prediction information of the channel. The first data is the element value at the corresponding position in the local prediction information or the global prediction information of the channel. Processing each element at each position in the target prediction information of each channel in the same way can obtain the target prediction information.
[0191] Optionally, the implementation process of determining the target first data corresponding to the multiple first data based on the first data corresponding to the local prediction information and the global prediction information respectively includes: taking the weighted average value of the first data corresponding to the local prediction information and the global prediction information respectively as the target first data. Of course, in practical applications, the maximum value or the minimum value of the first data corresponding to the local prediction information and the global prediction information respectively can also be used as the target first data, and the embodiments of the present application do not limit this.
[0192] In some other embodiments, the encoding end may fuse the local prediction information corresponding to the current image and the global prediction information corresponding to the current image according to relevant feature fusion techniques to obtain the target prediction information of the current image. For example, the local prediction information corresponding to the current image and the global prediction information corresponding to the current image are input into a fusion network to obtain the target prediction information of the current image output by the fusion network.
[0193] Step 303: Determine the residual information of the current image based on the target prediction information of the current image and the current image.
[0194] In some embodiments, if the reconstruction information of the first encoded image includes the image features of the reconstructed first encoded image, in this case, the encoding end can determine the residual information corresponding to the current image based on the target prediction information of the current image and the image features of the current image according to relevant algorithms.
[0195] In some other embodiments, if the reconstruction information of the first encoded image includes the reconstructed first encoded image, in this case, the encoding end can determine the residual information of the current image based on the target prediction information of the current image and the current image according to relevant algorithms.
[0196] Step 304: Encode the residual information of the current image and the motion information of the current image into a bitstream.
[0197] In some embodiments, the encoding end may encode the residual information of the current image and the motion information of the current image into a bitstream according to relevant algorithms, so that the subsequent decoding end can decode the current image based on the residual information of the current image and the motion information of the current image in the bitstream.
[0198] It should be noted that the residual information of the current image and the motion information of the current image may be encoded into different bitstreams respectively, or the residual information of the current image and the motion information of the current image may be encoded into the same bitstream. The embodiments of the present application do not limit this.
[0199] In the embodiment of the present application, the target prediction information of the current image is determined based on the global reference tensor corresponding to the current image and the motion information of the current image. Since the global reference tensor is determined based on multiple previously encoded images before the current image, that is, the embodiment of the present application can generate a global reference tensor according to multiple previously encoded images before the current image, so that the global reference tensor can represent the image information available for reference in the current image in multiple previously encoded images before the current image, and has cross-temporal globality. In this way, it can be ensured that the target prediction information corresponding to the current image obtained based on the global reference tensor corresponding to the current image and the motion information of the current image is more accurate, so that the data bit amount of the residual information corresponding to the current image determined subsequently is smaller, thereby effectively improving the encoding and decoding efficiency. In addition, by generating a global reference tensor from multiple previously encoded images before the current image, it can be ensured that the encoding end has more information available for reference by the current image when encoding the current image without caching multiple previously encoded images, thereby greatly reducing the caching pressure of the encoding end while improving the encoding and decoding efficiency.
[0200] Since the decoding end also needs to determine the global motion information corresponding to the current image in subsequent steps, but the decoding end cannot obtain the motion information of the original current image when decoding, and can only obtain the reconstructed motion information of the current image. That is to say, there may be a certain deviation between the motion information reconstructed by the decoding end and the motion information corresponding to the original current image. Therefore, the encoding end determines the global motion information corresponding to the current image through the reconstructed motion information of the current image, so as to ensure that the global motion information corresponding to the current image adopted by the encoding end and the decoding end subsequently is the same, so that the target prediction information of the current image obtained by the encoding end and the decoding end is more accurate, thereby effectively improving the encoding and decoding efficiency and performance. The encoding end directly determines the global prediction information corresponding to the current image based on the motion information of the current image and the global reference tensor corresponding to the current image, which can effectively reduce the computational amount of image compression and improve the encoding and decoding efficiency. Since the motion information of the current image can represent to a certain extent the motion of the target block in the current image relative to the encoded images, the encoding end can directly determine the global prediction information corresponding to the current image based on the motion information of the current image and the global reference tensor corresponding to the current image, which can effectively reduce the computational amount of image compression while ensuring the accuracy of the target prediction information determined subsequently, and improve the encoding and decoding efficiency.
[0201] Figure 8 is a flowchart of a decoding method provided by an embodiment of the present application. This method is applied to the decoding end. Please refer to Figure 8 , and this method includes the following steps.
[0202] Step 801: Obtain the reconstructed motion information of the current image based on the bitstream.
[0203] Among them, the reconstructed motion information of the current image includes a motion vector, which is used to indicate the offset of the target block in the current image from its position in the reference image of the current image. The reference image is the reconstructed image of a decoded image adjacent to the current image. In this case, the reference image of the current image is the reconstructed image of the first decoded image in the following text, or rather, the reference image of the current image is the reconstructed first decoded image.
[0204] It should be noted that in the embodiments of the present application, the reconstructed first decoded image and the reconstructed image of the first decoded image have the same meaning.
[0205] In some embodiments, the current image includes at least one target block, which may be a coding unit (CU) in the HEVC standard. In other embodiments, the target block may also be a pixel point in the current image or a set composed of multiple pixel points. In still other embodiments, the target block may also be an element in the tensor corresponding to the current image or a set composed of multiple elements. The embodiments of the present application do not limit the specific composition of the target block.
[0206] Step 802: Determine the global reference tensor corresponding to the current image, which is determined based on multiple decoded images before the current image.
[0207] In some embodiments, if the current image is in a video that includes at least one GOP, the decoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to the I-frame has all-zero characteristics.
[0208] When the current image is an I-frame in the Nth GOP in the video or is a P-frame in the video, the decoding end can determine the global reference tensor corresponding to the current image, where N is an integer greater than 1.
[0209] That is to say, when the current image is an I-frame in the first GOP in the video, the decoding end determines that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to the I-frame has all-zero characteristics. When the current image is an I-frame in the Nth GOP in the video, the decoding end determines the global reference tensor corresponding to the current image, or determines that the global reference tensor corresponding to the current image is 0.
[0210] It should be noted that when the current image is an I-frame in the GOP, if the decoding end determines that the global reference tensor corresponding to the current image is 0, that is, regardless of which GOP's I-frame in the video the current image is, the decoding end determines that the global reference tensor corresponding to the current image is 0. In this case, the global reference tensor is determined based on multiple decoded images before the current image, and these multiple decoded images are the decoded images in the GOP where the current image is located. In other words, the global reference tensor is updated in units of GOPs, updated within each GOP, and reset to 0 when crossing GOPs.
[0211] If the decoding end determines that the global reference tensor corresponding to the current image is 0 only when the current image is the I-frame in the first GOP of the video, and when the current image is the I-frame in the Nth GOP of the video, the decoding end determines the global reference tensor corresponding to the current image. In this case, the global reference tensor corresponding to the current image is determined based on all decoded images before the current image.
[0212] Optionally, when the current image is an I-frame in the GOP, the decoding end can decode the current image according to relevant I-frame decoding techniques.
[0213] In some other embodiments, if the current image is in the video, when the current image is the first frame image in the video, the decoding end can determine that the global reference tensor corresponding to the current image is 0, that is, the global reference tensor corresponding to the first frame image has the feature of all 0s. When the current image is the Mth frame in the video, the decoding end can determine the global reference tensor corresponding to the current image, where M is an integer greater than 1.
[0214] Optionally, when the current image is the first frame image in the GOP, the decoding end can decode the current image according to relevant I-frame decoding techniques.
[0215] In some embodiments, the decoding end determines the reconstruction information of the first decoded image, where the first decoded image is a decoded image adjacent to the current image. Based on the reconstruction information of the first decoded image, the global reference tensor corresponding to the first decoded image is updated to obtain the global reference tensor corresponding to the current image.
[0216] It should be noted that the first decoded image is usually the previous decoded image adjacent to the current image. For example, if the current image is the tth frame of the video, then the first decoded image is the (t - 1)th frame.
[0217] Exemplarily, when the current image is in a video that includes at least one GOP, if the current image is the first P frame in the GOP, the first decoded image can be the I frame in the GOP. If the current image is the second P frame in the GOP, the first decoded image is the first P frame in the GOP.
[0218] Exemplarily, if the current image is in a video, and when the decoding end determines that the global reference tensor corresponding to the current image is 0 when the current image is the first frame image in the video, and determines the global reference tensor corresponding to the current image when the current image is the Mth frame in the video. In this case, the global reference tensor corresponding to the Nth frame image in the video bitstream is obtained based on the global reference tensor corresponding to the decoded image of the (N - 1)th frame and the reconstruction information of the decoded image of the (N - 1)th frame, and the global reference tensor corresponding to the second frame image is the reconstruction information of the decoded image of the first frame.
[0219] In some embodiments, the reconstruction information of the first decoded image includes the reconstructed first decoded image or the image features of the reconstructed first decoded image.
[0220] In the case where the reconstruction information of the first decoded image includes the reconstructed first decoded image, in some embodiments, the decoding end stores the reconstructed first decoded image. Thus, the decoding end can directly determine the reconstruction information of the first decoded image.
[0221] In some other embodiments, the decoding end can reconstruct the first decoded image. That is, based on the bitstream, obtain the reconstruction motion information of the first decoded image, and then based on the global reference tensor corresponding to the first decoded image and the reconstruction motion information of the first decoded image, determine the target prediction information of the first decoded image; based on the bitstream, obtain the reconstruction residual information of the first decoded image; based on the target prediction information of the first decoded image and the reconstruction residual information of the first decoded image, reconstruct the first decoded image to obtain the reconstructed image of the first decoded image.
[0222] The implementation manner of determining the target prediction information of the first decoded image based on the global reference tensor corresponding to the first decoded image and the reconstruction motion information of the first decoded image is similar to the implementation manner of determining the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstruction motion information of the current image in the above encoding method; the implementation manner of obtaining the reconstructed image of the first decoded image based on the target prediction information of the first decoded image and the reconstruction residual information of the first decoded image is similar to the implementation manner of obtaining the reconstructed image of the current image based on the target prediction information of the current image and the reconstruction residual information of the current image in the above encoding method. For detailed content, please refer to the relevant content above and will not be elaborated here.
[0223] If the reconstruction information of the first decoded image includes the reconstructed first decoded image, the decoding end can reconstruct the first decoded image to obtain the reconstructed first decoded image, and based on the reconstructed first decoded image, determine the image features of the reconstructed first decoded image.
[0224] In some embodiments, the process of updating the global reference tensor corresponding to the current image based on the reconstruction information of the first decoded image includes: inputting the reconstruction information of the first decoded image and the global reference tensor corresponding to the first decoded image into the global reference tensor update network to obtain the global reference tensor corresponding to the current image output by the global reference tensor update network.
[0225] The implementation manner of determining the image features of the reconstructed first decoded image based on the reconstructed first decoded image is similar to the implementation manner of determining the image features of the reconstructed first encoded image based on the reconstructed first encoded image in the above encoding method; the implementation manner of inputting the reconstruction information of the first decoded image and the global reference tensor corresponding to the first decoded image into the global reference tensor update network to obtain the global reference tensor corresponding to the current image output by the global reference tensor update network is similar to the implementation manner of inputting the reconstruction information of the first encoded image and the global reference tensor corresponding to the first encoded image into the global reference tensor update network to obtain the global reference tensor corresponding to the current image output by the global reference tensor update network in the above encoding method. For detailed content, please refer to the corresponding content in the above text and will not be elaborated here.
[0226] Step 803: Determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image.
[0227] Next, the implementation manner of determining the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image will be introduced in detail through (1)-(2).
[0228] (1) Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image;
[0229] In some embodiments, the decoding end determines the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image according to relevant prediction techniques such as optical flow mapping (warping) and deformable convolution. The embodiments of the present application do not limit this.
[0230] In some other embodiments, the decoding end can determine the reconstructed motion information of the current image based on the reconstructed motion information of the current image, determine the global motion information corresponding to the current image based on the reconstructed motion information of the current image, where the global motion information is used to describe the position offset of the current image relative to the global reference tensor corresponding to the current image, and determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
[0231] In some embodiments, the decoding end can update the global motion information corresponding to the first decoded image based on the reconstructed motion information of the current image to obtain the global motion information corresponding to the current image.
[0232] The implementation manner of updating the global motion information corresponding to the first decoded image based on the reconstructed motion information of the current image to obtain the global motion information corresponding to the current image is similar to the implementation manner in the above encoding method of updating the global motion information corresponding to the first encoded image based on the reconstructed motion information of the current image to obtain the global motion information corresponding to the current image. For detailed content, please refer to the corresponding content in the above text and will not be elaborated here.
[0233] (2) Determine the target prediction information of the current image based on the global prediction information corresponding to the current image.
[0234] In some embodiments, the decoding end can directly determine the global prediction information corresponding to the current image as the target prediction information of the current image.
[0235] In some other embodiments, the decoding end determines the local prediction information corresponding to the current image based on the reconstructed motion information of the current image and the reconstructed information of the first decoded image, and determines the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image. For the detailed implementation manner, please refer to the corresponding content in the above encoding method and will not be elaborated here.
[0236] Step 804: Obtain the reconstructed residual information of the current image based on the bitstream, and obtain the reconstructed image of the current image based on the target prediction information of the current image and the reconstructed residual information of the current image.
[0237] In some embodiments, the decoding end can parse the bitstream to obtain the reconstructed residual information of the current image.
[0238] In some embodiments, if the reconstruction information of the first decoded image includes the image features of the reconstructed first decoded image, in this case, the decoding end determines the image features of the reconstructed image of the current image according to a related algorithm based on the target prediction information of the current image and the reconstruction residual information of the current image; and reconstructs the current image according to a related algorithm based on the image features of the reconstructed image of the current image to obtain the reconstructed image of the current image.
[0239] Optionally, the decoding end may input the target prediction information of the current image and the reconstruction residual information of the current image into a reconstruction network to obtain the image features of the reconstructed image of the current image output by the reconstruction network.
[0240] Exemplarily, Figure 9 is a schematic structural diagram of a reconstruction network provided by an embodiment of the present application. The reconstruction network includes a deformable convolutional layer (dconv) and three consecutive ResBlock sub-networks with the same configuration. Among them, [c in , h in , w in indicates that the number of channels of the input tensor of the reconstruction network is c in , the size of the input tensor of the reconstruction network in the x-axis direction is h in , and the size of the input tensor of the reconstruction network in the y-axis direction is w in ; dconv(ks, c in , c out , 2) indicates that the convolutional kernel size of the deformable convolutional layer is ks, the number of channels of the input tensor of the deformable convolutional layer is c in , the number of channels of the output tensor of the deformable convolutional layer is c out , and the step size of the convolutional kernel movement is 2. ResBlock(c in ) indicates that the number of channels of the output tensor of the ResBlock sub-network is c in , and the structure of the ResBlock sub-network is as Figure 5 shown. [c out , h in ×2, w in ×2] indicates that the number of channels of the output tensor of the reconstruction network is c out , the size of the output tensor of the reconstruction network in the x-axis direction is h in ×2, and the size of the output tensor of the reconstruction network in the y-axis direction is w in ×2.
[0241] In some other embodiments, if the reconstruction information of the first decoded image includes the reconstructed first decoded image, in this case, the decoding end can reconstruct the current image according to the relevant algorithm based on the target prediction information of the current image and the reconstruction residual information of the current image to obtain the reconstructed image of the current image.
[0242] During the decoding process, accurate target prediction information can be obtained through the reconstruction motion information of the current image and the global reference tensor corresponding to the current image. In this way, it can be ensured that the image quality of the reconstructed image of the current image obtained subsequently based on the target prediction information and the reconstruction residual information corresponding to the current image is better, thereby greatly improving the performance of encoding and decoding.
[0243] Next, Figures 10 - 12 the encoding and decoding method provided by the embodiments of the present application will be introduced in detail again. In Figures 10 - 12 it is assumed that the current image is x t , and both the first encoded image reconstructed by the encoding end and the first decoded image reconstructed by the decoding end are The global reference tensors corresponding to the first encoded image of the encoding end and the first decoded image of the decoding end are both The global motion information corresponding to the first encoded image of the encoding end and the global motion information corresponding to the first decoded image of the decoding end are
[0244] Please refer to Figure 10 , Figure 10 which is a flowchart of an encoding and decoding method provided by the embodiments of the present application. The reconstruction information of the first encoded image reconstructed by the encoding end includes the reconstructed first encoded image The reconstruction information of the first decoded image reconstructed by the decoding end includes the reconstructed first decoded image
[0245] During the encoding process:
[0246] Step 1: Input the reconstructed first encoded image and the global reference tensor corresponding to the first encoded image into the global reference tensor update network to obtain the global reference tensor ref corresponding to the current image output by the global reference tensor update network t global .
[0247] Step 2: Determine the motion information of the current image based on the current image x t and the reconstructed first encoded image and encode the motion information of the current image into the code stream to obtain the motion information code stream of the current image Furthermore, for this code stream Parse to obtain the reconstructed motion information of the current image
[0248] Step 3: Input the reconstructed motion information of the current image and the global motion information corresponding to the first encoded image into the global motion information update network to obtain the global motion information corresponding to the current image output by the global motion information update network
[0249] Step 4: Based on the global reference tensor ref corresponding to the current image t global and the global motion information corresponding to the current image Determine the global prediction information corresponding to the current image, and determine the global prediction information corresponding to the current image as the target prediction information of the current image
[0250] Step 5: Based on the target prediction information of the current image and the current image x t Determine the residual information of the current image, and encode the residual information of the current image into the bitstream to obtain the residual information bitstream of the current image
[0251] During the decoding process:
[0252] Step (1): Input the reconstructed first decoded image and the global reference tensor corresponding to the first decoded image into the global reference tensor update network to obtain the global reference tensor ref corresponding to the current image output by the global reference tensor update network t global .
[0253] Step (2): Parse the motion information bitstream of the current image to obtain the reconstructed motion information of the current image
[0254] Step (3): Input the reconstructed motion information of the current image and the global motion information corresponding to the first decoded image into the global motion information update network to obtain the global motion information corresponding to the current image output by the global motion information update network
[0255] Step (4): Based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image Determine the global prediction information corresponding to the current image, and determine the global prediction information corresponding to the current image as the target prediction information of the current image.
[0256] Step (5): Parse the residual information bitstream of the current image to obtain the reconstructed residual information of the current image
[0257] Step (6): Reconstruct the current image based on the target prediction information of the current image and the reconstructed residual information of the current image to obtain the reconstructed image of the current image
[0258] Please refer to Figure 11 , Figure 11 , which is a flowchart of another encoding and decoding method provided by an embodiment of the present application. The reconstruction information of the first encoded image reconstructed by the encoding end includes the image features of the reconstructed first encoded image , and the reconstruction information of the first decoded image reconstructed by the decoding end includes the image features of the reconstructed first decoded image .
[0259] During the encoding process:
[0260] Step 1: Input the reconstructed first encoded image into the image feature extraction network to obtain the image features of the reconstructed first encoded image output by the image feature extraction network , and input the current image x t into the image feature extraction network to obtain the image features of the current image x t output by the image feature extraction network.
[0261] Step 2: Input the image features of the reconstructed first encoded image and the global reference tensor corresponding to the first encoded image into the global reference tensor update network to obtain the global reference tensor ref corresponding to the current image output by the global reference tensor update network t global .
[0262] Step 3: Determine the motion information of the current image based on the image features of the current image x t and the image features of the reconstructed first encoded image , and encode the motion information of the current image into the bitstream to obtain the motion information bitstream of the current image Furthermore, parse the bitstream to obtain the reconstructed motion information of the current image
[0263] Step 4: Input the reconstruction motion information of the current image and the global motion information corresponding to the first encoded image into the global motion information update network to obtain the global motion information corresponding to the current image output by the global motion information update network
[0264] Step 5: Determine the global prediction information corresponding to the current image based on the global reference tensor ref corresponding to the current image t global and the global motion information corresponding to the current image Step 6: Determine the local prediction information corresponding to the current image based on the reconstruction motion information of the current image
[0265] Step 7: Fuse the local prediction information corresponding to the current image and the global prediction information corresponding to the current image to obtain the target prediction information of the current image and the image features of the reconstructed first encoded image Step 8: Determine the residual information of the current image based on the target prediction information of the current image and the image features of the current image x
[0266] Step 9: Encode the residual information of the current image into the bitstream to obtain the residual information bitstream of the current image
[0267] During the decoding process: t Step (1): Input the reconstructed first decoded image
[0268] into the image feature extraction network to obtain the image features of the reconstructed first decoded image output by the image feature extraction network
[0269] Step (2): Input the image features of the reconstructed first decoded image and the global reference tensor corresponding to the first decoded image into the global reference tensor update network to obtain the global reference tensor ref corresponding to the current image output by the global reference tensor update network
[0270] Step (3): Parse the motion information bitstream of the current image to obtain the reconstruction motion information of the current image Step (4): Input the reconstruction motion information of the current image t global .
[0271] Step (5): Input the reconstruction motion information of the current image into the global motion information update network to obtain the global motion information corresponding to the current image output by the global motion information update network
[0272] Step (6): Input the reconstruction motion information of the current image Global motion information corresponding to the first decoded image Input the global motion information into the global motion information update network to obtain the global motion information corresponding to the current image output by the global motion information update network
[0273] Step (5): Based on the global reference tensor ref corresponding to the current image t global and the global motion information corresponding to the current image Determine the global prediction information corresponding to the current image.
[0274] Step (6): Based on the reconstructed motion information of the current image and the image features of the reconstructed first decoded image Determine the local prediction information corresponding to the current image.
[0275] Step (7): Fuse the local prediction information corresponding to the current image and the global prediction information corresponding to the current image to obtain the target prediction information of the current image.
[0276] Step (8): Parse the bitstream of the residual information of the current image to obtain the reconstructed residual information of the current image
[0277] Step (9): Based on the target prediction information of the current image and the reconstructed residual information of the current image Reconstruct the image features of the current image and then, based on the image features of the reconstructed current image, reconstruct the current image to obtain the reconstructed image of the current image
[0278] Please refer to Figure 12 , Figure 12 which is the test result of the performance test of the encoding and decoding method provided in the embodiments of the present application using three standard test sequence sets, namely class B, class C, and class D. Among them, class B contains 5 1080p videos, class C contains 4 720p videos, and class D contains 4 videos with a size of 416x240. It can be easily seen from Figure 11 that the RD-Curve of the embodiments of the present application on the test set has an average improvement in compression efficiency of 14%, 13%, and 15% respectively in the three standard test sequence sets compared with the related technologies (that is, it can reduce the transmission bandwidth by 14%, 13%, and 15% compared with the related technologies). Figure 13
[0279] Please refer to Figure 13 , Figure 13 Flowchart of another encoding and decoding method provided by an embodiment of this application. The reconstruction information of the first encoded image reconstructed by the encoding end includes the reconstructed first encoded image The image features of the first decoded image reconstructed by the decoding end include the reconstructed first decoded image The image features of.
[0280] During the encoding process:
[0281] Step 1: Input the reconstructed first encoded image Into the image feature extraction network to obtain the image features of the reconstructed first encoded image output by the image feature extraction network Input the current image x t Into the image feature extraction network to obtain the image features of the current image x output by the image feature extraction network t The image features of.
[0282] Step 2: Input the image features of the reconstructed first encoded image And the global reference tensor corresponding to the first encoded image Into the global reference tensor update network to obtain the global reference tensor corresponding to the current image output by the global reference tensor update network
[0283] Step 3: Based on the image features of the current image x t And the image features of the reconstructed first encoded image Determine the motion information of the current image, and encode the motion information of the current image into the bitstream to obtain the motion information bitstream of the current image Furthermore, parse the bitstream To obtain the reconstructed motion information of the current image
[0284] Step 4: Based on the global reference tensor corresponding to the current image And the reconstructed motion information of the current image Determine the global prediction information corresponding to the current image.
[0285] Step 5: Based on the reconstructed motion information of the current image And the image features of the reconstructed first encoded image Determine the local prediction information corresponding to the current image.
[0286] Step 6: Fuse the local prediction information corresponding to the current image and the global prediction information corresponding to the current image to obtain the target prediction information of the current image.
[0287] Step 7: Based on the target prediction information of the current image and the current image xt Based on the image features of the current image, determine the residual information of the current image, and encode the residual information of the current image into a bitstream to obtain the residual information bitstream of the current image
[0288] During the decoding process:
[0289] Step (1): Input the reconstructed first decoded image into the image feature extraction network to obtain the image features of the reconstructed first decoded image output by the image feature extraction network of the current image
[0290] Step (2): Input the image features of the reconstructed first decoded image and the global reference tensor corresponding to the first decoded image into the global reference tensor update network to obtain the global reference tensor ref corresponding to the current image output by the global reference tensor update network t global .
[0291] Step (3): Parse the motion information bitstream of the current image to obtain the reconstructed motion information of the current image
[0292] Step (4): Based on the global reference tensor ref t global corresponding to the current image and the reconstructed motion information of the current image determine the global prediction information corresponding to the current image
[0293] Step (5): Based on the reconstructed motion information of the current image and the image features of the reconstructed first decoded image determine the local prediction information corresponding to the current image
[0294] Step (6): Fuse the local prediction information corresponding to the current image and the global prediction information corresponding to the current image to obtain the target prediction information of the current image
[0295] Step (7): Parse the residual information bitstream of the current image to obtain the reconstructed residual information of the current image
[0296] Step (8): Based on the target prediction information of the current image and the reconstructed residual information of the current image reconstruct the image features of the current image of the current image, and then based on the image features of the reconstructed current image, reconstruct the current image to obtain the reconstructed image of the current image
[0297] Figure 14 is a schematic diagram of the structure of a coding device provided in an embodiment of the present application. The coding device can be implemented by software, hardware or a combination of both to become part or all of the coding end, and the coding device can be Figure 1 The encoder 100 in FIG. Figure 14 The device includes: a first determination module 1401, a second determination module 1402, a third determination module 1403 and an encoding module 1404.
[0298] The first determination module 1401 is used to determine the global reference tensor corresponding to the current image and the motion information of the current image, and the global reference tensor is determined based on multiple encoded images before the current image. The detailed implementation process refers to the corresponding content in the above embodiments, which will not be repeated here.
[0299] The second determination module 1402 is used to determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image. The detailed implementation process refers to the corresponding content in the above embodiments, which will not be repeated here.
[0300] The third determination module 1403 is used to determine the residual information of the current image based on the target prediction information of the current image and the current image. The detailed implementation process refers to the corresponding content of each of the above embodiments, which will not be repeated here.
[0301] The encoding module 1404 is used to encode the residual information and the motion information of the current image into the bitstream. The detailed implementation process refers to the corresponding content in the above embodiments, which will not be repeated here.
[0302] Optionally, the first determining module 1401 is specifically configured to:
[0303] Determining reconstruction information of a first coded image, the first coded image being a coded image adjacent to the current image;
[0304] Based on the reconstruction information of the first encoded image, the global reference tensor corresponding to the first encoded image is updated to obtain the global reference tensor corresponding to the current image.
[0305] Optionally, the global reference tensor corresponding to the current image is determined based on all encoded images before the current image.
[0306] Optionally, the global reference tensor corresponding to the image of the Nth frame in the video bitstream is obtained based on the global reference tensor corresponding to the encoded image of the (N - 1)th frame and the reconstruction information of the encoded image of the (N - 1)th frame, and the global reference tensor corresponding to the image of the second frame is the reconstruction information of the encoded image of the first frame.
[0307] Optionally, the second determination module 1402 is specifically configured to:
[0308] Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the motion information of the current image;
[0309] Determine the target prediction information corresponding to the current image based on the global prediction information corresponding to the current image.
[0310] Optionally, the second determination module 1402 is specifically configured to:
[0311] Determine the reconstruction motion information of the current image based on the motion information of the current image;
[0312] Determine the global motion information corresponding to the current image based on the reconstruction motion information of the current image, where the global motion information includes a global motion vector, and the global motion vector indicates the offset of the target block in the current image from that in the global reference tensor corresponding to the current image;
[0313] Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
[0314] Optionally, the second determination module 1402 is specifically configured to:
[0315] Update the global motion information corresponding to the first encoded image based on the reconstruction motion information of the current image to obtain the global motion information corresponding to the current image.
[0316] Optionally, the second determination module 1402 is specifically configured to:
[0317] Determine the local prediction information corresponding to the current image based on the reconstruction motion information of the current image and the reconstruction information of the first encoded image;
[0318] Determine the target prediction information corresponding to the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image.
[0319] Optionally, the reconstruction information of the first encoded image includes the reconstructed first encoded image or the image features of the reconstructed first encoded image.
[0320] Optionally, the third determination module 1403 is specifically configured to:
[0321] Determine the residual information of the current image based on the target prediction information of the current image and the image features of the current image.
[0322] In the embodiments of the present application, the target prediction information of the current image is determined by the global reference tensor corresponding to the current image and the motion information of the current image. Since the global reference tensor is determined based on multiple previously encoded images before the current image, that is, the embodiments of the present application can generate a global reference tensor according to multiple previously encoded images before the current image, so that the global reference tensor can represent the image information available for reference in the current image in multiple previously encoded images before the current image, and has cross-temporal globality. In this way, it can be ensured that the target prediction information corresponding to the current image obtained based on the global reference tensor corresponding to the current image and the motion information of the current image is more accurate, so that the data bit amount of the residual information corresponding to the current image determined subsequently is smaller, thereby effectively improving the encoding and decoding efficiency. In addition, by generating a global reference tensor from multiple previously encoded images before the current image, it can be ensured that the encoding end has more information available for reference by the current image during encoding without caching multiple previously encoded images, thereby greatly reducing the caching pressure of the encoding end while improving the encoding and decoding efficiency.
[0323] Since the decoding end also needs to determine the global motion information corresponding to the current image in subsequent steps, but the decoding end cannot obtain the motion information of the original current image during decoding, and can only obtain the reconstructed motion information of the current image. That is to say, there may be a certain deviation between the motion information reconstructed by the decoding end and the motion information corresponding to the original current image. Therefore, the encoding end determines the global motion information corresponding to the current image through the reconstructed motion information of the current image, so as to ensure that the global motion information corresponding to the current image subsequently used by the encoding end and the decoding end is the same, so that the target prediction information of the current image obtained by the encoding end and the decoding end is more accurate, thereby effectively improving the encoding and decoding efficiency and performance. The encoding end directly determines the global prediction information corresponding to the current image based on the motion information of the current image and the global reference tensor corresponding to the current image, which can effectively reduce the computational amount of image compression and improve the encoding and decoding efficiency. Since the motion information of the current image can represent to a certain extent the motion of the target block in the current image relative to the encoded image, the encoding end can directly determine the global prediction information corresponding to the current image based on the motion information of the current image and the global reference tensor corresponding to the current image, which can effectively reduce the computational amount of image compression and improve the encoding and decoding efficiency while ensuring the accuracy of the target prediction information determined subsequently.
[0324] It should be noted that when the encoding device provided in the above embodiments performs encoding, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the encoding device provided in the above embodiments and the encoding method embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0325] Figure 15 FIG. 4 is a schematic structural diagram of a decoding device provided by an embodiment of the present application. The decoding device can be implemented as part or all of the decoding end by software, hardware, or a combination of both. Moreover, the decoding device can be the Figure 1 decoder 200 in Figure 15 . Referring to
[0326] FIG. 5, the device includes: a first parsing module 1501, a first determination module 1502, a second determination module 1503, a second parsing module 1504, and a reconstruction module 1505.
[0327] The first parsing module 1501 is configured to obtain the reconstructed motion information of the current image based on the bitstream. For the detailed implementation process, please refer to the corresponding content in the above embodiments and will not be elaborated here.
[0328] The first determination module 1502 is configured to determine the global reference tensor corresponding to the current image, and the global reference tensor is determined based on multiple decoded images before the current image. For the detailed implementation process, please refer to the corresponding content in the above embodiments and will not be elaborated here.
[0329] The second determination module 1503 is configured to determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image. For the detailed implementation process, please refer to the corresponding content in the above embodiments and will not be elaborated here.
[0330] The second parsing module 1504 is configured to obtain the reconstructed residual information of the current image based on the bitstream. For the detailed implementation process, please refer to the corresponding content in the above embodiments and will not be elaborated here.
[0331] Optionally, the first determination module 1502 is specifically configured to:
[0332] Determine the reconstruction information of the first decoded image, where the first decoded image is a decoded image adjacent to the current image;
[0333] Update the global reference tensor corresponding to the first decoded image based on the reconstruction information of the first decoded image, so as to obtain the global reference tensor corresponding to the current image.
[0334] Optionally, the global reference tensor corresponding to the current image is determined based on all decoded images before the current image.
[0335] Optionally, the global reference tensor corresponding to the image of the Nth frame in the video bitstream is obtained based on the global reference tensor corresponding to the decoded image of the (N - 1)th frame and the reconstruction information of the decoded image of the (N - 1)th frame, and the global reference tensor corresponding to the image of the second frame is the reconstruction information of the decoded image of the first frame.
[0336] Optionally, the second determination module 1503 is specifically configured to:
[0337] Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image;
[0338] Determine the target prediction information corresponding to the current image based on the global prediction information corresponding to the current image.
[0339] Optionally, the second determination module 1503 is specifically configured to:
[0340] Determine the global motion information corresponding to the current image based on the reconstructed motion information of the current image, where the global motion information includes a global motion vector, and the global motion vector indicates the offset of the target block in the current image from that in the global reference tensor corresponding to the current image;
[0341] Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
[0342] Optionally, the second determination module 1503 is specifically configured to:
[0343] Update the global motion information corresponding to the first decoded image based on the reconstructed motion information of the current image, so as to obtain the global motion information corresponding to the current image.
[0344] Optionally, the second determination module 1503 is specifically configured to:
[0345] Determine the local prediction information corresponding to the current image based on the reconstructed motion information of the current image and the reconstruction information corresponding to the first decoded image;
[0346] Determine the target prediction information corresponding to the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image.
[0347] Optionally, the reconstruction information of the first decoded image includes the reconstructed first decoded image or the image features of the reconstructed first decoded image.
[0348] Optionally, the reconstruction module 1505 is specifically configured to:
[0349] Determine the image features of the reconstructed image of the current image based on the target prediction information of the current image and the reconstruction residual information of the current image;
[0350] Determine the reconstructed image of the current image based on the image features of the reconstructed image of the current image.
[0351] During the decoding process, accurate target prediction information can be obtained through the reconstruction motion information of the current image and the global reference tensor corresponding to the current image. In this way, it can be ensured that the image quality of the reconstructed image of the current image obtained subsequently based on the target prediction information and the reconstruction residual information corresponding to the current image is better, thereby greatly improving the performance of encoding and decoding.
[0352] It should be noted that: when the decoding device provided in the above embodiment performs decoding, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the decoding device provided in the above embodiment and the decoding method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0353] The embodiment of the present application also provides an encoding device, where the encoding device includes: a processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the encoding device executes the above encoding method.
[0354] The embodiment of the present application also provides a decoding device, where the decoding device includes: a processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the decoding device executes the above decoding method.
[0355] The embodiment of the present application also provides an encoding and decoding system, where the encoding and decoding system includes the above encoding device and / or the above decoding device.
[0356] The embodiment of the present application also provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program. When the computer program runs on a computer or a processor, the computer or the processor executes the above encoding method or the above decoding method.
[0357] An embodiment of the present application further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the steps of the above encoding method are executed, or the steps of the above decoding method are executed.
[0358] An embodiment of the present application further provides a computer-readable storage medium, which is characterized in that a bitstream obtained according to the above encoding method is stored on the computer-readable storage medium.
[0359] An embodiment of the present application further provides a device for storing a bitstream, including at least one storage medium and a communication interface; the communication interface is used to receive or send a bitstream; the at least one storage medium is used to store the bitstream; the bitstream is encoded by an encoder according to the above image encoding method.
[0360] An embodiment of the present application further provides a method for storing a bitstream, including: receiving a bitstream through a communication interface; storing the bitstream in one or more storage mediums, and the bitstream is encoded by an encoder according to the above image encoding method.
[0361] An embodiment of the present application further provides a system for distributing a bitstream, including at least one storage medium and a video stream device; the at least one storage medium is used to store a bitstream, and the bitstream is encoded by an encoder according to the above image encoding method; the video stream device is used to send the bitstream in the at least one storage medium to the decoder in response to a request from the decoder.
[0362] An embodiment of the present application further provides a method for distributing a bitstream, including: receiving a first request; selecting a bitstream from at least one storage medium in response to the first request; sending the bitstream to a destination device; the at least one storage medium is used to store a bitstream, and the bitstream is encoded by an encoder according to the above image encoding method.
[0363] An embodiment of the present application further provides a system for processing a bitstream, including an image source device, an encoder, one or more storage media, and a destination device; the image source device is configured to provide image data; the encoder is configured to obtain the image data of the image source device through an interface, and encode the image data to obtain one or more bitstreams, where the bitstreams are encoded by the encoder according to the above image encoding method; the encoder is configured to store the one or more bitstreams in one or more storage media; or the encoder is configured to encapsulate the one or more bitstreams to obtain a transport bitstream; the encoder is configured to transmit the transport bitstream to the destination device through a communication link or a communication network; the destination device is configured to de-encapsulate the transport bitstream to obtain the one or more bitstreams; the destination device is configured to decode the one or more bitstreams to obtain decoded data.
[0364] In the above embodiment, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, it may be a non-transitory storage medium.
[0365] It should be understood that the "multiple" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; the "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, for the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily mean different.
[0366] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the current images involved in the embodiments of the present application are all obtained under sufficient authorization.
[0367] The above are the embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A coding method, characterized in that, The method includes: Determining a global reference tensor corresponding to the current image and motion information of the current image, where the global reference tensor is determined based on a plurality of previously encoded images before the current image; Determining target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image; Determining residual information of the current image based on the target prediction information of the current image and the current image; Encoding the residual information and the motion information of the current image into a bitstream.
2. The method according to claim 1, wherein The determining the global reference tensor corresponding to the current image includes: Determining reconstruction information of a first encoded image, where the first encoded image is an encoded image adjacent to the current image; Updating the global reference tensor corresponding to the first encoded image based on the reconstruction information of the first encoded image to obtain the global reference tensor corresponding to the current image.
3. The method according to claim 1 or 2, characterized in that, The global reference tensor corresponding to the current image is determined based on all previously encoded images before the current image.
4. The method according to claim 1, wherein The global reference tensor corresponding to the image of the Nth frame in the video bitstream is obtained based on the global reference tensor corresponding to the encoded image of the (N - 1)th frame and the reconstruction information of the encoded image of the (N - 1)th frame, and the global reference tensor corresponding to the second frame image is the reconstruction information of the encoded image of the first frame.
5. The method according to claim 2, wherein The determining the target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image includes: Determining global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the motion information of the current image; Determining the target prediction information of the current image based on the global prediction information corresponding to the current image.
6. The method according to claim 5, characterized in that, The determining the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the motion information of the current image includes: Determining reconstruction motion information of the current image based on the motion information of the current image; Determining global motion information corresponding to the current image based on the reconstruction motion information of the current image, where the global motion information includes a global motion vector, and the global motion vector indicates the offset of a target block in the current image from that in the global reference tensor corresponding to the current image; Determining global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
7. The method according to claim 6, wherein The determining the global motion information corresponding to the current image based on the reconstruction motion information of the current image includes: Updating the global motion information corresponding to the first encoded image based on the reconstruction motion information of the current image to obtain the global motion information corresponding to the current image.
8. The method according to any one of claims 5 to 7, characterized in that The determining the target prediction information of the current image based on the global prediction information corresponding to the current image includes: Determining local prediction information corresponding to the current image based on the reconstruction motion information of the current image and the reconstruction information of the first encoded image; Determine the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image.
9. The method according to any one of claims 2-8, characterized in that, The reconstruction information of the first encoded image includes the reconstructed first encoded image or the image features of the reconstructed first encoded image.
10. The method according to claim 1, wherein The determining the residual information of the current image based on the target prediction information of the current image and the current image includes: Determine the residual information of the current image based on the target prediction information of the current image and the image features of the current image.
11. A decoding method, characterized in that, The method includes: Obtain the reconstructed motion information of the current image based on the bitstream; Determine the global reference tensor corresponding to the current image, where the global reference tensor is determined based on a plurality of decoded images before the current image; Determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image; Obtain the reconstructed residual information of the current image based on the bitstream; Obtain the reconstructed image of the current image based on the target prediction information of the current image and the reconstructed residual information of the current image.
12. The method according to claim 11, wherein The determining the global reference tensor corresponding to the current image includes: Determine the reconstruction information of the first decoded image, where the first decoded image is a decoded image adjacent to the current image; Update the global reference tensor corresponding to the first decoded image based on the reconstruction information of the first decoded image to obtain the global reference tensor corresponding to the current image.
13. The method according to claim 11, wherein The global reference tensor corresponding to the current image is determined based on all decoded images before the current image.
14. The method according to claim 11, characterized in that, The global reference tensor corresponding to the image of the Nth frame in the video bitstream is obtained based on the global reference tensor corresponding to the decoded image of the (N - 1)th frame and the reconstruction information of the decoded image of the (N - 1)th frame, and the global reference tensor corresponding to the second frame image is the reconstruction information of the decoded image of the first frame.
15. The method according to claim 12, characterized in that, The determining the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image includes: Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image; Determine the target prediction information of the current image based on the global prediction information corresponding to the current image.
16. The method according to claim 15, wherein The determining the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image includes: Determine the global motion information corresponding to the current image based on the reconstructed motion information of the current image, where the global motion information includes a global motion vector, and the global motion vector indicates the offset of the target block in the current image from that in the global reference tensor corresponding to the current image; Determine the global prediction information corresponding to the current image based on the global reference tensor corresponding to the current image and the global motion information corresponding to the current image.
17. The method according to claim 16, wherein Determining the global motion information corresponding to the current image based on the reconstructed motion information of the current image includes: Updating the global motion information corresponding to the first decoded image based on the reconstructed motion information of the current image to obtain the global motion information corresponding to the current image.
18. The method according to any one of claims 15 to 17, characterized in that, Determining the target prediction information of the current image based on the global prediction information corresponding to the current image includes: Determining the local prediction information corresponding to the current image based on the reconstructed motion information of the current image and the reconstructed information corresponding to the first decoded image; Determining the target prediction information of the current image based on the local prediction information corresponding to the current image and the global prediction information corresponding to the current image.
19. The method according to any one of claims 12 - 18, characterized in that, The reconstructed information of the first decoded image includes the reconstructed first decoded image or the image features of the reconstructed first decoded image.
20. The method according to claim 11, wherein Obtaining the reconstructed image of the current image based on the target prediction information of the current image and the reconstructed residual information of the current image includes: Determining the image features of the reconstructed image of the current image based on the target prediction information of the current image and the reconstructed residual information of the current image; Determining the reconstructed image of the current image based on the image features of the reconstructed image of the current image.
21. A coding device, characterized in that, The apparatus includes: A first determination module, configured to determine the global reference tensor corresponding to the current image and the motion information of the current image, where the global reference tensor is determined based on multiple encoded images before the current image; A second determination module, configured to determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the motion information of the current image; A third determination module, configured to determine the residual information of the current image based on the target prediction information of the current image and the current image; An encoding module, configured to encode the residual information and the motion information of the current image into a bitstream.
22. A decoding device, characterized in that, The apparatus includes: A first parsing module, configured to obtain the reconstructed motion information of the current image based on the bitstream; A first determination module, configured to determine the global reference tensor corresponding to the current image, where the global reference tensor is determined based on multiple decoded images before the current image; A second determination module, configured to determine the target prediction information of the current image based on the global reference tensor corresponding to the current image and the reconstructed motion information of the current image; A second parsing module, configured to obtain the reconstructed residual information of the current image based on the bitstream; A reconstruction module, configured to obtain the reconstructed image of the current image based on the target prediction information of the current image and the reconstructed residual information of the current image.
23. A coding device, characterized in that, Includes: A processor, where the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the encoding device is caused to execute the method according to any one of claims 1 to 10.
24. A decoding device, characterized in that, Includes: A processor, the processor being coupled to a memory for storing programs or instructions, which when executed by the processor cause the decoding device to perform the method according to any one of claims 11 to 20.
25. A coding and decoding system, characterized in that, The encoding and decoding system includes an encoding device according to claim 23, and / or a decoding device according to claim 24.
26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program which, when running on a computer or a processor, causes the computer or the processor to perform the method according to any one of claims 1 to 10, or to perform the method according to any one of claims 11 to 20.
27. A computer program product, characterized in that, The computer program product contains computer instructions which, when executed by a computer or a processor, cause the steps of the method according to any one of claims 1 to 10 to be executed, or the steps of the method according to any one of claims 11 to 20 to be executed.
28. A computer-readable storage medium, characterized in that, A bitstream obtained by performing the method according to any one of claims 1-10 is stored on the computer-readable storage medium and is executed by one or more processors.
29. A device for storing a bitstream, characterized in that, Including at least one storage medium and a communication interface; The communication interface is used to receive or transmit a bitstream; The at least one storage medium is used to store the bitstream; The bitstream is encoded by an encoder according to any one of the encoding methods of claims 1 to 10.
30. A method for storing a bitstream, characterized in that, Including: Receiving a bitstream through a communication interface; Storing the bitstream into one or more storage media, the bitstream being an encoder Encoded according to any one of the encoding methods of claims 1 to 10.
31. A system for distributing a bitstream, characterized in that, Including at least one storage medium and a video stream device; The at least one storage medium is used to store a bitstream, the bitstream being an encoder Encoded according to any one of the encoding methods of claims 1 to 10; The video stream device is configured to, in response to a request from a decoder, cause the target bitstream in the at least one storage medium to be sent to the decoder.
32. A method for distributing a bitstream, characterized in that, Including: Receiving a first request; In response to the first request, selecting a target bitstream from at least one storage medium; Sending the target bitstream to a destination device; The at least one storage medium is used to store a bitstream, the bitstream being an encoder Encoded according to any one of the encoding methods of claims 1 to 10.
33. A system for processing a bitstream, characterized in that, Including an image source device, an encoder device, one or more storage media and a destination device; The image source device is used to provide image data; The encoder device is configured to obtain the image data of the image source device through an interface and encode the image data to obtain one or more bitstreams, the bitstreams being encoded by the encoder according to any one of the encoding methods of claims 1 to 10; The encoder device is configured to store the one or more bitstreams into one or more storage media; or, The encoder device is configured to encapsulate the one or more bitstreams to obtain a transport bitstream; The encoder device is configured to transmit the transport bitstream to the destination device through a communication link or a communication network; The target device is used to unpack the transmission bitstream to obtain the one or more bitstreams; The target device is used to decode the one or more bitstreams to obtain decoded data.