Video decoding method and apparatus, and video encoding method and apparatus

By acquiring the background content and camera information of the video sequence, and using the camera information to determine the encoding mode of the current block, only the identification information is encoded. This solves the problem of low encoding efficiency in stable background areas, achieves more efficient video encoding and decoding, and reduces storage and transmission costs.

WO2025246524A1PCT designated stage Publication Date: 2025-12-04HISENSE VISUAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/081095
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2025-03-06
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing video coding technologies fail to effectively utilize camera information when processing high-resolution video and cloud gaming content, resulting in low coding efficiency for stable background areas and increased storage and transmission costs.

Method used

By acquiring the background content and camera information of the video sequence, the camera information is used to determine whether the current block is a background block. When it is determined to be a background block, the target coding mode is adopted to encode only the identification information and not the residual data, thus generating coded data.

Benefits of technology

It reduces the amount of encoded data, improves the efficiency of video encoding and decoding, reduces storage and transmission costs, and enhances the quality of the video experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081095_04122025_PF_FP_ABST
    Figure CN2025081095_04122025_PF_FP_ABST
Patent Text Reader

Abstract

Some embodiments of the present disclosure relate to the technical field of video encoding and decoding, and provide a video decoding method and apparatus, and a video encoding method and apparatus. The decoding method comprises: acquiring encoding data of a current block; determining an encoding mode of the current block on the basis of the encoding data of the current block; when the encoding mode of the current block is a target encoding mode, acquiring current reconstructed background content and camera information of a current frame, the current reconstructed background content being background content constructed on the basis of a reconstructed block of a decoded background block in a current video sequence, the current frame being a video frame in which the current block is, and the current video sequence being a video sequence in which the current frame is located; and reconstructing the current block on the basis of the camera information of the current frame and the current reconstructed background content, so as to acquire a reconstructed block of the current block. Some embodiments of the present disclosure are used to improve the efficiency of video encoding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Video decoding method, video encoding method and device

[0001] The present disclosure claims priority to a Chinese patent application with the application number 202410693335X, the title of which is "A video decoding method, a video encoding method and device", filed on May 30, 2024, with the Chinese Patent Office, and a Chinese patent application with the application number 2024106923273, the title of which is "A video decoding method, a video encoding method and device", filed on May 30, 2024, with the Chinese Patent Office, the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD

[0002] Some embodiments of the present disclosure relate to the field of video coding technology. More specifically, the present disclosure relates to a video decoding method, a video encoding method and device. BACKGROUND

[0003] For the field of digital video, the development of video coding technology is of great importance. The main purpose of video coding technology is to remove redundant information in video by using compression algorithms, so as to reduce the size of video files, facilitate transmission on the network and save space on storage devices.

[0004] With the gradual popularity of high-resolution video formats such as 4K and 8K, and the rapid development of the virtual reality and gaming industries, video content is becoming more and more complex, and its data volume is also increasing dramatically. Among these video contents, there are often a large number of background parts with little change, such as the vast sky in game videos, the stable background scene in monitoring videos, and the background blackboard in remote education scenes. Although these background contents have little change in vision, the encoding method of calculating and referencing the residual of the block still produces a lot of encoding data, so the efficiency of video coding needs to be further improved. SUMMARY

[0005] Exemplary embodiments of the present disclosure provide a video decoding method, a video encoding method and device for improving the efficiency of video coding.

[0006] Some embodiments of the present disclosure provide technical solutions as follows:

[0007] In a first aspect, some embodiments of the present disclosure provide a video decoding method, comprising:

[0008] obtaining the encoding data of the current block;

[0009] determining the encoding mode of the current block according to the encoding data of the current block;

[0010] In a case where the coding mode of the current block is the target coding mode, obtain the current reconstructed background content and camera information of the current frame; the current reconstructed background content is background content constructed according to reconstructed blocks of background blocks that have completed decoding in a current video sequence, the current frame is a video frame to which the current block belongs, and the current video sequence is a video sequence to which the current frame belongs;

[0011] Reconstruct the current block according to the camera information of the current frame and the current reconstructed background content, to obtain a reconstructed block of the current block.

[0012] In a second aspect, some embodiments of the present disclosure provide a video coding method, including:

[0013] Obtain camera information of each video frame of a current video sequence;

[0014] Encode the camera information of each video frame of the current video sequence, to obtain camera information coding data;

[0015] Determine whether the current block is a background block; the current block is any image block obtained by block division on a current frame of a current video sequence;

[0016] In a case where the current block is a background block, determine whether the current reconstructed background content includes background content corresponding to the current block; the current reconstructed background content is background content constructed according to reconstructed blocks of background blocks that have completed encoding in the current video sequence;

[0017] In a case where the current reconstructed background content includes the background content corresponding to the current block, select a coding mode of the current block from all coding modes in a preset coding mode set, the preset coding mode set including a target coding mode;

[0018] In a case where the selected coding mode of the current block is the target coding mode, encode identification information of the target coding mode, to obtain coding data of the current block;

[0019] Generate coding data of the current video sequence according to the camera information coding data and coding data of each image block of each video frame of the current video sequence.

[0020] In a third aspect, some embodiments of the present disclosure provide an image decoding apparatus, including:

[0021] A memory configured to store a computer program;

[0022] A processor configured to, when the computer program is invoked, cause the video coding apparatus to implement the image decoding method of the first aspect.

[0023] In a fourth aspect, some embodiments of the present disclosure provide a video coding apparatus, including:

[0024] a memory configured to store the computer program;

[0025] a processor configured to, when invoking the computer program, cause the video decoding apparatus to implement the image encoding method of the second aspect.

[0026] In a fifth aspect, some embodiments of the present disclosure provide a computer readable storage medium, having stored thereon a computer program, which, when executed by a computing device, causes the computing device to implement the method of the first aspect or the second aspect.

[0027] In a sixth aspect, some embodiments of the present disclosure provide a computer program product, which, when running on a computer, causes the computer to implement the method of the first aspect or the second aspect.

[0028] According to the above technical solution, the video decoding method provided by some embodiments of the present disclosure first determines the encoding mode of the current block according to the encoding data of the current block after obtaining the encoding data of the current block, and in the case that the encoding mode of the current block is the target encoding mode, the reconstructed block of the current block is obtained by reconstructing the current block according to the camera information of the current frame and the reconstructed background content of the current frame. The current reconstructed background content is the background content constructed according to the reconstructed blocks of the background blocks that have been decoded in the current video sequence. The video decoding method provided by some embodiments of the present disclosure can reconstruct the current block according to the camera information of the current frame and the current reconstructed background content, so for the image block whose encoding mode is the target encoding mode, there is no need to add the residual data of the image block in the encoding data. Therefore, some embodiments of the present disclosure can reduce the data amount of the encoding data and improve the efficiency of video coding and decoding. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the embodiments of some embodiments of the present disclosure or the implementation manners in the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art according to these drawings.

[0030] FIG. 1 shows a structural block diagram of a hybrid coding framework according to some embodiments of the present disclosure;

[0031] FIG. 2 shows a schematic diagram of a frame difference image according to some embodiments of the present disclosure;

[0032] FIG. 3 shows a step flowchart of a video encoding method according to some embodiments of the present disclosure;

[0033] FIG. 4 shows a flow chart of steps of a video encoding method in some embodiments of the present disclosure;

[0034] FIG. 5 shows a flow chart of reconstructing a current block in some embodiments of the present disclosure;

[0035] FIG. 6 shows a schematic diagram of calculating pixel bit depth in some embodiments of the present disclosure;

[0036] FIG. 7 shows a schematic diagram of mapping points of pixels in background content in some embodiments of the present disclosure;

[0037] FIG. 8 shows a schematic diagram of positions of mapping points in some embodiments of the present disclosure;

[0038] FIG. 9 shows a flow chart of steps of a video encoding method in some embodiments of the present disclosure;

[0039] FIG. 10 shows a flow chart of steps of a video decoding method in some embodiments of the present disclosure;

[0040] FIG. 11 shows a flow chart of steps of a video encoding method in some embodiments of the present disclosure;

[0041] FIG. 12 shows a flow chart of steps of a video encoding method in some embodiments of the present disclosure;

[0042] FIG. 13 shows a flow chart of steps of a video encoding method in some embodiments of the present disclosure;

[0043] FIG. 14 shows a schematic diagram of a mapping area of a current block in background content in some embodiments of the present disclosure;

[0044] FIG. 15 shows a flow chart of steps of a video decoding method in some embodiments of the present disclosure;

[0045] FIG. 16 shows a flow chart of steps of a video encoding method in some embodiments of the present disclosure. DETAILED DESCRIPTION

[0046] For the purpose of making the objects and embodiments of the present disclosure more clear, the following will describe the exemplary embodiments of the present disclosure in detail with reference to the accompanying drawings. Obviously, the described exemplary embodiments are only some of the embodiments of the present disclosure, but not all of the embodiments of the present disclosure.

[0047] It should be noted that the brief description of the terms in the present disclosure is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present disclosure. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.

[0048] The terms "including," "containing," "having," and "including" and any variations thereof are intended to cover a non-exclusive inclusion, such that a product or process that comprises a list of components or steps does not include only those components or steps but can include other components or steps not expressly listed or inherent to such product or process.

[0049] References in the specification to "some implementations", "some embodiments", etc. indicate that the implementation or embodiment described can include a particular feature, structure, or characteristic, but every implementation or embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same implementation or embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an implementation or embodiment, it is submitted that it is within the knowledge of one of ordinary skill in the art to effect such feature, structure, or characteristic in connection with other implementations or embodiments, whether or not explicitly described herein.

[0050] Some embodiments of the present disclosure relate to a hybrid codec framework, and the following is a description of key components of the hybrid codec framework.

[0051] Referring to FIG. 1, in some embodiments, the hybrid codec framework includes an encoding control module 101. The encoding control module 101 is used to manage and regulate the overall encoding process, including the setting of encoding parameters, the selection and adjustment of encoding modes, the monitoring of the encoding process, and the coordination with other related parts, etc.

[0052] Referring to FIG. 1, in some embodiments, the hybrid codec framework further includes a quantization / transform module 102. The quantization / transform module 102 is used to transform and quantize image blocks in the video encoding process. Among them, transform coding refers to the process of converting image data into frequency domain representation through transform technology. Since transform can reduce the redundancy in the spatial domain, transform can make the encoder more effective in representing and compressing image data. The transform algorithm used by transform can be wavelet transform, discrete cosine transform (DCT), integer transform, etc. Quantization is a method for further compressing data after transform coding and prediction coding, mainly by mapping a signal interval to a signal value to reduce the amount of information required to be recorded. Prediction coding and transform coding do not bring distortion to the image, and quantization as a lossy compression technology is the main source of distortion in video coding. In order to reduce the distortion caused by quantization while maintaining good video quality, it is necessary to design reasonable quantization methods and quantization parameters. The quantization methods mainly include uniform quantization, non-uniform quantization and adaptive quantization, etc.

[0053] Referring to FIG. 1, in some embodiments, the hybrid codec framework further comprises an intra prediction module 103. The basic idea of Intra Prediction is to exploit the spatial correlation of neighboring pixels to remove the spatial redundancy of video. In video coding, neighboring pixels refer to the reconstructed pixels of already coded image blocks around the current block. Specifically, Intra Prediction exploits the already reconstructed pixels in the current frame to derive the prediction values for the current block. Therefore, the derivation scheme of the prediction block is one of the key techniques of Intra Prediction.

[0054] Intra Prediction is achieved by dividing a video frame into multiple image blocks. The video coding process based on Intra Prediction includes: first, dividing the current frame block into multiple image blocks, usually pixel blocks or macroblocks. These image blocks are called Prediction Units (PUs). Then, for each Prediction Unit, the encoder predicts the pixel values of each pixel within the unit according to the reconstructed blocks of already coded image blocks around the unit. Intra Prediction is usually based on the local structure of the already coded pixels around the unit. Common Intra Prediction modes include vertical, horizontal, DC (Direct Current), angular, and average, etc. The encoder will choose the prediction mode that minimizes the amount of residual data, where the residual is the difference between the original pixel and the predicted pixel.

[0055] Referring to FIG. 1, in some embodiments, the hybrid codec framework further comprises an inter prediction module 104. The basic idea of Inter Prediction is to exploit the temporal correlation between a sequence of consecutive frames to remove the temporal redundancy of video. When coding the current block, the information of already coded frames is used to predict the current image information. In a sequence of consecutive video frames, the correlation between adjacent image frames is strong, partly due to the continuity of object motion, and partly due to the lack of change in the background. Since the correlation between frames is often greater than the correlation between adjacent pixels within a frame, especially between video frames close in time, using temporal redundancy can more obviously compress the redundant information in the video than using spatial redundancy.

[0056] Inter Prediction is also achieved by dividing a video frame into multiple image blocks. The video coding process based on Inter Prediction includes: first, dividing the current frame into multiple Prediction Units. Then, the pixel values within each Prediction Unit are predicted using the motion relationship between the image region of the current frame and the reference frame (usually the previous frame of the current frame).

[0057] Referring to FIG. 1, in some embodiments, the hybrid codec framework further comprises a motion estimation module 105. Motion estimation is one of the key technologies in video coding, which uses the motion information between adjacent frames to reduce redundant data. Motion estimation is used to determine the position of each image block in adjacent video frames. Specifically, motion estimation is to calculate the direction and amplitude of the movement of objects in the image by analyzing the changes of pixels in the sequence of consecutive images. These motion information can be represented by motion vectors. Through accurate estimation of motion, the amount of video data can be greatly reduced, and at the decoding end, the current frame can be reconstructed based on the motion information and the reference frame, achieving efficient video coding and transmission. Common motion estimation methods include block matching method, etc.

[0058] Referring to FIG. 1, in some embodiments, the hybrid codec framework further comprises a motion compensation module 106. Motion compensation refers to the process of predicting and compensating the current frame based on the motion vectors obtained by the motion estimation module. After the motion estimation module determines the motion state of the object, the motion compensation module constructs a predicted frame based on the motion state of the object, and then encodes the difference between the actual frame and the predicted frame, thereby reducing the amount of data to be transmitted. The motion compensation module can better utilize the temporal correlation in the video, reduce the inter-frame redundancy information, and further improve the coding efficiency, while at the decoding end, the original video picture can be accurately restored based on the compensated information and the existing reference data.

[0059] Referring to FIG. 1, in some embodiments, the hybrid codec framework further comprises an inverse quantization / inverse transform module 107. Inverse quantization refers to the process of converting the quantized signal or data back to the original signal or data during video decoding. Inverse transform refers to the process of restoring the data processed by the transform to the original space or value range during video decoding.

[0060] Referring to FIG. 1, in some embodiments, the hybrid codec framework further comprises a filter control module 108. The filter control module is used to manage and control the loop filtering, including the setting of filter parameters, the selection and adjustment of filter algorithms, etc.

[0061] Referring to FIG. 1, in some embodiments, the hybrid codec framework further comprises a loop filter module 109. Loop filtering is a key technology to remove compression distortion, which can significantly improve the subjective and objective quality of reconstructed video and improve the efficiency of video compression. The prediction module and the transform quantization are both block-based coding technologies, so there will be a situation that the pixels adjacent to the block boundary fluctuate greatly, and some distortion effects still exist, such as block effect, ringing effect, color deviation, and image blur, etc. In order to eliminate this block effect, the reconstructed image after coding needs to be filtered, and the influence of noise is also eliminated.

[0062] Referring to FIG. 1, in some embodiments, the hybrid codec framework further includes an entropy coding module 110. The entropy coding module is a process that utilizes statistical properties of data to further encode the coding control parameters, quantized transform coefficients, intra prediction data, motion data, and filter control parameters, etc. into binary numbers to reduce the size of data for storage and transmission. Entropy coding methods include Huffman coding, arithmetic coding, Shannon coding, etc. which assign shorter coding sequences to symbols with higher occurrence frequencies according to the occurrence frequencies of data, thereby further improving compression efficiency.

[0063] Video codec standards such as H.265 and H.266 are technical specifications designed to decompress video data, maintaining video quality while reducing storage space and transmission bandwidth requirements through various algorithms and techniques, which is particularly important in cloud gaming to alleviate bandwidth pressure and improve transmission efficiency. They also strive to improve video quality, reduce transmission latency, and achieve universality on various devices and platforms, providing users with a better video viewing experience and important technical support for cloud gaming platforms and other applications. The next section describes the current widely used video coding standards.

[0064] Currently, one of the widely used video coding and decoding standards is High Efficiency Video Coding (HEVC / H.265), which was jointly developed by ITU-T and ISO / IEC in 2013. Compared with the H.264 standard, H.265 can provide higher compression rate, i.e. better video quality and smaller file size. H.265 plays an important role in the field of cloud gaming. Through higher compression rate, H.265 can reduce the bandwidth required in the process of video transmission, thereby alleviating the bandwidth pressure of cloud gaming platforms. In addition, H.265 can also provide higher quality video content, which helps to improve the gaming experience of players on cloud gaming platforms.

[0065] Another widely used video coding and decoding standard is the Versatile Video Coding (VVC / H.266), which aims to provide more efficient video compression rate than the related technology standards. The H.266 standard is jointly developed by the International Telecommunication Union (ITU-T) and the International Organization for Standardization (ISO), and was officially released in July 2020. The H.266 standard is a further improvement and optimization of the H.265 / HEVC (High Efficiency Video Coding) standard. Compared with previous standards, H.266 can increase the video compression rate by 50% while maintaining the same video quality. This means that the size of the video file encoded using H.266 will be smaller under the same video quality, so less bandwidth will be required for network transmission, which can effectively alleviate the bandwidth pressure of the cloud game platform. In addition, H.266 also has faster encoding and decoding speed, which can reduce the encoding delay and improve the response speed of the cloud game platform, thus improving the game experience of the players. This is particularly important for game types that require real-time interaction. In general, as a new generation of video coding and decoding standard, H.266 has higher compression rate and faster encoding and decoding speed, and is expected to provide more efficient video transmission solutions for cloud game platforms and further promote the development of cloud games.

[0066] The encoding efficiency of cloud game content encoded by H.265 / HEVC cannot be guaranteed. That is, if H.265 is used to encode cloud game content, there will be a problem of large bandwidth occupation. Although H.266 greatly improves the compression rate of encoding compared to H.265, the time cost is also increased synchronously, so if H.266 is used to encode cloud game content, the low latency requirement cannot be guaranteed. Therefore, no matter which related technology encoding standard is used, it cannot well meet the encoding needs of cloud game content. Moreover, these standards do not consider the specific optimization of cloud game content, and are universal encoding standards. Therefore, it is meaningful to study the optimization of special encoding tools for cloud game content from the encoding needs of cloud game content.

[0067] Some embodiments of the present disclosure relate to encoding cloud game content using auxiliary information of cloud game. First, the auxiliary information of cloud game is described below.

[0068] One of the important auxiliary information of cloud game is depth information. Depth information is used to represent the depth of the object captured by the camera in the camera coordinate system in the current frame, which can be a depth map corresponding to the current frame. Depth information can be used to guide the boundary, division process and motion compensation process of the object.

[0069] One of the important auxiliary information of cloud gaming is motion vector. Motion vector is used to represent the displacement of objects in the current frame relative to the previous video frame, which is equivalent to optical flow information. Motion vector can be used to guide the partitioning process and the motion compensation process.

[0070] One of the important auxiliary information of cloud gaming is camera information. Camera information can include the spatial pose matrix (extrinsic matrix) of the camera and the intrinsic matrix of the camera. Through the spatial pose matrix of the camera and the intrinsic matrix of the camera and the depth information, the real space point information corresponding to each pixel can be calculated.

[0071] One of the important auxiliary information of cloud gaming is skybox. Skybox is a technique used in games and virtual scenes to simulate the sky. This technique surrounds the scene with a large box or sphere, and uses texture images or program-generated images to present environmental information, in order to increase the atmosphere of the scene and enhance the immersion of cloud gaming. In the design of cloud gaming, skybox has no other use except to provide a certain visual completion effect. Therefore, in the current mainstream game market, major game manufacturers tend to design skybox simply and focus more on the production of game content. In addition, skybox has a characteristic: when the virtual camera of the game is translated, the content of the skybox does not change, only when the virtual camera is rotated, the content of the skybox will change. As shown in FIG. 2, the current frame 21 and the second video frame 22 in FIG. 2 are two video frames in the game video with an interval of 100 frames, and the image 23 is the frame difference image of the current frame 21 and the second video frame 22. It can be seen that under such a large interval of 100 frames, the difference of the sky part 230 is still very clean, and the change of the skybox is very small. Taking advantage of the characteristic that the skybox does not change when the virtual camera is translated, the redundancy in the video encoding data can be further eliminated, and the encoding efficiency can be improved.

[0072] With the gradual popularization of high-resolution video formats such as 4K and 8K, and the rapid development of the virtual reality and gaming industries, video content is becoming increasingly complex, and its data volume is also increasing dramatically. Among these video contents, there are often a large number of background parts that change very little, such as the vast sky in game videos, the stable background scenes in surveillance videos, and the background blackboard in remote education scenes, etc. Although these background contents change little in vision, a large amount of data is still needed for encoding to maintain high-quality images. In the face of this challenge, the progress of video coding technology has shown great potential. By improving the coding algorithm, more effective compression strategies can be implemented for these less changing backgrounds. For example, larger image blocks and longer prediction periods can be used, or coarser quantization methods can be applied to these regions to reduce the amount of data required while maintaining good visual effects. However, this part of the background content has a high correlation with camera information, but the related technology does not make good use of information other than spatiotemporal pixel information for encoding. In addition, in the current field of video coding, although the encoder of the related technology performs well in processing dynamic and rich video content, it still needs to make complex predictions and residual coding when encoding video content with less changing background, which also consumes unnecessary coding resources and data bandwidth. The existence of this problem not only affects the coding efficiency, but also increases the cost of storing and transmitting high-quality video content. In today's digital era of growing data volume, this inefficient coding method puts a huge pressure on network bandwidth and storage resources. Therefore, solving this problem not only improves the overall performance of video coding, but also brings a smoother video experience to users and saves resources for Internet services.

[0073] In summary, solving the coding problem of video content with less changing background has important practical significance and far-reaching influence. It not only improves the efficiency of video coding and reduces the cost of data storage and transmission, but also promotes the progress of video technology and provides strong technical support for future video applications such as remote education, game content coding, and video surveillance. This will be an important milestone in the development of digital video technology.

[0074] The encoder standard of the related technology does not design corresponding coding tools for stable background video sequences with strong camera pose correlation to improve coding performance. Stable background video sequences refer to video sequences with partially stable backgrounds in video frame. The background region has a strong correlation with camera pose information. For example, the sky region in game sequences, the background region in surveillance sequences, and the background blackboard in remote education, etc. When the background region is encoded using coding tools, it is usually not fully utilized with stable background features that are strongly related to the camera. Based on this, some embodiments of the present disclosure propose a coding scheme based on camera information, which can improve the coding efficiency of stable background video sequences.

[0075] Referring to FIG. 3, the method for encoding a video based on camera information provided by some embodiments of the present disclosure includes the following steps:

[0076] S301: Obtain background content of a current video sequence and camera information of each video frame of the current video sequence.

[0077] The background content of the current video sequence can be any background content that is strongly related to the camera pose. For example, the background content of the current video sequence can be a skybox in a game video, a monitoring scene content in a monitoring video, etc. For example, if the current video sequence is a video sequence in a cloud game video, the skybox obtained through a game rendering engine is the background content of the current video sequence. For another example, if the current video sequence is a video sequence in a monitoring video, the monitoring scene image captured by at least one camera pose can be used as the background content of the current video sequence.

[0078] In some embodiments, the camera information of a video frame includes camera intrinsic information of the video frame and camera extrinsic information of the video frame. The camera intrinsic information of the video frame can include a vertical field of view angle. The camera extrinsic information can include extrinsic quaternion of the camera.

[0079] S302: Encode the background content to obtain background content encoding data.

[0080] In some embodiments of the present disclosure, the encoding manner for encoding the background content is not limited, and the background content can be encoded by using any encoding manner. For example, the background content can be encoded by using an encoding manner specified in a video encoding standard such as H.265 / HEVC, H.266 / VCC, etc. to obtain the background content encoding data.

[0081] In some embodiments, the background content of the current video sequence can be encoded as an independent image, and the encoded data is transmitted to a decoding end.

[0082] S303: Encode the camera information of each video frame of the current video sequence to obtain camera information encoding data.

[0083] In some embodiments of the present disclosure, the encoding manner for encoding the camera information of each video frame is not limited, and the saving form of the camera information of each video frame in the camera information encoding data is also not limited. The camera information of each video frame of the current video sequence can be obtained by decoding the camera information encoding data.

[0084] S304: Determine whether a current block is an image block corresponding to the background content.

[0085] In some embodiments, the current block is any image block obtained by block partitioning a current frame of a current video sequence.

[0086] In some embodiments of the present disclosure, the current block can be a Coding Tree Unit (CTU) or a Coding Unit (CU).

[0087] In some embodiments, whether the current block is the image block corresponding to the background content can be determined based on the depth information of the pixels. If the depth of each pixel of the current block is the depth of the background content, the current block is the image block corresponding to the background content; if the depth of one or more pixels of the current block is not the depth of the background content, the current block is not the image block corresponding to the background content. For example, in normalized depth, only the depth of the pixels in the background region corresponding to the background content is 1, and the depth of the pixels in the image region corresponding to the foreground content is less than 1, so whether the current block is the image block corresponding to the background content can be determined by determining whether the depth values of the pixels of the current block are all 1. If the depth values of the pixels of the current block are all 1, the current block is the image block corresponding to the background content; if the depth values of one or more pixels of the current block are not 1 (less than 1), the current block is not the image block corresponding to the background content.

[0088] In some embodiments, whether the current block is the image block corresponding to the background content can be determined based on the motion vectors of the pixels. For example, in a surveillance video sequence, if the camera pose does not change, the motion vectors of the pixels in the background region corresponding to the surveillance scene will not change, so a reference frame can be obtained, wherein the camera pose of the reference frame is the same as that of the current frame; then, whether the motion vectors of the pixels of the current block are all 0 is determined according to the reference frame; if the motion vectors of the pixels of the current block are all 0, it is determined that the current block is the image block corresponding to the background content; if the motion vectors of one or more pixels of the current block are not 0, it is determined that the current block is not the image block corresponding to the background content. For another example, in a game video sequence, when the camera pose is the same or the camera only undergoes a pure translation motion (without rotation motion), the pixels in the background region corresponding to the background content such as the sky box will not change, so a reference frame can be obtained, wherein the camera pose of the reference frame is the same as that of the current frame or only undergoes a translation motion; then, whether the motion vectors of the pixels of the current block are all 0 is determined according to the reference frame; if the motion vectors of the pixels of the current block are all 0, it is determined that the current block is the image block corresponding to the background content; if the motion vectors of one or more pixels of the current block are not 0, it is determined that the current block is not the image block corresponding to the background content.

[0089] In some embodiments, the current block can be determined to be the image block corresponding to the background content by using the frame difference image. For example, in a monitoring video sequence, if the camera pose does not change, the pixels in the background region corresponding to the monitoring scene do not change, and thus a reference frame can be obtained, the camera pose of the reference frame being the same as that of the current frame; then, the frame difference image of the current frame and the reference frame is calculated, and it is determined whether the pixel values of each pixel of the frame difference image block corresponding to the current block in the frame difference image are all 0; if the pixel values of each pixel of the frame difference image block are all 0, it is determined that the current block is the image block corresponding to the background content; if the pixel values of one or more pixels of the frame difference image block are not 0, it is determined that the current block is not the image block corresponding to the background content. For another example, in a game video sequence, when the camera pose is the same or the camera undergoes only a pure translation motion, the pixels in the background region corresponding to the background content such as the sky box do not change, and thus a reference frame can be obtained, the camera pose of the reference frame being the same as that of the current frame or only undergoing a translation motion; then, the frame difference image of the current frame and the reference frame is calculated, and it is determined whether the pixel values of each pixel of the frame difference image block corresponding to the current block in the frame difference image are all 0; if the pixel values of each pixel of the frame difference image block are all 0, it is determined that the current block is the image block corresponding to the background content; if the pixel values of one or more pixels of the frame difference image block are not 0, it is determined that the current block is not the image block corresponding to the background content.

[0090] In some embodiments, in the above embodiments, a reference frame of the current frame can also be obtained by performing three-dimensional warping (3D warping) on a video frame in which the camera pose undergoes a compound motion (translation motion + rotation motion) according to the motion information of the rotation motion.

[0091] In some embodiments, the current block can be determined to be the image block corresponding to the background content according to the camera information. For example, the image region corresponding to the background content can be determined according to the camera information; a ray passing through each pixel of the current block is drawn from the camera position, and it is determined whether the ray passing through each pixel of the current block is projected into the image region corresponding to the background content; if the rays passing through each pixel of the current block are all projected into the image region corresponding to the background content, it is determined that the current block is the image block corresponding to the background content; if one or more rays passing through each pixel of the current block are not projected into the image region corresponding to the background content, it is determined that the current block is not the image block corresponding to the background content.

[0092] In step S304, if the current block is the image block corresponding to the background content, step S305 is performed.

[0093] S305, selecting the coding mode of the current block from all coding modes in the preset coding mode set.

[0094] The preset coding mode set includes the target coding mode.

[0095] In some embodiments, based on the current block being the image block corresponding to the background content, the coding mode of the current block is determined as the target coding mode.

[0096] In some embodiments, the target coding mode can include an identification syntax element (bg_skip), and when the coding mode of the current block is the target coding mode, i.e., when the current block is the image block corresponding to the background content, the value of the identification syntax element is set to 1, i.e., bg_skip: 1. When the coding mode of the current block is not the target coding mode, i.e., when the current block is not the image block corresponding to the background content, the value of the identification syntax element is set to 0, i.e., bg_skip: 0.

[0097] In some embodiments, selecting the coding mode of the current block from all coding modes in the preset coding mode set includes: obtaining the rate-distortion cost (RDcost) of each coding mode in the preset coding mode set, and selecting the coding mode with the minimum rate-distortion cost as the coding mode of the current block.

[0098] In some embodiments of the present disclosure, the rate-distortion cost of any coding mode refers to the rate-distortion cost of encoding the current block based on the coding mode.

[0099] In some embodiments, in addition to the target coding mode, the preset coding mode set can also include one or more coding modes specified in video coding standards such as H.265 / HEVC, H.266 / VCC, etc.

[0100] In step S305, when the selected coding mode of the current block is the target coding mode, step S306 is performed:

[0101] S306, encode the identification information of the target coding mode to obtain the encoding data of the current block.

[0102] That is, when the current block is encoded by the target coding mode, the encoding data of the current block only includes the identification information of the target coding mode, and does not include the residual data of the current block.

[0103] In some embodiments, when the value of a specified data bit in the encoding data of the current block is used to identify the coding mode of the current block, the value of the specified data bit can be set to the value corresponding to the target coding mode.

[0104] In some embodiments, encoding the identification information of the target coding mode to obtain the coding data of the current block comprises setting the value of the identification syntax element to 1.

[0105] The steps S304 to S307 are repeated to complete the encoding of all image blocks of the current video sequence. After obtaining the coding data of all image blocks of the current video sequence, the step S307 is executed:

[0106] S307, generating the coding data of the current video sequence according to the background content coding data, the camera information coding data, and the coding data of each image block of each video frame of the current video sequence.

[0107] The video encoding method provided by some embodiments of the present disclosure first obtains the background content of the current video sequence and the camera information of each video frame of the current video sequence, then encodes the background content to obtain the background content coding data, encodes the camera information of each video frame of the current video sequence to obtain the camera information coding data, determines whether the current block is an image block corresponding to the background content, and encodes the identification information of the target coding mode to obtain the coding data of the current block when the selected coding mode of the current block is the target coding mode, and generates the coding data of the current video sequence according to the background content coding data, the camera information coding data, and the coding data of each image block of each video frame of the current video sequence. Since the video encoding method provided by some embodiments of the present disclosure does not need to add the residual data of the image block corresponding to the background content in the video sequence to the coding data, some embodiments of the present disclosure can reduce the data amount of the coding data and improve the efficiency of video encoding and decoding.

[0108] As an extension and refinement of the above-mentioned embodiments, some embodiments of the present disclosure provide another video encoding method. Referring to FIG. 4, the video encoding method comprises the following steps:

[0109] S401, obtaining the background content of the current video sequence and the camera information of each video frame of the current video sequence.

[0110] The camera information of the video frame comprises camera extrinsic parameter information of the video frame and camera intrinsic parameter information of the video frame.

[0111] S402, encoding the background content to obtain the background content coding data.

[0112] S403, encoding the camera intrinsic parameter information of any video frame of the current video sequence to obtain camera intrinsic parameter coding data.

[0113] The camera intrinsic parameter information is sequence-level information, and the coding data of the current video sequence only includes one camera intrinsic parameter coding data.

[0114] Since the camera intrinsic information is stable and unchanged in the whole video sequence, the camera intrinsic information can be coded as sequence-level information, so that the current video sequence only includes one camera intrinsic coding data in the coding data, thereby further reducing the data amount of the coding data of the video sequence.

[0115] S404, respectively, the camera extrinsic information of each video frame of the current video sequence is coded to obtain the camera extrinsic coding data of each video frame of the current video sequence.

[0116] The camera extrinsic information is frame-level information, and the camera extrinsic coding data of each video frame is included in the coding data of the current video sequence.

[0117] Since the camera poses of different video frames may not be the same, the camera extrinsic information of different video frames may not be the same, so the extrinsic information of the video frame needs to be coded as frame-level information.

[0118] S405, determine whether the current block is the image block corresponding to the background content.

[0119] The current block is any image block obtained by block division on the current frame of the current video sequence.

[0120] In step S405, if the current block is not the image block corresponding to the background content, step S406 is executed:

[0121] S406, set the rate-distortion cost of the target coding mode to the maximum value of the rate-distortion cost.

[0122] In some embodiments, the maximum value of the rate-distortion cost can be infinity.

[0123] If it is determined in step S405 that the current block is the image block corresponding to the background content or after step S406 (setting the rate-distortion cost of the target coding mode to the maximum value of the rate-distortion cost) is executed, step S is executed:

[0124] S407, obtain the rate-distortion cost of each coding mode in the preset coding mode set.

[0125] The preset coding mode set includes the target coding mode and other coding modes. The rate-distortion cost of any coding mode is the rate-distortion cost based on coding the current block by using the coding mode.

[0126] It should be noted that if step S406 is executed, the rate-distortion cost of the target coding mode obtained in step S407 is the maximum value of the rate-distortion cost.

[0127] S408, encode the current block by selecting the coding mode with the minimum rate-distortion cost in the preset coding mode set to obtain the encoding data of the current block.

[0128] When the selected coding mode of the current block is the target coding mode, the step of encoding the current block by selecting the coding mode with the minimum rate-distortion cost in the preset coding mode set to obtain the encoding data of the current block comprises the step of encoding the identification syntax element corresponding to the target coding mode to obtain the encoding data of the current block.

[0129] It should be noted that if the step S406 is performed, the target coding mode will not be selected as the coding mode of the current block in the step S408. Only when the current block is the image block corresponding to the background content, the coding mode of the current block can be the target coding mode. Therefore, the step of determining the coding mode of the current block refers to selecting the coding mode of the current block from the preset coding mode set when the current block is the image block corresponding to the background content, and selecting the coding mode of the current block from the coding modes other than the target coding mode in the preset coding mode set when the current block is not the image block corresponding to the background content. In some embodiments, when the current block is the image block corresponding to the background content, the rate-distortion cost of the target coding mode is generally the minimum value among all coding modes. Therefore, if the current block is the image block corresponding to the background content, the coding mode of the current block is the target coding mode.

[0130] The image blocks obtained by performing the block division on each video frame of the current sequence are sequentially taken as the current block in some embodiments of the present disclosure, and the steps S405 to S408 are repeatedly performed to obtain the encoding data of the image blocks obtained by performing the block division on each video frame of the current sequence, and then the step S409 is performed. In some embodiments, the specific implementation scheme can comprise the following steps: after performing the step S408, it is judged whether the current block is the last image block in the image blocks obtained by performing the block division on each video frame of the current sequence, if not, the step S405 is returned, and if yes, the step S409 is performed.

[0131] S409, generating the encoding data of the current video sequence according to the background content encoding data, the camera intrinsic parameter encoding data, the camera extrinsic parameter encoding data of each video frame of the current video sequence, and the encoding data of each image block of each video frame of the current video sequence.

[0132] Since the current block can also be a reference block or part of a reference block of other image blocks in a subsequent encoding process, the video encoding method provided by the present disclosure further comprises reconstructing the current block to obtain a reconstructed block of the current block. When the encoding mode of the current block is the target encoding mode, reconstructing the current block to obtain a reconstructed block of the current block can comprise decoding the background content encoding data to obtain a reconstructed background content, and reconstructing the current block according to the reconstructed background content and camera information of the current frame to obtain a reconstructed block of the current block.

[0133] The decoding scheme for decoding the background content encoding data depends on the encoding scheme for encoding the background content in step S402, so the scheme for decoding the background content encoding data can comprise determining the decoding scheme corresponding to the encoding scheme for encoding the background content in step S402, and decoding the background content encoding data by the decoding scheme to obtain a reconstructed background content. For example, the background content is encoded by the encoding scheme specified in H.265 / HEVC in step S402 to obtain the background content encoding data, and the background content encoding data is decoded by the decoding scheme specified in H.265 / HEVC to obtain a reconstructed background content. For another example, the background content is encoded by the encoding scheme specified in H.266 / VCC in step S402 to obtain the background content encoding data, and the background content encoding data is decoded by the decoding scheme specified in H.266 / VCC to obtain a reconstructed background content.

[0134] In some embodiments, referring to FIG. 5, reconstructing the current block according to the reconstructed background content and camera information of the current frame to obtain a reconstructed block of the current block comprises:

[0135] S501, determining the position coordinates of the mapping points of each pixel of the current block in the reconstructed background content according to the camera information of the current frame.

[0136] In some embodiments, the camera information of the current frame comprises a vertical field of view angle of the current frame and an extrinsic quaternion of the current frame. The determination of the position coordinates of the mapping points of each pixel of the current block in the reconstructed background content according to the camera information of the current frame comprises steps a to e:

[0137] Step a, obtaining the depth of each pixel in the current block according to the vertical field of view angle.

[0138] In some embodiments, obtaining the depth of each pixel in the current block according to the vertical field of view angle comprises obtaining the depth of each pixel in the current block by the following formula (1): z = cot(FOV ver / 2)*0.5 (1)

[0139] wherein z is the depth of each pixel in the current block, FOV is the field of view of the camera. ver is the vertical field of view.

[0140] For example, referring to FIG. 6, the perpendicular line from the position O of the camera to the current frame 60 is the line segment OC, the triangle OCD is a right triangle, and the angle DOC is half of the vertical field of view of the camera, the angle DCO is a right angle, and the length of the line segment DC is half of the height of the current frame 60 (0.5 under the normalized height), so the length of the line segment OC can be calculated by the above formula (1). The length of the line segment OC is the depth of each pixel in the current block.

[0141] Step b, obtaining the position coordinates of each pixel in the current block in the camera coordinate system according to the depth of each pixel in the current block and the pixel coordinates of each pixel in the current block.

[0142] In some embodiments, the pixel coordinates of the pixels in the video frame are coordinates in a two-dimensional coordinate system with the upper right corner of the video frame as the origin, and the camera coordinate system is a three-dimensional coordinate system, and the origin is located at the geometric center of the current frame in the orthographic projection of the plane where the video frame is located, so the coordinates of each pixel in the current block in the camera coordinate system can be calculated by the following formula (2): (x, y, z) = (u-0.5, v-0.5, z) (2)

[0143] wherein (u, v) is the normalized coordinates of the pixels in the current block, (x, y, z) is the normalized coordinates of the pixels in the current block in the camera coordinate system, and z is the depth of the pixels in the current block.

[0144] Step c, obtaining the direction of the ray corresponding to each pixel in the current block in the background coordinate system according to the position coordinates of each pixel in the current block in the camera coordinate system and the extrinsic quaternion of the current frame.

[0145] wherein the ray corresponding to any pixel is a ray starting from the coordinate origin of the camera coordinate system and passing through the pixel, and the background coordinate system is the coordinate system of the reconstructed background content.

[0146] In some embodiments, obtaining the direction of the ray corresponding to each pixel in the current block in the background coordinate system according to the position coordinates of each pixel in the current block in the camera coordinate system and the extrinsic quaternion of the current frame comprises steps c1 to c4:

[0147] Step c1, obtaining the vector from the camera to each pixel in the current block in the camera coordinate system according to the coordinates of each pixel in the current block in the camera coordinate system.

[0148] Since the camera is located at the coordinate origin in the camera coordinate system, if a pixel in the current block has coordinates (x1, y1, z) in the camera coordinate system, the vector from the camera to the pixel in the camera coordinate system is v, where v = (x1, y1, z) (3)

[0149] Step c2, extending the vector from the camera to each pixel in the current block in the camera coordinate system into a quaternion to obtain the quaternion of each pixel in the current block in the camera coordinate system.

[0150] The vector from the camera to a pixel in the camera coordinate system is v = (x v ,y v ,z v ), and the corresponding quaternion of the vector from the camera to the pixel is v q , where v q = (0, x v ,y v ,z v ) (4)

[0151] Step c3, obtaining the quaternion of each pixel in the current block in the background coordinate system according to the extrinsic quaternion of the current frame and the quaternion of each pixel in the current block in the camera coordinate system.

[0152] In some embodiments, obtaining the quaternion of each pixel in the current block in the background coordinate system according to the extrinsic quaternion and the quaternion of each pixel in the current block in the camera coordinate system includes obtaining the quaternion of each pixel in the current block in the coordinate system of the background content by the following formula (5): v′ q = q·v q ·q * (5)

[0153] where v′ q is the quaternion of the pixel in the current block in the background coordinate system, q is the extrinsic quaternion, v q is the quaternion of the pixel in the current block in the camera coordinate system, and q * is the conjugate quaternion of the extrinsic quaternion.

[0154] Step c4, obtaining the direction of the ray from the camera to each pixel in the current block in the coordinate system of the background content according to the quaternion of each pixel in the current block in the coordinate system of the background content.

[0155] Step d, obtaining the intersection point of the ray corresponding to each pixel in the current block and the reconstructed background content according to the direction of the ray corresponding to each pixel in the current block in the background coordinate system.

[0156] Step e, determining the position coordinates of the intersection of the ray corresponding to each pixel in the current block and the reconstructed background content as the position coordinates of the mapping point of each pixel in the current block in the reconstructed background content.

[0157] For example, referring to FIG. 7, the intersection of the ray passing through pixel A in the current block 71 from the camera (point O) and the skybox 700 is point B, and thus the position coordinates of point B are determined as the position coordinates of the mapping point of pixel A in the skybox 700. In addition, since the direction of the ray passing through pixel A from point O has been obtained, the position coordinates of point B can be obtained according to the expression of the ray passing through pixel A from point O and the skybox 700.

[0158] S502, obtaining the reconstructed pixel value of each pixel in the current block according to the position coordinates of the mapping point of each pixel in the current block in the reconstructed background content and the pixel value of the reconstructed background content.

[0159] In some embodiments, obtaining the reconstructed pixel value of each pixel in the current block according to the position coordinates of the mapping point of each pixel in the current block in the reconstructed background content and the pixel value of the reconstructed background content comprises steps 1 and 2:

[0160] Step 1, performing pixel interpolation on the mapping point of each pixel in the current block in the reconstructed background content according to the pixel value of the reconstructed background content and the position coordinates of the mapping point of each pixel in the current block in the reconstructed background content of the current video sequence, thereby obtaining the interpolated pixel value of the mapping point of each pixel in the current block in the reconstructed background content.

[0161] The following takes the preset pixel precision of 1 / 16 precision 8-tap luminance and 1 / 32 precision 4-tap chrominance sub-pixel interpolation as an example to illustrate step 1.

[0162] Referring to FIG. 8, the position coordinates of the mapping point 81 of a certain pixel in the current block in the reconstructed background content of the current video sequence are (i+4 / 16, j+2 / 16), and the pixel coordinates of the upper left corner of the integral pixel 82 of the mapping point 81 are (i, j). Since the 1 / 16 precision 8-tap luminance coefficient and the 1 / 32 precision 4-tap chrominance coefficient are as shown in Table 1 and Table 2 respectively:

[0163] Table 1

[0164] Table 2

[0165] Since the integral pixels participating in the interpolation are integral pixels with pixel coordinates (i-3, j), (i-2, j), (i-1, j), (i, j), (i+1, j), (i+2, j), (i+3, j), (i+4, j), and the corresponding interpolation coefficients are -1, 4, -10, 58, 17, -5, 1, 0 according to Table 1, the luminance value of the pixel with pixel coordinates (x, y) is denoted as f(x, y), and the interpolation luminance value of the interpolation point 73 in the row is denoted as f(i+4 / 16, j), then the following equation is obtained: f(i+4 / 16, j) = -1×f(i-3, j) + 4×f(i-2, j) -10×f(i-1, j) + 58×f(i, j) + 17×f(i+1, j) -5×f(i+2, j) +1×f(i+3, j) + 0×f(i+4, j) (6)

[0166] The same row interpolation operation is performed on the j-3th to j+4th rows to obtain the interpolation pixel values of the interpolation points in the j-3th to j+4th rows.

[0167] It should be noted that, when performing the row interpolation operation, if the interpolation integral pixel exceeds the image boundary, the luminance value at the image boundary can be used as the luminance value of the interpolation integral pixel for the row interpolation operation.

[0168] After the row interpolation operation on the j-3th to j+4th rows is completed, there are four sub-pixels in the integral pixel row above and below the mapping point 81, and the corresponding interpolation coefficients are -1, 2, -5, 62, 8, -3, 1, 0 according to Table 1. Therefore, the luminance value of the pixel with pixel coordinates (x, y) is denoted as f(x, y), and the following equation is obtained: f(i+4 / 16, j+2 / 16) = -1×f(i+4 / 16-3, j) + 4×f(i+4 / 16-2, j) -10×f(i+4 / 16-1, j) + 58×f(i+4 / 16, j) + 17×f(i+4 / 16+1, j) -5×f(i+4 / 16+2, j) +1×f(i+4 / 16+3, j) + 0×f(i+4 / 16+4, j) (7)

[0169] The luminance value is expanded by 2 6 times during the row interpolation and column interpolation, and the overall luminance value is expanded by 2 12 times. Therefore, after f(i+4 / 16, j+2 / 16) is obtained by the above equation (7), a right shift operation of 12 bits is performed on the final obtained 2

[0170] Further, based on Table 2, the interpolation chrominance of the mapping point 81 is calculated in a similar way as the interpolation luminance of the mapping point 81, and the reconstructed pixel value of the mapping point 81 is obtained according to the interpolation luminance of the mapping point 81 and the interpolation chrominance of the mapping point 81. For the sake of brevity, the details are not described here.

[0171] S503, obtaining the reconstructed block of the current block according to the reconstructed pixel value of each pixel of the current block.

[0172] In some embodiments, the pixel value of each pixel of the current block is assigned as the corresponding reconstructed pixel value, thereby obtaining the reconstructed block of the current block.

[0173] In some embodiments, the video encoding method provided by some embodiments of the present disclosure can also only add the background content encoding data to the encoding data of the current video sequence when the current video sequence includes the image block with the target encoding mode, and not add the background content encoding data to the encoding data of the current video sequence when the current video sequence includes the image block without the target encoding mode. In addition, only the camera information encoding data of the video frame including the image block with the target encoding mode is added to the encoding data of the current video sequence, instead of adding the camera information encoding data of all video frames to the encoding data of the current video sequence. Referring to FIG. 9, another video encoding method provided by some embodiments of the present disclosure includes the following steps:

[0174] S901, obtaining the background content of the current video sequence and the camera information of each video frame of the current video sequence.

[0175] S902, determining whether the current block is the image block corresponding to the background content.

[0176] The current block is any image block obtained by block division on the current frame of the current video sequence.

[0177] In step S902, if the current block is not the image block corresponding to the background content, step S903 is performed:

[0178] S903, setting the rate-distortion cost of the target encoding mode as the maximum value of the rate-distortion cost.

[0179] The target encoding mode is to encode the first identification information corresponding to the target encoding mode to obtain the encoding data of the current block.

[0180] If the current block is determined to be the image block corresponding to the background content in step S902 or after step S903 is performed, the video encoding method performs the following steps:

[0181] S904, obtain the rate-distortion cost of each encoding mode in the preset encoding mode set.

[0182] The preset encoding mode set includes the target encoding mode, and the rate-distortion cost of any encoding mode is the rate-distortion cost of encoding the current block based on the encoding mode.

[0183] S905, select the encoding mode with the minimum rate-distortion cost in the preset encoding mode set to encode the current block to obtain the encoding data of the current block.

[0184] After the encoding of all the encoding units of all the video frames of the current video sequence is completed through the above-mentioned loop execution of steps S902 to S905, step S906 is further executed:

[0185] S906, determine whether the current video sequence includes an image block with the target encoding mode as the encoding mode.

[0186] In step S906, if the current video sequence does not include an image block with the target encoding mode as the encoding mode, step S907 is executed:

[0187] S907, generate the encoding data of the current video sequence according to the encoding data of each image block of each video frame in the current video sequence.

[0188] In step S906, if the current video sequence includes an image block with the target encoding mode as the encoding mode, steps S908 to S911 are executed:

[0189] S908, encode the background content of the current video sequence to obtain background content encoding data.

[0190] S909, determine the video frame in the current video sequence that includes an image block with the target encoding mode as the encoding mode.

[0191] S910, encode the camera information of the video frame in the current video sequence that includes an image block with the target encoding mode as the encoding mode to obtain camera information encoding data.

[0192] For example, if the current video sequence includes video frame A, video frame B, video frame C, video frame D, and video frame E, and only video frame B and video frame C include an image block with the target encoding mode as the encoding mode, the camera information of video frame B and video frame C is obtained and encoded to obtain the camera information encoding data, but the camera information of video frame A, video frame D, and video frame E is not obtained and encoded.

[0193] S911、generate the encoded data of the current video sequence according to the encoded data of each image block of each video frame in the current video sequence, the encoded data of the background content, and the encoded data of the camera information.

[0194] For the above-mentioned video encoding method, some embodiments of the present disclosure further provide a corresponding video decoding method. Referring to FIG. 10, the video decoding method comprises the following steps:

[0195] S1001, obtain the encoded data of a current block.

[0196] In some embodiments, obtaining the encoded data of the current block can comprise receiving the encoded data of the current block sent by a media resource server such as a cloud game server or a monitoring server.

[0197] In some embodiments, obtaining the encoded data of the to-be-decoded video sequence can comprise reading the encoded data of the current block from a preset storage space.

[0198] S1002, determine the encoding mode of the current block according to the encoded data of the current block.

[0199] In some embodiments, the encoded data of the current block comprises an identification syntax element corresponding to a target encoding mode and a value thereof, and determining the encoding mode of the current block as the target encoding mode comprises: when the value of the identification syntax element corresponding to the target encoding mode is a first preset value, determining the encoding mode of the current block as the target encoding mode.

[0200] For example, the identification syntax element corresponding to the target encoding mode can be bg_skip, and the first preset value can be 1, that is, when the identification syntax element corresponding to the target encoding mode and the value thereof are “bg_skip: 1”, it is determined that the encoding mode of the current block is the target encoding mode, and when the identification syntax element corresponding to the target encoding mode and the value thereof are “bg_skip: 0”, it is determined that the encoding mode of the current block is not the target encoding mode.

[0201] In some embodiments, when the value of the identification syntax element corresponding to the target encoding mode is 1, it indicates that the current block is an image block corresponding to the background content. When the value of the identification syntax element corresponding to the target encoding mode is 0, it indicates that the current block is not an image block corresponding to the background content.

[0202] In step S1002, if it is determined that the encoding mode of the current block is the target encoding mode, steps S1003 and S1004 are performed.

[0203] S1003, obtain the reconstructed background content and the camera information of the current frame.

[0204] The reconstructed background content is background content of the current video sequence reconstructed according to background content coding data of the current video sequence, and the background content coding data of the current video sequence is coding data obtained by encoding the background content of the current video sequence.

[0205] S1004, reconstructing the current block according to the camera information of the current frame and the reconstructed background content to obtain a reconstructed block of the current block.

[0206] Some embodiments of the present disclosure provide a video decoding method. After obtaining the coding data of the current block, the video decoding method determines the coding mode of the current block according to the coding data of the current block, and in the case that the coding mode of the current block is a target coding mode, the video decoding method obtains the background content of the current video sequence reconstructed according to the coding data obtained by encoding the background content of the current video sequence and the camera information of the current frame, and reconstructs the current block according to the camera information of the current frame and the reconstructed background content to obtain a reconstructed block of the current block. Since the video decoding method provided by some embodiments of the present disclosure can reconstruct the current block according to the camera information of the current frame and the reconstructed background content, for the image block with the target coding mode, the residual data of the image block does not need to be added in the coding data, and therefore some embodiments of the present disclosure can reduce the data amount of the coding data and improve the efficiency of video coding and decoding.

[0207] As an extension and refinement of the above-mentioned embodiments, some embodiments of the present disclosure provide another video decoding method. Referring to FIG. 11, the video decoding method includes the following steps:

[0208] S1101, obtaining the coding data of the current block.

[0209] S1102, determining the coding mode of the current block according to the coding data of the current block.

[0210] In step S1102, if the coding mode of the current block is a target coding mode, the following step is performed:

[0211] S1103, determining whether the reconstructed background content is saved in the background content storage space corresponding to the current video sequence.

[0212] If the reconstructed background content is saved in the background content storage space, step S1104 is performed:

[0213] S1104, reading the reconstructed background content from the background content storage space.

[0214] If the reconstructed background content is not saved in the background content storage space, steps S1105 to S1107 are performed:

[0215] S1105、read the background content coding data from the coding data of the current video sequence.

[0216] S1106、decode the background content coding data to obtain the reconstructed background content.

[0217] In some embodiments, the encoding manner of encoding the background content can be carried in the background content coding data, and step S116 includes: decoding the background content coding data by a decoding manner corresponding to the encoding manner to obtain the reconstructed background content.

[0218] S1107、save the reconstructed background content to the background content storage space.

[0219] S1108、determine whether the camera information of the current frame is saved in the camera information storage space corresponding to the current frame.

[0220] If the camera information of the current frame is saved in the camera information storage space corresponding to the current frame, step S1109 is performed.

[0221] S1109、read the camera information of the current frame from the camera information storage space corresponding to the current frame.

[0222] If the camera information of the current frame is not saved in the camera information storage space corresponding to the current frame, steps S1110 and S1111 are performed.

[0223] S1110、obtain the camera information of the current frame according to the coding data of the current video sequence.

[0224] In some embodiments, the camera information of the current frame includes camera intrinsic information of the current frame and camera extrinsic information of the current frame, and obtaining the camera information of the current frame according to the coding data of the current video sequence includes steps ① to ④.

[0225] Step ①, read camera extrinsic coding data of the current frame from the coding data of the current video sequence.

[0226] The camera extrinsic coding data of the current frame is coding data obtained by encoding camera extrinsic information of the current frame.

[0227] Step ②, decode the camera extrinsic coding data of the current frame to obtain the camera extrinsic information of the current frame.

[0228] Step ③, read camera intrinsic coding data from the coding data of the current video sequence.

[0229] The camera intrinsic coding data is coding data obtained by encoding camera intrinsic information of any video frame of the current video sequence.

[0230] Step IV, decoding the camera intrinsic parameter coding data to obtain camera intrinsic parameter information of the current frame.

[0231] S1111, saving the camera information of the current frame to the camera information storage space corresponding to the current frame.

[0232] S1112, determining position coordinates of mapping points of each pixel of the current block in the reconstructed background content according to the camera information of the current frame.

[0233] S1113, obtaining reconstructed pixel values of each pixel of the current block according to the position coordinates of the mapping points of each pixel of the current block in the reconstructed background content and pixel values of the reconstructed background content.

[0234] S1114, obtaining a reconstructed block of the current block according to the reconstructed pixel values of each pixel of the current block.

[0235] The implementation manners of steps S1112-S114 can be the same as those of steps S501-S503, and thus repeated description is omitted here.

[0236] In step S1102, if the coding mode of the current block is not the target coding mode, the coding data of the current block is decoded through a decoding mode corresponding to the coding mode of the current block to obtain a reconstructed block of the current block.

[0237] Some embodiments of the present disclosure also provide another video encoding method, which is shown in FIG. 12 and includes the following steps:

[0238] S1201, obtaining camera information of each video frame of a current video sequence.

[0239] In some embodiments, the camera information of the video frame includes camera intrinsic parameter information of the video frame and camera extrinsic parameter information of the video frame, the camera intrinsic parameter information of the video frame can include a vertical field of view, and the extrinsic parameter information of the camera can include extrinsic parameter quaternions of the camera.

[0240] The obtaining of the camera information of each video frame of the current video sequence can be reading the camera information of each video frame from a specified location or receiving the camera information of each video frame sent by a specified device. For example, if the current video sequence is a video sequence in a cloud game video, the camera information of each video frame of the current video sequence can be obtained through a game rendering engine. For another example, if the current video sequence is a video sequence in a monitoring video, the camera information of each video frame of the current video sequence can be received from a monitoring device.

[0241] S1202, encoding the camera information of each video frame of the current video sequence to obtain camera information coding data.

[0242] Some embodiments of the present disclosure do not limit the encoding manner of the camera information of each video frame, nor the saving form of the camera information corresponding to each video frame in the camera information encoding data, but are subject to the camera information of each video frame of the current video sequence being able to be acquired by decoding the camera information encoding data.

[0243] In S1203, it is determined whether the current block is a background block.

[0244] The current block is any image block obtained by block division on a current frame of a current video sequence.

[0245] In some embodiments of the present disclosure, the current block can be a coding tree unit (CTU) or a coding unit (CU).

[0246] In some embodiments, the depth information of the pixels can be used to determine whether the current block is a background block. If the depth of each pixel of the current block is the depth of the background content, it is determined that the current block is a background block; if the depth of one or more pixels of the current block is not the depth of the background content, it is determined that the current block is not an image block corresponding to the background content. For example, in normalized depth, only the depth of the background content is 1, and the depth of the pixels in the image region corresponding to the foreground content is less than 1, so whether the current block is an image block corresponding to the background content can be determined by determining whether the depth values of the pixels of the current block are all 1. If the depth values of the pixels of the current block are all 1, it is determined that the current block is a background block; if the depth values of one or more pixels of the current block are not 1 (less than 1), it is determined that the current block is not a background block.

[0247] In some embodiments, the motion vector of the pixels can be used to determine whether the current block is a background block. For example, in a surveillance video sequence, if the camera pose does not change, the motion vector of the pixels in the background region corresponding to the background content of the surveillance scene does not change, and thus a reference frame can be obtained, wherein the camera pose of the reference frame is the same as that of the current frame; then, it is determined whether the motion vector of each pixel of the current block is 0 according to the reference frame; if the motion vector of each pixel of the current block is 0, it is determined that the current block is a background block; if the motion vector of one or more pixels of the current block is not 0, it is determined that the current block is not a background block. For another example, in a game video sequence, when the camera pose is the same or the camera only undergoes a pure translation motion (without rotation motion), the pixels in the background region corresponding to the background content such as the sky box do not change, and thus a reference frame can be obtained, wherein the camera pose of the reference frame is the same as that of the current frame or only undergoes a translation motion; then, it is determined whether the motion vector of each pixel of the current block is 0 according to the reference frame; if the motion vector of each pixel of the current block is 0, it is determined that the current block is a background block; if the motion vector of one or more pixels of the current block is not 0, it is determined that the current block is not a background block.

[0248] In some embodiments, the frame difference image can be used to determine whether the current block is a background block. For example, in a surveillance video sequence, if the camera pose does not change, the pixels in the background region corresponding to the background content of the surveillance scene do not change, and thus a reference frame can be obtained, wherein the camera pose of the reference frame is the same as that of the current frame; then, the frame difference image of the current frame and the reference frame is calculated, and it is determined whether the pixel value of each pixel of the frame difference image block corresponding to the current block in the frame difference image is 0; if the pixel value of each pixel of the frame difference image block is 0, it is determined that the current block is a background block; if the pixel value of one or more pixels of the frame difference image block is not 0, it is determined that the current block is not a background block. For another example, in a game video sequence, when the camera pose is the same or the camera only undergoes a pure translation motion, the pixels in the background region corresponding to the background content such as the sky box do not change, and thus a reference frame can be obtained, wherein the camera pose of the reference frame is the same as that of the current frame or only undergoes a translation motion; then, the frame difference image of the current frame and the reference frame is calculated, and it is determined whether the pixel value of each pixel of the frame difference image block corresponding to the current block in the frame difference image is 0; if the pixel value of each pixel of the frame difference image block is 0, it is determined that the current block is a background block; if the pixel value of one or more pixels of the frame difference image block is not 0, it is determined that the current block is not a background block.

[0249] In some embodiments, in the above embodiments, a video frame in which the camera pose undergoes a compound motion (translation motion + rotation motion) can also be obtained, and a reference frame of the current frame is obtained by performing a three-dimensional warping (3D warping) on the video frame according to the motion information of the rotation motion.

[0250] In step S1203, if it is determined that the current block is a background block, step S1204 is performed.

[0251] S1204, determining whether the current reconstructed background content includes the background content corresponding to the current block.

[0252] The current reconstructed background content is background content constructed according to reconstructed blocks of background blocks that have been encoded in the current video sequence. In the process of encoding the current video sequence, the encoding end device also performs background content reconstruction on the image blocks that are background blocks.

[0253] In step S1204, if the current reconstructed background content includes the background content corresponding to the current block, step S1205 is performed.

[0254] S1205, selecting an encoding mode of the current block from all encoding modes in the preset encoding mode set.

[0255] The preset encoding mode set includes the target encoding mode.

[0256] In some embodiments, selecting an encoding mode of the current block from all encoding modes in the preset encoding mode set includes: obtaining rate-distortion costs of each encoding mode in the preset encoding mode set, and selecting an encoding mode with the minimum rate-distortion cost as the encoding mode of the current block.

[0257] In some embodiments of the present disclosure, the rate-distortion cost of any encoding mode refers to the rate-distortion cost of encoding the current block based on the encoding mode.

[0258] In some embodiments, in addition to the target encoding mode, the preset encoding mode set can also include one or more encoding modes specified in video coding standards such as H.265 / HEVC, H.266 / VCC, etc.

[0259] In step S1205, if the selected encoding mode of the current block is the target encoding mode, step S1206 is performed.

[0260] S1206, encoding the identification information of the target encoding mode to obtain the encoding data of the current block.

[0261] That is, when the current block is encoded by the target encoding mode, the encoding data of the current block only includes the identification information of the target encoding mode, and does not include the residual data of the current block.

[0262] In some embodiments, when the value of a specified data bit in the encoding data of the current block is used to identify the encoding mode of the current block, the value of the specified data bit can be set to the value corresponding to the target encoding mode.

[0263] In some embodiments, the identification information of the target coding mode includes background identification information (bg_skip) and background encoded identification information (bg_encoded); when the current block is an image block corresponding to background content, the background identification information can be set to 1, i.e., bg_skip = 1; when the current block is not an image block corresponding to background content, the background identification information can be set to 0, i.e., bg_skip = 0; when the current reconstructed background content includes the background content corresponding to the current block, the background encoded identification information can be set to 1, i.e., bg_encoded = 1; when the current reconstructed background content does not include the background content corresponding to the current block, the background encoded identification information can be set to 0, i.e., bg_encoded = 0.

[0264] In some embodiments, encoding the identification information of the target coding mode includes: encoding the background identification information and the background encoded identification information according to whether the current block is an image block corresponding to background content and whether the current reconstructed background content includes the background content corresponding to the current block.

[0265] In some embodiments, if the current block is not an image block corresponding to background content, in step S1205, the current block is encoded by selecting, from the preset coding mode set, a coding mode other than the target coding mode as the coding mode of the current block.

[0266] In some embodiments, if the current block is an image block corresponding to background content and the current reconstructed background content does not include the background content corresponding to the current block, in step S1204, the current block is reconstructed by the encoding end device for background content, and the current reconstructed background content is updated.

[0267] The steps S12203 to S1206 are repeated to complete the encoding of all image blocks of the current video sequence. After obtaining the encoding data of all image blocks of the current video sequence, step S1207 is further executed:

[0268] S1207, generating the encoding data of the current video sequence according to the camera information encoding data and the encoding data of each image block of each video frame of the current video sequence.

[0269] The video coding method provided by some embodiments of the present disclosure comprises: obtaining camera information of each video frame of a current video sequence; encoding the camera information of each video frame of the current video sequence to obtain camera information encoding data; determining whether a current block is a background block; in the case that the current block is a background block, determining whether the current reconstructed background content includes background content corresponding to the current block; in the case that the current reconstructed background content includes the background content corresponding to the current block, selecting an encoding mode of the current block from all encoding modes of a preset encoding mode set; when the selected encoding mode of the current block is a target encoding mode, encoding identification information of the target encoding mode to obtain encoding data of the current block; and finally, after completing the encoding of all image blocks of the current video sequence, generating encoding data of the current video sequence according to the camera information encoding data and the encoding data of each image block of each video frame of the current video sequence. Since the video coding method provided by some embodiments of the present disclosure does not need to add residual data of an image block belonging to a background block and having corresponding background content in the current reconstructed background content to the encoding data, some embodiments of the present disclosure can reduce the data amount of the encoding data and improve the efficiency of video coding and decoding.

[0270] As an extension and refinement of the above-mentioned embodiments, some embodiments of the present disclosure provide another video coding method, which comprises the following steps, as shown in FIG. 13:

[0271] S1301, obtaining camera information of each video frame of a current video sequence.

[0272] S1302, encoding camera intrinsic information of any video frame of the current video sequence to obtain camera intrinsic encoding data.

[0273] The camera intrinsic information is sequence-level information, and the encoding data of the current video sequence only includes one camera intrinsic encoding data.

[0274] Since the camera intrinsic information is stable and unchangeable in the entire video sequence, the camera intrinsic information can be encoded as sequence-level information, so that the encoding data of the current video sequence only includes one camera intrinsic encoding data, thereby further reducing the data amount of the encoding data of the video sequence.

[0275] S1303, respectively encoding camera extrinsic information of each video frame of the current video sequence to respectively obtain camera extrinsic encoding data of each video frame of the current video sequence.

[0276] The camera extrinsic information is frame-level information, and the encoding data of the current video sequence includes camera extrinsic encoding data of each video frame.

[0277] Since the camera poses of different video frames can not be the same, the camera extrinsic information of different video frames can not be the same, and thus the extrinsic information of the video frames needs to be encoded as frame-level information.

[0278] In S1304, it is determined whether the current block is a background block. The current block is any image block obtained by block division on a current frame of a current video sequence.

[0279] In S1304, if it is determined that the current block is not a background block, S1305 is performed.

[0280] In S1305, the rate-distortion cost of the target coding mode is set to a maximum value of the rate-distortion cost.

[0281] In S1304, if it is determined that the current block is a background block, S1306 is performed.

[0282] In S1306, it is determined whether the current reconstructed background content includes the background content corresponding to the current block. The current reconstructed background content is background content constructed according to reconstructed blocks of background blocks that have been encoded in the current video sequence.

[0283] In some embodiments, determining whether the current reconstructed background content includes the background content corresponding to the current block includes steps 1-7.

[0284] In step 1, the mapping region of the current block in the reconstructed background content is determined according to camera information of the current video frame.

[0285] In some embodiments, determining the mapping region of the current block in the reconstructed background content according to the camera information of the current video frame includes: determining mapping points of each corner point (top-left corner point, bottom-left corner point, top-right corner point, bottom-right corner point) of the current block in the reconstructed background content according to the camera information of the current frame, and determining the mapping region of the current block in the reconstructed background content according to the mapping points of each corner point of the current block in the reconstructed background content.

[0286] The implementation of determining the mapping points of each corner point of the current block in the reconstructed background content according to the camera information of the current frame can refer to the implementation of S501, and will not be repeated here to avoid redundancy.

[0287] For example, referring to FIG. 14, the top-left corner point, bottom-left corner point, top-right corner point, and bottom-right corner point of the current block 140 are a, b, c, and d, respectively. The mapping points of the top-left corner point a, bottom-left corner point b, top-right corner point c, and bottom-right corner point d of the current block in the reconstructed background content are A, B, C, and D, respectively, according to the camera information of the current frame. Therefore, it can be determined that the mapping region of the current block in the reconstructed background content is the region where the quadrilateral ABCD is located.

[0288] Step 2, determining whether each pixel in the mapping region has completed reconstruction.

[0289] As described in the foregoing example, when it is determined that the mapping region of the current block in the reconstructed background content is the region in which the quadrilateral ABCD is located, each pixel in the quadrilateral ABCD can be traversed to determine whether each pixel in the quadrilateral ABCD has completed reconstruction.

[0290] In step 2, if each pixel in the mapping region has completed reconstruction, step 3 is performed.

[0291] Step 3, determining that the current reconstructed background content includes the background content corresponding to the current block.

[0292] In step 2, if one or more pixels in the mapping region have not completed reconstruction, step 4 is performed.

[0293] Step 4, determining that the current reconstructed background content does not include the background content corresponding to the current block.

[0294] In step S1306, if the current reconstructed background content does not include the background content corresponding to the current block, step S1305 is returned.

[0295] After step S1305 is performed, or step S1306 determines that the current reconstructed background content includes the background content corresponding to the current block, the following step is further performed.

[0296] S1307, obtaining the rate-distortion cost of each coding mode in the preset coding mode set.

[0297] The preset coding mode set includes the target coding mode, and the rate-distortion cost of any coding mode is the rate-distortion cost of encoding the current block based on the coding mode.

[0298] It should be noted that if step S1305 is performed, the rate-distortion cost of the target coding mode obtained in step S1307 is the maximum value of the rate-distortion cost.

[0299] S1308, selecting the coding mode with the minimum rate-distortion cost in the preset coding mode set to encode the current block to obtain the encoding data of the current block.

[0300] When the selected coding mode of the current block is the target coding mode, selecting the coding mode with the minimum rate-distortion cost in the coding mode set to encode the current block to obtain the encoding data of the current block includes: encoding the identification information of the target coding mode to obtain the encoding data of the current block.

[0301] In some embodiments, the identification information of the target coding mode includes background identification information (bg_skip) and background encoded identification information (bg_encoded); when the current block is an image block corresponding to background content, the background identification information can be set as 1, i.e., bg_skip = 1; when the current block is not an image block corresponding to background content, the background identification information can be set as 0, i.e., bg_skip = 0; when the current reconstructed background content includes background content corresponding to the current block, the background encoded identification information can be set as 1, i.e., bg_encoded = 1; when the current reconstructed background content does not include background content corresponding to the current block, the background encoded identification information can be set as 0, i.e., bg_encoded = 0.

[0302] In some embodiments, encoding the identification information of the target coding mode includes encoding the background identification information and the background encoded identification information according to whether the current block is an image block corresponding to background content and whether the current reconstructed background content includes background content corresponding to the current block. For example, in step S1304, if it is determined that the current block is an image block corresponding to background content, the background identification information is set as 1; if it is determined that the current block is not an image block corresponding to background content, the background identification information is set as 0. In step S1305, if the current reconstructed background content includes background content corresponding to the current block, the background encoded identification information is set as 1; if the current reconstructed background content does not include background content corresponding to the current block, the background encoded identification information is set as 0.

[0303] In some embodiments, if the current block is an image block corresponding to background content and the current reconstructed background content includes background content corresponding to the current block, the rate-distortion cost of the target coding mode is generally the minimum value, and therefore, in step S1308, the target coding mode is selected as the coding mode of the current block.

[0304] In some embodiments, if the current block is not an image block corresponding to background content, in step S1205, other coding modes except the target coding mode are selected from the preset coding mode set as the coding mode of the current block to encode the current block.

[0305] It should be noted that if step S1305 is performed (the rate-distortion cost of the target coding mode is set to the maximum value of the rate-distortion cost), the target coding mode will not be selected as the coding mode of the current block in step S1308, and the target coding mode can only be selected as the coding mode of the current block when step S1306 determines that the current reconstructed background content includes the background content corresponding to the current block and step S1308 is directly performed. Therefore, the essence of the selection scheme of the coding mode is to select the coding mode of the current block from all coding modes in the preset coding mode set when the current block is a background block and the current reconstructed background content includes the background content corresponding to the current block, and to select the coding mode of the current block from the coding modes other than the target mode in the preset coding mode set when the current block is not a background block and / or the current reconstructed background content does not include the background content corresponding to the current block.

[0306] In step S1306, if the current reconstructed background content does not include the background content corresponding to the current block, after step S1308 is performed to select the coding mode with the minimum rate-distortion cost in the preset coding mode set to encode the current block and obtain the encoding data of the current block, steps S1309 and S1310 are performed. That is, the current block is a background block, but the current reconstructed background content does not include the background content corresponding to the current block, so after the encoding data of the current block is obtained by encoding the current block according to the selected coding mode, steps S1309 and S1310 are further performed.

[0307] S1309, decoding the encoding data of the current block to obtain a reconstructed block of the current block.

[0308] Since in the case where the current reconstructed background content does not include the background content corresponding to the current block, the rate-distortion cost of the target coding mode is first set to the maximum value of the rate-distortion cost, and then the coding mode with the minimum rate-distortion cost is selected from the preset coding mode set as the coding mode of the current block, in the case where the current reconstructed background content does not include the background content corresponding to the current block, the coding mode of the current block is not the target coding mode, and the decoding mode for decoding the encoding data of the current block depends on the encoding mode for encoding the current block. Therefore, in some embodiments, the decoding mode corresponding to the encoding mode for encoding the current block in step S1308 can be determined, and the encoding data of the current block is decoded by the decoding mode to obtain the reconstructed block of the current block.

[0309] S1310, updating the current reconstructed background content according to the reconstructed block of the current block.

[0310] In some embodiments, updating the current reconstructed background content according to the reconstructed block of the current block comprises:

[0311] According to the reconstructed block of the current block, the background content corresponding to the current block is obtained, and the background content corresponding to the current block is updated into the current reconstructed background content.

[0312] After the current reconstructed background content is updated according to the reconstructed block of the current block, the current reconstructed background content is increased by the background content corresponding to the current block.

[0313] After the image blocks obtained by block division of each video frame of the current sequence are taken as the current block in some embodiments of the present disclosure one by one, and steps S1304 to S1310 are repeatedly executed, step S1311 is executed. In some embodiments, the specific implementation scheme can include: after step S1308 or S1310 is executed, it is judged whether the current block is the last image block in the image blocks obtained by block division of each video frame of the current sequence, if not, returning to step S1304, if yes, step S1311 is executed.

[0314] S1311, according to the camera internal parameter coding data, the camera external parameter coding data of each video frame of the current video sequence, and the coding data of each image block of each video frame of the current video sequence, the coding data of the current video sequence is generated.

[0315] Similarly, in the above-mentioned video coding scheme, the current block can also be used as a reference block or part of a reference block of other image blocks in the subsequent coding process. Therefore, in some embodiments, the video coding method provided by some embodiments of the present disclosure further includes reconstructing the current block to obtain the reconstructed block of the current block. When the coding mode of the current block is the target coding mode, the implementation manner of reconstructing the current block to obtain the reconstructed block of the current block can include: reconstructing the current block according to the current reconstructed background content and the camera information of the current frame to obtain the reconstructed block of the current block. The implementation manner of reconstructing the current block according to the current reconstructed background content and the camera information of the current frame to obtain the reconstructed block of the current block can refer to the embodiment shown in the above-mentioned FIG. 5, and to avoid redundancy, it will not be repeated here.

[0316] In some embodiments, the video encoding method provided by the above embodiments can also only add the camera information encoding data of the video frame including the image block of the target encoding mode in the encoding data of the current video sequence, instead of adding the camera information encoding data of all video frames in the encoding data of the current video sequence. The specific implementation can be that in the above embodiments, step S1302 is not performed, and after the encoding of all image blocks of all video frames of the current video sequence is completed, it is determined that the current video sequence includes the video frame including the image block of the target encoding mode, and the camera information of the video frame including the image block of the target encoding mode in the current video sequence is encoded to obtain the camera information encoding data, and the encoding data of each image block of each video frame in the current video sequence and the camera information encoding data are used to generate the encoding data of the current video sequence.

[0317] In some embodiments, if it is determined in step S1306 that the current reconstructed background content does not include the background content corresponding to the current block, after the current block is encoded by selecting the encoding mode with the minimum rate-distortion cost in the preset encoding mode set in step S1308 to obtain the encoding data of the current block, the video encoding method provided by some embodiments of the present disclosure further includes:

[0318] Adding background identification information in the encoding data of the current block, the background identification information being used to identify that the current block is a background block.

[0319] In some embodiments, the background identification information can be bg_skip: 1.

[0320] In some embodiments, bg_skip can be a Boolean variable, when the value of bg_skip in the encoding data of the image block is 1, it indicates that the image block is a background block, and when the value of bg_skip in the encoding data of the image block is 0, it indicates that the image block is not a background block.

[0321] In some other embodiments, the encoding data of the image block can also indicate identification information not carrying whether the image block is a background block, and the decoding end determines whether the image block is a background block according to the same background detection algorithm as the encoding end.

[0322] For the above video encoding method, some embodiments of the present disclosure further provide a corresponding video decoding method, which is shown in FIG. 15 and includes the following steps:

[0323] S1501, obtaining the encoding data of the current block.

[0324] S1502, determining the encoding mode of the current block according to the encoding data of the current block.

[0325] In some embodiments, the identification syntax element and the value corresponding to the target coding mode are included in the coding data of the current block, wherein the identification syntax element corresponding to the target coding mode comprises a background identification syntax element and a background coded identification syntax element. In some embodiments, determining that the coding mode of the current block is the target coding mode comprises: when the value of the background identification syntax element is 1 and the value of the background coded identification syntax element is 1, determining that the coding mode of the current block is the target coding mode.

[0326] For example, the identification syntax element corresponding to the target coding mode can be bg_skip and bg_encoded, when the identification syntax element and the value corresponding to the target coding mode are "bg_skip: 1 and bg_encoded: 1", it is determined that the coding mode of the current block is the target coding mode, when the identification syntax element and the value corresponding to the target coding mode are "bg_skip: 0" or "bg_skip: 0", it is determined that the coding mode of the current block is not the target coding mode.

[0327] In some embodiments, the background identification syntax element is used to indicate whether the current block is an image block corresponding to background content, and the background coded identification syntax element is used to indicate whether the current reconstructed content includes background content constructed by a reconstructed block corresponding to the current block when the current block is an image block corresponding to background content. In some embodiments, when the value of the background identification syntax element is 1, it indicates that the current block is an image block corresponding to background content; when the value of the background identification syntax element is 0, it indicates that the current block is not an image block corresponding to background content. In some embodiments, when the value of the background coded identification syntax element is 1, it indicates that the current reconstructed content includes background content constructed by a reconstructed block corresponding to the current block; when the value of the background coded identification syntax element is 0, it indicates that the current reconstructed content does not include background content constructed by a reconstructed block corresponding to the current block.

[0328] In step S1502, if it is determined that the coding mode of the current block is the target coding mode, steps S1503 and S1504 are executed.

[0329] S1503, obtaining current reconstructed background content and camera information of the current frame.

[0330] Wherein, the current reconstructed background content is background content constructed by a reconstructed block of a background block that has been decoded in the current video sequence, the current frame is a video frame to which the current block belongs, and the current video sequence is a video sequence to which the current frame belongs.

[0331] S1504, reconstructing the current block according to the camera information of the current frame and the current reconstructed background content to obtain a reconstructed block of the current block.

[0332] The video decoding method provided in some embodiments of this disclosure, after obtaining the encoded data of the current block, first determines the encoding mode of the current block based on the encoded data of the current block. If the encoding mode of the current block is the target encoding mode, it obtains the camera information of the current frame based on the currently reconstructed background content, and then reconstructs the current block based on the camera information of the current frame and the reconstructed background content to obtain the reconstructed block of the current block. The currently reconstructed background content is the background content constructed based on the reconstructed block of the background block that has already been decoded in the current video sequence. Since the video decoding method provided in some embodiments of this disclosure can reconstruct the current block based on the camera information of the current frame and the currently reconstructed background content, for image blocks with the encoding mode of the target encoding mode, it is not necessary to add the residual data of the image block to the encoded data. Therefore, some embodiments of this disclosure can reduce the amount of encoded data and improve the efficiency of video encoding and decoding.

[0333] As an extension and refinement of the above embodiments, some embodiments of this disclosure also provide another video decoding method. Referring to FIG16, the video decoding method includes the following steps:

[0334] S1601. Obtain the encoded data of the current block.

[0335] S1602. Determine the encoding mode of the current block based on the encoding data of the current block.

[0336] In some embodiments, the encoded data of the current block includes an identifier syntax element corresponding to the target encoding mode and its value, wherein the identifier syntax element corresponding to the target encoding mode includes a background identifier syntax element and a background encoded identifier syntax element. In some embodiments, determining that the encoding mode of the current block is the target encoding mode includes: when the value of the background identifier syntax element is 1 and the value of the background encoded identifier syntax element is 1, determining that the encoding mode of the current block is the target encoding mode.

[0337] For example, the identifier syntax element corresponding to the target encoding mode can be bg_skip and bg_encoded. When the identifier syntax element corresponding to the target encoding mode and its value are "bg_skip:1 and bg_encoded:1", the encoding mode of the current block is determined to be the target encoding mode. When the identifier syntax element corresponding to the target encoding mode and its value are "bg_skip:0" or "bg_skip:0", the encoding mode of the current block is determined not to be the target encoding mode.

[0338] In step S1602, if the encoding mode of the current block is the target encoding mode, then the following steps are performed:

[0339] S1603. Determine whether the camera information of the current frame is stored in the camera information storage space corresponding to the current frame.

[0340] In step S1603, if the camera information of the current frame is stored in the camera information storage space corresponding to the current frame, steps S1604 to S1606 are performed.

[0341] S1604, read the camera information of the current frame from the camera information storage space corresponding to the current frame.

[0342] In step S1604, if the camera information of the current frame is not stored in the camera information storage space corresponding to the current frame, step S1605 is performed.

[0343] S1605, obtain the camera information of the current frame according to the encoded data of the current video sequence.

[0344] In some embodiments, the camera information of the current frame includes camera intrinsic information of the current frame and camera extrinsic information of the current frame, and obtaining the camera information of the current frame according to the encoded data of the current video sequence includes steps ① to ④.

[0345] Step ①, read camera extrinsic encoded data of the current frame from the encoded data of the current video sequence.

[0346] The camera extrinsic encoded data of the current frame is encoded data obtained by encoding the camera extrinsic information of the current frame.

[0347] Step ②, decode the camera extrinsic encoded data of the current frame to obtain the camera extrinsic information of the current frame.

[0348] Step ③, read camera intrinsic encoded data from the encoded data of the current video sequence.

[0349] The camera intrinsic encoded data is encoded data obtained by encoding the camera intrinsic information of any video frame of the current video sequence.

[0350] Step ④, decode the camera intrinsic encoded data to obtain the camera intrinsic information of the current frame.

[0351] S1606, save the camera information of the current frame to the camera information storage space corresponding to the current frame.

[0352] S1607, determine the position coordinates of the mapping points of each pixel of the current block in the current reconstructed background content according to the camera information of the current frame.

[0353] S1608, obtain the reconstructed pixel values of each pixel of the current block according to the position coordinates of the mapping points of each pixel of the current block in the current reconstructed background content and the pixel values of the current reconstructed background content.

[0354] S1609. Obtain the reconstructed block of the current block according to the reconstructed pixel values of the respective pixels of the current block.

[0355] The implementation manners of steps S1607-S1609 can be the same as those of steps S501-S503, and thus the description is not repeated here.

[0356] In the above step S1602, if the coding mode of the current block is not the target coding mode, the following step is performed:

[0357] S1610. Decode the coding data of the current block according to the coding mode of the current block to obtain the reconstructed block of the current block.

[0358] S1611. Determine whether the current block is a background block.

[0359] In some embodiments, the determination of whether the current block is a background block comprises: determining whether the background identification information is included in the coding data of the current block; if the background identification information is included in the coding data of the current block, determining that the current block is a background block; and if the background identification information is not included in the coding data of the current block, determining that the current block is not a background block.

[0360] For example, the determination of whether the background identification information is included in the coding data of the current block comprises: when the value of bg_skip in the coding data of the image block is 1, determining that the background identification information is included in the coding data of the current block; and when the value of bg_skip in the coding data of the image block is 0, determining that the background identification information is not included in the coding data of the current block.

[0361] In some embodiments, the determination of whether the current block is a background block comprises: determining the value of the background identification information in the coding data of the current block; if the value of the background identification information is 1 (bg_skip: 1), determining that the current block is a background block; and if the value of the background identification information is 0 (bg_skip: 0), determining that the current block is not a background block.

[0362] In some embodiments, the determination of whether the current block is a background block comprises: determining, based on a background detection algorithm, whether the reconstructed block of the current block belongs to a background block; if the reconstructed block of the current block is a background block, determining that the current block is a background block; and if the reconstructed block of the current block is not a background block, determining that the current block is not a background block.

[0363] That is, the decoding end determines whether the current block belongs to a background block by using the same background detection algorithm as the encoding end.

[0364] It should be noted that when the decoding end determines whether the current block belongs to the background block through the same background detection algorithm as the encoding end, and the background detection algorithm needs to use auxiliary information such as depth value and motion vector, the encoding data of the current block also needs to carry the encoding data of the auxiliary information such as depth value and motion vector.

[0365] In step S1611, if it is determined that the current block is not a background block, the decoding process of the current block ends, and if it is determined that the current block is a background block, step S1612 is further executed.

[0366] S1612, updating the current reconstructed background content according to the reconstructed block of the current block.

[0367] In some embodiments, updating the current reconstructed background content according to the reconstructed block of the current block includes:

[0368] According to the reconstructed block of the current block, the background content corresponding to the current block is obtained, and the background content corresponding to the current block is fused into the current reconstructed background content.

[0369] After updating the current reconstructed background content according to the reconstructed block of the current block, the current reconstructed background content is increased by the background content corresponding to the current block.

[0370] In some embodiments, some embodiments of the present disclosure provide a video encoding device, which comprises:

[0371] a memory configured to store a computer program;

[0372] a processor configured to, when the computer program is invoked, cause the video encoding device to implement the video encoding method of any of the above embodiments.

[0373] In some embodiments, some embodiments of the present disclosure provide a video decoding device, which comprises:

[0374] a memory configured to store a computer program;

[0375] a processor configured to, when the computer program is invoked, cause the video decoding device to implement the video decoding method of any of the above embodiments.

[0376] In some embodiments, some embodiments of the present disclosure provide a computer readable storage medium, which stores a computer program, and when the computer program is executed by a computing device, the computing device implements the video encoding method or the video decoding method of any of the above embodiments.

[0377] In some embodiments, the present disclosure provides a computer program product, which, when executed on a computer, causes the computer to implement the video encoding method or the video decoding method of any of the above embodiments.

[0378] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present disclosure, and not to limit them; although the foregoing embodiments of the present disclosure have been described in detail, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present disclosure.

[0379] For the convenience of explanation, the above description has been made in combination with specific embodiments. However, the above exemplary discussion is not intended to exhaust or limit the embodiments to the specific forms disclosed above. Various modifications and variations can be derived according to the above teachings. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.

Claims

1. A method of video decoding, the method comprising: The method comprises: obtaining encoding data of a current block; determining an encoding mode of the current block according to the encoding data of the current block; obtaining current reconstructed background content and camera information of a current frame based on the encoding mode of the current block being a target encoding mode, wherein the current reconstructed background content is background content constructed according to reconstructed blocks of background blocks that have been decoded in a current video sequence, the current frame is a video frame to which the current block belongs, and the current video sequence is a video sequence to which the current frame belongs; reconstructing the current block according to the camera information of the current frame and the current reconstructed background content to obtain a reconstructed block of the current block.

2. The method of claim 1, wherein, The determination of the encoding mode of the current block according to the encoding data of the current block comprises: determining a value of background identification information and a value of background encoded identification information based on the encoding data of the current block; determining that the encoding mode of the current block is the target encoding mode based on the value of the background identification information being 1 and the value of the background encoded identification information being 1; determining that the encoding mode of the current block is not the target encoding mode based on the value of the background identification information being 0 or the value of the background encoded identification information being 0.

3. The method according to claim 1 or 2, characterized in that, The method further comprises: decoding the encoding data of the current block according to the encoding mode of the current block to obtain a reconstructed block of the current block based on the value of the background identification information being 1 and the value of the background encoded identification information being 0; updating the current reconstructed background content according to the reconstructed block of the current block.

4. The method of claim 2, wherein, The background identification information is used to indicate whether the current block is an image block corresponding to background content, and the background encoded identification information is used to indicate whether the current reconstructed background content includes background content constructed by a reconstructed block corresponding to the current block when the current block is a background block.

5. The method of claim 1, wherein, The camera information of the current frame comprises camera intrinsic information of the current frame and camera extrinsic information of the current frame, and the obtaining of the camera information of the current frame according to the encoding data of the current video sequence comprises: reading camera extrinsic encoding data of the current frame from the encoding data of the current video sequence, wherein the camera extrinsic encoding data of the current frame is encoding data obtained by encoding the camera extrinsic information of the current frame; decoding the camera extrinsic encoding data of the current frame to obtain the camera extrinsic information of the current frame; reading camera intrinsic encoding data from the encoding data of the current video sequence, wherein the camera intrinsic encoding data is encoding data obtained by encoding camera intrinsic information of any video frame of the current video sequence; decoding the camera intrinsic encoding data to obtain the camera intrinsic information of the current frame.

6. The method of claim 1, wherein, The reconstruction of the current block according to the camera information of the current frame and the current reconstructed background content to obtain the reconstructed block of the current block comprises: determining position coordinates of mapping points of each pixel of the current block in the current reconstructed background content according to the camera information of the current frame; obtaining reconstructed pixel values of each pixel of the current block according to the position coordinates of the mapping points of each pixel of the current block in the current reconstructed background content and pixel values of the current reconstructed background content; and obtaining the reconstructed block of the current block according to the reconstructed pixel values of each pixel of the current block.

7. The method of claim 6, wherein, The camera information of the current frame comprises a vertical field of view angle of the current frame and an extrinsic parameter quaternion of the current frame. The depth of each pixel in the current block is obtained according to the vertical field of view angle. The position coordinates of each pixel in the current block in the camera coordinate system are obtained according to the depth of each pixel in the current block and the pixel coordinates of each pixel in the current block. The direction of the ray corresponding to each pixel in the current block in the background coordinate system is obtained according to the position coordinates of each pixel in the current block in the camera coordinate system and the extrinsic parameter quaternion of the current frame. The intersection of the ray corresponding to each pixel in the current block and the current reconstructed background content is obtained according to the direction of the ray corresponding to each pixel in the current block in the background coordinate system. The position coordinates of the intersection of the ray corresponding to each pixel in the current block and the current reconstructed background content are determined as the position coordinates of the mapping point of each pixel in the current block in the current reconstructed background content.

8. The method of claim 6, wherein, The reconstructed pixel value of each pixel in the current block is obtained according to the position coordinates of the mapping point of each pixel in the current block in the current reconstructed background content and the pixel value of the current reconstructed background content. The mapping point of each pixel in the current block in the current reconstructed background content is pixel-interpolated according to the pixel value of the current reconstructed background content and the position coordinates of the mapping point of each pixel in the current block in the current reconstructed background content, so as to obtain the interpolated pixel value of the mapping point of each pixel in the current block in the current reconstructed background content. The interpolated pixel value of the mapping point of each pixel in the current block in the current reconstructed background content is determined as the reconstructed pixel value of each pixel in the current block.

9. A method of video coding, the method comprising: The method comprises the following steps: Camera information of each video frame of the current video sequence is obtained. The camera information of each video frame of the current video sequence is encoded to obtain camera information encoding data. It is determined whether the current block is a background block. The current block is any image block obtained by block division on the current frame of the current video sequence. Based on the current block being a background block, it is determined whether the current reconstructed background content comprises the background content corresponding to the current block. The current reconstructed background content is background content constructed according to the reconstructed blocks of the background blocks that have been encoded in the current video sequence. Based on the current reconstructed background content comprising the background content corresponding to the current block, a target encoding mode is selected from a preset encoding mode set as the encoding mode of the current block. The identification information of the target encoding mode is encoded to obtain the encoding data of the current block.

10. The method of claim 9, wherein, The camera information encoding data and the encoding data of each image block of each video frame of the current video sequence are used to generate the encoding data of the current video sequence. The method further comprises the following steps: Based on the current reconstructed background content not comprising the background content corresponding to the current block, the encoding mode of the current block is selected from the encoding modes other than the target encoding mode in the preset encoding mode set. encode the current block based on the selected encoding mode to obtain encoded data of the current block; decode the encoded data of the current block to obtain a reconstructed block of the current block; update the current reconstructed background content according to the reconstructed block of the current block.

11. The method of claim 9, wherein, The method further comprises: selecting an encoding mode of the current block from other encoding modes in the preset encoding mode set except the target encoding mode based on that the current block is not a background block; encoding the current block based on the selected encoding mode to obtain encoded data of the current block.

12. The method according to claims 9-11, characterized by, The selecting of the encoding mode of the current block comprises: selecting an encoding mode with a minimum rate-distortion cost as the encoding mode of the current block. The selecting of the encoding mode of the current block from other encoding modes in the preset encoding mode set except the target encoding mode comprises: setting the rate-distortion cost of the target encoding mode as a maximum value of the rate-distortion cost, and then selecting the encoding mode of the current block from all encoding modes in the preset encoding mode set.

13. The method of claim 9, wherein, The determining of whether the current reconstructed background content comprises background content corresponding to the current block comprises: determining a mapping area of the current block in the reconstructed background content according to camera information of the current video frame; determining whether each pixel in the mapping area has been reconstructed; if each pixel in the mapping area has been reconstructed, determining that the current reconstructed background content comprises the background content corresponding to the current block; if one or more pixels in the mapping area have not been reconstructed, determining that the current reconstructed background content does not comprise the background content corresponding to the current block.

14. The method of claim 9, wherein, The camera information of the video frame comprises camera extrinsic parameter information of the video frame and camera intrinsic parameter information of the video frame; and the encoding of the camera information of each video frame in the current video sequence to obtain camera information encoded data comprises: encoding camera intrinsic parameter information of any video frame in the current video sequence to obtain camera intrinsic parameter encoded data; respectively encoding camera extrinsic parameter information of each video frame in the current video sequence to respectively obtain camera extrinsic parameter encoded data of each video frame in the current video sequence.

15. The method of claim 10, wherein, The identification information of the target encoding mode comprises background identification information and background encoded identification information; and the encoding of the identification information of the target encoding mode comprises: setting a value of the background identification information as 1; setting a value of the background encoded identification information as 1.

16. The method of claim 9, wherein, After the encoding of the identification information of the target encoding mode to obtain the encoded data of the current block, the method further comprises: reconstructing the current block according to the current reconstructed background content and the camera information of the current frame to obtain a reconstructed block of the current block.

17. A method of video decoding, comprising: comprises: obtaining encoded data of the current block; determining an encoding mode of the current block according to the encoded data of the current block; The encoding mode of the current block is a target encoding mode, and the reconstructed background content and camera information of the current frame are obtained; the reconstructed background content is background content of a current video sequence reconstructed according to background content encoding data of the current video sequence, and the background content encoding data of the current video sequence is encoding data obtained by encoding the background content of the current video sequence; the current frame is a video frame to which the current block belongs, and the current video sequence is a video sequence to which the current frame belongs; The current block is reconstructed according to the camera information of the current frame and the reconstructed background content, to obtain a reconstructed block of the current block.

18. The method of claim 17, wherein, Determining the encoding mode of the current block according to the encoding data of the current block includes: Determining a value of identification information of the current block based on the encoding data of the current block; Based on the value of the identification information of the current block being 1, it is determined that the encoding mode of the current block is the target encoding mode.

19. The method of claim 17, wherein, The camera information of the current frame includes camera intrinsic parameter information of the current frame and camera extrinsic parameter information of the current frame; and the camera information of the current frame is obtained from the encoding data of the current video sequence, including: Reading camera extrinsic parameter encoding data of the current frame from the encoding data of the current video sequence, the camera extrinsic parameter encoding data of the current frame being encoding data obtained by encoding the camera extrinsic parameter information of the current frame; Decoding the camera extrinsic parameter encoding data of the current frame to obtain the camera extrinsic parameter information of the current frame; Reading camera intrinsic parameter encoding data from the encoding data of the current video sequence, the camera intrinsic parameter encoding data being encoding data obtained by encoding the camera intrinsic parameter information of any video frame of the current video sequence; Decoding the camera intrinsic parameter encoding data to obtain the camera intrinsic parameter information of the current frame.

20. The method of claim 17, wherein, The current block is reconstructed according to the camera information of the current frame and the reconstructed background content, to obtain a reconstructed block of the current block, including: Determining position coordinates of mapping points of each pixel of the current block in the reconstructed background content according to the camera information of the current frame; Obtaining reconstructed pixel values of each pixel of the current block according to the position coordinates of the mapping points of each pixel of the current block in the reconstructed background content and pixel values of the reconstructed background content; Obtaining the reconstructed block of the current block according to the reconstructed pixel values of each pixel of the current block.

21. The method of claim 20, wherein, The camera information of the current frame includes a vertical field of view angle of the current frame and extrinsic quaternion of the current frame; and the position coordinates of the mapping points of each pixel of the current block in the reconstructed background content are determined according to the camera information of the current frame, including: Obtaining depths of each pixel in the current block according to the vertical field of view angle; Obtaining position coordinates of each pixel in the current block in a camera coordinate system according to the depths of each pixel in the current block and pixel coordinates of each pixel in the current block; Obtaining directions of rays corresponding to each pixel in the current block in a background coordinate system according to the position coordinates of each pixel in the current block in the camera coordinate system and the extrinsic quaternion of the current frame; any pixel corresponds to a ray starting from a coordinate origin of the camera coordinate system and passing through the pixel, and the background coordinate system is a coordinate system of the reconstructed background content. According to a direction of a ray corresponding to each pixel in the current block in the background coordinate system, an intersection point of the ray corresponding to each pixel in the current block and the reconstructed background content is obtained; Position coordinates of the intersection point of the ray corresponding to each pixel in the current block and the reconstructed background content are determined as position coordinates of a mapping point of each pixel in the current block in the reconstructed background content.

22. The method of claim 20, wherein, The method further comprises: According to the position coordinates of the mapping point of each pixel in the current block in the reconstructed background content and pixel values of the reconstructed background content, a reconstructed pixel value of each pixel in the current block is obtained. According to the pixel values of the reconstructed background content and the position coordinates of the mapping point of each pixel in the current block in the reconstructed background content, a pixel interpolation is performed on the mapping point of each pixel in the current block in the reconstructed background content to obtain an interpolated pixel value of the mapping point of each pixel in the current block in the reconstructed background content.

23. A method of video coding, the method comprising: The method further comprises: The background content of the current video sequence and camera information of each video frame of the current video sequence are obtained. The background content is encoded to obtain background content encoding data. The camera information of each video frame of the current video sequence is encoded to obtain camera information encoding data. It is determined whether the current block is an image block corresponding to the background content. Based on the current block being the image block corresponding to the background content, it is determined that an encoding mode of the current block is a target encoding mode, and identification information of the target encoding mode is encoded to obtain encoding data of the current block. According to the background content encoding data, the camera information encoding data and encoding data of each image block of each video frame of the current video sequence, encoding data of the current video sequence is generated.

24. The method of claim 23, wherein, The method further comprises: Based on the current block not being the image block corresponding to the background content, an encoding mode of the current block is selected from a preset encoding mode set, wherein the encoding mode of the current block is not the target encoding mode. Based on the encoding mode of the current block, the current block is encoded to obtain encoding data of the current block.

25. The method of claim 24, wherein, The encoding mode of the current block is an encoding mode with a minimum rate-distortion cost in the preset encoding mode set. The method further comprises: A rate-distortion cost of the target encoding mode is set as a maximum value of rate-distortion costs. An encoding mode with a minimum rate-distortion cost is selected from the preset encoding mode set as the encoding mode of the current block.

26. The method of claim 23, wherein, The camera information of each video frame of the current video sequence comprises camera extrinsic parameter information of the video frame and camera intrinsic parameter information of the video frame. The camera intrinsic parameter information of any video frame of the current video sequence is encoded to obtain camera intrinsic parameter encoding data. Camera extrinsic information of each video frame of the current video sequence is encoded respectively to obtain camera extrinsic encoding data of each video frame of the current video sequence.

27. The method of claim 23, wherein, After encoding the identification information of the target coding mode to obtain the encoding data of the current block, the method further comprises: decoding the background content encoding data to obtain reconstructed background content; reconstructing the current block according to the reconstructed background content and camera information of the current frame to obtain a reconstructed block of the current block.

28. An apparatus for video decoding, the apparatus comprising: comprises: a memory configured to store a computer program; a processor configured to, when the computer program is invoked, enable the video decoding device to implement the video decoding method in any one of claims 1-8.

29. An apparatus for video encoding, the apparatus comprising: comprises: a memory configured to store a computer program; a processor configured to, when the computer program is invoked, enable the video decoding device to implement the video decoding method in any one of claims 9-16.

30. An apparatus for video decoding, the apparatus comprising: comprises: a memory configured to store a computer program; a processor configured to, when the computer program is invoked, enable the video decoding device to implement the video decoding method in any one of claims 17-22.

31. An apparatus for video encoding, the apparatus comprising: comprises: a memory configured to store a computer program; a processor configured to, when the computer program is invoked, enable the video decoding device to implement the video decoding method in any one of claims 23-27.

Citation Information

Patent Citations

  • Grouping method in frequency domain distributed video coding

    CN101835044A

  • Video processing method, system and device, readable storage medium and electronic equipment

    CN112004114A

  • Adaptive motion estimation and mode decision apparatus and method for H.264 video codec

    US20060039470A1

  • Systems and methods for low complexity encoding and background detection

    US20150264367A1

  • Video encoding / decoding method, device and system, and storage medium

    WO2022257134A1