Video encoding method, device, computing device, and storage medium

By using the generated information of the video generation tool to guide the encoding block division, different strategies are adopted for the background and prospects, the problem of high computing in the existing technology is solved, and an efficient video encoding process is realized.

CN117201796BActive Publication Date: 2025-07-22SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311021953.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-14
Publication Date
2025-07-22
Estimated Expiration
2043-08-14

AI Technical Summary

Technical Problem

The existing video encoding technology requires multiple attempts when dividing encoding blocks, resulting in large amounts of calculations and high time consumption, making it difficult to adapt to the needs of high data volume and high transmission rates.

Method used

Use the generated information provided by the video generation tool, such as the location and material information of the foreground object, to guide the division of coded blocks, and use different division methods to target the background and foreground areas to reduce invalid calculations.

Benefits of technology

By reducing the calculation amount of encoding block division, the efficiency and transmission rate of video encoding are improved, and the picture clarity and quality are maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117201796B_ABST
    Figure CN117201796B_ABST
Patent Text Reader

Abstract

The present application provides a video encoding method, a video encoding device, a computing device, and a computer-readable storage medium. The method includes: obtaining generation information of a video; dividing a first frame of the video according to the generation information to obtain a plurality of coding blocks, the plurality of coding blocks of the first frame including a current block, and a target block being included in the plurality of coding blocks of the first frame or a second frame adjacent to the first frame of the video; predicting predicted pixel information of the target block according to current pixel information of the current block; and encoding the video according to the predicted pixel information. According to the present application, it is possible to quickly implement the division of coding blocks with less computation, thereby accelerating the video encoding process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and particularly to a video encoding method, a video encoding device, a computing device, and a computer-readable storage medium. Background Art

[0002] In the field of video encoding and compression technologies, there are intra-frame prediction and inter-frame prediction technologies. Both of these technologies need to divide video frames into multiple coding blocks for better predictive encoding. However, the coding block division process in the prior art often requires multiple attempts to obtain the optimal division method, which requires a large amount of computation, resulting in an increase in computational cost and time consumption, and it is becoming increasingly difficult to adapt to the current situation where the amount of video data is constantly increasing and the transmission rate requirement is getting higher and higher. Therefore, there is an urgent need in this field for a technology that can quickly implement the coding block division of video frames with less computation, thereby accelerating the video encoding process. Summary of the Invention

[0003] To this end, this application is committed to providing a video encoding method, a video encoding device, a computing device, and a computer-readable storage medium, which can quickly implement the coding block division of video frames with less computation, thereby accelerating the video encoding process.

[0004] In one aspect, this application provides a video encoding method, including: obtaining the generation information of a video; dividing the first frame of the video according to the generation information to obtain multiple coding blocks, where the multiple coding blocks of the first frame include a current block, and a target block is included in the multiple coding blocks of the first frame or in a second frame adjacent to the first frame of the video; predicting the predicted pixel information of the target block according to the current pixel information of the current block; and encoding the video according to the predicted pixel information.

[0005] According to this aspect, determining the coding block division method of video frames through the generation information of the video is beneficial to directly and quickly perform optimal division according to the nature of video content (such as material, foreground and background, etc.), avoiding the increase in computation caused by repeated attempts, greatly improving the encoding efficiency of generative videos, and enabling generative videos to have clearer pictures and higher transmission rates.

[0006] In a particular embodiment of this application, the generation information includes the position information and material information of the foreground object in the first frame.

[0007] According to this embodiment, the position information and material information of the foreground object often cannot be directly obtained from the video screen itself, but the video generation tool can often provide a distinction between the foreground object and the background object, thereby determining the position information of the foreground object. In addition, the generation tool can determine the material information of the foreground when rendering the surface color and material of the foreground object. Based on these two types of information, it can help determine different division strategies for the coding block, making the subsequent prediction and coding process more accurate and efficient.

[0008] In a particular embodiment of the present application, the first frame of the video is divided according to the generated information to obtain multiple coding blocks, including: determining the foreground area and the background area in the first frame according to the position information of the foreground object; dividing the foreground area according to a first mode, and dividing the background area according to a second mode, wherein the coding blocks in the first mode are smaller than the coding blocks in the second mode.

[0009] According to this embodiment, different division methods are used for the foreground area and the background area, which is beneficial to utilize the difference in the change range of the displayed content between the foreground area and the background area, and adopt different division methods to improve the prediction accuracy and compression efficiency.

[0010] In a particular embodiment of the present application, the first frame of the video is divided according to the generated information to obtain multiple coding blocks, and it also includes: determining the edge area of the foreground object in the first frame according to the position information of the foreground object; and dividing the edge area according to the first mode.

[0011] According to this embodiment, a more detailed coding block division method is adopted for the edge area, which can utilize the law of changes that are more likely to occur in the edge area, more clearly retain the display details of the edge of the foreground object, and maintain the accuracy and realism of the image content in the edge area as much as possible.

[0012] In a particular embodiment of the present application, the first frame of the video is divided according to the generated information to obtain multiple coding blocks, including: determining the foreground area and the background area in the first frame according to the position information of the foreground object; judging whether the foreground object is a rigid object or a flexible object according to the material information of the foreground object; if the foreground object is a rigid object, dividing the foreground area according to the third mode; if the foreground object is a flexible object, dividing the foreground area according to the fourth mode, wherein the coding blocks in the third mode are larger than the coding blocks in the fourth mode.

[0013] According to this embodiment, different division methods are adopted for rigid objects and flexible objects, which is conducive to utilizing the different display rules of the two objects in the video screen and compressing the video data more efficiently, while maintaining the display details of the flexible objects as much as possible and improving the clarity and restoration of the picture.

[0014] In a particular embodiment of the present application, the video includes a virtual scene video generated by a video generation tool, and generation information of the video is obtained, including: obtaining the generation information from the video generation tool.

[0015] According to this embodiment, virtual scene videos abound in today's video products, including a large number of virtual scene videos generated through manual production and special effects technology in games, movies, and TV dramas. Encoding acceleration for virtual scene videos is conducive to greatly promoting the dissemination and wide use of video content. In addition, obtaining generation information from the video generation tool for producing the virtual scene video is conducive to directly and quickly obtaining information helpful for video encoding from the source, improving the accuracy and timeliness of information acquisition.

[0016] In a particular embodiment of the present application, the video generation tool includes a renderer or an AI tool.

[0017] According to this embodiment, renderers and AI tools are common video generation tools, which can provide a large amount of hidden information about video content, thus facilitating quickly determining the partitioning method by referring to this information during coding block partitioning.

[0018] In a particular embodiment of the present application, the coding block includes one or more of a macroblock, a sub-block, a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU).

[0019] According to this embodiment, there are various specific forms of coding blocks in the art, and the coding block partitioning method provided by the present application can be adopted, which is conducive to accelerating coding block partitioning at each coding level, thus greatly improving the coding compression rate as a whole.

[0020] On the other hand, the present application provides a video coding device, including: an acquisition module for acquiring generation information of the video; a partitioning module for partitioning a first frame of the video according to the generation information to obtain a plurality of coding blocks, where the plurality of coding blocks of the first frame include a current block, and a target block is included in the plurality of coding blocks of the first frame or a second frame adjacent to the first frame of the video; a prediction module for predicting predicted pixel information of the target block according to current pixel information of the current block; and an encoding module for encoding the video according to the predicted pixel information.

[0021] On the other hand, the present application provides a computing device, characterized in that the computing device includes a processor and a memory, and the processor is configured to execute a computer program stored in the memory to implement the above-mentioned video coding method.

[0022] On the other hand, the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and the computer program is used to execute the above-mentioned video coding method.

[0023] On the other hand, the present application provides a computer program product, including program code, which enables a computer to implement the above video encoding method when the computer runs the computer program product.

[0024] Any of the above provided video encoding devices, computing devices, computer-readable storage media or computer program products are all used to execute the video encoding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding solutions in the corresponding method provided above, which will not be elaborated here. Description of the Drawings

[0025] Hereinafter, specific embodiments of the present application will be described in detail in conjunction with the drawings, where:

[0026] Figure 1 Schematic diagram of an application scenario of a video encoding method according to an embodiment of the present application;

[0027] Figure 2 Showing according to Figure 1 Schematic diagram of inter-frame prediction of a video encoding method according to an embodiment;

[0028] Figure 3 Showing according to Figure 1 Schematic diagram of an encoding block partitioning method of a video encoding method according to an embodiment;

[0029] Figure 4 Schematic diagram of a flowchart of a video encoding method according to another embodiment of the present application;

[0030] Figure 5 Showing according to Figure 4 Schematic diagram of a partitioning step of a video encoding method according to an embodiment;

[0031] Figure 6 Showing according to Figure 4 Another schematic diagram of a partitioning step of a video encoding method according to an embodiment;

[0032] Figure 7 Showing according to Figure 4 Another schematic diagram of a partitioning step of a video encoding method according to an embodiment;

[0033] Figure 8 Showing according to Figure 4 Another schematic diagram of a partitioning step of a video encoding method according to an embodiment;

[0034] Figure 9 Showing according to Figure 4 Another schematic diagram of a partitioning step of a video encoding method according to an embodiment;

[0035] Figure 10 Showing according to Figure 4Another schematic diagram of the partitioning step of the video encoding method of the embodiment;

[0036] Figure 11 Schematic diagram showing the structure of a video encoding apparatus according to an embodiment of the present application;

[0037] Figure 12 Schematic diagram showing the structure of a computing device according to an embodiment of the present application. Detailed implementation manners

[0038] In order to enable those skilled in the art to more clearly understand the concepts and ideas of the present application, the present application will be described in detail below in conjunction with specific embodiments. It should be understood that the embodiments given herein are only a part of all possible embodiments of the present application. After reading the specification of the present application, those skilled in the art are capable of making improvements, modifications, or substitutions to part or all of the following embodiments, and these improvements, modifications, or substitutions are also included within the scope of protection required by the present application.

[0039] In this document, the terms "a", "an" and other similar words do not intend to indicate that there is only one such thing, but rather indicate that the relevant description is only directed to one of such things, and such things may have one or more. In this document, the terms "comprise", "include" and other similar words are intended to represent logical relationships, and should not be regarded as representing spatial structural relationships. For example, "A includes B" is intended to indicate that logically B belongs to A, rather than indicating that B is located inside A in terms of space. Additionally, the meanings of the terms "comprise", "include" and other similar words should be regarded as open-ended rather than closed-ended. For example, "A includes B" is intended to indicate that B belongs to A, but B does not necessarily constitute all of A, and A may also include other elements such as C, D, E, etc.

[0040] In this document, the terms "first", "second" and other similar words do not intend to imply any order, quantity, or importance, but are only used to distinguish different elements. In this document, the terms "embodiment", "the present embodiment", "an embodiment", "one embodiment" do not indicate that the relevant description is only applicable to a specific embodiment, but rather indicate that these descriptions may also be applicable to one or more other embodiments. Those skilled in the art should understand that in this document, any description made for a certain embodiment can be substituted, combined, or otherwise combined with the relevant descriptions in one or more other embodiments, and the new embodiments generated by such substitution, combination, or other combination are easily conceivable by those skilled in the art and fall within the scope of protection of the present application.

[0041] In various embodiments of the present application, video may refer to various technologies that capture, record, process, store, transmit, and reproduce a series of static images in the form of electrical signals. For example, video utilizes the principle of the human eye's visual persistence. By playing a series of pictures (also called frames), it gives people the feeling of motion. In some embodiments, the video may include generative video. Generative video may refer to a video different from traditional video. Traditional video is generated by video capture devices such as cameras, mobile phones with camera functions, video recorders, etc., by shooting scenes in the real world; generative video, in contrast to traditional video, is not generated by a video capture device. It is usually generated by computer software such as renderers, AI, etc., and is composed of corresponding continuous pictures generated according to user needs. The most common examples are videos such as games and movies.

[0042] In various embodiments of the present application, video encoding may refer to a method of converting a file in the original video format into another video format file through compression technology. Video is a sequence of continuous images composed of continuous frames. Since the similarity between consecutive frames is extremely high, in order to facilitate storage and transmission, it is necessary to encode and compress the original video to remove redundancy in the spatial and temporal dimensions. Video image data has strong correlations, that is, there is a large amount of redundant information. The redundant information can be divided into spatial redundant information and temporal redundant information. Compression technology is to remove the redundant information in the data. Compression technology includes intra-frame image data compression technology, inter-frame image data compression technology, and entropy coding compression technology. In this article, "video compression" and "video encoding" can be interchanged in some scenarios.

[0043] In the art, the video images of games and most movies are generated by a computer through a renderer. Especially in non-local rendering scenarios such as cloud rendering, after the renderer generates a video, it needs to be compressed by a video encoder and then pushed to the terminal for decoding and display. The non-local rendering scenario generally follows the following workflow: 1) The renderer generates video images according to the set program; 2) The video images are encoded (compressed) by the video encoder, greatly reducing the network transmission bandwidth; 3) The compressed encoded video stream is transmitted to the terminal user device through the network; 4) The terminal user device decodes the compressed encoded video stream through the decoder, and after decoding, the video images generated by the renderer are displayed through the display.

[0044] The amount of data in the video before compression is extremely large. Taking 1080P, which is widely used in current scenarios, as an example, each pixel contains luminance and chrominance, represented by 12 bits in total, and the frame rate is 60 (that is, 60 frames are output per second). In this way, the required network transmission bandwidth is about 1.39 Gbps. The bandwidth requirement for 4K videos will be even greater. It is very difficult to meet such a large Internet bandwidth requirement. A key technology in video stream applications is video coding, also known as video compression. The device that completes the video compression function is the video encoder. Similar to text compression, the purpose is to remove redundant data in the video data and reduce the amount of data after encoding and compression. Video compression is sometimes a lossy compression. Although lossy video compression will introduce a certain degree of video image distortion, these distortions are usually difficult for the human eye to detect.

[0045] In a video, there is usually a strong spatial correlation among pixels within the adjacent range of a single image (also known as a single frame). There is a strong temporal correlation among consecutive image data. Video compression coding technology mainly adopts intra-frame prediction and inter-frame prediction methods, using the already encoded pixels of an image to predict neighboring pixels in space and time, thereby effectively compressing the redundancy of video image pixels in the time domain and spatial domain.

[0046] As the video image size increases, such as 4K and 8K images, in order to perform better compression coding, the video encoder usually divides the image into many small blocks for compression coding, such as the 16x16 pixel-sized macroblocks in AVC (advanced video coding). To more efficiently compress different textures and motion change details in the video scene, HEVC (high efficiency video coding) defines coding tree units (CTUs) with a unit of 64x64 pixels and coding units (CUs) of different sizes for images. The CU division in the VVC (versatile video coding) standard is more flexible and rich.

[0047] During the video coding process, it is necessary to try different CU division depths. As the video image size increases, the requirements for video coding latency increase, and the predictions between CUs are interdependent, resulting in increasing pressure on the video encoder.

[0048] A technology in this field, when compressing in units of CTUs, for each CU partition depth, namely the four partition methods of 64x64, 32x32, 16x16, and 8x8, motion estimation is respectively performed to obtain motion vectors, and based on the motion vectors, inter-frame prediction, discrete cosine transform (DCT), quantization, bit number evaluation entropy coding, and distortion calculation are carried out, and then comparison is performed layer by layer according to the rate-distortion optimized (RDO) algorithm. For a 16x16 pixel region, there are four partition methods: 4 8x8 CUs, 2 4x8 CUs, 2 8x4 CUs, or 1 16x16 CU. Motion estimation is respectively performed on these four partition methods to obtain motion vectors, and then prediction is performed based on the motion vectors, and finally the one with the minimum rate distortion is selected as the partition. The above-mentioned similar method is used to evaluate various different partition methods upwards.

[0049] This technology can be used for traditional videos and generative videos. The problem is that it is necessary to try all CU partition methods, and then perform motion estimation, prediction calculation, DCT, quantization, and entropy coding processes. Finally, the final partition method is obtained through rate distortion calculation. The amount of calculation is huge, and finally only one partition method is selected, resulting in a large amount of wasted calculation. Especially for 4K and 8K videos, the latest VCC coding standard is difficult to meet the needs of real-time video encoding and decoding.

[0050] Another technology in this field is used for generative video coding. The rendering engine will provide the motion vector of each pixel to the video encoder. The encoder respectively calculates the similarity of the motion vectors of each pixel within different partition methods, such as calculating the variance of the 256 pixel motion vectors within a 16x16 pixel block, the variance of the 128 pixel motion vectors within a 16x8 pixel block, and the variance of the 64 pixel motion vectors within an 8x8 pixel block. Finally, the block size of the block with the minimum variance is selected as the final partition method.

[0051] This technology can be used for generative videos. The problem is that although it is not necessary to perform motion estimation and RDO calculation processes, different CU partition attempts still need to be made, the amount of calculation is still huge, and there is still wasted calculation.

[0052] Some embodiments of this application are dedicated to solving such a problem: for generative videos, how to reduce the attempt of the CU (or other blocks) partition size by means of the information provided by the renderer, so as to reduce the waste of the encoder computing power.

[0053] In some embodiments of the present application, for generative videos, while minimizing the loss of video quality caused by compression, the waste of computing power of the video encoder is reduced, and at the same time, the encoding delay of the encoder is reduced to meet the requirements of high-quality and low-latency encoding. Specifically, current video encoders need to try different CU partitioning methods and finally select one method as the final partitioning method. A large amount of calculations are required for different CU partitionings. Since only one partitioning method is selected for the final video bitstream, there are a large number of ineffective calculations in the encoder. In fact, for the background area that occupies the main range in the video frame, the motion vectors of two consecutive frames are usually the same, and the motion vectors of foreground objects may also be the same. In theory, it is not necessary to try all partitioning methods. Especially for generative videos, rendering tools and other frame generation tools know which pixels belong to the background area and which belong to the foreground area. Therefore, the embodiments of the present application assist the video encoder in encoding through the information of the renderer, minimizing ineffective calculations to the greatest extent.

[0054] Some embodiments of the present application are applied to generative video encoders. Using video generators such as renderers and AI tools, videos and video-related information are generated, including but not limited to layer information, the coordinates of foreground moving object pixels in the frame, the materials of foreground objects such as rigid bodies and flexible object information, etc. The video encoder performs different partitioning processes on the background area and foreground moving objects according to the layer and moving object coordinate information, that is, quickly realizes inter-frame mode encoding, thereby saving the computing power of the video encoder and reducing the encoding delay of the video encoder.

[0055] In some embodiments of the present application, the CU partitioning for inter-frame prediction is performed with a fixed CU size according to the background area and foreground moving objects. Compared with trying four different CU partitioning methods and selecting one final partitioning method through the RDO algorithm, the embodiments of the present application can reduce the waste of computing power by 3 times. At the same time, for the background area, which is often not an important area of the video, it has little impact on the final user experience. In addition, in some embodiments of the present application, different CU partitioning methods are used for foreground moving objects according to rigid and flexible materials. The motion vectors of rigid objects are usually the same, and the motion vectors of different parts of flexible objects are usually different. For rigid objects, large CU partitioning methods such as 64x64 and 32x32 can be used, and for flexible objects, 8x8 CU partitioning is used to save computing power to the greatest extent while ensuring the quality of the video frame.

[0056] Figure 1 The schematic diagram of the application scenario of the video encoding method according to an embodiment of the present application is shown.

[0057] This embodiment includes a video generator and a video encoder. The video encoder quickly compresses the generated video according to the video generated by the video generator and related information, specifically involving the quick partitioning of intra-frame or inter-frame coding blocks. In this embodiment, a video generation tool generates a video and related generation information, and both the video and the generation information are input into the video encoder. The video encoder can be set and adjusted according to encoder parameters. The video encoder receives the input from the video generation tool, compresses the received video according to the content or parameters of the generation information, and finally generates a compressed bitstream for data transmission or further processing. In this embodiment, the video used for video coding is the video generated by the video generation tool, rather than, for example, a natural scene video directly captured by a shooting device. The encoding process of the video encoder not only requires the original video data but also the generation information that can represent information related to the video content. The information about the video content provided by the generation information plays an important role in accelerating the video coding process. The video encoder can perform compression coding on the video through intra-frame prediction and inter-frame prediction. Both intra-frame and inter-frame prediction require partitioning the video frames into coding blocks.

[0058] As an example, Figure 2 shows a Figure 1 schematic diagram of the inter-frame prediction process in the video coding method according to the Figure 2 embodiment. As shown, two consecutive frames of the video are shown. The dashed boxes in the figure are only for illustration and not the content of the frames. In these two consecutive frames, only the position of object 1 has changed. Then, the pixels in the area of object 1 in the second frame (the dashed square area) can be predicted by the pixels in the area of object 1 in the previous frame (the dashed square area). The difference between the corresponding pixels in the front and back areas, plus the motion vector information of object 1, can be used to perform predictive coding on the second frame. When performing inter-frame prediction on the latter frame, it is necessary to search in the previous frame to find the matching area. This search process is the motion estimation (ME), and the position change of the object or pixel between the front and back frames is the motion vector (MV).

[0059] As an example, Figure 3 shows a Figure 1 schematic diagram of the coding block partitioning process in the video coding method according to the Figure 3As shown, in this embodiment, a CTU with 64x64 pixels is defined for an image. Each image in the video is divided into several non-overlapping CTUs. Inside the CTU, a cyclic hierarchical structure based on a quadtree is adopted. For example, the top layer is a 64x64 pixel unit, which can be divided into the next layer, divided into 4 CU with a size of 32x32 pixels. Each 32x32 pixel size can be further divided into 16x16 coding blocks until the smallest 8x8 coding unit.

[0060] In this embodiment, both intra-frame prediction and inter-frame prediction can be performed in units of CU in the order from left to right and from top to bottom. Subtract the predicted value from the original pixel value of the video within the CU to obtain the residual. After DCT, the transformation result is quantized, and the quantization value and related syntax elements, such as the CU division depth, the motion vector corresponding to the pixel block, etc., are entropy encoded to obtain the compressed video stream. At the same time, the quantization result is inverse quantized and inverse discrete cosine transformed to obtain the decoded residual value, which is then added to the predicted value to obtain the pixel reconstruction value. The pixel reconstruction value, CU division depth, etc. of the previous CU block need to be used as the input for the prediction of the next CU. Therefore, in some cases, all CUs cannot be compressed and encoded in parallel simultaneously.

[0061] Figure 4 The flowchart showing a video encoding method according to an embodiment of the present application is shown.

[0062] According to this embodiment, the video encoding method includes steps S410 to S440, and each step is described in detail below.

[0063] S410. Obtain the generation information of the video.

[0064] In this embodiment, the generation information may refer to the information related to subsequent encoding operations generated during the generation process of the generative video. In some embodiments, the generation information of the video is different from the information possessed by the video image itself (such as the pixel value information of the image or the image texture, color, etc. that can be directly reflected by the pixel value). The generation information is provided by the tool for generating the video and is information that cannot or is difficult to determine only based on the image content or pixel value information of the video.

[0065] As an example, the generation information includes the position information and material information of the foreground object in the first frame. For example, the generation information includes layer information and the coordinate information of the foreground moving object pixels in the picture.

[0066] In this example, the position information of the foreground object may refer to the position or coordinate information of the foreground object in the video picture, and the material information of the foreground object may refer to the material given to the fictional foreground object by the video producer during video production, such as metal, wood, human skin, etc.

[0067] As an example, the video includes a virtual scene video generated by a video generation tool. At this time, in order to obtain the generation information of the video, the generation information can be obtained from the video generation tool.

[0068] In this example, the video generation tool may refer to a tool for generating a video according to manual video production means such as user modeling, rendering, and color correction of a video. Such a video generation tool is different from shooting, video recording, or filming tools for shooting actual objective objects. The virtual scene video may refer to a video in which the displayed objects do not actually exist, but are fictionalized by manual image editing means such as drawing, modeling, and rendering. Of course, the virtual scene video generated by the video generation tool may include actually existing objects. For example, a video in which virtual objects and real objects are mixed in an AR (augmented reality) glasses also belongs to the virtual scene video defined in this application.

[0069] In this example, in order to obtain the generation information of the video, relevant information or data can be directly exported from the video generation tool for subsequent video encoding. For example, some video generation tools can provide the three-dimensional coordinate information of the objects, from which it can be determined which objects are foreground objects and which objects are background objects. For another example, some video generation tools can provide the material information of the objects, so that the hardness or rigidity can be judged according to the material information, and different coding block division methods are adopted for different rigid materials for encoding.

[0070] As an example, the video generation tool includes a renderer or an AI (artificial intelligence) tool.

[0071] In this example, the renderer may refer to a software or hardware tool that converts a 3D (3-dimension) model into a 2D (2-dimension) image. It can convert the geometric shapes, textures, lighting, and materials and other information in the 3D model into a 2D image for easy display on the screen or printing. The renderer is usually used in fields such as movies, games, and architectural design to create realistic visual effects. For example, the renderer may refer to a software for processing and displaying three-dimensional graphics. It can draw a three-dimensional scene on a two-dimensional display device, such as a monitor, mobile phone, or TV. Its main function is to convert a digital model into an image that can be displayed by the device.

[0072] In this example, the AI tool may refer to a software tool that automatically generates a video using artificial intelligence technology. The tool can intelligently generate a video work through deep learning technology and algorithms according to the user's drawing operations or according to the picture materials, video materials, added text, voiceovers, background music, etc. provided by the user.

[0073] S420. Divide the first frame of the video according to the generated information to obtain a plurality of coding blocks. The plurality of coding blocks of the first frame include the current block, and the target block is included in the plurality of coding blocks of the first frame or the second frame adjacent to the first frame of the video.

[0074] In this embodiment, a frame of the video may refer to an image in a continuous image sequence that constitutes the video. The second frame adjacent to the first frame may refer to two video frames that are closely adjacent in the continuous image sequence of the video, or may refer to two video frames separated by one or more frames. These two video frames have temporal or spatial correlations and can be used for prediction and coding.

[0075] In this embodiment, a coding block may refer to a small block obtained by dividing a video frame during the compression coding process of the video for better coding. Each small block is an independent unit and is used for different coding methods. In the art, compression includes inter-frame prediction and intra-frame prediction. The small blocks required for inter-frame and intra-frame prediction are coding blocks. Coding blocks can have different levels. A large coding block can contain many small coding blocks. These coding blocks, whether large or small, are the coding blocks described in this application.

[0076] As an example, the coding block includes one or more of a macroblock, a sub-block, a coding tree unit CTU, a coding unit CU, a prediction unit PU (prediction unit), and a transform unit TU (transform unit).

[0077] In this example, the macroblock and the sub-block are, for example, the names of coding blocks in the AVC coding standard, and the CTU, CU, PU, and TU are, for example, the names of coding blocks in the HEVC and VVC standards. The division, prediction, and coding of these coding blocks can all adopt the video coding method provided by the embodiments of this application. Those skilled in the art should know that there are other coding standards in the art. The concepts in these standards that are the same or similar to the coding blocks in this application are within the protection scope of this application. This application is not limited to the several coding standards listed above.

[0078] S430. Predict the predicted pixel information of the target block according to the current pixel information of the current block.

[0079] In this embodiment, the current pixel information may refer to the known pixel information in the current block, including the pixel values of each pixel point in the current block. The predicted pixel information may refer to the pixel information predicted for the target block according to the current block, including the pixel values of each pixel point in the predicted block.

[0080] In this embodiment, to predict the predicted pixel information of the target block based on the current pixel information of the current block, a well-known prediction method in the art can be adopted. For example, based on the similarity and redundancy between the current block and the target block in space and time, the pixel value of each pixel point of the target block is predicted according to the known pixel information of the current block.

[0081] S440. Encode the video according to the predicted pixel information.

[0082] In this embodiment, to encode the video according to the predicted pixel information, techniques well-known in the art can be used, such as means like DCT, quantization, entropy coding, etc. Specifically, the original pixel value (current pixel information) within the encoding block range can be subtracted from the predicted value (predicted pixel information) to obtain a residual. After DCT, the transformed result is quantized, and the quantization value and related syntax elements, such as the CU partition depth, the motion vector corresponding to the pixel block, etc., are then entropy-encoded to obtain the compressed video stream.

[0083] As an example, Figure 5 shows a schematic diagram of the partitioning step in the video encoding method according to Figure 4 the embodiment. As Figure 5 shown, S420 can be specifically implemented as: partitioning the first frame of the video to obtain a plurality of encoding blocks, and the plurality of encoding blocks of the first frame include the current block and the target block. In this example, the encoding method is intra-frame prediction, and both the current block and the target block are in the same frame. The pixel information of the target block at an adjacent or other position is predicted according to the pixel information of the current block in this frame. Intra-frame prediction is mainly based on the spatial correlation of the picture content in the same frame, and adjacent encoding blocks usually have spatially correlated pixel information.

[0084] As an example, Figure 6 shows a schematic diagram of the partitioning step in the video encoding method according to Figure 4 the embodiment. As Figure 6 shown, S420 can be specifically implemented as: partitioning the first frame of the video to obtain a plurality of encoding blocks, and the plurality of encoding blocks of the first frame include the current block, and the second frame includes the target block. In this example, the encoding method is inter-frame prediction, and the current block and the target block are in two adjacent or neighboring frames respectively. The pixel information of the target block in one of the frames is predicted according to the current block in the other frame. Inter-frame prediction is mainly based on the temporal correlation of the picture content in two temporally adjacent frames, and adjacent two frames in the video frame sequence usually have temporally correlated pixel information.

[0085] As an example, Figure 7 shows a schematic diagram of the partitioning step in the video encoding method according to Figure 4 the embodiment. As Figure 7As shown, S420 may be specifically implemented as: determining a foreground region and a background region in the first frame according to the position information of the foreground object; dividing the foreground region according to a first mode and dividing the background region according to a second mode, where the coding blocks in the first mode are smaller than the coding blocks in the second mode.

[0086] In this example, the generated information includes the position information of the foreground object. According to this position information, the position and boundary of the foreground object in the current frame can be determined, so as to delimit the foreground region and the background region. Due to the shape limitation of the coding blocks, the foreground region used to divide the coding blocks may not be a region that exactly fits the contour of the foreground object, but a region that includes the contour of the foreground object. For example Figure 7 As shown, the foreground object is circular, and the foreground region is a square region that contains the circle, delimited by a black square. For the foreground region, coding blocks with a smaller area (fewer pixels) can be divided. For the background region, coding blocks with a larger area (more pixels) can be divided. This is because the amount of movement or change of the foreground object is usually large and needs to be divided in detail, while the content of the background region changes less, so the division can be less detailed. For example, for the background region, the coding blocks (such as CUs) are directly divided according to a fixed mode, such as all divided into 64x64-sized CUs for prediction and coding. In some embodiments, for the background region, a user interface may be provided, and the user can specify the coding block division method. For example, in order to further improve the video quality, the coding blocks can be divided into smaller sizes such as 32x32 or 16x16.

[0087] As an example, Figure 8 is shown according to Figure 4 the schematic diagram of the division step in the video coding method according to the embodiment. As Figure 8 shown, S420 may also be specifically implemented as: determining the edge region of the foreground object in the first frame according to the position information of the foreground object; dividing the edge region according to a first mode.

[0088] In this example, the generated information still includes the position information of the foreground object. According to this position information, the position and edge lines of the foreground object in the current frame can be determined, so as to determine the edge region of the foreground object according to the position and size of the edge lines. As Figure 7 shown, the foreground object is circular, and its edge line is a circular line. The coding region along this circular line can be Figure 7The rectangular area between the large black square and the small black square. For this rectangular area, a more detailed partitioning method can be adopted. For example, smaller (with fewer pixels) coding blocks can be used for partitioning, while for other areas, including the area outside the large black square and the area inside the small black square, a less detailed partitioning method can be adopted. For example, larger (with more pixels) coding blocks can be used for partitioning. For example, for the edge of the foreground area, such as the edge of the area where the person is located, it is partitioned according to a fixed pattern, such as only in the 8x8 pattern.

[0089] As an example, Figure 9 shows according to Figure 4 the two partitioning methods adopted in the partitioning step of the video coding method according to the Figure 9 embodiment. As shown in

[0090] As shown in Figure 9 the left side, the foreground object is a rigid object, so the third mode is used to partition the foreground area. As shown in Figure 9 the right side, the foreground object is a flexible object, so the fourth mode is used to partition the foreground area. It can be seen from the figure that both the third mode and the fourth mode are partitioning modes for the foreground object. Therefore, compared with the background area (i.e., the area outside the black square), the partitioning of both modes is more detailed. And between the third and fourth modes, the partitioning of the fourth mode is more detailed, using smaller (with fewer pixels) coding blocks to partition the circular foreground object in the figure, and the partitioning of the third mode is coarser, using larger (with more pixels) coding blocks to partition the foreground object. This is because the motion directions of the various parts of a rigid object are relatively consistent, and its spatial and temporal correlations are stronger. Therefore, in order to improve the compression efficiency, larger coding blocks can be used. While the various parts of a flexible object usually have inconsistent motion directions or motion vectors, and its spatial and temporal correlations are weaker. Therefore, in order to retain as many image details as possible and make the picture restoration degree higher, smaller coding blocks can be used. For example, for the inside of the foreground area, if the foreground area is made of a flexible material, it is partitioned according to a fixed pattern of smaller coding blocks, such as only in the 8x8 pattern; if the foreground area is made of a rigid material, it is partitioned according to a fixed pattern of larger coding blocks, such as only in the 64x64 pattern.

[0091] As an example, Figure 10 shows according toFigure 4 Schematic diagram of the partitioning step in the video encoding method of the embodiment. As Figure 10 shown, the foreground object is in the shape of a human body with a relatively complex edge shape. For such a foreground object, multiple partitioning methods can be adopted. First, the foreground area can be determined, that is, the black rectangular frame in the figure. For the part outside the rectangular frame, the coarsest background area coding block partitioning method can be adopted, such as coding blocks of size 64x64. For the image area in the foreground area that involves the edge of the human body contour, the most detailed partitioning method can be adopted, such as coding blocks of size 8x8. For the part in the foreground area that does not involve the edge of the foreground object, a partitioning method with a medium level of detail can be adopted, such as coding blocks of size 32x32. For the inside of the foreground object, considering the rigidity of the human body, a partitioning method with a medium level of detail can be adopted, such as coding blocks of size 32x32. For other parts, a flexible partitioning strategy can be adopted, such as using 24x32, 8x32, 16x16, etc.

[0092] Specifically, the encoder selects the smallest area aligned with 64x64 to enclose the foreground moving object according to the foreground moving object information provided by the video generator, as Figure 10 shown by the black rectangular area. If there are multiple foreground moving objects, multiple rectangular areas are generated. For the background area outside the rectangular area, compared with the foreground moving object, its importance is relatively low, and all are partitioned according to the largest coding block specification, that is, 64x64, and predicted and coded bitstream is generated according to 64x64. For the area within the rectangular area, further partitioning methods of coding blocks such as 32x32, 16x16, 24x32, 8x32, etc. (allowed by the video coding protocol) that do not contain any foreground moving objects can be carried out. For the foreground moving object, if it is a rigid object, the movement is often consistent, and the largest block partitioning that does not contain the edge of the foreground moving object can also be carried out, such as the ways allowed by the video coding protocol such as 64x64 and 32x32. For the foreground moving object, if it is a flexible object, the movement is often inconsistent, and it is partitioned according to the smallest 8x8 method. For the edge area of the foreground moving object, which is image detail information, it is partitioned into coding blocks according to the smallest 8x8.

[0093] Based on the foregoing Figure 4 method embodiment, the embodiment of the present application also provides a video encoding device 1100, and its structural schematic diagram is as Figure 11 shown. This device is used to execute each step in the foregoing Figure 4 one.

[0094] According to this embodiment, the video encoding device 1100 includes an acquisition module 1110, a partitioning module 1120, a prediction module 1130, and an encoding module 1140. The acquisition module 1110 is configured to acquire the generation information of the video. The partitioning module 1120 is configured to partition the first frame of the video according to the generation information to obtain a plurality of coding blocks. The plurality of coding blocks of the first frame include a current block, and a target block is included in the plurality of coding blocks of the first frame or in a second frame adjacent to the first frame of the video. The prediction module 1130 is configured to predict the predicted pixel information of the target block according to the current pixel information of the current block. The encoding module 1140 is configured to encode the video according to the predicted pixel information.

[0095] In one embodiment, the generation information includes the position information and the material information of the foreground object in the first frame.

[0096] In one embodiment, the partitioning module 1120 is further configured to:

[0097] Determine the foreground area and the background area in the first frame according to the position information of the foreground object;

[0098] Partition the foreground area according to a first mode and partition the background area according to a second mode, wherein the coding blocks in the first mode are smaller than the coding blocks in the second mode.

[0099] In one embodiment, the partitioning module 1120 is further configured to:

[0100] Determine the edge area of the foreground object in the first frame according to the position information of the foreground object;

[0101] Partition the edge area according to a first mode.

[0102] In one embodiment, the partitioning module 1120 is further configured to:

[0103] Determine the foreground area and the background area in the first frame according to the position information of the foreground object;

[0104] Judge whether the foreground object belongs to a rigid object or a flexible object according to the material information of the foreground object;

[0105] If the foreground object belongs to a rigid object, partition the foreground area according to a third mode;

[0106] If the foreground object belongs to a flexible object, partition the foreground area according to a fourth mode, wherein the coding blocks in the third mode are larger than the coding blocks in the fourth mode.

[0107] In one embodiment, the video includes a virtual scene video generated by a video generation tool. Acquiring the generation information of the video includes:

[0108] Obtain generation information from a video generation tool.

[0109] In one embodiment, the video generation tool includes a renderer or an AI tool.

[0110] In one embodiment, the coding block includes one or more of a macroblock, a subblock, a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU).

[0111] It should be noted that Figure 11 When the video coding device 1100 provided in the illustrated embodiment executes the method, only the division of the above functional modules is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the video coding device 1100 provided in the above embodiment and Figure 4 the illustrated video coding method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0112] Figure 12 FIG. is a schematic hardware structure diagram of a computing device 1200 provided by an embodiment of the present application.

[0113] See Figure 12 , the computing device 1200 includes a processor 1210, a memory 1220, a communication interface 1230, and a bus 1240. The processor 1210, the memory 1220, and the communication interface 1230 are connected to each other through the bus 1240. The processor 1210, the memory 1220, and the communication interface 1230 may also be connected by other connection methods other than the bus 1240.

[0114] Among them, the memory 1220 may be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical memory, hard disk, etc.

[0115] Among them, the processor 1210 can be a general-purpose processor, which can execute specific steps and / or operations by reading and executing the content stored in a memory (such as the memory 1220). For example, the general-purpose processor can be a central processing unit (CPU). The processor 1210 can include at least one circuit to execute Figure 4 all or part of the steps of the video encoding method provided by the illustrated embodiment.

[0116] Among them, the communication interface 1230 includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing interconnection of components inside the computing device 1200, and interfaces for implementing interconnection between the computing device 1200 and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc. The communication interface 1230 can be externally connected to an input device and an output device. For example, the input device can be a microphone or a microphone array for capturing voice input signals; it can be a communication network connector for receiving the collected input signals from the cloud or other devices; it can also include, for example, a keyboard, a mouse, etc. The output device can output various information to the outside, including the determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0117] Among them, the bus 1240 can be of any type and is a communication bus for implementing the interconnection of the processor 1210, the memory 1220, and the communication interface 1230, such as a system bus.

[0118] The above components can be respectively arranged on independent chips, or at least partially or entirely arranged on the same chip. Whether to arrange each component on different chips or integrate them on one or more chips often depends on the needs of product design. The specific implementation forms of the above components are not limited in the embodiments of the present application.

[0119] Figure 12 The illustrated computing device 1200 is only exemplary. During implementation, the computing device 1200 may further include other components, which will not be listed one by one herein.

[0120] An embodiment of the present application can also be a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the video encoding method according to various embodiments of the present application described above in this specification.

[0121] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0122] The concepts, principles, and ideas of the present application have been described in detail above in conjunction with specific embodiments (including examples and instances). Those skilled in the art should understand that the embodiments of the present application are not limited to the several forms given above. After reading the present application document, those skilled in the art may make any possible improvements, substitutions, and equivalent forms to the steps, methods, devices, and components in the above embodiments. These improvements, substitutions, and equivalent forms should be regarded as falling within the scope of the present application. The protection scope of the present application is only subject to the claims.

Claims

1. A video encoding method, characterized in that, The method includes: A video encoder obtains a first video and generation information generated by a video generation tool; the first video is a virtual scene video that has been generated by the video generation tool; the generation information includes three-dimensional coordinate information and material information of objects in the first video, as well as layer information; foreground objects and background objects are distinguished according to the three-dimensional coordinate information of the objects, and position information of the foreground objects is determined; According to the generation information, the first frame of the first video is divided to obtain a plurality of coding blocks, including: determining a foreground area and a background area in the first frame according to the position information and layer information of the foreground objects; dividing the foreground area into a plurality of first coding blocks according to a first mode, and dividing the background area into a plurality of second coding blocks according to a second mode, wherein the area of the first coding block is smaller than that of the second coding block; judging whether the foreground object belongs to a rigid object or a flexible object according to the material information of the object; if the foreground object belongs to a rigid object, dividing the foreground object into a plurality of third coding blocks according to a third mode; if the foreground object belongs to a flexible object, dividing the foreground object into a plurality of fourth coding blocks according to a fourth mode, wherein the area of the third coding block is larger than that of the fourth coding block; the foreground area does not include the foreground object; Predict pixel information of a target block according to the plurality of coding blocks of the first frame; Encode and generate a second video of the virtual scene according to the pixel information of the target block.

2. The video encoding method according to claim 1, wherein The step of dividing the first frame of the video into a plurality of coding blocks according to the generation information further includes: Determining the position and edge lines of the foreground object in the current frame according to the position information of the foreground object, and determining the edge area of the foreground object in the first frame according to the position and size of the edge lines; the edge area of the foreground object is an area within a certain range where the edge lines transition to the foreground area; Dividing the edge area according to the first mode.

3. The video encoding method according to claim 1 or 2, wherein The target block includes a first target block and a second target block. The step of predicting pixel information of the target block according to the plurality of coding blocks of the first frame includes: The plurality of coding blocks of the first frame include a current block, and the plurality of coding blocks of the first frame include a first target block; a second frame adjacent to the first frame of the first video includes a second target block; the current block is different from the first target block; Predicting the predicted pixel information of the target block according to the current pixel information of the current block; wherein, the current pixel information refers to the known pixel information in the current block, including the pixel value of each pixel point in the current block; the predicted pixel information refers to the pixel information predicted for the target block according to the current block, including the pixel value of each pixel point in the predicted block; including: Performing intra-frame prediction, where the current block and the first target block are in the same frame, and predicting the pixel information of the first target block at an adjacent or other position according to the pixel information of the current block in the frame; and / or Perform inter-frame prediction. The current block and the second target block are in two adjacent or neighboring frames respectively. Predict the pixel information of the second target block in the second frame based on the current block in the first frame.

4. The video encoding method according to claim 3, wherein Encoding the second video of the virtual scene according to the pixel information of the target block includes: Subtract the pixel information of the first target block from the current pixel information of the current block to obtain a first residual, and subtract the pixel information of the second target block from the current pixel information of the current block to obtain a second residual. The first residual and the second residual are subjected to DCT, and then the transformation result is quantized. The quantization value and related syntax elements are then entropy-encoded to obtain a compressed video stream. Generate the second video of the virtual scene according to the compressed video stream.

5. The video encoding method according to claim 1, wherein The video includes a virtual scene video generated by a video generation tool. Obtaining the generation information of the video includes: Obtain the generation information from the video generation tool.

6. The video encoding method according to claim 4, wherein The video generation tool includes a renderer or an AI tool.

7. The video encoding method according to claim 1, wherein The encoded block includes one or more of a macroblock, a sub-block, a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU).

8. A video encoding device, characterized in that, The apparatus includes: An acquisition module, configured to acquire, for a video encoder, a first video generated by a video generation tool and generation information; the first video is a virtual scene video that has been generated by the video generation tool; the generation information includes three-dimensional coordinate information of an object in the first video, material information of the object, and layer information; A partitioning module, configured to distinguish a foreground object and a background object according to the three-dimensional coordinate information of the object, and determine the position information of the foreground object; partition the first frame of the video according to the generation information to obtain a plurality of encoded blocks, including: determining a foreground area and a background area in the first frame according to the position information of the foreground object and the layer information; partitioning the foreground area into a plurality of first encoded blocks according to a first mode, and partitioning the background area into a plurality of second encoded blocks according to a second mode, where the area of the first encoded block is smaller than that of the second encoded block; judging whether the foreground object belongs to a rigid object or a flexible object according to the material information of the object; if the foreground object belongs to a rigid object, partitioning the foreground object into a plurality of third encoded blocks according to a third mode; if the foreground object belongs to a flexible object, partitioning the foreground object into a plurality of fourth encoded blocks according to a fourth mode, where the area of the third encoded block is larger than that of the fourth encoded block; the foreground area does not include the foreground object; A prediction module, configured to predict the pixel information of a target block according to a plurality of encoded blocks in the first frame; An encoding module, configured to encode according to the pixel information of the target block to obtain a compressed video stream, and generate a second video of the virtual scene according to the compressed video stream.

9. The video encoding device according to claim 8, wherein The partitioning module further includes: Determine the position and edge lines of the foreground object in the current frame according to the position information of the foreground object, and determine the edge region of the foreground object in the first frame according to the position and size of the edge lines; the edge region of the foreground object is the region within a certain range where the edge lines transition to the foreground region; Divide the edge region according to the first mode.

10. The video coding device according to claim 8 or 9, characterized in that, The target block includes a first target block and a second target block, and the prediction module includes: Among the multiple coding blocks of the first frame is a current block, and among the multiple coding blocks of the first frame is a first target block; in the second frame adjacent to the first frame of the video is a second target block; the current block is different from the first target block; Predict the predicted pixel information of the target block according to the current pixel information of the current block; where the current pixel information refers to the known pixel information in the current block, including the pixel values of each pixel point in the current block; the predicted pixel information refers to the pixel information predicted for the target block according to the current block, including the pixel values of each pixel point in the predicted block; including: performing intra-frame prediction, where the current block and the first target block are both in the same frame, and predicting the pixel information of the first target block at adjacent or other positions according to the current pixel information of the current block in the frame; and / or Performing inter-frame prediction, where the current block and the second target block are in two adjacent or neighboring frames respectively, and predicting the pixel information of the second target block in the second frame according to the current block in the first frame.

11. The video encoding device according to claim 10, wherein The coding module includes: Subtract the pixel information of the first target block from the current pixel information of the current block to obtain a first residual, subtract the pixel information of the second target block from the current pixel information of the current block to obtain a second residual, perform DCT on the first residual and the second residual, then quantize the transformation result, the quantization value and related syntax elements, and then perform entropy coding to obtain a compressed video stream, and generate the second video of the virtual scene according to the compressed video stream.

12. A computing device, characterized in that, The computing device includes a processor and a memory, and the processor is configured to execute a computer program stored in the memory to implement the video coding method according to any one of claims 1 to 7.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program runs on the processor, it causes the processor to execute the video coding method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Animation generation method, animation playing method and related devices

    CN112037311A

  • Video coding method and system based on region of interest

    CN114745549A

  • Ai-assisted programmable hardware video codec

    US20200374534A1