Video coding method, device and equipment and computer storage medium

By combining AI feature extraction and block prediction mode, the problem of high bitrate in existing video coding technology is solved, and more flexible block prediction and lower coding bitrate are achieved.

CN121814955APending Publication Date: 2026-04-07CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing video coding technologies, fixed or finitely variable block structures result in high video coding bitrates, especially when coding regions with complex textures, which require more residual data compensation and thus increase the bitrate.

Method used

The semantic features of video data are extracted using an AI feature extraction tool, and block prediction is performed based on an AI model. Rate distortion cost is calculated through the first and second block prediction modes, and the mode with the lowest rate distortion cost is selected for encoding. The video encoded stream is generated by combining integer discrete cosine transform, quantization and entropy coding.

Benefits of technology

It reduces the bitrate of video encoding, improves the flexibility and efficiency of block prediction, and reduces unnecessary data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814955A_ABST
    Figure CN121814955A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding method and device, equipment and a computer storage medium. The method comprises the following steps: carrying out blocking and coding prediction on acquired video data to obtain a first prediction block and a first coding rate; and performing block prediction on the video data based on the AI model to obtain a second prediction block. And respectively calculating the quadratic sum coding rate of the pixel difference values between the first prediction block and the target block and between the second prediction block and the target block to obtain a first rate distortion cost and a second rate distortion cost. And under the condition that the first rate distortion cost is greater than the second rate distortion cost, obtaining a video coding stream. According to the embodiment of the invention, through the first block prediction mode and the AI model memory block prediction, the first rate-distortion cost and the second rate-distortion cost corresponding to the two block prediction modes can be obtained, and the block prediction mode with the minimum rate-distortion cost is selected from the first block prediction mode and the second block prediction mode for coding. The rate distortion cost of the block prediction mode is minimum, so that the video coding rate can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video encoding and transmission, and particularly relates to a video encoding method, apparatus, device, and computer storage medium. Background Technology

[0002] Currently, users typically use their devices to watch different types of video data during multimedia entertainment. The reason users can see video data is because the video data is encoded and transmitted to their multimedia devices.

[0003] In existing technologies, traditional video coding techniques are limited by framework constraints, such as fixed coding block size. For example, using large coding blocks to predict small detail areas will result in loss of detail and require higher bitrate compensation, thus leading to higher bitrates in video coding. Summary of the Invention

[0004] This invention provides a video encoding method, apparatus, device, computer storage medium, and computer program product that can reduce the bitrate of video encoding.

[0005] In a first aspect, embodiments of the present invention provide a video encoding method, the method comprising: The acquired video data is divided into blocks and encoded according to the first block prediction mode to obtain the first prediction block and the first coding bitrate. The rate distortion cost of the first block prediction mode is less than a preset threshold. AI feature extraction tools were used to extract semantic features from video data, and the extraction bitrate of semantic features was recorded. The video data is segmented and predicted based on an AI model to obtain the second prediction block; Based on the first prediction block, the extracted bitrate, and the video data, the second coding bitrate of the second prediction block is determined; The coding rate of the sum of squares of the pixel differences between the first prediction block, the second prediction block and the target block is calculated respectively to obtain the first rate distortion cost and the second rate distortion cost; If the cost of first rate distortion is greater than the cost of second rate distortion, the video coded stream is obtained by encoding based on the residual matrix corresponding to the difference between the pixels of the second prediction block and the target block.

[0006] In one feasible implementation, before dividing the acquired video data into blocks and predicting encoding according to the first block prediction mode to obtain the first prediction block and the first encoding bitrate, the method further includes: Video data is divided into first target blocks corresponding to at least one block prediction mode, resulting in multiple first target blocks; Each type of first target block is predicted separately, resulting in multiple first prediction blocks; Based on each first target block, each first prediction block, and the coding rate corresponding to the first prediction block, multiple initial first rate distortion costs corresponding to the first target block are obtained; The initial first rate distortion cost with the smallest value is selected from multiple initial first rate distortion costs as the first rate distortion cost, and the block prediction mode corresponding to the first rate distortion cost is taken as the first prediction mode.

[0007] In one feasible implementation, prediction is performed for each type of first target block to obtain multiple first prediction blocks, including: Inter-frame prediction and intra-frame prediction are performed for each type of first target block to obtain multiple first prediction blocks.

[0008] In one feasible implementation, the semantic features include semantic tags. Based on an AI model, the video data is segmented and predicted to obtain a second prediction block, which includes: Obtain the weight information of semantic tags; Determine the corresponding initial block division method based on the weight information; Based on the initial block division method, the video data is divided into corresponding second target blocks; The second target block is encoded and predicted to obtain the second prediction block.

[0009] In one feasible implementation, the video data is divided into corresponding second target blocks based on the initial block division method, including: The video data is divided according to the initial block division method to obtain the third target block; When the pixel density of the semantic boundary of the third target block is less than or equal to a preset density threshold, the third target block is used as the second target block. When the semantic boundary pixel density of the third target block is greater than a preset density threshold, the third target block is divided into multiple fourth target blocks; the fourth target blocks are used as the second target blocks.

[0010] In one feasible implementation, determining the second coding bitrate of the second prediction block based on the first prediction block, the extracted bitrate, and the video data includes: Obtain the number of pixels in the video data and the number of pixels in the first prediction block; The prediction block size ratio is obtained by dividing the square of the number of pixels in the first prediction block by the square of the number of pixels in the video data. Multiply the predicted block size ratio by the extraction rate to obtain the second coding rate.

[0011] In one feasible implementation, the sum of squares of the pixel differences between the first prediction block, the second prediction block, and the target block is calculated to obtain the first rate-distortion cost and the second rate-distortion cost, including: The first pixel difference is obtained based on the difference between the pixels of the first target block and the pixels of the first prediction block corresponding to the first block prediction mode. The first distortion value is obtained by dividing the sum of squares of the first pixel differences by the number of pixels in the first target block; The control information bitrate is determined based on the prediction mode type corresponding to the first prediction block; The difference between the pixel matrix corresponding to the first target block and the pixel matrix corresponding to the first prediction block is used as the target residual matrix; Determine the residual data bitrate based on the values ​​in the target residual matrix; The control information bitrate and the residual data bitrate are used as the coding bitrates corresponding to the first prediction block; The first rate-distortion cost is obtained by adding the product of the first distortion value, the first preset weight, and the coding rate corresponding to the first prediction block. The second pixel difference is obtained based on the difference between the pixels in the second target block and the corresponding pixels in the second prediction block; The second distortion value is obtained by summing the squares of the second pixel differences and dividing by the number of pixels in the second target block; The second rate-distortion cost is obtained by adding the product of the second preset weight and the second coding rate.

[0012] In one feasible implementation, encoding is performed based on the residual matrix corresponding to the difference between the pixels of the second predicted block and the target block to obtain a video encoded stream, including: The residual matrix is ​​transformed from the spatial domain to the frequency domain to obtain the transformation matrix; Divide the values ​​in the transformation matrix by the preset quantization step size to obtain the quantization matrix; Entropy encoding is performed on the coefficients in the quantization matrix to obtain the video encoded stream.

[0013] In one feasible implementation, the residual matrix is ​​transformed from the spatial domain to the frequency domain to obtain the transformation matrix, including: Obtain the forward transform kernel matrix and the transpose matrix of the forward transform kernel matrix based on the floating-point discrete cosine transform (DCT) kernel; The transformation matrix is ​​obtained by multiplying the product of the forward transformation kernel matrix and the residual matrix by the transpose matrix.

[0014] In one feasible implementation, entropy encoding is performed on the coefficients in the quantization matrix to obtain a video encoded stream, including: The coefficients in the quantization matrix are sorted according to their numerical values ​​to obtain a one-dimensional sequence; Entropy coding is performed on the one-dimensional sequence based on the number of consecutive zero coefficients and the magnitude and sign of the non-zero coefficients to obtain the video coded stream.

[0015] Secondly, embodiments of the present invention provide a video encoding apparatus, the apparatus comprising: The acquired video data is divided into blocks and encoded according to the first block prediction mode to obtain the first prediction block and the first coding bitrate. The rate distortion cost of the first block prediction mode is less than a preset threshold. AI feature extraction tools were used to extract semantic features from video data, and the extraction bitrate of semantic features was recorded. The video data is segmented and predicted based on an AI model to obtain the second prediction block; Based on the first prediction block, the extracted bitrate, and the video data, the second coding bitrate of the second prediction block is determined; The coding rate of the sum of squares of the pixel differences between the first prediction block, the second prediction block and the target block is calculated respectively to obtain the first rate distortion cost and the second rate distortion cost; If the cost of first rate distortion is greater than the cost of second rate distortion, the video coded stream is obtained by encoding based on the residual matrix corresponding to the difference between the pixels of the second prediction block and the target block.

[0016] Thirdly, embodiments of the present invention provide a video encoding apparatus, the apparatus including a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the video encoding method as described in the first aspect.

[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the video encoding method as described in the first aspect.

[0018] Fifthly, embodiments of the present invention provide a computer program product, including a computer program that, when executed by a processor, implements the video encoding method as described in the first aspect.

[0019] This invention provides a video encoding method, apparatus, device, computer storage medium, and computer program product. The method involves segmenting and encoding the acquired video data according to a first block prediction mode to obtain a first prediction block and a first encoding bitrate. Semantic features of the video data are extracted using an AI feature extraction tool, and the extraction bitrate of the semantic features is recorded. Block prediction is performed on the video data based on an AI model to obtain a second prediction block. A second encoding bitrate of the second prediction block is determined based on the first prediction block, the extraction bitrate, and the video data. The encoding bitrate of the sum of the squares of the pixel differences between the first prediction block, the second prediction block, and the target block is calculated to obtain a first rate-distortion cost and a second rate-distortion cost. If the first rate-distortion cost is greater than the second rate-distortion cost, encoding is performed based on the residual matrix corresponding to the pixel differences between the second prediction block and the target block to obtain a video encoded stream. In this embodiment of the invention, block prediction is performed using a first block prediction mode and an AI model, respectively, which can obtain the first rate-distortion cost and the second rate-distortion cost corresponding to the two block prediction modes. The block prediction mode with the lowest rate-distortion cost is selected for encoding. Since this block prediction mode has the lowest rate-distortion cost, the video encoding bitrate can be reduced. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating a video encoding method provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for determining a first prediction mode according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating a method for obtaining a second prediction block according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating another method for obtaining a second prediction block provided by an embodiment of the present invention; Figure 5 This is a flowchart illustrating a method for obtaining a second coding rate according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating a method for obtaining a first rate distortion cost and a second rate distortion cost according to an embodiment of the present invention. Figure 7 This is a flowchart illustrating a method for obtaining a video encoded stream according to an embodiment of the present invention; Figure 8This is a flowchart illustrating a method for obtaining a transformation matrix according to an embodiment of the present invention; Figure 9 This is a flowchart illustrating another method for obtaining a video encoded stream provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a video encoding device provided in an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of a video encoding device provided in an embodiment of the present invention. Detailed Implementation

[0022] The features and exemplary embodiments of various aspects of the present invention will now be described in detail. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely intended to explain the present invention and not to limit the present invention. For those skilled in the art, the present invention can be practiced without some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present invention by illustrating examples of the invention.

[0023] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0024] Before describing the technical solutions provided by the embodiments of the present invention, in order to facilitate understanding of the embodiments of the present invention, the present invention will specifically explain the problems existing in the related technologies: Currently, users typically use their devices to watch different types of video data during multimedia entertainment. The reason users can see video data is because the video data is encoded and transmitted to their multimedia devices.

[0025] In existing technologies, traditional video coding techniques are commonly used to encode video data. However, these techniques, such as H.264, H.265, and H.266, typically use fixed or finitely variable block sizes when dividing video data into multiple target blocks. For example, a video image can only be divided into 4x4, 8x8, or 16x16 target blocks. This fixed block division method lacks flexibility. When encoding and predicting regions with complex textures, using larger target blocks leads to larger prediction errors, requiring more residual data for compensation and increasing the bit rate.

[0026] In summary, existing technologies achieve relatively high bitrates for video encoding.

[0027] To address the problems of the prior art, embodiments of the present invention provide a video encoding method, apparatus, device, computer storage medium, and computer program product.

[0028] This invention embodiment performs block prediction and encoding prediction on the acquired video data according to a first block prediction mode to obtain a first prediction block and a first encoding bitrate. Semantic features of the video data are extracted using an AI feature extraction tool, and the extraction bitrate of the semantic features is recorded. Block prediction is performed on the video data based on the AI ​​model to obtain a second prediction block. A second encoding bitrate of the second prediction block is determined based on the first prediction block, the extraction bitrate, and the video data. The encoding bitrate of the sum of squares of the pixel differences between the first prediction block, the second prediction block, and the target block is calculated to obtain a first rate-distortion cost and a second rate-distortion cost. If the first rate-distortion cost is greater than the second rate-distortion cost, encoding is performed based on the residual matrix corresponding to the pixel differences between the second prediction block and the target block to obtain the video encoded stream. This invention embodiment performs block prediction using both the first block prediction mode and the AI ​​model, obtaining the first and second rate-distortion costs corresponding to the two block prediction modes respectively. The block prediction mode with the lowest second rate-distortion cost is selected for encoding. Since this block prediction mode has the lowest rate-distortion cost, the video encoding bitrate can be reduced.

[0029] The video encoding method provided in the embodiments of the present invention will be introduced first below.

[0030] Figure 1 A flowchart illustrating a video encoding method according to an embodiment of the present invention is shown. Figure 1 As shown, the method includes steps S110-S160.

[0031] S110: The acquired video data is divided into blocks and encoded according to the first block prediction mode to obtain the first prediction block and the first coding bitrate.

[0032] The block prediction mode represents the method of dividing video data into blocks and encoding and predicting these blocks. The rate-distortion cost of the first block prediction mode is less than a preset threshold. The preset threshold is not fixed and can be modified according to user needs. The first coding bitrate represents the amount of data used per unit time during the encoding process of the first prediction block.

[0033] In one embodiment, the block prediction mode with the lowest rate-distortion cost among multiple block prediction modes can be selected as the first block prediction mode. The acquired video data is then divided into blocks of corresponding sizes according to the first block prediction mode, and coding prediction operations are performed on these blocks to obtain the first prediction block and the first coding bitrate.

[0034] In one example, the block prediction mode with the lowest rate-distortion cost among multiple block prediction modes, such as H.264, H.265, and H.266, can be selected as the first block prediction mode. The acquired video data is then divided into blocks of corresponding sizes based on the first block prediction mode, and coding prediction operations are performed on these blocks to obtain the first prediction block and the first coding bitrate.

[0035] S120: Uses AI feature extraction tools to extract semantic features from video data and records the extraction bitrate of semantic features.

[0036] Among them, semantic features represent high-level information extracted from video images that corresponds to human cognitive logic, such as information about objects, actions, and scenes.

[0037] In one embodiment, AI feature extraction tools can be used to perform spatiotemporal feature extraction, feature vectorization, and other operations to extract semantic features from video data. The semantic importance of these features is then determined based on their semantic importance, and the resulting bitrate is mapped to a corresponding recommendation bitrate, which serves as the extraction bitrate for the semantic features.

[0038] S130: Based on the AI ​​model, the video data is divided into blocks for prediction to obtain the second prediction block.

[0039] In one embodiment, an AI model can be used to determine target and background regions in video data based on semantic features. The target region is divided into smaller blocks and the blocks are encoded and predicted, while the background region is divided into larger blocks and the blocks are encoded and predicted, thereby obtaining a second prediction block.

[0040] S140: Based on the first prediction block, the extracted bitrate, and the video data, determine the second coding bitrate of the second prediction block.

[0041] In one embodiment, the second coding bitrate corresponding to the number of pixels in the first prediction block, the extraction bitrate, and the number of pixels in the video data can be determined based on the correspondence between the number of pixels in the first prediction block, the extraction bitrate, the number of pixels in the video data, and the second coding bitrate.

[0042] S150: Calculate the sum of squares of the pixel differences between the first prediction block, the second prediction block and the target block, respectively, to obtain the first rate-distortion cost and the second rate-distortion cost.

[0043] In one embodiment, a first pixel difference can be determined based on the difference between pixels in the first predicted block and pixels in the corresponding target block, and a first distortion value can be determined based on the sum of squares of the first pixel differences and the number of pixels in the first predicted block. Then, based on the correspondence between the first distortion value and the coding rate of the target block corresponding to the first predicted block, and the first rate-distortion cost, a first rate-distortion cost corresponding to the first distortion value and the coding rate of the target block corresponding to the first predicted block is determined. Similarly, a second pixel difference can be determined based on the difference between pixels in the second predicted block and pixels in the corresponding target block, and a second distortion value can be determined based on the sum of squares of the second pixel differences and the number of pixels in the second predicted block. Finally, based on the correspondence between the second distortion value and the coding rate of the target block corresponding to the second predicted block, and the second rate-distortion cost, a second distortion value corresponding to the coding rate of the target block corresponding to the second predicted block is determined.

[0044] S160: If the first rate distortion cost is greater than the second rate distortion cost, the video coded stream is obtained by encoding based on the residual matrix corresponding to the difference between the pixels of the second prediction block and the target block.

[0045] In this embodiment, when the cost of the first rate-distortion (RDD) is greater than the cost of the second RDD, integer discrete cosine transform (DCT), quantization, and entropy coding operations can be performed based on the residual matrix corresponding to the difference between the pixels of the second prediction block and the target block to obtain the video encoded stream. Specifically, integer DCT is used to concentrate most of the information in the matrix at a certain location. Quantization is used to bring the high-frequency coefficients in the quantization matrix close to zero, thereby discarding details that are not sensitive to the human eye (mainly high-frequency components) and significantly reducing the amount of data. Entropy coding is used to compress the data, achieving the most efficient lossless compression.

[0046] In this embodiment of the invention, block prediction is performed using a first block prediction mode and an AI model, respectively, which can obtain the first rate-distortion cost and the second rate-distortion cost corresponding to the two block prediction modes. The block prediction mode with the lowest rate-distortion cost is selected for encoding. Since this block prediction mode has the lowest rate-distortion cost, the video encoding bitrate can be reduced.

[0047] In one embodiment, before performing block segmentation and coding prediction on the acquired video data according to the first block prediction mode to obtain the first prediction block and the first coding bitrate, as follows: Figure 2 As shown, the video encoding method may further include steps S210-S240.

[0048] S210: Divide the video data into first target blocks corresponding to at least one block prediction mode to obtain multiple first target blocks.

[0049] In one embodiment, video data can be divided into first target blocks corresponding to each segmentation rule according to each segmentation prediction mode, thereby obtaining multiple first target blocks.

[0050] In one example, the video data is a 128*128 video image. The first block prediction mode divides the video image into multiple 4*4 target blocks, while the second block prediction mode divides the video image into multiple 16*16 target blocks. By dividing the video data according to the first and second block prediction modes respectively, two different sizes of first target blocks can be obtained.

[0051] S220: Predict each type of first target block separately to obtain multiple first prediction blocks.

[0052] In one embodiment, after obtaining multiple first target blocks, prediction encoding can be performed based on the semantic features of each first target block to obtain multiple first prediction blocks.

[0053] In one embodiment, predicting each first target block separately to obtain multiple first predicted blocks may include step S221.

[0054] S221: Perform inter-frame prediction and intra-frame prediction for each type of first target block to obtain multiple first prediction blocks.

[0055] In one embodiment, coded video data adjacent to the video data and already encoded can be acquired. The target block with the smallest difference between its pixel value and that of the first target block is identified. Since the pixel difference is the smallest, it indicates that the target block is most similar to the first target block, which is equivalent to finding the best prediction for the current block. The pixel values ​​in this target block are then used as the pixel values ​​at the corresponding positions in the first prediction block to obtain the first prediction block corresponding to the inter-frame prediction.

[0056] In one embodiment, the value of the pixel at the target position in the first target block can be used as the value of the pixel corresponding to the texture pattern of the pixel at the target position in the non-target pixel, so as to obtain the first prediction block corresponding to the intra-frame prediction.

[0057] S230: Based on each first target block, each first prediction block, and the coding rate corresponding to the first prediction block, obtain multiple initial first rate distortion costs corresponding to the first target block.

[0058] In one embodiment, multiple pixel differences can be determined based on the differences between the pixels of each first prediction block and the corresponding pixels of the first target block; that is, each first prediction block and the first target block correspond to one pixel difference. Multiple distortion values ​​are determined based on the sum of the squares of each pixel difference and the number of pixels in each first prediction block; that is, the sum of the squares of one pixel difference and the number of pixels in one first prediction block determine one distortion value. Then, based on the correspondence between each distortion value and the coding rate of the target block corresponding to each first prediction block, and the first rate distortion cost, the first rate distortion cost corresponding to each distortion value and the coding rate of the target block corresponding to each first prediction block is determined, resulting in multiple initial first rate distortion costs.

[0059] S240: Select the initial first rate distortion cost with the smallest value from multiple initial first rate distortion costs as the first rate distortion cost, and take the block prediction mode corresponding to the first rate distortion cost as the first prediction mode.

[0060] In one embodiment, after obtaining multiple initial first rate distortion costs, the multiple initial first rate distortion costs can be compared sequentially, and the initial first rate distortion cost with the smallest value can be selected as the first rate distortion cost. Then, the block prediction mode corresponding to the first rate distortion cost can be used as the first prediction mode.

[0061] This invention, in its embodiments, divides video data into first target blocks using at least one block prediction mode, and calculates the corresponding rate-distortion cost for each block. This allows for the determination of the block prediction mode with the lowest rate-distortion cost, which is then used as the first block prediction mode, improving the accuracy of determining the first block prediction mode. In one embodiment, semantic features include semantic tags. Based on an AI model, block prediction is performed on the video data to obtain a second predicted block, such as... Figure 3 As shown, steps S131-S134 may be included.

[0062] S131: Obtain the weight information of semantic labels.

[0063] Semantic tags represent meaningful category names or identifiers for a region, object, or pixel in a video. The magnitude of weight information corresponds to the way video data is segmented.

[0064] In this embodiment, an AI model is used to determine the semantic labels corresponding to the extracted semantic features. The weight information corresponding to the semantic label is then retrieved according to a pre-defined mapping table.

[0065] S132: Determine the corresponding initial block division method based on the weight information.

[0066] After obtaining the weight information, the initial segmentation method can be determined based on the magnitude of the weight information. In one embodiment, a larger weight information value indicates that the video data needs to be divided into smaller blocks. For example, when the weight information value is greater than a preset threshold, the corresponding segmentation method is to divide the video data into smaller blocks.

[0067] In one example, the weight information value of 8 is greater than the preset weight threshold of 5, indicating that the video data needs to be divided into smaller blocks, and it can be determined that the way to divide the video data is 4*4.

[0068] In another example, the weight information value is 4, indicating that the video data needs to be divided into larger blocks, and the method of dividing the video data is determined to be 32*32.

[0069] S133: Divide the video data into corresponding second target blocks based on the initial block division method.

[0070] In one embodiment, after obtaining the initial segmentation method, the video data can be divided into corresponding second target blocks according to the segmentation rules of the initial segmentation method using a segmentation component.

[0071] S134: Encode and predict the second target block to obtain the second prediction block.

[0072] In one embodiment, after obtaining the second target block, an AI model can be used to encode and predict the second target block based on the context information of the surrounding adjacent blocks of the current block and the relevant region of the previous frame, so that the pixel difference (residual) between the generated second predicted block and the second target block is smaller, thereby improving the prediction accuracy and finally obtaining the second predicted block.

[0073] This invention, through determining semantic tags and weight information based on semantic feature information, and determining the initial block segmentation method based on the weight information, compared to the fixed preset block segmentation method in the prior art, divides the second target block by the initial block segmentation method corresponding to the weight information, and then predicts the second target block to obtain the second prediction block. This allows the block size of the second target block and the second prediction block to change with the change of weight information, improving the flexibility of block segmentation and avoiding the phenomenon in related technologies where the block size is fixed, which requires the transmission of unnecessary data, thereby reducing the video encoding bitrate.

[0074] In one embodiment, the video data is divided into corresponding second target blocks based on the initial block division method, such as... Figure 4 As shown, it may include steps S1331-S1333.

[0075] S1331: Divide the video data according to the initial block division method to obtain the third target block.

[0076] In one embodiment, after obtaining the initial segmentation method, the video data can be divided into corresponding third target blocks according to the segmentation rules of the initial segmentation method using a segmentation component.

[0077] S1332: When the pixel density of the semantic boundary of the third target block is less than or equal to a preset density threshold, the third target block is used as the second target block.

[0078] Among them, semantic boundaries represent the dividing lines between different semantic objects.

[0079] In this embodiment, the detection component can be used to detect whether the pixel density of the semantic boundary of the third target block is less than or equal to a preset density threshold. If the pixel density of the semantic boundary of the third target block is less than or equal to the preset density threshold, it indicates that the semantic objects in the third target block are the same semantic objects, that is, the third target block does not need to be further divided, and the third target block as a whole can be used as the second target block for prediction encoding. Therefore, the third target block is used as the second target block.

[0080] S1333: When the semantic boundary pixel density of the third target block is greater than the preset density threshold, the third target block is divided into multiple fourth target blocks; the fourth target blocks are used as the second target blocks.

[0081] In one embodiment, when the detection component detects that the semantic boundary pixel density of the third target block is greater than a preset density threshold, the semantic objects in the third target block are characterized as including different semantic objects. The third target block is then divided into multiple fourth target blocks such that each semantic object in the third target block corresponds to one fourth target block. The fourth target blocks are then used as the second target blocks.

[0082] This invention improves the accuracy of determining the second target block by determining whether the third target block needs to be further subdivided by determining that the pixel density of the semantic boundary of the third target block is less than or equal to a preset density threshold.

[0083] In one embodiment, the second coding bitrate of the second prediction block is determined based on the first prediction block, the extracted bitrate, and the video data, such as... Figure 5 As shown, steps S141-S143 may be included.

[0084] S141: Obtain the number of pixels in the video data and the number of pixels in the first prediction block.

[0085] In one embodiment, the number of pixels in the video data and the number of pixels in the first prediction block can be obtained using an acquisition component.

[0086] S142: Divide the square of the number of pixels in the first prediction block by the square of the number of pixels in the video data to obtain the prediction block size ratio.

[0087] In one embodiment, a computing component can be used to calculate the square of the number of pixels in the first prediction block and the square of the number of pixels in the video data, respectively, and then the square of the number of pixels in the first prediction block can be divided by the square of the number of pixels in the video data to obtain the prediction block size ratio.

[0088] S143: Multiply the predicted block size ratio by the extraction code rate to obtain the second coding code rate.

[0089] In one embodiment, after calculating the predicted block size ratio, the calculation component can be used to multiply the predicted block size ratio by the extraction code rate to obtain the second coding code rate.

[0090] In one example, the expression for calculating the second coding rate is shown in formula (1).

[0091] Where a represents the second coding bitrate, b represents the extraction bitrate, c represents the number of pixels in the first prediction block, and d represents the square of the number of pixels in the video data.

[0092] This invention obtains the number of pixels in the video data and the number of pixels in the first prediction block by dividing the square of the number of pixels in the first prediction block by the square of the number of pixels in the video data to obtain the prediction block size ratio. Finally, the prediction block size ratio is multiplied by the extraction bitrate to determine the second coding bitrate required for predicting a block of the same size as the first prediction block in this prediction mode. This ensures that both the second coding bitrate and the first coding bitrate are the bitrates required to transmit prediction blocks of the same size, facilitating subsequent comparison of the first rate-distortion cost and the second rate-distortion cost, thereby improving the accuracy of determining the minimum rate-distortion cost.

[0093] In one embodiment, the sum of squares of the pixel differences between the first prediction block, the second prediction block, and the target block is calculated to obtain the first rate-distortion cost and the second rate-distortion cost, respectively. Figure 6 As shown, steps S151-S1510 may be included.

[0094] S151: Obtain the first pixel difference based on the difference between the pixels of the first target block and the pixels of the first prediction block corresponding to the first block prediction mode.

[0095] In one embodiment, a computing component can be used to calculate the difference between the pixels of the first target block corresponding to the first block prediction mode and the pixels of the first prediction block to obtain the first pixel difference.

[0096] S152: Divide the sum of squares of the first pixel differences by the number of pixels in the first target block to obtain the first distortion value.

[0097] In one embodiment, the first distortion value can be obtained by dividing the sum of squares of the first pixel differences by the number of pixels in the first target block using a calculation component.

[0098] S153: Determine the control information bit rate based on the prediction mode type corresponding to the first prediction block.

[0099] Among them, the control information bitrate represents the bitrate corresponding to the number of bits consumed by the encoding mode, syntax elements, and metadata required for decoding.

[0100] In one embodiment, a detection tool can be used to detect the prediction mode type corresponding to the first prediction block. If the prediction mode type is inter-frame prediction mode, the motion vector difference generated by the pixel displacements of adjacent frames produced by inter-frame prediction, the reference frame index, and the prediction direction are encoded to obtain the control information bitrate. If the prediction mode type is intra-frame prediction mode, the prediction mode index generated by intra-frame prediction is encoded to obtain the control information bitrate.

[0101] S154: Use the difference between the pixel matrix corresponding to the first target block and the pixel matrix corresponding to the first prediction block as the target residual matrix.

[0102] In one embodiment, the difference between the pixel matrix corresponding to the first target block and the pixel matrix corresponding to the first prediction block can be calculated using a computing component, and used as the target residual matrix.

[0103] S155: Determine the residual data bitrate based on the values ​​in the target residual matrix.

[0104] In one embodiment, the target residual matrix can be transformed and quantized to make the values ​​in the target residual matrix become quantized transformation coefficients. Then, entropy-based coding operations can be performed on the quantized transformation coefficients to obtain the residual coefficient code rate, i.e., the residual data code rate.

[0105] S156: Use the control information bitrate and residual data bitrate as the coding bitrate corresponding to the first prediction block.

[0106] In one embodiment, after obtaining the control information bitrate and the residual data bitrate, the control information bitrate and the residual data bitrate can be added together using a computing component to obtain the coding bitrate corresponding to the first prediction block.

[0107] S157: Add the product of the first distortion value and the first preset weight and the coding rate corresponding to the first prediction block to obtain the first rate distortion cost.

[0108] In one embodiment, a computing component can be used to calculate the product of the first preset weight and the coding rate corresponding to the first prediction block, and then the first distortion value can be added to the product of the first preset weight and the coding rate corresponding to the first prediction block to obtain the first rate distortion cost.

[0109] S158: Obtain the second pixel difference based on the difference between the pixel in the second target block and the corresponding pixel in the second prediction block.

[0110] In one embodiment, the difference between a pixel in the second target block and the corresponding pixel in the second prediction block can be calculated using a computing component to obtain the second pixel difference.

[0111] S159: The second distortion value is obtained by summing the squares of the second pixel differences and dividing by the number of pixels in the second target block.

[0112] In one embodiment, the sum of squares of the second pixel differences can be calculated using a computing component, and then the sum of squares of the second pixel differences can be divided by the number of pixels in the second target block to obtain the second distortion value.

[0113] S1510: Add the second distortion value to the product of the second preset weight and the second coding rate to obtain the second rate distortion cost.

[0114] In one embodiment, a computing component can be used to calculate the product of a second preset weight and a second coding rate, and then the second distortion value can be added to the product of the second preset weight and the second coding rate to obtain a second rate distortion value.

[0115] In one embodiment, the expression for calculating the rate distortion cost is shown in Equation (2).

[0116] Wherein, G represents the rate-distortion cost, I represents the distortion value, O represents the preset weight, and P represents the coding rate.

[0117] In this embodiment of the invention, by obtaining a first coding bitrate and a second coding bitrate, the first distortion value is added to the product of a first preset weight and the first coding bitrate, and the second distortion value is added to the product of a second preset weight and the second coding bitrate. Since the distortion value reflects image quality and the coding bitrate reflects compression efficiency, the first rate-distortion value and the second rate-distortion value are determined based on the distortion value and the coding bitrate, thus achieving a balance between image quality and compression efficiency.

[0118] In one embodiment, encoding is performed based on the residual matrix corresponding to the difference between the pixels of the second predicted block and the target block to obtain a video encoded stream, such as... Figure 7 As shown, steps S161-S163 may be included.

[0119] S161: Transform the residual matrix from the spatial domain to the frequency domain to obtain the transformation matrix.

[0120] In one embodiment, a floating-point Discrete Cosine Transform (DCT) kernel can be used to convert the residual matrix from the spatial domain to the frequency domain, resulting in a transformation matrix.

[0121] In one embodiment, the residual matrix is ​​transformed from the spatial domain to the frequency domain to obtain the transformation matrix, such as... Figure 8 As shown, it may include steps S1611 and S1612.

[0122] S1611: Obtain the forward transformation kernel matrix and the transpose of the forward transformation kernel matrix based on the floating-point discrete cosine transform (DCT) kernel.

[0123] In one embodiment, the elements of each row and column of the forward transformation kernel matrix can be calculated sequentially using formula (3) corresponding to the floating-point discrete cosine transform kernel, thereby obtaining the forward transformation kernel matrix. Then, the forward transformation kernel matrix is ​​transposed to obtain the transpose matrix of the forward transformation kernel matrix.

[0124] in, The element in the u-th row and x-th column of the matrix represents the element in the matrix. Characterized by the normalization coefficient, when u=0, = When u>0, = . The number of rows or columns of the residual matrix is ​​also the number of rows or columns of the forward transformation kernel matrix.

[0125] S1612: Multiply the product of the forward transformation kernel matrix and the residual matrix by the transpose matrix to obtain the transformation matrix.

[0126] In one embodiment, after obtaining the forward transformation kernel matrix and the transpose matrix, the forward transformation kernel matrix can be multiplied by the residual matrix using a computing component, and the product can then be multiplied by the transpose matrix to obtain the transformation matrix.

[0127] S162: Divide the values ​​in the transformation matrix by the preset quantization step size to obtain the quantization matrix.

[0128] The quantization step size represents the fineness of the quantization process and is used to control the balance between compression ratio and image quality.

[0129] In one embodiment, after obtaining the transformation matrix, the values ​​in the transformation matrix can be divided by a preset quantization step size using a computing component to obtain the quantization matrix.

[0130] S163: Entropy encoding is performed on the coefficients in the quantization matrix to obtain the video encoded stream.

[0131] In one embodiment, the quantization matrix can be converted into a one-dimensional matrix, and then the coefficients of the one-dimensional matrix can be entropy-modified to obtain the video encoded stream.

[0132] In one embodiment, entropy encoding is performed on the coefficients in the quantization matrix to obtain the video encoded stream, such as... Figure 9 As shown, it may include steps S1631 and S1632.

[0133] S1631: Sort the coefficients in the quantization matrix according to their numerical values ​​to obtain a one-dimensional sequence.

[0134] S1632: Based on the number of consecutive zero coefficients and the magnitude and sign of non-zero coefficients in the one-dimensional sequence, entropy coding is performed on the one-dimensional sequence to obtain the video coded stream.

[0135] In one example, the quantization matrix is The coefficients in the quantization matrix can be sorted by value to obtain a one-dimensional sequence (2,1,1,0). Then, the occurrence frequency of each symbol in the sequence is counted. For example, the non-zero coefficient 2 appears once, the non-zero coefficient 1 appears twice, and the zero coefficient 0 appears once. That is, the probability of 2 is 0.25, the probability of 1 is 0.5, and the probability of 0 is 0.25. Using the Huffman coding algorithm, the code length is allocated from highest to lowest probability: 1 (highest probability) is assigned code 0, 2 is assigned code 10, and 0 is assigned code 11. Following the order of the one-dimensional sequence, the final code is 100011, with a total code length of 6 bits.

[0136] The embodiments of the present invention, by sequentially performing transformation, quantization, and entropy encoding operations on the residual matrix, can solve the correlation redundancy problem of the residual matrix through quantization, solve the precision redundancy problem of the residual matrix through quantization, and solve the sign distribution redundancy problem of the residual matrix through entropy encoding, thereby improving the quality of the residual matrix compression encoding.

[0137] Figure 10 This is a schematic diagram of a video encoding apparatus 1000 provided in an embodiment of the present invention. Figure 10 As shown, the device 1000 includes a block prediction module 1001, an extraction module 1002, a determination module 1003, and a calculation module 1004.

[0138] Block prediction module 1001 is used to block and code prediction the acquired video data according to the first block prediction mode to obtain the first prediction block and the first coding bitrate, wherein the rate distortion cost of the first block prediction mode is less than a preset threshold. The extraction module 1002 is used to extract semantic features from video data using AI feature extraction tools and record the extraction bitrate of semantic features; The block prediction module 1001 is also used to perform block prediction on video data based on an AI model to obtain a second prediction block. The determination module 1003 is used to determine the second coding bitrate of the second prediction block based on the first prediction block, the extracted bitrate, and the video data; The calculation module 1004 is used to calculate the sum of squares of the pixel differences between the first prediction block, the second prediction block and the target block, respectively, and obtain the first rate distortion cost and the second rate distortion cost. The determination module 1003 is used to encode the video encoded stream based on the residual matrix corresponding to the difference between the pixels of the second prediction block and the target block when the first rate distortion cost is greater than the second rate distortion cost.

[0139] The embodiments of the present invention obtain the first rate-distortion cost and the second rate-distortion cost corresponding to the two block prediction modes respectively by using the first block prediction mode and the AI ​​model memory block prediction. The block prediction mode with the smallest rate-distortion cost is selected for encoding. Since this block prediction mode has the smallest rate-distortion cost, the video encoding bitrate can be reduced.

[0140] In one embodiment, the block prediction module 1001 is specifically used for: Video data is divided into first target blocks corresponding to at least one block prediction mode, resulting in multiple first target blocks; Each type of first target block is predicted separately, resulting in multiple first prediction blocks; Based on each first target block, each first prediction block, and the coding rate corresponding to the first prediction block, multiple initial first rate distortion costs corresponding to the first target block are obtained; The initial first rate distortion cost with the smallest value is selected from multiple initial first rate distortion costs as the first rate distortion cost, and the block prediction mode corresponding to the first rate distortion cost is taken as the first prediction mode.

[0141] In one embodiment, the block prediction module 1001 is specifically used for: Inter-frame prediction and intra-frame prediction are performed for each type of first target block to obtain multiple first prediction blocks.

[0142] In one embodiment, the video encoding apparatus 1000 includes an acquisition module for acquiring weight information of semantic tags; Determine the corresponding initial block division method based on the weight information; Based on the initial block division method, the video data is divided into corresponding second target blocks; The second target block is encoded and predicted to obtain the second prediction block.

[0143] In one embodiment, the block prediction module 1001 is specifically used for: The video data is divided according to the initial block division method to obtain the third target block; When the pixel density of the semantic boundary of the third target block is less than or equal to a preset density threshold, the third target block is used as the second target block. When the semantic boundary pixel density of the third target block is greater than a preset density threshold, the third target block is divided into multiple fourth target blocks; the fourth target blocks are used as the second target blocks.

[0144] In one embodiment, the video encoding apparatus 1000 includes an acquisition module for acquiring the number of pixels in the video data and the number of pixels in the first prediction block; The prediction block size ratio is obtained by dividing the square of the number of pixels in the first prediction block by the square of the number of pixels in the video data. Multiply the predicted block size ratio by the extraction rate to obtain the second coding rate.

[0145] In one embodiment, the computing module 1004 is specifically used for: The first pixel difference is obtained based on the difference between the pixels of the first target block and the pixels of the first prediction block corresponding to the first block prediction mode. The first distortion value is obtained by dividing the sum of squares of the first pixel differences by the number of pixels in the first target block; The control information bitrate is determined based on the prediction mode type corresponding to the first prediction block; The difference between the pixel matrix corresponding to the first target block and the pixel matrix corresponding to the first prediction block is used as the target residual matrix; Determine the residual data bitrate based on the values ​​in the target residual matrix; The control information bitrate and the residual data bitrate are used as the coding bitrates corresponding to the first prediction block; The first rate-distortion cost is obtained by adding the product of the first distortion value, the first preset weight, and the coding rate corresponding to the first prediction block. The second pixel difference is obtained based on the difference between the pixels in the second target block and the corresponding pixels in the second prediction block; The second distortion value is obtained by summing the squares of the second pixel differences and dividing by the number of pixels in the second target block; The second rate-distortion cost is obtained by adding the product of the second preset weight and the second coding rate.

[0146] In one embodiment, the video encoding apparatus 1000 includes a conversion module for converting the residual matrix from the spatial domain to the frequency domain to obtain a conversion matrix; Divide the values ​​in the transformation matrix by the preset quantization step size to obtain the quantization matrix; Entropy encoding is performed on the coefficients in the quantization matrix to obtain the video encoded stream.

[0147] In one embodiment, the video encoding apparatus 1000 includes an acquisition module, which is used to acquire a forward transform kernel matrix and a transpose matrix of the forward transform kernel matrix based on a floating-point discrete cosine transform (DCT) kernel. The transformation matrix is ​​obtained by multiplying the product of the forward transformation kernel matrix and the residual matrix by the transpose matrix.

[0148] In one embodiment, the video encoding apparatus 1000 includes a sorting module for sorting the coefficients in the quantization matrix according to their numerical values ​​to obtain a one-dimensional sequence. Entropy coding is performed on the one-dimensional sequence based on the number of consecutive zero coefficients and the magnitude and sign of the non-zero coefficients to obtain the video coded stream.

[0149] Figure 11 A schematic diagram of the hardware structure for video encoding provided in an embodiment of the present invention is shown.

[0150] The video encoding device may include a processor 1101 and a memory 1102 storing computer program instructions.

[0151] Specifically, the processor 1101 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of the present invention.

[0152] Memory 1102 may include mass storage for data or instructions. For example, and not limitingly, memory 1102 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 1102 may include removable or non-removable (or fixed) media. Where appropriate, memory 1102 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 1102 is non-volatile solid-state memory.

[0153] Memory 1102 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this disclosure.

[0154] The processor 1101 implements any of the video encoding methods described in the above embodiments by reading and executing computer program instructions stored in the memory 1102.

[0155] In one example, the video encoding device may also include a communication interface 1103 and a bus 1104. For example, Figure 11 As shown, the processor 1101, memory 1102, and communication interface 1103 are connected through bus 1104 and complete communication with each other.

[0156] The communication interface 1103 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of the present invention.

[0157] Bus 1104 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 1104 may include one or more buses. Although specific buses are described and illustrated in embodiments of the invention, the invention contemplates any suitable bus or interconnect. Additionally, in conjunction with the video encoding methods described in the above embodiments, embodiments of the invention also provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by the processor, they implement any of the video encoding methods described in the above embodiments.

[0158] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the video encoding methods described in the above embodiments. This invention further provides a vehicle, including a control device and a terminal. The terminal is used to implement the video encoding methods described in the above embodiments, and the control device is used to implement the video encoding methods described in the above embodiments.

[0159] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.

[0160] The functional blocks shown in the above structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0161] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0162] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0163] The above are merely specific embodiments of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A video encoding method, characterized in that, The method includes: The acquired video data is divided into blocks and encoded according to the first block prediction mode to obtain the first prediction block and the first coding bitrate. The rate distortion cost of the first block prediction mode is less than a preset threshold. The semantic features of the video data are extracted using an AI feature extraction tool, and the extraction bitrate of the semantic features is recorded. The video data is segmented and predicted based on an AI model to obtain a second prediction block. Based on the first prediction block, the extracted bitrate, and the video data, the second coding bitrate of the second prediction block is determined; The first rate-distortion cost and the second rate-distortion cost are obtained by calculating the sum of squares of the pixel differences between the first prediction block, the second prediction block and the target block, respectively. If the first rate-distortion cost is greater than the second rate-distortion cost, the video encoded stream is obtained by encoding based on the residual matrix corresponding to the difference between the pixels of the second prediction block and the target block.

2. The method according to claim 1, characterized in that, Before performing block segmentation and encoding prediction on the acquired video data according to the first block prediction mode to obtain the first prediction block and the first encoding bitrate, the method further includes: The video data is divided into first target blocks corresponding to at least one block prediction mode to obtain multiple first target blocks; Each type of first target block is predicted separately, resulting in multiple first prediction blocks; Based on each first target block, each first prediction block, and the coding rate corresponding to the first prediction block, multiple initial first rate distortion costs corresponding to the first target block are obtained; The initial first rate distortion cost with the smallest value is selected from the plurality of initial first rate distortion costs as the first rate distortion cost, and the block prediction mode corresponding to the first rate distortion cost is selected as the first prediction mode.

3. The method according to claim 2, characterized in that, The process of predicting each type of first target block yields multiple first predicted blocks, including: Inter-frame prediction and intra-frame prediction are performed for each type of first target block to obtain multiple first prediction blocks.

4. The method according to claim 1, characterized in that, The semantic features include semantic tags, and the step of segmenting and predicting the video data based on an AI model to obtain a second prediction block includes: Obtain the weight information of the semantic tags; The corresponding initial block division method is determined based on the weight information; Based on the initial segmentation method, the video data is divided into corresponding second target blocks; The second target block is encoded and predicted to obtain the second predicted block.

5. The method according to claim 4, characterized in that, The step of dividing the video data into corresponding second target blocks based on the initial block division method includes: The video data is divided according to the initial block division method to obtain the third target block; When the pixel density of the semantic boundary of the third target block is less than or equal to a preset density threshold, the third target block is used as the second target block; When the semantic boundary pixel density of the third target block is greater than a preset density threshold, the third target block is divided into multiple fourth target blocks; the fourth target blocks are used as the second target block.

6. The method according to claim 4, characterized in that, The step of determining the second coding bitrate of the second prediction block based on the first prediction block, the extracted bitrate, and the video data includes: Obtain the number of pixels in the video data and the number of pixels in the first prediction block; The prediction block size ratio is obtained by dividing the square of the number of pixels in the first prediction block by the square of the number of pixels in the video data. Multiply the predicted block size ratio by the extraction code rate to obtain the second coding code rate.

7. The method according to claim 4, characterized in that, The step of calculating the sum of squares of the pixel differences between the first prediction block, the second prediction block, and the target block, respectively, and obtaining the first rate-distortion cost and the second rate-distortion cost, includes: The first pixel difference is obtained based on the difference between the pixels of the first target block corresponding to the first block prediction mode and the pixels of the first prediction block. The first distortion value is obtained by dividing the sum of squares of the first pixel differences by the number of pixels in the first target block; The control information code rate is determined based on the prediction mode type corresponding to the first prediction block; The difference between the pixel matrix corresponding to the first target block and the pixel matrix corresponding to the first prediction block is used as the target residual matrix; The residual data bitrate is determined based on the values ​​in the target residual matrix. The control information code rate and the residual data code rate are used as the coding code rate corresponding to the first prediction block; The first rate-distortion cost is obtained by adding the product of the first distortion value, the first preset weight, and the coding rate corresponding to the first prediction block. The second pixel difference is obtained based on the difference between the pixels in the second target block and the corresponding pixels in the second prediction block; The second distortion value is obtained by summing the squares of the second pixel differences and dividing by the number of pixels in the second target block; The second rate-distortion cost is obtained by adding the second distortion value to the product of the second preset weight and the second coding rate.

8. The method according to claim 1, characterized in that, The step of encoding based on the residual matrix corresponding to the difference between the pixels of the second predicted block and the target block to obtain a video encoded stream includes: The residual matrix is ​​transformed from the spatial domain to the frequency domain to obtain the transformation matrix; Divide the values ​​in the transformation matrix by the preset quantization step size to obtain the quantization matrix; The coefficients in the quantization matrix are entropy encoded to obtain the video encoded stream.

9. The method according to claim 8, characterized in that, The step of converting the residual matrix from the spatial domain to the frequency domain to obtain the transformation matrix includes: Obtain the forward transformation kernel matrix and the transpose of the forward transformation kernel matrix based on the floating-point discrete cosine transform (DCT) kernel; The transformation matrix is ​​obtained by multiplying the product of the forward transformation kernel matrix and the residual matrix by the transpose matrix.

10. The method according to claim 8, characterized in that, The step of entropy encoding the coefficients in the quantization matrix to obtain the video encoded stream includes: The coefficients in the quantization matrix are sorted according to their numerical values ​​to obtain a one-dimensional sequence; The one-dimensional sequence is entropy encoded based on the number of consecutive zero coefficients and the magnitude and sign of the non-zero coefficients to obtain the video encoded stream.

11. A video encoding apparatus, characterized in that, The device includes: The acquired video data is divided into blocks and encoded according to the first block prediction mode to obtain the first prediction block and the first coding bitrate. The rate distortion cost of the first block prediction mode is less than a preset threshold. The semantic features of the video data are extracted using an AI feature extraction tool, and the extraction bitrate of the semantic features is recorded. The video data is segmented and predicted based on an AI model to obtain a second prediction block. Based on the first prediction block, the extracted bitrate, and the video data, the second coding bitrate of the second prediction block is determined; The first rate-distortion cost and the second rate-distortion cost are obtained by calculating the sum of squares of the pixel differences between the first prediction block, the second prediction block and the target block, respectively. If the first rate-distortion cost is greater than the second rate-distortion cost, the video encoded stream is obtained by encoding based on the residual matrix corresponding to the difference between the pixels of the second prediction block and the target block.

12. A video encoding device, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the video encoding method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the video encoding method as described in any one of claims 1-10.

14. A computer program product, characterized in that, When the instructions in the computer program product are executed by a processor in an electronic device, the electronic device causes the electronic device to perform the video encoding method as described in any one of claims 1-10.