Method and electronic device for processing video frame images

By obtaining the complexity information of the video frame image block and the complexity information of the sub-image block, determining the division method and adjusting it, the problems of slow speed and inaccurate results of the AV1 encoder when dividing video frames are solved, and the encoding efficiency and quality are improved.

CN114745541BActive Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110016716.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-07
Publication Date
2025-05-13
Estimated Expiration
2041-01-07

AI Technical Summary

Technical Problem

When the AV1 encoder divides video frames into image blocks, there are problems such as low division speed and inaccurate division results.

Method used

By obtaining the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks divided by the image block to be encoded, the division method is determined, including coarse-grained division and fine-grained division, and determining whether the division method meets the preset conditions based on the encoded data and complexity information, and then adjusting the division method.

Benefits of technology

The video frame division speed is improved, the accuracy of division results is improved, and the encoding speed and quality balance is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114745541B_ABST
    Figure CN114745541B_ABST
Patent Text Reader

Abstract

Provided is a method and electronic device for processing a video frame image, comprising: obtaining an image block to be encoded from the video frame image; obtaining complexity information of the image block to be encoded and complexity information of multiple sub-image blocks divided from the image block to be encoded; determining a division method of the image block to be encoded based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, wherein the division method includes a coarse-grained division method and a fine-grained division method; dividing and encoding the image block to be encoded based on the division method to obtain the encoding data of the image block to be encoded; determining whether the division method meets a preset condition based on the encoding data of the image block to be encoded, the complexity information of the multiple sub-image blocks and the complexity information of the image block to be encoded. The present disclosure can improve the encoding speed, thereby achieving a balance between the encoding speed and the encoding quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method for processing a video frame image, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the rapid development of the digital video application industry chain, the limitations of the previous generation video compression solution VP9 have become increasingly prominent as video applications continue to develop in the direction of high definition, high frame rate, and high compression rate. Therefore, the Alliance for Open Media (AOMedia) developed AOMedia Video 1 (AV1), a video coding format to improve the VP9 solution. The goal of AV1 is to improve compression efficiency while maintaining the same or minimal image quality loss and coding bit rate.

[0003] AV1 is a block-based coding scheme, which means that the video frame is first divided into multiple non-overlapping image blocks, and each image block is encoded using the same or different coding methods to obtain the code stream of each image block. Generally, image areas with more details or high complexity should be divided into more and smaller image blocks to ensure image quality. Image areas with less details or lower complexity should be divided into larger image blocks to compress the video smaller.

[0004] However, the current AV1 encoder still has the defects of low division speed and inaccurate division results when dividing video frames into image blocks. To this end, the present disclosure provides a video processing method to increase the division speed and improve the accuracy of the division results.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0006] An embodiment of the present disclosure provides a method for processing a video frame image, comprising: obtaining an image block to be encoded from the video frame image; obtaining complexity information of the image block to be encoded and complexity information of multiple sub-image blocks divided from the image block to be encoded; determining a division method of the image block to be encoded based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, the division method including a coarse-grained division method and a fine-grained division method; dividing and encoding the image block to be encoded based on the division method to obtain encoding data of the image block to be encoded; determining whether the division method meets a preset condition based on the encoding data of the image block to be encoded, the complexity information of the multiple sub-image blocks and the complexity information of the image block to be encoded; if the division method meets the preset condition, stopping the search for the division method of the image block to be encoded; if the division method does not meet the preset condition, continuing the search for the division method of the image block to be encoded.

[0007] For example, the dividing and encoding of the image block to be encoded based on the division method to obtain the encoding data of the image block to be encoded also includes: when the division method is a coarse-grained division method, encoding the image block to be encoded as a whole to obtain the encoding data of the image block to be encoded; when the division method is a fine-grained division method, dividing the image block to be encoded into multiple fine-grained sub-image blocks, and encoding the multiple sub-image blocks respectively to obtain the encoding data of the image block to be encoded.

[0008] For example, when the division method does not satisfy the preset conditions, continuing to search for the division method of the image block to be encoded also includes: when the division method is a coarse-grained division method, dividing the image block to be encoded into a plurality of fine-grained sub-image blocks, and encoding the plurality of sub-image blocks respectively to obtain encoding data of the image block to be encoded; when the division method is a fine-grained division method, encoding the image block to be encoded as a whole to obtain encoding data of the image block to be encoded; and determining the encoding data of the image block to be encoded based on the encoding data obtained by encoding the image block to be encoded as a whole and the encoding data obtained by encoding the plurality of sub-image blocks respectively.

[0009] For example, in the case where the division method is a coarse-grained division method, determining whether the division method meets the preset conditions based on the encoding data of the image block to be encoded and the complexity information of the multiple sub-image blocks also includes: determining the encoding features of the image block to be encoded based on the encoding data of the image block to be encoded; obtaining a preset quantization parameter, and determining whether the division method meets the preset conditions based on the encoding features of the image block to be encoded, the complexity information of the multiple sub-image blocks and at least two of the quantization parameters.

[0010] For example, determining whether the division method meets the preset conditions also includes: determining a first value based on the coding characteristics of the image block to be encoded and the quantization parameter; when the first value is less than a first threshold, determining that the division method meets the preset conditions; when the first value is greater than the first threshold, determining a second value based on the coding characteristics of the image block to be encoded, the quantization parameter, and complexity information of the multiple sub-image blocks; when the second value is less than a second threshold, determining that the division method meets the preset conditions; when the second value is greater than the second threshold, determining that the division method does not meet the preset conditions.

[0011] For example, in the case where the division method is a fine-grained division method, determining whether the division method meets the preset conditions based on the encoding data of the image block to be encoded and the complexity information of the multiple sub-image blocks also includes: determining the rate-distortion loss of each sub-image block of the image block to be encoded based on the encoding data of the image block to be encoded; and determining whether the division method meets the preset conditions based on the complexity information of the multiple sub-image blocks and the rate-distortion loss of each sub-image block.

[0012] For example, determining whether the division method meets the preset conditions also includes: determining a third value based on the complexity information of the multiple sub-image blocks and the rate-distortion loss of each sub-image block, and when the third value is greater than a third threshold, determining that the division method meets the preset conditions; when the third value is less than the third threshold, determining the prediction mode of each sub-image block of the image block to be encoded based on the encoding data of the image block to be encoded; determining a fourth value based on the complexity information of the multiple sub-image blocks and the prediction mode of each sub-image block; when the fourth value is greater than a fourth threshold, determining that the division method meets the preset conditions; when the fourth value is less than the fourth threshold, determining that the division method does not meet the preset conditions.

[0013] For example, the method of determining the division method of the image block to be encoded based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks also includes: based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, using a division method analysis model to determine the division method of the image block to be encoded, wherein the input of the division method analysis model is: the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks; the output of the division method analysis model is: the probability of mutual correlation between the multiple sub-image blocks.

[0014] For example, the method of determining the division method of the image block to be encoded also includes: when the probability of mutual association between the multiple sub-image blocks is greater than a preset association threshold, determining that the division method is a fine-grained division method; when the probability of mutual association between the multiple sub-image blocks is less than a preset association threshold, determining that the division method is a coarse-grained division method.

[0015] For example, the complexity information includes one or more of the following: intra-frame coarse prediction loss of a video frame, inter-frame coarse prediction loss of a video frame, prediction rate-distortion loss of a video frame, the sum of predicted absolute errors between an image block to be encoded and an image block reconstructed based on encoding data, the sum of squared predicted differences between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean absolute difference between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean squared error between an image block to be encoded and an image block reconstructed based on encoding data, spatial information and / or temporal information.

[0016] For example, the coding features include one or more of the following: coding mode, coding division depth, rate-distortion loss, the sum of absolute errors between the image block to be encoded and the image block reconstructed based on the encoding data, the sum of squared differences between the image block to be encoded and the image block reconstructed based on the encoding data, the average absolute difference between the image block to be encoded and the image block reconstructed based on the encoding data, and the average square error between the image block to be encoded and the image block reconstructed based on the encoding data.

[0017] For example, the encoded data is a code stream generated by encoding the image block to be encoded.

[0018] For example, determining the encoding data of the image block to be encoded based on the encoding data obtained by encoding the image block to be encoded as a whole and the encoding data obtained by encoding multiple sub-image blocks separately also includes: determining a first rate-distortion loss of the image block to be encoded according to the encoding data obtained by encoding the image block to be encoded as a whole; determining a second rate-distortion loss of the image block to be encoded according to the encoding data obtained by encoding multiple sub-image blocks separately; when the first rate-distortion loss is lower than the second rate-distortion loss, using the encoding data obtained by encoding the image block to be encoded as a whole as the encoding data of the image block to be encoded; when the first rate-distortion loss is higher than the second rate-distortion loss, using the encoding data obtained by encoding the multiple sub-image blocks separately as the encoding data of the image block to be encoded.

[0019] For example, obtaining the complexity information of the image block to be encoded also includes: using a first complexity analysis model to obtain the complexity information of the image block to be encoded, wherein the input of the first complexity analysis model is the pixel value of each pixel in the image block to be encoded; and the output of the first complexity analysis model is the complexity of the image block to be encoded.

[0020] For example, the method of obtaining the complexity information of the image block to be encoded also includes: using a second complexity analysis model to obtain the complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided, wherein the input of the second complexity analysis model is the pixel value of each pixel in the plurality of fine-grained sub-image blocks into which the image block to be encoded is divided; and the output of the second complexity analysis model is one or more of the sum of the complexities of the fine-grained sub-image blocks into which the image block to be encoded is divided, the mean of the complexity, and the weighted sum of the complexity.

[0021] An embodiment of the present disclosure provides a method for processing a video frame image, comprising: obtaining accelerated encoding enabling information input by a user, wherein the accelerated encoding enabling information indicates: in a process of determining a division method of an image block to be encoded in a video frame image, using complexity information of the image block to be encoded and complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided to determine a division method of the image block to be encoded, the division method including a coarse-grained division method and a fine-grained division method; encoding the video frame image based on the accelerated encoding enabling information.

[0022] An embodiment of the present disclosure provides a method for processing a video frame image, comprising: obtaining an image block to be encoded from the video frame image; obtaining complexity information of the image block to be encoded and complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided; determining a division method of the image block to be encoded based on the complexity information of the image block to be encoded and the complexity information of the plurality of sub-image blocks; when the division method is a coarse-grained division method, encoding the image block to be encoded as a whole to obtain encoding data of the image block to be encoded; when the division method is a fine-grained division method, dividing the image block to be encoded into a plurality of fine-grained sub-image blocks, and encoding the plurality of sub-image blocks respectively to obtain encoding data of the image block to be encoded.

[0023] An embodiment of the present disclosure provides a method for processing a video frame image, comprising: obtaining an image block to be encoded from the video frame image; obtaining complexity information of the image block to be encoded; based on the complexity information of the image block to be encoded, determining a division method of the image block to be encoded, the division method comprising a coarse-grained division method and a fine-grained division method; dividing and encoding the image block to be encoded based on the division method to obtain encoding data of the image block to be encoded; based on the encoding data of the image block to be encoded, determining encoding features of the image block to be encoded; based on the encoding features of the image block to be encoded and the complexity information of the image block to be encoded, determining whether the division method meets a preset condition; if the division method meets the preset condition, stopping the search for the division method of the image block to be encoded; if the division method does not meet the preset condition, continuing the search for the division method of the image block to be encoded.

[0024] An embodiment of the present disclosure provides an electronic device, which includes: one or more processors; and one or more memories, wherein the memories store computer-readable codes, and when the computer-readable codes are executed by the one or more processors, the above method is executed.

[0025] According to another embodiment of the present disclosure, a computer-readable storage medium is provided, on which instructions are stored. When the instructions are executed by a processor, the processor executes the above method.

[0026] According to another aspect of the present disclosure, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various aspects or various optional implementations of the above-mentioned various aspects.

[0027] Therefore, the embodiment of the present disclosure can obtain a more accurate division method by combining the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks into which the image block to be encoded is divided. At the same time, the embodiment of the present disclosure can determine whether the image block to be encoded needs to jump out of the search for the division method at the current size by comparing the encoding features, quantization parameters, and complexity information of multiple sub-image blocks of the image block to be encoded, thereby improving the encoding speed. The embodiment of the present disclosure combines the complexity of the image block to be encoded with the quantization parameters and / or the complexity information of multiple sub-image blocks, which can realize a rapid judgment of the division method and quickly jump out of the division process of the image block to be encoded at the current size, so as to improve the encoding speed as much as possible without affecting the encoding bit rate and the user's subjective experience, thereby achieving a balance between the encoding speed and the encoding quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solution of the embodiment of the present disclosure, the following briefly introduces the drawings required for describing the embodiment. The drawings described below are only exemplary embodiments of the present disclosure.

[0029] Figure 1A It is a structural schematic diagram of an AV1 video encoding framework according to an embodiment of the present disclosure.

[0030] Figure 1B It is a schematic diagram of image block division in AV1 video encoding according to an embodiment of the present disclosure.

[0031] Figure 2A This is a first flow chart of a method for processing a video frame image according to an embodiment of the present disclosure.

[0032] Figure 2B is a second flow chart of the method for processing video frame images according to an embodiment of the present disclosure.

[0033] Figure 2C is a third flow chart of the method for processing video frame images according to an embodiment of the present disclosure.

[0034] Figure 2D It is a schematic diagram of dividing a video frame image according to an embodiment of the present disclosure.

[0035] Figure 3A is another schematic diagram of a method for processing a video frame image according to an embodiment of the present disclosure.

[0036] Figure 3B It is a schematic diagram of determining whether a division method satisfies a preset condition when the division method is a coarse-grained division method according to an embodiment of the present disclosure.

[0037] Figure 3C It is a schematic diagram of determining whether a division method satisfies a preset condition when the division method is a fine-grained division method according to an embodiment of the present disclosure.

[0038] Figure 3D It is a schematic diagram comparing the advantages and disadvantages of a coarse-grained partitioning method and a fine-grained partitioning method according to an embodiment of the present disclosure.

[0039] Figure 4A It is a schematic diagram of an interface of a method for processing video frame images according to an embodiment of the present disclosure.

[0040] Figure 4B It is a schematic flowchart of a method for processing a video frame image according to an embodiment of the present disclosure.

[0041] Figure 5 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0042] Figure 6 A schematic diagram showing the architecture of an exemplary computing device according to an embodiment of the present disclosure.

[0043] Figure 7 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the present disclosure more obvious, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described here.

[0045] In this specification and the accompanying drawings, substantially the same or similar steps and elements are represented by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance or ranking.

[0046] Several concepts related to the present disclosure are introduced below.

[0047] The quantization parameter is the serial number of the quantization step size Qstep. For luminance (Luma) coding, the quantization step size Qstep has a total of 52 values, QP takes values ​​0 to 51, and for chroma (Chroma) coding, QP takes values ​​0 to 39.

[0048] Capture frame rate: the number of video frames captured per second, in fps (frame per second).

[0049] Encoding frame rate: the number of video frames encoded per second, in fps (frame per second).

[0050] Original Image: The original encoded image that is input to the encoder.

[0051] Reconstructed image: The reconstructed image output by the decoder after encoding is completed.

[0052] Rate-Distortion Cost (RDCost): is a method to measure bit rate and distortion in video coding. Rate-distortion cost indicates the minimum distortion loss achieved at a given bit rate.

[0053] Spatial Information (SI): It represents the amount of spatial detail in a frame. The more complex the scene is, the higher the SI value is.

[0054] Temporal Information (TI): It characterizes the temporal variation of a video sequence. Sequences with higher degrees of motion usually have higher TI values.

[0055] Cloud Technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which is used on demand and is flexible and convenient. At present, the background services of the technical network system require a large amount of computing and storage resources, such as video websites, image websites and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, which needs to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data need strong system backing support, which can only be achieved through cloud computing.

[0056] At present, cloud technology is mainly divided into cloud basic technology and cloud application. Cloud basic technology can be further subdivided into: cloud computing, cloud storage, database and big data, etc.; cloud application can be further subdivided into: medical cloud, cloud Internet of Things, cloud security, cloud call, private cloud, public cloud, hybrid cloud, cloud gaming, cloud education, cloud conferencing, cloud social networking and artificial intelligence cloud services, etc.

[0057] The method for processing video frames according to the present disclosure may involve cloud computing and cloud storage under cloud technology.

[0058] Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, allowing various application systems to obtain computing power, storage space and information services as needed. The network that provides resources is called a "cloud". From the user's perspective, the resources in the "cloud" are infinitely scalable and can be obtained at any time, used on demand, expanded at any time, and paid for by use.

[0059] In the present disclosure, since determining the complexity of the current image block (and / or the multiple sub-image blocks into which the current image block is divided, and the granularity of the sub-image blocks is lower than that of the current image block) involves large-scale calculations and requires huge computing power and storage space, in the present disclosure, the terminal device can obtain sufficient computing power and storage space through cloud computing technology, and then execute the determination of the complexity of the original image block involved in the present disclosure, and generate the encoding of the original image block according to the complexity of the current image block (and / or the multiple fine-grained sub-image blocks into which the current image block is divided), and then generate the encoded data (code stream) of the video.

[0060] Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.

[0061] In the present disclosure, video frame images can be stored on the "cloud". When it is necessary to determine the complexity of the current image block (and / or the multiple fine-grained sub-image blocks into which the current image block is divided), and to determine the encoding strategy based on the complexity of the current image block (and / or the multiple fine-grained sub-image blocks into which the current image block is divided), the current image block (and / or the multiple fine-grained sub-image blocks into which the current image block is divided) of the current video frame, or image blocks at corresponding positions of multiple frames before and after the current video frame, can be pulled from the cloud storage device to reduce the storage pressure of the terminal device.

[0062] From a business perspective, the video encoding method involved in the present disclosure can be applied to business scenarios related to video encoding, such as video uploading, web conferencing, and online training (for example, online conferencing); from a technical principle perspective, the technical solution of the present disclosure can be directly applied to the AV1 video coding standard to improve the coding flexibility and coding speed when performing video encoding based on the AV1 video coding standard. The present disclosure uses the AV1 video coding standard as an example for illustration, and the technical solution of the present disclosure can also be applied to other known video coding standards, such as MPEG (Moving Picture Experts Group), HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), or can also be applied to other more advanced video coding standards, which the present disclosure does not limit.

[0063] The following first describes the working principle of the AV1 video coding framework corresponding to the AV1 video coding standard.

[0064] Figure 1A It is a structural schematic diagram of an AV1 video encoding framework according to an embodiment of the present disclosure. Figure 1B It is a schematic diagram of image block division in AV1 video encoding according to an embodiment of the present disclosure.

[0065] AV1's coding system uses a hybrid coding framework, which includes a mixture of multiple modules. In AV1, each module compresses the data redundancy of different aspects of the image from different angles and methods. As a result, AV1 can achieve relatively high performance.

[0066] like Figure 1A As shown, AV1 can take a certain video frame of a video as an input image and then divide the image into multiple image blocks. Figure 1A The block division method is schematically shown by white horizontal and vertical lines. Those skilled in the art should understand that AV1 can also use other block division methods for the video frame. The AV1 encoder will then process the image block as a unit. For example, the processing of the image block may include inter-frame prediction, intra-frame prediction, transformation (e.g., discrete cosine (DCT) transformation), quantization, entropy coding, loop filtering, film granularity synthesis, etc., so as to obtain compressed encoded data (e.g., bitstream).

[0067] The present disclosure mainly relates to the block division technology in the AV1 video coding standard. The block division technology is, for example, to divide an image into a plurality of rectangular image blocks, and then encode and decode the image in units of image blocks. In the current AV1 video coding standard, the size of the largest image block is 128 (pixels) x 128 (pixels), and the size of the smallest image block is 4 (pixels) x 4 (pixels). The size of the largest image block can be further divided into four equal parts or two equal parts. Figure 1B As shown, the four equal sub-image blocks ( Figure 1B The sub-image blocks marked with R in the figure can be further divided recursively, and each sub-image block can be further divided into fine-grained sub-image blocks according to at most nine division methods.

[0068] Therefore, for complex and diverse image content, different division methods can enable the AV1 encoder to perform the most effective encoding for image blocks of different sizes and complexities. Furthermore, different prediction modes (for example, directional prediction mode, recursive filtering mode, cross-component prediction mode, smoothing prediction mode, etc.) and processing methods can be used for different image blocks, thereby further improving the efficiency and quality of encoding.

[0069] Generally, for image blocks with complex scenes and more details (eg, higher complexity), a smaller partition size should be used, while for image blocks with simple scenes and fewer details (eg, lower complexity), a larger partition size should be used.

[0070] The following two solutions are currently proposed for how to find the optimal way to divide image blocks more quickly to increase the coding efficiency of the AV1 encoder.

[0071] Solution 1: When encoding a video, traverse all the division methods from top to bottom (for example, from the coarsest division method to the finest division method) or from bottom to top (for example, from the finest division method to the coarsest division method), and then compare the encoding effects of these division methods to determine the optimal division method. Solution 1 can usually find the optimal division method accurately. However, when encoding a video, if the recursive block division is adopted from top to bottom, for blocks with complex scenes, the optimal division size of the block is usually small, so many layers of recursion are required to find the optimal division method. If the recursive block division is adopted from bottom to top, for blocks with simple scenes, the optimal division size of the block is usually large, so if the division is from small-sized image blocks to large-sized image blocks, multiple layers of recursion are required to find the optimal division method. Although the encoding quality and compression efficiency of searching for the optimal block division method from top to bottom and bottom to top are relatively good, the encoding of the video with complex scenes by Solution 1 will bring a large speed loss.

[0072] Solution 2: When encoding a video, before determining whether to divide the current image block into smaller image blocks (fine-grained image blocks) each time, analyze the complexity of the current image block, and determine the division granularity of the current image block based on the complexity of the current image block. If the complexity of the current image block is high, it is encoded in a fine-grained division manner (for example, divided into four square sub-blocks for encoding). If the complexity is low, it is encoded in a coarse-grained division manner (for example, directly encoding without division). However, although Solution 2 can quickly determine the division method to increase the encoding speed, in many cases it is likely to be inaccurate to determine the division method based solely on the complexity of the current image block, resulting in a large quality loss, and thus a relatively poor user experience.

[0073] Therefore, it is necessary to further improve the block partitioning technology so as to further improve the partitioning speed and the accuracy of the partitioning results.

[0074] Figure 2A is a flowchart of a method 21 for processing a video frame image according to an embodiment of the present disclosure. Figure 2B is a flowchart of a method 22 for processing a video frame image according to an embodiment of the present disclosure. Figure 2C is a flowchart of a method 23 for processing a video frame image according to an embodiment of the present disclosure. Figure 2D It is a schematic diagram of dividing a video frame image according to an embodiment of the present disclosure.

[0075] Methods 21 to 23 can be performed by a user terminal. The user terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. Methods 21 to 23 can also be performed by a network server. The network server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The above-mentioned terminals and servers can be directly or indirectly connected via wired or wireless communications, which is not limited in this disclosure. Methods 21 to 23 can also be performed by a combination of a user terminal and a network server, which is not limited in this disclosure.

[0076] like Figure 2A As shown, the method 21 for processing a video frame image according to an embodiment of the present disclosure includes the following steps.

[0077] In step S211, an image block to be encoded is obtained from the video frame image. Figure 1A As shown, the image block to be encoded can be any one of the image blocks.

[0078] In step S212, complexity information of the image block to be encoded and complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided are obtained.

[0079] For example, the granularity of the multiple sub-image blocks into which the image block to be encoded is divided is smaller than that of the image block to be encoded.

[0080] Optionally, the complexity information includes one or more of the following: intra-frame coarse prediction loss of a video frame, inter-frame coarse prediction loss of a video frame, prediction rate distortion loss of a video frame, the sum of predicted absolute errors between an image block to be encoded and an image block reconstructed based on encoding data, the sum of squared predicted differences between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean absolute difference between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean squared error between an image block to be encoded and an image block reconstructed based on encoding data, spatial information and / or temporal information. It is worth noting that the complexity information may also include less or more information, and the present disclosure is not limited thereto. Reference will be made later to FIG. 3A to FIG. 3D How to obtain the complexity information is further described, and the disclosure will not go into details here.

[0081] Among them, spatial information (SI) can refer to the temporal information (TI) of a video frame image. Spatial information (SI) represents the amount of spatial detail of a frame. The more complex the scene is, the higher the SI value. Temporal information (TI) represents the temporal change of a video frame sequence. A sequence with a higher degree of motion usually has a higher TI value.

[0082] The predicted rate-distortion loss of the video frame image is the relationship between the bit rate and the distortion in the video encoding process, which indicates the minimum distortion loss that can be achieved under a given bit rate.

[0083] like Figure 2D As shown, the image block of the to-be-encoded block can be divided into four sub-image blocks (e.g., a first sub-image block, a second sub-image block, a third sub-image block, and a fourth sub-image block). Based on the complexity information of the first sub-image block, the complexity information of the second sub-image block, the complexity information of the third sub-image block, and the complexity information of the fourth sub-image block, one or more of the sum of the complexities of the fine-grained sub-image blocks into which the to-be-encoded image block is divided, the complexity mean, and the complexity weighted sum can be obtained and used as the complexity information of the multiple sub-image blocks. The present disclosure is not limited to this.

[0084] Will refer to FIG. 3A to FIG. 3DHow to obtain the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks into which the image block to be encoded is divided is further described, and the present disclosure will not go into details here.

[0085] In step S213, based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, a division method of the image block to be encoded is determined, and the division method includes a coarse-grained division method and a fine-grained division method.

[0086] Optionally, the division method of the image block to be encoded determined in step S213 is a tendency division method, which indicates that, based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, the coarse-grained division method may be better than the fine-grained division method, or the fine-grained division method may be better than the coarse-grained division method.

[0087] The complexity information of each fine-grained sub-image block divided by the image block to be encoded has a certain reference significance for the division of the image block to be encoded. Through the complexity information of each fine-grained sub-image block divided by the image block to be encoded, a more accurate division method (for example, a tendency division method) can be obtained.

[0088] Wherein, determining the partitioning method of the image block to be encoded can be implemented, for example, using a function / module related to the search for the partitioning method in AV1. For example, the function / module can be "partition_search" or "av1_prune_partitions_before_search", etc. or a combination thereof. The present disclosure is not limited to this, and will include more or fewer functional modules as the relevant video coding standards evolve.

[0089] Optionally, the above coarse-grained division method indicates: the image block to be encoded is encoded as a whole to obtain the encoded data of the image block to be encoded. The encoded data is a code stream obtained by encoding the image block to be encoded. Figure 1B As shown in the squares filled with diagonal lines.

[0090] Optionally, the above-mentioned fine-grained division method indicates: dividing the image block to be encoded into a plurality of fine-grained sub-image blocks, and encoding the plurality of sub-image blocks respectively to obtain the encoded data of the image block to be encoded. The fine-grained division method can be as follows: Figure 1B As shown in the white squares in , each sub-image block can be further divided into fine-grained sub-image blocks in up to nine ways. Figure 1B As shown, the four equal sub-image blocks ( Figure 1B The sub-image blocks marked with R in the figure can be further recursively divided into fine-grained sub-image blocks.

[0091] Of course, with the further evolution of the AV1 video coding standard and other video coding standards, more division methods may emerge. The fine-grained division method and coarse-grained division method described in the present disclosure should be understood as relative concepts and include any possible division method.

[0092] The division method in the present disclosure can be determined by a trained division method analysis model based on the coding features of the coded surrounding image blocks and the complexity information of the image block to be coded. FIG. 3A to FIG. 3D How to determine the division method of the image block to be encoded through the division method analysis model is further described in the disclosure, and will not be repeated here.

[0093] In step S214, the image block to be encoded is divided and encoded based on the division method to obtain encoding data of the image block to be encoded.

[0094] Wherein, in the case where the division method is a coarse-grained division method, the division and encoding of the image block to be encoded based on the division method to obtain the encoding data of the image block to be encoded also includes: encoding the image block to be encoded as a whole to obtain the encoding data of the image block to be encoded.

[0095] Among them, in the case that the division method is a fine-grained division method, the division and encoding of the image block to be encoded based on the division method to obtain the encoding data of the image block to be encoded also includes: dividing the image block to be encoded into a plurality of fine-grained sub-image blocks, and encoding the plurality of sub-image blocks respectively to obtain the encoding data of the image block to be encoded.

[0096] The division of the image blocks to be encoded based on the division method can be implemented, for example, by selecting a related function / module using the division method in AV1. For example, the function / module can be "rd_pick_partition". The present disclosure is not limited to this, and will include more or fewer function modules as the relevant video coding standards evolve.

[0097] In step S215, based on the encoding data of the image block to be encoded, the complexity information of the multiple sub-image blocks and the complexity information of the image block to be encoded, it is determined whether the division method meets a preset condition.

[0098] For example, assuming that the complexity information refers to the predicted rate distortion loss of the video frame, the coding feature obtained based on the coded data is the actual rate distortion loss of the video frame. If the predicted rate distortion loss is much smaller than the actual rate distortion loss, it means that the division method does not meet the preset conditions.

[0099] For example, step S215 may also include the following steps. First, based on the coding data of the image block to be encoded, the coding features (for example, coding mode, coding division depth, rate distortion loss, etc.) of the image block to be encoded in the case of the division method are obtained. Then, based on the coding features of the image block to be encoded and the quantization parameter / complexity information of multiple sub-image blocks, calculations are performed to obtain the ratio of a certain coding feature to the quantization parameter, or the ratio of a certain coding feature to the complexity information of multiple sub-image blocks. Alternatively, calculations may also be performed based on the coding features of the image block to be encoded and the quantization parameter / complexity information of multiple sub-image blocks to obtain a certain value related to the two / three. If the ratio of the two or a certain value calculated based on the two meets the preset threshold, it means that the current division method meets the preset conditions, otherwise the division method does not meet the preset conditions.

[0100] The complexity information of multiple sub-image blocks and the complexity information of the image block to be encoded are obtained in step S212, which occurs before the image block to be encoded is encoded, so it is also called pre-analysis information. The encoded data based on the image block to be encoded is obtained through encoding, and the encoding features based on the division method can be further obtained from the encoded data. By jointly analyzing the encoding features and the pre-analysis information, it can be determined whether the division method is optimal (that is, whether the division method meets the preset conditions). If the encoding effect predicted based on the pre-analysis information is significantly different from the actual encoding effect obtained after encoding, it means that the judgment of the division method is not accurate. In this case, the encoded data based on the division method may cause greater losses.

[0101] The following is an example to describe how to determine the difference between the coding effect predicted based on the pre-analysis information and the coding effect actually obtained after coding.

[0102] For example, in the case where the division method is a coarse-grained division method, determining whether the division method meets the preset conditions based on the encoding data of the image block to be encoded and the complexity information of the multiple sub-image blocks also includes: determining the encoding features of the image block to be encoded based on the encoding data of the image block to be encoded; obtaining a preset quantization parameter, and determining whether the division method meets the preset conditions based on the encoding features of the image block to be encoded, the complexity information of the multiple sub-image blocks and at least two of the quantization parameters.

[0103] Optionally, when the division method is a coarse-grained division method, determining whether the division method meets the preset conditions also includes: determining a first value based on the coding characteristics of the image block to be encoded and the quantization parameter; when the first value is less than a first threshold, determining that the division method meets the preset conditions; when the first value is greater than the first threshold, determining a second value based on the coding characteristics of the image block to be encoded, the quantization parameter, and complexity information of the multiple sub-image blocks; when the second value is less than a second threshold, determining that the division method meets the preset conditions; when the second value is greater than the second threshold, determining that the division method does not meet the preset conditions.

[0104] For example, in the case where the division method is a fine-grained division method, determining whether the division method meets the preset conditions based on the encoding data of the image block to be encoded and the complexity information of the multiple sub-image blocks also includes: determining the rate-distortion loss of each sub-image block of the image block to be encoded based on the encoding data of the image block to be encoded; and determining whether the division method meets the preset conditions based on the complexity information of the multiple sub-image blocks and the rate-distortion loss of each sub-image block.

[0105] Optionally, in the case where the division method is a fine-grained division method, determining whether the division method meets the preset conditions also includes: determining a third value based on the complexity information of the multiple sub-image blocks and the rate-distortion loss of each sub-image block, and when the third value is greater than a third threshold, determining that the division method meets the preset conditions; when the third value is less than the third threshold, determining the prediction mode of each sub-image block of the image block to be encoded based on the encoding data of the image block to be encoded; determining a fourth value based on the complexity information of the multiple sub-image blocks and the prediction mode of each sub-image block; when the fourth value is greater than the fourth threshold, determining that the division method meets the preset conditions; when the fourth value is less than the fourth threshold, determining that the division method does not meet the preset conditions.

[0106] In step S216, when the division method satisfies a preset condition, the search for the division method of the image block to be encoded is stopped.

[0107] Among them, stopping the search for the partitioning method of the image block to be encoded can be implemented, for example, using a prune-related function / module in AV1. For example, the function / module can be "av1_prune_partitions_by_max_min_bsize", "partition_search_skippable", "partition_search_breakout", etc. or a combination thereof. The present disclosure is not limited to this, and will include more or less functional modules as the relevant video coding standards evolve.

[0108] For example, when the division method is a coarse-grained division method and the division method satisfies a preset condition, the coded data obtained by encoding the image block to be encoded as a whole is determined as the coded data of the image block to be encoded. At this point, further evaluation / search of the coarse-grained division method and the fine-grained division method of the image block to be encoded can be stopped, and there is no need to compare other performance indicators (for example, rate-distortion loss) of the coarse-grained division and the fine-grained division. The determined coded data of the image block to be encoded is the final coded data of the image block to be encoded, and there is no need to further judge other coding methods.

[0109] For example, when the division method is a fine-grained division method and the division method satisfies a preset condition, it is determined that the image block to be encoded should be divided into fine-grained divisions, and further evaluation / search of the coarse-grained division and fine-grained division of the image block to be encoded is stopped, and there is no need to compare other performance indicators (for example, rate-distortion loss) of the coarse-grained division method and the fine-grained division method. Next, the multiple sub-image blocks after division are respectively used as the image blocks to be encoded to determine the division method of each sub-image block. The determination of the division method of each sub-image block is similar to the determination of the division method of the image block to be encoded, and will not be repeated here.

[0110] In step S217, when the division method does not satisfy the preset condition, the search for the division method of the image block to be encoded continues.

[0111] Optionally, when the division mode is a coarse-grained division mode, continuing to search for the division mode of the image block to be encoded further includes dividing the image block to be encoded into a plurality of fine-grained sub-image blocks, encoding the plurality of sub-image blocks respectively to obtain the encoding data of the image block to be encoded (that is, dividing and encoding the image block to be encoded in a fine-grained division mode). Then, based on the encoding data obtained by encoding the image block to be encoded as a whole and the encoding data obtained by encoding the plurality of sub-image blocks respectively, the encoding data of the image block to be encoded is determined.

[0112] Optionally, when the division method is a fine-grained division method, continuing to search for the division method of the image block to be encoded also includes encoding the image block to be encoded as a whole to obtain the encoding data of the image block to be encoded (that is, dividing and encoding the image block to be encoded in a coarse-grained division method).

[0113] Then, based on the coded data obtained by encoding the image block to be encoded as a whole (i.e., the coded data based on the coarse-grained division method) and the coded data obtained by encoding multiple sub-image blocks respectively (i.e., the coded data based on the fine-grained division method), the coded data of the image block to be encoded is determined. For example, the coded data with the best coding feature is selected from the coded data based on the coarse-grained division method and the coded data based on the fine-grained division method. The coding feature can optionally be rate-distortion loss.

[0114] Continuing to search for the partitioning mode of the image block to be encoded can be implemented, for example, using functions / modules related to "partition_search" in AV1. The present disclosure is not limited thereto, and will include more or fewer functional modules as relevant video coding standards evolve.

[0115] Optionally, the above-mentioned determining the encoding data of the image block to be encoded based on the encoding data obtained by encoding the image block to be encoded as a whole and the encoding data obtained by encoding multiple sub-image blocks separately also includes: determining a first rate-distortion loss of the image block to be encoded according to the encoding data obtained by encoding the image block to be encoded as a whole; determining a second rate-distortion loss of the image block to be encoded according to the encoding data obtained by encoding multiple sub-image blocks separately; when the first rate-distortion loss is lower than the second rate-distortion loss, using the encoding data obtained by encoding the image block to be encoded as a whole as the encoding data of the image block to be encoded; when the first rate-distortion loss is higher than the second rate-distortion loss, using the encoding data obtained by encoding the multiple sub-image blocks separately as the encoding data of the image block to be encoded.

[0116] In the case where the coded data obtained by respectively encoding the plurality of sub-image blocks is used as the coded data of the image block to be encoded (that is, in the case where the coded data based on the fine-grained division method is used as the coded data of the image block to be encoded), the plurality of divided sub-image blocks are respectively used as the image blocks to be encoded to determine the division method of each sub-image block. The determination of the division method of each sub-image block is similar to the determination of the division method of the image block to be encoded, and will not be repeated here.

[0117] Scheme 1 iterates continuously to obtain the optimal block division method, which brings a lot of speed loss. Compared with Scheme 1, Method 21 can determine whether the division method meets the encoding requirements through the complexity information of the image block to be encoded and the encoding features obtained after the image block to be encoded is encoded (for example, prediction mode, encoding mode, encoding division depth, rate distortion loss, the absolute error sum between the image block to be encoded and the image block reconstructed based on the encoding data, the square sum of the difference between the image block to be encoded and the image block reconstructed based on the encoding data, the average absolute difference between the image block to be encoded and the image block reconstructed based on the encoding data, the average square error between the image block to be encoded and the image block reconstructed based on the encoding data, etc.). And when the preset conditions are met, the division search process under the current size can be quickly jumped out, which greatly improves the encoding efficiency.

[0118] Scheme 2 determines the division method only according to the complexity of the image block to be encoded and does not further determine the encoded data obtained based on the division method, thereby causing a large amount of quality loss. Compared with Scheme 2, Method 21 uses the complexity information of the image block to be encoded and the encoding features obtained after the image block to be encoded is encoded (for example, prediction mode, encoding mode, encoding division depth, rate distortion loss, the absolute error sum between the image block to be encoded and the image block reconstructed based on the encoding data, the square sum of the difference between the image block to be encoded and the image block reconstructed based on the encoding data, the average absolute difference between the image block to be encoded and the image block reconstructed based on the encoding data, the average square error between the image block to be encoded and the image block reconstructed based on the encoding data, etc.), and can determine whether the division method meets the encoding requirements. And if the preset conditions are not met, the division search process can be continued, and the accuracy of the encoding is guaranteed by comparing the first rate distortion loss and the second rate distortion loss.

[0119] Therefore, method 21 can obtain a more accurate division method by combining the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks divided from the image block to be encoded. At the same time, method 21 can determine whether to jump out of the search for the division method of the current size by comparing the encoding features of the image block to be encoded with the complexity information of the image block to be encoded (and / or the complexity information of the multiple sub-image blocks divided from the image block to be encoded), thereby improving the encoding speed. Method 21 combines the complexity information of the image block to be encoded (and / or the complexity information of the multiple sub-image blocks divided from the image block to be encoded) with the encoding features, which can realize the rapid judgment of the division method and quickly jump out of the division process of the image block to be encoded at the current size, so as to improve the encoding speed as much as possible without affecting the encoding bit rate and the user's subjective experience.

[0120] like Figure 2B As shown, the method 22 for processing a video frame image according to an embodiment of the present disclosure includes the following steps.

[0121] In step S221, an image block to be encoded is obtained from the video frame image. Step S221 is similar to step S211, and thus will not be described in detail.

[0122] In step S222, the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks into which the image block to be encoded is divided are obtained. Step S222 is similar to step S212, and thus will not be described in detail.

[0123] In step S223, based on the complexity information of the image block to be encoded and the complexity information of the plurality of sub-image blocks, a division method of the image block to be encoded is determined. Step S223 is similar to step S213, and thus will not be described in detail.

[0124] In step S224, when the division method is a coarse-grained division method, the image block to be encoded is encoded as a whole to obtain encoding data of the image block to be encoded.

[0125] In step S225, when the division method is a fine-grained division method, the image block to be encoded is divided into a plurality of fine-grained sub-image blocks, and the plurality of sub-image blocks are respectively encoded to obtain encoding data of the image block to be encoded.

[0126] Step S224 and step S225 are similar to step S214 and thus will not be described in detail.

[0127] The complexity information of each fine-grained sub-image block divided by the image block to be encoded has a certain reference significance for the division of the image block to be encoded. Method 22 can obtain a more accurate division method through the complexity information of each fine-grained sub-image block divided by the image block to be encoded.

[0128] like Figure 2C As shown, the method 23 for processing a video frame image according to an embodiment of the present disclosure includes the following steps.

[0129] In step S231, an image block to be encoded is obtained from the video frame image. Step S231 is similar to step S211, and thus will not be described in detail.

[0130] In step S232, the complexity information of the image block to be encoded is obtained. Step S232 is similar to the step S212 of obtaining the complexity information of the image block to be encoded, and thus will not be described in detail.

[0131] In step S233, based on the complexity information of the image block to be encoded, a division method of the image block to be encoded is determined, and the division method includes a coarse-grained division method and a fine-grained division method.

[0132] Optionally, the complexity information includes one or more of the following: intra-frame coarse prediction loss of a video frame, inter-frame coarse prediction loss of a video frame, prediction rate distortion loss of a video frame, the sum of predicted absolute errors between an image block to be encoded and an image block reconstructed based on encoding data, the sum of predicted squared differences between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean absolute difference between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean squared error between an image block to be encoded and an image block reconstructed based on encoding data, spatial information, and / or temporal information. It is worth noting that the complexity information may also include less or more information, and the present disclosure is not limited thereto.

[0133] In step S234, the image block to be encoded is divided and encoded based on the division method to obtain encoding data of the image block to be encoded. Step S234 is similar to step S214, so it is not repeated here.

[0134] In step S235, the encoding feature of the image block to be encoded is determined based on the encoding data of the image block to be encoded.

[0135] For example, the encoded data is a code stream generated by encoding the image block to be encoded.

[0136] For example, the coding features include one or more of the following: coding mode, coding division depth, rate-distortion loss, the sum of absolute errors between the image block to be encoded and the image block reconstructed based on the encoding data, the sum of squared differences between the image block to be encoded and the image block reconstructed based on the encoding data, the average absolute difference between the image block to be encoded and the image block reconstructed based on the encoding data, and the average square error between the image block to be encoded and the image block reconstructed based on the encoding data.

[0137] In step S236, based on the encoding features of the image block to be encoded and the complexity information of the image block to be encoded, it is determined whether the division method meets a preset condition.

[0138] The complexity information is obtained in step S232, which occurs before the image block to be encoded is encoded, so it is also called pre-analysis information. The encoded data based on the image block to be encoded is obtained through encoding, and the encoding features based on the division method can be further obtained from the encoded data. By jointly analyzing the encoding features and the pre-analysis information, it can be determined whether the division method is optimal (that is, whether the division method meets the preset conditions). If the encoding effect predicted based on the pre-analysis information is significantly different from the actual encoding effect obtained after encoding, it means that the judgment of the division method is not accurate. In this case, the encoded data based on the division method may cause greater losses.

[0139] For example, assuming that the complexity information refers to the predicted rate distortion loss of the video frame, the coding feature obtained based on the coded data is the actual rate distortion loss of the video frame. If the predicted rate distortion loss is much smaller than the actual rate distortion loss, it means that the division method does not meet the preset conditions.

[0140] For example, the coding feature of the image block to be encoded may be compared with the quantization parameter. If the ratio of the two or a value calculated based on the two meets the preset threshold, it means that the current division method meets the preset condition, otherwise the division method does not meet the preset condition. The present disclosure does not limit this.

[0141] In step S237, when the division method satisfies the preset condition, the search for the division method of the image block to be encoded is stopped. Step S237 is similar to step S216, so it is not described again.

[0142] In step S238, if the division method does not meet the preset condition, the division method of the image block to be encoded is continued to be searched. Step S238 is similar to step S217, so it is not repeated here.

[0143] Scheme 1 iterates continuously to obtain the optimal block division method, which results in a large amount of speed loss. Compared with Scheme 1, method 23 can determine whether the division method meets the encoding requirements (e.g., the rate-distortion loss requirements) through the complexity information of the image block to be encoded (e.g., the predicted rate-distortion loss) and the encoding characteristics of the image block to be encoded (e.g., the actual rate-distortion loss). And when the preset conditions are met, the division search process under the current size can be quickly jumped out, which greatly improves the encoding efficiency.

[0144] Scheme 2 determines the division method only according to the complexity of the image block to be encoded and does not further determine the encoded data obtained based on the division method, which results in a large amount of quality loss. Compared with Scheme 2, method 23 uses pre-analysis information (such as predicted encoding effect) and encoding features of the image block to be encoded (such as actual encoding effect) to determine whether the division method meets the encoding requirements. And if the preset conditions are not met, the division search process can be continued to ensure the accuracy of the encoding.

[0145] Therefore, method 23 can determine whether to jump out of the search of the division method of the current size by comparing the coding feature of the image block to be coded with the complexity information (i.e., pre-analysis information) of the image block to be coded, thereby improving the coding speed. Method 23 combines the complexity of the image block to be coded with the coding feature, and can realize a fast judgment of the division method and quickly jump out of the division process of the image block to be coded at the current size, so as to improve the coding speed as much as possible without affecting the coding bit rate and the user's subjective experience.

[0146] According to another aspect of the present disclosure, a device for processing a video frame image is also provided. The device for processing a video frame image includes: a first acquisition module, configured to acquire an image block to be encoded from the video frame image; a second acquisition module, configured to acquire complexity information of the image block to be encoded and complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided; a first determination module, configured to determine a division method of the image block to be encoded based on the complexity information of the image block to be encoded and the complexity information of the plurality of sub-image blocks, the division method including a coarse-grained division method and a fine-grained division method; an encoding module, configured to determine the image block to be encoded based on the division method Perform division and encoding to obtain the encoding data of the image block to be encoded; a second determination module is configured to determine whether the division method meets the preset conditions based on the encoding data of the image block to be encoded, the complexity information of the multiple sub-image blocks and the complexity information of the image block to be encoded; a stop division method search module is configured to: when the division method meets the preset conditions, stop searching for the division method of the image block to be encoded; a continue division method search module is configured to: when the division method does not meet the preset conditions, continue searching for the division method of the image block to be encoded.

[0147] According to another aspect of the present disclosure, a device for processing video frame images is also provided. The device for processing video frame images includes: a first acquisition module, configured to acquire an image block to be encoded from the video frame image; a second acquisition module, configured to acquire complexity information of the image block to be encoded and complexity information of multiple sub-image blocks divided from the image block to be encoded; a determination module, configured to: determine the division method of the image block to be encoded based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks; the first encoding module, configured to: when the division method is a coarse-grained division method, encode the image block to be encoded as a whole to obtain the encoding data of the image block to be encoded; the second encoding module, configured to: when the division method is a fine-grained division method, divide the image block to be encoded into multiple fine-grained sub-image blocks, and encode the multiple sub-image blocks respectively to obtain the encoding data of the image block to be encoded.

[0148] According to yet another aspect of the present disclosure, a device for processing a video frame image is provided. The device for processing video frame images includes: a first acquisition module, configured to acquire an image block to be encoded from the video frame image; a second acquisition module, configured to acquire complexity information of the image block to be encoded; a first determination module, configured to determine a division method of the image block to be encoded based on the complexity information of the image block to be encoded, wherein the division method includes a coarse-grained division method and a fine-grained division method; an encoding module, configured to divide and encode the image block to be encoded based on the division method to obtain encoding data of the image block to be encoded; a second determination module, configured to determine an encoding feature of the image block to be encoded based on the encoding data of the image block to be encoded; a third determination module, configured to determine whether the division method meets a preset condition based on the encoding feature of the image block to be encoded and the complexity information of the image block to be encoded; a stop division method search module, configured to stop searching for the division method of the image block to be encoded if the division method meets the preset condition; and a continue division method search module, configured to continue searching for the division method of the image block to be encoded if the division method does not meet the preset condition.

[0149] Figure 3A is another schematic diagram of the method 30 for processing a video frame image according to an embodiment of the present disclosure. Figure 3B It is a schematic diagram of determining whether a division method satisfies a preset condition when the division method is a coarse-grained division method according to an embodiment of the present disclosure. Figure 3C It is a schematic diagram of determining whether a division method satisfies a preset condition when the division method is a fine-grained division method according to an embodiment of the present disclosure. Figure 3DIt is a schematic diagram comparing the advantages and disadvantages of a coarse-grained partitioning method and a fine-grained partitioning method according to an embodiment of the present disclosure.

[0150] The first complexity analysis model, the second complexity analysis model, the encoding data analysis model and the partitioning method analysis model described below can all be artificial intelligence models, especially artificial intelligence-based neural network models. Typically, artificial intelligence-based neural network models are implemented as acyclic graphs in which neurons are arranged in different layers. Typically, a neural network model includes an input layer and an output layer, which are separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating outputs in the output layer. The network nodes are fully connected to the nodes in the adjacent layers via edges, and there are no edges between the nodes in each layer. The data received at the nodes of the input layer of the neural network is propagated to the nodes of the output layer via any one of the hidden layer, activation layer, pooling layer, convolution layer, etc. The input and output of the neural network model can take various forms, and the present disclosure is not limited to this.

[0151] The first complexity analysis model, the second complexity analysis model, the coding data analysis model and the division method analysis model described below may not be artificial intelligence models, but other types of computing models, and the present disclosure does not limit this.

[0152] like Figure 3A As shown, the method 30 for processing a video frame image according to an embodiment of the present disclosure includes the following steps.

[0153] First, in step S301, the complexity information of the image block to be encoded is obtained using a first complexity analysis model, wherein the input of the first complexity analysis model is the pixel value of each pixel in the image block to be encoded; and the output of the first complexity analysis model is the complexity of the image block to be encoded.

[0154] For example, step S301 may correspond to the above steps S212, S222 and S232 of obtaining the complexity information of the image block to be encoded, wherein the complexity information of the image block to be encoded is obtained using a first complexity analysis model.

[0155] In step S302, the complexity information of the multiple sub-image blocks into which the image block to be encoded is divided is obtained by using a second complexity analysis model. The input of the second complexity analysis model is the pixel value of each pixel in the multiple fine-grained sub-image blocks into which the image block to be encoded is divided; the output of the second complexity analysis model is one or more of the sum of the complexities, the mean complexity, and the weighted sum of the complexity of each fine-grained sub-image block into which the image block to be encoded is divided.

[0156] For example, step S302 may correspond to the above-mentioned step S212 and step S222 of obtaining complexity information of multiple sub-image blocks into which the image block to be encoded is divided, wherein a second complexity analysis model is used to obtain complexity information of multiple sub-image blocks into which the image block to be encoded is divided.

[0157] As described above, the complexity information may include one or more of the following: intra-frame coarse prediction loss of a video frame, inter-frame coarse prediction loss of a video frame, prediction rate-distortion loss of a video frame, the sum of predicted absolute errors between an image block to be encoded and an image block reconstructed based on encoding data, the sum of squared predicted differences between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean absolute difference between an image block to be encoded and an image block reconstructed based on encoding data, the predicted mean squared error between an image block to be encoded and an image block reconstructed based on encoding data, spatial information and / or temporal information.

[0158] The image block to be encoded is located in the video frame image, and the image block reconstructed based on the encoded data is located in the reconstructed video frame image, and the two include pixels at corresponding positions. The above-mentioned sum of absolute errors, sum of squared differences, mean absolute difference and mean squared error can be calculated / predicted based on the difference between the pixels at corresponding positions of the two image blocks.

[0159] Therefore, the first complexity analysis model and / or the second complexity analysis model can be further designed according to the contents included in the complexity information. When the complexity information includes the above-mentioned multiple contents, the first complexity analysis model and / or the second complexity analysis model can include corresponding prediction models based on the above-mentioned multiple contents, and design corresponding pooling layers, activation layers and output layers based on the outputs of each prediction model, so as to minimize the weighted sum of the multiple complexity information. The present disclosure is not limited to this.

[0160] The first complexity analysis model and / or the second complexity analysis model can be trained using historical videos, or the complexity analysis model can be trained in real time during the encoding process of a specific video. Of course, the first complexity analysis model and / or the second complexity analysis model can also be trained by combining the two.

[0161] In step S303, based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, a division method analysis model is used to determine the division method of the image block to be encoded, wherein the input of the division method analysis model is: the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks; the output of the division method analysis model is: the probability of mutual correlation between the multiple sub-image blocks.

[0162] In step S304, a division method is determined based on the output of the division method analysis model. Optionally, the division method of the image block to be encoded further includes: when the probability of mutual association between the multiple sub-image blocks is greater than a preset association threshold, determining the division method as a fine-grained division method; when the probability of mutual association between the multiple sub-image blocks is less than a preset association threshold, determining the division method as a coarse-grained division method.

[0163] For example, step S303 and step S304 may correspond to the determination of the division method of the image blocks to be encoded in the above step S213, step S223 and step S233. The division method analysis model is used to determine the division method of the image blocks to be encoded.

[0164] In step S305, when the division method is a coarse-grained division method, division and encoding are performed in the coarse-grained division method.

[0165] Step S305 may correspond to the above steps S214, S224 and S234: when the division mode is a coarse-grained division mode, the image block to be encoded is divided and encoded based on the division mode to obtain the encoding data of the image block to be encoded. For example, the image block to be encoded is encoded as a whole to obtain the encoding data of the image block to be encoded.

[0166] The coded data is a code stream generated by encoding the image block to be encoded.

[0167] In step S306, when the division method is a coarse-grained division method, it is determined whether the division method meets a preset condition based on the encoding data of the image block to be encoded and the complexity information of the multiple sub-image blocks.

[0168] Step S306 may correspond to the above steps S215 and S236, which determine whether the division method meets a preset condition based on the encoding data of the image block to be encoded, the complexity information of the plurality of sub-image blocks and the complexity information of the image block to be encoded.

[0169] Determining whether the division method meets the preset conditions in step S306 also includes: determining the encoding characteristics of the image block to be encoded based on the encoding data of the image block to be encoded; obtaining a preset quantization parameter, and determining whether the division method meets the preset conditions based on the encoding characteristics of the image block to be encoded, the complexity information of the multiple sub-image blocks and at least two of the quantization parameters.

[0170] The coding features include one or more of the following: coding mode, coding division depth, rate-distortion loss, the sum of absolute errors between the image block to be encoded and the image block reconstructed based on the coding data, the sum of squared differences between the image block to be encoded and the image block reconstructed based on the coding data, the average absolute difference between the image block to be encoded and the image block reconstructed based on the coding data, and the average square error between the image block to be encoded and the image block reconstructed based on the coding data.

[0171] For example, a coding data analysis model may be used to extract coding features from the coding data, wherein the input of the coding data analysis model is the coding data of the coded image block, and the output of the coding data analysis model is the coding features of the coded image block.

[0172] The coding data analysis model can be further designed according to the content included in the coding feature. When the coding feature includes the above-mentioned multiple contents, the coding feature can include a corresponding calculation model based on the above-mentioned multiple contents, and design the corresponding pooling layer, activation layer and output layer based on the output of each calculation model, so as to minimize the weighted sum of the multiple coding features. The present disclosure is not limited to this.

[0173] The coding data analysis model can be trained using historical videos, or it can be trained in real time during the encoding process of a specific video. Of course, the coding data analysis model can also be trained by combining the two.

[0174] For example, see Figure 3B , the step S306 of determining whether the division method meets the preset conditions also includes the following steps.

[0175] In step S3061, a first value is determined based on the encoding feature of the image block to be encoded and the quantization parameter.

[0176] Optionally, the coding feature of the image block to be coded in step S3061 is any one of the sum of absolute errors between the image block to be coded and the image block reconstructed based on the coding data, the sum of squared differences between the image block to be coded and the image block reconstructed based on the coding data, the average absolute difference between the image block to be coded and the image block reconstructed based on the coding data, and the average square error between the image block to be coded and the image block reconstructed based on the coding data. The first value is the ratio between the coding feature and the quantization parameter.

[0177] The quantization parameter may be a trained value or a value preset by the encoder, which is not limited in the present disclosure.

[0178] In step S3062, when the first value is less than the first threshold, it is determined that the division method meets the preset condition. Optionally, the first threshold in step S3062 may be a trained value or a value preset by the encoder, which is not limited in the present disclosure.

[0179] In step S3063, when the first value is greater than a first threshold, a second value is determined based on the encoding feature of the image block to be encoded, the quantization parameter, and complexity information of the multiple sub-image blocks.

[0180] When the first value is greater than the first threshold, it is necessary to further determine whether the division method meets the preset condition.

[0181] Optionally, the complexity information of the multiple sub-image blocks in step S3063 is the maximum difference or average difference between the sum of squares of prediction differences between the first sub-image block and the first sub-image block reconstructed based on the encoding data, the sum of squares of prediction differences between the second sub-image block and the second sub-image block reconstructed based on the encoding data, the sum of squares of prediction differences between the third sub-image block and the third sub-image block reconstructed based on the encoding data, and the sum of squares of prediction differences between the fourth sub-image block and the fourth sub-image block reconstructed based on the encoding data.

[0182] Optionally, the encoding feature of the image block to be encoded in step S3063 is a division depth, wherein the division depth may indicate the number of iterations performed when a certain image block is iterated from an image block of maximum size (an image block of coarsest granularity) to the size of an image block for encoding.

[0183] Optionally, the second value is a ratio of complexity information of the multiple sub-image blocks to the encoding feature of the image block to be encoded and / or the quantization parameter.

[0184] In step S3064, when the second value is less than the second threshold, it is determined that the division method meets the preset condition; when the second value is greater than the second threshold, it is determined that the division method does not meet the preset condition.

[0185] Optionally, the second threshold in step S3064 may be a trained value or a value preset by an encoder, which is not limited in the present disclosure.

[0186] In step S307, when the division method is a fine-grained division method, division and encoding are performed in the fine-grained division method.

[0187] Step S307 may correspond to the above steps S214, S225 and S234: when the division mode is a fine-grained division mode, the image block to be encoded is divided and encoded based on the division mode to obtain the encoded data of the image block to be encoded. For example, the image block to be encoded is divided into a plurality of fine-grained sub-image blocks, and the plurality of sub-image blocks are respectively encoded to obtain the encoded data of the image block to be encoded.

[0188] The coded data is a code stream generated by encoding the image block to be encoded.

[0189] In step S308, when the division method is a fine-grained division method, it is determined whether the division method meets a preset condition based on the encoding data of the image block to be encoded and the complexity information of the multiple sub-image blocks.

[0190] Step S308 may correspond to the above step S215 and step S235, which determines whether the division method meets the preset condition based on the encoding data of the image block to be encoded, the complexity information of the multiple sub-image blocks and the complexity information of the image block to be encoded.

[0191] Determining whether the division method meets the preset conditions in step S308 also includes: determining the rate-distortion loss of each sub-image block of the image block to be encoded based on the encoding data of the image block to be encoded; and determining whether the division method meets the preset conditions based on the complexity information of the multiple sub-image blocks and the rate-distortion loss of each sub-image block.

[0192] For example, a coding data analysis model may be used to extract coding features from the coding data, wherein the input of the coding data analysis model is the coding data of the coded image block, and the output of the coding data analysis model is the coding features of the coded image block.

[0193] The coding data analysis model can be further designed according to the content included in the coding feature. When the coding feature includes the above-mentioned multiple contents, the coding feature can include a corresponding calculation model based on the above-mentioned multiple contents, and design the corresponding pooling layer, activation layer and output layer based on the output of each calculation model, so as to minimize the weighted sum of the multiple coding features. The present disclosure is not limited to this.

[0194] The coding data analysis model can be trained using historical videos, or it can be trained in real time during the encoding process of a specific video. Of course, the coding data analysis model can also be trained by combining the two.

[0195] For example, refer to Figure 3C, the step S308 of determining whether the division method meets the preset conditions also includes the following steps.

[0196] In step S3081, a third value is determined based on the complexity information of the plurality of sub-image blocks and the rate-distortion loss of each sub-image block.

[0197] Optionally, the complexity information of the multiple sub-image blocks in step S3081 is the predicted rate-distortion loss of the video frame. The third value in step S3081 is the ratio between the maximum difference of the rate-distortion loss of each sub-image block and the complexity information of the multiple sub-image blocks. The present disclosure is not limited to this. For example, the third value in step S3081 may also be the ratio between the maximum difference of the rate-distortion loss of each sub-image block and the quantization parameter. For another example, the third value in step S3081 may also be the maximum difference of the rate-distortion loss of each sub-image block.

[0198] In step S3082, when the third value is greater than the third threshold, it is determined that the division method meets the preset condition. Optionally, the third threshold in step S3082 may be a trained value or a value preset by the encoder, which is not limited in the present disclosure.

[0199] In step S3083, when the third value is less than the third threshold, a fourth value is determined based on the complexity information of the multiple sub-image blocks and the prediction mode of each sub-image block.

[0200] Optionally, the fourth value in step S3083 is the difference between the prediction modes of the sub-image blocks, or is the ratio between the difference between the prediction modes of the sub-image blocks and the complexity information of the multiple sub-image blocks. The present disclosure is not limited thereto.

[0201] In step S3084, when the fourth value is greater than the fourth threshold, it is determined that the division method meets the preset condition; when the fourth value is less than the fourth threshold, it is determined that the division method does not meet the preset condition.

[0202] Optionally, the fourth threshold in step S3084 may be a trained value or a value preset by an encoder, which is not limited in the present disclosure.

[0203] In step S309, when it is determined that the division method does not meet the preset conditions, it is determined that the encoding process needs to be reset and the comparison of the advantages and disadvantages of the division methods is performed, and based on this, step S310 is entered. In step S310, the search for the division method of the image block to be encoded is continued. For example, by resetting the encoding process, the encoded data under the division method different from the division method is obtained. Then, the advantages and disadvantages of the coarse-grained division method and the fine-grained division method are compared to select a better division method.

[0204] In step S309, when it is determined that the division method meets the preset condition, the search for the division method of the image block to be encoded is stopped.

[0205] Steps S309 to S310 may correspond to the above-mentioned steps S216-S217 and steps S237-S238.

[0206] For example, when the division method does not satisfy the preset conditions, continuing to search for the division method of the image block to be encoded also includes: when the division method is a coarse-grained division method, dividing the image block to be encoded into a plurality of fine-grained sub-image blocks, and encoding the plurality of sub-image blocks respectively to obtain the encoding data of the image block to be encoded; when the division method is a fine-grained division method, encoding the image block to be encoded as a whole to obtain the encoding data of the image block to be encoded.

[0207] Then, based on the coded data obtained by encoding the image block to be encoded as a whole and the coded data obtained by encoding a plurality of sub-image blocks respectively, the coded data of the image block to be encoded is determined.

[0208] For example, refer to Figure 3D , step S310 may further include the following steps.

[0209] In step S3101, a first rate-distortion loss of the image block to be encoded is determined based on the encoded data obtained by encoding the image block to be encoded as a whole; and a second rate-distortion loss of the image block to be encoded is determined based on the encoded data obtained by encoding multiple sub-image blocks separately.

[0210] In step S3102, when the first rate-distortion loss is lower than the second rate-distortion loss, the coded data obtained by encoding the image block to be encoded as a whole is used as the coded data of the image block to be encoded.

[0211] In step S3103, when the first rate-distortion loss is higher than the second rate-distortion loss, the coded data obtained by respectively encoding the plurality of sub-image blocks is used as the coded data of the image block to be encoded.

[0212] Optionally, if it is finally determined that the coded data obtained by respectively encoding the plurality of sub-image blocks should be used as the coded data of the image block to be encoded, the plurality of divided sub-image blocks can be respectively used as the image blocks to be encoded to determine the division method of each sub-image block. The determination of the division method of each sub-image block is similar to the determination of the division method of the image block to be encoded, and will not be described in detail here.

[0213] In step S311, a more optimal division method is analyzed.

[0214] For example, a coding data analysis model can be used to extract coding features from the coded data obtained by encoding the image block to be encoded as a whole, and coding features from the coded data obtained by encoding multiple sub-image blocks separately, and analyze these coding features. For example, the difference in coding features brought about by these two division methods can be further analyzed, so as to provide a reference for the prejudgment of the division methods of other video frames.

[0215] Then, in step S312, the parameters in the above model are adjusted.

[0216] For example, in step S312, based on the analysis result in step S311, the parameters in the first complexity analysis model, the second complexity analysis model, the coding data analysis model, and the partitioning method analysis model are adjusted.

[0217] When the complexity information of the image block to be encoded is obtained by using the first complexity analysis model, the first rate-distortion loss and the second rate-distortion loss can be used to adjust the parameters in the first complexity analysis model. Of course, other encoding features can also be used to adjust the parameters in the first complexity analysis model, and the present disclosure is not limited to this.

[0218] For example, when the complexity information of the plurality of sub-image blocks into which the image block to be encoded is divided is obtained by using the second complexity analysis model, the first rate-distortion loss and the second rate-distortion loss can be used to adjust the parameters in the second complexity analysis model. Of course, other encoding features can also be used to adjust the parameters in the second complexity analysis model, and the present disclosure is not limited thereto.

[0219] For example, when the coding features are obtained using the coding data analysis model, the first rate distortion loss and the second rate distortion loss are used to adjust the parameters in the coding data analysis model. Of course, other coding features can also be used to adjust the parameters in the coding data analysis model, and the present disclosure is not limited thereto.

[0220] For example, when the partitioning method analysis model is used to determine the partitioning method of the image block to be encoded, the first rate-distortion loss and the second rate-distortion loss are used to adjust the parameters in the partitioning method analysis model. Of course, other encoding features can also be used to adjust the parameters in the partitioning method analysis model, and the present disclosure is not limited thereto.

[0221] Optionally, if it is finally determined that the coded data obtained by respectively encoding the plurality of sub-image blocks should be used as the coded data of the image block to be encoded, the plurality of divided sub-image blocks can be respectively used as the image blocks to be encoded to determine the division method of each sub-image block. The determination of the division method of each sub-image block is similar to the determination of the division method of the image block to be encoded, and will not be described in detail here.

[0222] Therefore, method 30 can obtain a more accurate division method by combining the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks divided by the image block to be encoded. At the same time, method 30 can determine whether to jump out of the search for the division method of the current size by comparing the encoding features of the image block to be encoded with the complexity information of the image block to be encoded (and / or the complexity information of the multiple sub-image blocks divided by the image block to be encoded), thereby improving the encoding speed. Method 30 combines the complexity information of the image block to be encoded (and / or the complexity information of the multiple sub-image blocks divided by the image block to be encoded) with the encoding features, which can realize the rapid judgment of the division method and quickly jump out of the division process of the image block to be encoded at the current size, so as to improve the encoding speed as much as possible without affecting the encoding bit rate and the user's subjective experience.

[0223] Figure 4A 40 is a schematic diagram of an interface of a method 40 for processing a video frame image according to an embodiment of the present disclosure. Figure 4B is a schematic flowchart of a method 40 for processing a video frame image according to an embodiment of the present disclosure.

[0224] refer to Figure 4B According to an embodiment of the present disclosure, the method 40 for processing a video frame image may include the following steps.

[0225] In step S401, accelerated coding enabling information input by a user is obtained, wherein the accelerated coding enabling information indicates: in the process of determining a division method of an image block to be encoded in a video frame image, the division method of the image block to be encoded is determined by utilizing complexity information of the image block to be encoded and complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided, and the division method includes a coarse-grained division method and a fine-grained division method.

[0226] For example, in Figure 4AIn the example, if the user clicks whether to accelerate encoding, the display of the acceleration option will be triggered. At this time, if the user selects the pre-analyzed acceleration mode, then, next, in step S402, the video frame image is encoded based on the accelerated encoding enabling information. For example, the encoder will use the above-mentioned methods 21 to 23, and method 30 to accelerate the encoding process.

[0227] In addition, the user may input parameters corresponding to relevant acceleration options in the command line panel to enable the above-mentioned methods 21 to 23 and method 30 to accelerate the encoding process.

[0228] According to yet another aspect of the present disclosure, an electronic device is provided. Figure 5 A schematic diagram of an electronic device 2000 according to an embodiment of the present disclosure is shown.

[0229] like Figure 5 As shown, the electronic device 2000 may include one or more processors 2010 and one or more memories 2020. The memory 2020 stores a computer-readable code, and when the computer-readable code is run by the one or more processors 2010, the method described above may be executed.

[0230] The processor in the embodiments of the present disclosure may be an integrated circuit chip having the ability to process signals. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc., and may be an X86 architecture or an ARM architecture.

[0231] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.

[0232] For example, the method or device according to the embodiment of the present disclosure may also be Figure 6 The architecture of the computing device 3000 shown in FIG. Figure 6As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method for determining the driving risk of a vehicle provided by the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 6 One or more components of a computing device are shown.

[0233] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided. Figure 7 A schematic diagram 4000 of a storage medium according to the present disclosure is shown.

[0234] like Figure 7 As shown, the computer storage medium 4020 stores computer readable instructions 4010. When the computer readable instructions 4010 are executed by the processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous connection dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM). It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0235] The embodiments of the present disclosure also provide a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method according to the embodiments of the present disclosure.

[0236] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0237] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.

[0238] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for processing a video frame image, comprising: Obtaining an image block to be encoded from the video frame image; Acquire complexity information of the image block to be encoded and complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided; Based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, determine a division method of the image block to be encoded, wherein the division method includes a coarse-grained division method and a fine-grained division method; When the division method is a coarse-grained division method, Encoding the image block to be encoded as a whole to obtain encoded data of the image block to be encoded; Determining a coding feature of the image block to be encoded based on the coding data of the image block to be encoded; Obtaining a preset quantization parameter, and determining whether the division method meets a preset condition based on the encoding feature of the image block to be encoded, the complexity information of the multiple sub-image blocks, and at least two of the quantization parameters; In the case where the division method is a fine-grained division method, Dividing the image block to be encoded into a plurality of fine-grained sub-image blocks, and encoding the plurality of sub-image blocks respectively to obtain encoding data of the image block to be encoded; Determining rate-distortion loss of each sub-image block of the image block to be encoded based on the encoded data of the image block to be encoded; Determining whether a division method satisfies a preset condition based on complexity information of the plurality of sub-image blocks and rate-distortion loss of each sub-image block; When the division method satisfies a preset condition, stopping the search for the division method of the image block to be encoded; When the division method does not satisfy the preset condition, the search for the division method of the image block to be encoded continues.

2. The method of claim 1, wherein: When the division mode does not satisfy the preset condition, continuing to search for the division mode of the image block to be encoded further includes: In the case where the division mode is a coarse-grained division mode, the image block to be encoded is divided into a plurality of fine-grained sub-image blocks, and the plurality of sub-image blocks are respectively encoded to obtain encoding data of the image block to be encoded; In the case where the division mode is a fine-grained division mode, encoding the image block to be encoded as a whole to obtain encoding data of the image block to be encoded; The encoding data of the image block to be encoded is determined based on the encoding data obtained by encoding the image block to be encoded as a whole and the encoding data obtained by encoding a plurality of sub-image blocks respectively.

3. The method of claim 1, wherein: The determining whether the division method meets the preset condition also includes: Determining a first value based on the encoding feature of the image block to be encoded and the quantization parameter; When the first value is less than a first threshold, determining that the division method meets a preset condition; When the first value is greater than a first threshold, determining a second value based on the encoding feature of the image block to be encoded, the quantization parameter, and complexity information of the multiple sub-image blocks; When the second value is less than a second threshold, determining that the division method meets a preset condition; When the second value is greater than the second threshold, it is determined that the division method does not meet the preset condition.

4. The method of claim 1, wherein: The determining whether the division method meets the preset condition includes: determining a third value based on the complexity information of the plurality of sub-image blocks and the rate-distortion loss of each of the sub-image blocks, and determining that the division method meets a preset condition when the third value is greater than a third threshold; In a case where the third value is less than a third threshold, determining a prediction mode of each sub-image block of the image block to be encoded based on the encoding data of the image block to be encoded; Determining a fourth value based on complexity information of the plurality of sub-image blocks and a prediction mode of each of the sub-image blocks; When the fourth value is greater than a fourth threshold, determining that the division method satisfies a preset condition; When the fourth value is less than the fourth threshold, it is determined that the division method does not meet the preset condition.

5. The method of claim 1, wherein: Determining the division method of the image block to be encoded based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks also includes: determining the division method of the image block to be encoded by using a division method analysis model based on the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks, wherein: The input of the division method analysis model is: the complexity information of the image block to be encoded and the complexity information of the multiple sub-image blocks; The output of the segmentation method analysis model is: the probability of the multiple sub-image blocks being associated with each other, The method of determining the division of the image blocks to be encoded further includes: When the probability of the mutual association between the plurality of sub-image blocks is greater than a preset association threshold, determining that the division method is a fine-grained division method; When the probability of the mutual association between the multiple sub-image blocks is less than a preset association threshold, it is determined that the division method is a coarse-grained division method.

6. The method according to any one of claims 1 to 5, wherein: The complexity information includes one or more of the following items: intra-frame coarse prediction loss of a video frame, inter-frame coarse prediction loss of a video frame, prediction rate-distortion loss of a video frame, a sum of predicted absolute errors between an image block to be encoded and an image block reconstructed based on encoding data, a sum of squared predicted differences between an image block to be encoded and an image block reconstructed based on encoding data, a predicted mean absolute difference between an image block to be encoded and an image block reconstructed based on encoding data, a predicted mean squared error between an image block to be encoded and an image block reconstructed based on encoding data, spatial information and / or temporal information; The coding feature includes one or more of the following: coding mode, coding division depth, rate distortion loss, the sum of absolute errors between the image block to be coded and the image block reconstructed based on the coding data, the sum of squared differences between the image block to be coded and the image block reconstructed based on the coding data, the average absolute difference between the image block to be coded and the image block reconstructed based on the coding data, and the average square error between the image block to be coded and the image block reconstructed based on the coding data; The coded data is a code stream generated by encoding the image block to be encoded.

7. The method of claim 2, wherein: The determining of the encoding data of the image block to be encoded based on the encoding data obtained by encoding the image block to be encoded as a whole and the encoding data obtained by encoding the plurality of sub-image blocks separately comprises: Determining a first rate-distortion loss of the image block to be encoded according to encoded data obtained by encoding the image block to be encoded as a whole; Determining a second rate-distortion loss of the image block to be encoded according to encoded data obtained by respectively encoding the plurality of sub-image blocks; When the first rate-distortion loss is lower than the second rate-distortion loss, using the coded data obtained by encoding the image block to be encoded as a whole as the coded data of the image block to be encoded; When the first rate-distortion loss is higher than the second rate-distortion loss, the coded data obtained by respectively encoding the plurality of sub-image blocks is used as the coded data of the image block to be encoded.

8. The method of claim 1, wherein: The obtaining the complexity information of the image block to be encoded further comprises: obtaining the complexity information of the image block to be encoded using a first complexity analysis model, wherein: The input of the first complexity analysis model is the pixel value of each pixel in the image block to be encoded; The output of the first complexity analysis model is the complexity of the image block to be encoded; The obtaining of the complexity information of the image block to be encoded further comprises: using a second complexity analysis model to obtain the complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided, wherein: The input of the second complexity analysis model is the pixel value of each pixel in a plurality of fine-grained sub-image blocks into which the image block to be encoded is divided; The output of the second complexity analysis model is one or more of the sum of the complexities of each fine-grained sub-image block into which the image block to be encoded is divided, the complexity mean, and the complexity weighted sum.

9. A method for processing a video frame image, comprising: Acquiring accelerated coding enabling information input by a user, wherein the accelerated coding enabling information indicates that: in the process of determining the division mode of the image block to be encoded in the video frame image, the division mode of the image block to be encoded is determined by using complexity information of the image block to be encoded and complexity information of a plurality of sub-image blocks into which the image block to be encoded is divided, and the division mode includes a coarse-grained division mode and a fine-grained division mode; Based on the accelerated encoding enabling information, the video frame image is encoded using the method described in any one of claims 1-8.

10. An electronic device, comprising: one or more processors; and One or more memories, wherein the memories store computer readable codes, and when the computer readable codes are executed by the one or more processors, the method according to any one of claims 1 to 9 is executed.

Citation Information

Patent Citations

  • Video processing method, video processing device, intelligent equipment and storage medium

    CN112104867A