Video processing method, video processing apparatus, smart device, and storage medium
Patent Information
- Application Number
- MYPI2023002739
- Authority / Receiving Office
- MY · MY
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-09
- Filing Date
- 2021-11-03
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2041-11-03
AI Technical Summary
In the video encoding process, it is a challenge to accurately determine the appropriate encoding mode to improve encoding efficiency and speed. It is difficult to quickly and accurately determine the optimal division of data blocks with existing technology, especially when the scene complexity is high or relatively large. low data block.
By performing scene complexity analysis on the target data block in the video frame to be encoded, it is divided into multiple sub-data blocks, and the scene complexity is analyzed separately. The joint statistical model is used to determine the optimal encoding mode based on the data block and sub-block indicator information. thereby choosing an appropriate encoding strategy.
It realizes the adaptive selection of encoding mode according to the scene complexity of the video frame, improves the video encoding speed and efficiency, is suitable for video encoding in different scenes, and ensures the best balance between encoding quality and bit rate.
Abstract
Description
Video processing method, video processing device, intelligent device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 9, 2020, with application number 202011239333.1 and titled “A video processing method, video processing device, smart device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a video processing method, a video processing apparatus, an intelligent device, and a computer-readable storage medium. Background Art
[0003] Currently, video coding technology is widely used in scenarios such as video conversations and video on demand. For example, video coding is used to process conversational videos in video conversations and to process on-demand videos in video on demand. Video coding technology compresses and encodes videos according to encoding modes. This compression method effectively saves storage space and improves transmission efficiency.
[0004] Accurately determining a more appropriate encoding mode during the video encoding process and encoding the video according to the encoding mode can accelerate the entire video encoding process and improve video encoding efficiency. Therefore, how to more accurately determine the encoding mode for video frame encoding during the video encoding process has become a hot topic of current research.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide a video processing method, apparatus, device, and storage medium, which can more accurately select a suitable encoding mode for video frame encoding.
[0007] In one aspect, an embodiment of the present application provides a video processing method, which is executed by a smart device. The video processing method includes:
[0008] Obtaining a target video frame from a video to be encoded, and determining a target data block to be encoded from the target video frame;
[0009] Perform scene complexity analysis on the target data block to obtain data block indicator information;
[0010] Divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2;
[0011] Determining a coding mode for a target data block according to the data block indicator information and the sub-block indicator information;
[0012] The target data block is encoded according to the determined encoding mode.
[0013] On the other hand, an embodiment of the present application provides a video processing device, the video processing device comprising:
[0014] An acquisition unit, configured to acquire a target video frame from the video to be encoded, and determine a target data block to be encoded from the target video frame;
[0015] A processing unit is used to perform scene complexity analysis on a target data block to obtain data block index information; divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2; determine an encoding mode for the target data block based on the data block index information and the sub-block index information; and encode the target data block according to the determined encoding mode.
[0016] On the other hand, an embodiment of the present application provides a smart device, comprising:
[0017] a processor adapted to implement a computer program; and
[0018] A memory stores a computer program, which is loaded and executed by a processor to implement the above-mentioned video processing method.
[0019] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is read and executed by a processor of a computer device, the above-mentioned video processing method is implemented.
[0020] In another aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the above-described video processing method.
[0021] BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] FIG1a is a schematic diagram of a recursive segmentation process of a video frame provided in an embodiment of the present application;
[0024] FIG1b is a schematic diagram of a video frame division result provided in an embodiment of the present application;
[0025] FIG2 is a schematic structural diagram of an encoder provided in an embodiment of the present application;
[0026] FIG3 is a schematic diagram of an associated data block provided in an embodiment of the present application;
[0027] FIG4 is a schematic diagram of the architecture of a video processing system provided in an embodiment of the present application;
[0028] FIG5 is a flow chart of a video processing method provided in an embodiment of the present application;
[0029] FIG6 is a flow chart of another video processing method provided in an embodiment of the present application;
[0030] FIG7 is a flow chart of another video processing method provided in an embodiment of the present application;
[0031] FIG8 is a schematic structural diagram of a video processing device provided in an embodiment of the present application;
[0032] FIG9 is a schematic structural diagram of a smart device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] The embodiments of the present application involve cloud technology (Cloud Technology), and the embodiments of the present application can use cloud technology to implement the video processing process. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to realize the calculation, storage, processing, and sharing of data. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the application of cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all types of industry data require strong system backing support, which can only be achieved through cloud computing.
[0035] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, these resources appear infinitely scalable and can be accessed, used, expanded, and paid for at any time. Cloud computing infrastructure providers establish a cloud computing resource pool (referred to as a cloud platform), commonly referred to as an IaaS (Infrastructure as a Service) platform. Various types of virtual resources are deployed within this pool for external clients to choose from. The cloud computing resource pool primarily includes computing devices (virtualized machines, including operating systems), storage devices, and network devices. Based on logical functionalities, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of the IaaS layer. PaaS is the platform on which software (such as databases and web containers) runs. SaaS refers to a wide variety of business software (such as web portals, SMS mass senders, etc.). Generally speaking, SaaS and PaaS are upper layers relative to IaaS. Cloud computing also refers to the delivery and usage model of IT (Internet Technology) infrastructure, which means obtaining required resources through the network in an on-demand, easily scalable manner. In a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining required services through the network in an on-demand, easily scalable manner. Such services can be IT and software, Internet-related, or other services. Cloud computing is the product of the integration of the development of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.
[0036] Cloud computing can be used in the field of cloud conferencing. Cloud conferencing is an efficient, convenient, and low-cost conferencing format based on cloud computing technology. Users simply use a simple, easy-to-use internet interface to quickly and efficiently share voice, data files, and videos with teams and clients around the world. The cloud conferencing service provider handles the complex technical aspects of data transmission and processing. Currently, domestic cloud conferencing services primarily utilize a SaaS model, encompassing phone, internet, and video services. Cloud-based video conferencing is known as cloud conferencing. In the cloud conferencing era, data transmission, processing, and storage are all handled by the video conferencing vendor's computer resources. Users no longer need to purchase expensive hardware or install complex software. Simply open a browser and log in to the appropriate interface to conduct efficient remote meetings. Cloud conferencing systems support dynamic multi-server cluster deployment and provide multiple high-performance servers, significantly improving conferencing stability, security, and availability. In recent years, video conferencing has gained widespread popularity due to its ability to significantly improve communication efficiency, continuously reduce communication costs, and enhance internal management. It has been widely adopted in various sectors, including government, military, transportation, finance, carriers, education, and enterprises. There is no doubt that after the application of cloud computing, video conferencing will have greater appeal in terms of convenience, speed and ease of use, and will surely trigger a new wave of video conferencing applications.
[0037] The embodiments of the present application relate to video, which is a sequence of video frames consisting of continuous video frames. Due to the persistence of vision effect of the human eye, when a video frame sequence is played at a certain rate, what we see is a video of continuous action. Since the similarity between continuous video frames is extremely high, there is a large amount of redundant information within each video frame and between continuous video frames. Therefore, before storing or transmitting the video, it is often necessary to encode the video using video encoding technology to remove redundant information in the video in dimensions such as space and time, so as to save storage space and improve video transmission efficiency.
[0038] Video encoding technology, also known as video compression technology, refers to the technology of compressing and encoding each video frame in a video according to a coding mode. Specifically, during the video encoding process, each video frame in the video needs to be recursively divided into data blocks of various sizes, and then each divided data block is input into an encoder for encoding. The recursive division process involved in the embodiment of the present application is described below with reference to Figure 1a. Figure 1a is a schematic diagram of a recursive division process of a video frame provided in an embodiment of the present application. As shown in FIG1a , a video frame 10 is composed of a plurality of data blocks 101 of a first size. The recursive division method of the video frame 10 may include: ① directly inputting the data blocks 101 of the first size into the encoder for encoding, that is, directly inputting the data blocks of the first size into the encoder for encoding without dividing the data blocks 101 of the first size; ② dividing the data blocks 101 of the first size into four data blocks 102 of the second size of the same size, without further dividing the data blocks 102 of the second size, and inputting the data blocks 102 of the second size into the encoder for encoding; ③ further dividing the data blocks 102 of the second size to obtain four data blocks 103 of the third size of the same size, without further dividing the data blocks 103 of the third size, and inputting the data blocks 103 of the third size into the encoder for encoding; and so on. The data blocks 103 of the third size may also be further divided, thereby recursively dividing the video frame 10 into data blocks of multiple sizes. The division result of the video frame 10 can be seen in FIG1b . FIG1b is a schematic diagram of the division result of a video frame provided in an embodiment of the present application. As shown in FIG. 1 b , the divided video frame 10 is composed of data blocks of three sizes, namely, a data block 101 of a first size, a data block 102 of a second size, and a data block 103 of a third size.
[0039] In the actual video encoding process, the division strategies of the multiple data blocks contained in the video frame are different, the encoding modes for encoding the video frames are also different, and the encoding speeds of the video frames are also different. An embodiment of the present application provides a video processing solution, which provides two different data block division strategies, namely a top-down data block division strategy and a bottom-up data block division strategy. Both of these data block division strategies can determine the optimal division method for each data block in the video frame. The optimal division method can be understood as that after the data blocks are divided according to the data block division method, the encoding quality when encoding the divided data blocks is better and the encoding bit rate is lower. The encoding bit rate can refer to the amount of data transmitted per unit time (for example, 1 second) when the encoded video data stream is transmitted. The more data transmitted per unit time, the higher the encoding bit rate. The above two data block division strategies are introduced in detail below.
[0040] (1) Top-down data block partitioning strategy
[0041] In the process of encoding the video, each data block in the video frame is divided from top to bottom and recursively. Specifically, the top-down and recursive division may refer to the process of dividing the data block from the maximum size of the video frame to the minimum size of the data block in the video frame until the optimal division method of the data block is found. For example, the maximum size of the data block in the video frame is 64px (Pixel) × 64px, and the minimum size of the data block in the video frame is 4px × 4px. The maximum size of the data block is 64px × 64px and the minimum size of the data block is 4px × 4px. This is only for example and does not constitute a limitation on the embodiments of the present application. For example, the maximum size of the data block can also be 128px × 128px, and the minimum size of the data block can also be 8px × 8px, and so on. A data block of 64px × 64px can be divided into four data blocks of 32px × 32px. The rate-distortion cost (RDCost) for encoding each of the four 32px×32px data blocks is calculated separately. The sum of the RDCosts for the four 32px×32px data blocks is calculated (i.e., the first sum RDCost of the four 32px×32px data blocks is calculated). The RDCost for directly encoding the 64px×64px data block is also calculated. If the first sum RDCost is greater than the RDCost of the 64px×64px data block, the data block is partitioned into 64px×64px, meaning there is no need to partition the 64px×64px data block. If the first sum RDCost is less than or equal to the RDCost of the 64px×64px data block, the data block is partitioned into 32px×32px. Subsequently, the 32px×32px data block can be further divided into four 16px×16px data blocks for encoding, and so on, until the optimal partitioning method for the data blocks is found. The top-down data block partitioning strategy can usually determine the optimal partitioning method more accurately. For data blocks with low scene complexity, the optimal partitioning size of the data block is relatively large. Therefore, for data blocks with low scene complexity, the top-down data block partitioning strategy can quickly determine the optimal partitioning method of the data block. For data blocks with high scene complexity, the optimal partitioning size of the data block is relatively small. Therefore, for data blocks with high scene complexity, the process of determining the optimal partitioning method of the data block through the top-down data block partitioning strategy requires a lot of time and cost, which in turn affects the encoding speed of the data block.
[0042] The optimal partition size for a data block can be understood as the size of the data block that results in better encoding quality and a lower bitrate when encoding the resulting data block. Generally speaking, data blocks with lower scene complexity have a larger optimal partition size, while data blocks with higher scene complexity have a smaller optimal partition size. The scene complexity of a data block can be measured using methods such as spatial information (SI) or temporal information. Spatial information can be used to characterize the amount of spatial detail in a data block. The more elements a data block contains, the higher the value of its spatial information and the higher its scene complexity. For example, if data block A contains five elements (a cat, a dog, a tree, a flower, and a sun), and data block B contains two elements (a cat and a dog), and data block A contains more elements than data block B, then the value of data block A's spatial information is higher than that of data block B, and the scene complexity of data block A is higher than that of data block B. Temporal information can be used to characterize the temporal variation of a data block. The higher the degree of motion of a target data block in the target video frame currently being processed relative to a reference data block in a reference video frame of the target video frame, the larger the value of the target data block's temporal information, and the higher the scene complexity of the target data block. The reference video frame is a video frame that precedes the target video frame in the coding order in the video frame sequence, and the position of the target data block in the target video frame is the same as the position of the reference data block in the reference video frame. For example, the target data block contains an element (such as a car element), and the reference data block also contains the same element. The greater the displacement of the car element in the target data block relative to the car element in the reference data block, the higher the degree of motion of the target data block in the target video frame relative to the reference data block in the reference video frame, the larger the value of the target data block's temporal information, and the higher the scene complexity of the target data block.
[0043] (2) Bottom-up video frame division strategy
[0044] During video encoding, each data block in a video frame is recursively partitioned from the bottom up. Specifically, the bottom-up recursive partitioning may refer to a process of partitioning the data blocks from the minimum size of the video frame to the maximum size of the data blocks in the video frame until the optimal partitioning method is found. For example, the minimum size of a data block in a video frame is 4px×4px, and the maximum size of a data block in a video frame is 64px×64px. Each of the four 4px×4px data blocks is encoded separately, and then an 8px×8px data block composed of the four 4px×4px data blocks is encoded. The RDCost for encoding each of the four 4px×4px data blocks is calculated separately, and the sum of the RDCosts of the four 4px×4px data blocks is calculated (i.e., the second sum RDCost of the four 4px×4px data blocks is calculated); and the RDCost for encoding the 8px×8px data block is calculated. If the second sum RDCost is less than or equal to the RDCost of the 8px×8px data block, the data block is partitioned into 4px×4px. If the second sum RDCost is greater than the RDCost of the 8px×8px data block, the data block is partitioned into 8px×8px. Subsequently, encoding can be continued for a 16px×16px data block composed of four 8px×8px data blocks, and so on, until the optimal partitioning method for the data block is found. The bottom-up video frame partitioning strategy is generally able to accurately determine the optimal partitioning method. For data blocks with high scene complexity, the optimal partitioning size is relatively small. Therefore, the bottom-up video frame partitioning strategy can quickly determine the optimal partitioning method for these data blocks. For data blocks with low scene complexity, the optimal partitioning size is relatively large. Therefore, determining the optimal partitioning method for these data blocks using the bottom-up video frame partitioning strategy is time-consuming, which in turn affects the encoding speed of the data block.
[0045] It can be seen that the top-down data block division strategy and the bottom-up video frame division strategy can usually more accurately determine the optimal division method for each data block in the video frame. The top-down data block division strategy is more suitable for data blocks with lower scene complexity, and the bottom-up video frame division strategy is more suitable for data blocks with higher scene complexity.
[0046] On this basis, the embodiment of the present application provides a further video processing scheme, which performs scene complexity analysis on the target data block to be encoded in the encoded video frame; divides the target data block into multiple sub-data blocks, and performs scene complexity analysis on each sub-data block respectively; determines the optimal division method of the target data block through the scene complexity analysis result of the target data block, and the scene complexity analysis result of each sub-data block in the multiple sub-data blocks obtained by division, and then determines the encoding mode of the target data block, so that the target data block can be encoded according to the determined encoding mode. The scene complexity analysis in the embodiment of the present application can be performed using any scene complexity analysis method in the relevant technology, such as the sum of squared error (SSE) method, the sum of absolute difference (SAD) method, etc. The optimal division method mentioned in the embodiment of the present application refers to a data block division method that can improve the encoding speed to a certain extent under the premise of achieving a balance between encoding quality and encoding bit rate. This video processing solution can formulate an encoding mode that is adapted to the scene complexity of the target data block according to the scene complexity of the target data block, effectively improving the encoding speed of the target data block, thereby improving the video encoding speed; and this video processing solution is more universal and applicable to any video encoding scenario. It can determine an encoding mode suitable for data blocks with lower scene complexity, and can also determine an encoding mode suitable for data blocks with higher scene complexity.
[0047] Based on the above description, an embodiment of the present application provides an encoder that can be used to implement the above-mentioned video processing solution. The video encoder can be an AV1 (Alliance for Open Media Video 1) standard encoder. Figure 2 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application. As shown in Figure 2, the encoder 20 includes a scene complexity analysis module 201, an adaptive partitioning decision module 202, and an early partitioning termination module 203.
[0048] The scene complexity analysis module 201 can be used to perform scene complexity analysis on the target data block to be encoded in the video frame to be encoded to obtain data block index information. The scene complexity analysis module 201 can also be used to divide the target data block into N sub-data blocks, and perform scene complexity analysis on each of the N sub-data blocks to obtain sub-block index information, where N is an integer greater than or equal to 2. In actual encoding scenarios, the value of N is generally 4, which is not limited in the embodiments of the present application. Dividing the target data block into N sub-data blocks can refer to dividing the target data block into 4 sub-data blocks of the same size.
[0049] The adaptive partitioning decision module 202 can be configured to determine an encoding mode for the target data block based on the data block indicator information and the sub-block indicator information. The encoding mode can include either a first mode or a second mode. That is, the adaptive partitioning decision module 202 can determine the encoding mode for the target data block as the first mode or the second mode based on the data block indicator information and the sub-block indicator information. The first mode is an encoding mode that divides the data block into N sub-data blocks and encodes each of the N sub-data blocks separately; the second mode is an encoding mode that directly encodes the target data block without dividing it. If the adaptive partitioning decision module 202 cannot determine an encoding mode for the target data block based on the data block indicator information and the sub-block indicator information, the adaptive partitioning decision module 202 can also be configured to obtain associated block indicator information for M associated data blocks related to the target data block and determine an attempted encoding order for the target data block based on the associated block indicator information, where M is a positive integer. In actual encoding scenarios, the value of M is generally 3, but this is not limited in this embodiment of the present application. Figure 3 is a schematic diagram of an associated data block provided in an embodiment of the present application. As shown in FIG3 , the M associated data blocks associated with the target data block may refer to associated data block 302 located at the left of target data block 301, associated data block 303 located at the top of target data block 301, and associated data block 304 located at the upper left of target data block 301. The attempted coding order may include either a first attempted coding order or a second attempted coding order, that is, the adaptive partitioning decision module 202 may determine the attempted coding order for the target data block as the first attempted coding order or the second attempted coding order based on the associated block indicator information. The first attempted coding order refers to an encoding order that first attempts to encode the target data block according to the first mode, and then attempts to encode the target data block according to the second mode; the second attempted coding order refers to an encoding order that first attempts to encode the target data block according to the second mode, and then attempts to encode the target data block according to the first mode.
[0050] The early partition termination module 203 is configured to set a partition termination condition when the adaptive partition decision module 202 is unable to determine the coding mode for the target data block, and determine the coding mode for the target data block based on the set partition termination condition. Specifically, if the adaptive partition decision module 202 determines that the attempted coding sequence for the target data block is the first attempted coding sequence, the early partition termination module 203 obtains coding information for the N sub-data blocks obtained by encoding the target data block according to the first mode. If the coding information of the N sub-data blocks satisfies the first partition termination condition (i.e., the fifth condition below), the early partition termination module 203 determines that the coding mode for the target data block is the first mode. If the adaptive partition decision module 202 determines that the attempted coding sequence for the target data block is the second attempted coding sequence, the early partition termination module 203 obtains coding information for the target data block obtained by encoding the target data block according to the second mode. If the coding information of the target data block satisfies the second partition termination condition (i.e., the sixth condition below), the early partition termination module 203 determines that the coding mode for the target data block is the second mode.
[0051] The encoder 20 shown in Figure 2 can formulate a coding mode that is compatible with the scene complexity of the target data block according to the scene complexity of the target data block, effectively improving the coding speed of the target data block, thereby improving the video coding speed. In addition, the early division termination module 203 in the encoder 20 can determine the coding mode of the target data block when the coding information meets the division termination condition, and terminate the further division of the target data block, thereby further improving the coding speed of the target data block under the premise of achieving a balance between the coding speed and the coding bit rate of the target data block, and further improving the video coding efficiency. The video processing scheme provided in the embodiment of the present application and the specific application scenario of the encoder 20 provided in the embodiment shown in Figure 2 will be introduced below in conjunction with the video processing system shown in Figure 4.
[0052] Figure 4 is a schematic diagram of the architecture of a video processing system provided in an embodiment of the present application. As shown in Figure 4, the video processing system includes P terminals (e.g., a first terminal 401, a second terminal 402, etc.) and a server 403, where P is an integer greater than 1. Any of the P terminals can be a device with a camera function, such as a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart wearable device, etc., but is not limited thereto. Any of the P terminals can support the installation and operation of various applications, including but not limited to social applications (e.g., instant messaging applications, audio conversation applications, video conversation applications, etc.), audio and video applications (e.g., audio and video on demand applications, audio and video players, etc.), game applications, etc. The server can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services, which is not limited in this embodiment of the present application. The P terminals can be directly or indirectly connected to the server 403 via wired or wireless communication, which is not limited in this embodiment of the present application. The video processing solution provided in the embodiment of the present application is introduced below using a video conversation scenario as an example.
[0053] (1) The encoders 20 are deployed in P terminals respectively, and the P terminals perform video processing.
[0054] P users participate in a video conversation using video conversation applications running on P terminals, respectively. For example, user 1 participates in a video conversation using a video conversation application running on first terminal 401, user 2 participates in a video conversation using a video conversation application running on second terminal 402, and so on. Server 403 is configured to transmit target videos generated by the P terminals during the video conversation. The target videos may include conversation videos generated by the P terminals during the video conversation. First terminal 401 is any one of the P terminals. Here, the processing of the target video by first terminal 401 is described as an example. The processing of the target video by the other terminals among the P terminals, excluding first terminal 401, is the same as the processing by first terminal 401.
[0055] The target video is composed of multiple continuous video frames, and each video frame contained in the target video includes multiple data blocks to be encoded. The encoder 20 is deployed in the first terminal 401. The first terminal 401 calls the encoder 20 to analyze the scene complexity of each data block in each video frame contained in the target video, determines the encoding mode for each data block, and encodes each data block according to the determined encoding mode, and finally obtains the encoded target video. The first terminal 401 can send the encoded target video to the server 403, and the server 403 transmits the encoded target video to other terminals participating in the video session; alternatively, the first terminal 401 can also directly transmit the encoded target video to other terminals participating in the video session; the other terminals receive the encoded target video, parse and play the target video, so as to realize a video session in which P terminals participate.
[0056] (2) The encoder 20 is deployed in the server 403, and the server 403 performs video processing.
[0057] P users participate in a video conversation using video conversation applications running on P terminals, respectively. For example, user 1 participates in a video conversation using the video conversation application running on first terminal 401, while user 2 participates in a video conversation using the video conversation application running on second terminal 402. Server 403 is configured to process and transmit target videos generated by the P terminals during the video conversation. The target videos may include conversation videos generated by the P terminals during the video conversation.
[0058] The target video can be a conversation video generated by the first terminal 401 during the video conversation. The first terminal 401 transmits the target video to the server 403. The target video is composed of multiple continuous video frames, and each video frame contained in the target video includes multiple data blocks to be encoded. The encoder 20 is deployed in the server 403. The server 403 calls the encoder 20 to analyze the scene complexity of each data block in each video frame contained in the target video, determine the encoding mode for each data block, and encode each data block according to the determined encoding mode, and finally obtain the encoded target video. The server 403 transmits the encoded target video to the other terminals participating in the video conversation; the other terminals receive the encoded target video, parse and play the target video, so as to realize the video conversation in which P terminals participate.
[0059] In an embodiment of the present application, each terminal or server participating in a video session can call an encoder to encode a target video generated during the video session, and the encoding mode of each data block contained in the target video by each terminal or server is determined according to the scene complexity of each data block, which effectively improves the encoding speed of the data block, thereby improving the encoding speed of the target video. While ensuring the quality of the target video, the encoding speed of the target video is greatly accelerated, thereby improving the smoothness of the target video, improving the session quality of the video session to a certain extent, optimizing the video session effect, and improving the user experience.
[0060] FIG5 is a flow chart of a video processing method provided in an embodiment of the present application. The video processing method can be executed by an intelligent device, which can be a user terminal or a server. The user terminal can be a device with a camera function, such as a smart phone, a tablet computer, or a smart wearable device. The intelligent device can be, for example, any terminal or server in the video processing system shown in FIG4 . The video processing method includes the following steps S501 to S505:
[0061] S501: Obtain a target video frame from a video to be encoded, and determine a target data block to be encoded from the target video frame. The video to be encoded consists of multiple consecutive video frames. The target video frame is obtained from the video to be encoded, and the target video frame is any video frame in the video to be encoded. The target video frame contains multiple data blocks, which may include already encoded data blocks and data blocks to be encoded. The target data block to be encoded is determined from the target video frame, and the target data block is any data block to be encoded in the target video frame.
[0062] S502, perform scene complexity analysis on the target data block to obtain data block index information. The data block index information obtained by performing scene complexity analysis on the target data block may include but is not limited to at least one of the following: an estimated distortion parameter of the target data block, a spatial information parameter of the target data block, and a time information parameter of the target data block. Among them, the estimated distortion (Distortion, Dist) parameter is obtained by estimating the degree of distortion during intra-frame prediction coding or inter-frame prediction coding. The estimated distortion parameter of the target data block can be used to measure the degree of distortion of the reconstructed target data block compared to the original target data block. The original target data block refers to the target data block that has not been encoded, and the reconstructed target data block refers to the encoded target data block obtained after performing intra-frame prediction coding or inter-frame prediction coding on the target data block. The spatial information parameter of the target data block may refer to the value of the spatial information calculated for the target data block; the time information parameter of the target data block may refer to the value of the time information calculated for the target data block.
[0063] S503 , dividing the target data block into N sub-data blocks, and performing scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2.
[0064] The sub-block indicator information may include N sub-block indicator data, the i-th sub-data block is any sub-data block among the N sub-data blocks, the i-th sub-block indicator data is any sub-block indicator data among the N sub-block indicator data, the i-th sub-data block corresponds to the i-th sub-block indicator data, the i-th sub-block indicator data is obtained by performing scene complexity analysis on the i-th sub-data block, i∈[1,N].
[0065] The i-th sub-block indicator data may include but is not limited to at least one of the following: a distortion estimation parameter of the i-th sub-data block, a spatial information parameter of the i-th sub-data block, and a time information parameter of the i-th sub-data block. Among them, the estimated distortion parameter of the i-th sub-data block can be used to measure the degree of distortion of the reconstructed i-th sub-data block compared to the original i-th sub-data block. The original i-th sub-data block refers to the i-th sub-data block that has not been encoded, and the reconstructed i-th sub-data block refers to the encoded i-th sub-data block obtained by performing intra-frame prediction encoding or inter-frame prediction encoding on the i-th sub-data block. The spatial information parameter of the i-th sub-data block may refer to the value of the spatial information calculated for the i-th sub-data block; the time information parameter of the i-th sub-data block may refer to the value of the time information calculated for the i-th sub-data block.
[0066] S504: Determine a coding mode for the target data block according to the data block index information and the sub-block index information.
[0067] In one embodiment, the data block index information and the sub-block index information can be input into a joint statistical model, and the data block index information and the sub-block index information can be calculated using the joint statistical model to obtain an output value obtained after the joint statistical model calculates the data block index information and the sub-block index information, and the coding mode for the target data block is determined based on the output value of the joint statistical model. The joint statistical model can be trained using the data block index information of a data block with a determined coding mode and the sub-block index information of N sub-data blocks obtained by dividing the data block with a determined coding mode. In an embodiment, the joint statistical model can perform weighted calculation on the data block index information and the sub-block index information to obtain an output value, wherein the weighting factor can be obtained by training the relevant information of the data block with a determined coding mode and its sub-blocks.
[0068] If the output value of the joint statistical model satisfies the first condition, that is, the output value is greater than the first division threshold, then the encoding mode for the target data block is determined to be the first mode. The fact that the output value of the joint statistical model satisfies the first condition indicates that the correlation between the N sub-data blocks obtained by dividing the target data block is weak, and it is more inclined to divide the target data block into N sub-data blocks and then encode each sub-data block separately. If the output value of the joint statistical model satisfies the second condition, that is, the output value is less than the second division threshold, then the encoding mode for the target data block is determined to be the second mode. The fact that the output value of the joint statistical model satisfies the second condition indicates that the correlation between the N sub-data blocks obtained by dividing the target data block is strong, and it is more inclined not to divide the target data block, but to directly encode the target data block. It should be noted that the first division threshold and the second division threshold can be trained during the training process of the joint statistical model.
[0069] S505: Encode the target data block according to the determined encoding mode. If the encoding mode is the first mode, encoding the target data block according to the determined encoding mode may refer to encoding the target data block according to the first mode, specifically, dividing the target data block into N sub-data blocks and inputting each sub-data block into an encoder for encoding. If the encoding mode is the second mode, encoding the target data block according to the determined encoding mode may refer to encoding the target data block according to the second mode, specifically, directly inputting the target data block into an encoder for encoding.
[0070] In an embodiment of the present application, the scene complexity analysis result (i.e., data block index information and sub-block index information) of the scene complexity analysis of the target data block is input into a joint statistical model for calculation, and the encoding mode of the target data block is determined according to the output value calculated by the joint statistical model. The video processing solution provided by the embodiment of the present application is applicable to any video scene, and can adaptively adjust the encoding mode of the target data block according to the scene complexity of the target data block to be encoded in any video, thereby determining a coding mode that is suitable for the scene complexity of the target data block, effectively improving the encoding speed of the target data block, thereby improving the video encoding speed, and being able to obtain the best balance between the encoding speed of the video and the encoding bit rate of the video.
[0071] FIG6 is a flow chart of another video processing method provided in an embodiment of the present application. The video processing method can be executed by an intelligent device, which can be a user terminal or a server. The user terminal can be a device with a camera function, such as a smart phone, a tablet computer, or a smart wearable device. The intelligent device can be, for example, any terminal or server in the video processing system shown in FIG4 . The video processing method includes the following steps S601 to S608:
[0072] S601 : Acquire a target video frame from a video to be encoded, and determine a target data block to be encoded from the target video frame.
[0073] S602: Perform scene complexity analysis on the target data block to obtain data block indicator information.
[0074] S603: Divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2.
[0075] The execution process of step S601 in the embodiment of the present application is the same as the execution process of step S501 in the embodiment shown in Figure 5, the execution process of step S602 is the same as the execution process of step S502 in the embodiment shown in Figure 5, and the execution process of step S603 is the same as the execution process of step S503 in the embodiment shown in Figure 5. For the specific execution process, please refer to the description of the embodiment shown in Figure 5 and will not be repeated here.
[0076] S604: Input the data block index information and the sub-block index information into the joint statistical model.
[0077] S605: Obtain an output value obtained after the joint statistical model calculates the data block index information and the sub-block index information.
[0078] S606: If the output value satisfies the third condition, perform scene complexity analysis on the M associated data blocks related to the target data block to obtain associated block indicator information of the M associated data blocks, where M is a positive integer.
[0079] S607: Determine the encoding mode for the target data block according to the associated block indicator information.
[0080] In S606 to S607, if the output value of the joint statistical model satisfies the third condition, i.e., the output value is less than or equal to the first partitioning threshold and greater than or equal to the second partitioning threshold, it indicates that the correlation between the N sub-data blocks obtained by partitioning the target data block is between strong and weak, and there is no obvious tendency to encode each sub-data block separately after partitioning the target data block, nor is there an obvious tendency to encode the target data block directly without partitioning the target data block. In this case, M associated data blocks related to the target data block are determined, and scene complexity analysis is performed on the M associated data blocks to obtain associated block indicator information for the M associated data blocks. Based on the associated block indicator information for the M associated data blocks, an encoding mode for the target data block is determined, where M is a positive integer. The associated block indicator information for the M associated data blocks may include a first number of associated data blocks in the M associated data blocks that are divided into multiple sub-data blocks for encoding.
[0081] In one embodiment, if the first number satisfies the fourth condition, i.e., the first number is greater than or equal to the first number threshold, then this indicates that the scene complexity of the M associated data blocks associated with the target data block is relatively high. Therefore, the target data block may be first divided into N sub-data blocks to be analyzed (i.e., N sub-data blocks), and scene complexity analysis may be performed on the N sub-data blocks to be analyzed to obtain sub-data block indicator information for the N sub-data blocks to be analyzed. The encoding mode for the target data block may then be further determined based on the sub-data block indicator information. It should be noted that the first number threshold may be set based on empirical values. For example, if the number of associated data blocks associated with the target data block is 3, the first number threshold may be set to 2. Taking Figure 3 as an example, the three associated data blocks associated with target data block 301 are associated data block 302, associated data block 303, and associated data block 304. Associated data blocks 302 and 303 are divided into multiple sub-data blocks for encoding, while associated data block 304 is not divided and is directly encoded. That is, the first number of associated data blocks divided into multiple sub-data blocks for encoding among the three associated data blocks is two. If the first number meets the fourth condition, it can be determined that the scene complexity of the three associated data blocks associated with target data block 301 is relatively high, and it is possible to first attempt to divide target data block 301 into N sub-data blocks to be analyzed. The indicator information of the sub-data blocks to be analyzed can include a second number of sub-data blocks to be analyzed among the N sub-data blocks to be analyzed that meet the further division condition. Specifically, the index information of the sub-data block to be analyzed may refer to the second number of output values that meet the first condition among the N output values obtained by calculating the data block index information of each of the N sub-data blocks to be analyzed and the sub-block indication information of the N sub-data blocks into which each data block to be analyzed is divided using a joint statistical model.
[0082] In one embodiment, if the second number satisfies the fifth condition, i.e., the second number is greater than or equal to the second number threshold, the encoding mode for the target data block can be determined to be the first mode. If the second number does not satisfy the fifth condition, i.e., the second number is less than the second number threshold, direct encoding of the target data block is attempted, and the encoding mode for the target data block is determined based on the encoding information of the target data block obtained by directly encoding the target data block, and based on the encoding information of the N sub-data blocks obtained by dividing the target data block into N sub-data blocks and encoding each sub-data block separately. It should be noted that the second number threshold can also be set based on an empirical value. For example, if the target data block is divided into four sub-data blocks to be analyzed, the second number threshold can be set to 3. If all four sub-data blocks to be analyzed meet the further division condition, i.e., the second number of sub-data blocks to be analyzed that meet the further division condition is four, and the second number satisfies the fifth condition, the encoding mode for the target data block can be determined to be the first mode.
[0083] In one embodiment, if the first number does not meet the fourth condition, that is, the first number is less than the first number threshold, it indicates that the scene complexity of the M associated data blocks related to the target data block is relatively low. You can first try to encode the target data block directly without dividing the data block, and obtain the encoding information of the target data block obtained by encoding the target data block. The encoding information of the target data block may include the encoding distortion parameter of the target data block. It should be noted that the distortion estimation parameter of the target data block may refer to an estimated value of the degree of distortion during the encoding process of the target data block, and the encoding distortion parameter of the target data block can be used to measure the actual degree of distortion of the encoded target data block (that is, the reconstructed target data block) compared to the target data block before encoding (that is, the original target data block). Furthermore, the encoding parameter is calculated based on the encoding distortion parameter of the target data block and the quantization parameter (Quatization Parameter, QP) for quantizing the target data block. The calculation process of the encoding parameter can be referred to the following formula 1:
[0084] Code=Dist / QP 2 Formula 1
[0085] As shown in the above formula 1, Code represents the encoding parameter, Dist represents the encoding distortion parameter of the target data block, and QP represents the quantization parameter.
[0086] In one embodiment, if the encoding parameter satisfies the sixth condition, that is, the encoding parameter is less than the third division threshold, then the encoding mode for the target data block can be determined to be the second mode. In another embodiment, if the encoding parameter does not satisfy the sixth condition, that is, the encoding parameter is greater than or equal to the third division threshold, an attempt is made to divide the target data block into N sub-data blocks, and each sub-data block is encoded separately, and the encoding mode for the target data block is determined based on the encoding information of the target data block obtained by directly encoding the target data block, and based on the encoding information of the N sub-data blocks obtained by dividing the target data block into N sub-data blocks and encoding each sub-data block separately. It should be noted that the third division threshold can be trained during the training process of the joint statistical model.
[0087] In the above two embodiments, the encoding information of the target data block may further include a first rate-distortion loss parameter for the target data block; and the encoding information of the N sub-data blocks may include second rate-distortion loss parameters for the N sub-data blocks. The second rate-distortion loss parameter is calculated based on the third rate-distortion loss parameter for each of the N sub-data blocks. For example, the second rate-distortion loss parameter may be the sum of the third rate-distortion loss parameters for each of the N sub-data blocks. The first rate-distortion loss parameter is calculated based on the encoding rate and the encoding distortion parameter according to the encoding mode for directly encoding the target data block. For example, the first rate-distortion loss parameter may be the ratio of the encoding rate to the encoding distortion parameter. The first rate-distortion loss parameter can be used to measure the encoding effect of directly encoding the target data block. The smaller the first rate-distortion loss parameter, the better the encoding effect of the target data block. A good encoding effect of the target data block may mean that the encoding rate of the target data block is low while ensuring that the distortion of the encoded target data block is lower than that of the target data block before encoding. Similarly, the second rate-distortion loss parameter can be used to measure the encoding effect of dividing the target data block into N sub-data blocks and encoding each sub-data block separately. The smaller the second rate-distortion loss parameter, the better the encoding effect of the target data block. If the first rate-distortion loss parameter is greater than or equal to the second rate-distortion loss parameter, the encoding mode for the target data block is determined to be the first mode; if the first rate-distortion loss parameter is less than the second rate-distortion loss parameter, the encoding mode for the target data block is determined to be the second mode.
[0088] S608: Encode the target data block according to the determined encoding mode.
[0089] The execution process of step S608 in the embodiment of the present application is the same as the execution process of step S505 in the embodiment shown in FIG5 . For the specific execution process, please refer to the description of the embodiment shown in FIG5 , which will not be repeated here.
[0090] In the embodiment of the present application, the scene complexity of the associated data blocks related to the target data block can also reflect the scene complexity of the target data block to a certain extent. Therefore, the embodiment of the present application jointly determines the encoding mode of the target data block according to the scene complexity of the target data block (i.e., the output value of the joint statistical model) and the scene complexity of the associated data blocks related to the target data block (i.e., the associated block indication information). The video processing solution provided by the embodiment of the present application is universal and can adaptively adjust the encoding mode of the target data block according to the scene complexity of the target data block to be encoded in any video, thereby determining the encoding mode that is suitable for the scene complexity of the target data block, effectively improving the encoding speed of the target data block, and then improving the video encoding speed, and can effectively improve the video encoding speed without losing the video encoding efficiency, and obtain the best balance between the video encoding speed and the video encoding bit rate.
[0091] FIG7 is a flow chart of another video processing method provided in an embodiment of the present application. The video processing method can be executed by an intelligent device, which can be a user terminal or a server. The user terminal can be a device with a camera function, such as a smart phone, a tablet computer, or a smart wearable device. The intelligent device can be, for example, any terminal or server in the video processing system shown in FIG4 . The video processing method includes the following steps S701 to S720:
[0092] S701 , obtaining a target video frame from a video to be encoded, and determining a target data block to be encoded from the target video frame.
[0093] S702: Perform scene complexity analysis on the target data block to obtain data block indicator information.
[0094] S703: Divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information.
[0095] S704: Input the data block index information and the sub-block index information into the joint statistical model for calculation to obtain an output value of the joint statistical model.
[0096] S705: Determine whether the output value satisfies a first condition. If the output value satisfies the first condition, execute step S706; if the output value does not satisfy the first condition, execute step S707.
[0097] S706: If the output value satisfies the first condition, the encoding mode for the target data block is determined to be the first mode. After step S706 is completed, step S720 is executed.
[0098] S707: If the output value does not meet the first condition, determine whether the output value meets the second condition. If the output value meets the second condition, execute step S708; if the output value does not meet the second condition, execute step S709.
[0099] S708: If the output value satisfies the second condition, the encoding mode for the target data block is determined to be the second mode. After step S708 is completed, step S720 is executed.
[0100] S709: If the output value does not meet the second condition, determine whether the output value meets the third condition. If the output value meets the third condition, execute step S710.
[0101] S710: If the output value satisfies a third condition, determine a first number of associated data blocks that are divided into a plurality of sub-data blocks for encoding among the M associated data blocks related to the target data block.
[0102] S711, determine whether the first quantity meets the fourth condition. If the first quantity meets the fourth condition, execute step S712; if the first quantity does not meet the fourth condition, execute step S716.
[0103] S712: If the first number satisfies the fourth condition, preferentially try to divide the target sub-data block into N sub-data blocks, and encode each sub-data block separately.
[0104] S713: Determine a second number of sub-data blocks that meet the further division condition in the N sub-data blocks.
[0105] S714: Determine whether the second number satisfies the fifth condition. If the second number satisfies the fifth condition, determine that the encoding mode for the target data block is the first mode, and execute step S720. If the second number does not satisfy the fifth condition, execute step S715.
[0106] At step S715, an attempt is made to directly encode the target data block without dividing the target data, and an encoding mode for the target data block is determined based on the encoding information. Specifically, encoding information for the target data block obtained by directly encoding the target data block is obtained, as well as encoding information for each of the N sub-data blocks obtained by dividing the target data block into N sub-data blocks and encoding each sub-data block separately. The encoding mode for the target data block is determined based on the encoding information for the target data block and the encoding information for the N sub-data blocks. The specific execution process can be seen in the embodiment shown in FIG6. After step S715 is completed, step S720 is executed.
[0107] S716: If the first number does not satisfy the fourth condition, preferentially try not to divide the target data block, directly encode the target data block, and determine the encoding distortion parameter of the target data block.
[0108] S717: Calculate and obtain a coding parameter according to the coding distortion parameter and the quantization parameter.
[0109] S718: Determine whether the encoding parameters meet the sixth condition. If the encoding parameters meet the sixth condition, determine that the encoding mode for the target data block is the second mode, and execute step S720; if the encoding parameters do not meet the sixth condition, execute step S719.
[0110] At step S719, an attempt is made to divide the target sub-data block into N sub-data blocks, and each sub-data block is encoded separately. The encoding mode for the target data block is determined based on the encoding information. Specifically, encoding information for the target data block obtained by directly encoding the target data block is obtained, as well as encoding information for the N sub-data blocks obtained by dividing the target data block into N sub-data blocks and encoding each sub-data block separately. The encoding mode for the target data block is determined based on the encoding information for the target data block and the encoding information for the N sub-data blocks. The specific execution process can be seen in the embodiment shown in FIG6. After step S719 is completed, step S720 is executed.
[0111] S720: Encode the target data block according to the determined encoding mode. After encoding the target data block according to the determined encoding mode, another data block to be encoded other than the target data block can be determined from the video frame to be encoded, and the encoding mode for the data block to be encoded is determined based on the scene complexity of the data block to be encoded. The data block to be encoded is encoded according to the determined encoding mode. The process of determining the encoding mode for the data block to be encoded based on the scene complexity of the data block to be encoded is the same as the process of determining the encoding mode for the target data block based on the scene complexity of the target data block.
[0112] In the embodiment of the present application, the scene complexity of the associated data blocks related to the target data block can also reflect the scene complexity of the target data block to a certain extent. Therefore, the embodiment of the present application jointly determines the encoding mode of the target data block based on the scene complexity of the target data block (i.e., the output value of the joint statistical model) and the scene complexity of the associated data blocks related to the target data block (i.e., the associated block indication information). The embodiment of the present application sets 6 judgment conditions and conducts a multi-angle comprehensive analysis of the scene complexity of the target data block, so that the encoding mode of the target data block finally determined is highly adaptable to the scene complexity of the target data block, effectively improving the encoding speed of the target data block, and thus improving the video encoding speed, and can effectively improve the video encoding speed without losing the video encoding quality and video encoding efficiency while ensuring the clarity and smoothness of the encoded video.
[0113] Please refer to Figure 8, which is a structural diagram of a video processing device provided in an embodiment of the present application. The video processing device 80 in the embodiment of the present application can be set in an intelligent device, which can be a smart terminal or a server. The video processing device 80 can be used to execute the corresponding steps in the video processing method shown in Figure 5, Figure 6 or Figure 7. The video processing device 80 includes the following units.
[0114] The acquisition unit 801 is configured to acquire a target video frame from a video to be encoded, and determine a target data block to be encoded from the target video frame;
[0115] Processing unit 802 is used to perform scene complexity analysis on the target data block to obtain data block index information; divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2; determine the encoding mode for the target data block based on the data block index information and the sub-block index information; and encode the target data block according to the determined encoding mode.
[0116] In one embodiment, the data block indicator information includes any one or more of a distortion estimation parameter of the target data block, a spatial information parameter of the target data block, and a temporal information parameter of the target data block.
[0117] The sub-block indicator information includes: N sub-block indicator data, wherein the i-th sub-block indicator data among the N sub-block indicator data includes: any one or more of the distortion estimation parameters of the i-th sub-data block among the N sub-data blocks, the spatial information parameters of the i-th sub-data block, and the time information parameters of the i-th sub-data block, i∈[1,N].
[0118] In one embodiment, the processing unit 802 is specifically configured to:
[0119] Inputting the data block indicator information and the sub-block indicator information into the joint statistical model;
[0120] Obtaining the output value obtained after the joint statistical model calculates the data block indicator information and the sub-block indicator information;
[0121] The encoding mode for the target data block is determined according to the output value.
[0122] In one embodiment, the processing unit 802 is specifically configured to:
[0123] If the output value satisfies the first condition, the encoding mode is determined to be the first mode;
[0124] Divide the target data block into N sub-data blocks, and input each sub-data block into the encoder for encoding;
[0125] The output value meeting the first condition means that the output value is greater than the first division threshold.
[0126] In one embodiment, the processing unit 802 is specifically configured to:
[0127] If the output value satisfies the second condition, determining the encoding mode to be the second mode;
[0128] Input the target data block into the encoder for encoding;
[0129] The output value meeting the second condition means that the output value is less than the second division threshold.
[0130] In one embodiment, the processing unit 802 is specifically configured to:
[0131] If the output value satisfies the third condition, then performing scene complexity analysis on the M associated data blocks related to the target data block to obtain associated block indicator information of the M associated data blocks, where M is a positive integer;
[0132] Determining a coding mode for a target data block according to associated block indicator information;
[0133] The output value meeting the third condition means that the output value is less than or equal to the first classification threshold, and the output value is greater than or equal to the second classification threshold.
[0134] In one embodiment, the processing unit 802 is specifically configured to:
[0135] Acquire associated block indicator information, where the associated block indicator information includes: a first number of associated data blocks divided into a plurality of sub-data blocks for encoding in the M associated data blocks;
[0136] If the first number satisfies the fourth condition, dividing the target data block into N sub-data blocks to be analyzed;
[0137] Performing scene complexity analysis on the N sub-data blocks to be analyzed, and determining sub-data block indicator information of the N sub-data blocks to be analyzed;
[0138] Determine the encoding mode for the target data block according to the indicator information of the sub-data block to be analyzed;
[0139] The first quantity meeting the fourth condition means that the first quantity is greater than or equal to a first quantity threshold.
[0140] In one embodiment, the processing unit 802 is specifically configured to:
[0141] Obtaining index information of the sub-data blocks to be analyzed, the index information of the sub-data blocks to be analyzed including: a second number of the sub-data blocks to be analyzed that meet the further division condition among the N sub-data blocks to be analyzed;
[0142] If the second number satisfies the fifth condition, determining the encoding mode to be the first mode;
[0143] If the second number does not satisfy the fifth condition, inputting the target data block into the encoder for encoding, and determining an encoding mode for the target data block according to encoding information of the target data block obtained by encoding;
[0144] Among them, the second quantity meets the fifth condition means that the second quantity is greater than or equal to the second quantity threshold; the second quantity does not meet the fifth condition means that the second quantity is less than the second quantity threshold.
[0145] In one embodiment, the processing unit 802 is further configured to:
[0146] If the first number does not satisfy the fourth condition, inputting the target data block into the encoder for encoding;
[0147] Determining a coding mode for the target data block according to coding information of the target data block obtained by coding;
[0148] The fact that the first quantity does not satisfy the fourth condition means that the first quantity is less than the first quantity threshold.
[0149] In one embodiment, the coding information of the target data block includes a coding distortion parameter of the target data block; the processing unit 802 is specifically configured to:
[0150] A coding parameter is calculated based on a coding distortion parameter of the target data block and a quantization parameter for quantizing the target data block;
[0151] If the encoding parameter satisfies the sixth condition, determining the encoding mode to be the second mode;
[0152] If the encoding parameter does not satisfy the sixth condition, the target data block is divided into N sub-data blocks, and each of the N sub-data blocks is input into the encoder for encoding;
[0153] Determining a coding mode for the target data block according to the coding information of the target data block and the coding information of the N sub-data blocks obtained by coding;
[0154] The fact that the coding parameter satisfies the sixth condition means that the coding parameter is less than the third division threshold; the fact that the coding parameter does not satisfy the sixth condition means that the coding parameter is greater than or equal to the third division threshold.
[0155] In one embodiment, the encoding information of the target data block further includes a first rate-distortion loss parameter of the target data block; the encoding information of the N sub-data blocks includes second rate-distortion loss parameters of the N sub-data blocks, where the second rate-distortion loss parameter is calculated based on a third rate-distortion loss parameter of each of the N sub-data blocks; the processing unit 802 is specifically configured to:
[0156] If the first coding rate-distortion loss parameter is greater than or equal to the second coding rate-distortion loss parameter, determining that the coding mode for the target data block is the first mode;
[0157] If the first coding rate-distortion loss parameter is smaller than the second coding rate-distortion loss parameter, the coding mode for the target data block is determined to be the second mode.
[0158] According to one embodiment of the present application, the various units in the video processing device 80 shown in Figure 8 can be individually or entirely combined into one or several other units to constitute, or one (or more) of the units can be further divided into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the function of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the video processing device 80 can also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, the video processing device 80 shown in Figure 8 can be constructed by running a computer program (including program code) capable of executing the steps involved in the corresponding method shown in Figure 5, Figure 6 or Figure 7 on a general-purpose computing device such as a general-purpose computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory medium (RAM), and a read-only memory medium (ROM), to achieve the video processing method of the embodiment of the present application. The computer program can be recorded on a computer-readable storage medium, for example, and loaded into any terminal (such as the first terminal 401 or the second terminal 402, etc.) or the server 403 of the video processing system shown in FIG4 through the computer-readable storage medium and run therein.
[0159] In an embodiment of the present application, by performing scene complexity analysis on the target data block to be encoded in the target video frame, analysis results of the entire target data block to be encoded and analysis results after dividing the target data block to be encoded into multiple sub-data blocks are obtained respectively, and the encoding mode is selected based on these analysis results. It is possible to more accurately determine a suitable encoding mode for the current target data block to be encoded, effectively improve the encoding speed of the target data block, and improve the video encoding efficiency.
[0160] Please refer to Figure 9, which is a schematic diagram of the structure of a smart device provided in an embodiment of the present application. The smart device 90 includes at least a processor 901 and a memory 902. The processor 901 and the memory 902 may be connected via a bus or other means.
[0161] Processor 901 may be a central processing unit (CPU). Processor 901 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or the like. The PLD may be a field-programmable gate array (FPGA), a generic array logic (GAL), or the like.
[0162] The memory 902 may include a volatile memory, such as a random-access memory (RAM); the memory 902 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the memory 902 may also include a combination of the above types of memory.
[0163] Memory 902 is used to store computer programs, which include computer instructions. Processor 901 is used to execute computer instructions. Processor 901 (also known as CPU (Central Processing Unit)) is the computing and control core of smart device 90. It is suitable for implementing one or more computer instructions, and is specifically suitable for loading and executing one or more computer instructions to implement corresponding method processes or corresponding functions.
[0164] The smart device 90 may be any terminal (e.g., the first terminal 401 or the second terminal 402, etc.) or the server 403 in the video processing system shown in FIG4 ; the memory 902 stores a computer program, which includes one or more computer instructions; the processor 901 loads and executes the one or more computer instructions to implement the corresponding steps in the method embodiments shown in FIG5 , FIG6 , or FIG7 ; in a specific implementation, the computer instructions in the memory 902 are loaded by the processor 901 and the following steps are executed:
[0165] Obtaining a target video frame from a video to be encoded, and determining a target data block to be encoded from the target video frame;
[0166] Perform scene complexity analysis on the target data block to obtain data block indicator information;
[0167] Divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2;
[0168] Determining a coding mode for a target data block according to the data block indicator information and the sub-block indicator information;
[0169] The target data block is encoded according to the determined encoding mode.
[0170] In one embodiment, the data block indicator information includes: any one or more of a distortion estimation parameter of the target data block, a spatial information parameter of the target data block, and a temporal information parameter of the target data block;
[0171] The sub-block indicator information includes: N sub-block indicator data, wherein the i-th sub-block indicator data among the N sub-block indicator data includes: any one or more of the distortion estimation parameters of the i-th sub-data block among the N sub-data blocks, the spatial information parameters of the i-th sub-data block, and the time information parameters of the i-th sub-data block, i∈[1,N].
[0172] In one embodiment, when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically performed:
[0173] Inputting the data block indicator information and the sub-block indicator information into the joint statistical model;
[0174] Obtaining the output value obtained after the joint statistical model calculates the data block indicator information and the sub-block indicator information;
[0175] The encoding mode for the target data block is determined according to the output value.
[0176] In one embodiment, when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically performed:
[0177] If the output value satisfies the first condition, the encoding mode is determined to be the first mode;
[0178] Divide the target data block into N sub-data blocks, and input each sub-data block into the encoder for encoding;
[0179] The output value meeting the first condition means that the output value is greater than the first division threshold.
[0180] In one embodiment, when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically performed:
[0181] If the output value satisfies the second condition, determining the encoding mode to be the second mode;
[0182] Input the target data block into the encoder for encoding;
[0183] The output value meeting the second condition means that the output value is less than the second division threshold.
[0184] In one embodiment, when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically performed:
[0185] If the output value satisfies the third condition, then performing scene complexity analysis on the M associated data blocks related to the target data block to obtain associated block indicator information of the M associated data blocks, where M is a positive integer;
[0186] Determining a coding mode for a target data block according to associated block indicator information;
[0187] The output value meeting the third condition means that the output value is less than or equal to the first classification threshold, and the output value is greater than or equal to the second classification threshold.
[0188] In one embodiment, when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically performed:
[0189] Acquire associated block indicator information, where the associated block indicator information includes: a first number of associated data blocks divided into a plurality of sub-data blocks for encoding in the M associated data blocks;
[0190] If the first number satisfies the fourth condition, dividing the target data block into N sub-data blocks to be analyzed;
[0191] Performing scene complexity analysis on the N sub-data blocks to be analyzed, and determining sub-data block indicator information of the N sub-data blocks to be analyzed;
[0192] Determine the encoding mode for the target data block according to the indicator information of the sub-data block to be analyzed;
[0193] The first quantity meeting the fourth condition means that the first quantity is greater than or equal to a first quantity threshold.
[0194] In one embodiment, when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically performed:
[0195] Obtaining index information of the sub-data blocks to be analyzed, the index information of the sub-data blocks to be analyzed including: a second number of the sub-data blocks to be analyzed that meet the further division condition among the N sub-data blocks to be analyzed;
[0196] If the second number satisfies the fifth condition, determining the encoding mode to be the first mode;
[0197] If the second number does not satisfy the fifth condition, inputting the target data block into the encoder for encoding, and determining an encoding mode for the target data block according to encoding information of the target data block obtained by encoding;
[0198] Among them, the second quantity meets the fifth condition means that the second quantity is greater than or equal to the second quantity threshold; the second quantity does not meet the fifth condition means that the second quantity is less than the second quantity threshold.
[0199] In one embodiment, the computer instructions in the memory 902, when loaded by the processor 901, further perform the following steps:
[0200] If the first number does not satisfy the fourth condition, inputting the target data block into the encoder for encoding;
[0201] Determining a coding mode for the target data block according to coding information of the target data block obtained by coding;
[0202] The fact that the first quantity does not satisfy the fourth condition means that the first quantity is less than the first quantity threshold.
[0203] In one embodiment, the coding information of the target data block includes coding distortion parameters of the target data block; when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically executed:
[0204] A coding parameter is calculated based on a coding distortion parameter of the target data block and a quantization parameter for quantizing the target data block;
[0205] If the encoding parameter satisfies the sixth condition, determining the encoding mode to be the second mode;
[0206] If the encoding parameter does not satisfy the sixth condition, the target data block is divided into N sub-data blocks, and each of the N sub-data blocks is input into the encoder for encoding;
[0207] Determining a coding mode for the target data block according to the coding information of the target data block and the coding information of the N sub-data blocks obtained by coding;
[0208] The fact that the coding parameter satisfies the sixth condition means that the coding parameter is less than the third division threshold; the fact that the coding parameter does not satisfy the sixth condition means that the coding parameter is greater than or equal to the third division threshold.
[0209] In one embodiment, the encoding information of the target data block further includes a first rate-distortion loss parameter of the target data block; the encoding information of the N sub-data blocks includes second rate-distortion loss parameters of the N sub-data blocks, where the second rate-distortion loss parameter is calculated based on a third rate-distortion loss parameter of each of the N sub-data blocks; and when the computer instructions in the memory 902 are loaded by the processor 901, the following steps are specifically executed:
[0210] If the first coding rate-distortion loss parameter is greater than or equal to the second coding rate-distortion loss parameter, determining that the coding mode for the target data block is the first mode;
[0211] If the first coding rate-distortion loss parameter is smaller than the second coding rate-distortion loss parameter, the coding mode for the target data block is determined to be the second mode.
[0212] In an embodiment of the present application, by performing scene complexity analysis on the target data block to be encoded in the target video frame, analysis results of the entire target data block to be encoded and analysis results after dividing the target data block to be encoded into multiple sub-data blocks are obtained respectively, and the encoding mode is selected based on these analysis results. It is possible to more accurately determine a suitable encoding mode for the current target data block to be encoded, effectively improve the encoding speed of the target data block, and improve the video encoding efficiency.
[0213] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video processing methods provided in the various optional embodiments described above.
[0214] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0215] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present application are still within the scope covered by the present application.
Claims
1. A video processing method, performed by a smart device, comprising: Acquire a target video frame from a video to be encoded, and determine a target data block to be encoded from the target video frame; Performing scene complexity analysis on the target data block to obtain data block indicator information; Divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2; Determining a coding mode for the target data block according to the data block indicator information and the sub-block indicator information; The target data block is encoded according to the determined encoding mode.
2. The method according to claim 1, wherein The data block indicator information includes: any one or more of a distortion estimation parameter of the target data block, a spatial information parameter of the target data block, and a temporal information parameter of the target data block; The sub-block indicator information includes: N sub-block indicator data, wherein the i-th sub-block indicator data among the N sub-block indicator data includes: any one or more of the distortion estimation parameters of the i-th sub-data block among the N sub-data blocks, the spatial information parameters of the i-th sub-data block, and the time information parameters of the i-th sub-data block, i∈[1,N].
3. The method according to claim 1, wherein The determining, according to the data block indicator information and the sub-block indicator information, a coding mode for the target data block includes: inputting the data block indicator information and the sub-block indicator information into a joint statistical model; Obtaining an output value obtained after the joint statistical model calculates the data block indicator information and the sub-block indicator information; An encoding mode for the target data block is determined according to the output value.
4. The method according to claim 3, wherein: The determining of the encoding mode for the target data block according to the output value includes: If the output value satisfies a first condition, determining that the encoding mode is the first mode; The encoding of the target data block according to the determined encoding mode includes: Dividing the target data block into the N sub-data blocks, and inputting each sub-data block into an encoder for encoding; The output value meeting the first condition means that the output value is greater than a first division threshold.
5. The method according to claim 3, wherein: The determining of the encoding mode for the target data block according to the output value includes: If the output value satisfies the second condition, determining that the encoding mode is the second mode; The encoding of the target data block according to the determined encoding mode includes: Inputting the target data block into an encoder for encoding; The output value meeting the second condition means that the output value is less than a second division threshold.
6. The method of claim 3, wherein: The determining of the encoding mode for the target data block according to the output value includes: If the output value satisfies the third condition, performing scene complexity analysis on M associated data blocks related to the target data block to obtain associated block indicator information of the M associated data blocks, where M is a positive integer; Determining a coding mode for the target data block according to the associated block indicator information; The output value meeting the third condition means that the output value is less than or equal to a first classification threshold, and the output value is greater than or equal to a second classification threshold.
7. The method according to claim 6, wherein: The determining of the encoding mode for the target data block according to the associated block indicator information includes: Acquire the associated block indicator information, where the associated block indicator information includes: a first number of associated data blocks divided into a plurality of sub-data blocks for encoding in the M associated data blocks; If the first number satisfies a fourth condition, dividing the target data block into N sub-data blocks to be analyzed; Performing scene complexity analysis on the N to-be-analyzed sub-data blocks to determine to-be-analyzed sub-data block indicator information of the N to-be-analyzed sub-data blocks; Determining a coding mode for the target data block according to the indicator information of the sub-data block to be analyzed; The first quantity meeting the fourth condition means that the first quantity is greater than or equal to a first quantity threshold.
8. The method of claim 7, wherein: The determining of the encoding mode for the target data block according to the indicator information of the sub-data block to be analyzed includes: Acquire the index information of the sub-data blocks to be analyzed, wherein the index information of the sub-data blocks to be analyzed includes: a second number of the sub-data blocks to be analyzed that meet the further division condition among the N sub-data blocks to be analyzed; If the second number satisfies a fifth condition, determining that the encoding mode is the first mode; If the second number does not satisfy the fifth condition, inputting the target data block into an encoder for encoding, and determining an encoding mode for the target data block according to encoding information of the target data block obtained by encoding; The second quantity meeting the fifth condition means that the second quantity is greater than or equal to a second quantity threshold; the second quantity not meeting the fifth condition means that the second quantity is less than the second quantity threshold.
9. The method of claim 7, wherein: The step of determining the encoding mode for the target data block according to the associated block indicator information further includes: If the first number does not satisfy the fourth condition, inputting the target data block into an encoder for encoding; Determining a coding mode for the target data block according to the coding information of the target data block obtained by coding; The fact that the first quantity does not satisfy the fourth condition means that the first quantity is less than the first quantity threshold.
10. The method according to claim 8 or 9, wherein The encoding information of the target data block includes an encoding distortion parameter of the target data block; The determining of the encoding mode of the target data block according to the encoding information of the target data block obtained by encoding includes: Calculating a coding parameter according to a coding distortion parameter of the target data block and a quantization parameter for quantizing the target data block; If the encoding parameter satisfies the sixth condition, determining that the encoding mode is the second mode; If the encoding parameter does not satisfy the sixth condition, dividing the target data block into the N sub-data blocks, and inputting each of the N sub-data blocks into the encoder for encoding; Determining a coding mode for the target data block according to the coding information of the target data block and the coding information of the N sub-data blocks obtained by coding; The fact that the coding parameter satisfies the sixth condition means that the coding parameter is less than the third division threshold; the fact that the coding parameter does not satisfy the sixth condition means that the coding parameter is greater than or equal to the third division threshold.
11. The method according to claim 10, wherein: The encoding information of the target data block further includes a first rate-distortion loss parameter of the target data block; the encoding information of the N sub-data blocks includes a second rate-distortion loss parameter of the N sub-data blocks, where the second rate-distortion loss parameter is calculated based on a third rate-distortion loss parameter of each of the N sub-data blocks; The determining the encoding mode of the target data block according to the encoding information of the target data block and the encoding information of the N sub-data blocks includes: If the first rate-distortion loss parameter is greater than or equal to the second rate-distortion loss parameter, determining that the encoding mode for the target data block is the first mode; If the first coding rate-distortion loss parameter is smaller than the second coding rate-distortion loss parameter, the coding mode for the target data block is determined to be the second mode.
12. A video processing device, comprising: an acquisition unit, configured to acquire a target video frame from a video to be encoded, and determine a target data block to be encoded from the target video frame; A processing unit is used to perform scene complexity analysis on the target data block to obtain data block index information; divide the target data block into N sub-data blocks, and perform scene complexity analysis on each sub-data block to obtain sub-block index information, where N is an integer greater than or equal to 2; determine an encoding mode for the target data block based on the data block index information and the sub-block index information; and encode the target data block according to the determined encoding mode.
13. A smart device comprising: a processor suitable for implementing a computer program; as well as, A memory storing a computer program, wherein the computer program, when executed by the processor, implements the video processing method according to any one of claims 1 to 11.
14. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is read and executed by a processor, the video processing method according to any one of claims 1 to 11 is implemented.