Video encoding method and apparatus
By using the absolute difference and SATD cost model of Hadamar transform specific to prediction methods in video encoding, selecting the target mode from multiple prediction modes for encoding, the problem of insufficient encoding efficiency and complexity in the prior art is solved, and a more efficient and simplified encoding process is achieved.
Patent Information
- Application Number
- CN202111357945.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-11-16
AI Technical Summary
The existing video encoding technology has shortcomings in encoding efficiency and complexity, especially when choosing a prediction mode, it is difficult to accurately estimate the cost of rate distortion, resulting in high encoding complexity and low efficiency.
By using the prediction method-specific Hadamma transform absolute difference and SATD cost model, the target prediction mode is obtained from multiple prediction modes of multiple prediction modes, and the encoded video frames are encoded based on the target mode. This method combines intra prediction and inter prediction, and uses different SATD cost models to select the target mode to be closer to the estimated value of rate distortion optimization technology.
By estimating the cost of rate distortion more accurately, the rate distortion optimization process is reduced, the encoding complexity is reduced, and the encoding efficiency is improved.
Smart Images

Figure CN114071137B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to video coding technologies, and specifically, to a video coding method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] The High Efficiency Video Coding standard H.265 / HEVC is another collaborative effort between the International Telecommunication Union ITU-T and the International Organization for Standardization ISO / IEC after the previously introduced H.264 / AVC. The High Efficiency Video Coding standard H.265 / HEVC proposes a large number of advanced coding technologies, promotes the development of video technologies, and brings about a code rate saving of about 50% under the same coding quality to meet the needs of new video applications and more efficiently transmit video data.
[0003] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0004] The present disclosure provides a video coding method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0005] According to an aspect of the present disclosure, there is provided a video coding method, including: determining a current prediction unit of a current coding unit of a video frame to be coded; obtaining, based on a Hadamard transform absolute difference and SATD cost model specific to a prediction mode, a target prediction mode corresponding to the current prediction unit from multiple prediction modes of multiple prediction methods, where the multiple prediction methods include intra prediction and inter prediction, the intra prediction corresponds to a first SATD cost model, the inter prediction corresponds to a second SATD cost model different from the first SATD cost model, and where the SATD cost model specific to the prediction mode is used to obtain multiple SATD costs for predicting the current prediction unit using multiple prediction modes corresponding to the prediction mode, and each of the multiple SATD costs includes a residual bit cost related to the prediction mode, the residual bit cost indicating the number of bits required to encode a prediction residual obtained after predicting the current prediction unit; and encoding the video frame to be coded based on the target prediction mode.
[0006] According to another aspect of the present disclosure, there is provided a video encoding apparatus, including: a determination unit configured to determine a current prediction unit of a current coding unit of a video frame to be encoded; an acquisition unit configured to obtain, based on a Hadamard transform absolute difference and SATD cost model corresponding to a prediction mode, a target prediction mode corresponding to the current prediction unit from a plurality of prediction modes of the prediction mode, where the prediction mode is any one of intra prediction and inter prediction, the intra prediction corresponds to a first SATD cost model, the inter prediction corresponds to a second SATD cost model different from the first SATD cost model, and where the SATD cost model corresponding to the prediction mode is used to obtain a plurality of SATD costs for predicting the current prediction unit using a plurality of prediction modes corresponding to the prediction mode, each of the plurality of SATD costs includes a residual bit cost related to the prediction mode, the residual bit cost related to the prediction mode indicates the number of bits required to encode a prediction residual obtained after predicting the current prediction unit, and a video frame encoding unit configured to encode the video frame to be encoded based on the target prediction mode.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the method according to the above.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable the computer to implement the method according to the above.
[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the method according to the above.
[0010] According to one or more embodiments of the present disclosure, by performing a plurality of feature extraction operations arranged in sequence on a target image and performing multi-classification based on the features extracted by the last feature extraction operation among the plurality of feature extraction operations, since the features extracted by each of the plurality of feature extraction operations can be used to distinguish the target image between a first classification and at least a second classification different from the first classification, that is, performing binary classification of the target image relative to the first classification, the features extracted have clear boundaries for the first classification, and using them for multi-classification makes the multi-classification results accurate.
[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0013] Figure 1 A schematic diagram showing an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure;
[0014] Figure 2 A flowchart showing a video coding method according to an embodiment of the present disclosure;
[0015] Figure 3 A flowchart showing the process of determining a current prediction unit of a current coding unit of a video frame to be coded in a video coding method according to an embodiment of the present disclosure;
[0016] Figure 4 A flowchart showing the process of determining a processing method for a current coding unit based on at least a first SATD cost model and a second SATD cost model according to the size of the current coding unit in a video coding method according to an embodiment of the present disclosure;
[0017] Figure 5 A flowchart showing the process of determining a processing method for a current coding unit based on at least a first SATD cost model and a second SATD cost model according to the size of the current coding unit in a video coding method according to an embodiment of the present disclosure;
[0018] Figure 6 A flowchart showing the process of determining a current prediction unit of a current coding unit of a video frame to be coded in a video coding method according to an embodiment of the present disclosure;
[0019] Figure 7 A flowchart showing the process of determining a current coding unit in a video coding method according to an embodiment of the present disclosure;
[0020] Figure 8 A flowchart showing the process of obtaining a target prediction mode corresponding to a current prediction unit from multiple prediction modes of a prediction method based on a SATD cost model corresponding to the prediction method in a video coding method according to an embodiment of the present disclosure;
[0021] Figure 9 A flowchart showing a process of obtaining a target prediction mode corresponding to a current prediction unit from multiple prediction modes of a prediction mode based on a SATD cost model corresponding to the prediction mode in a video encoding method according to an embodiment of the present disclosure
[0022] Figure 10 A flowchart showing a process of obtaining a target prediction mode during an inter-frame prediction process in a video encoding method according to an embodiment of the present disclosure;
[0023] Figure 11 A block diagram showing the structure of a video encoding device according to an embodiment of the present disclosure; and
[0024] Figure 12 A block diagram showing an exemplary electronic device capable of implementing an embodiment of the present disclosure. Detailed implementation manners
[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0027] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0028] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0029] Figure 1 A schematic diagram showing an exemplary system 100 in which various methods and devices described herein can be implemented according to an embodiment of the present disclosure. Refer to Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0030] In embodiments of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of video encoding methods.
[0031] In certain embodiments, the server 120 can also provide other services or software applications that can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, such as provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0032] In Figure 1 In the depicted configuration, the server 120 can include one or more components that implement the functions performed by the server 120. These components can include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 can in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which can be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0033] Users can use the client devices 101, 102, 103, 104, 105, and / or 106 to receive encoded video. The client devices can provide an interface that enables the users of the client devices to interact with the client devices. The client devices can also output information to the users via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure can support any number of client devices.
[0034] Client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices can include head-mounted displays (such as smart glasses) and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0035] Network 110 can be any type of network known to those skilled in the art, and it can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0036] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 can run one or more services or software applications that provide the functions described below.
[0037] The computing unit in server 120 can run one or more operating systems including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0038] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0039] In some embodiments, server 120 can be a server of a distributed system or a server incorporating blockchain. Server 120 can also be a cloud server or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to address the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0040] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and object files. The data repository 130 can reside in various locations. For example, the data repository used by server 120 can be local to server 120 or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. The data repository 130 can be of different types. In certain embodiments, the data repository used by server 120 can be a database, such as a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.
[0041] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0042] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and apparatuses described in this disclosure.
[0043] Referring to Figure 2 , a video encoding method 200 according to some embodiments of this disclosure includes:
[0044] Step S210: Determine the current prediction unit of the current coding unit of the video frame to be encoded;
[0045] Step S220: Based on a Hadamard transform absolute difference sum (SATD) cost model specific to the prediction mode, obtain the target prediction mode corresponding to the current prediction unit from multiple prediction modes of multiple prediction methods; and
[0046] Step S230: Encode the video frame to be encoded based on the target prediction mode.
[0047] Wherein, in step S220, the multiple prediction methods include intra prediction and inter prediction, the intra prediction corresponds to a first SATD cost model, the inter prediction corresponds to a second SATD cost model different from the first SATD cost model, and wherein, the SATD cost model specific to the prediction mode is used to obtain multiple SATD costs for predicting the current prediction unit using multiple prediction modes corresponding to the prediction mode, and each SATD cost among the multiple SATD costs includes a residual bit cost related to the prediction mode, and the residual bit cost indicates the number of bits required to encode the prediction residual obtained after predicting the current prediction unit.
[0048] According to one or more embodiments of this disclosure, by being able to estimate the SATD cost model under different prediction modes, obtain the SATD costs corresponding to different prediction modes of different prediction methods, and make the obtained SATD costs include the residual bit costs corresponding to the prediction modes, where the residual bit costs indicate the number of bits required to encode the prediction residual obtained after predicting the current prediction unit. Make the obtained SATD costs closer to the rate-distortion costs obtained using the Rate-Distortion Optimization (RDO) technique. Therefore, estimate the rate-distortion costs under different prediction modes based on the SATD costs obtained by this solution, make the estimated rate-distortion costs more accurate, so that the target mode selected for rate-distortion optimization based on the SATD costs can be more accurate, reduce the rate-distortion optimization process, reduce the encoding complexity, and improve the encoding efficiency.
[0049] In related technologies, a relevant SATD cost that does not include the residual bit cost is used as an estimated value for estimating the rate-distortion cost, so as to make a rough selection from multiple prediction modes (including multiple prediction modes of intra prediction and multiple prediction modes of inter prediction), obtain multiple modes with a smaller relevant SATD cost, perform rate-distortion optimization, obtain a prediction mode with the minimum rate-distortion cost, and use it as the prediction mode for subsequent coding.
[0050] Among them, the relevant SATD cost is calculated using the following estimation model (1):
[0051] (1)
[0052] is the Lagrange multiplier used in the SATD cost. is the number of bits required for mode coding, and refers to the value obtained by performing Hadamard transform on the residual signal and then summing the absolute values of the coefficients, representing the amplitude of the residual in the frequency domain. Assume that a certain residual signal is , where is calculated using the following formula (2):
[0053] (2)
[0054] Among them, M is the size of the transform unit TU, is the normalized Hadamard matrix of size.
[0055] Although the relevant SATD cost calculated using the above estimation model (1) can replace a part of the rate-distortion cost, there is still a certain gap with the rate-distortion cost, and it does not include the bits required for encoding the prediction residual.
[0056] According to some embodiments of the present disclosure, a residual bit cost corresponding to the prediction mode is adopted, and the residual bit cost indicates the number of bits required for encoding the prediction residual obtained after predicting the current prediction unit. The obtained SATD cost is closer to the rate-distortion cost obtained using the Rate Distortion Optimization (RDO) technology. Therefore, based on the SATD cost obtained by this solution, the rate-distortion costs in different prediction modes are estimated, making the estimated rate-distortion cost more accurate, so that the target mode selected for rate-distortion optimization based on the SATD cost can be more accurate, reducing the rate-distortion optimization process, reducing the coding complexity, and improving the coding efficiency.
[0057] In some embodiments, such as Figure 3As shown in the figure, determining the current prediction unit of the current coding unit of the video frame to be coded and the target prediction method corresponding to the current coding unit includes:
[0058] Step S310: Obtain the current coding unit; and
[0059] Step S320: Based on the size of the current coding unit, at least utilize the first SATD cost model and the second SATD cost model to determine the processing method for the current coding unit, where the processing method is one of multiple processes including at least intra prediction processing and inter prediction processing; and
[0060] Step S330: Based on the processing method, determine the current prediction unit.
[0061] By using the SATD cost models corresponding to different prediction methods, the SATD cost when predicting the coding unit using different prediction methods can be estimated to determine the prediction method of the coding unit, making the method for obtaining the prediction method simpler and the coding more efficient and accurate.
[0062] In some embodiments, the current coding unit includes a coding tree unit CTU or a coding unit CU divided based on the coding tree unit.
[0063] In some embodiments, the size of the current coding unit is one of 64×64, 32×32, 16×16, and 8×8.
[0064] Among them, when the size of the current coding unit is 64×64, 32×32, or 16×16, it is necessary to determine the division depth of the current coding unit to determine whether it needs to be divided. When the size of the current coding unit is 8×8, since it is the size of the smallest coding unit, it is not divided, and inter prediction or intra prediction is performed on the multiple prediction units obtained therefrom.
[0065] In some embodiments, when it is determined that the processing method is one of the two processing methods including intra prediction processing and inter prediction processing, the current coding unit is not divided, and prediction is performed based on the prediction units divided from the current coding unit.
[0066] In some embodiments, when it is determined that the processing method is the division processing, the current coding unit needs to be divided
[0067] In some embodiments, the processing method is one of including intra prediction processing, inter prediction processing, and division processing.
[0068] In some embodiments, such as Figure 4As shown, based on the size of the current coding unit, determining the processing method for the current coding unit by at least using the first SATD cost model and the second SATD cost model includes:
[0069] Step S410: In response to the size of the current coding unit being equal to the size of the smallest coding unit, obtain a first SATD cost set by using the first SATD cost model and obtain a second SATD cost set by using the second SATD cost model;
[0070] Step S420: According to the ascending order of the elements in the first SATD cost set and the second SATD cost set, obtain a first preset number of third SATD costs; and
[0071] Step S430: Determine the processing method based on the first preset number of third SATD costs.
[0072] Among them, in step S410, a first SATD cost set is obtained by using the first SATD cost model and adopting multiple prediction modes of intra prediction, where the current coding unit is used as a prediction unit, and a second SATD cost set is obtained by using the second SATD cost model and adopting multiple prediction modes of inter prediction, where the current coding unit is used as a prediction unit.
[0073] When the size of the current coding unit is equal to the size of the smallest coding unit, directly estimate the prediction method of the current coding unit through the SATD cost models corresponding to intra prediction and inter prediction, further reduce the rate-distortion optimization process, reduce the coding complexity, and improve the coding efficiency.
[0074] In some embodiments, the first SATD cost and the second SATD cost can be obtained based on the following estimation model (3):
[0075] (3)
[0076] Among them, is the Lagrange multiplier used in the SATD cost; is the number of bits required for mode coding; SATD refers to the value obtained by performing a Hadamard transform on the residual signal and then summing the absolute values of each coefficient, representing the amplitude of the residual in the frequency domain, and is calculated using the above formula (2); is an estimated value of the number of bits required for coding the prediction residual obtained after predicting the current prediction unit, which is related to the prediction method; in response to being different, the estimation model (3) is transformed into SATD cost models corresponding to different prediction methods to obtain the corresponding SATD costs.
[0077] By comparing the estimation model (3) with the estimation model (1), it can be seen that according to the technical solution of the present disclosure, in the estimation model (3) for obtaining the SATD cost model according to the present disclosure, compared with the estimation model (1) in the related art, it estimates the bits required for encoding the prediction residual, so that the obtained SATD cost ( ), compared with the related SATD cost ( ) obtained by using the estimation model (1) in the related art, is closer to the rate-distortion cost, improves the accuracy of the estimated SATD cost, thus better simulating the rate-distortion cost, reducing the number of target modes that need to perform rate-distortion optimization, and reducing the coding complexity.
[0078] In some embodiments, it is obtained by using the following formula (4):
[0079] (4)
[0080] wherein, QP is a quantization parameter, which can be preset for each video frame to be encoded; SATD refers to the value obtained by performing Hadamard transform on the residual signal and then summing the absolute values of each coefficient, representing the amplitude of the residual in the frequency domain, and is obtained by using the above-mentioned formula (2); k is a constant, corresponding to different prediction methods, and the value of k is different.
[0081] In some embodiments according to the present disclosure, by taking different values of k, SATD cost models corresponding to different prediction methods are obtained.
[0082] In some embodiments, as Figure 5 shown, based on the size of the current coding unit, determining the processing method for the current coding unit by using at least the first SATD cost model and the second SATD cost model includes:
[0083] Step S510: In response to determining that the size of the current coding unit is greater than the size of the smallest coding unit, obtaining a fourth SATD cost set by using the first SATD cost model, obtaining a fifth SATD cost set by using the second SATD cost model, and obtaining a partitioned SATD cost set by using the SATD cost;
[0084] Step S520: Obtaining a second preset number of sixth SATD costs according to the ascending order of the elements in the fourth SATD cost set, the fifth SATD cost set, and the partitioned SATD cost set; and
[0085] Step S530: Determining the processing method based on the second preset number of sixth SATD costs.
[0086] Among them, in step S510, a fourth SATD cost set is obtained by using the first SATD cost model and adopting multiple prediction modes of intra prediction, where the current coding unit is used as the prediction unit; a fifth SATD cost set is obtained by using the second SATD cost model and adopting multiple prediction modes of inter prediction, where the current coding unit is used as the prediction unit; and a partition SATD cost set is obtained by using the partition SATD cost model and adopting multiple prediction modes of intra prediction and multiple prediction modes of inter prediction, where four sub-units obtained by partitioning the current coding unit are used as four prediction units, and each element in the partition SATD cost set includes a residual bit cost based on partition prediction, and the residual bit cost based on partition prediction indicates the number of bits required to encode the prediction residual obtained by prediction after partitioning the current coding unit.
[0087] When the size of the current coding unit is larger than the size of the minimum coding unit, directly estimate whether the current coding unit needs to be further partitioned through the SATD cost model and the partition estimation model corresponding to intra prediction and inter prediction, further reduce the rate-distortion optimization process, reduce the coding complexity, and improve the coding efficiency.
[0088] In some embodiments according to the present disclosure, the partition SATD cost is obtained by using the partition SATD cost model shown as follows (5):
[0089] (5),
[0090] Among them, is the total SATD after partitioning the current coding unit into four sub-units, and is obtained through the following formula (6),
[0091] (6),
[0092] Among them, , , and are the SATDs of the four sub-units respectively, and are obtained through the formula (2) as described above; is the Lagrange multiplier used in the SATD cost; is the number of bits required for mode coding; and among them, is the residual bit cost based on partition prediction, and is obtained through the following formula (7):
[0093] (7),
[0094] Wherein, QP is a quantization parameter, which can be preset for each video frame to be encoded; is the sum of SATD after dividing the current coding unit into four sub-units, representing the amplitude in the frequency domain of the residual based on the division, and is obtained by using the formula (6) as described above; is a constant corresponding to the division process.
[0095] In some embodiments, as Figure 6 shown, determining the current prediction unit of the current coding unit of the video frame to be encoded further includes:
[0096] Step S610: In response to determining that the processing mode is one of intra-frame prediction processing and inter-frame prediction processing, determine the prediction mode of the current coding unit, wherein the prediction mode corresponds to the determined processing mode; and
[0097] Step S620: Determine the sub-unit corresponding to the prediction mode obtained by dividing the current coding unit as the current prediction unit.
[0098] When it is determined that the processing mode is inter-frame prediction processing or intra-frame prediction processing, the current coding unit is not further divided, reducing the coding complexity.
[0099] For example, for a 64×64 coding unit, when the above process is used to determine that the establishment mode is intra-frame prediction processing, it is not further divided and directly undergoes intra-frame prediction processing, wherein the current prediction unit is determined by obtaining the prediction unit corresponding to the intra-frame processing.
[0100] In some embodiments, referring to Figure 7 , the method 200 according to the present disclosure may further include:
[0101] Step S710: In response to the processing mode being the division process, divide the current coding unit into four sub-units; and
[0102] Step S720: Determine one of the four sub-units as the updated current coding unit, and re-execute the step of determining the current prediction unit of the current coding unit of the video frame to be encoded.
[0103] When it is necessary to perform a division process on the current coding unit, the sub-unit of the current coding unit is determined as the current coding unit to perform the step of determining the current prediction unit as introduced above with reference to Figures 1 - 6 until it is determined that the current coding unit does not undergo a division process.
[0104] According to some embodiments of the present disclosure, the step of determining the prediction unit of the current coding unit is a loop-recursive process. By continuously performing the loop-recursive execution of the step of determining the prediction unit of the current coding unit until the size of the current coding unit is equal to the size of the smallest coding unit (8×8), it is determined to perform further intra prediction or inter prediction on the current coding unit with the size equal to the size of the smallest coding unit (8×8). Among them, according to whether the prediction method is intra prediction or inter prediction, the division of the prediction unit of the current coding unit is determined respectively.
[0105] For example, for a current coding unit of 64×64, since its size is larger than that of the smallest coding unit (8×8), the process such as Figure 3 , Figure 5 , Figure 6 and Figure 7 is performed to execute the step of determining the prediction unit of the current coding unit; when it is determined to perform division processing on it, it is divided into 4 sub-units of 32×32, and one of the 4 sub-units of 32×32 is used as the current coding unit, and the process such as Figure 3 , Figure 5 , Figure 6 and Figure 7 is performed to execute the step of determining the prediction unit of the current coding unit. For a current coding unit of 32×32, in the step of determining the prediction unit of the current coding unit, when it is determined to perform division processing on it, it is divided into 4 sub-units of 16×16, and one of the 4 sub-units of 16×16 is used as the current coding unit, and the process such as Figure 3 , Figure 5 , Figure 6 and Figure 7 is performed to execute the step of determining the prediction unit of the current coding unit. For a current coding unit of 16×16, in the step of determining the prediction unit of the current coding unit, when it is determined to perform division processing on it, it is divided into 4 sub-units of 8×8, and one of the 4 sub-units of 8×8 is used as the current coding unit. Since the sub-unit of 8×8 is the current coding unit with the smallest size, the process such as Figure 3 , Figure 4 , Figure 6 and Figure 7 is performed to execute the step of determining the prediction unit of the current coding unit.
[0106] According to some embodiments of the present disclosure, the step of determining the prediction unit of the current coding unit is a loop-recursive process. By continuously performing the loop-recursive execution of the step of determining the prediction unit of the current coding unit until it is determined that the current coding unit (the size may be larger than 8×8) does not perform division processing, it is determined to perform further intra prediction or inter prediction on the current coding unit (the size may be larger than 8×8).
[0107] When it is determined that the current coding unit is not to be partitioned, that is, intra prediction or inter prediction is performed on the current coding unit, according to the prediction method, the current prediction unit of the current coding unit is determined.
[0108] For example, in response to the prediction method being intra prediction, the current coding unit is partitioned into 2N×2N or N×N to obtain prediction units. In response to the prediction method being inter prediction, the current coding unit is partitioned into various partitioning forms including 2N×2N partitioning, 2N×N partitioning, N×2N partitioning, N×N partitioning, 2N×nU partitioning, 2N×nD partitioning, nL×2N partitioning, and nR×2N partitioning to obtain prediction units.
[0109] The following further introduces the process of obtaining the target prediction mode corresponding to the current prediction unit from multiple prediction modes of the prediction method for the current prediction unit through a (SATD) cost model corresponding to the prediction method.
[0110] It should be noted that the prediction processes for intra prediction and inter prediction are different. For intra prediction, the current coding unit is partitioned into 2N×2N or N×N. For each prediction unit of each partition, by comparing the SATD costs of 35 prediction modes in intra prediction, the target prediction mode of intra prediction is obtained for the rate-distortion optimization process. For inter prediction, the current coding unit is partitioned into various partitioning forms including 2N×2N partitioning, 2N×N partitioning, N×2N partitioning, N×N partitioning, 2N×nU partitioning, 2N×nD partitioning, nL×2N partitioning, and nR×2N partitioning. For each inter prediction mode obtained from each partition, the corresponding prediction mode is used for prediction and it is determined whether rate-distortion optimization is required. Among them, during the inter prediction process, it is further determined whether the rate-distortion optimization process of intra prediction is required. The following separately introduces inter prediction and intra prediction.
[0111] In some embodiments, in response to the prediction method being the intra prediction, as Figure 8 shown, obtaining the target prediction mode corresponding to the current prediction unit from multiple prediction modes of the prediction method based on the Hadamard transform absolute difference sum (SATD) cost model corresponding to the prediction method includes:
[0112] Step S810: Use the relevant SATD cost model to obtain the first prediction mode corresponding to the current prediction unit among multiple prediction modes of the intra prediction;
[0113] Step S820: Obtain a second prediction mode corresponding to the current prediction unit among multiple prediction modes for intra-frame prediction based on the first SATD cost model;
[0114] Step S830: In response to the first prediction mode and the second prediction mode being the same prediction mode among multiple prediction modes for intra-frame prediction, determine the same prediction mode as the target prediction mode, and
[0115] Step S840: In response to the first prediction mode and the second prediction mode being two different prediction modes among multiple prediction modes for intra-frame prediction, determine both the first prediction mode and the second prediction mode as target prediction modes.
[0116] Among them, in step S810, the relevant SATD cost model is used to obtain multiple relevant SATD costs for predicting the current prediction unit using multiple prediction modes for intra-frame prediction. Each relevant SATD cost among the multiple relevant SATD costs does not include the residual bit cost related to the prediction method. And among the multiple relevant SATD costs, the relevant SATD cost corresponding to the first prediction mode is less than the relevant SATD cost corresponding to any prediction mode different from the first prediction mode. In step S820, among the multiple SATD costs obtained based on the first SATD cost model, the SATD cost corresponding to the second prediction mode is less than the SATD cost corresponding to any prediction mode different from the second prediction mode
[0117] The first prediction mode is obtained from multiple relevant SATD costs that do not include the residual bit cost related to the prediction method based on the relevant SATD cost model. The first prediction mode is the prediction mode that is most likely to be the target prediction mode based on the relevant SATD cost model. Similarly, the second prediction mode is obtained based on the first SATD cost model, which includes the residual bit cost related to the prediction method. The second prediction mode is the prediction mode that is most likely to be the target prediction mode based on the first SATD cost model. Compare the prediction modes that are most likely to be the target prediction mode obtained by the two methods. When the prediction modes that are most likely to be the target prediction mode obtained by the two methods are the same mode, it indicates that the probability of this prediction mode that is most likely to be the target prediction mode being the target prediction mode is very high. Based on it, subsequent rate-distortion optimization is performed with high accuracy. When the prediction modes that are most likely to be the target prediction mode obtained by the two methods are not the same mode, both the first prediction mode and the second prediction mode are determined as target prediction modes to perform the rate-distortion optimization process, thereby obtaining coding parameters.
[0118] It can be seen that according to the encoding method of the present disclosure, for intra prediction, the optimal prediction mode among the 35 prediction modes in the intra prediction process of the prediction unit can be obtained by only performing rate-distortion optimization through 1 to 2 prediction modes, greatly reducing the process of rate-distortion optimization.
[0119] In step S810, the estimation mode (1) is used as the relevant SATD cost model to obtain the first prediction mode corresponding to the current prediction unit among the multiple prediction modes of the intra prediction.
[0120] In some embodiments, in response to the prediction mode being the inter prediction, as Figure 9 shown, based on the Hadamard transform absolute difference sum (SATD) cost model corresponding to the prediction mode, obtaining the target prediction mode corresponding to the current prediction unit from the multiple prediction modes of the prediction mode includes:
[0121] Step S910: Use the second SATD cost model to sequentially traverse the multiple prediction modes of the current coding unit in the inter prediction to determine one or more inter target prediction modes among the multiple prediction modes of the inter prediction;
[0122] Step S920: Obtain the intra target prediction mode among the multiple prediction modes of the intra prediction when the prediction mode is the intra prediction; and
[0123] Step S930: Based on the one or more inter target prediction modes and the intra target prediction mode, obtain the target prediction mode.
[0124] Wherein, in step S910, for each inter target prediction mode among the one or more inter target prediction modes, the SATD cost corresponding to the inter target prediction mode is smaller than the SATD cost corresponding to the prediction mode traversed before the target prediction mode among the multiple prediction modes of the inter prediction; in step S930, among the multiple SATD costs obtained by using the first SATD cost model, the SATD cost corresponding to the intra target prediction mode is smaller than the SATD cost corresponding to any prediction mode different from the intra target prediction mode among the multiple prediction modes of the intra prediction.
[0125] During the inter prediction process, since the inter prediction process is a different prediction process from the intra prediction process, for different divisions of prediction units, corresponding inter prediction modes (such as skip mode, merge mode, or AMVP mode) are used for corresponding prediction, and the target prediction mode that needs to perform rate-distortion optimization is obtained by sequentially traversing each inter mode.
[0126] See Figure 10, an exemplary introduction to the process of the inter-frame prediction mode according to an example is provided. As Figure 10 shown, the process of sequentially traversing multiple prediction modes of the current coding unit in the inter-frame prediction using the second SATD cost model includes:
[0127] First, in step1: Use the second SATD cost model to traverse the inter-frame 2N×2N mode of the prediction unit obtained by dividing the current coding unit by 2N×2N, and obtain the SATD cost of the inter-frame 2N×2N mode J SATD 0, and use it as the minimum SATD cost J SATD min ;
[0128] Next, in step2: Use the second SATD cost model to traverse the inter-frame N×2N mode of the prediction unit obtained by dividing the current coding unit by N×2N, and obtain the SATD cost of the inter-frame N×2N mode J SATD 1, compare the SATD cost of this inter-frame N×2N mode with the minimum SATD cost (the SATD cost of the inter-frame 2N×2N mode):
[0129] If the SATD cost of this inter-frame N×2N mode J SATD 1 is less than the minimum SATD cost J SATD min , then perform rate-distortion optimization on this inter-frame N×2N mode, obtain the rate-distortion cost of the inter-frame N×2N mode, and compare the rate-distortion cost of this inter-frame N×2N mode with the minimum SATD cost, and use the smaller one of the two as the new minimum SATD cost J SATD min ,
[0130] If the SATD cost of this inter-frame N×2N mode J SATD 1 is greater than the minimum SATD cost J SATD min , then skip the rate-distortion optimization and continue to use the SATD cost of the inter-frame 2N×2N mode J SATD 0 as the minimum SATD cost J SATD min ;
[0131] Next, in step 3: Use the second SATD cost model to traverse the inter-frame 2N×N mode of the prediction unit divided according to 2N×N of the current coding unit to obtain the SATD cost of the inter-frame 2N×N mode. J SATD 2. Compare the SATD cost of this inter-frame 2N×N mode J SATD 2 with the minimum SATD cost determined in the aforementioned step 2. J SATD min :
[0132] If the SATD cost of this inter-frame 2N×N mode J SATD 2 is less than the minimum SATD cost J SATD min then perform rate-distortion optimization on this inter-frame 2N×N mode to obtain the rate-distortion cost of the inter-frame 2N×N mode, and compare the rate-distortion cost of this inter-frame 2N×N mode with the minimum SATD cost, and take the smaller one of the two as the new minimum SATD cost. J SATD min ,
[0133] If the SATD cost of this inter-frame 2N×N mode J SATD 2 is greater than the minimum SATD cost, skip the rate-distortion optimization, and continue to determine the minimum SATD cost in the aforementioned step 2 as the minimum SATD cost. J SATD min ;
[0134] Next, in step 4: Use the second SATD cost model to traverse the 2N×nU mode divided according to the AMVP mode prediction of the current coding unit... and so on until the traversal of various modes of inter-frame prediction is completed.
[0135] Finally, in step x: Obtain the intra-frame target prediction mode of intra-frame prediction, and the SATD cost of this target prediction mode is smaller than that of any prediction mode different from the intra-frame target prediction mode among the multiple prediction modes of intra-frame prediction. Compare the minimum SATD cost J SATD min obtained in the aforementioned inter-frame prediction process with the SATD cost J SATD x:
[0136] If the SATD cost of the intra-frame target prediction mode of this intra-frame prediction J SATDx is less than the minimum SATD cost J SATD min , the rate - distortion optimization process of intra - prediction is performed; if the SATD cost of the intra - target prediction mode of intra - prediction J SATD x is greater than the minimum SATD cost J SATD min , the rate - distortion optimization process of this intra - prediction is skipped, and the intra - rate - distortion optimization process of this coding unit is ended.
[0137] Obviously, from the above examples, it can be concluded that according to the method of the present disclosure, the rate - distortion optimization processes of some inter - predictions and intra - predictions can be effectively skipped, thereby simplifying the variable coding method and improving the coding efficiency.
[0138] In some embodiments, in the above - mentioned Figure 10 stepx, the method for obtaining the intra - target prediction mode of intra - prediction is obtained by the process of steps S810 - S840 as shown in Figure 8 , which will not be elaborated here.
[0139] In some embodiments, encoding the video frame to be encoded based on the target prediction mode includes: obtaining the target prediction mode with a smaller rate - distortion cost through rate - distortion optimization based on the target prediction mode, and obtaining the target coding parameters obtained when performing rate - distortion optimization for this target prediction mode, and encoding the video frame to be encoded based on the target coding parameters.
[0140] According to another aspect of the present disclosure, a video coding device is further provided, as shown in Figure 11As shown, a video encoding device 1100 includes: a determination unit 1110 configured to determine a current prediction unit of a current coding unit of a video frame to be encoded; and an acquisition unit 1120 configured to obtain, based on a Hadamard transform absolute difference sum (SATD) cost model corresponding to a prediction mode, a target prediction mode corresponding to the current prediction unit from multiple prediction modes of the prediction mode, so as to encode the video frame to be encoded, where the prediction mode is any one of intra prediction and inter prediction, the intra prediction corresponds to a first SATD cost model, the inter prediction corresponds to a second SATD cost model different from the first SATD cost model, and where the SATD cost model corresponding to the prediction mode is used to obtain multiple SATD costs for predicting the current prediction unit using multiple prediction modes corresponding to the prediction mode, and each SATD cost in the multiple SATD costs includes a residual bit cost related to the prediction mode, and the residual bit cost related to the prediction mode indicates the number of bits required to encode a prediction residual obtained after predicting the current prediction unit.
[0141] In some embodiments, the determination unit 1110 includes: a coding unit acquisition unit configured to acquire the current coding unit; a processing mode determination unit configured to determine, based on the size of the current coding unit, at least using the first SATD cost model and the second SATD cost model, a processing mode for the current coding unit, where the processing mode is one of multiple processes including at least intra prediction processing and inter prediction processing; and a prediction unit determination unit configured to determine the current prediction unit based on the processing mode.
[0142] In some embodiments, the processing mode determination unit includes: a first calculation unit configured to, in response to the size of the current coding unit being equal to the size of the smallest coding unit, use the first SATD cost model and obtain a first set of SATD costs using multiple prediction modes of the intra prediction, where the current coding unit is used as a prediction unit, use the second SATD cost model and obtain a second set of SATD costs using multiple prediction modes of the inter prediction, where the current coding unit is used as a prediction unit; a first acquisition unit configured to obtain a first preset number of third SATD costs in ascending order of the elements in the first set of SATD costs and the second set of SATD costs; and a first determination subunit configured to determine the processing mode based on the first preset number of third SATD costs.
[0143] In some embodiments, the processing method determining unit includes: a second calculation unit configured to, in response to determining that the size of the current coding unit is greater than the size of the minimum coding unit, obtain a fourth SATD cost set by using the first SATD cost model and adopting a plurality of prediction modes of intra prediction, wherein the current coding unit is used as a prediction unit, obtain a fifth SATD cost set by using the second SATD cost model and adopting a plurality of prediction modes of inter prediction, wherein the current coding unit is used as a prediction unit, and obtain a plurality of partition SATD costs to obtain a partition SATD cost set by using a partition SATD cost model and adopting a plurality of prediction modes of intra prediction and a plurality of prediction modes of inter prediction, wherein when four sub-units obtained by partitioning the current coding unit are used as four prediction units, and wherein each element in the partition SATD cost set includes a residual bit cost based on partition prediction, and the residual bit cost based on partition prediction indicates the number of bits required for a prediction residual obtained by performing prediction after partitioning the current coding unit; and a second obtaining unit configured to obtain a second preset number of sixth SATD costs according to the ascending order of the elements in the fourth SATD cost set, the fifth SATD cost set, and the partition SATD cost set; and a second determining subunit configured to determine the processing method based on the second preset number of sixth SATD costs, wherein the processing method is one of a plurality of processes including a partitioning process, an intra prediction process, and an inter prediction process.
[0144] In some embodiments, the determining unit 1110 further includes: a prediction method determining unit configured to, in response to determining that the processing method is one of an intra prediction process and an inter prediction process, determine the prediction method of the current coding unit, wherein the prediction method corresponds to the determined processing method; and a prediction unit determining unit configured to determine the sub-unit corresponding to the prediction method obtained by partitioning the current coding unit as the current prediction unit.
[0145] In some embodiments, the apparatus 1100 further includes: a partitioning unit configured to, in response to the processing method being the partitioning process, partition the current coding unit into four sub-units; and a re-determining unit configured to determine one of the four sub-units as the updated current coding unit and re-determine the current prediction unit of the current coding unit.
[0146] In some embodiments, the obtaining unit 1120 includes: a related SATD cost obtaining unit configured to, in response to the prediction method being intra prediction, obtain a first prediction mode corresponding to the current prediction unit among a plurality of prediction modes of the intra prediction by using a related SATD cost model, where the related SATD cost model is used to obtain a plurality of related SATD costs for predicting the current prediction unit by using a plurality of prediction modes of the intra prediction, each of the plurality of related SATD costs does not include the residual bit cost related to the prediction method, and where the related SATD cost corresponding to the first prediction mode among the plurality of related SATD costs is less than the related SATD cost corresponding to any prediction mode different from the first prediction mode; a first SATD cost model calculation unit configured to obtain a second prediction mode corresponding to the current prediction unit among a plurality of prediction modes of the intra prediction by using the first SATD cost model, where, among the plurality of SATD costs obtained by using the first SATD cost model, the SATD cost corresponding to the second prediction mode is less than the SATD cost corresponding to any prediction mode different from the second prediction mode; a third determination subunit configured to, in response to the first prediction mode and the second prediction mode being the same prediction mode among the plurality of prediction modes of the intra prediction, determine the same prediction mode as the target prediction mode; and a fourth determination subunit configured to, in response to the first prediction mode and the second prediction mode being two different prediction modes among the plurality of prediction modes of the intra prediction, determine both the first prediction mode and the second prediction mode as target prediction modes.
[0147] In some embodiments, the obtaining unit includes: a second SATD cost model calculation unit configured to, in response to the prediction mode being the inter prediction, sequentially traverse a plurality of prediction modes of the current coding unit in the inter prediction by using the second SATD cost model to determine one or more inter target prediction modes among the plurality of prediction modes of the inter prediction, wherein, for each of the one or more inter target prediction modes, the SATD cost corresponding to the inter target prediction mode is smaller than the SATD cost corresponding to a prediction mode traversed before the target prediction mode among the plurality of prediction modes of the inter prediction; an intra target prediction mode obtaining unit configured to obtain an intra target prediction mode among the plurality of prediction modes of the intra prediction when the prediction mode is the intra prediction, wherein, among the plurality of SATD costs obtained by using the first SATD cost model, the SATD cost corresponding to the intra target prediction mode is smaller than the SATD cost corresponding to any prediction mode different from the intra target prediction mode among the plurality of prediction modes of the intra prediction; and a fifth determination subunit configured to obtain the target prediction mode based on the one or more inter target prediction modes and the intra target prediction mode.
[0148] According to another aspect of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program, and the computer program, when executed by the at least one processor, implements the method according to the above.
[0149] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to the above.
[0150] According to another aspect of the present disclosure, there is also provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method according to the above.
[0151] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.
[0152] Reference Figure 10, a block diagram of an electronic device 1000 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0153] As Figure 12 shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the electronic device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0154] A plurality of components in the electronic device 1200 are connected to the I / O interface 1205, including: an input unit 1206, an output unit 1207, a storage unit 1208, and a communication unit 1209. The input unit 1206 can be any type of device that can input information into the electronic device 1200. The input unit 1206 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 1207 can be any type of device that can present information, and can include but are not limited to a display, a speaker, an object / audio output terminal, a vibrator, and / or a printer. The storage unit 1208 can include but are not limited to magnetic disks, optical disks. The communication unit 1209 allows the electronic device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0155] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the method 200 described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute method 200 in any other suitable manner (e.g., by means of firmware).
[0156] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0157] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0158] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0160] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0161] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.
[0162] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0163] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. A video encoding method, comprising: Determining a current prediction unit of a current coding unit of a video frame to be encoded; Obtaining a target prediction mode corresponding to the current prediction unit from multiple prediction modes of multiple prediction methods based on a Hadamard transform absolute difference and SATD cost model specific to a prediction method, wherein, The multiple prediction methods include intra prediction and inter prediction, the intra prediction corresponds to a first SATD cost model, the inter prediction corresponds to a second SATD cost model different from the first SATD cost model, and wherein, The SATD cost model specific to the prediction method is used to obtain multiple SATD costs for predicting the current prediction unit using multiple prediction modes corresponding to the prediction method, and each SATD cost among the multiple SATD costs includes a residual bit cost related to the prediction method, and the residual bit cost indicates the number of bits required to encode the prediction residual obtained after predicting the current prediction unit; and Encoding the video frame to be encoded based on the target prediction mode; Wherein, obtaining the target prediction mode corresponding to the current prediction unit from multiple prediction modes of the prediction method based on the Hadamard transform absolute difference and SATD cost model corresponding to the prediction method includes: In response to the prediction method being the intra prediction, obtaining a first prediction mode corresponding to the current prediction unit among multiple prediction modes of the intra prediction using a relevant SATD cost model, wherein the relevant SATD cost model is used to obtain multiple relevant SATD costs for predicting the current prediction unit using multiple prediction modes of the intra prediction, and each relevant SATD cost among the multiple relevant SATD costs does not include a residual bit cost related to the prediction method, and wherein the relevant SATD cost corresponding to the first prediction mode among the multiple relevant SATD costs is less than the relevant SATD cost corresponding to any prediction mode different from the first prediction mode; Obtaining a second prediction mode corresponding to the current prediction unit among multiple prediction modes of the intra prediction using the first SATD cost model, wherein, among the multiple SATD costs obtained using the first SATD cost model, the SATD cost corresponding to the second prediction mode is less than the SATD cost corresponding to any prediction mode different from the second prediction mode; In response to the first prediction mode and the second prediction mode being the same prediction mode among multiple prediction modes of the intra prediction, determining the same prediction mode as the target prediction mode; and In response to the first prediction mode and the second prediction mode being two different prediction modes among multiple prediction modes of the intra prediction, determining both the first prediction mode and the second prediction mode as target prediction modes.
2. The method according to claim 1, wherein Determining a current prediction unit of a current coding unit of a video frame to be encoded includes: Obtaining the current coding unit; Based on the size of the current coding unit, determine a processing method for the current coding unit by using at least the first SATD cost model and the second SATD cost model, where the processing method is one of multiple processes including at least intra prediction processing and inter prediction processing; and Based on the processing method, determine the current prediction unit.
3. The method according to claim 2, wherein Determining a processing method for the current coding unit based on the size of the current coding unit by using at least the first SATD cost model and the second SATD cost model includes: In response to the size of the current coding unit being equal to the size of the smallest coding unit, Use the first SATD cost model and adopt multiple prediction modes of intra prediction to obtain a first set of SATD costs, where the current coding unit serves as a prediction unit, Use the second SATD cost model and adopt multiple prediction modes of inter prediction to obtain a second set of SATD costs, where the current coding unit serves as a prediction unit; According to the ascending order of the elements in the first set of SATD costs and the second set of SATD costs, obtain a first preset number of third SATD costs; and Based on the first preset number of third SATD costs, determine the processing method.
4. The method according to claim 2, wherein, Determining a processing method for the current coding unit based on the size of the current coding unit by using at least the first SATD cost model and the second SATD cost model includes: In response to determining that the size of the current coding unit is greater than the size of the smallest coding unit, Use the first SATD cost model and adopt multiple prediction modes of intra prediction to obtain a fourth set of SATD costs, where the current coding unit serves as a prediction unit, Use the second SATD cost model and adopt multiple prediction modes of inter prediction to obtain a fifth set of SATD costs, where the current coding unit serves as a prediction unit, and Use a partition SATD cost model and adopt multiple prediction modes of intra prediction and multiple prediction modes of inter prediction to obtain a partition SATD cost set, where four sub-units obtained by partitioning the current coding unit serve as four prediction units, and where each element in the partition SATD cost set includes a residual bit cost based on partition prediction, and the residual bit cost based on partition prediction indicates the number of bits required to encode the prediction residual obtained by performing prediction after partitioning the current coding unit; and According to the ascending order of the elements in the fourth set of SATD costs, the fifth set of SATD costs, and the partition SATD cost set, obtain a second preset number of sixth SATD costs; and Based on the second preset number of sixth SATD costs, determine the processing method, where the processing method is one of multiple processes including partition processing, intra prediction processing, and inter prediction processing.
5. The method according to claim 2, wherein, The determining of the current prediction unit of the current coding unit of the video frame to be encoded further includes: In response to determining that the processing method is one of intra prediction processing and inter prediction processing, determine the prediction method of the current coding unit, where the prediction method corresponds to the determined processing method; and Determine the sub-unit corresponding to the prediction method obtained by dividing the current coding unit as the current prediction unit.
6. The method according to claim 4, further comprising: In response to the processing method being the division processing, divide the current coding unit into four sub-units; And Determine one of the four sub-units as the updated current coding unit, and re-determine the current prediction unit of the current coding unit.
7. The method according to claim 1, wherein The obtaining, from multiple prediction modes of the prediction method, the target prediction mode corresponding to the current prediction unit based on the Hadamard transform absolute difference and SATD cost model corresponding to the prediction method includes: In response to the prediction method being the inter prediction, use the second SATD cost model to sequentially traverse multiple prediction modes of the current coding unit in the inter prediction to determine one or more inter prediction target prediction modes among the multiple prediction modes of the inter prediction, where for each inter prediction target prediction mode among the one or more inter prediction target prediction modes, the SATD cost corresponding to this inter prediction target prediction mode is smaller than the SATD cost corresponding to the prediction mode traversed before this target prediction mode among the multiple prediction modes of the inter prediction; Obtain the intra prediction target prediction mode among the multiple prediction modes of the intra prediction when the prediction method is the intra prediction, where among the multiple SATD costs obtained by using the first SATD cost model, the SATD cost corresponding to the intra prediction target prediction mode is smaller than the SATD cost corresponding to any prediction mode different from the intra prediction target prediction mode among the multiple prediction modes of the intra prediction; and Based on the one or more inter prediction target prediction modes and the intra prediction target prediction mode, obtain the target prediction mode.
8. A video coding device, comprising: A determination unit configured to determine the current prediction unit of the current coding unit of the video frame to be coded; An obtaining unit configured to obtain, based on the Hadamard transform absolute difference and SATD cost model specific to the prediction method, the target prediction mode corresponding to the current prediction unit from multiple prediction modes of multiple prediction methods, where The multiple prediction methods include intra prediction and inter prediction, the intra prediction corresponds to a first SATD cost model, the inter prediction corresponds to a second SATD cost model different from the first SATD cost model, and where The SATD cost model specific to the prediction method is used to obtain multiple SATD costs for predicting the current prediction unit by using multiple prediction modes corresponding to the prediction method, and each SATD cost among the multiple SATD costs includes a residual bit cost related to the prediction method, and the residual bit cost indicates the number of bits required to encode the prediction residual obtained after predicting the current prediction unit; and A video frame encoding unit, configured to encode the video frame to be encoded based on the target prediction mode; Wherein, the obtaining unit includes: A related SATD cost obtaining unit, configured to, in response to the prediction method being intra prediction, obtain a first prediction mode corresponding to the current prediction unit among multiple prediction modes of the intra prediction by using a related SATD cost model, wherein the related SATD cost model is used to obtain multiple related SATD costs of predicting the current prediction unit by using multiple prediction modes of the intra prediction, each related SATD cost in the multiple related SATD costs does not include a residual bit cost related to the prediction method, and wherein the related SATD cost corresponding to the first prediction mode in the multiple related SATD costs is less than the related SATD cost corresponding to any prediction mode different from the first prediction mode; A first SATD cost model calculation unit, configured to obtain a second prediction mode corresponding to the current prediction unit among multiple prediction modes of the intra prediction by using the first SATD cost model, wherein, among multiple SATD costs obtained by using the first SATD cost model, the SATD cost corresponding to the second prediction mode is less than the SATD cost corresponding to any prediction mode different from the second prediction mode; A third determination subunit, configured to, in response to the first prediction mode and the second prediction mode being the same prediction mode among multiple prediction modes of the intra prediction, determine the same prediction mode as the target prediction mode; and A fourth determination subunit, configured to, in response to the first prediction mode and the second prediction mode being two different prediction modes among multiple prediction modes of the intra prediction, determine both the first prediction mode and the second prediction mode as the target prediction mode.
9. The device according to claim 8, wherein, The determination unit includes: An encoding unit obtaining unit, configured to obtain the current encoding unit; A processing method determination unit, configured to determine a processing method for the current encoding unit based on the size of the current encoding unit, at least by using the first SATD cost model and the second SATD cost model, wherein the processing method is one of multiple processes including at least intra prediction processing and inter prediction processing; and A prediction unit determination unit, configured to determine the current prediction unit based on the processing method.
10. The device according to claim 9, wherein, The processing method determination unit includes: A first calculation unit, configured to, in response to the size of the current encoding unit being equal to the size of the smallest encoding unit, obtain a first SATD cost set by using the first SATD cost model and multiple prediction modes of the intra prediction, wherein the current encoding unit is used as the prediction unit, obtain a second SATD cost set by using the second SATD cost model and multiple prediction modes of the inter prediction, wherein the current encoding unit is used as the prediction unit; A first obtaining unit, configured to obtain a first preset number of third SATD costs according to the ascending order of elements in the first SATD cost set and the second SATD cost set; and A first determining subunit, configured to determine the processing method based on the first preset number of third SATD costs.
11. The device according to claim 9, wherein, The processing method determining unit includes: A second calculating unit, configured to, in response to determining that the size of the current coding unit is greater than the size of the smallest coding unit, Obtain a fourth SATD cost set by using the first SATD cost model and multiple prediction modes of intra prediction, where the current coding unit is used as a prediction unit, Obtain a fifth SATD cost set by using the second SATD cost model and multiple prediction modes of inter prediction, where the current coding unit is used as a prediction unit, and Obtain a partition SATD cost set by using a partition SATD cost model and multiple prediction modes of intra prediction and multiple prediction modes of inter prediction, where when four sub-units obtained by partitioning the current coding unit are used as four prediction units, and where each element in the partition SATD cost set includes a residual bit cost based on partition prediction, and the residual bit cost based on partition prediction indicates the number of bits required to encode the prediction residual obtained by performing prediction after partitioning the current coding unit; and A second obtaining unit, configured to obtain a second preset number of sixth SATD costs according to the ascending order of elements in the fourth SATD cost set, the fifth SATD cost set, and the partition SATD cost set; and A second determining subunit, configured to determine the processing method based on the second preset number of sixth SATD costs, where the processing method is one of multiple processes including a partitioning process, an intra prediction process, and an inter prediction process.
12. The apparatus according to claim 9, wherein, The determining unit further includes: A prediction method determining unit, configured to, in response to determining that the processing method is one of an intra prediction process and an inter prediction process, determine the prediction method of the current coding unit, where the prediction method corresponds to the determined processing method; and A prediction unit determining unit, configured to determine the sub-unit corresponding to the prediction method obtained by partitioning the current coding unit as the current prediction unit.
13. The device according to claim 11, wherein It further includes: A partitioning unit, configured to, in response to the processing method being the partitioning process, partition the current coding unit into four sub-units; And A re-determining unit, configured to determine one of the four sub-units as the updated current coding unit and re-determine the current prediction unit of the current coding unit.
14. The apparatus according to claim 8, wherein, The obtaining unit includes: A second SATD cost model calculation unit, configured to, in response to the prediction mode being an inter-frame prediction, sequentially traverse multiple prediction modes of the current coding unit in the inter-frame prediction by using the second SATD cost model, so as to determine one or more inter-frame target prediction modes among the multiple prediction modes of the inter-frame prediction, wherein, for each inter-frame target prediction mode among the one or more inter-frame target prediction modes, the SATD cost corresponding to this inter-frame target prediction mode is smaller than the SATD cost corresponding to the prediction mode traversed before this target prediction mode among the multiple prediction modes of the inter-frame prediction; An intra-frame target prediction mode acquisition unit, configured to acquire an intra-frame target prediction mode among multiple prediction modes of the intra-frame prediction when the prediction mode is the intra-frame prediction, wherein, among multiple SATD costs obtained by using the first SATD cost model, the SATD cost corresponding to the intra-frame target prediction mode is smaller than the SATD cost corresponding to any prediction mode different from the intra-frame target prediction mode among the multiple prediction modes of the intra-frame prediction; and A fifth determination subunit, configured to obtain the target prediction mode based on the one or more inter-frame target prediction modes and the intra-frame target prediction mode.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Method and device for selecting macro block pattern
CN101304529A
Method and system for fast mode decision for high efficiency video coding
US20160127725A1
Image encoding / decoding method and device, and recording medium storing bitstream
US20200366900A1