Video encoding method, apparatus, computer readable medium, and electronic device

By comparing the current motion vector prediction with the collected motion vector prediction, unnecessary information acquisition is skipped, and the optimal motion vector is selected for encoding. This solves the problems of complexity and low efficiency in existing video coding protocols, achieves efficient video coding, and reduces the requirements for machine performance.

CN115314716BActive Publication Date: 2026-05-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-12-24
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing video encoding protocols are complex, have low encoding efficiency, and require high machine performance. Ordinary machines cannot achieve real-time encoding, which limits the application of encoding protocols.

Method used

By comparing the current motion vector prediction with the relevant information of the collected motion vector predictions, it can be determined whether the acquisition of relevant information can be terminated in advance, skipping unnecessary information acquisition processes, and selecting the optimal motion vector prediction for encoding.

Benefits of technology

It improves the efficiency of video encoding, reduces the requirements for machine performance, saves computing resources, and expands the application scope of encoding protocols.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115314716B_ABST
    Figure CN115314716B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video coding method, device, computer readable medium and electronic device. The video coding method comprises: comparing first related information of a current motion vector predictor (MVP) obtained and first related information of collected MVPs to obtain a comparison result; if it is determined that the acquisition of related information of the current MVP can be exited in advance according to the comparison result, skipping the acquisition of second related information of the current MVP, and sequentially acquiring first related information and second related information of other MVPs after the current MVP; selecting an optimal MVP corresponding to a current coding block according to the first related information and the second related information of the multiple MVPs corresponding to the current coding block acquired; and performing coding processing on the current coding block based on the optimal MVP. The technical solution of the embodiments of the present application can improve the efficiency of video coding and reduce the performance requirement of the machine.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is based on and claims priority to Chinese Patent Application No. 2021104907766, filed on May 6, 2021, entitled "Video Coding Method, Apparatus, Computer-Readable Medium and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of video coding technology, and more specifically, to a video coding method, apparatus, computer-readable medium, and electronic device. Background Technology

[0003] The future trend of video development is towards high definition, high frame rate, and high compression ratio. This requires continuous upgrades to video compression standards, and related video compression standards have already reached a high level in terms of compression ratio. However, current encoding protocols are too complex, have low encoding efficiency, and place high demands on machine performance. Ordinary machines cannot yet achieve real-time encoding capabilities, which limits the application of these encoding protocols. Summary of the Invention

[0004] The embodiments of this application provide a video encoding method, apparatus, computer-readable medium, and electronic device, which can at least to some extent improve the efficiency of video encoding, reduce the performance requirements of the machine, and enable lossless compression performance.

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0006] According to one aspect of the embodiments of this application, a video coding method is provided, comprising: comparing the first relevant information of the current motion vector prediction value (MVP) and the first relevant information of the collected MVPs to obtain a comparison result, wherein the current motion vector prediction value and the collected MVPs are MVPs corresponding to the current coding block; if it is determined according to the comparison result that the acquisition of relevant information of the current MVP can be exited in advance, then the acquisition of second relevant information of the current MVP is skipped, and the first and second relevant information of other MVPs after the current MVP are acquired in sequence; selecting the optimal MVP corresponding to the current coding block according to the first and second relevant information of the multiple MVPs corresponding to the current coding block; and encoding the current coding block based on the optimal MVP.

[0007] According to one aspect of the embodiments of this application, a video encoding apparatus is provided, comprising: a comparison unit, configured to compare the first relevant information of the acquired current motion vector prediction value (MVP) with the first relevant information of the collected MVPs to obtain a comparison result, wherein the current motion vector prediction value and the collected MVPs are MVPs corresponding to the current coding block; a skip unit, configured to skip the acquisition of the second relevant information of the current MVP if it is determined according to the comparison result that the acquisition of the relevant information of the current MVP can be exited in advance, and sequentially acquire the first relevant information and second relevant information of other MVPs after the current MVP; a selection unit, configured to select the optimal MVP corresponding to the current coding block according to the acquired first relevant information and second relevant information of multiple MVPs corresponding to the current coding block; and an encoding unit, configured to encode the current coding block based on the optimal MVP.

[0008] In some embodiments of this application, based on the foregoing scheme, the selection unit is further configured to: determine the optimal MVP corresponding to each combination of inter-frame prediction mode and reference frame as candidate MVPs based on the relevant information of each MVP corresponding to the current coding block; determine the optimal reference frame corresponding to each inter-frame prediction mode based on the relevant information of each candidate MVP; determine the optimal inter-frame prediction mode corresponding to the current coding block based on the relevant information of each candidate MVP corresponding to the optimal reference frame; and determine the optimal MVP corresponding to the current coding block based on the candidate MVP corresponding to the optimal inter-frame prediction mode.

[0009] In some embodiments of this application, based on the foregoing scheme, the skipping unit is further configured to: if it is determined according to the comparison result that the acquisition of relevant information of the current MVP should not be prematurely terminated, then continue to collect the first relevant information of the current MVP; after collecting the first relevant information of the current MVP, acquire the second relevant information of the current MVP.

[0010] In some embodiments of this application, based on the foregoing scheme, the skip unit is configured to: if the current inter-frame prediction mode is a combined reference frame mode, then obtain the optimal combined mode type, optimal interpolation method and optimal motion mode corresponding to the current MVP.

[0011] In some embodiments of this application, based on the foregoing scheme, the skipping unit is further configured to: if relevant information corresponding to all MVPs of the current inter-frame prediction mode has been obtained, then obtain the MVP corresponding to other inter-frame prediction modes as the current MVP, and continue to obtain relevant information corresponding to the current MVP.

[0012] In some embodiments of this application, based on the foregoing scheme, the first relevant information corresponding to the current motion vector prediction value MVP includes: the motion vector corresponding to the current MVP, the current reference frame corresponding to the current MVP, and a portion of the bit count corresponding to the current MVP.

[0013] In some embodiments of this application, based on the foregoing scheme, the inter-frame prediction mode corresponding to the current MVP includes NEWMV, and the skip unit is further configured to: perform motion estimation based on the current MVP to obtain the optimal motion vector corresponding to the current MVP; and determine the optimal motion vector as the motion vector corresponding to the current MVP.

[0014] In some embodiments of this application, based on the foregoing scheme, the inter-frame prediction mode corresponding to the current MVP does not include NEWMV, and the skip unit is further configured to: determine the current MVP as the motion vector corresponding to the current MVP.

[0015] In some embodiments of this application, based on the foregoing scheme, the skip unit is configured as follows: if, according to the comparison result, it is determined that the motion vector of the current MVP is the same as the motion vector of the collected MVP, the reference frame of the current MVP is the same as the reference frame of the collected MVP, and the partial bit count of the current MVP is greater than the partial bit count of the collected MVP, then it is determined that it is necessary to exit the acquisition of relevant information of the current MVP in advance.

[0016] In some embodiments of this application, based on the foregoing scheme, the skip unit is configured as follows: if, according to the comparison result, it is determined that the forward reference frame of the current MVP is the same as the forward reference frame of the collected MVP, the backward reference frame of the current MVP is the same as the backward reference frame of the collected MVP, the motion vector corresponding to the forward reference frame of the current MVP is the same as the motion vector corresponding to the forward reference frame of the collected MVP, the motion vector corresponding to the backward reference frame of the current MVP is the same as the motion vector corresponding to the backward reference frame of the collected MVP, and the partial bit count of the current MVP is greater than the partial bit count of the collected MVP, then it is determined that it is necessary to exit the acquisition of relevant information of the current MVP in advance.

[0017] In some embodiments of this application, based on the foregoing scheme, the comparison unit is further configured to: determine whether the current inter-frame prediction mode meets a predetermined condition, wherein the predetermined condition is generated based on the inter-frame prediction modes of GLOBALMV and GLOBAL_GLOBALMV; and if the current inter-frame prediction mode meets the predetermined condition, perform a process of comparing the first relevant information of the current motion vector prediction value MVP obtained with the first relevant information of the collected MVP.

[0018] In some embodiments of this application, based on the foregoing scheme, the skip unit is configured to: if it is determined according to the comparison result that the acquisition of relevant information of the current MVP should not be prematurely terminated and the current inter-frame prediction mode meets the predetermined conditions, then the first relevant information of the current MVP is collected.

[0019] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the video encoding method as described in the above embodiments.

[0020] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the video encoding method as described in the above embodiments.

[0021] In some embodiments of this application, during the prediction mode selection stage, it is first determined whether the acquisition of relevant information for the current MVP can be prematurely terminated based on the first relevant information of the current motion vector prediction value (MVP). If premature termination is achieved, the acquisition of the second relevant information for the current MVP is skipped. Therefore, for some MVPs, it is unnecessary to acquire the corresponding second relevant information, and the selection results of the optimal MVP and optimal inter-frame prediction mode are not affected even without acquiring the second relevant information for these MVPs. Furthermore, since acquiring the second relevant information requires a huge amount of computation, the computational load of video encoding is greatly reduced without adding new computations, and the accuracy is very high. While ensuring lossless compression performance, the efficiency of video encoding can be significantly improved, saving overall computational resources and reducing the performance requirements of the machine. This allows even machines with lower performance to perform video encoding quickly, thus expanding the application scope of the encoding protocol.

[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0024] Figure 1A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown;

[0025] Figure 2 This diagram illustrates the placement of the video encoding and decoding devices in a streaming system.

[0026] Figure 3A A schematic diagram of a standard coding framework according to an embodiment of this application is shown;

[0027] Figure 3B A schematic diagram illustrating the type of encoding unit segmentation under the AV1 encoding protocol according to an embodiment of this application is shown;

[0028] Figure 4 A schematic diagram showing the corresponding positions of motion vector prediction values ​​under different single-reference frame modes of the AV1 encoding protocol according to an embodiment of this application is illustrated.

[0029] Figure 5 A flowchart illustrating the optimal outcome selection process corresponding to any combination of prediction mode and reference frame in related technologies is shown.

[0030] Figure 6 A flowchart of a video encoding method according to an embodiment of this application is shown;

[0031] Figure 7 A flowchart illustrating the selection of the optimal MVP corresponding to the current coding block according to an embodiment of this application is shown;

[0032] Figure 8 An embodiment according to this application is shown. Figure 6 Flowchart of the steps following step 610;

[0033] Figure 9 A flowchart illustrating the optimal result selection process corresponding to any combination of prediction mode and reference frame according to an embodiment of this application is shown.

[0034] Figure 10 A schematic diagram of a diamond-shaped search template according to an embodiment of this application is shown;

[0035] Figure 11 A schematic diagram of a two-point search according to an embodiment of this application is shown;

[0036] Figure 12 A partial schematic diagram of a raster scanning method for searching location points according to an embodiment of this application is shown;

[0037] Figure 13 A schematic diagram of a large diamond search template according to an embodiment of this application is shown;

[0038] Figure 14 A schematic diagram of a small diamond search template according to an embodiment of this application is shown;

[0039] Figure 15 A schematic diagram illustrating the process of determining early exit conditions in a single reference frame mode according to an embodiment of this application is shown.

[0040] Figure 16 A schematic diagram illustrating the process of determining early exit conditions in a combined reference frame mode according to an embodiment of this application is shown.

[0041] Figure 17 A block diagram of a video encoding apparatus according to an embodiment of this application is shown;

[0042] Figure 18 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0043] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0044] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0045] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0046] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0047] The following is an explanation of some of the proper nouns used in this application.

[0048] ME: Motion Estimation, which is defined in protocols such as HEVC (High Efficiency Video Coding).

[0049] MV: motion vector, is a vector used to mark the positional relationship between the current block and the reference block during inter-frame prediction.

[0050] MVP: Motion Vector Prediction, which is the initial position of the MVP derived from neighboring blocks.

[0051] MVD: Motion Vector Difference. MVD = MV-MVP. By encoding the difference between the predicted and actual values ​​of MV, the number of bits consumed can be reduced.

[0052] rdcost: Rate Distortion Cost, used for selecting the best among multiple options.

[0053] SAD: Sum of Absolute Difference, which only reflects the time-domain difference of the residuals and cannot effectively reflect the size of the bitstream.

[0054] SATD: Sum of Absolute Transformed Difference, is a method for calculating distortion by summing the absolute values ​​of the residual signal after Hadamard transformation. It involves performing Hadamard transformation on the residual signal and then summing the absolute values ​​of each element. Compared to SAD, it is more computationally complex but also more accurate.

[0055] SSE: Represents the sum of squares of the errors between the original pixel and the reconstructed pixel. It requires the process of transforming, quantizing, inverse quantizing, and inverse transforming the residual signal. The estimated code is the same as the actual encoded code. The selected mode saves the most code, but also has the greatest computational complexity.

[0056] Video encoding refers to the method of converting a file in an original video format into a file in another video format through compression technology.

[0057] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.

[0058] like Figure 1As shown, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, a network 150. For instance, system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via network 150. Figure 1 In this embodiment, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission. In practical applications, the first terminal device 110 and the second terminal device 120 can exist as nodes in a blockchain, and data transmission can take place within the blockchain. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block.

[0059] For example, the first terminal device 110 can encode video data (e.g., a video image stream captured by the terminal device 110) to transmit it to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to recover the video data, and display video images based on the recovered video data.

[0060] In one embodiment of this application, system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission over network 150 to the other terminal device. Each of the third terminal device 130 and the fourth terminal device 140 may also receive encoded video data transmitted by the other terminal device, decode the encoded video data to recover the video data, and display the video images on an accessible display device based on the recovered video data.

[0061] exist Figure 1In the embodiments disclosed herein, the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140 may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited to these. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 150 refers to any number of networks that transmit encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140, including, for example, wired and / or wireless communication networks. Communication network 150 may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 150 may be irrelevant to the operation of this application.

[0062] In one embodiment of this application, Figure 2 The illustration shows the placement of video encoding and decoding devices in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television (television), storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0063] The streaming system may include an acquisition subsystem 213, which may include a video source 201 such as a digital camera, which creates an uncompressed video image stream 202. In an embodiment, the video image stream 202 includes samples captured by a digital camera. The video image stream 202 is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data 204 (or encoded video bitstream 204). The video image stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data 204 (or encoded video bitstream 204) is depicted as a thin line to emphasize the lower data volume of the encoded video data 204 (or encoded video bitstream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as Figure 2Client subsystems 206 and 208 can access streaming server 205 to retrieve copies 207 and 209 of encoded video data 204. Client subsystem 206 may include, for example, a video decoding device 210 in electronic device 230. Video decoding device 210 decodes the incoming copy 207 of the encoded video data and produces an output video picture stream 211 that can be displayed on display 212 (e.g., a screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video streams) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In embodiments, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.

[0064] It should be noted that electronic devices 220 and 230 may include other components not shown in the figures. For example, electronic device 220 may include a video decoding device, and electronic device 230 may also include a video encoding device.

[0065] In one embodiment of this application, taking the international video coding standards HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding), as well as the Chinese national video coding standard AVS (Audio Video Coding Standard), as examples, after an input video frame image is received, the video frame image is divided into several non-overlapping processing units according to a block size. Each processing unit will perform a similar compression operation. This processing unit is called a CTU (Coding Tree Unit) or LCU (Largest Coding Unit). The CTU can be further subdivided into more refined units to obtain one or more basic coding units CU. The CU is the most basic element in a coding process, and the CU corresponds to the CB (Coding Block).

[0066] Figure 3A A schematic diagram of a standard coding framework according to an embodiment of this application is shown. The coding process under this standard coding framework is as follows: Current frame F n The image signal and the reference frame F n-1The predicted image signals (inter-frame or intra-frame) obtained from the prediction are interpolated to obtain residual signals. These residual signals are then transformed and quantized to obtain quantization coefficients. These coefficients are then used for entropy coding to obtain an encoded bitstream, and for inverse quantization and inverse transform to obtain the reconstructed residual signal. The predicted image signal and the reconstructed residual signal are superimposed to generate an image signal. This image signal is then input to the intra-frame prediction selection module and the intra-frame prediction module for intra-frame prediction processing. Simultaneously, it is filtered (usually through a loop filter) to output the reconstructed frame F'. n The image signal is used to reconstruct frame F'. n The image signal can be used as a reference image for the next frame to perform motion estimation (ME) and motion compensation (MC) prediction. Then, based on the results of the motion compensation prediction and the prediction results, the predicted image signal for the next frame is obtained, and the above process is repeated until the encoding is completed.

[0067] Figure 3B This diagram illustrates the types of coding unit segmentation in the AV1 (Alliance for Open Media Video 1) coding protocol. There are 10 main types: NONE, Split, HORZ (horizontal bisection), VERT (vertical bisection), HORZ_4 (horizontal quadrupling), HORZ_A (first horizontal trisection), HORZ_B (second horizontal trisection), VERT_A (first vertical trisection), VERT_B (second vertical trisection), and VERT_4 (vertical quadrupling).

[0068] Figure 3B In the illustrated embodiment, the segmentation type of the coding unit corresponds to 22 coding block sizes, namely: 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, 32×16, 32×32, 32×64, 64×32, 64×64, 64×128, 128×64, 128×128, 4×16, 16×4, 8×32, 32×8, 16×64, and 64×16.

[0069] Each coding unit contains two prediction types: intra-frame prediction mode and inter-frame prediction mode. First, within the same prediction type, different prediction modes are compared to find the optimal segmentation mode. Then, intra-frame and inter-frame prediction modes are compared to find the optimal prediction mode for the current coding unit. Simultaneously, transform units are applied to the coding unit; each coding unit corresponds to multiple transform types, and the optimal transform type is found. Finally, a frame of image is divided into coding units.

[0070] Intra-frame prediction modes include the following: mean prediction based on reference pixels above and to the left (DC_PRED), prediction combining horizontal and vertical interpolation (SMOOTH_PRED), vertical interpolation prediction (SMOOTH_V_PRED), horizontal interpolation prediction (SMOOTH_H_PRED), prediction in the direction of minimum gradient (PEATH_PRED), and prediction in eight different principal directions: vertical prediction (V_PRED), horizontal prediction (H_PRED), 45-degree angle prediction (D45_PRED), 67-degree angle prediction (D67_PRED), 113-degree angle prediction (D113_PRED), 135-degree angle prediction (D135_PRED), 157-degree angle prediction (D157_PRED), and 203-degree angle prediction (D203_PRED). Each principal direction includes six angular offsets: plus or minus 3 degrees, plus or minus 6 degrees, and plus or minus 9 degrees. In some cases, intra-frame prediction modes may also include palette prediction mode and intra-block copy prediction.

[0071] Inter-frame prediction modes include single-reference frame mode and combined reference frame mode. Single-reference frame mode includes four types: NEARESTMV, NEARMV, GLOBALMV, and NEWMV; combined reference frame mode includes eight types: NEAREST_NEARESTMV, NEAR_NEARMV, NEAREST_NEWMV, NEW_NEARESTMV, NEAR_NEWMV, NEW_NEARMV, GLOBAL_GLOBALMV, and NEW_NEWMV. NEARESTMV and NEARMV refer to the prediction block's MV being derived from surrounding block information, without needing to transmit the MVD; while NEWMV requires transmitting the MVD, and GLOBALMV means the block's MV information is derived from global motion. Therefore, inter-frame prediction modes including NEARESTMV, NEARMV, and NEWMV all rely on MVP derivation. For example, the patterns NEWMV, NEW_NEWMV, NEAREST_NEWMV, NEW_NEARESTMV, NEAR_NEWMV, and NEW_NEARMV all contain NEWMV, and these patterns all rely on MVP derivation.

[0072] For a given reference frame, the AV1 standard will calculate 4 MVPs according to the rules.

[0073] The MVP derivation process under the AV1 encoding protocol can be as follows: Scan the block information of the left 1 / 3 / 5 columns and the top 1 / 3 / 5 rows in a skipping manner. First, select the blocks that use the same reference frame and remove duplicate MVs. If the number of non-repeating MVs is less than 8, relax the requirement to reference frames in the same direction and continue to add MVs. If it is still less than 8, fill it with global motion vectors. After selecting 8 MVs, sort them according to importance and select the 4 most important MVs.

[0074] Figure 4 This diagram illustrates the corresponding positions of motion vector prediction values ​​under different single-reference frame modes of the AV1 encoding protocol according to an embodiment of this application. Please refer to... Figure 4 It displays a dynamic list of reference frames, where Ref1, Ref2, and Ref3 are the reference frames, and the corresponding MV records for each reference frame are recorded in the corresponding column. For example, the four MV1 records in the column corresponding to Ref1 are the four MV records selected based on the reference frame Ref1. Figure 4 As can be seen, among the four most important prediction modes, the 0th prediction mode is designated as the MVP, the 1st to 3rd prediction modes are designated as the MVP, and the 0th to 2nd prediction modes are designated as the MVP.

[0075] Reference Frame Type value meaning INTRA_FRAME 0 Intra prediction, inter_intra LAST_FRAME 1 The point of view (PoC) is less than the closest reference frame in the current frame; forward reference... LAST2_FRAME 2 The point of view (POC) is less than the second closest reference frame in the current frame. (Forward reference) LAST3_FRAME 3 The point of view (POC) is less than the third closest reference frame in the current frame; forward reference... GOLDEN_FRAME 4 The point of view (POC) is less than the I-frame or GPB frame corresponding to the current frame, similar to a long-term reference frame. BWDREF_FRAME 5 The point of view (PoC) is greater than the closest reference frame in the current frame, and the backward reference is... ALTREF2_FRAME 6 The point of view (POC) is greater than the second closest reference frame in the current frame, and the backward reference frame is also greater. ALTREF_FRAME 7 The point of view (POC) is greater than the third closest reference frame in the current frame, and the backward reference frame is also greater.

[0076] Table 1

[0077] Each prediction mode corresponds to a different reference frame. The reference frames and their corresponding meanings under the AV1 encoding protocol are shown in Table 1. Among them, Poc (picture order count) is the image sequence number, INTRA_FRAME represents that the intra-frame prediction mode does not use reference frames, and values ​​1 to 7 are the 7 reference frames corresponding to the inter-frame prediction mode.

[0078] For the four single-reference frame modes mentioned above, each mode corresponds to seven reference frames: LAST_FRAME, LAST2_FRAME, LAST3_FRAME, GOLDEN_FRAME, BWDREF_FRAME, ALTREF2_FRAME, and ALTREF_FRAME. Therefore, there are a total of 4 × 7 = 28 combinations of single-reference frame modes and reference frames.

[0079] For the 8 aforementioned combined reference frame modes, the 16 reference frame combinations corresponding to each combined reference frame mode are respectively {LAST_FRAME, ALTREF_FRAME}, {LAST2_FRAME, ALTREF_FRAME}, {LAST3_FRAME, ALTREF_FRAME}, {GOLDEN_FRAME, ALTREF_FRAME}, {LAST_FRAME, BWDREF_FRAME}, {LAST2_FRAME, BWDREF_FRAME}, {LAST3_FRAME, BWDREF_FRAME}, {GOLDEN_FRAME, BWDREF_FRAME}, {LAST_FRAME, ALTREF2_FRAME}, {LAST2_FRAME, ALTREF2_FRAME}, {LAST3_FRAME, ALTREF2_FRAME}, {GOLDEN_FRAME, ALTREF2_FRAME}, {LAST_FRAME, LAST2_FRAME}, {LAST_FRAME, LAST3_FRAME}, {LAST_FRAME, GOLDEN_FRAME}, {BWDREF_FRAME, ALTREF_FRAME}. Therefore, there are a total of 8 × 16 = 128 combinations of combined reference frame modes and reference frames.

[0080] Therefore, there are a total of 156 (4 × 7 + 8 × 16) combinations of inter-frame prediction modes and reference frames. For any combination of an inter-frame prediction mode and a reference frame, the current combination corresponds to a maximum of 3 mvps. The prediction information of the coding module is found through four processes: motion estimation for the current mvp (Note: Motion estimation is only performed when the prediction mode contains NEWMV), selection of the optimal combined mode type, selection of the optimal interpolation method, and selection of the optimal motion mode.

[0081] Figure 5 The flowchart showing the optimal result selection process corresponding to any combination of a prediction mode and a reference frame in the related art is shown. The above process is specifically as Figure 5 shown, and includes the following steps:

[0082] Step 310, N = 0, and obtain the number of mvps ref_set.

[0083] Among them, N is used for counting, and ref_set is the number of mvps corresponding to the current inter-frame prediction mode and the current reference frame.

[0084] Step 320, determine whether N < ref_set holds. If so, execute Step 330; otherwise, execute Step 340.

[0085] Step 330, obtain MVP, N+1.

[0086] That is, get the MVP as the current MVP and increment N by 1.

[0087] Step 340: Obtain relevant information corresponding to the motion vector prediction value MVP for the next prediction mode.

[0088] Step 350: Determine whether the current prediction mode contains NEWMV. If so, execute step 360 first, then execute step 390; otherwise, execute step 390 directly.

[0089] Step 360: Perform motion estimation, that is, search for the optimal motion vector corresponding to the current MVP.

[0090] Step 390: Determine whether the current prediction mode is a combined reference frame mode. If so, execute step 3100 first, then execute step 3110; otherwise, execute step 3110 directly.

[0091] Step 3100: Select the best type.

[0092] There are four types of combined reference frame modes: AVERAGE, DISTWTD, WEDGE, and DIFFWTD. This step involves selecting the optimal one from these four types as the optimal combined mode type. The predicted pixels of the two reference frames are then fused together. Each combined mode type corresponds to a predicted pixel fusion method, which is specified in the AV1 protocol and will not be detailed further.

[0093] Step 3110: Select the best interpolation method under BestMv.

[0094] There are nine interpolation methods, and this step involves selecting the optimal one from these nine methods. The specific implementation details for each interpolation method are defined by the AV1 protocol and will not be elaborated here.

[0095] Step 3120: Select the optimal exercise mode.

[0096] After completing step 3120, re-execute step 320.

[0097] The motion modes corresponding to single-reference frame mode and combined reference frame mode are different. Single-reference frame mode corresponds to four motion modes: SIMPLE, OBMC, WARPED, and SIMPLE(inter_intra); combined reference frame mode corresponds to only one motion mode: SIMPLE. The specific implementation of each motion mode is defined by the AV1 protocol and will not be detailed here. Motion mode selection requires transforming, quantizing, inverse quantizing, and inverse transforming the residual signal, and calculating the integrity rate distortion cost; therefore, it has the highest computational complexity.

[0098] Therefore, it can be seen that the computational complexity of the prediction process for a coding unit is very high. A coding unit can have a maximum of 376 MVPs. Among them, the single reference frame mode has a maximum of 56 MVPs (7×3+7×3+7+7, where NEARESTMV and GLOBALMV each have 7 MVPs, and NEWMV and NEARMV each have 7×3 MVPs), and the combined reference frame mode has a maximum of 320 MVPs (except for NEAREST_NEARESTMV and GLOBAL_GLOBALMV, which have 16 MVPs, the other 6 combined reference frame modes each have 16×3 MVPs).

[0099] In related techniques, optimization is only performed on multiple MVPs for a single reference frame of NEARMV. The rate-distortion cost of each MVP is estimated, and MVPs are then discarded based on a threshold. However, this method, which relies on estimating the rate-distortion cost of each MVP, is simplistic and may be inaccurate. It is only useful when the differences between rate-distortion costs are significant; conversely, if the differences are small, MVPs cannot be discarded, increasing computational burden. Therefore, this technique requires substantial computation, has low coding efficiency, and low accuracy.

[0100] The inventors of this application discovered that, regardless of whether it is a single reference frame mode or a combined reference frame mode, once the predicted values ​​(motion vector and reference frame) sent to interpolation are the same, the result obtained from selecting the best interpolation method from the nine interpolation methods during the interpolation method selection process is also the same. Furthermore, the number of bits consumed by the generated distortion and residual coefficients, the number of bits consumed by the transform type, and the number of bits consumed by the TU segmentation type are also the same when performing a complete rate-distortion cost process for several motion modes. Therefore, once the predicted pixels are the same, among all the factors affecting the rate-distortion cost of the current MVP, the number of bits consumed before interpolation is particularly critical. Therefore, the judgment of interpolation and motion mode can be exited in advance by judging the predicted value and the number of bits consumed before interpolation.

[0101] Therefore, this application first provides a video encoding method. The video encoding method provided in this application can be applied to any scenario requiring video encoding, such as live streaming platforms, online conferencing platforms, and short video platforms.

[0102] For example, in a live streaming scenario, when the streamer is a game streamer, the encoded video data generated by executing the video encoding method provided in this application embodiment can include the streamer's live footage and the game footage; when the streamer is an entertainment streamer or a shopping streamer, the encoded video data generated by executing the video encoding method provided in this application embodiment can include the streamer's live footage.

[0103] In one embodiment of this application, the video encoding method is applied to a short video platform. After a user uploads a short video and obtains the original video data using their terminal device, they encode the original video data by executing the video encoding method provided in this embodiment to generate encoded video data. The encoded video data is then transmitted to the short video platform via the network. Subsequently, when a user requests to watch a short video uploaded by the same user, the short video platform sends the corresponding encoded video data to the terminal device of the user via the network. Finally, the terminal device decodes and renders the video to enable playback, thus realizing the entire process of short video sharing.

[0104] Therefore, since the embodiments of this application can improve the efficiency of video encoding, the overhead of computing resources is saved, the latency of video transmission is reduced, and the user experience is improved; it can also reduce the performance requirements of the machine, enabling more machines with lower performance to perform video encoding quickly, and expanding the application scope of the encoding protocol.

[0105] It should be noted that although the video encoding method in this embodiment is executed by the terminal device, and correspondingly, the video decoding device is generally located in the terminal device, in other embodiments of this application, the video encoding method may also be executed by a server or a cluster of servers, such as by a live streaming platform or a short video platform. This application does not impose any limitations on this, and the scope of protection of this application should not be limited as a result.

[0106] The video encoding method provided in this application can be applied to cloud technology fields such as cloud gaming, cloud education, and cloud conferencing.

[0107] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.

[0108] Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.

[0109] Cloud gaming, also known as gaming on demand, is an online gaming technology based on cloud computing. It enables thin clients with relatively limited graphics processing and data processing capabilities to run high-quality games. In cloud gaming, the game does not reside on the player's terminal but runs on a cloud server. The cloud server renders the game scene as a video and audio stream, which is then transmitted to the player's terminal via the network. The player's terminal does not need powerful graphics processing and data processing capabilities; it only needs basic streaming media playback capabilities and the ability to receive player input commands and send them to the cloud server.

[0110] Cloud Computing Education (CCEDU) refers to an education platform service based on cloud computing business models. On the cloud platform, all educational institutions, training institutions, enrollment service agencies, publicity agencies, industry associations, management agencies, industry media, legal structures, etc., are centrally integrated into a resource pool. These resources can be displayed and interacted with each other, communicate on demand, and reach agreements, thereby reducing education costs and improving efficiency.

[0111] Cloud conferencing is an efficient, convenient, and low-cost form of meeting based on cloud computing technology. Users only need to use an internet interface to quickly and efficiently share voice, data files, and video with teams and clients around the world. The complex technologies such as data transmission and processing during the meeting are handled by the cloud conferencing service provider.

[0112] Currently, cloud conferencing in China mainly focuses on services based on the SaaS (Software as a Service) model, including telephone, internet, and video services. Video conferencing based on cloud computing is called cloud conferencing.

[0113] In the era of cloud conferencing, data transmission, processing, and storage are all handled by the computer resources of the video conferencing vendors. Users no longer need to purchase expensive hardware or install cumbersome software. They can simply open a browser, log in to the corresponding interface, and conduct efficient remote meetings.

[0114] Cloud conferencing systems support dynamic multi-server cluster deployment and provide multiple high-performance servers, significantly improving meeting stability, security, and availability. In recent years, video conferencing has gained popularity due to its ability to greatly improve communication efficiency, continuously reduce communication costs, and upgrade internal management, and is widely used in various fields such as transportation, logistics, finance, telecommunications, education, and enterprises. Undoubtedly, with the application of cloud computing, video conferencing will be even more attractive in terms of convenience, speed, and ease of use, inevitably triggering a new surge in video conferencing applications.

[0115] The implementation details of the technical solutions in the embodiments of this application are described in detail below:

[0116] Figure 6 A flowchart of a video encoding method according to an embodiment of this application is shown. This video encoding method can be executed by a device with computing capabilities, such as a server or a terminal device like a smartphone. (Refer to...) Figure 6 As shown, the video encoding method includes at least the following steps:

[0117] In step 610, the first relevant information of the current motion vector prediction value MVP is compared with the first relevant information of the collected MVP to obtain the comparison result.

[0118] The current motion vector prediction and the collected MVP are the MVPs corresponding to the current coding block.

[0119] In one embodiment of this application, before comparing the first relevant information of the acquired current motion vector prediction value MVP with the first relevant information of the collected MVP, the method further includes:

[0120] Get the current MVP;

[0121] Retrieve the first relevant information corresponding to the current MVP.

[0122] In one embodiment of this application, the first relevant information corresponding to the current motion vector prediction value MVP includes: the motion vector corresponding to the current MVP, the current reference frame corresponding to the current MVP, and a portion of the bit count corresponding to the current MVP.

[0123] Specifically, the partial bit count corresponding to the current MVP includes the MVD (Motion Vector Deviation) consumption bits corresponding to the current MVP, the reference frame consumption bits corresponding to the current MVP, and the MVP index consumption bits. That is, the partial bit count is the sum of the MVD consumption bits, the reference frame consumption bits, and the MVP index consumption bits. Among them, the MVD consumption bits are 0 when no motion estimation is performed.

[0124] Please continue to refer to Figure 5 MVD and its bit consumption are generated through motion estimation. When motion estimation is not performed, the bit consumption of MVD is 0, and at this time, the current MVP and its corresponding current reference frame are also determined. Therefore, the bit consumption of MVD, the bit consumption of the reference frame corresponding to the current MVP, and the bit consumption of the MVP index are all bit consumption numbers that can be obtained before the interpolation method is optimized.

[0125] The first relevant information of the collected MVP corresponds to other MVPs whose relevant information was acquired before the current MVP. The first relevant information of the collected MVP also includes the motion vectors, current reference frame and part of the bit count of the other MVPs. The first relevant information of the other MVPs is acquired in the same way as the current MVP.

[0126] Step 610 is able to Figure 5 This step is performed before the BestMv interpolation method selection step; the specific execution method will be introduced in later content.

[0127] In one embodiment of this application, the inter-frame prediction mode corresponding to the current MVP includes NEWMV, and the motion vector corresponding to the current MVP is obtained through the following process:

[0128] Motion estimation is performed based on the current MVP to obtain the optimal motion vector corresponding to the current MVP;

[0129] The optimal motion vector is determined as the motion vector corresponding to the current MVP.

[0130] As mentioned earlier, the patterns NEWMV, NEW_NEWMV, NEAREST_NEWMV, NEW_NEARESTMV, NEAR_NEWMV, and NEW_NEARMV all contain NEWMV. These patterns all rely on MVP derivation, which means that motion estimation is required.

[0131] In one embodiment of this application, the inter-frame prediction mode corresponding to the current MVP does not include NEWMV, and the motion vector corresponding to the current MVP is obtained through the following process:

[0132] The current MVP is determined as the motion vector corresponding to the current MVP.

[0133] If the inter-frame prediction mode does not include NEWMV, then motion estimation is not required, and the current MVP is used as the motion vector.

[0134] In one embodiment of this application, before comparing the first relevant information of the acquired current motion vector prediction value MVP with the first relevant information of the collected MVP, the method further includes:

[0135] Determine whether the current inter-frame prediction mode meets the predetermined conditions, wherein the predetermined conditions are generated based on the inter-frame prediction modes of GLOBALMV and GLOBAL_GLOBALMV.

[0136] If the current inter-frame prediction mode meets the predetermined conditions, the process of comparing the first relevant information of the current motion vector prediction value MVP obtained with the first relevant information of the collected MVP is executed.

[0137] Specifically, if the current inter-frame prediction mode is not GLOBALMV or not GLOBAL_GLOBALMV, then the predetermined conditions are met.

[0138] Otherwise, if the current inter-frame prediction mode is GLOBALMV or GLOBAL_GLOBALMV, and the current mode warpmotion type is not rotation, not zoom, or not affine, then the predetermined conditions are met.

[0139] In this embodiment of the application, predetermined conditions are set for the use of the solution of this application, taking into account special inter-frame prediction modes, so as to make the video encoding process more accurate.

[0140] In step 620, if it is determined from the comparison result that it is possible to exit the acquisition of relevant information of the current MVP in advance, then the acquisition of the second relevant information of the current MVP is skipped, and the first and second relevant information of other MVPs after the current MVP are acquired in sequence.

[0141] The relevant information of the current MVP includes the first relevant information and the second relevant information. If you exit the acquisition of the relevant information of the current MVP in advance, it means that you exit before you have acquired all the relevant information of the current MVP. Specifically, at this time, only the first relevant information has been acquired, and the second relevant information has not been acquired.

[0142] For an MVP, the second set of relevant information is obtained after the first set of relevant information is obtained.

[0143] Specifically, in Figure 5 In the embodiment, both the BestMv interpolation method selection and motion mode selection processes can obtain the second relevant information. However, these processes are computationally complex. For example, motion mode selection requires transforming, quantizing, inverse quantizing, and inverse transforming the residual signal, as well as calculating the integrity rate distortion cost. Therefore, these processes are one of the important reasons for the low video coding efficiency and high machine performance requirements. By skipping the acquisition of the second relevant information of the current MVP in step 620, the video coding efficiency can be greatly improved.

[0144] For other MVPs following the current MVP, the process should be the same as for the current MVP, determining whether to skip obtaining the second relevant information based on the comparison results.

[0145] In one embodiment of this application, the video encoding method further includes:

[0146] If, based on the comparison results, it is determined that the motion vector of the current MVP is the same as the motion vector of the collected MVP, the reference frame of the current MVP is the same as the reference frame of the collected MVP, and the number of bits in the current MVP is greater than the number of bits in the collected MVP, then it is determined that the acquisition of relevant information of the current MVP needs to be terminated in advance.

[0147] For cases where the current inter-frame prediction mode is single reference frame mode, the motion vector, reference frame, partial bit count of the current MVP, and the collected MVP information corresponding to the single reference frame mode are compared.

[0148] MVP selection is performed by calculating rate-distortion cost. The lower the rate-distortion cost, the more likely the corresponding MVP is to be selected. The rate-distortion cost is determined by the sum of distortion and all bits consumed by the MVP. As mentioned earlier, if the motion vectors and reference frames of two MVPs are the same, then the bit consumption for distortion, interpolation method, residual coefficients, transform type, and TU segmentation type are also the same. Therefore, the MVP selection result is determined by the number of bits consumed by the two MVPs before interpolation, i.e., the partial bit count. Thus, if the motion vector and reference frame of the current MVP are the same as those of the collected MVPs, and the partial bit count of the current MVP is greater than that of the collected MVPs, it means that the current MVP is definitely not the optimal MVP. In this case, the acquisition of relevant information for the current MVP is stopped, i.e., the acquisition of the second relevant information for the current MVP is skipped.

[0149] In one embodiment of this application, the video encoding method further includes:

[0150] If, based on the comparison results, it is determined that the forward reference frame of the current MVP is the same as the forward reference frame of the collected MVPs, the backward reference frame of the current MVP is the same as the backward reference frame of the collected MVPs, the motion vector corresponding to the forward reference frame of the current MVP is the same as the motion vector corresponding to the forward reference frame of the collected MVPs, and the motion vector corresponding to the backward reference frame of the current MVP is the same as the motion vector corresponding to the backward reference frame of the collected MVPs, and the number of bits in the current MVP is greater than the number of bits in the collected MVPs, then it is determined that the acquisition of relevant information of the current MVP needs to be terminated in advance.

[0151] In the case where the current inter-frame prediction mode is the combined reference frame mode, due to the different encoding methods, in addition to comparing the number of bits of the current MVP and the collected MVP, it is also necessary to compare the forward reference frame, the backward reference frame, the motion vector corresponding to the forward reference frame, and the motion vector corresponding to the backward reference frame of both.

[0152] Figure 8 An embodiment according to this application is shown. Figure 6 The flowchart for steps following step 610 is shown below. Please refer to [link / reference]. Figure 8 After comparing the first relevant information of the current motion vector prediction value MVP with the first relevant information of the collected MVP, the following steps are also included:

[0153] In step 810, if it is determined based on the comparison results that the acquisition of relevant information of the current MVP should not be prematurely terminated, then the first relevant information of the current MVP will continue to be collected.

[0154] If, based on the comparison results, it is determined not to prematurely exit the acquisition of relevant information for the current MVP, the first relevant information of the current MVP is collected so that the first relevant information of the current MVP can be used to compare with the first relevant information of other MVPs in the future. The first relevant information of the MVP collected in step 610 is collected in this way.

[0155] In one embodiment of this application, the step of continuing to collect the first relevant information of the current MVP if it is determined based on the comparison result that the acquisition of relevant information of the current MVP should not be prematurely terminated includes:

[0156] If, based on the comparison results, it is determined that the acquisition of relevant information for the current MVP should not be prematurely terminated and the current inter-frame prediction mode meets the predetermined conditions, then the first relevant information for the current MVP will continue to be collected.

[0157] The predetermined conditions here are consistent with those mentioned in the previous embodiments. That is, if the current inter-frame prediction mode is not GLOBALMV or not GLOBAL_GLOBALMV, then the predetermined conditions are met; otherwise, if the current inter-frame prediction mode is GLOBALMV or GLOBAL_GLOBALMV, and the current mode warp motion type is not rotation, not zoom, not affine, then the predetermined conditions are met.

[0158] It should be noted that although the above embodiments perform conditional checks before collecting the first relevant information of the current MVP, and only collect the first relevant information of the current MVP when the corresponding conditions are met, it is easy to understand that the first relevant information of the current MVP can also be collected without conditional checks, although in this case more first relevant information needs to be checked. Therefore, the advantages of using the above embodiments also include: reducing the amount of first relevant information collected and used for comparison, thereby saving computational overhead and improving video encoding efficiency.

[0159] In step 820, after collecting the first relevant information of the current MVP, the second relevant information of the current MVP is obtained.

[0160] If the comparison results indicate that we should not prematurely exit the acquisition of relevant information for the current MVP, it means that the current MVP may be selected as the optimal MVP. Therefore, we need to continue to acquire the second relevant information for the current MVP to provide data support for the selection of the optimal MVP.

[0161] In one embodiment of this application, obtaining the second relevant information of the current MVP includes:

[0162] If the current inter-frame prediction mode is the combined reference frame mode, then obtain the optimal combined mode type, optimal interpolation method, and optimal motion mode corresponding to the current MVP.

[0163] As mentioned earlier, if the current prediction mode is a combined reference frame mode, then the optimal one needs to be selected from the four types of combined reference frame modes as the optimal combined mode type. The optimal interpolation method can be obtained through... Figure 5 In the embodiment, the bestMv interpolation method is selected through a step, and the optimal motion mode can be obtained through... Figure 5 The step of selecting the optimal motion pattern in the embodiment is obtained.

[0164] In one embodiment of this application, obtaining the second relevant information of the current MVP includes:

[0165] If the current inter-frame prediction mode is single reference frame mode, then obtain the optimal interpolation method and optimal motion mode corresponding to the current MVP.

[0166] It is worth mentioning that, in addition to the optimal combination mode type, optimal interpolation method, and optimal motion mode, the second relevant information may also include information calculated by obtaining the optimal interpolation method and optimal motion mode, as well as distortion, residual coefficient bit consumption, transformation type bit consumption, TU segmentation type bit consumption, interpolation method bit consumption, motion mode bit consumption, etc.

[0167] In one embodiment of this application, after obtaining the second relevant information of the current MVP, the method further includes:

[0168] If all relevant information corresponding to the current inter-frame prediction mode has been obtained, then the MVP corresponding to other inter-frame prediction modes is obtained as the current MVP, and relevant information corresponding to the current MVP is obtained.

[0169] Video encoding requires prediction mode and MVP selection. Therefore, for each inter-frame prediction mode, the MVP and related information must be obtained.

[0170] Please continue to refer to the following. Figure 6 In step 630, the optimal MVP corresponding to the current coding block is selected based on the first and second related information of the multiple MVPs corresponding to the current coding block.

[0171] Figure 7 A flowchart illustrating the selection of the optimal MVP corresponding to the current coding block according to an embodiment of this application is shown. Please refer to... Figure 7 As shown, the process of selecting the optimal MVP corresponding to the current coding block may include the following steps:

[0172] Step 710: Based on the relevant information of each MVP corresponding to the current coding block, determine the optimal MVP corresponding to each combination of inter-frame prediction mode and reference frame, and use them as candidate MVPs.

[0173] Each combination of inter-frame prediction mode and reference frame corresponds to an optimal MVP.

[0174] For example, the combination of the current inter-frame prediction mode and the current reference frame corresponds to several MVPs, and the optimal MVP is selected from these MVPs.

[0175] The optimal MVP can be selected by calculating the rate-distortion cost corresponding to each MVP.

[0176] Step 720: Based on the relevant information of each candidate MVP, determine the optimal reference frame corresponding to each inter-frame prediction mode.

[0177] Based on the relevant information of an inter-frame prediction mode and candidate MVPs corresponding to different reference frames, the optimal reference frame corresponding to the inter-frame prediction mode is determined.

[0178] The optimal reference frame can be selected by calculating the rate-distortion cost of each reference frame. The rate-distortion cost of the reference frame needs to be calculated based on the relevant information of the corresponding candidate MVP. For example, the relevant information of the candidate MVP corresponding to the reference frame may include the number of bits consumed by the reference frame.

[0179] Step 730: Determine the optimal inter-frame prediction mode corresponding to the current coding block based on the relevant information of the candidate MVPs corresponding to each optimal reference frame.

[0180] The optimal inter-frame prediction mode can be selected by calculating the rate-distortion cost of each inter-frame prediction mode. The rate-distortion cost of the inter-frame prediction mode needs to be calculated based on the relevant information of the candidate MVP corresponding to the optimal reference frame.

[0181] Step 740: Determine the optimal MVP corresponding to the current coding block based on the candidate MVP corresponding to the optimal inter-frame prediction mode.

[0182] It should be pointed out that, Figure 7 This is merely one embodiment of this application. The selection of the optimal MVP may include more than just... Figure 7 The steps shown in the illustrated embodiment may also include other processes, such as selecting the optimal CU segmentation type and selecting the optimal intra-frame prediction mode and inter-frame prediction mode. The final selected optimal MVP corresponding to the current coding block is the MVP corresponding to the optimal combination of motion vector, reference frame, and prediction mode.

[0183] Please continue to refer to the following. Figure 6 In step 640, the current coding block is encoded based on the optimal MVP.

[0184] The encoding process for the current block is based on the optimal MVP and related information such as the prediction mode, motion vector, reference frame, interpolation method, and motion mode corresponding to the optimal MVP. This can achieve better compression performance and reduce information redundancy in the video.

[0185] Below, in conjunction with Figure 9 The embodiments of this application are further described below. Figure 9 A flowchart illustrating the optimal result selection process corresponding to any combination of prediction mode and reference frame according to an embodiment of this application is shown. Please refer to... Figure 9 Specifically, it includes the following steps:

[0186] In step 310, N = 0, and the number of MVPs is obtained (ref_set).

[0187] Among them, N is used for counting, and ref_set is the number of mvps corresponding to the current inter-frame prediction mode and the current reference frame.

[0188] In step 320, it is judged whether N < ref_set holds. If so, step 330 is executed; otherwise, step 340 is executed.

[0189] In step 330, obtain mvp, N + 1.

[0190] That is, obtain mvp as the current mvp and increment N by 1.

[0191] In step 340, obtain the relevant information corresponding to the predicted motion vector value for the next prediction mode.

[0192] In step 350, it is judged whether the current prediction mode contains NEWMV. If so, first execute step 360, and then execute step 370; otherwise, directly execute step 370.

[0193] The modes NEWMV, NEW_NEWMV, NEAREST_NEWMV, NEW_NEARESTMV, NEAR_NEWMV, NEW_NEARMV all contain NEWMV. For example, NEW_NEWMV means both directions contain NEWMV, NEAREST_NEWMV means the backward direction contains NEWMV, and so on for the others.

[0194] In step 360, perform motion estimation. <s

[0195] If the current prediction mode contains NEWMV, motion estimation needs to be performed to search for the optimal motion vector corresponding to the current mvp.

[0196] There are many motion estimation methods, which are divided into two parts: integer-pixel motion estimation and fractional-pixel motion estimation. For example, integer-pixel motion estimation can adopt methods such as TZ search, nstep, diamond, hexagon, etc., and fractional-pixel can adopt methods such as diamond, full search, etc.

[0197] Figure 10 Shows a schematic diagram of a diamond search template according to an embodiment of the present application; Figure 11 Shows a schematic diagram of a two-point search according to an embodiment of the present application; Figure 12 Shows a partial schematic diagram of searching for position points in a raster scan manner according to an embodiment of the present application.

[0198] Please refer to Figures 10-12 , and a specific process of TZ search can be like this:

[0199] (1) Determine the search starting point.

[0200] The current MVP is used as the starting point for the search. There is also the position (0,0). The rate-distortion cost of the corresponding motion vectors is compared. The motion vector with the smaller cost is used as the final starting point for the search.

[0201] (2) Starting with a step size of 1, according to... Figure 10 The diamond-shaped search template shown searches within the search window, with the step size increasing in integer powers of 2, selecting the point with the minimum rate-distortion cost as the search result for this step.

[0202] (3) If the step size corresponding to the optimal point obtained in step (2) is 1, then start the 2-point search, according to Figure 11 Add points around the current point that have not yet been searched: add two points AC at position 1, two points BD at position 3, two points EG at position 6, and two points FH at position 8. Other positions, such as the four positive directions (left, right, up, and down) of positions 2, 4, 5, and 7, have already been calculated and do not need to be added.

[0203] (4) If the step size corresponding to the optimal point obtained in step (3) is greater than 5, then start searching for all points every 5 rows and 5 columns using a raster scan method, such as... Figure 12 As shown.

[0204] (5) Using the optimal point obtained in step (4) as the new search starting point, repeat steps (2) to (3), each time using the new optimal point as the search starting point, until the starting point and the optimal point obtained by the search no longer change. The MV obtained at this time is recorded as the optimal motion vector of the integer pixel motion estimation.

[0205] The diamond search algorithm, also known as diamond search, has two different matching templates: large diamond and small diamond.

[0206] Figure 13 A schematic diagram of a large diamond search template according to an embodiment of this application is shown; Figure 14 A schematic diagram of a small diamond search template according to an embodiment of this application is shown.

[0207] Please see Figure 13 and Figure 14 As can be seen, the large rhombus has 9 search points, while the small rhombus only has 5. The rhombus search first uses the large rhombus search template with a larger step size for a coarse search, and then uses the small rhombus search template for a fine search. The search steps are as follows:

[0208] Step 1: First, using the center point of the search window as the center and the large diamond search template as the template, calculate the rate-distortion cost of the center point and the eight points around it, for a total of nine points, and compare them to find the point with the smallest rate-distortion cost.

[0209] Step 2: If the center point of the search is the point with the minimum rate-distortion cost, then jump to step 3 and use the small diamond search template; otherwise, return to step 1 for the search.

[0210] Step 3: Using a small diamond search template with only 5 search points, calculate the rate-distortion cost of these 5 points, and take the point with the smallest rate-distortion cost as the best matching point, i.e. the optimal motion vector.

[0211] The motion estimation used in step 360 can be arbitrary, and this application embodiment does not impose any restrictions on it.

[0212] In one embodiment of this application, the method further includes:

[0213] Determine the corresponding motion estimation method based on the encoded speed gear command;

[0214] Motion estimation is performed according to the determined motion estimation method.

[0215] The encoding speed setting command can be a command submitted by the user in real time, or a command preset in the program.

[0216] In this embodiment, the encoding speed is freely adjustable according to the encoding speed level instruction, thereby improving the user experience.

[0217] For example, the speed gear instruction can include slow and fast speeds. When encoding slow speeds, the tz search or nstep method is used, while when encoding fast speeds, the diamond search or hexagonal search is used.

[0218] In one embodiment of this application, before determining the corresponding motion estimation method based on the encoded speed gear instruction, the method further includes:

[0219] Get the current device's CPU utilization or device model;

[0220] The corresponding encoding speed level is determined based on the obtained CPU utilization or device model, and the encoding speed level instruction is generated accordingly.

[0221] For example, if the CPU utilization of the current device is too high, or if the encoding performance of the current device is judged to be poor based on the device model, a lower encoding speed setting can be set to reduce the burden on the current device.

[0222] In this embodiment, the encoding speed level is determined based on the current device's CPU utilization or device model, thus allowing the encoding speed level to be matched with the current device's status.

[0223] Please continue to refer to Figure 9In step 370, it is determined whether the condition for early exit is met. If it is, the exit is performed early and step 320 is executed again; otherwise, step 380 is executed.

[0224] The judgment conditions for single reference frame mode and combined reference frame mode are different in this step, and they need to be processed separately. The specific implementation process is as follows:

[0225] Step 1: Determine the entry conditions.

[0226] The condition is true if the current mode is not GLOBALMV or not GLOBAL_GLOBALMV;

[0227] In other cases, if the current mode is GLOBALMV or GLOBAL_GLOBALMV, and the current mode warpmotion type is not rotation, zoom, or affine, then the condition is met.

[0228] Otherwise, the condition is not met.

[0229] If the condition in Step 1 is met, then proceed to Step 2.

[0230] Step 2: Determine the comparison information for the current MVP.

[0231] mv[0] represents the motion vector corresponding to the forward reference frame of the current MVP, mv[1] represents the motion vector corresponding to the backward reference frame of the current MVP, ref_frame[0] represents the forward reference frame of the current MVP, and ref_frame[1] represents the backward reference frame of the current MVP.

[0232] rate_mv represents the number of bits consumed by the current MVP's MVD. If no motion estimation is performed, the number of bits consumed by MVD is 0. head_rate represents the number of bits consumed by the current MVP's reference frame.

[0233] Step 3: Determine the conditions for early exit.

[0234] This step distinguishes between single-reference-frame mode and combined-reference-frame mode. The combined-reference-frame mode requires comparing two reference frames and two motion vectors. Details are as follows:

[0235] If the current mode is single-reference frame mode, then the MVP information collected in the single-reference frame mode is traversed, and the current MVP information is compared with the information of each MVP. If the motion vector and the reference frame are the same, and the rate value of the current MVP is larger, then the condition for early exit is met. Specifically, as follows: Figure 15 As shown.

[0236] Figure 15A schematic diagram illustrating the process for determining early exit conditions in a single-reference-frame mode according to an embodiment of this application is shown. Please refer to... Figure 15 The process for determining the early exit condition in single reference frame mode includes the following steps:

[0237] Step 1510, initial t=0, skip=0.

[0238] Initialize t and skip to 0.

[0239] Step 1520: Obtain the current MVP information.

[0240] Step 1530, t < single_mode_mvp_num.

[0241] That is, determine whether t < single_mode_mvp_num is true. If it is true, proceed to step 1540; otherwise, proceed to step 1550.

[0242] single_mode_mvp_num represents the number of MVPs in the single reference frame mode that the current coding unit has collected.

[0243] Step 1550, end the judgment.

[0244] If the early exit condition is met or the value of t reaches single_mode_mvp_num, the judgment ends.

[0245] Step 1540, mv[0].as_int==single_mode_mvp_info[t].mv.as_int&&ref_frame[0]==single_mode_mvp_info[t].ref&&rate_mv+head_rate>=single_mode_mvp_info[t].rate.

[0246] That is, determine whether `mv[0].as_int == single_mode_mvp_info[t].mv.as_int && ref_frame[0] == single_mode_mvp_info[t].ref && rate_mv + head_rate >= single_mode_mvp_info[t].rate` is true. If it is true, proceed to step 1560; otherwise, proceed to step 1570.

[0247] In this step, the motion vector, reference frame, and a portion of the bit count of the current MVP are compared with the corresponding information of the MVPs that have been collected.

[0248] Step 1560, skip=1.

[0249] Set the value of skip to 1. skip = 1 means that the condition for early exit is met. After executing step 1560, execute step 1550.

[0250] Step 1570, t = t + 1.

[0251] Increment t by 1.

[0252] After step 1570, step 1530 is executed again.

[0253] If the current mode is the combined reference frame mode, then the MVP information collected by traversing the combined modes is compared with each MVP information. If the two motion vectors and two reference frames are the same, and the rate value of the current MVP is larger, then the early exit condition is met. Specifically, as follows... Figure 16 As shown.

[0254] Figure 16 A schematic diagram illustrating the process for determining early exit conditions in a combined reference frame mode according to an embodiment of this application is shown. Please refer to... Figure 16 The process for determining the early exit condition in combined reference frame mode includes the following steps:

[0255] Step 1610, initial t=0, skip=0.

[0256] Initialize t and skip to 0.

[0257] Step 1620: Obtain the current MVP information.

[0258] Step 1630, t < comp_mode_mvp_num.

[0259] That is, determine whether t < comp_mode_mvp_num is true. If it is true, proceed to step 1640; otherwise, proceed to step 1650.

[0260] comp_mode_mvp_num represents the number of MVPs that the current coding unit has collected for the combined reference frame modes.

[0261] Step 1650, end the judgment.

[0262] If the early exit condition is met or the value of t reaches comp_mode_mvp_num, the judgment ends.

[0263] Step 1640, mv[0].as_int==comp_mode_mvp_info[t].mv[0].as_int&&ref_frame[0]==comp_mode_mvp_info[t].ref[0]&&mv[1].as_int==comp_ mode_mvp_info[t].mv[1].as_int&&ref_frame[1]==comp_mode_mvp_info[t].ref[1]&&rate_mv+head_rate>comp_mode_mvp_info[t].rate.

[0264] That is, determine whether `mv[0].as_int == comp_mode_mvp_info[t].mv[0].as_int && ref_frame[0] == comp_mode_mvp_info[t].ref[0] && mv[1].as_int == comp_mode_mvp_info[t].mv[1].as_int && ref_frame[1] == comp_mode_mvp_info[t].ref[1] && rate_mv + head_rate > comp_mode_mvp_info[t].rate` is true. If it is true, proceed to step 1660; otherwise, proceed to step 1670.

[0265] In this step, the forward reference frame, the motion vector corresponding to the forward reference frame, the backward reference frame, the motion vector corresponding to the backward reference frame, and a portion of the bit count are compared with the corresponding information of the collected MVPs.

[0266] Step 1660, skip=1.

[0267] Set the value of skip to 1. skip = 1 means that the condition for early exit is met. After executing step 1660, execute step 1650.

[0268] Step 1670, t = t + 1.

[0269] Increment t by 1.

[0270] After step 1670, step 1630 is executed again.

[0271] It should be pointed out that, although Figure 15 and Figure 16 The middle part of the bit count includes the MVD bit count and the reference frame bit count of the current MVP, but it is easy to understand that the middle part of the bit count may also include the MVP index bit count.

[0272] Step 4: Exit early.

[0273] if Figure 15 and Figure 16 If skip is equal to 1, the game will exit early.

[0274] For single-reference frame mode, early exit will skip interpolation method selection and motion mode selection.

[0275] For combined reference frame mode, early exit will skip the combination mode type selection, interpolation method selection, and motion mode selection.

[0276] Please continue to refer to Figure 9 In step 380, information is collected.

[0277] This step involves collecting the first relevant information for the current MVP, specifically used to store the motion vector, reference frame, and number of bits corresponding to each MVP within the current coding unit.

[0278] Within a coding unit, there can be a maximum of 56 MVPs for a single reference frame mode (7x3+7x3+7+7), and a maximum of 320 MVPs for a combined reference frame mode. The two must be stored separately.

[0279] Specifically, for motion vectors: if the current mode contains NEWMV, motion estimation is required, and the stored motion vector is the optimal motion vector after motion estimation; if the current mode does not contain NEWMV, the stored motion vector is the current MVP.

[0280] For reference frames: If the current mode is single reference frame mode, only the forward reference frame needs to be stored; if the current mode is combined reference frame mode, both the forward and backward reference frames need to be stored.

[0281] For bit count: This refers to the total number of bits consumed before interpolation, including the bits consumed by the reference frame, the bits consumed by the MVP index, and the bits consumed by MVD. If no motion estimation is performed, the MVD bit count is 0.

[0282] The specific implementation process for this step is as follows:

[0283] Step 1: Define the data structure used to store the data.

[0284] Define a structure MV to store the x-coordinates and y-coordinates of the motion vector.

[0285]

[0286] Among them, int16_t defines x and y as 16-bit unsigned short integers, and uint32_t defines asint as 32-bit unsigned integers.

[0287] Define a structure called single_mode_mv_ref to store the motion vector, reference frame, and number of bits for single-reference frame mode.

[0288]

[0289] Here, int is used to define ref, and rate is an integer.

[0290] Define a structure `comp_mode_mv_ref` to store the motion vector, reference frame, and bit count of the combined reference frame mode.

[0291]

[0292] Among them, int is used to define ref[2], and rate is an integer.

[0293] Define an array to store the data and record the number of data points collected.

[0294] int single_mode_mvp_num;

[0295] single_mode_mv_ref single_mode_mvp_info

[56] ;

[0296] int comp_mode_mvp_num;

[0297] comp_mode_mv_ref comp_mode_mvp_info

[320] ;

[0298] Wherein, `int` is used to define the data types of `single_mode_mvp_num` and `comp_mode_mvp_num` as integers. `single_mode_mvp_num` represents the number of MVPs for the single reference frame mode that the current coding unit has collected. `single_mode_mvp_info` is used to store the motion vectors, reference frames, and bit counts for the single reference frame mode. The maximum number of MVPs corresponding to a single reference frame mode is 56, i.e., `nearestmv(7) + nearmv(7*3) + globalmv(7) + newmv(7*3)`. `comp_mode_mvp_num` represents the number of MVPs for the combined reference frame mode that the current coding unit has collected. `comp_mode_mvp_info` is used to store the motion vectors, reference frames, and bit counts for the combined reference frame mode. The maximum number of MVPs is 320, specifically:

[0299] nearest_nearestmv(16)+global_globalmv(16)+near_nearmv(16*3)+new_newmv(16*3)+nearest_newmv(16*3)+new_nearestmv(16*3)+near_newmv(16*3)+new_nearmv(16*3).

[0300] Initialize the count to start from 0:

[0301] single_mode_mvp_num = 0;

[0302] comp_mode_mvp_num = 0.

[0303] Step 2: Confirm the conditions for collecting information.

[0304] If the current MVP does not exit prematurely, and the mode is not GLOBALMV or not GLOBAL_GLOBALMV, then the condition is met;

[0305] Otherwise, if the current MVP does not exit prematurely, and the mode is GLOBALMV or GLOBAL_GLOBALMV, and the current mode warp motion type is not rotation, not zoom, not affine, then the condition is met.

[0306] Otherwise, the condition is not met.

[0307] Step 3: If the conditions are met, save the data.

[0308] If it is a single reference frame mode, the data is saved in the following way:

[0309] single_mode_mvp_info[single_mode_mvp_num].mv.as_int=mv[0].as_int;

[0310] single_mode_mvp_info[single_mode_mvp_num].ref=ref_frame[0];

[0311] single_mode_mvp_info[single_mode_mvp_num].rate=head_rate+rate_mv;

[0312] single_mode_mvp_num++;

[0313] Otherwise, if it is a combined reference frame mode, the data is saved in the following way:

[0314] comp_mode_mvp_info[comp_mode_mvp_num].mv[0].as_int=mv[0].as_int;

[0315] comp_mode_mvp_info[comp_mode_mvp_num].mv[1].as_int=mv[1].as_int;

[0316] comp_mode_mvp_info[comp_mode_mvp_num].ref[0]=ref_frame[0];

[0317] comp_mode_mvp_info[comp_mode_mvp_num].ref[1]=ref_frame[1];

[0318] comp_mode_mvp_info[comp_mode_mvp_num].rate=rate_mv+head_rate+1;

[0319] comp_mode_mvp_num++;

[0320] Both single_mode_mvp_num and comp_mode_mvp_num start counting from 0. The count increases by 1 for each new MVP's first related information added. Therefore, this value indicates how many MVPs' first related information have been stored.

[0321] Similar to the above, mv[0] represents the motion vector corresponding to the forward reference frame of the current MVP, mv[1] represents the motion vector corresponding to the backward reference frame of the current MVP, ref_frame[0] represents the forward reference frame of the current MVP, ref_frame[1] represents the backward reference frame of the current MVP, rate_mv represents the number of bits consumed by the current MVP's MVD (if no motion estimation is performed, the number of bits consumed by MVD is 0); head_rate represents the number of bits consumed by the reference frame of the current MVP. In addition, it is also necessary to collect the number of bits consumed by the MVP index.

[0322] Bit count estimation is related to the context model of entropy coding and is defined by the coding protocol. Different coding protocol standards may have different entropy coding models, but they will all estimate the corresponding number of bits.

[0323] Please continue to refer to Figure 9In step 390, it is determined whether the current prediction mode is a combined reference frame mode. If so, step 3100 is executed first, and then step 3110 is executed; otherwise, step 3110 is executed directly.

[0324] In step 3100, type selection is performed.

[0325] There are four types of combined reference frame modes: AVERAGE, DISTWTD, WEDGE, and DIFFWTD. This step involves selecting the optimal one from these four types as the optimal combined mode type. The predicted pixels of the two reference frames are then fused together. Each combined mode type corresponds to a predicted pixel fusion method, which is specified in the AV1 protocol and will not be detailed further.

[0326] In step 3110, the best interpolation method is selected.

[0327] The purpose of interpolation is that if the optimal motion vector contains sub-pixels, the predicted pixel cannot be obtained directly. It is necessary to first obtain the reference data corresponding to the integer pixel position of the optimal motion vector, and then interpolate based on the sub-pixel coordinates to finally obtain the predicted pixel.

[0328] The interpolation calculation process involves first performing horizontal interpolation, followed by vertical interpolation. AV1 has designed three interpolation methods for pixel-level interpolation: REG, SMOOTH, and SHARP. All filter kernels have 8 taps. The main difference between the three interpolation methods is the different coefficients of the filter kernels.

[0329] Because horizontal and vertical can be combined arbitrarily, a total of 9 interpolation methods are obtained, namely: REG_REG, REG_SMOOTH, REG_SHARP, SMOOTH_REG, SMOOTH_SMOOTH, SMOOTH_SHARP, SHARP_REG, SHARP_SMOOTH, and SHARP_SHARP.

[0330] This step iterates through 9 interpolation methods and estimates the rate distortion cost. The interpolation method corresponding to the minimum rate distortion cost is the optimal interpolation method. This step selects the optimal interpolation method from the 9 methods.

[0331] In step 3120, the motion mode is selected.

[0332] After completing step 3120, re-execute step 320.

[0333] The motion modes corresponding to single-reference frame mode and combined reference frame mode are different. Single-reference frame mode corresponds to 4 motion modes: SIMPLE, OBMC, WARPED, and SIMPLE(inter_intra); combined reference frame mode corresponds to only one motion mode: SIMPLE.

[0334] The optimal motion mode is written into the bitstream, telling the decoder which motion mode to use to recover and reconstruct the data during decoding. inter_intra and SIMPLE are both SIMPLE modes, but they are very different. During decoding, the reference frame information in the syntax can be used to determine whether it is the first SIMPLE mode or inter_intra. Since the same flag is used, one bit can be saved.

[0335] All four motion modes incur a full rate distortion penalty, which involves the entire reconstruction process of transform, quantization, inverse quantization, and inverse transform. The difference between the four motion modes lies in the method of acquiring the predicted pixels. The specific process of this step is described below:

[0336] (1) Obtain the predicted pixels.

[0337] For SIMPLE mode: the predicted value obtained after interpolating the predicted pixels.

[0338] For OBMC mode: The predicted pixels obtained after interpolation are further processed. The predicted pixels of neighboring blocks are obtained based on the mv of neighboring blocks, and then fused with the predicted value of the current block after interpolation according to certain rules to obtain a new predicted value.

[0339] For WARPED mode: An affine transformation mv is constructed using the three available positions (left, top, and top right), followed by a small-range motion search. Finally, interpolation is performed to obtain the predicted pixels.

[0340] For SIMPLE (inter_intra) mode: The predicted pixels obtained after interpolation are processed a second time. First, intra-frame prediction is performed for four intra-frame modes: DC, V, H, and SMOOTH, to obtain the optimal intra-frame predicted pixel. Then, the intra-frame and inter-frame predicted pixels are fused to obtain a new predicted value.

[0341] (2) Calculation of integrity rate distortion.

[0342] Based on the input and predicted pixels, residual pixels are obtained. Then, transform and TU depth partitioning are performed to obtain the distortion, residual bit consumption, transform type bit consumption, and TU segmentation type bit consumption. Combined with the previously obtained reference frame bit consumption, MVP index bit consumption, MVD bit consumption, interpolation method bit consumption, and motion mode bit consumption, the rate-distortion cost is obtained, which is calculated using the following formula:

[0343] rdcost = dist + rate × λ

[0344] Where dist represents distortion, rate is the sum of all bits consumed by the current MVP, and λ is the Lagrange daily number.

[0345] By execution Figure 9 The steps shown in the example can obtain the optimal MVP index, optimal motion vector, optimal interpolation method, optimal motion mode, optimal transform type, and optimal TU partitioning type corresponding to the current prediction mode and the current reference frame; then, by comparing the rate-distortion cost (rdcost) of different reference frames, the optimal reference frame is found; and then, by comparing different mode combinations, the optimal prediction information of the current coding block is found.

[0346] The video coding method provided in this application collects the motion vector, reference frame information, and bit consumption of each MVP. During prediction, it pre-determines whether the current MVP is useful, thus skipping interpolation method selection and motion mode selection, thereby significantly reducing the computational load of video coding. This scheme does not add new computations, has very high accuracy, can speed up the process by more than 6%, and maintains lossless compression performance. Therefore, it can significantly improve the efficiency of video coding, save overall computational resources, reduce the performance requirements of the machine, and enable even lower-performance machines to perform video coding quickly, thus expanding the application scope of the coding protocol.

[0347] The following describes an apparatus embodiment of this application, which can be used to execute the entity risk identification method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the video encoding method described above.

[0348] Figure 17 A block diagram of a video encoding apparatus according to an embodiment of this application is shown.

[0349] Reference Figure 17 As shown, a video encoding apparatus 1700 according to an embodiment of this application includes: a comparison unit 1710, a skip unit 1720, a selection unit 1730, and an encoding unit 1740.

[0350] The comparison unit 1710 compares the first relevant information of the current motion vector prediction value (MVP) with the first relevant information of the collected MVPs to obtain a comparison result. The current motion vector prediction value and the collected MVPs are the MVPs corresponding to the current coding block. The skip unit 1720 skips the acquisition of the second relevant information of the current MVP if it is determined from the comparison result that the acquisition of the relevant information of the current MVP can be exited in advance, and then acquires the first and second relevant information of other MVPs after the current MVP in sequence. The selection unit 1730 selects the optimal MVP corresponding to the current coding block based on the first and second relevant information of the multiple MVPs corresponding to the current coding block. The encoding unit 1740 encodes the current coding block based on the optimal MVP.

[0351] In some embodiments of this application, based on the foregoing scheme, the selection unit 1730 is further configured to: determine the optimal MVP corresponding to each combination of inter-frame prediction mode and reference frame as candidate MVPs based on the relevant information of each MVP corresponding to the current coding block; determine the optimal reference frame corresponding to each inter-frame prediction mode based on the relevant information of each candidate MVP; determine the optimal inter-frame prediction mode corresponding to the current coding block based on the relevant information of each candidate MVP corresponding to the optimal reference frame; and determine the optimal MVP corresponding to the current coding block based on the candidate MVP corresponding to the optimal inter-frame prediction mode.

[0352] In some embodiments of this application, based on the foregoing scheme, the skipping unit 1720 is further configured to: if it is determined according to the comparison result that the acquisition of relevant information of the current MVP should not be prematurely terminated, then continue to collect the first relevant information of the current MVP; after collecting the first relevant information of the current MVP, acquire the second relevant information of the current MVP.

[0353] In some embodiments of this application, based on the aforementioned scheme, the skip unit 1720 is configured to: if the current inter-frame prediction mode is a combined reference frame mode, then obtain the optimal combined mode type, optimal interpolation method and optimal motion mode corresponding to the current MVP.

[0354] In some embodiments of this application, based on the foregoing scheme, the skipping unit 1720 is further configured to: if all relevant information corresponding to the current inter-frame prediction mode has been obtained, then obtain the MVP corresponding to other inter-frame prediction modes as the current MVP, and continue to obtain relevant information corresponding to the current MVP.

[0355] In some embodiments of this application, based on the foregoing scheme, the first relevant information corresponding to the current motion vector prediction value MVP includes: the motion vector corresponding to the current MVP, the current reference frame corresponding to the current MVP, and a portion of the bit count corresponding to the current MVP.

[0356] In some embodiments of this application, based on the foregoing scheme, the inter-frame prediction mode corresponding to the current MVP includes NEWMV, and the skip unit 1720 is further configured to: perform motion estimation based on the current MVP to obtain the optimal motion vector corresponding to the current MVP; and determine the optimal motion vector as the motion vector corresponding to the current MVP.

[0357] In some embodiments of this application, based on the foregoing scheme, the inter-frame prediction mode corresponding to the current MVP does not include NEWMV, and the skip unit 1720 is further configured to: determine the current MVP as the motion vector corresponding to the current MVP.

[0358] In some embodiments of this application, based on the foregoing scheme, the skip unit 1720 is configured as follows: if, according to the comparison result, it is determined that the motion vector of the current MVP is the same as the motion vector of the collected MVP, the reference frame of the current MVP is the same as the reference frame of the collected MVP, and the partial bit count of the current MVP is greater than the partial bit count of the collected MVP, then it is determined that it is necessary to exit the acquisition of relevant information of the current MVP in advance.

[0359] In some embodiments of this application, based on the foregoing scheme, the skip unit 1720 is configured as follows: if, according to the comparison result, it is determined that the forward reference frame of the current MVP is the same as the forward reference frame of the collected MVP, the backward reference frame of the current MVP is the same as the backward reference frame of the collected MVP, the motion vector corresponding to the forward reference frame of the current MVP is the same as the motion vector corresponding to the forward reference frame of the collected MVP, the motion vector corresponding to the backward reference frame of the current MVP is the same as the motion vector corresponding to the backward reference frame of the collected MVP, and the partial bit count of the current MVP is greater than the partial bit count of the collected MVP, then it is determined that it is necessary to exit the acquisition of relevant information of the current MVP in advance.

[0360] In some embodiments of this application, based on the foregoing scheme, the comparison unit 1710 is further configured to: determine whether the current inter-frame prediction mode meets a predetermined condition, wherein the predetermined condition is generated based on the inter-frame prediction modes of GLOBALMV and GLOBAL_GLOBALMV; and if the current inter-frame prediction mode meets the predetermined condition, perform a process of comparing the first relevant information of the current motion vector prediction value MVP obtained with the first relevant information of the collected MVP.

[0361] In some embodiments of this application, based on the foregoing scheme, the skip unit 1720 is configured to: if it is determined according to the comparison result that the acquisition of relevant information of the current MVP should not be prematurely terminated and the current inter-frame prediction mode meets the predetermined conditions, then the first relevant information of the current MVP is collected.

[0362] Figure 18 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0363] It should be noted that, Figure 18 The computer system 1800 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0364] like Figure 18 As shown, the computer system 1800 includes a Central Processing Unit (CPU) 1801, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1802 or programs loaded from storage portion 1808 into Random Access Memory (RAM) 1803, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 1803. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via bus 1804. An Input / Output (I / O) interface 1805 is also connected to bus 1804.

[0365] The following components are connected to I / O interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to I / O interface 1805 as needed. Removable media 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1810 as needed so that computer programs read from them can be installed into storage section 1808 as needed.

[0366] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by central processing unit (CPU) 1801, it performs various functions defined in the system of this application.

[0367] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0368] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0369] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0370] In one aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0371] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0372] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0373] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0374] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A video encoding method, characterized in that, include: The first relevant information of the current motion vector prediction value MVP is compared with the first relevant information of the collected MVPs to obtain the comparison result. The current motion vector prediction value and the collected MVPs are the MVPs corresponding to the current coding block. If it is determined from the comparison results that it is possible to exit the acquisition of relevant information of the current MVP in advance, then the acquisition of the second relevant information of the current MVP is skipped, and the first and second relevant information of other MVPs after the current MVP are acquired in sequence. Based on the first and second related information of the multiple MVPs corresponding to the current coding block, select the optimal MVP corresponding to the current coding block; The current encoding block is encoded based on the optimal MVP.

2. The video encoding method according to claim 1, characterized in that, The step of selecting the optimal MVP corresponding to the current coding block based on the first and second related information of multiple MVPs corresponding to the current coding block includes: Based on the relevant information of each MVP corresponding to the current coding block, determine the optimal MVP corresponding to each combination of inter-frame prediction mode and reference frame, and use them as candidate MVPs; Based on the relevant information of each candidate MVP, determine the optimal reference frame corresponding to each inter-frame prediction mode; Based on the relevant information of the candidate MVPs corresponding to each of the optimal reference frames, the optimal inter-frame prediction mode corresponding to the current coding block is determined; The optimal MVP corresponding to the current coding block is determined based on the candidate MVP corresponding to the optimal inter-frame prediction mode.

3. The video encoding method according to claim 2, characterized in that, After comparing the first relevant information of the current motion vector prediction value MVP with the first relevant information of the collected MVP, the method further includes: If, based on the comparison results, it is determined not to prematurely exit the acquisition of relevant information for the current MVP, then the first relevant information for the current MVP continues to be collected. After collecting the first relevant information of the current MVP, the second relevant information of the current MVP is obtained.

4. The video encoding method according to claim 3, characterized in that, The step of obtaining the second relevant information of the current MVP includes: If the current inter-frame prediction mode is the combined reference frame mode, then obtain the optimal combined mode type, optimal interpolation method, and optimal motion mode corresponding to the current MVP.

5. The video encoding method according to claim 3, characterized in that, After obtaining the second relevant information of the current MVP, the method further includes: If relevant information corresponding to all MVPs of the current inter-frame prediction mode has been obtained, then the MVPs corresponding to other inter-frame prediction modes are obtained as the current MVP, and relevant information corresponding to the current MVP is obtained.

6. The video encoding method according to claim 1, characterized in that, The first relevant information corresponding to the current motion vector prediction value (MVP) includes: the motion vector corresponding to the current MVP, the current reference frame corresponding to the current MVP, and a portion of the bit count corresponding to the current MVP.

7. The video encoding method according to claim 6, characterized in that, The inter-frame prediction mode corresponding to the current MVP includes NEWMV, and the motion vector corresponding to the current MVP is obtained through the following process: Motion estimation is performed based on the current MVP to obtain the optimal motion vector corresponding to the current MVP; The optimal motion vector is determined as the motion vector corresponding to the current MVP.

8. The video encoding method according to claim 6, characterized in that, The inter-frame prediction mode corresponding to the current MVP does not include NEWMV, and the motion vector corresponding to the current MVP is obtained through the following process: The current MVP is determined as the motion vector corresponding to the current MVP.

9. The video encoding method according to claim 1, characterized in that, The video encoding method further includes: If, based on the comparison results, it is determined that the motion vector of the current MVP is the same as the motion vector of the collected MVP, the reference frame of the current MVP is the same as the reference frame of the collected MVP, and the number of bits in the current MVP is greater than the number of bits in the collected MVP, then it is determined that the acquisition of relevant information of the current MVP needs to be terminated in advance.

10. The video encoding method according to claim 1, characterized in that, The video encoding method further includes: If, based on the comparison results, it is determined that the forward reference frame of the current MVP is the same as the forward reference frame of the collected MVPs, the backward reference frame of the current MVP is the same as the backward reference frame of the collected MVPs, the motion vector corresponding to the forward reference frame of the current MVP is the same as the motion vector corresponding to the forward reference frame of the collected MVPs, and the motion vector corresponding to the backward reference frame of the current MVP is the same as the motion vector corresponding to the backward reference frame of the collected MVPs, and the partial bit count of the current MVP is greater than the partial bit count of the collected MVPs, then it is determined that the acquisition of relevant information of the current MVP needs to be terminated in advance.

11. The video encoding method according to claim 3, characterized in that, Before comparing the first relevant information of the current motion vector prediction value MVP with the first relevant information of the collected MVP, the method further includes: Determine whether the current inter-frame prediction mode meets a predetermined condition, wherein the predetermined condition is generated based on the inter-frame prediction modes of GLOBALMV and GLOBAL_GLOBALMV. If the current inter-frame prediction mode meets the predetermined conditions, the process of comparing the first relevant information of the current motion vector prediction value MVP obtained with the first relevant information of the collected MVP is executed.

12. The video encoding method according to claim 11, characterized in that, If, based on the comparison result, it is determined not to prematurely exit the acquisition of relevant information for the current MVP, then the first relevant information for the current MVP continues to be collected, including: If, based on the comparison results, it is determined that the acquisition of relevant information of the current MVP will not be prematurely terminated and the current inter-frame prediction mode meets the predetermined conditions, then the first relevant information of the current MVP will continue to be collected.

13. A video encoding device, characterized in that, include: The comparison unit is used to compare the first relevant information of the current motion vector prediction value MVP with the first relevant information of the collected MVPs to obtain a comparison result. The current motion vector prediction value and the collected MVPs are the MVPs corresponding to the current coding block. The skip unit is used to skip the acquisition of the second relevant information of the current MVP if it is determined from the comparison result that the acquisition of the relevant information of the current MVP can be exited in advance, and the first and second relevant information of other MVPs after the current MVP can be acquired in sequence. The selection unit is used to select the optimal MVP corresponding to the current coding block based on the first and second related information of the multiple MVPs corresponding to the current coding block. An encoding unit is used to encode the current encoding block based on the optimal MVP.

14. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video encoding method as described in any one of claims 1 to 12.

15. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the video encoding method as described in any one of claims 1 to 12.