Method, apparatus, device and medium for selecting reference frame

By grading and filtering the candidate reference frames of the target encoding unit in the AV1 video encoding standard, the inefficiency problem caused by its computational complexity is solved, and a faster encoding process is achieved.

CN114286089BActive Publication Date: 2025-06-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111131774.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2025-06-17
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

The calculation process of the optimal reference frame in the AV1 video encoding standard is complicated, resulting in inefficient video encoding.

Method used

By acquiring multiple candidate reference frames of the target encoding unit, scoring them based on quality scoring information, and filtering out the optimal reference frame based on the scoring results.

Benefits of technology

The determination process of the optimal reference frame is simplified, and the cost of calculating the rate distortion of all candidate reference frames is avoided, which significantly speeds up the encoding speed of the target encoding unit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114286089B_ABST
    Figure CN114286089B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device, and storage medium for selecting a reference frame, belonging to the field of video coding. The method includes: obtaining m candidate reference frames of a target coding unit, where the target coding unit is one of multiple coding units of a video frame; m is an integer greater than 1; scoring the m candidate reference frames based on the quality score information of the m candidate reference frames, where the quality score information is used to indicate the coding quality of the target coding unit for inter-frame prediction through the m candidate reference frames; and screening out the optimal reference frame of the target coding unit according to the scoring results of the m candidate reference frames. The above technical solution simplifies the process of determining the optimal reference frame and greatly speeds up the coding speed of the target coding unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video coding, and particularly to a method, apparatus, device, and medium for selecting reference frames. Background Art

[0002] Video coding is a technology for compressing video to reduce the data volume of video files. For example, inter-frame coding is a technology that uses the correlation between video image frames to perform compression coding on the current video frame.

[0003] The inter-frame coding in AV1 (the first-generation video coding standard developed by the Alliance for Open Media) includes 4 single-reference-frame prediction modes and 8 combined-reference-frame prediction modes. Under each single-reference-frame prediction mode, there are 7 reference frames, and under each combined-reference-frame prediction mode, there are 16 reference-frame combinations. For inter-frame coding, there are a total of 28 candidate reference frames and 128 candidate reference-frame combinations. Frame-by-frame optimization is performed within each candidate reference-frame combination, and finally, the optimal reference frame is selected together with the 28 candidate reference frames.

[0004] The calculation process for AV1 to determine the optimal reference frame is very complex, resulting in low video coding efficiency. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for selecting reference frames, which can improve video coding efficiency. The technical solutions are as follows:

[0006] According to one aspect of this application, a method for selecting a reference frame is provided. The method includes:

[0007] Obtain m candidate reference frames of a target coding unit, where the target coding unit is one of multiple coding units of a video frame; m is an integer greater than 1;

[0008] Score the m candidate reference frames based on the quality score information of the m candidate reference frames, where the quality score information is used to indicate the coding quality of the target coding unit for inter-frame prediction through the m candidate reference frames;

[0009] According to the scoring results of the m candidate reference frames, screen out the optimal reference frame of the target coding unit.

[0010] According to another aspect of this application, a device for selecting a reference frame is provided. The device includes:

[0011] An obtaining module, configured to obtain m candidate reference frames of a target coding unit, where the target coding unit is one of multiple coding units of a video frame; m is an integer greater than 1;

[0012] A scoring module, configured to score m candidate reference frames based on the quality scoring information of the m candidate reference frames; the quality scoring information is used to indicate the coding quality of the target coding unit for inter-frame prediction through the m candidate reference frames.

[0013] A screening module, configured to screen out the optimal reference frame of the target coding unit according to the scoring results of the m candidate reference frames.

[0014] According to one aspect of the present application, there is provided a computer device, including: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the method for selecting a reference frame as described above.

[0015] According to another aspect of the present application, there is provided a computer-readable storage medium, the storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the method for selecting a reference frame as described above.

[0016] According to another aspect of the present application, there is provided a computer program product, the computer program product stores computer instructions, the computer instructions are stored in a computer-readable storage medium, the processor reads the computer instructions from the computer-readable storage medium, and the computer instructions are loaded and executed by the processor to implement the method for selecting a reference frame as described above.

[0017] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:

[0018] By scoring the m candidate reference frames before inter-frame prediction and screening out the optimal reference frame for inter-frame prediction according to the scoring results, the rate-distortion cost when calculating all the m candidate reference frames for inter-frame prediction is avoided, the process of determining the optimal reference frame is simplified, and the coding speed of the target coding unit is greatly accelerated. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic diagram of a video coding framework provided by an exemplary embodiment of the present application;

[0021] Figure 2 It is a schematic diagram of the segmentation type of a coding unit provided by an exemplary embodiment of the present application;

[0022] Figure 3It is a schematic diagram of a direction prediction mode provided by an exemplary embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of a method for determining a single reference frame provided by an exemplary embodiment of the present application;

[0024] Figure 5 It is a flowchart of inter-frame prediction based on a predicted motion vector provided by an exemplary embodiment of the present application;

[0025] Figure 6 It is a flowchart of a method for selecting a reference frame provided by an exemplary embodiment of the present application;

[0026] Figure 7 It is a flowchart of calculating the rate-distortion cost of n candidate reference frames provided by an exemplary embodiment of the present application;

[0027] Figure 8 It is a schematic diagram of the positions of a target coding unit and adjacent coding units provided by an exemplary embodiment of the present application;

[0028] Figure 9 It is a schematic diagram of the order of executing the segmentation type of a target coding unit provided by an exemplary embodiment of the present application;

[0029] Figure 10 It is a schematic diagram of the reference relationship within a group of pictures provided by an exemplary embodiment of the present application;

[0030] Figure 11 It is a flowchart of a method for selecting a reference frame provided by an exemplary embodiment of the present application;

[0031] Figure 12 It is a structural block diagram of a reference frame selection device provided by an exemplary embodiment of the present application;

[0032] Figure 13 It is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0033] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0034] To explain the present application, the related technologies related to the present application will be introduced below.

[0035] AV1: The first-generation video coding standard developed by the Alliance for Open Media (AOM). AV1 maintains the traditional coding framework. Schematically, Figure 1It is a schematic diagram of an encoding framework for inter-frame encoding. The process of inter-frame encoding is described as follows:

[0036] First, the current frame 101 is divided into several coding tree units (CTUs) of size 128*128. Each coding tree unit is further divided depthwise to obtain multiple coding units (CUs). Then, for each CU, a prediction is made to obtain a predicted value. Among them, the prediction for each CU includes inter-frame prediction and intra-frame prediction. In inter-frame prediction, first, motion estimation (ME) is obtained between the current frame 101 and the reference frame 102, and then motion compensation (MC) is obtained based on the motion estimation. Based on MC, the predicted value is obtained. The predicted value is subtracted from the input data to obtain a residual. The residual is transformed and quantized to obtain residual coefficients. The residual coefficients are sent to the entropy coding module to output a bitstream. At the same time, after the residual coefficients are inverse quantized and inverse transformed, the residual value of the reconstructed image is obtained. After adding it to the predicted value, the reconstructed image is obtained. After the reconstructed image is filtered, the reconstructed frame 103 is obtained. The reconstructed frame 103 is put into the reference frame queue as the reference frame for the next frame, and thus the encoding is carried out sequentially backward.

[0037] Target coding unit: Each CU contains two prediction coding types: intra-frame encoding and inter-frame encoding. The CUs under each segmentation type are compared among different prediction modes (see the detailed description of intra-frame prediction and inter-frame prediction below) within the same prediction type (intra-frame prediction or inter-frame prediction) to find the optimal prediction mode corresponding to each of the two prediction types. Then, they are compared among different prediction types to find the optimal prediction mode of the target coding unit. At the same time, a TU (Transform Unit) transform is performed on the CU. Each CU corresponds to multiple transform types, and the optimal transform type is found from them. Finally, an image frame is divided into individual CUs.

[0038] In one embodiment, there are ten segmentation types for the target coding unit obtained by depthwise division of the CTU. Schematically, Figure 2Shows the types of CU partitioning, including the NONE type 201, SPLIT type 202, HORZ type 203, VERT type 204, HORZ_4 type 205, HORZ_A type 206, HORZ_B type 207, VERT_A type 208, VERT_B type 209, and VERT_4 type 210. The Coding Tree Unit (CTU) size is 128*128. The CTU is further divided into four equal parts, two equal parts, or one equal part according to the above ten partitioning types. The sub-blocks of the four equal parts can be further recursively divided. Each sub-block is further divided into smaller units according to the nine partitioning types except the NONE type mentioned above until the target coding unit is finally obtained. Optionally, the size of the target coding unit has the following cases: 4*4, 4*8, 8*4, 8*8, 8*16, 16*8, 16*16, 16*32, 32*16, 32*32, 32*64, 64*32, 64*128, 128*64, 128*128, 4*16, 16*4, 8*32, 32*8, 16*64, 64*16.

[0039] Intra prediction: includes the direction prediction mode (assuming that there is directional texture in the video frame, and better matching coding units can be obtained by predicting along the direction). Figure 3 Shows a schematic diagram of the direction prediction mode. In AV1, there are 8 main directions and offset directions based on the 8 main directions. Each of the 8 main directions (203°, horizontal, 157°, 135°, 113°, vertical, 67°, and 45°) has 6 offset angles (±3°, ±6°, and ±9°). That is, the direction prediction mode includes a total of 56 prediction directions. In addition, intra coding also includes the palette coding mode and the copy coding unit coding mode, which will not be elaborated here.

[0040] Inter prediction: includes 4 single-reference frame prediction modes, namely NEARESTMV, NEARMV, GLOBALMV, and NEWMV, and 8 combined-reference frame prediction modes, namely NEAREST_NEARESTMV, NEAR_NEARMV, NEAREST_NEWMV, NEW_NEARESTMV, NEAR_NEWMV, NEW_NEARMV, GLOBAL_GLOBALMV, and NEW_NEWMV.

[0041] Among them, the NEARESTMV mode and the NEARMV mode mean that the motion vector (MV) of the target coding unit is derived from the motion vectors of surrounding coding units, and the motion vector difference (MVD) does not need to be transmitted for inter-frame coding; while NEWMV means that the MVD needs to be transmitted, and the GLOBALMV mode means that the MV information of the predicted coding unit is derived from global motion. The NEARESTMV mode, the NEARMV mode, and the NEWMV mode are related to the derivation of the motion vector prediction (MVP) of the target coding unit.

[0042] Derivation of the MVP of the target coding unit: For a given reference frame, the AV1 standard will calculate 4 mvps according to the rules (this is the content of the AV1 protocol here). Skip-scan the coding units in the left 1 / 3 / 5 columns and the upper 1 / 3 / 5 rows in a certain way, and preferentially select the coding units that use the same reference frame to remove duplicates from the MVs; in the case where the number of non-duplicate MVs in the coding units in the left 1 / 3 / 5 columns and the upper 1 / 3 / 5 rows is less than 8, select the coding units that use the same-direction reference frame and continue to add MVs; in the case where the added MVs are still less than 8, then fill them with the global motion vector; after selecting 8 MVs, sort them according to importance and select the most important 4 MVs. Among them, the 0th MV is NEARESTMV, and the 1st to 3rd correspond to NEARMV. NEWMV uses one of the MVs from 0 to 2 as the MVP. Schematically, as Figure 4 shown, 4 important MVs are selected within a single reference frame, from top to bottom are the 0th MV, the 1st MV, the 2nd MV, and the 3rd MV. In one embodiment, in each single-reference-frame prediction mode, 7 reference frames can be selected, as shown in Table 1 below.

[0043] Table 1

[0044]

[0045] There are 7 reference frames in each of the 4 single-reference-frame prediction modes: LAST_FRAME, LAST2_FRAME, LAST3_FRAME, GOLDEN_FRAME, BWDREF_FRAME, ALTREF2_FRAME, and ALTREF_FRAME respectively.

[0046] Under each of the 8 combined reference frame prediction modes, there are 16 reference frame combinations, namely {LAST_FRAME, ALTREF_FRAME}, {LAST2_FRAME, ALT REF_FRAME}, {LAST3_FRAME, ALTREF_FRAME}, {GOLDEN_FRAME, ALTREF_FRAME}, {LAST_FRAME, BWDREF_FRAME}, {LAST2_FRAME, BWDREF_FRAME}, {LAST3_FRAME, BWDREF_FRAME}, {GOLDEN_FRAME, BWDREF_FRAME}, {LAST_FRAME, ALTREF2_FRAME}, {LAST2_FRAME, ALTREF2_FRAME}, {LAST3_FRAME, ALTREF2_FRAME}, {GOLDEN_FRAME, ALTREF2_FRAME}, {LAST_FRAME, LAST2_FRAME}, {LAST_FRAME, LAST3_FRAME}, {LAST_FRAME, GOLDEN_FRAME}, {BWDR EF_FRAME, ALTREF_FRAME}.

[0047] That is, there are a total of 28 (7 * 4) candidate reference frames and 128 (16 * 8) candidate reference frame combinations corresponding to inter-frame prediction. Each reference frame combination corresponds to a maximum of 3 MVPs, and then motion estimation (only when the prediction mode contains NEWMV will motion estimation be performed), inter-frame optimization, interpolation method optimization, that is, motion mode optimization are carried out for the current MVP. Schematically, Figure 5 The process of inter-frame prediction based on MVP is shown.

[0048] Step 501, set N = 0 and obtain the number of MVPs ref_set;

[0049] That is, set the initial value of N to 0. N represents the Nth MVP, so that each MVP can be traversed.

[0050] Step 502, is N less than ref_set?

[0051] When N is not less than the number of MVPs ref_set, execute step 510; when N is less than ref_set, execute step 503.

[0052] Step 503, obtain the MVP, N = N + 1;

[0053] Obtain the current MVP and execute N = N + 1.

[0054] Step 504, does the prediction mode contain NEWMV?

[0055] When the prediction mode contains NEWMV, execute Step 505; when the prediction mode does not contain NEWMV, execute Step 506.

[0056] Step 505, motion estimation;

[0057] Perform motion estimation on the MVP;

[0058] Step 506, combine reference frame prediction mode?

[0059] When the prediction mode is the combined reference frame prediction mode, execute Step 507; when the prediction mode is not the combined reference frame prediction mode, execute Step 508.

[0060] Step 507, inter-frame decision;

[0061] Select a better-performing reference frame under the current MVP.

[0062] Step 508, select the optimal interpolation method;

[0063] Select the optimal interpolation method under the optimal MVP.

[0064] Step 509, select the optimal motion mode;

[0065] Select the optimal motion mode under the current MVP.

[0066] Step 510, end.

[0067] End the inter-frame prediction.

[0068] Based on the above introduction of related technologies, the computational complexity of encoding the target coding unit is extremely large. Especially in the NEWMV mode, motion estimation is also performed, making the encoding speed of the target coding unit very slow. This application speeds up the encoding of the target coding unit through inter-frame prediction by eliminating some candidate reference frames before inter-frame prediction and not fully executing the prediction mode during the inter-frame prediction process.

[0069] Next, the implementation environment of this application will be introduced:

[0070] Optionally, the method for selecting a reference frame provided by an exemplary embodiment of the present application is applied to a terminal, which includes but is not limited to a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc. Schematically, when the terminal is implemented as a vehicle-mounted terminal, the method provided by the embodiments of the present application can be applied to a vehicle-mounted scenario, that is, a reference frame is selected on the vehicle-mounted terminal as a part of the Intelligent Traffic System (ITS). The intelligent traffic system effectively integrates advanced scientific and technological means (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing, strengthening the connection among vehicles, roads, and users, thereby forming an integrated transportation system that ensures safety, improves efficiency, improves the environment, and saves energy.

[0071] Optionally, the method for selecting a reference frame provided by an exemplary embodiment of the present application is applied to a server, that is, a reference frame is selected by the server, and the encoded bitstream is sent to a terminal or another server. It should be noted that the above server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0072] In some embodiments, the above server can also be implemented as a node in a blockchain system. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. Blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.

[0073] To reduce the computational complexity of inter-frame predictive coding for a target coding unit, Figure 6 The following shows a schematic diagram of the method for selecting a reference frame provided by an exemplary embodiment of the present application. Taking the application of this method to a terminal as an example for elaboration, this method includes:

[0074] Step 620, obtaining m candidate reference frames of the target coding unit;

[0075] The target coding unit is one of multiple coding units of the video frame, and m is an integer greater than 1.

[0076] The video frame is divided into CTUs based on the size of 128*128, and further divided into Figure 2 The type of CU segmentation shown is divided to obtain a target coding unit (CU). Each target coding unit corresponds to m candidate reference frames. Optionally, in the AV1 protocol, each target coding unit corresponds to 7 candidate reference frames. Combined with reference to Table 1, the 7 candidate reference frames are: LAST_FRAME (the forward reference frame closest to the current frame), LAST2_FRAME (the second forward reference frame closest to the current frame), LAST3_FRAME (the third forward reference frame closest to the current frame), GOLDEN_FRAME (long-term reference frame), BWDREF_FRAME (the backward reference frame closest to the current frame), ALTREF2_FRAME (the second backward reference frame closest to the current frame), and ALTREF_FRAME (the third backward reference frame closest to the current frame).

[0077] In one embodiment, the terminal obtains 7 candidate reference frames of the target coding unit.

[0078] Step 640, scoring the m candidate reference frames based on the quality score information of the m candidate reference frames;

[0079] The quality score information is used to indicate the encoding quality of the target coding unit through inter-frame prediction using m candidate reference frames.

[0080] In the above-mentioned related technical introduction, the process of predictive coding by the target coding unit can be summarized as follows: Under each segmentation type of the target coding unit, intra-frame prediction and inter-frame prediction are selected. Under inter-frame prediction, the prediction mode is selected. Inter-frame prediction includes a single reference frame prediction mode and a combined reference frame prediction mode. The single reference frame prediction mode includes four prediction modes: NEARESTMV, NEARMV, NEWMV and GLOBALMV. Each prediction mode corresponds to 7 candidate reference frames (see step 620). The combined reference frame prediction mode includes eight combined reference frame prediction modes (see related technology for details). Combined with the 7 candidate reference frames, there are a total of 16 candidate reference frame combinations. Under each prediction mode, the candidate reference frames (or candidate reference frame combinations) are selected.

[0081] Therefore, based on the above-mentioned optimization under inter-frame prediction, the m candidate reference frames can be scored according to the quality score information of the m candidate reference frames. In this application, the quality score information can be simply summarized into 6 types of quality score information:

[0082] Optimal reference frame information of adjacent coding units;

[0083] · Optimal reference frame information of target coding units of different segmentation types;

[0084] · Distortion degree information of m candidate reference frames;

[0085] · Information of the m candidate reference frames that belong to a preset reference frame set;

[0086] · Optimal reference frame information in the previous prediction mode;

[0087] · Candidate coding unit information of the current frame.

[0088] The above six types of quality scoring information will be elaborated one by one in the following embodiments.

[0089] Based on the above six types of quality scoring information, the terminal scores the m candidate reference frames.

[0090] Optionally, the terminal sets the scoring weights of the six types of quality scoring information to be the same. For example, for each candidate reference frame, each type of quality scoring is uniformly scored using the method of sorce = sorce[0] + weight[0], where sorce[0] is the initial score of the candidate reference frame, and each type of scoring uses the same scoring weight weight[0] to score the candidate reference frame; Optionally, the terminal sets some of the scoring weights of the six types of quality scoring information to be the same. For example, the array weight[6] = {10, 10, 15, 30, 20, 5} is set to represent the scoring weights of the six types of quality scoring information respectively; Optionally, the terminal sets the scoring weights corresponding to the six types of quality scoring information to be completely different.

[0091] Optionally, the terminal sets the initial scores of the seven candidate reference frames to be the same. For example, the array sorce[0, 0, 0, 0, 0, 0, 0] is set to be the initial scores of each candidate reference frame before scoring; Optionally, the terminal sets some of the initial scores of the seven candidate reference frames to be the same; Optionally, the terminal sets the initial scores of the seven candidate reference frames to be completely different. For example, the array sorce[0, 5, 10, 15, 20, 25, 30] is set to represent the initial scores of the seven candidate reference frames respectively.

[0092] It should be noted that the scoring of the m candidate reference frames is implemented based on some or all of the six types of quality scoring information.

[0093] Step 660, according to the scoring results of the m candidate reference frames, filter out the optimal reference frame of the target coding unit.

[0094] According to the above scoring of the m candidate reference frames, the scoring results of the m candidate reference frames can be obtained. Then, based on the scoring results of the m candidate reference frames, the optimal reference frame of the target coding unit is selected.

[0095] In one embodiment, step 660 may include the following steps:

[0096] S1: Eliminate the candidate reference frames with scoring results lower than the scoring threshold among the m candidate reference frames, and obtain n candidate reference frames; n is an integer greater than 1 and less than m;

[0097] In one embodiment, the terminal eliminates the candidate reference frames with scoring results lower than the scoring threshold according to the scoring results of the m candidate reference frames. Optionally, the scoring threshold is the average score of the m candidate reference frames.

[0098] S2: Calculate the rate-distortion cost of all or part of the candidate reference frames among the n candidate reference frames during inter-frame prediction;

[0099] The terminal calculates the rate-distortion cost of the n candidate reference frames during inter-frame prediction. Optionally, the terminal calculates the rate-distortion cost of the n candidate reference frames in each prediction mode.

[0100] Inter-frame prediction includes k prediction modes. The prediction mode is classified based on the predicted motion vector MVP of the target coding unit; k is an integer greater than 1; for the specific derivation of the prediction mode, please refer to the derivation of the MVP of the target coding unit in the related art.

[0101] Schematically, Figure 7 shows the flowchart for calculating n candidate reference frames provided by an exemplary embodiment of the present application.

[0102] Step 701, sort the scoring results of the n candidate reference frames from high to low to obtain a sorting result;

[0103] The terminal sorts the scoring results of the n candidate reference frames from high to low to obtain a sorting result.

[0104] Step 702, for the i-th prediction mode, perform inter-frame prediction on the j-th candidate reference frame based on the sorting result, and calculate the rate-distortion cost of the j-th candidate reference frame;

[0105] where j is a positive integer, and the initial value of j is 1.

[0106] For the i-th prediction mode, the terminal performs inter-frame prediction according to the sorting result of the n candidate reference frames, and calculates the rate-distortion cost corresponding to the candidate reference frame.

[0107] Schematically, the formula for calculating the rate-distortion cost is:

[0108] rdcost = dist + bit × λ;

[0109] Where λ is a constant, bit represents bits, and dist represents distortion, which records the difference between the pixel values of the target coding unit and the predicted pixel values of the target coding unit. dist can be calculated by any one of SAD (Sum of Absolute Difference), SATD (Sum of Absolute Transformed Difference), and SSE (The Sum of Squares Due to Error).

[0110] Step 703: When the rate-distortion cost of the j-th candidate reference frame is lower than the i-th cost threshold, update j + 1 to j, and re-execute the step: for the i-th prediction mode, perform inter-frame prediction on the j-th candidate reference frame based on the sorting result, and calculate the rate-distortion cost of the j-th candidate reference frame;

[0111] The terminal performs inter-frame prediction according to the sorting result. When the rate-distortion cost of the j-th candidate reference frame is lower than the i-th cost threshold, the terminal obtains the candidate reference frame. The terminal considers the j-th candidate reference frame as a candidate reference frame with better performance. After that, the terminal updates j + 1 to j again and re-executes step 702.

[0112] Step 704: When the rate-distortion cost of the j-th candidate reference frame is not lower than the i-th cost threshold, perform inter-frame prediction for the (i + 1)-th prediction mode, update j to 1, update i + 1 to i, and re-execute the step of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result and calculating the rate-distortion cost of the j-th candidate reference frame for the i-th prediction mode.

[0113] The terminal performs inter-frame prediction according to the sorting result. When the rate-distortion cost of the j-th candidate reference frame is not lower than the i-th cost threshold, the terminal performs inter-frame prediction for the (i + 1)-th prediction mode, updates j to 1, updates i + 1 to i, and re-executes the step of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result and calculating the rate-distortion cost of the j-th candidate reference frame.

[0114] S3: Determine the candidate reference frame with the minimum rate-distortion cost as the optimal reference frame of the target coding unit.

[0115] In the i-th prediction mode, the terminal calculates the rate-distortion cost of the candidate reference frame whose rate-distortion cost is lower than the i-th cost threshold. The terminal determines the candidate reference frame with the minimum rate-distortion cost among all prediction modes as the optimal reference frame of the target coding unit.

[0116] It should be noted that the above steps S1, S2, and S3 actually implement the optimal reference frame selection for each prediction mode under k prediction modes, compare the rate-distortion costs of the optimal reference frames under different prediction modes, and determine the optimal reference frame corresponding to the minimum rate-distortion cost as the final optimal reference frame of the target coding unit.

[0117] In summary, the above method scores m candidate reference frames before inter-frame prediction, and filters out the optimal reference frame for inter-frame prediction according to the scoring results, avoiding calculating the rate-distortion costs when performing inter-frame prediction on all m candidate reference frames, simplifying the determination process of the optimal reference frame, and greatly accelerating the encoding speed of the target coding unit.

[0118] When performing inter-frame prediction, the above method can further determine whether the current candidate reference frame is the optimal reference frame according to whether the rate-distortion cost corresponding to n candidate reference frames in each prediction mode is less than the cost threshold. When the current candidate reference frame is less than the cost threshold, the current candidate reference frame enters the competition sequence of the optimal candidate reference frames. When the current candidate reference frame is not less than the cost threshold, the prediction of the next mode is performed, and the above judgment steps are repeated. The optimal reference frames of all prediction modes under k prediction modes are obtained, and the optimal reference frame of the target coding unit is finally determined according to the rate-distortion costs of the optimal reference frames of all prediction modes. The above method does not need to calculate the rate-distortion costs of all candidate reference frames in each prediction mode, further simplifies the determination process of the optimal reference frame, and speeds up the encoding speed of the target coding unit again.

[0119] Next, the scoring of m candidate coding units based on the 6 quality scoring information of the above m candidate reference frames will be introduced in detail.

[0120] The first possible implementation: Based on Figure 6 In the shown embodiment, step 640 may be replaced by: when there is a first candidate reference frame among the m candidate reference frames that is the optimal reference frame of an adjacent coding unit, determining the score of the first candidate reference frame as the first score;

[0121] Wherein, the adjacent coding unit is a coding unit encoded by inter-frame prediction in a video frame, and the adjacent coding unit is adjacent to the target coding unit.

[0122] Schematically, Figure 8The figure shows a schematic diagram of the positions of adjacent coding units and a target coding unit, where E is the target coding unit, and A, B, C, and D are adjacent coding units. It should be noted that in the related art, there are 22 possible sizes for the target coding unit, and the smallest size of the target coding unit is 4*4. When encoding the target coding unit, the adjacent coding units have been split into coding units of size 4*4 for storage, and only need to obtain A, B, C, and D as shown in Figure 8 for judgment. Figure 8 The figure shows that the size of the target coding unit is 16*16. Similarly, the relationship between target coding units of other sizes and adjacent coding units can be obtained.

[0123] In one embodiment, the terminal sequentially judges four adjacent coding units. If an adjacent coding unit is encoded by inter-frame prediction, the terminal obtains the optimal reference frame corresponding to the adjacent coding unit. When there is a first candidate reference frame among the m candidate reference frames that is the optimal reference frame of the adjacent coding unit, the terminal determines the score of the first candidate reference frame as the first score.

[0124] In one embodiment, if the optimal reference frames corresponding to the four adjacent coding units are completely different, the four candidate reference frames corresponding to the four adjacent coding units are respectively determined as the first score; in one embodiment, if at least two of the four adjacent coding units have the same candidate reference frame as the optimal reference frame, the candidate reference frame is determined as the first candidate reference frame, and the score of the first candidate reference frame is determined as the first score. In one embodiment, regardless of whether there is an overlap in the optimal reference frames corresponding to the four adjacent coding units, the optimal reference frames corresponding to the four adjacent coding units are all determined as the first score.

[0125] Optionally, the scoring weight weight[0] of the first score is 10.

[0126] Schematically, the scoring of the optimal reference frames corresponding to the four adjacent coding units is as follows:

[0127] sorce[ref_A] = sorce[ref_A] + weight[0];

[0128] sorce[ref_B] = sorce[ref_B] + weight[0];

[0129] sorce[ref_C] = sorce[ref_C] + weight[0];

[0130] sorce[ref_D] = sorce[ref_D] + weight[0];

[0131] ref_A, ref_B, ref_C, and ref_D are the first candidate reference frames corresponding to four adjacent coding units respectively, and weight[0] is the scoring weight of the first score.

[0132] The second possible implementation: Based on Figure 6 In the embodiment shown, step 640 may be replaced by: determining the optimal reference frame of the target coding unit under the first splitting type to the (x - 1)-th splitting type; counting the number of times each of the m candidate reference frames is elected as the optimal reference frame of the target coding unit under the first splitting type to the (x - 1)-th splitting type; based on the statistical result, determining the score of the second candidate reference frame among the m candidate reference frames whose elected times exceed the first threshold as the second score.

[0133] Among them, the target coding unit can be obtained based on p splitting types; currently, the screening of the target coding unit obtained based on the x-th splitting type is performed.

[0134] In one embodiment, the optimal reference frame corresponding to each splitting type may be different or the same, and the terminal counts the optimal reference frames corresponding to ten splitting types in related technologies respectively.

[0135] Schematically, Figure 9 shows the order of the splitting types for performing the target coding unit. In one embodiment, the current target coding unit performs HORZ_B splitting prediction 901. Then, before the HORZ_B splitting type, the target coding unit also performs NONE splitting prediction, HORZ splitting prediction, VERT splitting prediction, SPLIT splitting prediction, and HORZ_A splitting prediction. The terminal determines the optimal reference frame of the target coding unit under each splitting type before HORZ_B splitting.

[0136] Schematically, Table 2 shows the number of times the candidate reference frames are elected as the optimal reference frame under the first splitting type to the (x - 1)-th splitting type.

[0137] Table 2

[0138] Reference Frame Selected Times LAST_FRAME 6 LAST2_FRAME 0 LAST3_FRAME 1 GOLDEN_FRAME 3 BWDREF_FRAME 5 ALTREF2_FRAME 2 ALTREF_FRAME 2

[0139] Combined with reference Figure 9 and Table 2, the terminal determines the optimal reference frame of the target coding unit under each splitting type before HORZ_B splitting; the terminal counts the number of times each of the 7 candidate reference frames is elected as the optimal reference frame of the target coding unit under the splitting types before HORZ_B splitting; based on the statistical result, the terminal determines the score of the second candidate reference frame among the 7 candidate reference frames whose elected times exceed the first threshold as the second score. Optionally, the first threshold is 4, then the terminal determines the scores of the LAST_FRAME and BWDREF_FRAME reference frames as the second score.

[0140] Optionally, the scoring weight of the second scoring, weight[1]=10.

[0141] Illustratively, the LAST_FRAME and BWDREF_FRAME are scored as follows:

[0142] sorce[LAST_FRAME]=sorce[LAST_FRAME]+weight[1];

[0143] sorce[BWDREF_FRAME]=sorce[BWDREF_FRAME]+weight[1];

[0144] Wherein, weight[1] is the scoring weight of the second scoring.

[0145] The third possible implementation manner: Based on Figure 6 In the illustrated embodiment, step 640 may be replaced with: Based on the distortion degree of m candidate reference frames, determine the score of the third candidate reference frame with the smallest distortion degree among the m candidate reference frames as the third score.

[0146] Combined with reference to Table 1, the m candidate reference frames are divided into a forward reference frame cluster, a backward reference frame cluster, and a long-term reference frame according to the reference frame direction. That is, the forward reference frame cluster includes: LAST_FRAME, LAST2_FRAME, and LAST3_FRAME, the backward reference frame cluster includes BWDREF_FRAME, ALTREF2_FRAME, and ALTREF_FRAME, and the long-term reference frame is GOLDEN_FRAME.

[0147] In one embodiment, the scoring weight of the third scoring, weight[2]=15.

[0148] In one embodiment, the terminal determines the score of the first forward reference frame as the third score based on that the distortion degree of the first forward reference frame in the forward reference frame cluster is less than the distortion degrees of other forward reference frames, and the distortion degree of the first forward reference frame is not greater than the preset distortion degree. Illustratively, the scoring method of the terminal for the first forward reference frame is:

[0149] sorce[ref_list0]=sorce[ref_list0]+weight[2];

[0150] Wherein, weight[2] is the scoring weight of the third scoring, and ref_list0 indicates the first forward reference frame.

[0151] Based on the fact that the distortion degree of the first backward reference frame in the backward reference frame cluster is less than that of other backward reference frames, and the distortion degree of the first backward reference frame is not greater than the preset distortion degree, the terminal determines the score of the first backward reference frame as the third score.

[0152] Schematically, the way for the terminal to score the first backward reference frame is:

[0153] sorce[ref_list1] = sorce[ref_list1] + weight[2];

[0154] Among them, weight[2] is the scoring weight of the third score, and ref_list1 indicates the first backward reference frame.

[0155] When the distortion degree of the long-term reference frame is not equal to the preset distortion degree, and the distortion degrees of the first forward reference frame and the first backward reference frame are both not greater than the preset distortion degree, and the distortion degree of the long-term reference frame is less than the first distortion threshold, the terminal determines the score of the long-term reference frame as the third score; among them, the first distortion threshold is the sum of the distortion degrees of the first forward reference frame and the first backward reference frame.

[0156] Schematically, the way for the terminal to score the long-term reference frame is:

[0157] sorce[GOLDEN_FRAME] = sorce[GOLDEN_FRAME] + weight[2];

[0158] Among them, weight[2] is the scoring weight of the third score, and GOLDEN_FRAME indicates the long-term reference frame.

[0159] In one embodiment, the ways for the terminal to calculate the distortion degree of the candidate reference frame may include the following two cases:

[0160] T1: When there is a first candidate reference frame in the m candidate reference frames that is the optimal reference frame of the adjacent coding unit, the motion vector of the adjacent coding unit is used as the motion vector of the target coding unit; based on the motion vector of the target coding unit, the distortion degree of the first candidate reference frame is calculated; when the first candidate reference frame corresponds to the distortion degrees of the optimal reference frames of at least two adjacent coding units, the minimum distortion degree is used as the distortion degree of the first candidate reference frame.

[0161] Among them, the adjacent coding unit is a coding unit in the video frame that is encoded using inter-frame prediction, and the adjacent coding unit is adjacent to the target coding unit.

[0162] Combined reference Figure 8, for the target coding unit E, there are four adjacent coding units A, B, C, and D, and if the optimal prediction type corresponding to the adjacent coding unit is inter prediction, the terminal obtains the motion vector of each adjacent coding unit. The terminal calculates the distortion degree of the first candidate reference frame based on the motion vectors of the adjacent coding units. Schematically, the terminal calculates the distortion degree of the first candidate reference frame using the formula:

[0163]

[0164] where src represents the input data of the target coding unit, dst represents the predicted data under the motion vector corresponding to the first candidate reference frame, and sad only reflects the temporal difference of the residual and cannot effectively reflect the size of the bitstream. i, j, m, and n are variables used to locate the first candidate reference frame.

[0165] Optionally, the distortion degree of the first candidate reference frame can also be calculated through satd. After satd undergoes the hadama rd transform and then the absolute value summation, satd is also a way to calculate the distortion. It is to perform the had amard transform on the residual signal and then sum the absolute values of each element. Compared with sad, the computational complexity is more complex but the accuracy is also higher.

[0166] T2: In the case where there are other candidate reference frames among the m candidate reference frames that are not the optimal reference frames of the adjacent coding units except for the first candidate reference frame, the preset distortion degree is used as the distortion degree of the other candidate reference frames.

[0167] Among them, the adjacent coding unit is a coding unit that uses inter prediction for coding in the video frame, and the adjacent coding unit is adjacent to the target coding unit.

[0168] The fourth possible implementation method: Based on Figure 6 In the embodiment shown, step 640 can be replaced by: Based on the fact that there is a fourth candidate reference frame belonging to the preset reference frame set among the m candidate reference frames, the score of the fourth candidate reference frame is determined as the fourth score; the preset reference frame set includes at least one of the m candidate reference frames.

[0169] In one embodiment, the scoring weight weight[3] of the fourth score is 30.

[0170] In one embodiment, the developer selects at least one candidate reference frame from the m candidate reference frames to form the preset reference frame set. When performing inter prediction, the score of the candidate reference frame belonging to the preset reference frame set among the m candidate reference frames is determined as the fourth score. The developer can select and form the preset reference frame set from the m candidate reference frames according to the preference degree. The present application does not limit the composition method of the preset reference frame set.

[0171] Schematically, the fourth score is calculated using the following formula:

[0172] sorce = sorce + weight[3];

[0173] where weight[3] is the scoring weight for the fourth score.

[0174] The fifth possible implementation: Based on Figure 6 In the illustrated embodiment, step 640 may be replaced with: determining the optimal reference frame of the target coding unit in the first prediction mode to the (i - 1)th prediction mode; counting the number of times each of the m candidate reference frames is elected as the optimal reference frame of the target coding unit in the first prediction mode to the (i - 1)th prediction mode; based on the statistical results, determining the score of the fifth candidate reference frame among the m candidate reference frames whose number of elections exceeds the second number threshold as the fifth score.

[0175] Inter-frame prediction includes k prediction modes, and currently the inter-frame prediction in the ith prediction mode is being executed.

[0176] In one embodiment, the currently executed prediction mode is the NEWMV mode, and the target coding unit has executed the NEARESTMV mode, the NEARMV mode, and the GLOBALMV mode. Then the terminal determines the optimal reference frame of the target coding unit in the NEARESTMV mode, the NEARMV mode, and the GLOBALMV mode; the terminal further counts the number of times each of the m candidate reference frames is elected as the optimal reference frame of the target coding unit in the NEARESTMV mode, the NEARMV mode, and the GLOBALMV mode; based on the statistical results, the terminal determines the score of the fifth candidate reference frame among the m candidate reference frames whose number of elections exceeds the second number threshold as the fifth score. Optionally, the second number threshold is 0 or 1 or 2.

[0177] In one embodiment, the scoring weight weight[4] of the fifth score is 20.

[0178] Schematically, the terminal scores the fifth candidate reference frame as follows:

[0179] sorce[nearestmv_best_ref] = sorce[nearestmv_best_ref] + weight[4];

[0180] sorce[nearmv_best_ref] = sorce[nearmv_best_ref] + weight[4];

[0181] sorce[globalmv_best_ref] = sorce[globalmv_best_ref + weight[4];

[0182] Among them, nearestmv_best_ref is the optimal reference frame in the NEARESTMV mode, nearmv_best_ref is the optimal reference frame in the NEARMV mode, globalmv_best_ref is the optimal reference frame in the GLOBALMV mode, and weight[4] is the scoring weight of the fifth score.

[0183] The sixth possible implementation method: Based on Figure 6 In the shown embodiment, step 640 may be replaced by: Based on the number of candidate coding units in the current frame where the target coding unit is located being less than the first quantity threshold, determining the scores of the nearest forward reference frame, the farthest backward reference frame, and the long-term reference frame among the m candidate reference frames as the sixth score;

[0184] Among them, the candidate coding units are used to provide the motion vector of the target coding unit, the first quantity threshold corresponds to the frame type of the current frame where the target coding unit is located, and the frame type is divided based on the reference relationship within the frame group where the current frame is located.

[0185] The candidate coding units are used to provide the motion vector of the target coding unit. Optionally, the candidate coding units include the adjacent coding units of the target coding unit and the coding units obtained by using other segmentation types.

[0186] Schematically, Figure 10 shows a schematic diagram of the reference relationship within the frame group provided by an exemplary embodiment of the present application. From Figure 10 it can be obtained that poc16 refers to poc0; poc8 refers to poc0 and poc16; poc4 refers to poc0 and poc8; poc2 refers to poc0 and poc4; while poc1, poc3, poc5, poc7, poc9, poc11, poc13, and poc15 are not referred to.

[0187] In one embodiment, the 17 frames can be divided into different frame types according to the above reference relationship within the frame group, and Table 3 shows the frame types of each frame within the frame group.

[0188] Table 3

[0189]

[0190] Among them, the weight levels shown in Table 3 are 0-5 in sequence. For different weight levels, a first quantity threshold can be set for the video frames within each weight level. Schematically, use thr to represent the first quantity threshold, thr = param[slice_level]; param represents the number of candidate coding units of the target coding unit within the video frame. Schematically, param[6] = {5, 5, 5, 5, 4, 4}, and the embodiments of the present application do not limit the value of param.

[0191] When the number of candidate coding units of the current frame where the target coding unit is located is less than the first quantity threshold, determine the scores of the nearest forward reference frame (LAST_FRAME), the nearest backward reference frame (BWDREF_FRAME), and the farthest backward reference frame (ALTREF_FRAME) among the m candidate reference frames as the sixth score.

[0192] In one embodiment, the scoring weight weight[5] of the sixth score is 5.

[0193] Schematically, the manner in which the terminal scores the nearest forward reference frame, the nearest backward reference frame, and the farthest backward reference frame is as follows:

[0194] sorce[LAST_FRAME] = sorce[LAST_FRAME] + weight[5];

[0195] sorce[BWDREF_FRAME] = sorce[BWDREF_FRAME] + weight[5];

[0196] sorce[ALTREF_FRAME] = sorce[ALTREF_FRAME] + weight[5];

[0197] Among them, weight[5] is the scoring weight of the sixth score.

[0198] In one embodiment, based on Figure 6 In the embodiment of, step 640 can be replaced by: based on the q quality score information of the candidate reference frames scored among the m candidate reference frames, score the candidate reference frames to be scored q times to obtain q scoring results; by summing the q scoring results, obtain the final score of the candidate reference frames to be scored, where q is a positive integer.

[0199] Schematically, calculate the final score of each candidate reference frame to be scored through the following formula:

[0200] sorce = sorce[0] + weight[0] + weight[1] + weight[2] + … + weight[q - 1];

[0201] Among them, sorce[0] is the initial score of the candidate reference frame for scoring, and weight[0] to weight[q - 1] represent the scoring weights corresponding to q quality scoring information respectively.

[0202] It can be understood that the above candidate reference frame for scoring is at least one of m candidate reference frames. Correspondingly, there may also be candidate reference frames among the m candidate reference frames that do not undergo scoring.

[0203] It is worth noting that in the same scoring method, there may be a situation where the same candidate reference frame is selected multiple times. Then, the terminal can perform corresponding scoring according to the number of times the same candidate reference frame is selected in the same scoring method.

[0204] For example, in the fifth possible implementation manner, in the nearestmv_best_ref, nearmv_best_ref, and globalmv_best_ref modes of the target coding unit, there may be a situation where the same candidate reference frame is recognized as the optimal reference frame multiple times. Then, the corresponding fifth scoring can be performed on the candidate reference frame according to the number of times the same candidate reference frame is recognized as the optimal reference frame.

[0205] Another point worth noting is that the above first possible implementation manner to the sixth possible implementation manner are different scoring methods. This application does not limit the number, order, and scoring weights of the six scoring methods specifically used. That is, in specific implementation, for the same candidate reference frame, some or all of the six scoring methods can be adopted, the six scoring methods can be executed in any order, and the scoring values of the six scoring methods can be set to be exactly the same, partially the same, or completely different. This application does not limit the scoring values of the six scoring methods.

[0206] In summary, according to the above first embodiment to the sixth embodiment, the quality scoring information can be simply summarized into 6 types of quality scoring information: the optimal reference frame information of adjacent coding units, the optimal reference frame information of target coding units of different segmentation types, the distortion degree information of m candidate reference frames, the information of m candidate reference frames belonging to the preset reference frame set, the optimal reference frame information in the previous prediction mode, and the candidate coding unit information of the current frame. The target coding unit completes the scoring of m candidate reference frames according to the six types of quality scoring information. The above method provides a specific information source for scoring, simplifies the process of determining the optimal reference frame, and also speeds up the coding speed of the target coding unit.

[0207] In one embodiment, based on Figure 6 In the embodiment shown, step 660 may be replaced by Figure 11 the steps shown, Figure 11 FIG. shows a method for selecting a reference frame provided by an exemplary embodiment of the present application, the method comprising:

[0208] Step 1101, eliminating candidate reference frames with low scores according to the candidate reference frame scores;

[0209] The terminal sorts the scores of the m candidate reference frames from high to low, calculates the average value of the scores of the m candidate reference frames, eliminates the reference frames with scores lower than the average value, that is, the eliminated reference frames are not used for prediction, and records the candidate reference frame with the highest eliminated score as the ref_num-th candidate reference frame.

[0210] Step 1102, initially set the value of i to 0;

[0211] The terminal sets the initial value of i of the i-th candidate reference frame among the candidate reference frames after eliminating the m candidate reference frames to 0;

[0212] Step 1103, is i less than ref_num?

[0213] The terminal determines whether i is less than ref_num. If i is less than ref_num, execute step 1104; if i is not less than ref_num, execute step 1109.

[0214] Step 1104, obtaining the candidate reference frame of the current mode;

[0215] The terminal obtains the candidate reference frame of the current mode.

[0216] Step 1105, performing inter-frame prediction on the current candidate reference frame;

[0217] The terminal performs inter-frame prediction on the current candidate reference frame.

[0218] Step 1106, calculating the rate-distortion cost of the current candidate reference frame;

[0219] The terminal calculates the rate-distortion cost of the current candidate reference frame.

[0220] Step 1107, is the rate-distortion cost of the current candidate reference frame less than the rate-distortion cost threshold?

[0221] The rate-distortion cost threshold is a threshold set by the developer. If the rate-distortion cost of the current candidate reference frame is less than the rate-distortion cost threshold, execute step 1108; if the rate-distortion cost of the current candidate reference frame is not less than the rate-distortion cost threshold, execute step 1109.

[0222] Step 1108, i = i + 1;

[0223] The terminal updates i + 1 to i, and then executes Step 1103.

[0224] Step 1109, perform prediction in the next prediction mode;

[0225] The terminal performs prediction in the next prediction mode.

[0226] Figure 12 The structure block diagram of the reference frame selection device provided by an exemplary embodiment of the present application is shown. The device includes:

[0227] An acquisition module 1201, configured to acquire m candidate reference frames of a target coding unit, where the target coding unit is one of multiple coding units of a video frame; m is an integer greater than 1;

[0228] A scoring module 1202, configured to score the m candidate reference frames based on the quality scoring information of the m candidate reference frames; the quality scoring information is used to indicate the coding quality of the target coding unit for inter-frame prediction through the m candidate reference frames;

[0229] A screening module 1203, configured to screen out the optimal reference frame of the target coding unit according to the scoring results of the m candidate reference frames.

[0230] In one embodiment, the screening module 1203 is further configured to eliminate candidate reference frames with scoring results lower than a scoring threshold among the m candidate reference frames, to obtain n candidate reference frames; n is an integer greater than 1 and less than m.

[0231] In one embodiment, the screening module 1203 is further configured to calculate the rate-distortion cost of the n candidate reference frames during inter-frame prediction;

[0232] In one embodiment, the screening module 1203 is further configured to determine the candidate reference frame with the minimum rate-distortion cost as the optimal reference frame of the target coding unit.

[0233] In one embodiment, the inter-frame prediction includes k prediction modes; the prediction mode is classified based on the prediction motion vector of the target coding unit; k is an integer greater than 1.

[0234] In one embodiment, the screening module 1203 is further configured to sort the scoring results of the n candidate reference frames from high to low.

[0235] In one embodiment, the screening module 1203 is further configured to, for the i-th prediction mode, perform inter-frame prediction on the j-th candidate reference frame based on the sorting result, and calculate the rate-distortion cost of the j-th candidate reference frame; j is a positive integer; the initial value of j is 1.

[0236] In one embodiment, the screening module 1203 is further configured to, when the rate distortion cost of the j-th candidate reference frame is lower than the i-th cost threshold, update j to j + 1, re-execute the steps of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result for the i-th prediction mode, and calculating the rate distortion cost of the j-th candidate reference frame.

[0237] In one embodiment, the screening module 1203 is further configured to, when the rate distortion cost of the j-th candidate reference frame is not lower than the i-th cost threshold, perform inter-frame prediction for the (i + 1)-th prediction mode, update j to 1, update i + 1 to i, and re-execute the steps of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result for the i-th prediction mode and calculating the rate distortion cost of the j-th candidate reference frame.

[0238] In one embodiment, the scoring module 1202 is further configured to, when there is a first candidate reference frame among the m candidate reference frames that is the optimal reference frame of an adjacent coding unit, determine the score of the first candidate reference frame as the first score; wherein, the adjacent coding unit is a coding unit encoded by using inter-frame prediction in a video frame, and the adjacent coding unit is adjacent to the target coding unit.

[0239] In one embodiment, the target coding unit is obtained based on p segmentation types; the apparatus is currently used to encode the target coding unit obtained by the x-th segmentation type.

[0240] In one embodiment, the scoring module 1202 is further configured to determine the optimal reference frame of the target coding unit under the first segmentation type to the (x - 1)-th segmentation type.

[0241] In one embodiment, the scoring module 1202 is further configured to count the number of times each of the m candidate reference frames is elected as the optimal reference frame of the target coding unit under the first segmentation type to the (x - 1)-th segmentation type.

[0242] In one embodiment, the scoring module 1202 is further configured to, based on the statistical result, determine the score of a second candidate reference frame among the m candidate reference frames whose number of elected times exceeds the first number threshold as the second score.

[0243] In one embodiment, the scoring module 1202 is further configured to, based on the distortion degree of the m candidate reference frames, determine the score of a third candidate reference frame with the smallest distortion degree among the m candidate reference frames as the third score.

[0244] In one embodiment, the m candidate reference frames are divided into a forward reference frame cluster, a backward reference frame cluster, and a long-term reference frame according to the reference frame direction.

[0245] In one embodiment, the scoring module 1202 is further configured to determine the score of the first forward reference frame as the third score based on that the distortion degree of the first forward reference frame in the forward reference frame cluster is less than that of other forward reference frames and the distortion degree of the first forward reference frame is not greater than the preset distortion degree.

[0246] In one embodiment, the scoring module 1202 is further configured to determine the score of the first backward reference frame as the third score based on that the distortion degree of the first backward reference frame in the backward reference frame cluster is less than that of other backward reference frames and the distortion degree of the first backward reference frame is not greater than the preset distortion degree.

[0247] In one embodiment, the scoring module 1202 is further configured to determine the score of the long-term reference frame as the third score when the distortion degree of the long-term reference frame is not equal to the preset distortion degree, and the distortion degrees of the first forward reference frame and the first backward reference frame are both not greater than the preset distortion degree, and the distortion degree of the long-term reference frame is less than the first distortion threshold; wherein the first distortion threshold is the sum of the distortion degrees of the first forward reference frame and the first backward reference frame.

[0248] In one embodiment, the scoring module 1202 is further configured to use the motion vector of the adjacent coding unit as the motion vector of the target coding unit when there is a first candidate reference frame among the m candidate reference frames that is the optimal reference frame of the adjacent coding unit.

[0249] In one embodiment, the scoring module 1202 is further configured to calculate the distortion degree of the first candidate reference frame based on the motion vector of the target coding unit.

[0250] In one embodiment, the scoring module 1202 is further configured to use the minimum distortion degree as the distortion degree of the first candidate reference frame when the first candidate reference frame corresponds to the distortion degrees of the optimal reference frames of at least two adjacent coding units. Wherein, the adjacent coding unit is a coding unit encoded by inter-frame prediction in the video frame and is adjacent to the target coding unit.

[0251] In one embodiment, the scoring module 1202 is further configured to use the preset distortion degree as the distortion degree of other candidate reference frames when there are other candidate reference frames among the m candidate reference frames that are not the optimal reference frames of the adjacent coding units; wherein, the adjacent coding unit is a coding unit encoded by inter-frame prediction in the video frame and is adjacent to the target coding unit.

[0252] In one embodiment, the scoring module 1202 is further configured to determine a fourth score for a fourth candidate reference frame based on the existence of the fourth candidate reference frame in the m candidate reference frames belonging to a preset reference frame set; the preset reference frame set includes at least one of the m candidate reference frames.

[0253] In one embodiment, the inter-frame prediction includes k prediction modes, and the current device is configured to perform inter-frame prediction in the i-th prediction mode.

[0254] In one embodiment, the scoring module 1202 is further configured to determine the optimal reference frame of the target coding unit in the 1st prediction mode to the (i - 1)-th prediction mode.

[0255] In one embodiment, the scoring module 1202 is further configured to count the number of times each of the m candidate reference frames is elected as the optimal reference frame of the target coding unit in the 1st prediction mode to the (i - 1)-th prediction mode.

[0256] In one embodiment, the scoring module 1202 is further configured to determine a fifth score for a fifth candidate reference frame whose number of elections exceeds a second number threshold among the m candidate reference frames based on the statistical result.

[0257] In one embodiment, the scoring module 1202 is further configured to determine a sixth score for the nearest forward reference frame, the nearest backward reference frame, and the farthest backward reference frame among the m candidate reference frames based on the number of candidate coding units in the current frame where the target coding unit is located being less than a first number threshold. Wherein, the candidate coding units are used to provide motion vectors for the target coding unit, and the first number threshold corresponds to the frame type of the current frame where the target coding unit is located, and the frame type is divided based on the reference relationship within the frame group where the current frame is located.

[0258] In one embodiment, the scoring module 1202 is further configured to perform q scores on the candidate reference frames to be scored among the m candidate reference frames based on q quality score information of the candidate reference frames to be scored, and obtain q score results; by summing the q score results, the final score of the candidate reference frames to be scored is obtained, and q is a positive integer.

[0259] It should be noted that: for the reference frame selection device provided in the above embodiments, only the above-described functional modules are used for illustration. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the reference frame selection device provided in the above embodiments belongs to the same concept as the reference frame selection method embodiments, and the specific implementation process is detailed in the method embodiments, which will not be elaborated here.

[0260] In summary, the device provided in this embodiment scores m candidate reference frames before inter-frame prediction, and filters out the optimal reference frame for inter-frame prediction according to the scoring results, avoiding the rate-distortion cost when calculating all m candidate reference frames for inter-frame prediction, simplifying the determination process of the optimal reference frame, and greatly accelerating the encoding speed of the target coding unit.

[0261] Figure 13 FIG. shows a structural block diagram of a computer device 1300 provided by an exemplary embodiment of the present application. The computer device 1300 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The computer device 1300 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0262] Generally, the computer device 1300 includes a processor 1301 and a memory 1302.

[0263] The processor 1301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1301 may also include a main processor and a co-processor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the co-processor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1301 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0264] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may further include high-speed random access memory, as well as non-volatile memory, such as one or more magnetic disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 is used to store at least one instruction for being executed by the processor 1301 to implement the method for selecting a reference frame provided in the method embodiments of the present application.

[0265] In some embodiments, the computer device 1300 may further optionally include: a peripheral device interface 1303 and at least one peripheral device. The processor 1301, the memory 1302, and the peripheral device interface 1303 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1303 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, a positioning assembly 1308, and a power supply 1309.

[0266] The peripheral device interface 1303 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1301 and the memory 1302. In some embodiments, the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1301, the memory 1302, and the peripheral device interface 1303 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0267] The radio frequency circuit 1304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1304 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1304 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 1304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1304 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1304 may further include a circuit related to NFC (Near Field Communication), which is not limited in this application.

[0268] The display screen 1305 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1305 is a touch display screen, the display screen 1305 also has the ability to collect touch signals on or above the surface of the display screen 1305. The touch signals can be input as control signals to the processor 1301 for processing. At this time, the display screen 1305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1305, which is set on the front panel of the computer device 1300; in other embodiments, there may be at least two display screens 1305, which are respectively set on different surfaces of the computer device 1300 or are in a folding design; in other embodiments, the display screen 1305 may be a flexible display screen, which is set on the curved surface or the folding surface of the computer device 1300. Even, the display screen 1305 can be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 1305 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0269] The camera component 1306 is used to collect images or videos. Optionally, the camera component 1306 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which can be any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera component 1306 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. The two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0270] The audio circuit 1307 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1301 for processing, or input to the radio frequency circuit 1304 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively disposed at different parts of the computer device 1300. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signals from the processor 1301 or the radio frequency circuit 1304 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1307 may further include a headphone jack.

[0271] The positioning component 1308 is used to locate the current geographical location of the computer device 1300 to implement navigation or LBS (Location Based Service). The positioning component 1308 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, or Russia's Galileo system.

[0272] The power supply 1309 is used to supply power to each component in the computer device 1300. The power supply 1309 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 1309 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. The wired rechargeable battery is a battery charged through a wired line, and the wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0273] In some embodiments, the computer device 1300 further includes one or more sensors 1310. The one or more sensors 1310 include, but are not limited to: an acceleration sensor 1311, a gyroscope sensor 1312, a pressure sensor 1313, a fingerprint sensor 1314, an optical sensor 1315, and a proximity sensor 1316.

[0274] The acceleration sensor 1311 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the computer device 1300. For example, the acceleration sensor 1311 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1301 can control the display screen 1305 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1311. The acceleration sensor 1311 can also be used for collecting game or user movement data.

[0275] The gyroscope sensor 1312 can detect the body orientation and rotation angle of the computer device 1300. The gyroscope sensor 1312 can cooperate with the acceleration sensor 1311 to collect the 3D actions of the user on the computer device 1300. Based on the data collected by the gyroscope sensor 1312, the processor 1301 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0276] The pressure sensor 1313 can be disposed on the side frame of the computer device 1300 and / or the lower layer of the display screen 1305. When the pressure sensor 1313 is disposed on the side frame of the computer device 1300, it can detect the holding signal of the user on the computer device 1300, and the processor 1301 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 1313. When the pressure sensor 1313 is disposed on the lower layer of the display screen 1305, the processor 1301 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1305. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0277] The fingerprint sensor 1314 is used to collect the user's fingerprint. The processor 1301 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 1314, or the fingerprint sensor 1314 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 1301 authorizes the user to perform relevant sensitive operations, which include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 1314 can be set on the front, back, or side of the computer device 1300. When there are physical buttons or manufacturer logos on the computer device 1300, the fingerprint sensor 1314 can be integrated with the physical buttons or manufacturer logos.

[0278] The optical sensor 1315 is used to collect the ambient light intensity. In one embodiment, the processor 1301 can control the display brightness of the display screen 1305 according to the ambient light intensity collected by the optical sensor 1315. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1305 is increased; when the ambient light intensity is low, the display brightness of the display screen 1305 is decreased. In another embodiment, the processor 1301 can also dynamically adjust the shooting parameters of the camera module 1306 according to the ambient light intensity collected by the optical sensor 1315.

[0279] The proximity sensor 1316, also known as the distance sensor, is usually set on the front panel of the computer device 1300. The proximity sensor 1316 is used to collect the distance between the user and the front of the computer device 1300. In one embodiment, when the proximity sensor 1316 detects that the distance between the user and the front of the computer device 1300 is gradually decreasing, the processor 1301 controls the display screen 1305 to switch from the lit state to the off state; when the proximity sensor 1316 detects that the distance between the user and the front of the computer device 1300 is gradually increasing, the processor 1301 controls the display screen 1305 to switch from the off state to the lit state.

[0280] Those skilled in the art can understand that Figure 13 the structure shown in does not constitute a limitation on the computer device 1300, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0281] This application also provides a computer-readable storage medium, in which at least one instruction, at least one segment of program, code set or instruction set is stored, and the at least one instruction, the at least one segment of program, the code set or instruction set is loaded and executed by the processor to implement the reference frame selection method provided by the above method embodiment.

[0282] The present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for selecting a reference frame provided in the foregoing method embodiment.

[0283] The serial numbers of the foregoing embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0284] Those of ordinary skill in the art can understand that all or part of the steps for implementing the foregoing embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, or the like.

[0285] The foregoing are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for selecting a reference frame, characterized in that, The method includes: Obtaining m candidate reference frames of a target coding unit, where the target coding unit is one of multiple coding units of a video frame, and m is an integer greater than 1; Rating the m candidate reference frames based on the quality score information of the m candidate reference frames; the quality score information is used to indicate the coding quality of the target coding unit for inter-frame prediction through the candidate reference frames; the inter-frame prediction includes k prediction modes, and the prediction modes are classified based on the predicted motion vectors of the target coding unit; k is an integer greater than 1; Eliminating the candidate reference frames with rating results lower than the rating threshold among the m candidate reference frames to obtain n candidate reference frames; n is an integer greater than 1 and less than m; Sorting the rating results of the n candidate reference frames from high to low to obtain a sorting result; For the i-th prediction mode among the k prediction modes, performing inter-frame prediction on the j-th candidate reference frame based on the sorting result, and calculating the rate-distortion cost of the j-th candidate reference frame; j is a positive integer; the initial value of j is 1; When the rate-distortion cost of the j-th candidate reference frame is lower than the i-th cost threshold, assigning the result of j + 1 to j, and re-executing the step of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result and calculating the rate-distortion cost of the j-th candidate reference frame for the i-th prediction mode among the k prediction modes; When the rate-distortion cost of the j-th candidate reference frame is not lower than the i-th cost threshold, performing inter-frame prediction for the (i + 1)-th prediction mode, updating j to 1, assigning the result of i + 1 to i, and re-executing the step of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result and calculating the rate-distortion cost of the j-th candidate reference frame for the i-th prediction mode among the k prediction modes; Determining the candidate reference frame with the minimum rate-distortion cost as the optimal reference frame of the target coding unit.

2. The method according to claim 1, characterized in that, The rating of the m candidate reference frames based on the quality score information of the m candidate reference frames includes: When there is a first candidate reference frame among the m candidate reference frames that is the optimal reference frame of an adjacent coding unit, determining the rating of the first candidate reference frame as a first rating; Wherein, the adjacent coding unit is a coding unit encoded by inter-frame prediction in the video frame, and the adjacent coding unit is adjacent to the target coding unit.

3. The method according to claim 1, characterized in that, The target coding unit is obtained based on p segmentation types; the method is applied to encode the target coding unit obtained by the x-th segmentation type; p is an integer greater than 1; x is a positive integer and x is less than p; The rating of the m candidate reference frames based on the quality score information of the m candidate reference frames includes: Determining the optimal reference frame of the target coding unit under the 1st segmentation type to the (x - 1)-th segmentation type; Counting the number of times each of the m candidate reference frames is elected as the optimal reference frame of the target coding unit under the 1st segmentation type to the (x - 1)-th segmentation type; Based on the statistical results, determine the score of the second candidate reference frame among the m candidate reference frames whose selection times exceed the first threshold as the second score.

4. The method according to claim 1, characterized in that, The m candidate reference frames are divided into a forward reference frame cluster, a backward reference frame cluster, and a long-term reference frame according to the reference frame direction; The scoring of the m candidate reference frames based on the quality scoring information of the m candidate reference frames includes: Based on the fact that the distortion degree of the first forward reference frame in the forward reference frame cluster is less than the distortion degrees of other forward reference frames, and the distortion degree of the first forward reference frame is not greater than the preset distortion degree, determine the score of the first forward reference frame as the third score; Or; Based on the fact that the distortion degree of the first backward reference frame in the backward reference frame cluster is less than the distortion degrees of other backward reference frames, and the distortion degree of the first backward reference frame is not greater than the preset distortion degree, determine the score of the first backward reference frame as the third score; Or; Determine the first forward reference frame in the forward reference frame cluster, where the distortion degree of the first forward reference frame in the forward reference frame cluster is less than the distortion degrees of other forward reference frames; Determine the first backward reference frame in the backward reference frame cluster, where the distortion degree of the first backward reference frame in the backward reference frame cluster is less than the distortion degrees of other backward reference frames; In the case that the distortion degree of the long-term reference frame is not equal to the preset distortion degree, and the distortion degrees of the first forward reference frame and the first backward reference frame are both not greater than the preset distortion degree, and the distortion degree of the long-term reference frame is less than the first distortion threshold, determine the score of the long-term reference frame as the third score; where the first distortion threshold is the sum of the distortion degrees of the first forward reference frame and the first backward reference frame.

5. The method according to claim 4, characterized in that, The method further includes: In the case that there is a first candidate reference frame among the m candidate reference frames that is the optimal reference frame of an adjacent coding unit, use the motion vector of the adjacent coding unit as the motion vector of the target coding unit; Based on the motion vector of the target coding unit, calculate the distortion degree of the first candidate reference frame; In the case that the first candidate reference frame corresponds to the distortion degrees of at least two optimal reference frames of the adjacent coding units, use the minimum distortion degree as the distortion degree of the first candidate reference frame; Where the adjacent coding unit is a coding unit in the video frame that is encoded using inter-frame prediction, and the adjacent coding unit is adjacent to the target coding unit.

6. The method according to claim 4, characterized in that, The method further includes: In the case that there are other candidate reference frames among the m candidate reference frames that are not the optimal reference frames of the adjacent coding units except for the first candidate reference frame, use the preset distortion degree as the distortion degree of the other candidate reference frames; Where the adjacent coding unit is a coding unit in the video frame that is encoded using inter-frame prediction, and the adjacent coding unit is adjacent to the target coding unit.

7. The method according to claim 1, wherein The scoring of the m candidate reference frames based on the quality scoring information of the m candidate reference frames includes: Based on the fact that there is a fourth candidate reference frame among the m candidate reference frames that belongs to a preset reference frame set, determine the score of the fourth candidate reference frame as the fourth score; the preset reference frame set includes at least one of the m candidate reference frames.

8. The method according to claim 1, wherein The inter-frame prediction includes k prediction modes, and the method is applied to perform inter-frame prediction in the i-th prediction mode; The scoring of the m candidate reference frames based on the quality score information of the m candidate reference frames includes: Determine the optimal reference frame of the target coding unit in the 1st prediction mode to the (i - 1)-th prediction mode; Count the number of times each of the m candidate reference frames is elected as the optimal reference frame of the target coding unit in the 1st prediction mode to the (i - 1)-th prediction mode; Based on the statistical result, determine the score of a fifth candidate reference frame among the m candidate reference frames whose elected times exceed a second threshold as the fifth score.

9. The method according to claim 1, wherein The scoring of the m candidate reference frames based on the quality score information of the m candidate reference frames includes: Based on the number of candidate coding units of the current frame where the target coding unit is located being less than a first threshold, determine the scores of the nearest forward reference frame, the nearest backward reference frame, and the farthest backward reference frame among the m candidate reference frames as the sixth score; Wherein, the candidate coding unit is used to provide the motion vector of the target coding unit, the first threshold corresponds to the frame type of the current frame where the target coding unit is located, and the frame type is divided based on the reference relationship within the frame group where the current frame is located.

10. The method according to claim 1, wherein The scoring of the m candidate reference frames based on the quality score information of the m candidate reference frames includes: Based on the q quality score information of the candidate reference frames to be scored among the m candidate reference frames, score the candidate reference frames to be scored q times to obtain q scoring results; by summing the q scoring results, obtain the final score of the candidate reference frames to be scored, and q is a positive integer.

11. A reference frame selection device, wherein The device includes: An acquisition module, configured to acquire m candidate reference frames of a target coding unit, where the target coding unit is one of multiple coding units of a video frame, and m is an integer greater than 1; A scoring module, configured to score the m candidate reference frames based on the quality score information of the m candidate reference frames; the quality score information is used to indicate the coding quality of the target coding unit for inter-frame prediction through the candidate reference frames; the inter-frame prediction includes k prediction modes, and the prediction modes are classified based on the predicted motion vector of the target coding unit; k is an integer greater than 1; A screening module, configured to eliminate candidate reference frames with scoring results lower than a scoring threshold among the m candidate reference frames to obtain n candidate reference frames; n is an integer greater than 1 and less than m; Sort the scoring results of the n candidate reference frames from high to low to obtain a sorting result; For the i-th prediction mode among the k prediction modes, perform inter-frame prediction on the j-th candidate reference frame based on the sorting result, and calculate the rate-distortion cost of the j-th candidate reference frame; j is a positive integer; the initial value of j is 1; If the rate-distortion cost of the j-th candidate reference frame is lower than the i-th cost threshold, assign j + 1 to j, and re-execute the step of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result for the i-th prediction mode among the k prediction modes, and calculating the rate-distortion cost of the j-th candidate reference frame; If the rate-distortion cost of the j-th candidate reference frame is not lower than the i-th cost threshold, perform inter-frame prediction for the (i + 1)-th prediction mode, update j to 1, assign i + 1 to i, and re-execute the step of performing inter-frame prediction on the j-th candidate reference frame based on the sorting result for the i-th prediction mode among the k prediction modes, and calculating the rate-distortion cost of the j-th candidate reference frame; Determine the candidate reference frame with the minimum rate-distortion cost as the optimal reference frame of the target coding unit.

12. A computer device, wherein The computer device includes: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the method for selecting a reference frame as described in any one of claims 1 to 10.

13. A computer-readable storage medium, wherein The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the method for selecting a reference frame as described in any one of claims 1 to 10.

14. A computer program product, wherein The computer program product stores at least one computer instruction, the computer instruction is stored in a computer-readable storage medium, the processor reads the computer instruction from the computer-readable storage medium, and the computer instruction is loaded and executed by the processor to implement the method for selecting a reference frame as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Quick depth video coding method

    CN101986716A

  • Fast multi-composite frame video coding method and device and storage medium

    CN110278434A

  • Video coding method and device, electronic equipment and computer readable storage medium

    CN111263151A

  • Optimized search for reference frames in predictive video coding system

    US20120328018A1