AVS3-Based Affine Mode Screening Method, Device and Electronic Device
By filtering the affine mode of the encoding unit in AVS3 and writing the affine fusion candidate list, the problem of determining the prediction mode of the encoding unit in AVS3 is solved, and the effect of reducing the calculation complexity and improving the prediction mode efficiency is achieved.
Patent Information
- Application Number
- CN202111534864.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-15
AI Technical Summary
In AVS3, the prediction mode efficiency of determining the coding unit is low, mainly because the affine SKIP mode requires a large number of fine interpolation operations, which increases the complexity of the coding algorithm and reduces the parallel computing efficiency.
By selecting the reference airspace CU of the current encoding unit CU, and filtering out the affine control point in the affine mode in the reference airspace CU, writing it into the affine fusion candidate list, and inter prediction is performed.
By pre-filtering the candidate list of affine SKIP mode before the RDO stage of rate distortion optimization, the computational complexity of the affine SKIP mode is reduced and the efficiency of determining the prediction mode of the CU is improved.
Smart Images

Figure CN114554209B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video coding and decoding. Specifically, it relates to an affine mode screening method, device, and electronic device based on AVS3. Background Technique
[0002] In the third-generation audio and video standard AVS3, an affine model is newly added to predict motions such as scaling and rotation between coding unit (CU) blocks. A four- or six-parameter affine model is adopted in AVS3: the motion vectors (MVs) of two or three control points are encoded respectively. It is only used for CUs with a width and height greater than or equal to 8, and the motion vectors of each 4x4 or 8x8 block are derived through the control point MVs.
[0003] There are two types of affine SKIP modes in AVS3, the model-based affine mode and the control point-based affine mode. The CONTROL POINT BASED mode directly selects 2 or 3 control points in the spatial domain. The MODEL BASED mode first finds the available spatial MVs in the upper left, upper right, and lower left of the current CU block, and finds the corresponding available temporal spatial MVs in the lower right. Secondly, for each spatial MV, find the coding CU block where it is located, and use the control points of this CU block as the control points of the current CU block.
[0004] The affine SKIP mode requires a large number of fine interpolation operations, which will increase the complexity of the coding algorithm. And since the motion vectors of each smallest coding unit (SCU) obtained by the affine transformation are different, it will also reduce the parallel computing efficiency of the affine SKIP mode. Moreover, due to the large number of prediction modes, the complexity of calculating the rate-distortion cost of the CU in each prediction mode is high, resulting in a large amount of calculation for determining the CU prediction mode, and thus the efficiency of determining the CU prediction mode is low. Summary of the Invention
[0005] Embodiments of this application provide an affine mode screening method, device, and electronic device based on AVS3 to at least solve the technical problem of low efficiency in determining the prediction mode of the CU in the related art.
[0006] According to one aspect of the embodiments of the present application, there is provided an affine mode screening method based on AVS3, including: selecting a reference spatial CU of the current coding unit CU; wherein, the reference spatial CU is a coding unit in an available state; when there is a reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, taking the affine control points of the reference spatial CU whose target coding mode is an affine mode as the affine control points of the current CU, and writing the first affine model into the affine fusion candidate list; wherein, the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among the rate-distortion costs of multiple inter-frame prediction modes, and the first affine model includes an affine mode based on a model; when there are at least two reference spatial CUs in the reference spatial CUs whose target coding mode is an affine mode, writing the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes an affine mode based on control points; when there is no reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, writing the second affine model into the affine fusion candidate list; performing inter-frame prediction on the current CU according to the affine fusion candidate list.
[0007] According to another aspect of the embodiments of the present application, there is further provided an affine mode screening device based on AVS3, including: a selection unit, configured to select a reference spatial CU of the current coding unit CU; wherein, the reference spatial CU is a coding unit in an available state; a first writing unit, configured to, when there is a reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, take the affine control points of the reference spatial CU whose target coding mode is an affine mode as the affine control points of the current CU, and write the first affine model into the affine fusion candidate list; wherein, the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among the rate-distortion costs of multiple inter-frame prediction modes, and the first affine model includes an affine mode based on a model; a second writing unit, configured to, when there are at least two reference spatial CUs in the reference spatial CUs whose target coding mode is an affine mode, write the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes an affine mode based on control points; a third writing unit, configured to, when there is no reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, write the second affine model into the affine fusion candidate list; a prediction unit, configured to perform inter-frame prediction on the current CU according to the affine fusion candidate list.
[0008] According to still another aspect of the embodiments of the present application, there is further provided a computer-readable storage medium, in which a computer program is stored, and wherein the computer program is configured to execute the above-mentioned affine mode screening method based on AVS3 when running.
[0009] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the above-mentioned AVS3-based affine mode screening method through the computer program.
[0010] In the embodiments of the present application, by selecting a reference spatial CU of the current coding unit CU; wherein, the reference spatial CU is a coding unit in an available state; when there is a target coding mode of a reference spatial CU in the reference spatial CUs that is an affine mode, the affine control points of the reference spatial CU with the target coding mode of the affine mode are used as the affine control points of the current CU, and the first affine model is written into the affine fusion candidate list; wherein, the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes an affine mode based on a model; when there are at least two reference spatial CUs in the reference spatial CUs whose target coding mode is an affine mode, the first affine model and / or the second affine model are written into the affine fusion candidate list; wherein, the second affine model includes an affine mode based on control points; when there is no reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, the second affine model is written into the affine fusion candidate list; frame inter prediction is performed on the current CU according to the affine fusion candidate list. Since the candidate list of the affine SKIP mode is pre-screened before the rate-distortion optimization (RDO) stage, not only the computational complexity of the affine SKIP mode is reduced, but also the efficiency of determining the prediction mode of the CU is improved, solving the technical problem of low efficiency in determining the prediction mode of the CU in the related art. Description of the Drawings
[0011] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0012] Figure 1 is a schematic diagram of an application environment of an optional AVS3-based affine mode screening method according to an embodiment of the present invention;
[0013] Figure 2 is a schematic diagram of another optional application environment of an AVS3-based affine mode screening method according to an embodiment of the present invention;
[0014] Figure 3 is a schematic flowchart of an optional AVS3-based affine mode screening method according to an embodiment of the present invention;
[0015] Figure 4It is a schematic diagram of a spatial CU of a current CU according to an embodiment of the present invention;
[0016] Figure 5 It is another schematic diagram of a temporal CU of a current CU according to an embodiment of the present invention;
[0017] Figure 6 It is a schematic diagram of an affine motion vector generated by two control points for a current CU according to an embodiment of the present invention;
[0018] Figure 7 It is a schematic diagram of an affine motion vector generated by three control points for a current CU according to an embodiment of the present invention;
[0019] Figure 8 It is a schematic diagram of the structure of an optional affine mode screening device based on AVS3 according to an embodiment of the present invention;
[0020] Figure 9 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0021] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0023] According to one aspect of an embodiment of the present invention, there is provided an affine mode screening method based on AVS3. Optionally, as an alternative implementation manner, the above-mentioned affine mode screening method based on AVS3 can be but is not limited to being applied to, for example Figure 1In the hardware environment shown. The hardware environment includes: a terminal device 102 for human-computer interaction with the user, a network 104, and a server 106. Human-computer interaction can be carried out between the user 108 and the terminal device 102. An application client is screened in the affine mode based on AVS3 in the terminal device 102. The terminal device 102 includes a human-computer interaction screen 1022, a processor 1024, and a memory 1026. The human-computer interaction screen 1022 is used to present an interface for video frame processing; the processor 1024 is used to obtain the reference spatial CU of the current coding unit CU. The memory 1026 is used to store the reference spatial CU of the current coding unit CU.
[0024] In addition, the server 106 includes a database 1062 and a processing engine 1064. The database 1062 is used to store the reference spatial CU of the current coding unit CU and an affine fusion candidate list. The processing engine 1064 is used to select the reference spatial CU of the current coding unit CU according to an affine mode screening request based on AVS3; wherein, the above reference spatial CU is a coding unit in an available state; when there is a target coding mode of an affine mode in the above reference spatial CUs, the affine control points of the reference spatial CU with the target coding mode of the affine mode are used as the affine control points of the current CU, and a first affine model is written into the affine fusion candidate list; wherein, the above target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes an affine mode based on the model; when there are at least two reference spatial CUs with the target coding mode of the affine mode in the above reference spatial CUs, the first affine model and / or a second affine model are written into the affine fusion candidate list; wherein, the second affine model includes an affine mode based on control points; when there is no reference spatial CU with the target coding mode of the affine mode in the above reference spatial CUs, the second affine model is written into the affine fusion candidate list; frame-inter prediction is performed on the current CU according to the above affine fusion candidate list.
[0025] As another alternative embodiment, the above method for screening the affine mode based on AVS3 in the present application can be applied to Figure 2 in. As Figure 2 shown, human-computer interaction can be carried out between the user 202 and the user device 204. The user device 204 includes a memory 206 and a processor 208. In this embodiment, the terminal device 204 can but is not limited to refer to and execute the operations performed by the above terminal device 102, and frame-inter prediction is performed on the current CU as above.
[0026] Optionally, the above-mentioned terminal device 102 and user device 204 may be, but are not limited to, terminals such as mobile phones, tablet computers, laptop computers, and PC machines. The above-mentioned network 104 may include, but is not limited to, a wireless network or a wired network. Among them, the wireless network includes: WIFI and other networks that implement wireless communication. The above-mentioned wired network may include, but is not limited to: wide area network, metropolitan area network, local area network. The above-mentioned server 106 may include, but is not limited to, any hardware device that can perform calculations. The above-mentioned server may be a single server, or a server cluster composed of multiple servers, or a cloud server. The above is only an example, and this embodiment does not make any limitations in this regard.
[0027] Optionally, in one or more embodiments, as Figure 3 shown, the above-mentioned AVS3-based affine mode screening method includes:
[0028] S302, select a reference spatial CU of the current coding unit CU; wherein, the reference spatial CU is a coding unit in an available state.
[0029] In the embodiment of the present invention, as Figure 4 shown in Figure 5 accordance with the AVS3 standard, the reference frame image Col_pic of the current frame image Cur_pic; 1 co-located temporal T of the current CU (Cur) is located at the lower right of the current Cu adjacent to it, and 6 spatial CUs (A, B, D, G, C, F) adjacent to the current CU are a total of 7 reference CUs. For the CUs that are not encoded in the standard and have an encoding mode of intra mode, they are discarded.
[0030] S304, when there is a reference spatial CU in the reference spatial CUs whose target coding mode is the affine mode, use the affine control points of the reference spatial CU with the target coding mode of the affine mode as the affine control points of the current CU, and write the first affine model into the affine fusion candidate list; wherein, the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter prediction modes, and the first affine model includes the model-based affine mode;
[0031] In the embodiment of the present invention, for the model-based affine mode (MODEL BASED AFFINE MODE), as Figure 4 shown in
[0032] S306. When the target coding modes of at least two reference spatio-temporal CUs in the reference spatio-temporal CU are affine modes, write the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes the affine mode based on control points.
[0033] In the embodiments of the present invention, the affine mode based on control points (CONTROL POINT BASED AFFINE MODE), such as Figure 4 As shown, in this mode, 2 or 3 control points are selected from the spatio-temporal sets (A, B, D, G, C, F) adjacent to the current CU (Cur). Secondly, these 2 or 3 control points are used as the control points of the current CU block.
[0034] S308. When there is no reference spatio-temporal CU in the reference spatio-temporal CU whose target coding mode is an affine mode, write the second affine model into the affine fusion candidate list.
[0035] S310. Perform inter-frame prediction on the current CU according to the affine fusion candidate list.
[0036] Specifically, after obtaining the affine fusion candidate list, perform affine transformation interpolation on the current frame to predict the next frame of the current frame.
[0037] In the embodiments of the present application, by selecting the reference spatio-temporal CU of the current coding unit CU; wherein, the above reference spatio-temporal CU is a coding unit in an available state; when there is one reference spatio-temporal CU in the above reference spatio-temporal CU whose target coding mode is an affine mode, use the affine control points of the reference spatio-temporal CU with the target coding mode of affine mode as the affine control points of the above current CU, and write the first affine model into the affine fusion candidate list; wherein, the above target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes the affine mode based on the model; when the target coding modes of at least two reference spatio-temporal CUs in the above reference spatio-temporal CU are affine modes, write the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes the affine mode based on control points; when there is no reference spatio-temporal CU in the above reference spatio-temporal CU whose target coding mode is an affine mode, write the above second affine model into the affine fusion candidate list; perform inter-frame prediction on the above current CU according to the above affine fusion candidate list. Since the candidate list of the affine SKIP mode is pre-screened before the rate-distortion optimization (RDO) stage, it not only reduces the computational complexity of the affine SKIP mode, but also improves the efficiency of determining the prediction mode of the CU, and solves the technical problem of low efficiency in determining the prediction mode of the CU in the related art.
[0038] In one or more embodiments, in step S302 above, the selection of the reference spatial CU of the current coding unit CU includes: based on the AVS3 coding standard, taking one co-located temporal CU and six spatial CUs of the current CU as the reference spatial CU.
[0039] In the embodiments of the present invention, as Figure 4 with Figure 5 shown, according to the AVS3 standard, the reference frame image Col_pic of the current frame image Cur_pic; one co-located temporal T of the current CU (Cur) is located at the lower right adjacent to the current Cu, and six spatial CUs (A, B, D, G, C, F) adjacent to the current CU, a total of 7 CUs, are taken as the reference spatial CU of the current Cu. For the CUs that are not coded in the standard and have an intra coding mode, they are discarded.
[0040] In one or more embodiments, in step S304 above, when the target coding mode of one reference spatial CU in the reference spatial CUs is the affine mode, taking the affine control points of the reference spatial CU with the target coding mode of the affine mode as the affine control points of the current CU, and writing the first affine model into the affine fusion candidate list, includes:
[0041] When the target coding mode of one spatial CU among the six spatial CUs is the affine mode, taking the affine control points of the spatial CU with the target coding mode of the affine mode as the affine control points of the current CU to perform affine transformation interpolation on the current CU; writing the first affine model into the affine fusion candidate list.
[0042] It should be noted here that according to the AVS3 standard, as Figure 4 and Figure 5 shown, the co-located temporal CU (T) does not participate in the MODELBASED affine model calculation, and only six spatial CUs (A, B, D, G, C, F) are considered.
[0043] In one or more embodiments, in step S306 above, when the target coding modes of at least two reference spatial CUs in the reference spatial CUs are the affine mode, writing the first affine model and / or the second affine model into the affine fusion candidate list, includes:
[0044] When the target coding modes of at least two reference spatial CUs in the reference spatial CUs are the affine mode, perform the following operations:
[0045] When it is determined that the positional relationship between the CU with the above target coding mode being the affine mode and the current CU satisfies the preset rule, write the second affine model into the affine fusion candidate list as the target affine model; wherein, the above preset rule includes the reference spatial domain CUs corresponding to two control points or three control points adjacent to the current CU.
[0046] In the embodiment of the present invention, the above preset rule is that the positional relationship between the CU in the affine mode and the current CU can form an affine model of two control points or three control points, then use the CONTROL POINT BASED affine model. The control points of the CONTROL POINTBASED affine model are divided into four categories: upper left, lower left, upper right, and lower right. Among them, the upper left control points include {A, B, D}, the lower left control points include {F}, the upper right control points include {G, C}, and the lower right control points include {T}.
[0047] As Figure 6 and Figure 7 shown, Figure 6 For the current CU (Cur) including two control points, Figure 7 For the current CU (Cur) including three control points; the affine model of two control points (4 parameters) is as follows:
[0048]
[0049] Among them, (v 0x , v 0y ) is the motion vector of the upper left control point of the current CU, (v 1x , v 1y ) is the motion vector of the upper right control point of the current CU, and W is the width of the current CU.
[0050] The affine model of three control points (6 parameters) is as follows:
[0051]
[0052] Among them, (v 0x , v 0y ) is the motion vector of the upper left control point of the current CU, (v 1x , v 1y ) is the motion vector of the upper right control point of the current CU, (v 2x , v 2y ) is the motion vector of the lower left control point of the current CU, W is the width of the current CU, and H is the height of the current CU.
[0053] When it is determined that the positional relationship between the CU with the above-mentioned target coding mode being the affine mode and the current CU does not meet the preset rules, the first affine model is written into the affine fusion candidate list as the target affine model. That is to say, if there is a CU in the affine mode that cannot form an affine model with two control points or three control points with the current CU, this CU is used as the control point of the MODELBASED affine model.
[0054] In one or more embodiments, the set of reference spatial domain CUs corresponding to two control points or three control points in the above-mentioned preset rules includes at least one of the following:
[0055] A first subset of reference spatial domain CUs including the upper left CU adjacent to the current CU, the upper right CU adjacent to the current CU, the lower left CU adjacent to the current CU, or the lower right CU adjacent to the current CU; for example, Figure 4 As shown, the first subset of reference spatial domain CUs is {A / B / D, G / C, F / T}.
[0056] A second subset of reference spatial domain CUs including the upper CU adjacent to the current CU, the lower left CU adjacent to the current CU, and the lower right CU adjacent to the current CU; for example, Figure 4 As shown, the second subset of reference spatial domain CUs is {A / B / D / G / C, F, T}.
[0057] A third subset of reference spatial domain CUs including the upper left CU adjacent to the current CU and the right CU adjacent to the current CU. For example, Figure 4 As shown, the third subset of reference spatial domain CUs is {A / B / D, G / C / T}.
[0058] In one or more embodiments, in step S308 above, when the target coding mode of the reference spatial domain CU in which there is no reference spatial domain CU is the affine mode, writing the second affine model into the affine fusion candidate list includes:
[0059] When the target coding mode of the reference spatial domain CU in which there is no reference spatial domain CU is the affine mode, a fourth subset of reference spatial domain CUs including the upper left CU adjacent to the current CU, the upper right CU adjacent to the current CU, and the lower left CU adjacent to the current CU is determined; the affine control points of the fourth subset of reference spatial domain CUs are used as the affine control points of the current CU to perform affine transformation interpolation on the current CU; the second affine model is written into the affine fusion candidate list.
[0060] In an embodiment of the present invention, when the target coding mode of the reference spatial CU does not exist in the above-mentioned reference spatial CU and is the affine mode, the encoded CUs adjacent to the current CU in the upper left, upper right, and lower left are found according to the AVS3 standard and used as the affine control points of the current CU, and the CONTROL POINT BASED affine model is used. It should be noted that it is only required that the spatial CU used in this process is encoded, and it is not required that this CU is encoded in the affine mode.
[0061] Based on the above embodiments, in one or more embodiments, the above-mentioned affine mode screening method based on AVS3 includes the following steps:
[0062] 1) Select 7 CUs including 6 spatial CUs (F / G / C / A / B / D) and the temporal co-located CU (T) of the current CU (Cur). The positions of the 7 CUs are shown as Figure 4 、 Figure 5 shown.
[0063] 2) If there is only one CU among the 7 CUs whose optimal mode is the affine mode, directly use this CU as the control point of the MODELBASED affine model.
[0064] 3) If there are two or more CUs among the 7 CUs whose optimal mode is the affine mode, select the CONTROL POINTBASED affine model or the MODEL BASED affine model.
[0065] 4) If the candidate list depth is less than 2, add the above-mentioned CONTROL POINT BASED affine model in the spatial domain to the candidate list.
[0066] In the above step 1): According to the AVS3 standard, there are 7 relevant CUs including 1 temporal and 6 spatial CUs of the current CU. For the modes that are not encoded and have the intra mode in the standard, they should be discarded.
[0067] In the above step 2), it specifically includes the following steps:
[0068] a) According to the AVS3 standard, the temporal co-located CU (T) does not participate in the MODEL BASED affine model calculation, and only 6 spatial CUs are considered.
[0069] b) According to the AVS3 standard, use the affine control points of this CU as the affine control points of the current CU for affine transformation interpolation.
[0070] In the above step 2), it specifically includes the following steps:
[0071] a) If the positional relationship between the affine mode CU and the current CU can form a 4-point or 6-point affine model, then the CONTROL POINT BASED affine model is used.
[0072] b) According to the AVS3 standard, the control points of the CONTROL POINT BASED affine model are divided into four categories: upper left, lower left, upper right, and lower right. Among them, the upper left control points include {A, B, D}, the lower left control point includes {F}, the upper right control points include {G, C}, and the lower right control point includes {T}.
[0073] c) The CONTROL POINT BASED affine model is used when the affine mode CU is one of the following combinations.
[0074] {A / B / D, G / C, F / T};
[0075] {A / B / D / G / C, F, T};
[0076] {A / B / D, G / C / T};
[0077] d) If there is an affine mode CU that cannot form a 2-point or 6-point affine model with the current CU, then this CU is used as a control point of the MODEL BASED affine model.
[0078] The above step 4) specifically includes the following steps:
[0079] a) If the depth of the candidate list is less than 2, then according to the AVS3 standard, find the encoded CUs adjacent to the current CU in the upper left, upper right, and lower left as the affine control points of the current CU, and use the CONTROL POINT BASED affine model.
[0080] b) For the spatial domain CU used in this process, it only needs to be encoded, and it is not required that this CU uses the affine mode for encoding.
[0081] The embodiment of the present invention is based on the AVS3 standard, and through the encoded information, preprocesses the candidates provided by the affine merge (AFFINEMERGE) mode of the current CU block, reduces the number of interpolation and RDO operations, and improves the encoding efficiency of the video frame.
[0082] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0083] According to another aspect of the embodiments of the present application, there is also provided an AVS3-based affine mode screening device for implementing the above AVS3-based affine mode screening method. As Figure 8 shown, the device includes:
[0084] A selection unit 802, configured to select a reference spatial CU of the current coding unit CU; wherein, the above reference spatial CU is a coding unit in an available state.
[0085] In the embodiments of the present invention, as Figure 4 and Figure 5 shown, according to the AVS3 standard, the reference frame image Col_pic of the current frame image Cur_pic; one co-located temporal domain T of the current CU (Cur) is located at the lower right of the current Cu adjacent, and 7 reference CUs including 6 spatial CUs (A, B, D, G, C, F) adjacent to the current CU. CUs that are not coded in the standard and have an intra mode are discarded.
[0086] A first writing unit 804, configured to, when there is a reference spatial CU in the above reference spatial CUs whose target coding mode is an affine mode, use the affine control points of the reference spatial CU with the target coding mode of the affine mode as the affine control points of the current CU, and write the first affine model into the affine fusion candidate list; wherein, the above target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among the rate-distortion costs of multiple inter prediction modes, and the first affine model includes a model-based affine mode.
[0087] In the embodiments of the present invention, the model-based affine mode (MODEL BASED AFFINE MODE), as Figure 4 shown, first finds available spatial motion vectors MV in the upper left, upper right, and lower left of the current CU (Cur), and finds available temporal MV (such as the motion vector of T) in the lower right. Secondly, for the above available MVs, find the coding CU blocks where they are located, and use the control points of this CU block as the control points of the current CU block.
[0088] A second writing unit 806, configured to, when there are at least two reference spatial CUs in the above reference spatial CUs whose target coding mode is an affine mode, write the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes a control point-based affine mode.
[0089] In the embodiments of the present invention, the control point-based affine mode (CONTROL POINT BASED AFFINEMODE), as Figure 4As shown, in this mode, two or three control points are selected from the set of adjacent spatial domains (A, B, D, G, C, F) of the current CU (Cur). Secondly, these two or three control points are used as the control points of the current CU block.
[0090] The third writing unit 808 is configured to write the second affine model into the affine fusion candidate list when the target coding mode of the reference spatial domain CU does not exist in the above reference spatial domain CU and is the affine mode.
[0091] The prediction unit 810 is configured to perform inter-frame prediction on the current CU according to the above affine fusion candidate list. Specifically, after obtaining the affine fusion candidate list, perform affine transformation interpolation on the current frame to predict the next frame of the current frame.
[0092] In the embodiment of the present application, by selecting the reference spatial domain CU of the current coding unit CU; wherein, the above reference spatial domain CU is an available coding unit; when there is a reference spatial domain CU in the above reference spatial domain CU whose target coding mode is the affine mode, use the affine control points of the reference spatial domain CU with the target coding mode of the affine mode as the affine control points of the current CU, and write the first affine model into the affine fusion candidate list; wherein, the above target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes the affine mode based on the model; when there are at least two reference spatial domain CUs in the above reference spatial domain CU whose target coding mode is the affine mode, write the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes the affine mode based on the control points; when there is no reference spatial domain CU in the above reference spatial domain CU whose target coding mode is the affine mode, write the second affine model into the affine fusion candidate list; perform inter-frame prediction on the current CU according to the above affine fusion candidate list. Since the candidate list of the affine SKIP mode is pre-screened before the rate-distortion optimization (RDO) stage, it not only reduces the computational complexity of the affine SKIP mode, but also improves the efficiency of determining the prediction mode of the CU, and solves the technical problem of low efficiency in determining the prediction mode of the CU in the related art.
[0093] In one or more embodiments, the above selection unit 802 specifically includes:
[0094] The selection module is configured to use one co-located temporal CU and six spatial CUs of the current CU as the above reference spatial domain CU based on the AVS3 coding standard.
[0095] In one or more embodiments, the above first writing unit 804 specifically includes:
[0096] A first determining module is configured to, when a target coding mode of one of the six spatial domain CUs is an affine mode, use the affine control points of the spatial domain CU whose target coding mode is the affine mode as the affine control points of the current CU to perform affine transformation interpolation on the current CU;
[0097] The first writing module is used to write the first affine model into the affine fusion candidate list.
[0098] In one or more embodiments, the second writing unit 806 specifically includes:
[0099] The judgment module is configured to perform the following operations when the target coding mode of at least two reference spatial domain CUs in the reference spatial domain CU is an affine mode:
[0100] A second writing module is used to write the second affine model as the target affine model into the affine fusion candidate list when it is determined that the position relationship between the CU whose target coding mode is the affine mode and the current CU satisfies a preset rule; wherein the preset rule includes a reference spatial domain CU corresponding to two control points or three control points adjacent to the current CU;
[0101] The third writing module is used to write the first affine model as the target affine model into the affine fusion candidate list when it is determined that the position relationship between the CU whose target coding mode is the affine mode and the current CU does not meet the preset rule.
[0102] In one or more embodiments, the reference spatial domain CU set corresponding to two control points or three control points in the above preset rule includes at least one of the following:
[0103] A first reference spatial domain CU subset including an upper left CU adjacent to the current CU, an upper right CU adjacent to the current CU, a lower left CU adjacent to the current CU, or a lower right CU adjacent to the current CU;
[0104] A second reference spatial domain CU subset including an upper CU adjacent to the current CU, a lower left CU adjacent to the current CU, and a lower right CU adjacent to the current CU;
[0105] The third reference spatial domain CU subset includes the upper left CU adjacent to the current CU and the right CU adjacent to the current CU.
[0106] In one or more embodiments, the third writing unit specifically includes:
[0107] A second determination module is used to determine a fourth reference spatial domain CU subset including an upper left CU adjacent to the current CU, an upper right CU adjacent to the current CU, and a lower left CU adjacent to the current CU when the target coding mode of the reference spatial domain CU does not exist in the reference spatial domain CU is an affine mode;
[0108] A third determination module, configured to use the affine control points of the above-mentioned fourth reference spatial CU subset as the affine control points of the current CU, so as to perform affine transformation interpolation on the current CU;
[0109] A fourth writing module, configured to write the above-mentioned second affine model into the affine fusion candidate list.
[0110] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above-mentioned affine mode screening method based on AVS3. The electronic device may be Figure 1 the terminal device or server shown in the figure. This embodiment takes the electronic device as a server as an example for illustration. As Figure 9 shown, the electronic device includes a memory 902 and a processor 904. A computer program is stored in the memory 902, and the processor 904 is configured to execute the steps in any one of the above method embodiments through the computer program.
[0111] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.
[0112] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:
[0113] S1. Select a reference spatial CU of the current coding unit CU; wherein, the reference spatial CU is a coding unit in an available state;
[0114] S2. When there is a reference spatial CU with a target coding mode of affine mode among the reference spatial CUs, use the affine control points of the reference spatial CU with the target coding mode of affine mode as the affine control points of the current CU, and write the first affine model into the affine fusion candidate list; wherein, the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes an affine mode based on a model;
[0115] S3. When there are at least two reference spatial CUs with a target coding mode of affine mode among the reference spatial CUs, write the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes an affine mode based on control points;
[0116] S4. When there is no reference spatial CU with a target coding mode of affine mode among the reference spatial CUs, write the second affine model into the affine fusion candidate list;
[0117] S5. Perform inter-frame prediction on the current CU according to the affine fusion candidate list.
[0118] Optionally, those of ordinary skill in the art can understand that Figure 9 the structure shown is only schematic, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (MID), a PAD, and other terminal devices. Figure 9 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 9 , or have a different configuration from that shown in Figure 9 .
[0119] Among them, the memory 902 can be used to store software programs and modules, such as the program instructions / modules corresponding to the AVS3-based affine mode screening method and device in the embodiments of the present application. The processor 904 executes various functional applications and data processing by running the software programs and modules stored in the memory 902, that is, implements the above-mentioned AVS3-based affine mode screening method. The memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 902 may further include a memory remotely set relative to the processor 904, and these remote memories may be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 902 may specifically but not limitedly be used to store information such as the reference spatial CU of the current coding unit CU. As an example, as Figure 9 shown, the above memory 902 may but not limitedly include the selection unit 802, the first writing unit 804, the second writing unit 806, the third writing unit 808, and the prediction unit 810 in the above AVS3-based affine mode screening device. In addition, it may also include but not limited to other module units in the above AVS3-based affine mode screening device, which will not be elaborated in this example.
[0120] Optionally, the above transmission device 909 is used to receive or send data via a network. Specific examples of the above network may include a wired network and a wireless network. In one instance, the transmission device 909 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one instance, the transmission device 909 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0121] In addition, the above-mentioned electronic device further includes: a display 908 for displaying the reference spatial CU information of the current coding unit CU; and a connection bus 910 for connecting each module component in the above-mentioned electronic device.
[0122] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes in a form of network communication. Among them, the nodes may form a peer-to-peer (P2P) network, and any form of computing device, such as electronic devices like servers and terminals, can become a node in the blockchain system by joining the peer-to-peer network.
[0123] In one or more embodiments, the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned AVS3-based affine mode screening method. Among them, the computer program is set to execute the steps in any one of the above method embodiments when running.
[0124] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store a computer program for executing the following steps:
[0125] S1, select a reference spatial CU of the current coding unit CU; wherein, the reference spatial CU is a coding unit in an available state;
[0126] S2, when there is a target coding mode of a reference spatial CU in the reference spatial CUs that is an affine mode, use the affine control points of the reference spatial CU with the target coding mode of affine mode as the affine control points of the current CU, and write the first affine model into the affine fusion candidate list; wherein, the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes an affine mode based on a model;
[0127] S3, when there are at least two reference spatial CUs in the reference spatial CUs whose target coding mode is an affine mode, write the first affine model and / or the second affine model into the affine fusion candidate list; wherein, the second affine model includes an affine mode based on control points;
[0128] S4, when there is no reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, write the second affine model into the affine fusion candidate list;
[0129] S5, perform inter-frame prediction on the current CU according to the affine fusion candidate list.
[0130] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the relevant hardware of the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0131] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0132] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0133] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the units or modules can be in an electrical or other form.
[0134] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0136] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An affine mode screening method based on AVS3, characterized in that, Including: Selecting a reference spatial coding unit (CU) of the current CU; wherein, the reference spatial CU is a coding unit in an available state; When there is a reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, using the affine control points of the reference spatial CU with the target coding mode of the affine mode as the affine control points of the current CU, and writing a first affine model into an affine fusion candidate list; wherein, the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes an affine mode based on a model; When there are at least two reference spatial CUs in the reference spatial CUs whose target coding mode is an affine mode, writing the first affine model and / or a second affine model into the affine fusion candidate list; wherein, the second affine model includes an affine mode based on control points; When there is no reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, writing the second affine model into the affine fusion candidate list; Performing inter-frame prediction on the current CU according to the affine fusion candidate list.
2. The method according to claim 1, characterized in that, The selection of the reference spatial CU of the current CU includes: Based on the AVS3 coding standard, using a co-located temporal CU of the current CU and six spatial CUs as the reference spatial CU.
3. The method according to claim 2, characterized in that, When there is a reference spatial CU in the reference spatial CUs whose target coding mode is an affine mode, using the affine control points of the reference spatial CU with the target coding mode of the affine mode as the affine control points of the current CU, and writing the first affine model into the affine fusion candidate list, includes: When there is a spatial CU in the six spatial CUs whose target coding mode is an affine mode, using the affine control points of the spatial CU with the target coding mode of the affine mode as the affine control points of the current CU to perform an affine transformation difference on the current CU; Writing the first affine model into the affine fusion candidate list.
4. The method according to claim 2, characterized in that, When there are at least two reference spatial CUs in the reference spatial CUs whose target coding mode is an affine mode, writing the first affine model and / or the second affine model into the affine fusion candidate list, includes: When there are at least two reference spatial CUs in the reference spatial CUs whose target coding mode is an affine mode, perform the following operations: When it is determined that the positional relationship between the CU with the target coding mode of the affine mode and the current CU satisfies a preset rule, writing the second affine model as the target affine model into the affine fusion candidate list; wherein, the preset rule includes the reference spatial CUs corresponding to two control points or three control points adjacent to the current CU; When it is determined that the positional relationship between the CU with the target coding mode of the affine mode and the current CU does not satisfy the preset rule, writing the first affine model as the target affine model into the affine fusion candidate list.
5. The method according to claim 4, characterized in that, The set of reference spatial CUs corresponding to two control points or three control points in the preset rule includes at least one of the following: A first reference spatial CU subset including the upper left CU adjacent to the current CU, the upper right CU adjacent to the current CU, the lower left CU adjacent to the current CU, or the lower right CU adjacent to the current CU; A second reference spatial CU subset including the upper CU adjacent to the current CU, the lower left CU adjacent to the current CU, and the lower right CU adjacent to the current CU; A third reference spatial CU subset including the upper left CU adjacent to the current CU and the right CU adjacent to the current CU.
6. The method according to claim 2, characterized in that, When the target coding mode of the reference spatial CU in which there is no reference spatial CU is the affine mode, writing the second affine model into the affine fusion candidate list includes: When the target coding mode of the reference spatial CU in which there is no reference spatial CU is the affine mode, determining a fourth reference spatial CU subset including the upper left CU adjacent to the current CU, the upper right CU adjacent to the current CU, and the lower left CU adjacent to the current CU; Using the affine control points of the fourth reference spatial CU subset as the affine control points of the current CU to perform affine transformation difference on the current CU; Writing the second affine model into the affine fusion candidate list.
7. An affine mode screening device based on AVS3, characterized in that, Including: A selection unit for selecting a reference spatial CU of the current coding unit CU; wherein the reference spatial CU is a coding unit in an available state; A first writing unit for, when the target coding mode of a reference spatial CU in which there is one reference spatial CU is the affine mode, using the affine control points of the reference spatial CU with the target coding mode of the affine mode as the affine control points of the current CU and writing the first affine model into the affine fusion candidate list; wherein the target coding mode is the prediction mode corresponding to the minimum rate-distortion cost among multiple inter-frame prediction modes, and the first affine model includes an affine mode based on a model; A second writing unit for, when the target coding mode of at least two reference spatial CUs in the reference spatial CU is the affine mode, writing the first affine model and / or the second affine model into the affine fusion candidate list; wherein the second affine model includes an affine mode based on control points; A third writing unit for, when the target coding mode of the reference spatial CU in which there is no reference spatial CU is the affine mode, writing the second affine model into the affine fusion candidate list; A prediction unit for performing inter-frame prediction on the current CU according to the affine fusion candidate list.
8. The device according to claim 7, characterized in that, The selection unit specifically includes: A selection module for, based on the AVS3 coding standard, using a co-located temporal CU of the current CU and six spatial CUs as the reference spatial CU.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method described in any one of claims 1 to 6.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 6 through the computer program.
Citation Information
Patent Citations
Video processing method, coding end and decoding end
CN113194314A
Affine motion estimation method, device and equipment and storage medium
CN113630601A