An av1 video intra coarse mode decision optimization and hardware architecture method
Patent Information
- Application Number
- CN202310887654.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-07-19
AI Technical Summary
[0004](1)AV1的帧内模式决策情况繁多,所需决策时间较长;
[0040]本发明通过精简模式决策过程,将原来5+56种帧内决策模式(5种非方向决策模式+56种方向决策模式)优化为5+8+6种帧内决策模式(5种非方向决策模式+8种主方向决策模式+6种次方向决策模式),在保证压缩效果良好的前提下,减少计算复杂度;所有块在SATD过程进行hadamard变换时均基于4x4块,使得决策效率更高,进一步减少帧内决策时间。相比现有的一些帧内模式决策,本方案在决策速度更快且改进代价小,创新度较高,并为视频帧内模式决策算法在硬件上实现的相关研究提供了参考。
Smart Images

Figure CN117041565B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing and video compression technology, specifically relating to an intra-frame coarse-mode decision optimization and hardware architecture method for AV1 video. Background Technology
[0002] Released in 2018 by the Alliance for Open Media (AOMedia) (AV1), AV1 is a royalty-free and open-source video codec. Developed by AOMedia, a company comprised of many leading technology companies, AV1 aims to handle Ultra High Definition (UHD) 8K (7680×4320 pixels) and 4K video (3840×2160 pixels) with high encoding efficiency. This results in increased complexity compared to other codecs on the market, such as VP9, HEVC, and H.264.
[0003] During video encoding, the encoder divides the image to be encoded into coding tree units (CTUs) and encodes them one by one. In AV1, a coding tree unit (CTU) is defined as a Super Block (SB), for example, an SB of size 64x64. Intra-frame prediction aims to determine the optimal block partition for each SB and the optimal mode for each partitioned block, thereby achieving better compression. Therefore, it is necessary to make shape decisions on all modes for all possible block sizes to obtain the optimal mode. However, existing technologies mainly suffer from the following problems:
[0004] (1) AV1 has many intra-frame mode decision-making situations and requires a long decision-making time;
[0005] (2) The hardware implementation of AV1 intra-frame mode decision has low parallelism and incomplete pipeline, resulting in low implementation efficiency.
[0006] (3) In AV1 intra-frame mode decision-making, the computational complexity of prediction cost ranking is high. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention optimizes intra-frame prediction decisions, significantly reducing computational complexity and improving video coding efficiency. It also facilitates hardware implementation, greatly shortening implementation time and saving required hardware resources and space. The technical solution adopted in this invention is as follows:
[0008] A video frame coarse-pattern decision optimization method includes the following steps:
[0009] Step S1: Obtain pixel blocks of different sizes within the video frame;
[0010] Step S2: For pixel blocks of different sizes, perform intra-frame coarse-mode prediction. In intra-frame coarse-mode prediction, a first round of decision-making is performed based on a set of intra-frame prediction modes, and several best intra-frame prediction modes in the first round are selected. In the second round of decision-making, it is first determined whether there is a directional intra-frame prediction mode in the first round of decision-making. If there is, several best intra-frame prediction modes in the second round are selected as the final decision result based on several best intra-frame prediction modes in the first round and the existing directional intra-frame prediction modes. Otherwise, the result obtained from the first round of decision-making is directly used as the final decision result.
[0011] Further, in step S2, based on the intra-predictors corresponding to pixel blocks of different sizes, predicted pixel values are obtained. Residual values are calculated by comparing the predicted values with the original pixel values. SATD values are obtained based on the residual values, and the optimal intra-prediction mode is selected by sorting the SATD values. This invention does not perform Hadamard transform on the predicted values of rectangular blocks; it is based on the superposition of SATD values from 4x4 square blocks.
[0012] Furthermore, the SATD values are bitonically sorted, and the best intra-frame prediction modes are selected as the final decision results. In the bitonic sorting, the sequence of n SATD values is divided in half, assuming n = 2^k. Then, the 1st and n / 2+1th SATD values are compared. For ascending order, the smaller value is placed on top, and for descending order, the larger value is placed on top. Next, the 2nd and n / 2+2nd values are compared, and so on. For the two sequences of length n / 2 formed, since they are both bitonic sequences, the above process can be repeated for a total of k rounds. That is, the last round is a comparison of sequences of length 2, and the final sorting result is obtained.
[0013] Bitone sorting is a data-independent sorting method, meaning that the comparison order is independent of the data. It is particularly suitable for parallel computing, such as using GPUs and FPGAs.
[0014] A hardware architecture method for video intra-frame coarse-mode decision optimization is disclosed. Based on the aforementioned video intra-frame coarse-mode decision optimization method, the optimal intra-frame prediction mode is selected, and its corresponding hardware framework is constructed based on bitonic sorting. First, functions are written and the recursive and calling relationships between functions are established. The bitonic merging function calls the bitonic sorting function to perform large-step sorting. Then, it checks whether the recursive call is on the last step. If not, it continues to recursively call itself and increments the maximum number by one, resulting in log_2N recursions. In the bitonic sorting function, the input array is divided into several smaller arrays according to the different maximum numbers, and then the comparison function is called several times for sorting. Similarly, the bitonic sorting function also has a recursive calling relationship from the maximum number to 1. The comparison function receives data of variable length and calls the cell-level comparison function to sort pairwise, and finally returns to the bitonic sorting function that called it. Finally, the bitonic merging function executes to the last step and outputs the final sorted data. The written functions are compiled into gate circuits, and the function calling relationship is the execution order during compilation.
[0015] Furthermore, the pipelining design of the bitonic sorting function in hardware implementation is related to the number of pipesteps. `pipe_step` bitonic sorting functions are connected to form a single-stage pipelining. The total number of bitonic sorting functions is... There are N = log2 length, where length represents the length of the array to be sorted. When pipe_step is a factor of M, that is, when the number of sorting units contained in each level is the same, the pipeline time of each level is balanced.
[0016] Furthermore, the system-level optimization of RMD is mainly reflected in the pipeline and parallelism. In the pixel block processing flow, the calculation of SATD value, accumulation of SATD value and sorting process adopt a pipeline design, which saves a lot of area resources on the one hand and shortens the implementation time on the other hand. In addition, the parallel processing design is adopted between different modes of blocks of the same size, which not only shortens the implementation time, but also meets the timing requirements of each level of pipeline.
[0017] An AV1 video intra-frame coarse-mode decision optimization method is provided. Based on the aforementioned video intra-frame coarse-mode decision optimization method, the coding tree unit of the video frame is defined as a superblock. From the different pixel block partitioning modes that constitute each superblock, the optimal block partitioning mode is predicted. The intra-frame prediction modes include the flying direction intra-frame prediction mode and the directional intra-frame prediction mode, wherein the directional intra-frame prediction mode can be further subdivided into multiple auxiliary directional mode predictions.
[0018] Furthermore, the non-directional intra-frame prediction modes include DC mode, Paeth mode, Smooth mode, Smooth_vertical mode, and Smooth_horizontal mode;
[0019] The Paeth mode predicts blocks by performing sample-level copying from a reference sample, as shown in the following formula:
[0020] B = (L + MA)
[0021] p L =|BL|
[0022] p A =|BA|
[0023] p M =|BM|
[0024]
[0025] Where L represents the reference sample pixel value based on the vertical direction of the pixel block, A represents the reference sample pixel value based on the horizontal direction of the pixel block, M represents the pixel value based on a single reference sample of the pixel block, B is the base value, representing the lowest value among the three reference sample pixel values, and p L p A p M These are intermediate variables, Lmin(p) L ,p A ,p M ) = p L p L It is p L p A p M The minimum value in, Amin(p) L ,p A ,p M ) = p A p A It is p L p A p M The minimum value in, Mmin(p) L ,p A ,p M ) = p M p M It is p L p A p M The minimum value in the range, where P represents the pixel value of the predicted sample;
[0026] The smoothing mode makes predictions by applying bilinear interpolation in the vertical and horizontal directions based on the distance between the predicted sample and the reference sample. It includes a smoothing vertical mode, a smoothing horizontal mode, and a combination of the two smoothing modes, as shown in the following formula:
[0027]
[0028]
[0029] Among them, V (-1,M-1) This indicates that the smooth vertical application uses the interpolation obtained by using the reference sample pixel values corresponding to each column of A-RSV and L-RSV, V (N-1,-1) The smoothing level is represented by the interpolation obtained by using the reference sample pixel values of each row of L-RSV and A-RSV. L-RSV and A-RSV represent the reference sample pixel values in the vertical and horizontal directions of the pixel block, respectively. M represents the number of pixels in the height of the pixel block, N represents the number of pixels in the width of the pixel block, (x,y) represents the position of the predicted sample in the pixel block, and 255 represents the maximum pixel value.
[0030] The DC pattern is used to predict a block in which all samples are predicted as the average of the available RSV samples. This average is also called the DC value, and the formula is as follows:
[0031]
[0032] Furthermore, the intra-frame prediction mode is applicable to regions with defined edge and direction features, and uses the base angle B and the increment angle Δ to represent multiple prediction angles P based on defined prediction angles:
[0033] P = B + 3 × Δ
[0034] The directional mode requires all RSVs, but the utilization of RSVs depends on the predicted angle. Therefore, only L-RSVs are used for angles greater than 180°, only A-RSVs are used for angles less than 90°, and all three RSVs are used for angles between 180° and 90°.
[0035] Furthermore, angle prediction uses interpolation to define the prediction sample; for angles less than 90°, formula (10) is used; for angles greater than 180°, formula (10) is used; and for angles between 90° and 180°, formulas (9) and (10) are used:
[0036]
[0037]
[0038] Where x and y represent the horizontal and vertical coordinates of the predicted sample in the pixel block. This represents the horizontal reference sample pixel value with the same x-coordinate as the predicted pixel. S represents the reference sample pixel value in the vertical direction with the same ordinate as the predicted pixel. x and S y This indicates the parameter value obtained from a predefined table.
[0039] The advantages and beneficial effects of this invention are as follows:
[0040] This invention simplifies the mode decision-making process, optimizing the original 5+56 intra-frame decision modes (5 non-directional decision modes + 56 directional decision modes) into 5+8+6 intra-frame decision modes (5 non-directional decision modes + 8 primary directional decision modes + 6 secondary directional decision modes). This reduces computational complexity while maintaining good compression performance. All blocks are based on 4x4 blocks during the Hadamard transform in the SATD process, resulting in higher decision-making efficiency and further reducing intra-frame decision time. Compared to some existing intra-frame mode decision-making methods, this scheme offers faster decision-making speed, lower improvement costs, and higher innovation, providing a reference for related research on the hardware implementation of video intra-frame mode decision-making algorithms. Attached Figure Description
[0041] Figure 1 This is a flowchart illustrating the basic process of intra-frame coarse mode decision-making in this embodiment of the invention.
[0042] Figure 2 This is an example diagram of intra-frame prediction for 8x8 blocks in an embodiment of the present invention.
[0043] Figure 3 This is an example diagram of Paeth pattern prediction in an embodiment of the present invention.
[0044] Figure 4 This is an example diagram of Sooth pattern prediction in an embodiment of the present invention.
[0045] Figure 5 This is the original flowchart of intra-frame decision-making in AV1 in this embodiment of the invention.
[0046] Figure 6 This is a flowchart of intra-frame RMD optimization in an embodiment of the present invention.
[0047] Figure 7 This is a flowchart illustrating the specific process of intra-frame RMD in an embodiment of the present invention.
[0048] Figure 8 This is a flowchart of the bitone sorting process in an embodiment of the present invention.
[0049] Figure 9 This is a schematic diagram of the flow structure of the bi-tuning sorting system in an embodiment of the present invention.
[0050] Figure 10 This is a schematic diagram of system-level pipelined parallelism in an embodiment of the present invention. Detailed Implementation
[0051] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0052] Existing intra-frame mode decision-making is complex, mode cost sorting is slow, and the related hardware parallelism is low and the pipeline is incomplete, resulting in slow overall decision-making speed and low efficiency.
[0053] Therefore, the basic optimization process of the video frame coarse-mode decision optimization and hardware architecture method of the present invention is as follows: Figure 1 As shown, firstly, pixel blocks of different sizes are classified. Predicted values are obtained using intra-predictors corresponding to different pixel block sizes. Then, residual values are calculated by comparing the predicted values with the original pixels. The obtained residual values are then subjected to a Hadamard transform to obtain SATD (Sum of Absolute Transformed Difference) values. SATD values represent the quality of the relevant mode's performance. The required intra-frame mode can be determined using SATD values. Therefore, these SATD values are bitonic-sorted to obtain the four smallest SATD values. The four modes corresponding to these SATD values are the four optimal intra-frame decision modes. The purpose of this coarse-grained mode decision optimization method is to obtain the four optimal intra-frame mode decisions, preparing for subsequent mode decisions and improving overall decision efficiency. Compared to the original AV1 open-source method, this coarse-grained mode decision optimization method has made process optimizations. These optimizations make the implementation of intra-frame mode decisions more hardware-friendly and conducive to related hardware implementation. At the same time, these optimizations save a significant amount of computational complexity, thereby greatly shortening hardware implementation time and saving hardware resources and area. The hardware architecture method of this invention is mainly reflected in parallelism and pipeline stages. This hardware design method greatly improves the utilization of hardware resources and shortens the implementation time.
[0054] 1.1: Pixel Block Division
[0055] In AV1, the Coding Tree Unit (CTU) is defined as a Super Block (SB). In this scheme, the SB size is 64x64. The purpose of intra-frame prediction is to determine the optimal block partition for each SB and the optimal mode for each partitioned block, thereby achieving a better compression effect. Therefore, it is necessary to make decisions on all modes for all possible block sizes to obtain the optimal mode. For a 64x64 pixel block, the pixel blocks in this scheme are: 256 4x4 blocks, 128 4x8 blocks, 128 8x4 blocks, 64 4x16 blocks, 64 16x4 blocks, 64 8x8 blocks, 32 8x16 blocks, 32 16x8 blocks, 16 8x32 blocks, 16 32x8 blocks, 16 16x16 blocks, 8 16x32 blocks, 8 32x16 blocks, 4 16x64 blocks, 4 64x16 blocks, 4 32x32 blocks, 2 32x64 blocks, and 2 64x32 blocks. The four optimal patterns will be selected from these pixel blocks.
[0056] 1.2: Intra-frame mode prediction for AV1
[0057] This scheme includes 5 non-directional intra-prediction modes and 8 directional intra-prediction modes. The 8 directional intra-prediction modes can be further subdivided into 6 auxiliary directional prediction modes. The 5 non-directional intra-prediction modes are: DC mode, Paeth mode, Smooth mode, Smooth_vertical mode, and Smooth_horizontal mode.
[0058] like Figure 2 As shown, predicting pixel values requires a reference sample vector (RSV), used in an 8x8 block to execute all RSVs and neighboring samples for all modes, where M×N indicates the block size (height and width). The number of samples required for each RSV varies depending on the block size (M×N) and the prediction mode being executed. In the worst case, the vertical and horizontal sizes of L-RSV and A-RSV are twice the size of the prediction block, while M-RSV has only one sample. Different intra-frame prediction modes will use different reference pixel blocks and different calculation formulas. Intra-frame prediction modes include the following:
[0059] 1) Paeth mode
[0060] Paeth mode requires three pixels, namely the M-RSV, A-RSV, and L-RSV corresponding to the coordinates, as shown in the example image. Figure 3As shown. The Paeth model predicts blocks by performing sample-level copying from reference samples. The Paeth model evaluates three reference samples to determine which will be used as the predictor (copying) each sample. An example of this prediction model and the required reference samples from L-RSV, A-RSV, and M-RSV are shown. Equations (1)-(4) describe the application of the Paeth model, where it first calculates the base value B, as defined in (1). The base value is used to calculate PL, PA, and PM, as defined in (2)-(4). Then, the lowest value among these three variables (i.e., the closest to the base value B) is used, as defined in (5). For example, if pL is the lowest one, the predicted sample P is defined as a copy of pL.
[0061] B = (L + MA) (1)
[0062] p L =|BL| (2)
[0063] p A =|BA| (3)
[0064] p M =|BM| (4)
[0065]
[0066] Where L represents the reference sample pixel value on the left, M represents the reference sample pixel value at the top left, and A represents the reference sample pixel value at the top. Lmin(p L ,p A ,p M ) = p L p L It is p L p A p M The minimum value in, Amin(p) L ,p A ,p M ) = p A p A It is p L p A p M The minimum value in, Mmin(p) L ,p A ,p M ) = p M p M It is p L p A p M The minimum value in the range, where P represents the pixel value of the predicted sample.
[0067] 2) Smooth mode
[0068] Smoothing prediction modes effectively predict gradient regions by applying bilinear interpolation in the vertical and horizontal directions based on the distance between the predicted and reference samples. AV 1 has three smoothing modes: smoothing vertical mode, smoothing horizontal mode, and a combination of the two smoothing modes described above, called smoothing. Smoothing vertical application uses the interpolation of the A-RSV and L-RSV samples V[-1][M-1] for each column. Smoothing horizontal application uses the interpolation of the L-RSV and A-RSV samples V[N-1][-1] for each row, such as... Figure 4 As shown. Equations (6) and (7) define the vertical and horizontal predictions (Pv and Ph, respectively) used in the smooth vertical and smooth horizontal modes. The reference sample is multiplied by a weight W that depends on the block size and the position of the predicted sample. W is defined to achieve an accuracy of 1 / 256 pixels for a 64×64 block.
[0069]
[0070]
[0071] 3) Directional Mode
[0072] AV 1 defines 56 prediction angles, organized into eight basic angles: 45°, 67°, 90°, 113°, 135°, 157°, 180°, and 203°. The orientation pattern is applicable to regions with defined edge and orientation features. AV 1 uses the base angle (B) and increment angle (Δ) to represent the 56 prediction angles. The Δ angle can take integer values in the range [-3, 3]. Equation (8) defines the prediction angle P inference:
[0073] P = B + 3 × Δ (8)
[0074] Orientation mode requires all RSVs, but RSV utilization depends on the prediction angle. Therefore, angles greater than 180° use only L-RSV, angles less than 90° use only A-RSV, and angles between 180° and 90° use all three RSVs. All other angle predictions use interpolation to define the prediction sample. The calculation results are shown in (9) and (10). For angles less than 90°, (10) is used, while for angles greater than 180°, (10) is used. For angles between 90° and 180°, (9) and (10) are used. The AV1 reference software uses a predefined table to generate the correct values of Sx and Sy in the following equation:
[0075]
[0076]
[0077] 4) DC mode
[0078] The DC pattern is used to predict blocks in which all samples are predicted as the average of the available RSV samples. This average is also called the DC value and is defined in (11):
[0079]
[0080] 2.1: Intra-frame decision of AV1
[0081] The intra-frame decision-making process of AV1 is as follows: Figure 5 As shown, the decision-making process requires iterating through 61 patterns one by one. Each iteration requires passing through the predictor for the corresponding pattern to generate predicted pixel values, obtaining residual values based on reference pixels, and finally calculating the pattern cost based on SATD. Because there are many patterns to iterate through, the required resources and time are relatively large, making it impossible to complete the decision-making process under limited resources and time.
[0082] 2.2: Rough Mode Decision (RMD)
[0083] Because there are many AV1 modes, it is difficult to complete the decision-making process in a short time and with limited resources. Therefore, this invention proposes an intra-frame coarse-mode decision-making method for AV1. Figure 6 and Figure 7 As shown, the main innovation of coarse-pattern decision-making is that the decision-making process is divided into two rounds: the first round only considers 13 patterns, that is, ranking the calculated SATD costs of the corresponding patterns and selecting the four best ones; the second round first determines whether there is a directional pattern in the first round's decision. If the first round's decision result contains a directional pattern, then decisions are made on the four patterns obtained in the first round and the six auxiliary directions of the main direction, and then the four best patterns are selected; if the first round's decision result does not contain a directional pattern, then the result obtained in the first round is directly used as the final decision result.
[0084] 3.1: Bitonic Sort
[0085] Bitone sorting is a data-independent sorting method, meaning the comparison order is independent of the data. It is particularly suitable for parallel computing, such as using GPUs and FPGAs. A bitone sequence is a sequence that is first monotonically increasing and then monotonically decreasing (or first monotonically decreasing and then monotonically increasing).
[0086] Batcher's theorem states that any bitonic sequence A of length 2n is divided into two halves X and Y of equal length. Elements in X are compared one by one with elements in Y in the original order, that is, a[i] is compared with a[i+n] (i < n), the larger one is placed into the MAX sequence, and the smaller one is placed into the MIN sequence. Then the obtained MAX and MIN sequences are still bitonic sequences, and any element in the MAX sequence is not less than any element in the MIN sequence.
[0087] Taking ascending sorting as an example, the specific method is as follows: bisect a sequence (1…n), assuming n = 2^k, then compare the 1st element with the (n / 2+1)th element, place the smaller one at the upper position, next compare the 2nd element with the (n / 2+2)th element, place the smaller one at the upper position, and so on; then regard the result as two sequences of length (n / 2), since they are both bitonic sequences, the above process can be repeated; the process is repeated for k rounds in total, that is, the final round compares sequences of length 2, and the final sorting result can be obtained.
[0088] 3.2: Scala Design of Bitonic Sorting
[0089] Figure 8 The sequentially executed functions during the compilation process under Scala are provided. Various functions are first written in the class, and these functions have recursion that is call relationship. Second, the first function bitonalMerge (bitonic merge) is called in the class to implement the overall structure.
[0090] bitonalMerge calls the bitonalSort (bitonic sort) function to implement one major step of sorting, and then detects whether the recursion has reached the final step of sorting. If the final step is not reached, it continues to recursively call itself and increments the maximum number of steps max_step by 1. Thus, there are log_2N recursive calls. In the function bitonalSort, first, the input array is divided into several sub-arrays according to different values of max_step, then the comp (compare) function is called several times for sorting, and second, the function bitonalSort also has a recursive call relationship, recursing sequentially from max_step to 1. Both the comp function and the unitcomp (unit-level compare) function sort the input data. The comp function receives variable-length data and calls the unitcomp function to perform pairwise sorting, then returns the result to the bitonalSort function that calls it. After several sorting steps, bitonalMerge executes to the final step and outputs the final sorted data. In RTL, all these function executions are compiled into gate circuits, and there are no operations such as calls and recursion. The function call relationship here is actually the execution order during compiler compilation.
[0091] As Figure 9As shown, the system pipeline design is related to `pipe_step` (pipe number of pipeline iterations). `pipe_step` bitonalSorts are connected to form a single-level pipeline. There are a total of `bitonalSort`s in the system. There are N = log2 length. When pipe_step is a factor of M, that is, when the number of sorting units contained in each level is the same, the pipeline time of each level is balanced.
[0092] 4: System-level optimization
[0093] like Figure 10 As shown, the system-level optimization of RMD is mainly reflected in pipeline and parallelism. In the pixel block processing flow, the calculation of SATD cost, accumulation of SATD cost, and sorting process adopt a pipeline design, which saves a lot of area resources and shortens the implementation time. In addition, parallel processing design is adopted between different modes of blocks of the same size, which not only shortens the implementation time but also meets the timing requirements of each stage of the pipeline.
[0094] This invention replaces the inverse quantization and inverse transform processes in the original RMD process with a rate-distortion estimation method. Finally, we set the quantization parameter QP to {22, 27, 32, 37} and use BD-Rate to evaluate coding performance, with negative values representing performance gains. The formula for measuring the savings in RMD computation time is as follows:
[0095]
[0096] In the formula: T RMD_ORIG and T RMD_EST Table 1 shows the original total coding time and the coding time of the improved RMD algorithm, respectively. The results for full I-frames and LDP frames are listed in Table 1. Overall, the coding loss for full I-frames and LDP frames is 1.98% and 3.2%, respectively, while the optimized RMD process reduces computation time by 62.5% and 64%, respectively. Furthermore, we also present the estimated reduction in coding time for the entire coding process on the far right of Table 1. Compared to the original coding, the average coding time is reduced by approximately 27%.
[0097] Table 1. Encoding performance and time reduction under different configurations
[0098]
[0099] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hardware architecture method for intra-frame coarse-mode decision optimization in video, characterized in that... Includes the following steps: Step S1: Obtain pixel blocks of different sizes within the video frame; Step S2: For pixel blocks of different sizes, perform intra-frame coarse-mode prediction. In intra-frame coarse-mode prediction, firstly, a first round of decision-making is performed based on a set of intra-frame prediction modes, and several best intra-frame prediction modes in the first round are selected. In the second round of decision-making, it is first determined whether there is a directional intra-frame prediction mode in the first round of decision-making. If there is, several best intra-frame prediction modes in the second round are selected as the final decision result based on several best intra-frame prediction modes in the first round and the existing directional intra-frame prediction modes. Otherwise, the result obtained from the first round of decision-making is directly used as the final decision result. Based on the intra-predictor corresponding to pixel blocks of different sizes, the predicted pixel value is obtained. The predicted value and the original pixel value are calculated to obtain the residual value. The SATD value is obtained based on the residual value. The best intra-prediction mode is selected by sorting the SATD values. The SATD values are bitonically sorted, and the best intra-frame prediction modes are selected as the final decision results. In the bitonic sorting, the sequence of n SATD values is divided in half, assuming n=2^k. Then the 1st and n / 2+1th SATD values are compared. If they are in ascending order, the smaller one is placed on top, and if they are in descending order, the larger one is placed on top. Then the 2nd and n / 2+2nd are compared, and so on. For the two sequences of length n / 2 formed, the above process is repeated for a total of k rounds. That is, the last round is a comparison of sequences of length 2, and the final sorting result is obtained. The selection of the optimal intra-frame prediction mode is based on a hardware framework constructed using bitonic sorting. First, functions are written and their recursive and calling relationships are established. A bitonic merging function calls a bitonic sorting function to perform large-step sorting. Then, it checks if the recursive call is at the last step; if not, it continues recursively calling itself and increments the maximum count. In the bitonic sorting function, the input array is divided into smaller arrays based on the maximum count, and then a comparison function is called several times for sorting. Similarly, the bitonic sorting function also has a recursive calling relationship from the maximum count down to 1. The comparison function receives data of variable length and calls unit-level comparison functions for pairwise sorting, finally returning to the bitonic sorting function that called it. Finally, the bitonic merging function executes to the last step and outputs the final sorted data. The written functions are compiled into gate circuits, and the function calling relationships follow the compiler's execution order. In the pixel block processing flow, the calculation, accumulation, and sorting of SATD values are carried out using a pipeline design; in addition, parallel processing is used between different modes of blocks of the same size.
2. The hardware architecture method for intra-frame coarse-mode decision optimization according to claim 1, characterized in that: The pipelining design of the bitonic sorting function in hardware implementation is related to the number of pipesteps. `pipe_step` bitonic sorting functions are connected to form a single-stage pipelining. The total number of bitonic sorting functions is... indivual, Where length represents the length of the array to be sorted, and pipe_step is a factor of M, that is, when the number of sorting units contained in each level is the same, the pipeline time of each level is balanced.
3. A method for coarse-mode decision optimization within AV1 video frames, characterized in that: Based on the video frame coarse mode decision optimization hardware architecture method according to claim 1, the coding tree unit of the video frame is defined as a superblock. From the different pixel block partitioning modes that constitute each superblock, the optimal block partitioning mode is predicted. The intra-frame prediction mode includes non-directional intra-frame prediction mode and directional intra-frame prediction mode, wherein the directional intra-frame prediction mode can be further subdivided into multiple auxiliary directional mode predictions.
4. The AV1 video frame intra-coarse mode decision optimization method according to claim 3, characterized in that: The non-directional intra-frame prediction modes include DC mode, Paeth mode, Smooth mode, Smooth_vertical mode, and Smooth_horizontal mode. The Paeth mode predicts blocks by performing sample-level copying from a reference sample, as shown in the following formula: Where L represents the reference sample pixel value based on the vertical direction of the pixel block, A represents the reference sample pixel value based on the horizontal direction of the pixel block, M represents the pixel value based on a single reference sample of the pixel block, B is the base value, representing the lowest value among the three reference sample pixel values, and p L p A p M These are intermediate variables, Lmin(p) L ,p A ,p M )=p L p L It is p L p A p M The minimum value in, Amin(p) L ,p A ,p M )=p A p A It is p L p A p M The minimum value in, Mmin(p) L ,p A ,p M )=p M p M It is p L p A p M The minimum value in the range, where P represents the pixel value of the predicted sample; The Smooth mode makes predictions by applying bilinear interpolation in the vertical and horizontal directions based on the distance between the predicted sample and the reference sample. It includes a smoothed vertical mode, a smoothed horizontal mode, and a combination of the two smoothing modes, as shown in the following formula: in, This indicates that the smooth vertical application uses interpolation obtained by using the reference sample pixel values corresponding to each column of A-RSV and L-RSV. The smoothing level is represented by the interpolation obtained by using the reference sample pixel values of each row of L-RSV and A-RSV. L-RSV and A-RSV represent the reference sample pixel values in the vertical and horizontal directions of the pixel block, respectively. M represents the number of pixels in the height of the pixel block, N represents the number of pixels in the width of the pixel block, (x,y) represents the position of the predicted sample in the pixel block, and 255 represents the maximum pixel value. The DC pattern is used to predict a block in which all samples are predicted as the average of the available RSV samples. This average is also called the DC value, and the formula is as follows: 。 5. The AV1 video frame intra-coarse mode decision optimization method according to claim 3, characterized in that: The intra-frame prediction mode, based on a defined prediction angle, uses the base angle B and the increment angle Δ to represent multiple prediction angles, including prediction angle P: P = B + 3 × Δ Use only L-RSV for angles greater than 180°, use only A-RSV for angles less than 90°, and use all three RSVs for angles between 180° and 90°.
6. The AV1 video frame intra-coarse mode decision optimization method according to claim 5, characterized in that: Angle prediction uses interpolation to define the prediction sample; for angles less than 90°, formula (10) is used; for angles greater than 180°, formula (10) is used; for angles between 90° and 180°, formulas (9) and (10) are used: (9) (10) Where x and y represent the horizontal and vertical coordinates of the predicted sample in the pixel block. This represents the horizontal reference sample pixel value with the same x-coordinate as the predicted pixel. This represents the reference sample pixel value in the vertical direction, which shares the same ordinate as the predicted pixel. and This indicates the parameter value obtained from a predefined table.