Management method and apparatus for adaptation parameter set, and electronic device and storage medium
By adaptively correcting the ALF decision factor, increasing the probability of using new APS and storing it to the historical candidate set, the problem of insufficient filling degree of the ALF APS historical candidate set is solved, and the encoding efficiency and image quality of video encoding are improved.
Patent Information
- Application Number
- PCT/CN2024/131993
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-11-14
- Publication Date
- 2025-07-17
AI Technical Summary
In the existing ALF technology, the decision-making process of the filtering mode is independent and does not fully consider the video characteristics, resulting in insufficient filling of the ALF APS historical candidate set, affecting the encoding efficiency and image quality.
By obtaining the categories and characteristics of the video blocks, the decision factor is adaptively corrected to increase the probability of new APS usage in the filtering mode and storing it to the historical candidate set, optimizing the ALF decision process.
It improves the filling degree of the ALF APS historical candidate set, enhances the optimization effect of filtering technology, and improves encoding efficiency and image quality.
Smart Images

Figure CN2024131993_17072025_PF_FP_ABST
Abstract
Description
Adaptive parameter set management method and device, electronic device, and storage medium
[0001] This disclosure claims priority to Chinese patent application No. 202410047381.2, filed on January 11, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to the field of video coding technology, and in particular to a method and device for managing an adaptive parameter set, an electronic device, and a storage medium. Background Art
[0003] Adaptive loop filter (ALF) technology includes four types of filtering modes: chroma ALF, luma ALF, cross-component adaptive loop filter (CCALF) Cb (the difference between the blue component and luma), and CCALF Cr (the difference between the red component and luma). Each filtering mode corresponds to an adaptation parameter set (APS) containing filtering parameters.
[0004] Currently, when performing ALF processing on a video, each filtering mode determines the corresponding decision factor based on the current video block in the video, and determines multiple new APSs used by the current video block when performing ALF processing based on multiple decision factors. Then, the multiple historical APSs used by the historical video blocks in the ALF APS historical candidate set are replaced with the multiple new APSs used by the current video block, providing a filtering reference for subsequent video blocks in the video.
[0005] Summary of the Invention
[0006] On the one hand, a method for managing an adaptive parameter set is provided. The method for managing an adaptive parameter set includes: a management device for an adaptive parameter set (hereinafter referred to as the "management device") obtains a target video category and / or target video feature corresponding to a target video block. The management device determines a correction factor for each decision factor based on the target video category and / or target video feature, wherein one decision factor corresponds to a preset filtering mode, and the decision factor is used to indicate whether to filter the video block, and to indicate the use of a new APS, a historical APS in a historical candidate set, or any one of the preset APSs during filtering. The management device corrects each decision factor based on the correction factor of each decision factor to obtain a plurality of corrected decision factors, wherein the probability of one corrected decision factor indicating filtering is greater than the probability of one decision factor indicating filtering before correction, and the probability of one corrected decision factor indicating the use of a new APS during filtering is greater than the probability of one decision factor indicating the use of a new APS during filtering before correction. The management device stores the obtained at least one new APS into the historical candidate set based on each corrected decision factor.
[0007] In another aspect, a device for managing an adaptive parameter set is provided, which includes an acquisition module, a processing module, and a storage module.
[0008] The acquisition module is configured to acquire the target video category and / or target video features corresponding to the target video block. The processing module is configured to determine a correction factor for each decision factor based on the target video category and / or target video features, wherein one decision factor corresponds to a preset filtering mode and indicates whether the video block is to be filtered, and the probability of using any of the new APS, the historical APS in the historical candidate set, and the preset APS during filtering. The processing module is further configured to correct each decision factor based on its correction factor to obtain multiple corrected decision factors, wherein the probability of filtering indicated by one corrected decision factor is greater than the probability of filtering indicated by the decision factor before correction, and the probability of using the new APS indicated by one corrected decision factor during filtering is greater than the probability of using the new APS indicated by the decision factor before correction. The storage module is configured to store at least one new APS obtained based on each corrected decision factor in the historical candidate set.
[0009] In yet another aspect, an electronic device is provided. The electronic device includes a memory and a processor. The memory and the processor are coupled. The memory is configured to store a computer program. When the processor executes the computer program, the adaptive parameter set management method described in the above aspect is implemented.
[0010] In yet another aspect, a computer-readable storage medium is provided, wherein computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the method for managing the adaptive parameter set described in the above aspect is implemented.
[0011] In yet another aspect, a computer program product is provided, including computer program instructions, which, when executed by a processor, implement the method for managing the adaptive parameter set described in the above aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the technical solutions of the present disclosure, the following briefly introduces the drawings required for use in some embodiments of the present disclosure. Obviously, the drawings described below are only drawings of some embodiments of the present disclosure, and those skilled in the art can also derive other drawings based on these drawings.
[0013] FIG1 is a structural block diagram of a video encoder according to some embodiments of the present disclosure.
[0014] FIG2 is a structural block diagram of a video decoder according to some embodiments of the present disclosure.
[0015] FIG3 is a schematic diagram of a flow chart of loop filtering in the H.266 / VVC standard according to some embodiments of the present disclosure.
[0016] FIG4 is a schematic diagram illustrating an example of an internal implementation of an ALF APS according to some embodiments of the present disclosure.
[0017] FIG5 is a schematic diagram illustrating an example of ALF decision candidates according to some embodiments of the present disclosure.
[0018] FIG6 is a schematic diagram illustrating an example of managing an APS according to some embodiments of the present disclosure.
[0019] FIG7 is a schematic diagram of a communication system according to some embodiments of the present disclosure.
[0020] FIG8 is a flowchart of a method for managing an adaptive parameter set according to some embodiments of the present disclosure.
[0021] FIG9 is a schematic diagram illustrating an example of an ALF APS according to some embodiments of the present disclosure.
[0022] FIG10 is a schematic diagram illustrating another example of an ALF APS according to some embodiments of the present disclosure.
[0023] FIG11 is a schematic diagram illustrating an example of an ALF APS historical candidate set according to some embodiments of the present disclosure.
[0024] 12A-12B are schematic diagrams illustrating another example of managing an APS according to some embodiments of the present disclosure.
[0025] 13A-13B are schematic diagrams illustrating another example of managing an APS according to some embodiments of the present disclosure.
[0026] FIG14 is a schematic structural diagram of a device for managing an adaptive parameter set according to some embodiments of the present disclosure.
[0027] FIG15 is a schematic structural diagram of a management device for an adaptive parameter set according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0028] To help those skilled in the art better understand the technical solutions of the embodiments of the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below in conjunction with the drawings in the present disclosure. Obviously, the embodiments described are only some of the embodiments of the present disclosure, not all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0029] It should be noted that in this disclosure, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this disclosure as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs in this disclosure. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.
[0030] In the following, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features being referred to. Therefore, a feature defined with the terms "first," "second," etc., may explicitly or implicitly include one or more of such features.
[0031] In the description of this disclosure, unless otherwise specified, " / " means "or." For example, A / B can mean A or B. "And / or" herein is merely a description of an association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: only A, only B, and both A and B. Furthermore, "at least one" means one or more, and "a plurality" means two or more.
[0032] Before introducing in detail the method for managing an adaptive parameter set provided by an embodiment of the present disclosure, the implementation environment and application scenarios of the embodiment of the present disclosure are first introduced.
[0033] First, the application scenarios of the embodiments of the present disclosure are introduced.
[0034] Improving video coding compression efficiency has long been an effective way to enhance video service quality. Video coding and decoding technology has undergone continuous iteration and evolution over the past 30 years. The released H.266 / Versatile Video Coding (VVC) standard boasts industry-leading coding efficiency. ALF technology is a new filtering technology added to the H.266 / VVC standard. In ALF technology, the ALF APS is used to transmit ALF filter coefficients. The ALF APS can include filter coefficients corresponding to four filtering modes: luma ALF APS, chroma ALF APS, CCALF Cb APS, and CCALF Cr APS. In current ALF technology, the decision process for whether to adopt the new APS for each of the four filtering modes—luma ALF, chroma ALF, CCALF Cb, and CCALF Cr—is relatively independent, and is determined using the decision factors corresponding to each filtering mode. The H.266 / VVC coding framework, a next-generation video coding standard developed by the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) in a joint video project, includes modules such as intra-frame prediction, inter-frame prediction, transform, quantization, loop filtering, and entropy coding.
[0035] Figure 1 is a block diagram of the video encoder provided by one embodiment. As shown in Figure 1, the overall framework process of the encoding end is as follows. During encoding, the input video is divided into video blocks for subsequent processing. This description uses coding tree units (CTUs) as an example. Of course, it is understood that the implementation process is similar when the video blocks are of other types.
[0036] (1) The input video is first divided into frames and CTU partitioning is performed.
[0037] (2) The divided CTU is sent to the intra-frame / inter-frame prediction module for predictive coding. The intra-frame prediction module is mainly used to remove the spatial correlation of the image. It predicts the current pixel block through the encoded reconstructed CTU information to remove spatial redundant information. The inter-frame prediction module is mainly used to remove the temporal correlation of the image. It uses the encoded image as the reference image of the current frame to obtain the motion information of each CTU, thereby removing temporal redundancy.
[0038] (3) The resulting intra / inter prediction value is then subtracted from the original CTU to obtain a residual value. This residual value is then transformed and quantized to remove frequency domain correlation and perform lossy compression on the data. Transform coding transforms the image from a spatial domain signal to the frequency domain, concentrating energy in low-frequency regions. The quantization module can reduce the dynamic range of image coding.
[0039] (4) All coding parameters and residual values are then entropy-coded to form a binary stream for storage or transmission. The output data of the entropy coding module is the compressed bitstream of the original video. Figure 1 shows entropy coding as context-based adaptive binary arithmetic coding (CABAC) and header information coding.
[0040] (5) The intra-frame / inter-frame prediction value and the residual value after inverse quantization and inverse transformation are added to obtain the block reconstruction value, and finally form the reconstructed image.
[0041] (6) The reconstructed image is filtered through a loop filter and stored in the image cache as a reference image in the future.
[0042] H.266 / VVC uses a CTU-based hybrid coding framework. Distortion effects such as blocking artifacts, ringing artifacts, color deviation, and image blurring still exist in videos compressed using the H.266 / VVC standard. To reduce the impact of these distortions on video quality, H.266 / VVC uses in-loop filtering. In-loop filtering technologies in H.266 / VVC include luma mapping with chroma scaling (LMCS), deblocking filter (DBF), sample adaptive offset (SAO), and ALF. LMCS improves compression efficiency by reallocating codewords across the dynamic range; DBF reduces blocking artifacts; SAO improves ringing artifacts; and ALF reduces decoding errors.
[0043] Figure 2 is a structural block diagram of a video decoder provided by an embodiment. As shown in Figure 2, the framework process of the video decoding end is as follows.
[0044] (1) Analyze the code stream to obtain the prediction mode and get the intra / inter prediction value.
[0045] (2) Perform inverse transformation and inverse quantization on the residual value obtained from bitstream analysis.
[0046] Header information decoding and CABAC decoding are required during the parsing process.
[0047] (3) The intra / inter prediction value and the residual value after inverse quantization and inverse transformation are added to obtain the CTU reconstruction value, and finally form the reconstructed image.
[0048] (4) The reconstructed image is filtered through a loop filter and stored in the image cache as a reference image in the future.
[0049] Figure 3 is a schematic diagram of the loop filtering process in the H.266 / VVC standard. As shown in Figure 3, whether on the encoding or decoding end, the connection order of the various filter modules in the loop filtering technology is: LMCS module, DBF module, SAO module, and ALF module. In other words, the filtering order is: LMCS, DBF, SAO, and ALF.
[0050] Figure 4 is a schematic diagram of an example of the internal implementation of the ALF APS. The ALF APS is used to transmit ALF filter coefficients, which can also be called ALF filter parameters. In this disclosure, ALF filter coefficients and ALF filter parameters are referred to the same. The ALF APS can include filter coefficients corresponding to four filter modes: luma ALF APS, chroma ALF APS, CCALF Cb APS, and CCALF Cr APS. The ALF APS can include four modules: luma ALF, chroma ALF, CCALF Cb, and CCALF Cr. Whether the new APS is used in each module is independent of each other, and each module has a corresponding identifier to indicate whether the new APS is used.
[0051] alf_luma_filter_signal_flag: indicates whether there are luma ALF-related filter parameters in the current ALF APS. If the value of this identifier is equal to the first preset value, it is present; if it is equal to the second preset value, it is not present.
[0052] Exemplarily, the first preset value may be 1, and the second preset value may be 0.
[0053] alf_chroma_filter_signal_flag: indicates whether there are chroma ALF-related filter parameters in the current ALF APS. If the value of this identifier is equal to the first preset value, it means there are chroma ALF-related filter parameters, and if it is equal to the second preset value, it means there are no chroma ALF-related filter parameters.
[0054] alf_cc_cb_filter_signal_flag: indicates whether the CCALF Cb mode related filtering parameters exist in the current ALF APS. If the value of this identifier is equal to the first preset value, it exists; if it is equal to the second preset value, it does not exist.
[0055] alf_cc_cr_filter_signal_flag: indicates whether the current ALF APS contains CCALF Cr mode related filtering parameters. If the value of this identifier is equal to the first preset value, it is present; if it is equal to the second preset value, it is not present.
[0056] It should be noted that whether the current ALF APS contains filtering parameters related to each filtering mode refers to whether a new APS related to each filtering mode exists within the current ALF APS. In other words, whether a new APS related to each filtering mode is used. The new APS related to each filtering mode is based on the original frame and the reconstructed frame, using the concept of the Wiener filter, with the minimum mean squared error (MSE) between the original frame and the reconstructed frame as the optimization goal. The Wiener-Hoff formula related to the two is listed, and the ALF filter coefficients are obtained based on the results of the equation solution.
[0057] Whether each filtering mode within the ALF APS uses the corresponding new APS is independent of each other, and needs to be judged by the rate-distortion criterion when making decisions on each filtering mode of the ALF. If all four identifier decisions are the second preset value, the current video block does not use the new APS of the ALF. However, once there is an identifier with the first preset value, the new APS of the ALF of the current video block will be used. In this embodiment, using the new APS of the ALF of the current video block refers to filtering according to the new APS of the ALF to achieve image reconstruction. Furthermore, the new APS of the ALF can also be transmitted in the bitstream. Simply put, for the ALF APS transmitted to the bitstream, there is at least one filtering coefficient of the filtering mode inside it, and the filtering coefficient can be as well as Any one of these four filter coefficients, and there can be at most four filter coefficients, that is, all of the above four filter coefficients exist.
[0058] During the ALF operation, the ALF APS historical candidate set is also maintained. The ALF APS historical candidate set plays a major role in the acquisition of historical APS. When a new APS is determined to be used, the ALF APS historical candidate set is also updated according to the first-in-first-out rule. The acquisition of historical APS in ALF depends on the ALF APS historical candidate set. The number of historical sets of the current video block can be calculated through the historical candidate set. When obtaining the historical set candidate, it is necessary to determine whether the APS in the historical candidate set can be used by the current video block. The judgment conditions are:
[0059] (1) The time layer of the APS is lower than or equal to the time layer of the current video block;
[0060] (2) There is a filter set corresponding to the ALF mode in the APS.
[0061] Only the APS that meets the above two conditions will be selected as the historical APS of the current video block.
[0062] In addition to the restrictions on the rules for obtaining historical APS, the historical candidate sets will also be different under different encoding configurations.
[0063] In the all intra mode (AI) configuration, all frames are in full intra mode, so the historical APS cannot be used, and therefore no historical candidate set is formed.
[0064] In the random access (RA) configuration, the ALF APS historical candidate set is cleared every time an I frame passes, which limits the effectiveness of the historical APS.
[0065] In low-latency (LD) mode, all video blocks belong to the same temporal layer. Therefore, B / P frames are not restricted by temporal layers when retrieving the history set. In addition, there is no mechanism to clear the history candidate set in LD mode. Transmitted APSs remain in the ALF APS history candidate set unless they are overwritten by a new APS with the same identity (ID).
[0066] ALF decision refers to determining the ALF filter coefficients from the ALF decision candidate list. Figure 5 is a schematic diagram of an example of ALF decision candidate list provided by one embodiment. As shown in Figure 5, the ALF decision candidate list includes the following four options.
[0067] Option 1: When alf_enable_flag = 0, do not perform ALF (skip ALF).
[0068] When alf_enable_flag=1, there are the following three options.
[0069] Option 2, new APS (new ALF APS), requires transmitting the index of the new APS and the filter coefficients of the new APS.
[0070] Option 3, fixed filter sets (offine-trained filter sets), only transmits the index of the fixed filter set. Fixed filter sets are trained offline. When deciding to use a fixed filter set, you also need to determine the index of the fixed filter set.
[0071] Option 4, historical APS, transmits only the index of the historical APS. The historical APS is calculated online from the previous frame.
[0072] The ALF decision process uses the rate-distortion criterion. This process involves the use of decision factors. These factors must be calculated before encoding the current video block, using parameters such as the quantization parameter (QP) configured in the configuration. Luma and chroma differ in their component characteristics, leading to slightly different fixed function settings. Consequently, different ALF filter modes require different decision factors.
[0073] Since mainstream video compression is lossy, the rate-distortion criterion is used as a tool to measure compression performance and is the most widely used technology in video compression encoding. Almost all encoding modules use the rate-distortion criterion to determine the final encoding mode. The rate-distortion criterion is based on the principle of rate distortion optimization (RDO). Under the constraint of the encoding bit rate, the optimization problem is formulated by minimizing the distortion, as shown in Equation 1:
[0074] Where D is the distortion. Since the loop filter no longer performs time-domain and frequency-domain transformations, the distortion D is uniformly expressed as the minimum mean square error (MSE). R is the bitrate, and X can represent the partition structure, prediction mode, transform coefficients, and other factors. In the ALF module, when making slice-level decisions, X is determined within the range of ALF decision candidate identifiers shown in Figure 5, including skipping ALF, adopting a new APS, a fixed filter set, and the historical APS.
[0075] The constrained problem in the above formula can be converted into an unconstrained problem by introducing a decision factor (e.g., Lagrange factor (lambda, λ)), as shown in Formula 2:
[0076] Here, J represents the rate-distortion cost. This cost accurately describes the combined impact of the number of bits required for the current mode and the resulting distortion. For example, to determine the performance of two filter sets, one can simply compare their rate-distortion costs. A lower cost indicates better compression performance.
[0077] Generally, the Lagrangian factor used in video encoding has a fixed functional relationship with the QP of the current slice. The Lagrangian factor must be calculated before encoding the current slice, using parameters such as the configured QP. Due to the different characteristics of luma and chroma components, the settings for this fixed function vary slightly, resulting in different Lagrangian factors for different components.
[0078] Furthermore, existing video coding standards (such as H.265 / High Efficiency Video Coding (HEVC) and H.266 / VVC) all use YUV files as input. Compared to RGB files, YUV files can be compressed during signal acquisition. A pixel in an RGB file consists of three components: R, G, and B, while a pixel in a YUV file consists of three components: Y, U, and V. Because the human visual system is more sensitive to the Y component and less sensitive to the U and V components, formats such as 4:2:0, 4:2:2, and 4:4:4 Y:U:V data ratios have emerged.
[0079] To more comprehensively test the overall performance of the standard reference software, the H.266 / VVC standard's common test sequences include different video sequences with varying resolution, bit depth, and content characteristics. Table 1 shows the differences in these common test sequences. Video resolution measures the pixel width and pixel height of a video, helping to determine video quality and its clarity or fidelity. Video bit depth, often referred to as color depth or quantization depth, primarily represents the number of bits used to represent a single color component of a pixel. This determines the accuracy of color representation in a digital image. Video content characteristics often include texture complexity and motion complexity. Texture generally refers to the locally irregular but macroscopically regular nature of an image. Texture complexity is a spatial domain metric that characterizes content characteristics and includes metrics such as mean, variance, contrast, dissimilarity, correlation, and entropy. Motion generally refers to the variation in image content between video frames. Motion complexity is a temporal domain metric that characterizes content characteristics, including the amount of temporal variation (slow / dramatic) in a video sequence.
[0080] Table 1. Differences in characteristics of common test sequences for H.266 / VVC standards
[0081] In summary, ALF technology utilizes the concept of the Wiener filter, taking the mean square error (MSE) between the original and reconstructed frames as the optimization target. The Wiener-Hoff equation relating the two is formulated, and the final ALF filter coefficients are obtained based on the solution. This more directly reduces the MSE, thereby improving the image's peak signal-to-noise ratio (PSNR). In ALF technology, the ALF filter coefficients are transmitted using an ALF Adaptive Parameter Set (APS). The ALF APS includes filter coefficients corresponding to four filtering modes: luma ALF APS, chroma ALF APS, cross-component adaptive loop filtering (CCALF) (Cb APS), and CCALF Cr APS.
[0082] In current ALF technology, the decision process for adopting a new APS for each of the four filter modes—luma ALF, chroma ALF, CCALF Cb, and CCALF Cr—is relatively independent. Each filter mode is compared and judged based on its own performance. This decision is made using the decision factors corresponding to each filter mode. However, when transmitting ALF-related APSs, different filter modes share the same ALF APS identifier (ID). If all four APSs are determined to be transmitted for the same slice, then the four APSs share a single ID, maximizing the benefits of this APS ID. In contrast, if only one APS is transmitted within an ALF APS ID, then the ID corresponds to only one APS and may also overwrite several APSs previously associated with the ID. This may, in turn, reduce the number of APSs corresponding to certain ALF modes in the historical ALF APS candidate set. When the ALF APS is updated, if the number of APS types transmitted within the same ALF APS ID is small, the historical candidate set in the ALF decision process will be reduced, and the ALF function cannot be fully utilized.
[0083] For example, as shown in FIG6 , the ALF encoding process in the existing mechanism is as follows.
[0084] (1) The encoder calculates the new APS of the luma ALF and the new APS of the chroma ALF for the current slice.
[0085] (2) The encoder determines whether the brightness ALF decision is enabled.
[0086] If the luma ALF decision is not enabled, the encoder does not perform ALF filtering.
[0087] (3) If the luminance ALF decision is turned on, the encoder performs the luminance ALF decision.
[0088] Perform RDO decision without ALF, new APS, fixed filter set, or historical APS. If the decision result uses the new luma APS, alf_luma_filter_signal_flag needs to be set to 1, otherwise it should be set to 0.
[0089] (4) If the luma ALF decision is on, the encoder makes a chroma ALF decision.
[0090] Perform RDO decision without ALF, new APS, or historical APS. If the decision result uses the new chroma APS, alf_chroma_filter_signal_flag needs to be set to 1, otherwise it should be set to 0.
[0091] (5) The encoder reconstructs the image after the luminance ALF and chrominance ALF according to the decision results of the luminance ALF and chrominance ALF.
[0092] (6) The encoder calculates the new APS of CCALF Cb and CCALF Cr for the current slice.
[0093] (7) If the luminance ALF decision is on, the encoder makes a CCALF Cb decision.
[0094] Perform RDO decision when not performing ALF, new APS, or historical APS. If the decision result uses CCALF Cb new APS, alf_cc_cb_filter_signal_flag needs to be set to 1, otherwise it should be set to 0.
[0095] (8) If the luminance ALF decision is on, the encoder makes a CCALF Cr decision.
[0096] Perform RDO decision when not performing ALF, new APS, or historical APS. If the decision result uses the CCALF Cr new APS, alf_cc_cr_filter_signal_flag needs to be set to 1, otherwise it should be set to 0.
[0097] (9) The encoder reconstructs the image after CCALF action based on the decision results of CCALF Cb and CCALF Cr.
[0098] (10) The encoding end determines whether at least one of the identifiers of the multiple filtering modes is 1.
[0099] The identifiers of the multiple filter modes include: alf_luma_filter_signal_flag, alf_chroma_filter_signal_flag, alf_cc_cb_filter_signal_flag, alf_cc_cr_filter_signal_flag.
[0100] That is, alf_luma_filter_signal_flag∨alf_chroma_filter_signal_flag∨alf_cc_cb_filter_signal_flag∨alf_cc_cr_filter_signal_flag==1.
[0101] If at least one of the above four identifiers is 1, the encoding end executes (11).
[0102] (11) The encoding end transmits the current new APS.
[0103] That is, if at least one of the four identifiers is 1, the encoder transmits the current new APS to the bitstream.
[0104] It is understandable that in the existing mechanism, the decision factor corresponding to one of the four filtering modes may indicate that a new APS is not adopted, resulting in a small number of new APSs adopted by the video block. Consequently, after the multiple new APSs adopted by the video block are stored in the ALF APS historical candidate set, the number of APSs corresponding to a certain filtering mode in the ALF APS historical candidate set is reduced. This poor design of the ALF decision factor results in a large number of luma ALF APSs being explicitly transmitted, while the chroma-related ALF APSs (chroma ALF APS, CCALF Cb APS, and CCALF Cr APS) are missing, resulting in a low fill level in the ALF APS historical candidate set.
[0105] Furthermore, while ALF can effectively improve the subjective and objective quality of encoded and decoded video, existing mechanisms do not consider the impact of video characteristics on ALF decision factors during ALF decision-making and ALF APS update. Therefore, when optimizing the design of ALF decision factors, fully considering video classification or individual video characteristics can more specifically improve the efficiency of ALF APS utilization.
[0106] In order to solve the above problems, the embodiments of the present disclosure provide a method for managing an adaptive parameter set. The method for managing an adaptive parameter set provided by the embodiments of the present disclosure is applied to various filtering optimization scenarios related to video coding, including but not limited to: video transmission and storage, video conferencing systems, video surveillance systems, etc.
[0107] The present disclosure is described using H.266 / VVC as an example, and is applicable to any video encoding and decoding scheme that includes an ALF parameter set. Based on the ALF mechanism in the existing VVC, the ALF APS has the problem of insufficient fullness, especially for components related to chroma, including chroma APS, CCALF Cb APS, and CCALF Cr APS. Therefore, the present disclosure sets appropriate constraints and trigger conditions to recommend the transmission of luminance APS, chroma APS, CCALF Cb APS, and CCALF Cr APS as much as possible, thereby increasing the fullness of the ALF APS history candidate sets of different components, increasing the number of ALF APS history sets, and enhancing the use effect of ALF. For example, according to the decision result of the luminance ALF, such as when it has been confirmed that the luminance APS is to be transmitted, the Lagrangian factor in the ALF related to chroma can be adaptively adjusted based on the video features to optimize the ALF decision related to chroma.
[0108] That is, since there are many gaps in the APS historical candidate sets in different current filtering modes (such as luma ALF, chroma ALF, CCALF Cb, and CCALF Cr), the image optimization effect of the filtering technology during video encoding is not fully exerted. Therefore, the embodiment of the present disclosure proposes to adaptively correct the Lagrangian factors in different filtering modes based on the ALF decision results and / or the fullness of the relevant APS historical candidate sets, according to the video content and / or video features, so that the corrected Lagrangian factors are more inclined to perform filtering during the RDO decision process, and tend to use new APSs during filtering, ensuring that the number of new APSs of the video block under different filtering modes is large, and after all new APSs of the video block are added to the ALF APS historical candidate set, the number of APSs corresponding to a certain filtering mode in the ALF APS historical candidate set is avoided from being reduced, so as to improve the fullness of the ALF APS historical candidate set and maximize the optimization effect of the filtering technology, thereby further improving the coding efficiency and image quality.
[0109] It should be noted that the present disclosure is not limited to the H.266 / VVC standard.
[0110] The implementation environment of the embodiments of the present disclosure is introduced below.
[0111] As shown in FIG7 , which is a schematic diagram of a communication system provided by an embodiment of the present disclosure, the communication system may include: a management device 701 and a collection device 702. The management device 701 may perform wired / wireless communication with the collection device 702.
[0112] The acquisition device 702 can acquire videos and send the acquired videos to the management device 701 .
[0113] The management device 701 can obtain a video block from the video captured by the acquisition device 702 and determine the video category and video features corresponding to the video block. Next, the management device 701 can modify the decision factors for each filtering mode based on the video category and video features corresponding to the video block, so that the modified decision factors are more likely to favor filtering and the use of new APSs during filtering. The management device 701 can then store at least one new APS obtained based on each modified decision factor in the ALF APS historical candidate set.
[0114] It should be noted that the embodiments of the present disclosure do not limit the acquisition device 702. For example, the acquisition device 702 may be a driving recorder. For another example, the acquisition device 702 may be a mobile phone with a video capture function. For another example, the acquisition device 702 may be a surveillance camera.
[0115] In the embodiment of the present disclosure, the management device 701 may be a terminal, or the management device 701 may be a server.
[0116] The terminal may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook computer, or other device with transceiver functions. The embodiments of the present disclosure do not impose any particular restrictions on the specific form of the terminal. The terminal can interact with the user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device.
[0117] The server can be a single physical server, or a server cluster consisting of multiple servers. Alternatively, the server cluster can be a distributed cluster. Alternatively, the server can be a cloud server. The embodiments of this disclosure do not limit the specific implementation of the server.
[0118] After introducing the application scenario and implementation environment of the embodiment of the present disclosure, the management method of the adaptive parameter set provided by the embodiment of the present disclosure is described in detail below in combination with the above implementation environment.
[0119] An embodiment of the present disclosure provides a method for managing an adaptive parameter set. As shown in FIG8 , the method for managing an adaptive parameter set may include: S801 - S804.
[0120] S801: The management device obtains a target video category and / or target video features corresponding to a target video block.
[0121] In one implementation, the target video category is the video category of the target video block, and the target video feature is the video feature of the target video block.
[0122] That is, each time a video block is filtered, the present disclosure determines the video category and video features of the filtered video block, thereby providing a valuable reference for subsequent correction of decision factors for different filtering modes and improving the accuracy of the correction results.
[0123] Optionally, the target video block may be a video block in the video to be processed, the target video category may be the category of the video to be processed, and the target video feature may be the video feature of the video to be processed.
[0124] In other words, when filtering all video blocks in a video, the present disclosure only needs to determine the video category and video features of the entire video segment, and then use the video category and video features of the entire video segment as the video category and video features corresponding to each video block in the video. This can reduce the number of times the video category and video features are determined, and improve the efficiency of subsequent correction of decision factors.
[0125] It should be noted that the embodiment of the present disclosure does not limit the type of the target video block. The target video block can be any of the following types: sub-picture, tile, slice, coding tree unit (CTU), coding unit (CU).
[0126] That is to say, the embodiments of the present disclosure are applicable to different types of filtering objects, thereby increasing the correction range of subsequent decision factors.
[0127] S802: The management device determines a correction factor for each decision factor according to the target video category and / or target video features.
[0128] A decision factor corresponds to a preset filtering mode, and the decision factor is used to indicate whether to filter the video block and to indicate whether to use a new APS, a historical APS in a historical candidate set, or a preset APS (i.e., a fixed filtering set) during filtering.
[0129] In the embodiment of the present disclosure, the preset filtering mode is any one of the following ALFs: chroma ALF, luma ALF, CCALF Cb, and CCALF Cr.
[0130] That is to say, the embodiment of the present disclosure is applicable to the ALF process of the video, and modifies the decision factors of different ALF modes when the video undergoes ALF, so that the modified decision factors can provide richer APS for subsequent ALF and improve the accuracy of the video filtering results.
[0131] In one implementation, the decision factor may be λ in the RDO decision process.
[0132] That is, when filtering a video block, the management device can determine whether filtering is required for the video block and which APS to use when filtering the video block through RDO decision-making. In this way, the efficiency of video filtering can be improved and the accuracy of the video filtering results can be improved.
[0133] In one implementation, the management device stores multiple preset correction factors, multiple preset video categories, and / or multiple preset video features, with each preset video category corresponding to one preset correction factor, and each preset video feature corresponding to one preset correction factor. The management device may determine, based on a target video category, a preset video category identical to the target video category from the multiple preset video categories, and use the preset correction factors corresponding to the preset video categories identical to the target video category from the multiple preset video categories as correction factors for each decision factor, thereby determining the correction factor for each decision factor.
[0134] Optionally, the management device may determine, based on the target video feature, a preset video feature that is identical to the target video feature from a plurality of preset video features, and use the preset correction factor corresponding to the preset video feature that is identical to the target video feature in the plurality of preset video features as the correction factor for each decision factor to determine the correction factor for each decision factor.
[0135] In one implementation, the management device stores multiple preset correction factors, multiple preset video weights, multiple preset video categories, multiple preset video features, multiple preset category weights, and multiple preset feature weights, where each preset video weight corresponds to one preset correction factor, each preset category weight corresponds to one preset video category, and each preset feature weight corresponds to one preset video feature. The management device can determine, based on the target video category, a preset video category identical to the target video category from among the multiple preset video categories, and use the preset category weight corresponding to the preset video category identical to the target video category among the multiple preset video categories as the target category weight corresponding to the target video category.
[0136] Similarly, the management device can determine the preset video features that are the same as the target video features from multiple preset video features based on the target video features, and use the preset feature weights corresponding to the preset video features that are the same as the target video features in the multiple preset video features as the target feature weights corresponding to the target video category.
[0137] Next, the management device determines a target video weight for the target video block by summing the target category weight and the target feature weight. The management device can then determine, based on the target video weight, a preset video weight identical to the target video weight from a plurality of preset video weights, and use the preset correction factor corresponding to the preset video weight identical to the target video weight from the plurality of preset video weights as the correction factor for each decision factor to determine the correction factor for each decision factor.
[0138] S803: The management device modifies each decision factor according to the correction factor of each decision factor to obtain multiple modified decision factors.
[0139] A modified decision factor indicates a probability of performing filtering that is greater than the probability indicated by the decision factor before the modification, and a modified decision factor indicates a probability of using the new APS when filtering that is greater than the probability indicated by the decision factor before the modification.
[0140] In one implementation, the management device may determine the corrected decision factor by calculating the product between the decision factor and the corresponding correction factor.
[0141] The following describes the impact of the decision factors before and after correction on the APS used in the filtering mode with some examples.
[0142] For example, using the ALF filtering mode as an example, Figure 9 is a schematic diagram of the ALF APS provided by one embodiment, illustrating the APSs of different ALF modes determined by the decision factors before modification. A check mark (√) indicates the use of the new APS, and a cross (×) indicates the non-use of the new APS. This means that only the chroma ALF APS uses the new APS; the other three filtering modes do not. Figure 10 is a schematic diagram of the ALF APS provided by another embodiment, illustrating the APSs of different ALF modes determined by the decision factors after modification for a target video block. All four filtering modes use the new APS.
[0143] In other words, the ALF APS obtained in Figure 10 has three new APSs compared to Figure 9. These three new APSs can then be updated to the historical filtering parameters corresponding to the filtering mode, thereby increasing the richness of the historical candidate set.
[0144] S804: The management device stores the obtained at least one new APS into the historical candidate set according to each revised decision factor.
[0145] In one implementation, the management device may determine, based on each revised decision factor, multiple first filtering modes to be executed by the target video block from among multiple preset filtering modes through RDO decision making. Based on each revised decision factor, the management device may determine, from among the multiple first filtering modes, at least one second filtering mode that uses a new APS, and obtain the new APS used by each second filtering mode to obtain at least one new APS. The management device may then store the obtained at least one new APS in the historical candidate set.
[0146] In one implementation, the historical candidate set may be an ALF APS historical candidate set, which may include multiple historical APSs and multiple preset identifiers, one preset identifier corresponding to at least one historical APS, and different preset identifiers are sorted in chronological order. In the process of the management device storing the obtained at least one new APS into the historical candidate set, the management device may delete all historical APSs corresponding to the first identifier in the historical candidate set, where the first identifier is an identifier among the multiple preset identifiers, and the time corresponding to the first identifier is earlier than the time corresponding to any identifier among the multiple preset identifiers except the first identifier. Thereafter, the management device may store the at least one new APS into the historical candidate set, and update the correspondence between the preset identifier and the APS in the historical candidate set based on the time when the at least one new APS is stored into the historical candidate set, so that the time corresponding to the second identifier corresponding to the at least one new APS among the multiple preset identifiers is later than the time corresponding to any identifier among the multiple preset identifiers except the second identifier.
[0147] For example, Figure 11 is a schematic diagram of an example of an ALF APS historical candidate set provided by one embodiment. As shown in Figure 11, the ALF APS historical candidate set is also maintained during ALF operation and plays a primary role in acquiring historical APSs. When a new APS is determined to be in use, the ALF APS historical candidate set is also updated according to the first-in, first-out principle. Acquisition of historical APSs in the ALF relies on the ALF APS historical candidate set. It should be noted that historical APSs can also be referred to as historical filtering parameters.
[0148] The codec will construct the ALF APS historical candidate set in the same way. A maximum of N APSs can be maintained within the code stream, with IDs (also known as indexes) ranging from 0 to N-1. Each APS includes the historical APSs of four filtering modes: as well as In Figure 10, N is 8 as an example. Whenever a new APS is used, the old APS is overwritten by the new APS with the same ID number. The ID number of the new APS corresponds to the target video block, that is, the ID number of the new APS can be determined based on the target video block.
[0149] That is, the management device adds multiple new APSs used by the current video block in the video to the historical candidate set, thereby providing valuable references for filtering subsequent video blocks in the video and improving the accuracy of the filtering results of the entire video.
[0150] It can be understood that when making filtering decisions for different filtering modes, the decision factors in the filtering decision process are corrected according to the video category and video features corresponding to the video block, so that the corrected decision factors are more inclined to perform filtering in the filtering decision process, and tend to use new APS when filtering, ensuring that the video block has a large number of new APSs under different filtering modes, and after all the new APSs of the video block are added to the historical candidate set, avoiding reducing the number of APSs corresponding to a certain filtering mode in the historical candidate set, ensuring the fullness of the historical candidate set.
[0151] In some embodiments, when the management device obtains the target video category corresponding to the target video block, the management device may obtain the target video content of the target video block and determine at least one content feature of the target video content, where the content feature is an optical flow feature, a frequency domain feature, a spatial feature, or a depth feature. The management device may then classify the target video block based on the at least one content feature to determine the target video category.
[0152] It's important to note that the primary goal of video classification is to understand the content of a video and identify one or more key themes. Video classification algorithms can automatically analyze the semantic information contained in videos or video clips, automatically labeling, categorizing, and describing the videos, achieving accuracy comparable to human classification.
[0153] Video classification typically considers features such as optical flow, frequency domain, spatial, and depth features of video content. Optical flow features can reflect the motion of objects in a video. These features can be extracted by calculating the motion vectors of pixels between adjacent frames in the video to obtain the optical flow field. Frequency domain features can reflect periodic changes in a video, such as periodic motion. Frequency domain features can be obtained by performing a Fourier transform on the video. Spatial features can reflect the appearance and form of objects in a video. Spatial features, such as color, texture, and shape, can be extracted by processing each frame. Depth features can reflect the semantic information of objects in a video, such as their category and attributes. Deep learning techniques can be used to extract depth features from videos. By combining these features, a more comprehensive description of video content can be achieved, thereby improving the accuracy and robustness of video classification.
[0154] It can be understood that ALF can reduce redundant information in the video and remove noise and interference in the image or video by comparing and smoothing the current pixel with the surrounding pixels, so as to achieve the purpose of enhancing the video image quality. Based on video classification, loop filtering technology can be associated with the main features and motion information of the video, so as to implement more targeted or personalized filtering schemes for different categories of videos and improve video encoding and decoding performance. Preferably, the Lagrangian factor in the ALF can be adaptively corrected based on video classification, that is, the optimal Lagrangian factor correction parameter corresponding to each video category can be tested and determined in advance, and then the corresponding optimal Lagrangian factor correction parameter is selected according to the category to which the current video or video block (such as sub-picture, tile, slice, coding unit CTU or CU) belongs.
[0155] In some embodiments, the target video feature may include at least one of the following features: image resolution, texture complexity, motion complexity, and color histogram.
[0156] It should be noted that Lagrangian factor adjustment based on video features refers to extracting useful information from the video by analyzing video features, and constructing a fitting relationship between it and the optimal value of the Lagrangian factor (corresponding to the best video quality). Therefore, in actual video encoding, it is only necessary to extract the video features of the current video or video block (for example, sub-picture, tile, slice, coding unit CTU or CU) to obtain the optimal Lagrangian factor correction parameter corresponding to the coding object. In obtaining video features, the appropriate feature extraction method can be selected according to the specific application scenario and requirements. Common video features include image resolution, texture complexity, motion complexity (such as optical flow), color histogram, etc.
[0157] It is understandable that by determining different types of video features, valuable references can be provided for the correction of decision factors, thereby improving the accuracy of the corrected decision factors.
[0158] The following describes the management method of the adaptive parameter set provided by the embodiment of the present disclosure with reference to some examples.
[0159] For example, taking the modification of the decision factor based on the video category as an example, with reference to FIG12A and FIG12B , the method for managing the adaptive parameter set includes:
[0160] Operation 1: The management device calculates new APSs for luma ALF and chroma ALF.
[0161] The management device calculates the new APS of the brightness ALF for the current slice according to the principle of Wiener filtering. and the new APS for Chroma ALF
[0162] Operation 2: The management device adjusts λ within the brightness ALF.
[0163] Operation 2-a: The management device determines whether the λ correction condition of the luminance ALF is satisfied.
[0164] It should be noted that, in this embodiment, preferably, a Lagrangian factor correction condition is introduced, that is, when the correction condition is met, the Lagrangian factor is adjusted based on the video category, and the present disclosure does not limit the adoption of any correction condition.
[0165] If the λ correction condition of the brightness ALF is satisfied, the management device performs operation 2-b.
[0166] If the λ correction condition of the brightness ALF is not satisfied, the management device performs operation 2-d.
[0167] Operation 2-b: The management device determines a modification parameter k of the Lagrangian factor based on the category to which the video or video block belongs (ie, the video category).
[0168] Operation 2-c: The management device uses k to adjust λ of the brightness ALF.
[0169] As shown in Formula 3:
[0170] Operation 2-d: The management device makes a brightness ALF decision.
[0171] Make RDO decisions among the following candidate modes:
[0172] (1) Do not do ALF;
[0173] (2) New APS;
[0174] (3) Fixed filter set (i.e., preset APS);
[0175] (4) Historical APS;
[0176] If the decision result is "use new APS of luma ALF", alf_luma_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0177] Operation 3: The management device determines whether the brightness ALF decision is enabled.
[0178] If the brightness ALF decision is not enabled, the management device does not perform ALF filtering.
[0179] If the brightness ALF decision is on, the management device performs operation 4.
[0180] Operation 4: The management device adjusts λ within the chroma ALF.
[0181] Operation 4-a: The management device determines whether the λ correction condition of the chroma ALF is satisfied.
[0182] If the λ correction condition of the chrominance ALF is satisfied, the management device performs operation 4-b.
[0183] If the λ correction condition of the chroma ALF is not satisfied, the management device performs operation 4-d.
[0184] Operation 4-b: The management device determines a modification parameter k of the Lagrangian factor based on the category to which the video or video block belongs (ie, the video category).
[0185] Operation 4-c: The management device uses k to adjust λ of the chroma ALF.
[0186] As shown in formula 4:
[0187] Operation 4-d: The management device makes a chroma ALF decision.
[0188] Make RDO decisions among the following candidate modes:
[0189] (1) Do not do ALF;
[0190] (2) New APS;
[0191] (3) Fixed filter set;
[0192] (4) Historical APS;
[0193] If the decision result is "use new APS for chroma ALF", alf_luma_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0194] Operation 5: The management device performs luminance ALF and chrominance ALF reconstruction.
[0195] The management device reconstructs the image after the luminance and chrominance ALF effects based on the decision results of the luminance and chrominance ALF.
[0196] Operation 6: The management device calculates new APSs for CCALF Cb and CCALF Cr.
[0197] The management device calculates the new APS of CCALF Cb for the current slice according to the principle of Wiener filtering. and the new APS of CCALF Cr
[0198] Operation 7: The management device adjusts λ in CCALF Cb.
[0199] Operation 7-a: The management device determines whether the λ correction condition of CCALF Cb is satisfied.
[0200] If the λ correction condition of CCALF Cb is satisfied, the management device performs operation 7-b.
[0201] If the λ correction condition of CCALF Cb is not satisfied, the management device performs operation 7-d.
[0202] Operation 7-b: The management device determines a modification parameter k of the Lagrangian factor based on the category to which the video or video block belongs (ie, the video category).
[0203] Operation 7-c: The management device uses k to adjust λ of CCALF Cb.
[0204] As shown in Formula 5:
[0205] Operation 7-d: The management device makes a CCALF Cb decision.
[0206] Make RDO decisions among the following candidate modes:
[0207] (1) Do not do ALF;
[0208] (2) New APS;
[0209] (3) Fixed filter set;
[0210] (4) Historical APS;
[0211] If the decision result is "use new APS of CCALF Cb", alf_cc_cb_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0212] Operation 8: The management device adjusts λ in CCALF Cr.
[0213] Operation 8-a: The management device determines whether the λ correction condition of CCALF Cr is satisfied.
[0214] If the λ correction condition of CCALF Cr is satisfied, the management device performs operation 8-b.
[0215] If the λ correction condition of CCALF Cr is not satisfied, the management device performs operation 8-d.
[0216] Operation 8-b: The management device determines a modification parameter k of the Lagrangian factor based on the category to which the video or video block belongs (ie, the video category).
[0217] Operation 8-c: The management device uses k to adjust λ of CCALF Cr.
[0218] As shown in Formula 6:
[0219] Operation 8-d: The management device makes a CCALF Cr decision.
[0220] Make RDO decisions among the following candidate modes:
[0221] (1) Do not do ALF;
[0222] (2) New APS;
[0223] (3) Fixed filter set;
[0224] (4) Historical APS;
[0225] If the decision result is "use new APS of CCALF Cr", alf_cc_cr_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0226] Operation 9: The management device performs CCALF Cb and CCALF Cr reconstruction.
[0227] The management device reconstructs the image after the effects of CCALF Cb and CCALF Cr based on the decision results of CCALF Cb and CCALF Cr.
[0228] Operation 10: The management device determines whether at least one of the identifiers of the plurality of filtering modes is 1.
[0229] That is, alf_luma_filter_signal_flag∨alf_chroma_filter_signal_flag∨alf_cc_cb_filter_signal_flag∨alf_cc_cr_filter_signal_flag==1.
[0230] If at least one of the four identifiers is 1, the management device performs operation 11.
[0231] If the above four identifiers are all 0, the management device does not need to perform operation 11.
[0232] Operation 11: The management device transmits the current APS.
[0233] The management device transmits the APS of the current transmission target filter unit (ie, the target video block) into the bitstream.
[0234] If an APS with the same index already exists, the new APS will be used to overwrite the old one.
[0235] Exemplarily, taking the modification of decision factors based on video features as an example, with reference to FIG. 13A and FIG. 13B , the method for managing an adaptive parameter set includes:
[0236] Operation 1: The management device calculates new APSs for luma ALF and chroma ALF.
[0237] The management device calculates the new APS of the brightness ALF for the current slice according to the principle of Wiener filtering. and the new APS for Chroma ALF
[0238] Operation 2: The management device adjusts λ within the brightness ALF.
[0239] Operation 2-a: The management device determines whether the λ correction condition of the luminance ALF is satisfied.
[0240] It should be noted that, in this embodiment, preferably, a Lagrangian factor correction condition is introduced, that is, when the correction condition is met, the Lagrangian factor is adjusted based on the video features, and the present disclosure does not limit the adoption of any correction condition.
[0241] If the λ correction condition of the brightness ALF is satisfied, the management device performs operation 2-b.
[0242] If the λ correction condition of the brightness ALF is not satisfied, the management device performs operation 2-d.
[0243] Operation 2-b: The management device determines a modification parameter k of the Lagrangian factor based on video or video block characteristics (ie, video characteristics).
[0244] Operation 2-c: The management device uses k to adjust λ of the brightness ALF.
[0245] As shown in Formula 3:
[0246] Operation 2-d: The management device makes a brightness ALF decision.
[0247] Make RDO decisions among the following candidate modes:
[0248] (1) Do not do ALF;
[0249] (2) New APS;
[0250] (3) Fixed filter set (i.e., preset APS);
[0251] (4) Historical APS;
[0252] If the decision result is "use new APS of luma ALF", alf_luma_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0253] Operation 3: The management device determines whether the brightness ALF decision is enabled.
[0254] If the brightness ALF decision is not enabled, the management device does not perform ALF filtering.
[0255] If the brightness ALF decision is on, the management device performs operation 4.
[0256] Operation 4: The management device adjusts λ within the chroma ALF.
[0257] Operation 4-a: The management device determines whether the λ correction condition of the chroma ALF is satisfied.
[0258] If the λ correction condition of the chrominance ALF is satisfied, the management device performs operation 4-b.
[0259] If the λ correction condition of the chroma ALF is not satisfied, the management device performs operation 4-d.
[0260] Operation 4-b: The management device determines a modification parameter k of the Lagrangian factor based on video or video block characteristics (ie, video characteristics).
[0261] Operation 4-c: The management device uses k to adjust λ of the chroma ALF.
[0262] As shown in formula 4:
[0263] Operation 4-d: The management device makes a chroma ALF decision.
[0264] Make RDO decisions among the following candidate modes:
[0265] (1) Do not do ALF;
[0266] (2) New APS;
[0267] (3) Fixed filter set;
[0268] (4) Historical APS;
[0269] If the decision result is "use new APS for chroma ALF", alf_luma_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0270] Operation 5: The management device performs luminance ALF and chrominance ALF reconstruction.
[0271] The management device reconstructs the image after the luminance and chrominance ALF effects based on the decision results of the luminance and chrominance ALF.
[0272] Operation 6: The management device calculates new APSs for CCALF Cb and CCALF Cr.
[0273] The management device calculates the new APS of CCALF Cb for the current slice according to the principle of Wiener filtering. and the new APS of CCALF Cr
[0274] Operation 7: The management device adjusts λ in CCALF Cb.
[0275] Operation 7-a: The management device determines whether the λ correction condition of CCALF Cb is satisfied.
[0276] If the λ correction condition of CCALF Cb is satisfied, the management device performs operation 7-b.
[0277] If the λ correction condition of CCALF Cb is not satisfied, the management device performs operation 7-d.
[0278] Operation 7-b: The management device determines a modification parameter k of the Lagrangian factor based on video or video block characteristics (ie, video characteristics).
[0279] Operation 7-c: The management device uses k to adjust λ of CCALF Cb.
[0280] As shown in Formula 5:
[0281] Operation 7-d: The management device makes a CCALF Cb decision.
[0282] Make RDO decisions among the following candidate modes:
[0283] (1) Do not do ALF;
[0284] (2) New APS;
[0285] (3) Fixed filter set;
[0286] (4) Historical APS;
[0287] If the decision result is "use new APS of CCALF Cb", alf_cc_cb_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0288] Operation 8: The management device adjusts λ in CCALF Cr.
[0289] Operation 8-a: The management device determines whether the λ correction condition of CCALF Cr is satisfied.
[0290] If the λ correction condition of CCALF Cr is satisfied, the management device performs operation 8-b.
[0291] If the λ correction condition of CCALF Cr is not satisfied, the management device performs operation 8-d.
[0292] Operation 8-b: The management device determines a modification parameter k of the Lagrangian factor based on video or video block characteristics (ie, video characteristics).
[0293] Operation 8-c: The management device uses k to adjust λ of CCALF Cr.
[0294] As shown in Formula 6:
[0295] Operation 8-d: The management device makes a CCALF Cr decision.
[0296] Make RDO decisions among the following candidate modes:
[0297] (1) Do not do ALF;
[0298] (2) New APS;
[0299] (3) Fixed filter set;
[0300] (4) Historical APS;
[0301] If the decision result is "use new APS of CCALF Cr", alf_cc_cr_filter_signal_flag needs to be set to 1, otherwise it is set to 0.
[0302] Operation 9: The management device performs CCALF Cb and CCALF Cr reconstruction.
[0303] The management device reconstructs the image after the effects of CCALF Cb and CCALF Cr based on the decision results of CCALF Cb and CCALF Cr.
[0304] Operation 10: The management device determines whether at least one of the identifiers of the plurality of filtering modes is 1.
[0305] That is, alf_luma_filter_signal_flag∨alf_chroma_filter_signal_flag∨alf_cc_cb_filter_signal_flag∨alf_cc_cr_filter_signal_flag==1.
[0306] If at least one of the four identifiers is 1, the management device performs operation 11.
[0307] If the above four identifiers are all 0, the management device does not need to perform operation 11.
[0308] Operation 11: The management device transmits the current APS.
[0309] The management device transmits the APS of the current transmission target filter unit (ie, the target video block) into the bitstream.
[0310] If an APS with the same index already exists, the new APS will be used to overwrite the old one.
[0311] It is understandable that, in order to realize the above functions, the management device of the adaptive parameter set includes hardware structures and / or software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the algorithm steps of each example described in the embodiments of the present disclosure, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present disclosure.
[0312] The embodiment of the present disclosure can divide the management device of the adaptive parameter set into functional modules according to the above-mentioned method embodiment. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one functional module. The above-mentioned integrated module can be implemented in the form of hardware or software. It should be noted that the division of modules in the embodiment of the present disclosure is schematic and is only a logical function division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.
[0313] Figure 14 is a schematic diagram of the structure of an adaptive parameter set management device provided in an embodiment of the present disclosure. Adaptive parameter set management device 1400 can execute the adaptive parameter set management method shown in Figure 8 of the above method embodiment. As shown in Figure 14, adaptive parameter set management device 1400 includes: an acquisition module 1401, a processing module 1402, and a storage module 1403.
[0314] Acquisition module 1401 is configured to acquire a target video category and / or target video features corresponding to a target video block. Processing module 1402 is configured to determine a correction factor for each decision factor based on the target video category and / or target video features. Each decision factor corresponds to a preset filtering mode and indicates whether filtering is to be performed on the video block and, when filtering, to use a new APS, a historical APS from a historical candidate set, or a preset APS. Processing module 1402 is further configured to correct each decision factor based on its correction factor to obtain multiple corrected decision factors, wherein each corrected decision factor indicates a greater probability of performing filtering than the probability indicated by the previous decision factor, and each corrected decision factor indicates a greater probability of using the new APS during filtering than the probability indicated by the previous decision factor. Storage module 1403 is configured to store at least one new APS obtained based on each corrected decision factor in the historical candidate set.
[0315] Optionally, the historical candidate set includes multiple historical APSs and multiple preset identifiers, one preset identifier corresponds to at least one historical APS, and different preset identifiers are sorted in chronological order. Processing module 1402 is used to delete all historical APSs corresponding to the first identifier in the historical candidate set, where the first identifier is an identifier among multiple preset identifiers, and the time corresponding to the first identifier is earlier than the time corresponding to any identifier among the multiple preset identifiers except the first identifier. Storage module 1403 is also used to store at least one new APS in the historical candidate set. Processing module 1402 is also used to update the correspondence between the preset identifier and the APS in the historical candidate set according to the time when at least one new APS is stored in the historical candidate set, so that the time corresponding to the second identifier corresponding to at least one new APS among the multiple preset identifiers is later than the time corresponding to any identifier among the multiple preset identifiers except the second identifier.
[0316] Optionally, acquisition module 1401 is configured to acquire target video content of a target video block. Processing module 1402 is further configured to determine at least one content feature of the target video content, where the content feature is an optical flow feature, a frequency domain feature, a spatial feature, or a depth feature. Processing module 1402 is further configured to determine a target video category based on the at least one content feature.
[0317] Optionally, the target video feature includes at least one of the following features: image resolution, texture complexity, motion complexity, and color histogram.
[0318] Optionally, the target video block is a video block in the video to be processed, the target video category is the video category of the video to be processed, and the target video feature is the video feature of the video to be processed, or the target video category is the video category of the target video block, and the target video feature is the video feature of the target video block.
[0319] Optionally, the decision factor is a Lagrangian factor in a rate-distortion optimization (RDO) decision process.
[0320] Optionally, the preset filtering mode is any one of the following adaptive loop filters ALF: chroma ALF, luma ALF, inter-component adaptive loop filter CCALF blue component and luma difference Cb, CCALF red component and luma difference Cr.
[0321] Optionally, the target video block is any of the following types: a sub-picture, a tile, a slice, a tree coding unit CTU, or a coding unit CU.
[0322] In the case of implementing the functions of the above-mentioned integrated modules in hardware, the disclosed embodiments provide another electronic device structure for the adaptive parameter set management device involved in the above-mentioned embodiments. As shown in Figure 15, the adaptive parameter set management device 1500 includes: a memory 1501, a processor 1502, a communication interface 1503, and a bus 1504.
[0323] The memory 1501 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store dynamic information and instructions, an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0324] The processor 1502 may be a logic block, module, and circuit that implements or executes the various exemplary methods described in conjunction with the embodiments of the present disclosure. The processor 1502 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. The processor 1502 may also implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. The processor 1502 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP (digital signal processor) and a microprocessor, and the like.
[0325] The communication interface 1503 is used to connect to other devices via a communication network, such as Ethernet, wireless access network, or wireless local area network (WLAN).
[0326] In one implementation, the memory 1501 may exist independently of the processor 1502 and may be connected to the processor 1502 via a bus 1504 for storing instructions or program codes. When the processor 1502 calls and executes the instructions or program codes stored in the memory 1501, the adaptive parameter set management method provided in the embodiment of the present disclosure can be implemented.
[0327] In one implementation, the memory 1501 may also be integrated with the processor 1502 .
[0328] Bus 1504 can be an Extended Industry Standard Architecture (EISA) bus, for example. Bus 1504 can be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG15 shows bus 1504 using only a single bold line. This does not imply that there is only one bus or only one type of bus.
[0329] Some embodiments of the present disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium), which stores computer program instructions. When the computer program instructions are executed on a computer, the computer executes the adaptive parameter set management method as described in any of the above embodiments.
[0330] Exemplarily, the above-mentioned computer-readable storage media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes, etc.), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0331] An embodiment of the present disclosure provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer is enabled to execute the method for managing an adaptive parameter set as described in any one of the above embodiments.
[0332] The above is only a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or replacements within the technical scope disclosed in the present disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A method for managing an adaptive parameter set, comprising: Obtaining a target video category and / or target video features corresponding to a target video block; Determining a correction factor for each decision factor according to the target video category and / or the target video features, where one decision factor corresponds to one preset filtering mode, and the decision factor is used to indicate whether to filter a video block and, when filtering, indicate to use any one of a new APS, a historical APS in a historical candidate set, and a preset APS; Correcting each of the decision factors according to the correction factor of each decision factor to obtain a plurality of corrected decision factors, where the probability of a corrected decision factor indicating filtering is greater than the probability of the decision factor before correction indicating filtering, and the probability of a corrected decision factor indicating using a new APS when filtering is greater than the probability of the decision factor before correction indicating using a new APS when filtering; Storing at least one obtained new APS in the historical candidate set according to each corrected decision factor.
2. The method according to claim 1, wherein, The historical candidate set includes a plurality of historical APSs and a plurality of preset identifiers, where one preset identifier corresponds to at least one historical APS, and different preset identifiers are sorted in chronological order. The storing the at least one obtained new APS in the historical candidate set includes: Deleting all historical APSs corresponding to a first identifier in the historical candidate set, where the first identifier is an identifier among the plurality of preset identifiers, and the time corresponding to the first identifier is earlier than the time corresponding to any other identifier among the plurality of preset identifiers except the first identifier; Storing the at least one new APS in the historical candidate set and updating the correspondence between the preset identifier and the APS in the historical candidate set according to the time when the at least one new APS is stored in the historical candidate set, so that the time corresponding to a second identifier corresponding to the at least one new APS among the plurality of preset identifiers is later than the time corresponding to any other identifier among the plurality of preset identifiers except the second identifier.
3. The method according to claim 1 or 2, wherein Obtaining the target video category corresponding to the target video block includes: Obtaining the target video content of the target video block; Determining at least one content feature of the target video content, where the content feature is an optical flow feature, a frequency domain feature, a spatial feature, or a depth feature; Determining the target video category according to the at least one content feature.
4. The method according to claim 1 or 2, wherein The target video features include at least one of the following features: image resolution, texture complexity, motion complexity, color histogram.
5. The method according to claim 1 or 2, wherein The target video block is a video block in a video to be processed; The target video category is the video category of the video to be processed, and the target video features are the video features of the video to be processed, or The target video category is the video category of the target video block, and the target video features are the video features of the target video block.
6. The method according to claim 1 or 2, wherein The decision factor is the Lagrangian factor in the rate-distortion optimization (RDO) decision process.
7. The method according to claim 1 or 2, wherein The preset filtering mode is any one of the following adaptive loop filters (ALF): chrominance ALF, luma ALF, cross-component adaptive loop filter (CCALF) for the difference between the blue component and luma (Cb), and CCALF for the difference between the red component and luma (Cr).
8. The method according to claim 1 or 2, wherein The target video block is any one of the following types: sub-picture, tile, slice, coding tree unit (CTU), coding unit (CU).
9. An electronic device, comprising: A memory and a processor; The memory is coupled to the processor; The memory is used to store instructions executable by the processor; When the processor executes the instructions, it executes the method for managing an adaptive parameter set according to any one of claims 1-8.
10. A computer-readable storage medium, wherein, Computer instructions are stored on the computer-readable storage medium, and when the computer instructions run on a computer, the computer is caused to execute the method for managing an adaptive parameter set according to any one of claims 1-8.
Citation Information
Patent Citations
Adaptive parameter set management method and device, equipment and storage medium
CN120302033A
Video processing method and device, storage medium and electronic device
CN116433783A
Encoding method and apparatus, decoding method and apparatus, and devices therefor
US20230104806A1
Spatial resolution adaptation of in-loop and post-filtering of compressed video using metadata
WO2022073811A1