Perception-geometry joint code rate control method, system, and medium
By using a perception-geometric joint rate control method, the CTU-level rate allocation is dynamically adjusted, solving the geometric redundancy and visual perception problems in high-resolution video transmission, and achieving more efficient video encoding and transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies fail to adequately eliminate geometric redundancy and take into account the differences in perceptual characteristics of the human visual system in high-resolution video transmission, resulting in loss of coding quality and mismatched bandwidth allocation.
A combined perceptual-geometric rate control method is adopted. By calculating the perceptual-geometric joint weight of the CTU, the rate allocation is dynamically adjusted. Combined with the Gaussian decay model and the visual saliency gradient modulation mechanism, the CTU-level rate allocation is optimized.
It effectively eliminates structural redundancy in high-resolution videos, enhances the consistency of human visual perception, promotes efficient video transmission, and improves the fit between bitstream distribution and the human visual system.
Smart Images

Figure CN121397222B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of video coding, and more particularly, to a perceptual-geometric joint rate control method, system and medium. BACKGROUND
[0002] With the rapid development of dynamic spectrum combustion diagnosis and hazard gas leakage monitoring, the research on the demand of high-efficiency data transmission in these fields has become an important direction to promote the progress of related technologies. When monitoring the combustion process and gas leakage in real time, a large amount of spectral data often needs to be transmitted, which is usually high-resolution and contains a large amount of redundant features, which leads to an increase in communication costs and affects real-time data transmission. However, these high-resolution data are prone to significant geometric redundancy during transmission, especially in the data feature space, when the feature distribution of different regions has significant differences, the transmission of useful information will be limited.
[0003] Rate control (RC) is a crucial technical link in the video coding process, which controls the size of the encoded video file or the data transmission rate within a desired target range by intelligently allocating data bits while ensuring video quality. Currently, based on the international video coding standard VVC (Versatile Video Coding), the transmission of high-resolution spectral data still uses traditional rate control algorithms for traditional video, which leads to the problem of data sampling redundancy and bandwidth allocation imbalance during monitoring. At the same time, these methods fail to fully consider the perceptual characteristics of HVS (Human Visual System) for high-resolution video, i.e., the human visual system has higher tolerance for distortion of CTUs outside the core region of the field of view, and these perceptual redundancies lead to loss of coding quality in key visual regions during the rate control process.
[0004] In recent years, the rate control research for high-resolution videos presents a multi-dimensional optimization trend, mainly focusing on distortion compensation, visual perception modeling, and dynamic bit allocation strategies. For example, by optimizing the video bit allocation of important areas to improve the encoding quality of the video; or designing a weighted bit allocation algorithm based on CTU layer weights and developing a CTU row-level rate control algorithm to optimize the rate control performance of video transmission; there are also inter-frame / intra-frame bit competition methods modeled by game theory to reduce the bit waste of high-latitude areas; in addition, optimization is carried out for the VVC standard, such as implementing CTU-level fine allocation through salient region segmentation, or introducing the concept of virtual competitors to reduce Group of Pictures (GOP) level bit fluctuations and improve video transmission stability. Although the above research methods have made significant progress, there are still two key problems in dealing with the transmission of high-resolution video data: first, the existing research still does not fully exploit the spatial redundancy of high-latitude areas, and needs to further reduce the spatial geometric redundancy; second, the traditional rate allocation algorithm fails to accurately reflect the differences in perceptual weights caused by human visual preferences. These factors lead to the mismatch of traditional CTU-level bit allocation models, and a joint optimization framework that closely couples the VVC encoding characteristics and perceptual geometric redundancy constraints needs to be explored. SUMMARY
[0005] In view of the defects of the prior art and the need for improvement, the present application provides a perceptual-geometric joint rate control method, system and medium, which aims to effectively eliminate geometric redundancy and improve human visual perception consistency, and promote efficient transmission of high-resolution videos.
[0006] To achieve the above-mentioned purpose, according to one aspect of the present application, a perceptual-geometric joint rate control method is provided, comprising:
[0007] S1: Assign weights to each CTU in each frame in the video according to the VVC standard;
[0008] S2: For each I frame, calculate the perceptual-geometric joint weight of each CTU in the I frame, and update the weight of each CTU in the I frame according to
[0009] The calculation method of the perceptual-geometric joint weight of the CTU includes:
[0010] Taking as the visual center reference line of the maximum weight area, the visual angle of the CTU relative to the visual center reference line is calculated;
[0011] If is located in the visual core area, its perceptual-geometric joint weight is calculated according to ; otherwise, the perceptual-geometric joint weight of the CTU is calculated according to computing its perceptual-geometric joint weight ;
[0012] wherein, denotes the CTU weight calculated according to the VVC standard; is the projection plane height; is the vertical position of the top-left pixel of each CTU, is the CTU height; denotes the total number of CTUs in the current frame, denotes the perceptual-geometric joint weight of the CTU at the th position in the current frame; updates the weight of the CTU at the th position in the I frame after updating the weight; and is a parameter for controlling the weight decay rate, and ;
[0013] S3: allocates code rate to each CTU according to the weight of each CTU in each frame, and completes the code rate control.
[0014] Further, in S3, the code rate allocated to the CTU at the th position in the I frame is :
[0015]
[0016] wherein, is the number of bits not allocated when encoding the CTU at the th position in the current frame; denotes the weight of the CTU at the th position in the I frame after updating the weight, denotes the weight of the CTU at the th position in the I frame after updating the weight.
[0017] Further, S2 further comprises: for each B frame, calculating the perceptual-geometric joint weight of each CTU, and updating the weight of each CTU in the B frame according to the following expression:
[0018]
[0019] wherein, , denotes a code rate control experience parameter, denotes the code rate, denotes the number of pixels in a CTU; denotes the weight of the CTU at the th position in the B frame after updating the weight.
[0020] Furthermore, in S3, the first frame in B... The code rate allocated to each CTU for:
[0021]
[0022] in, This indicates the weighted B-frame after the weight update. The weight of each CTU, This indicates the weighted B-frame after the weight update. The weight of each CTU; This indicates a preset sliding window.
[0023] Furthermore, the angle of the CTU relative to the baseline of this field of view center. The calculation formula is as follows:
[0024] .
[0025] Furthermore, , ;
[0026] in, Indicates the height of the frame.
[0027] Furthermore, the field of view corresponding to the core area of the field of view is .
[0028] According to another aspect of the present invention, a perception-geometry joint rate control system is provided, comprising:
[0029] The weight initialization module is used to assign weights to each CTU in each frame of the video according to the VVC standard.
[0030] The perception-geometric joint weight construction module is used to calculate the perception-geometric joint weight of CTU;
[0031] The I-frame weight optimization module is used to calculate the joint perception-geometric weights of each CTU in each I-frame using the perception-geometric joint weight construction module, and according to... Update the weights of each CTU in the I-frame;
[0032] The bitrate allocation module is used to allocate bitrate to each CTU according to the weight of each CTU in each frame, thereby completing bitrate control;
[0033] The calculation method for the CTU's sensing-geometric joint weights includes:
[0034] by Calculate the angle of view of the CTU relative to the center baseline of the field of view of the region with the highest weight. ;
[0035] If is located in the core region of the view angle, then the perceptual-geometric joint weight of the CTU is calculated according to ; otherwise, the perceptual-geometric joint weight of the CTU is calculated according to . ;
[0036] represents the CTU weight calculated according to the VVC standard; is the height of the projection plane; is the vertical position of the top-left pixel of each CTU, is the height of the CTU; represents the total number of CTUs in the current frame, represents the perceptual-geometric joint weight of the th CTU in the current frame; updates the weight of the CTU with the view angle of in the updated I frame; and is a parameter for controlling the weight decay rate, and .
[0037] Further, the perceptual-geometric joint rate control system provided by the present application further comprises a B frame weight optimization module configured to calculate the perceptual-geometric joint weight of each CTU in each B frame by using the perceptual-geometric joint weight construction module, and update the weight of each CTU in the B frame according to the following expression:
[0038]
[0039] wherein, , represents a rate control experience parameter, represents the code rate, represents the number of pixels in a CTU; represents the weight of the CTU with the view angle of in the updated B frame.
[0040] According to still another aspect of the present application, there is provided a computer readable storage medium comprising a stored computer program, wherein the computer program, when executed by a processor, implements the perceptual-geometric joint rate control method provided by the present application.
[0041] Overall, the above technical solutions conceived by the present application can achieve the following beneficial effects:
[0042] (1) The application constructs a CTU-level spatial geometry visual field change weight model based on a Gaussian attenuation model, so that the spatio-temporal coupling effect of the distortion tolerance change gradient of the inside and outside of the visual field and the coding unit between the plane height can be accurately modeled, at the same time, by using the characteristic that there is a nonlinear correlation between the plane longitudinal visual angle and the visual attention, that is, the sensitivity of the human eye to the distortion of the image in the core area of the visual angle is strong, and the distortion tolerance of the CTU outside the core area of the visual angle is higher, a larger parameter is set for the CTU in the core area of the visual angle , so that the CTU weight in the area changes slowly to retain more image details, and a smaller parameter is set for the CTU outside the core area of the visual angle , so that the CTU weight outside the core area of the visual angle mutates to suppress redundancy, and the perception-geometry joint weight of the CTU calculated on the basis of accurate modeling can effectively eliminate the structural redundancy of high-resolution video, so that the code stream distribution is more consistent with the spatial perception characteristics of the human visual system. For I frame, on the basis of calculating the weight of each CTU in each frame according to the VVC standard, the application further calculates the perception-geometry joint weight of each CTU in the I frame, and updates the weight of the CTU in the I frame by using the perception-geometry joint weight, so that the code stream distribution result can effectively eliminate the geometric redundancy and improve the consistency of human visual perception, and promote the efficient transmission of high-resolution video.
[0043] (2) In the preferred scheme of the application, for B frame, on the basis of calculating the weight of each CTU in each frame according to the VVC standard, the application further calculates the perception-geometry joint weight of each CTU in the B frame, and updates the weight of the CTU in the B frame by using the perception-geometry joint weight, so that the code rate control result can be further optimized.
[0044] Overall, the application dynamically models the spatio-temporal coupling effect of the distortion tolerance change gradient of the inside and outside of the visual field and the coding unit between the plane height, and fuses the visual saliency gradient modulation mechanism, to construct a differentiable, perception-adaptive weight model; then the new weight model is used to improve the code rate allocation framework of VVC, to realize the perception-geometry collaborative optimization allocation of CTU-level code rate, effectively eliminate the structural redundancy of high-resolution video, and make the code stream distribution more consistent with the spatial perception characteristics of the human visual system. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The perception-geometry joint weight change heat map of the perception-geometry joint code rate control method provided by the embodiment of the application.
[0046] Figure 2 The I frame / B frame code rate control method flowchart provided by the embodiment of the application.
[0047] Figure 3 A flow chart of the sensing-geometry joint code rate allocation method provided for the embodiment of the present application.
[0048] Figure 4 A schematic diagram of the sensing-geometry joint code rate control system provided for the embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0050] In the present application, the terms "first", "second", etc. (if any) in the present application and the accompanying drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0051] CTU (Coding Tree Unit, coding tree unit) is the basic coding unit in modern video coding standards (such as HEVC / H.265 and VVC / H.266), in the video coding process, the encoder will present a square image block by dividing a frame of image into a square image block, these image blocks are CTU, common sizes are 64x64, 32x32 pixels, etc. The main task of code rate control is to determine how many bits to allocate to each frame according to the complexity of each frame under the premise of ensuring video quality; after determining how many total bits to allocate to a frame, based on the complexity of each CTU in the frame, the bits are reasonably allocated to each CTU in the frame, in order to achieve this, the code rate control algorithm calculates a weight value for each CTU, when allocating bits, it is no longer simply evenly distributed, but is distributed according to the weight proportion.
[0052] The traditional VVC standard takes the statistical characteristic value of the CTU of the coded I frame after Hadamard transform, or the reference characteristic value of the CTU of the coded B frame as the weight of the CTU, and performs code rate allocation based on the weight, thereby realizing code rate control, which will have significant geometric redundancy phenomenon, and does not fully consider the differences in the perception characteristics of HVS for high-resolution videos, which is not conducive to the transmission of high-resolution videos.
[0053] In order to effectively eliminate geometric redundancy and improve human visual perception consistency, and promote high-resolution video transmission efficiency, the present application provides a perceptual-geometric joint code rate control method, system and medium, the overall concept of which is to deeply analyze the spatio-temporal coupling effect of the change gradient of distortion tolerance inside and outside the field of view and the coding unit in the plane height, and accurately model it, on the basis of which, the nonlinear correlation between the longitudinal angle of the plane and the visual attention is utilized, the visual saliency gradient modulation mechanism is fused, the differentiable and perceptual adaptive weight model is constructed, and then the new weight model is utilized to improve the code rate allocation framework of VVC, so that the perceptual-geometric collaborative optimization allocation of CTU level code rate is realized.
[0054] Based on the above concept, the present application first constructs a CTU level spatial geometric field of view change weight model based on the Gaussian decay model, which is used to dynamically model the spatio-temporal coupling effect of the change gradient of distortion tolerance inside and outside the field of view and the coding unit in the plane height, so as to accurately quantify the weight importance change degree of the pixels inside and outside the field of view. The weight change model can be represented by the following weight importance mapping function based on the change of the plane height:
[0055]
[0056] Wherein, is the weight of the CTU; represents the longitudinal position of the top-left corner pixel of each CTU, is the projection plane height, is the CTU height; characterizes the field of view center reference line of the maximum weight area, which corresponds to the middle area in the plane; (the standard deviation) is a parameter for controlling the weight decay rate.
[0057] The core of the model is to adaptively fit the perceptual importance change in the longitudinal direction of the angle through the decay characteristics of the Gaussian distribution, to nonlinearly adjust the statistical characteristics or distortion parameters at the CTU level, so as to reduce the mismatch between the traditional code rate allocation weight and the spatial visual perception, and make the code rate allocation consistent with the field of view bias habit.
[0058] On the basis of the established CTU level adaptive plane height change weight model , the present application further utilizes the nonlinear correlation between the longitudinal angle and the visual attention: in the core area of the angle (usually corresponding to ), the sensitivity of the human eye to image distortion is strong, so the weight gradient change needs to be changed slowly to retain more image details, while in the outer area of the angle or , the weight gradient change needs to be changed suddenly to suppress redundancy. Accordingly, the present application changes the weight gradient change in the CTU level adaptive plane height change weight model Different parameters are used for different regions to control the weight decay rate, specifically, the parameter , controls the gradient change rate of the weight of the core region and the edge region respectively , the core region adopts a larger parameter , so that the weight slowly decreases with , which conforms to the high sensitivity characteristics of the HVS to the subtle distortion of the central visual field. The edge region adopts a smaller parameter , simulating the fast and dull response of human beings to the distortion of the periphery of the visual field. The core region CTU obtains a higher weight coefficient, driving the VVC rate-distortion optimizer to allocate more bits to it, preserving the texture details. The edge region weight sharply decreases, actively merging the coding unit nodes using the multi-type tree partitioning mechanism of VVC, and actively eliminating visual redundancy using the Skip / Merge mode of VVC, dynamically reducing the bit allocation of the edge region. The weight function is continuously derivable at , avoiding the rate allocation jump at the CTU boundary. Based on this, the present application fuses the visual saliency gradient modulation mechanism on the basis of the adaptive plane height change weight model of the CTU level, and constructs a differentiable and perceptually adaptive weight model; the differentiable and perceptually adaptive weight model can be represented by the following perceptual-geometric joint weight segmentation function:
[0059]
[0060] After calculating the perceptual-geometric joint weight of each CTU based on the above perceptual-geometric joint weight segmentation function, the normalized value of the perceptual-geometric joint weight in the height interval corresponding to the frame is enlarged to the interval (0, 255), and the perceptual-geometric joint weight change heat map of the CTU is as shown in Figure 1 . In the region within the absolute value 60°, that is, in the visual core region, the CTU weight value changes slowly, and outside the visual core region, the CTU weight value changes faster.
[0061] In the above perceptual-geometric joint weight segmentation function, the visual angle of the CTU determines the segmentation position, and the present application takes to represent the visual field center reference line of the maximum weight region, and accordingly, the position of the plane visual field center reference line is set to 0 degrees, and the longitudinal position angle of each CTU is calculated according to the longitudinal position of the center pixel point of the CTU in the upper and lower parts of the image, respectively changing from the middle of the image to the upper and lower areas. , Accordingly, the visual angle of each CTU can be calculated by the following calculation expression:
[0062]
[0063] Through a large number of experiments, it is shown that when setting , ( representing the height of the frame), the weight model can make the code rate control accuracy of encoding high-resolution video increase by 0.3% on average in AI / LDB / RA mode. The function promotes visual perception enhancement and realizes code rate allocation weight adjustment more in line with the preference of HVS.
[0064] The I frame in the video is a self-contained, independent complete picture, which can be decoded without relying on any other frame; the B frame is the frame type with the highest compression efficiency. It is encoded by simultaneously referring to the front and rear frames. Considering that the complexities of different video frames are different, the present application optimizes the allocation of the number of encoding bits of the I frame in the video based on the established differentiable, perceptually adaptive weight model, so that when the I frame is encoded, the code rate control can be adaptively adjusted, thereby better meeting the requirements of visual quality, and at the same time promoting the fact that the region needing higher quality can obtain more bits, and vice versa, thereby realizing fine management of code rate control. Specifically, the present application optimizes the CTU in the I frame through the following expression:
[0065]
[0066] wherein, the weight of the CTU with a perspective of in the updated I frame is updated; represents the CTU weight calculated according to the VVC standard, that is, the statistical characteristic value of the current encoding CTU after Hadamard transformation; represents the total number of CTUs in the current frame, represents the perceptual-geometric joint weight of the CTU in the current frame. In the above weight updating formula, the perceptual-geometric joint weight of the current CTU emphasizes the importance of the current CTU in the visual field perception of the overall image, and the geometric mean of all CTU weights of the current frame is calculated by , which provides a benchmark value, so that the current CTU weight is compared in a wider context, which helps to overcome the influence of extreme values.
[0067] After updating the weight of the CTU using the perceptual-geometric joint weight of the CTU, the present application proposes an I frame dynamic bit allocation regulation model, which reconstructs the CTU-level code rate allocation function as:
[0068]
[0069] wherein, represents the code rate allocated to the i-th CTU in the I frame; represents the code rate allocated to the i-th CTU in the I frame; is the number of bits not allocated when encoding the i-th CTU of the current frame; is the number of bits not allocated when encoding the i-th CTU of the current frame; represents the weight of the i-th CTU in the I frame after updating the weight, represents the weight of the i-th CTU in the I frame after updating the weight, represents the weight of the i-th CTU in the I frame after updating the weight, represents the weight of the i-th CTU in the I frame after updating the weight, is the sum of the remaining CTU weights.
[0070] In order to further optimize the code rate control and promote efficient transmission of high-resolution video, the present application can also optimize the code rate control of B frames on the basis of optimizing the code rate allocation of I frames.
[0071] When optimizing the code rate control of B frames, the present application aims to improve the perceptual experience from the visual preference, and according to the space-time characteristics of B frames, the present application adjusts the parameter reflecting the VVC encoded B frame rate distortion characteristics according to the perceptual preference, first, the CTU-level perceptual-geometric joint weight is used to optimize is , based on which the optimized pre-allocated code rate is obtained, and then the new weight is obtained, which is used for the B frame dynamic bit allocation control model, the CTU-level code rate allocation function is reconstructed, and efficient dynamic bit rate allocation and control are realized.
[0072] For B frames, the optimized code rate control parameter is:
[0073]
[0074] wherein, , represents the code rate control experience parameter, consistent with the VVC standard; represents the code rate. The optimized pre-allocated single-pixel code rate is obtained as follows:
[0075]
[0076] The code rate allocation weight of the i-th CTU of the current VVC encoded B frame is updated to as follows:
[0077]
[0078] wherein, represents the number of pixels in a CTU. Based on this new weight A B-frame dynamic bit allocation control model is proposed, and a CTU-level rate allocation function is reconstructed as follows:
[0079]
[0080] wherein, is the unallocated bit number of the current frame; is the sum of the remaining CTU weights; is a preset sliding window, used to overcome excessive bit fluctuations.
[0081] This design uses the proposed perceptual-geometric joint weight to optimize the rate allocation mechanism of the B-frame, to improve the human eye's perceptual visual quality in different areas, achieve efficient dynamic bit rate allocation and control, and thus improve the performance of high-resolution video coding.
[0082] After determining the target allocated bit number of the I-frame or the B-frame, the average code rate is calculated to calculate the encoding parameter value of the current coding unit, so as to obtain the quantization parameter value (QP) of the current coding unit, and complete the rate control process.
[0083] The above process of optimizing the rate control framework based on the perceptual-geometric joint weight can be represented as follows: Figure 2 wherein, the pgj_weight is the perceptual-geometric joint weight of the CTU.
[0084] The following is an embodiment.
[0085] Embodiment 1
[0086] A perceptual-geometric joint rate control method, as shown in Figure 3 , comprises the following steps:
[0087] S1: According to the VVC standard, allocate weights to each CTU in each frame in the video;
[0088] S2: For each I-frame, calculate the perceptual-geometric joint weight of each CTU therein, and update the weight of each CTU in the I-frame according to ; for each B-frame, calculate the perceptual-geometric joint weight of each CTU therein, and update the weight of each CTU in the B-frame according to ;
[0089] S3: According to the weight of each CTU in each frame, allocate rate to each CTU to complete the rate control.
[0090] The beneficial effects achieved by the present embodiment are further analyzed and verified in combination with the relevant comparative experimental results.
[0091] In this embodiment, the rate error (Re) and exceeding the limit rate (Er) indicators are used to verify the code rate control performance of the proposed perceptual-geometric joint rate control optimization method (hereinafter referred to as PGJRC), and the effectiveness of PGJRC is verified by comparing the benchmark VTM14RC and the optimal related work LiRC under the same encoding scenario. The smaller the reaction code rate control precision is, the higher the precision is; the smaller the Er value is, the more effectively the encoding process follows the bandwidth limit. In the case of setting the same target code rate, in the AI encoding mode, the Er of the PGJRC proposed in this embodiment and the VTM14RC algorithm is 34.4% and 53.1% respectively, all 0.1%. In the LDB / RA encoding mode, the Er of the proposed PGJRC and the LiRC, VTM14RC algorithm is 81.3% / 98.4% / 92.2%, 92.2% / 90.6% / 90.6%, 1.2% / 1.3% / 1.7%, 3.0% / 3.0% / 3.4% respectively, and the Er of the proposed PGJRC and the indicators are the lowest in average in various encoding modes, which verifies that the code rate control performance of PGJRC is better than the baseline and other methods under the limited bandwidth limit.
[0092] Further, the rate-distortion performance of video encoding is compared by using the code rate saving performance indicator Bjøntegaard delta bitrate (BDBR), which reflects the ability of the RC algorithm to balance video quality and transmission efficiency. BDBR represents the bit rate saving of one method compared with another method under the same image quality, and the smaller the BDBR value is, the more bit rate is saved. Compared with the VVC benchmark RC strategy VTM14RC, the PGJRC proposed in this embodiment achieves a certain rate-distortion (R-D) performance improvement. In the AI / LDB / RA encoding mode, the average gain of BDBR is 0.256%, 2.345%, and 0.982% respectively.
[0093] Further, the encoding time saving rate indicator is used to measure the complexity of the proposed algorithm PGJRC, and at the same time, compared with the optimal work LiRC and the benchmark VTM14RC in the same application scenario and encoding mode, the practical application value of the PGJRC algorithm is verified. The smaller, the more time saving, the lower the algorithm complexity. In the AI / LDB / RA coding mode, compared with VTM14RC and LiRC respectively, the complexity of the proposed PGJRC is reduced by 13.8% and 9.5% on average. The above complexity comparison results verify that the PGJRC has good practical application value.
[0094] Further, the performance of the PGJRC algorithm in improving the quality of the reconstructed frame of the high-resolution video is verified by using the structural similarity (SSIM) subjective quality evaluation index. The focus is to verify the degree of fit of the proposed PGJRC with the characteristics of the HVS by using the SSIM value of the core visual attention area. The closer the SSIM value is to 1, the better the subjective quality of the image. The local brightness pixel value of the core visual attention area of the reconstructed frame, that is, the area in the upper and lower ranges of the transverse center position of the reconstructed frame, is collected. Then, the SSIM of the local position of the core visual attention area is calculated by using the reconstructed frame and the original frame in the AI / LDB / RA coding mode of the PGJRC and VTM14RC respectively. The experimental test results show that, compared with VTM14RC, the SSIM value of the proposed PGJRC is improved in the three coding modes, which indicates that the perception intensity of the PGJRC to the visual core attention area is the highest, and the degree of fit of the PGJRC with the characteristics of the HVS is verified. It is verified that the PGJRC can effectively improve the perceived subjective quality of the reconstructed frame of the high-resolution video.
[0095] Overall, the embodiment dynamically models the spatio-temporal coupling effect of the change gradient of the distortion tolerance inside and outside the field of view and the height of the coding unit in the plane, and fuses the visual saliency gradient modulation mechanism to construct a differentiable, perceptually adaptive weight model. Then, the new weight model is used to improve the rate allocation framework of VVC, realizing the perceptual-geometric collaborative optimization allocation of the CTU-level code rate, effectively eliminating the structural redundancy of high-resolution video, and making the code stream distribution more consistent with the spatial perception characteristics of the human visual system.
[0096] Embodiment 2:
[0097] A perceptual-geometric joint code rate control system, as shown in Figure 4 , includes:
[0098] A weight initialization module is configured to assign weights to each CTU in each frame of a video according to the VVC standard.
[0099] A perceptual-geometric joint weight construction module is configured to calculate the perceptual-geometric joint weight of the CTU.
[0100] The I frame weight optimization module is configured to calculate the perceptual-geometric joint weight of each CTU in each I frame by using the perceptual-geometric joint weight construction module, and update the weight of each CTU in the I frame according to the perceptual-geometric joint weight of each CTU. update the weight of each CTU in the I frame;
[0101] The B frame weight optimization module is configured to calculate the perceptual-geometric joint weight of each CTU in each B frame by using the perceptual-geometric joint weight construction module, and update the weight of each CTU in the B frame according to the perceptual-geometric joint weight of each CTU. update the weight of each CTU in the B frame;
[0102] The code rate allocation module is configured to allocate the code rate to each CTU according to the weight of each CTU in each frame, and complete the code rate control.
[0103] Embodiment 3
[0104] A computer readable storage medium includes a stored computer program, and the computer program, when executed by a processor, implements the perceptual-geometric joint code rate control method provided in Embodiment 1.
[0105] Specifically, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash storage device, or other volatile solid-state storage devices.
[0106] Those skilled in the art will easily understand that the above description is only a preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for perceptual-geometry joint code rate control, characterized in that, The method comprises: S1: assigning weights to each CTU in each frame of the video according to the VVC standard; S2: for each I-frame, calculate the perceptual-geometric joint weight of each CTU, and update the weight of each CTU in the I-frame according to S2: for each I-frame, calculate the perceptual-geometric joint weight of each CTU, and update the weight of each CTU in the I-frame according to The calculation method of the perceptual-geometric joint weight of the CTU comprises: With a field of view center reference line of the maximum weight region, calculate the viewing angle of the CTU relative to the field of view center reference line ; like If it is located in the core area of the viewpoint, then according to... Calculate its perceptual-geometric joint weights Otherwise, according to Calculate its perceptual-geometric joint weights ; wherein, represents the CTU weight calculated according to the VVC standard; is the projection plane height; is the vertical position of the top-left pixel of each CTU, is the CTU height; represents the total number of CTUs in the current frame, represents the perceptual-geometric joint weight of the th CTU in the current frame; updates the weight of the CTU with the view angle of in the I frame after updating; and are parameters to control the weight decay rate, and ; S3: assigning code rates to each CTU according to the weights of the CTUs in each frame to complete code rate control. 2.The perceptual-geometry joint code rate control method of claim 1, wherein, In S3, the code rate allocated to the CTU in the I frame is the 1st wherein, is the number of bits not allocated when encoding the i-th CTU of the current frame; is the number of bits not allocated when encoding the i-th CTU of the current frame; represents the weight of the i-th CTU in the I frame after updating the weight, represents the weight of the i-th CTU in the I frame after updating the weight. represents the weight of the i-th CTU in the I frame after updating the weight. represents the weight of the i-th CTU in the I frame after updating the weight. 3.The perceptual-geometry joint code rate control method of claim 1, wherein, S2 further comprises: for each B frame, calculating the perceptual-geometric joint weight of each CTU in the B frame, and updating the weight of each CTU in the B frame according to the following expression: wherein, , represents a code rate control experience parameter, represents a code rate, represents the number of pixels in a CTU; represents the weight of the CTU with the view angle of in the updated B frame. 4.The perceptual-geometry joint code rate control method of claim 3, wherein, In S3, the code rate allocated to the CTU at the 1th row in the B frame in S2 is: wherein, represents the weight of the i-th CTU in the B frame after updating the weight, represents the weight of the i-th CTU in the B frame after updating the weight; represents a preset sliding window. 5. The perceptual-geometry aware rate control method of any one of claims 1-4, wherein, A CTU's angle of view relative to the field center reference line The formula for calculating the angle of view is as follows: 。 6. The perceptual-geometry aware rate control method of any of claims 1-4, wherein, , ; wherein, represents the height of the frame.
7. The perceptual-geometry aware rate control method of any of claims 1-4, wherein, The view angle range corresponding to the view angle core region is .
8. A perception-geometry joint code rate control system, comprising: The method comprises: A weight initialization module configured to assign weights to each CTU in each frame of the video according to the VVC standard; A perceptual-geometric joint weight construction module configured to calculate the perceptual-geometric joint weight of the CTU; The I frame weight optimization module is configured to calculate the perception-geometry joint weights of each CTU in each I frame by using the perception-geometry joint weight construction module, and update the weights of each CTU in the I frame according to the perception-geometry joint weights of each CTU in each I frame. update the weights of each CTU in the I frame. A code rate assignment module configured to assign code rates to each CTU according to the weights of the CTUs in each frame to complete code rate control; The calculation method of the perceptual-geometric joint weight of the CTU comprises: With a field of view center reference line of the maximum weight region, calculate the viewing angle of the CTU relative to the field of view center reference line ; like If it is located in the core area of the viewpoint, then according to... Calculate its perceptual-geometric joint weights Otherwise, according to Calculate its perceptual-geometric joint weights ; represents the CTU weight calculated according to the VVC standard; is the projection plane height; is the vertical position of the top-left pixel of each CTU, is the CTU height; represents the total number of CTUs in the current frame, represents the perceptual-geometric joint weight of the th CTU in the current frame; is the weight of the CTU with view in the updated I-frame; and is a parameter to control the weight decay rate, and .
9. The perception-geometry coalesced rate control system of claim 8, wherein, Further comprising: a B frame weight optimization module configured to calculate the perceptual-geometric joint weight of each CTU in each B frame by using the perceptual-geometric joint weight construction module, and update the weight of each CTU in the B frame according to the following expression: wherein, , represents a code rate control experience parameter, represents a code rate, represents the number of pixels in a CTU; represents the weight of the CTU with the view angle of in the updated B frame.
10. A computer-readable storage medium, characterized in that, The method comprises: A stored computer program; when the computer program is executed by a processor, the perceptual-geometric joint code rate control method of any one of claims 1-7 is implemented.
Citation Information
Patent Citations
VVC code rate control optimization method for 360-degree video compression coding
CN117135352A
Screen content video coding perception code rate control method and device based on DCT domain JND
CN120751139A