A Dynamic Bitrate Allocation Method and System Based on AI Multi-Token Prediction
Through the AI multi-word element prediction method, the encoding unit is divided into high-dynamic range video and the bit rate allocation is optimized, which solves the problem that the bit rate allocation in the prior art is not suitable for human eye perception, and improves the video quality and encoding efficiency.
Patent Information
- Application Number
- CN202510707800.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
When processing high dynamic range videos, it is difficult to accurately allocate the code rate to adapt to the perception characteristics of the human eye, resulting in a decline in the quality of the video after encoding, especially in scenarios where light and dark alternately alternating rapidly.
Through the AI multi-word prediction method, the light and dark alternate regions are divided into coding units, and the visual saliency is calculated in the nonlinear visual response space, attention propagation map is constructed, visual correlation intensity is determined, code rate sharing group is established, perceived redundancy analysis is performed, and bit rate allocation is optimized.
The subjective quality of the encoded video is improved, and the coding efficiency and quality are improved by accurately identifying the visual correlation area and sharing the code rate resources, and optimizing the code rate allocation.
Smart Images

Figure CN120263986B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image communication, and in particular to a dynamic bit rate allocation method and system based on AI multi-word prediction. Background Art
[0002] With the continuous development of video coding technology, AI-based video coding has gradually become a research hotspot. In the video coding process, properly allocating bitrates is crucial for improving coding efficiency and video quality. Traditional bitrate allocation methods, primarily based on rate-distortion optimization theory, allocate bitrates by analyzing the complexity of video content. However, this approach often ignores the temporal correlation between video frames, resulting in inaccurate bitrate allocation. This is particularly true for video clips with scene changes and intense motion.
[0003] In related technologies, convolutional neural networks can be used to extract features from video frames, combined with long short-term memory (LSTM) networks to model the temporal characteristics of videos. End-to-end training can then be used to achieve adaptive bitrate allocation. Compared to traditional methods, this approach can better capture the spatiotemporal characteristics of videos and improve the accuracy of bitrate allocation.
[0004] However, when processing high dynamic range (HDR) video content, since HDR video has a wider brightness range and richer color information, the impact of different brightness areas on human visual perception varies significantly. Existing methods often regard all brightness areas as equally important when allocating bitrates. This makes it difficult to accurately allocate bitrates that adapt to the human eye's perception characteristics in scenes with drastic alternations between light and dark, thereby reducing the subjective quality of the encoded video. Summary of the Invention
[0005] The present application provides a dynamic bitrate allocation method and system based on AI multi-word prediction, which is used to improve the accuracy of bitrate allocation that adapts to the human eye perception characteristics in scenes with drastic alternations between light and dark, thereby improving the subjective quality of the encoded video.
[0006] In a first aspect, the present application provides a dynamic bitrate allocation method based on AI multi-word prediction to determine the alternating light and dark areas in the high dynamic range video frame to be encoded;
[0007] The alternating light and dark areas are divided into multiple coding units, and the brightness values of the coding units are mapped to a nonlinear visual response space based on the visual perception characteristics of the human eye;
[0008] In the nonlinear visual response space, the visual saliency of each coding unit is calculated, and the attention propagation map of the coding unit is constructed based on the visual saliency;
[0009] Determine the visual correlation strength between encoding units based on the attention propagation map;
[0010] Cluster the coding units with visual association strength greater than a preset threshold into a bitrate sharing group;
[0011] Within each bitrate sharing group, select the coding unit with the highest visual saliency as the leading unit, and determine the visual difference degree between other coding units except the leading unit and the leading unit;
[0012] Establish a bitrate sharing coefficient according to the visual difference degree, and allocate the bitrate of the leading unit to other coding units through the bitrate sharing coefficient;
[0013] Perform perceptual redundancy analysis on the boundary regions between bitrate sharing groups to identify the transition regions;
[0014] Transfer the spatial redundancy bitrate of the transition region to the visual key region with visual saliency greater than the preset saliency to obtain an optimized bitrate allocation scheme;
[0015] Encode the high-dynamic range video frame according to the optimized bitrate allocation scheme.
[0016] By adopting the above technical solution, by analyzing the visual saliency of coding units in the non-linear visual response space and constructing an attention propagation map, the regions with strong visual correlation are accurately identified and clustered into bitrate sharing groups, enabling regions with similar visual characteristics to share bitrate resources. By establishing a bitrate sharing mechanism based on visual difference degree within the bitrate sharing group, the high-quality coding of the leading unit is ensured, and at the same time, a reasonable bitrate is allocated to other coding units according to the visual difference degree. Perform perceptual redundancy analysis on the boundary regions between bitrate sharing groups, and transfer the redundant bitrate in the identified transition region to the visual key region, which not only ensures the smooth transition of the boundary region, improves the coding quality of the visually important region, realizes the perceptual optimized allocation of bitrate resources, improves the accuracy of allocating bitrates adapted to the human eye perception characteristics, and thus improves the subjective quality of the encoded video.
[0017] Combined with some embodiments of the first aspect, in some embodiments, determining the light and dark alternating regions in the high-dynamic range video frame to be encoded specifically includes:
[0018] Calculate the luminance difference between adjacent pixel points in the high-dynamic range video frame to be encoded to obtain a luminance difference matrix;
[0019] Perform multi-scale decomposition on the luminance difference matrix to obtain luminance gradient information at different scales;
[0020] Construct a luminance change frequency map according to the luminance gradient information at different scales, and mark the regions with luminance change frequency greater than the first preset threshold in the luminance change frequency map as candidate light and dark alternating regions;
[0021] Calculate the spatial connectivity of candidate light and dark alternating regions, and determine the candidate light and dark alternating regions with a spatial connectivity greater than a second preset threshold as light and dark alternating regions.
[0022] By adopting the above technical solution, the multi-scale decomposition method is used to analyze the luminance gradient information in the high-dynamic-range video frame. By constructing a luminance change frequency map, the luminance change characteristics at different scales can be comprehensively captured. Combining with the spatial connectivity analysis, the light and dark alternating regions with significant luminance changes and good spatial continuity can be accurately identified, enabling the system to concentrate the limited bitrate resources on the light and dark transition regions that really need fine coding, improving the coding efficiency and reconstruction quality.
[0023] Combined with some embodiments of the first aspect, in some embodiments, in the non-linear visual response space, calculate the visual saliency of each coding unit, specifically including:
[0024] Extract the local contrast feature, local direction feature and local color feature of the coding unit;
[0025] Calculate the luminance saliency map according to the local contrast feature, calculate the texture saliency map according to the local direction feature, and calculate the chromaticity saliency map according to the local color feature;
[0026] Determine the weight coefficients of the luminance saliency map, texture saliency map and chromaticity saliency map based on the human eye visual perception characteristic curve;
[0027] Weightedly fuse the luminance saliency map, texture saliency map and chromaticity saliency map according to the corresponding weight coefficients to obtain the visual saliency of the coding unit.
[0028] By adopting the above technical solution, when calculating the visual saliency of the coding unit, three visual feature dimensions of local contrast, direction and color are comprehensively considered. By calculating the luminance saliency map, texture saliency map and chromaticity saliency map respectively, the saliency of the coding unit in different visual features is characterized. By introducing weight coefficients based on the human eye visual perception characteristics for feature fusion, the saliency calculation result is more in line with the human eye visual perception law, and the obtained visual saliency more accurately reflects the contribution degree of the coding unit to the visual quality, providing a reliable visual importance measure for subsequent attention propagation and bitrate allocation.
[0029] Combined with some embodiments of the first aspect, in some embodiments, determine the visual association strength between coding units based on the attention propagation map, specifically including:
[0030] Construct the adjacency matrix of the attention propagation map, and the matrix elements of the adjacency matrix represent the visual saliency difference between adjacent coding units;
[0031] Perform eigen decomposition on the adjacency matrix to obtain eigenvectors and eigenvalues;
[0032] Calculate the visual association propagation path between coding units based on the eigenvector and eigenvalue;
[0033] Calculate the visual association strength between coding units based on the length of the visual association propagation path and the visual saliency difference.
[0034] By adopting the above technical solution, by constructing an adjacency matrix based on the visual saliency difference and performing eigen - decomposition on it, the visual association structure information between coding units can be effectively extracted. The visual association propagation path calculated based on the eigenvector and eigenvalue not only considers the saliency difference between directly adjacent units but also can reflect the indirect visual association generated by multi - step propagation between distant coding units. This method of calculating the visual association strength by combining the propagation path length and saliency difference enables the system to more accurately identify groups of coding units that are closely related in visual perception, providing a reasonable basis for subsequent rate - sharing group division.
[0035] Combined with some embodiments of the first aspect, in some embodiments, establish a rate - sharing coefficient according to the visual difference degree, specifically including:
[0036] Calculate the Euclidean distance between the dominant unit and other coding units in terms of luminance, texture, and chrominance features;
[0037] Construct a feature similarity matrix based on the Euclidean distance and perform normalization on the feature similarity matrix;
[0038] According to the normalized feature similarity matrix, calculate the rate - sharing weight of each other coding unit with respect to the dominant unit;
[0039] Take the rate - sharing weight as the rate - sharing coefficient.
[0040] By adopting the above technical solution, by calculating the Euclidean distance between the dominant unit and other coding units in terms of luminance, texture, and chrominance features, constructing a feature similarity matrix and performing normalization, the degree of visual feature difference between coding units can be accurately quantified. The rate - sharing weight calculated based on the normalized feature similarity matrix can establish a quantitative relationship for rate allocation between coding units, making the rate allocation more accurate, ensuring that coding units with similar visual features obtain similar rate allocations, while coding units with larger visual feature differences obtain differentiated rate allocations, thereby improving the rate utilization efficiency while ensuring the coding quality.
[0041] Combined with some embodiments of the first aspect, in some embodiments, after encoding a high - dynamic - range video frame according to the optimized rate - allocation scheme, the method further includes:
[0042] Obtain the encoding quality evaluation index of the high-dynamic range video frame and the decoded video frame;
[0043] Analyze the luminance distribution of each bitrate sharing group in the decoded video frame to obtain the luminance distribution curve;
[0044] Calculate the peak-valley ratio of each bitrate sharing group based on the luminance distribution curve, and determine the dynamic range compression coefficient of the bitrate sharing group according to the peak-valley ratio;
[0045] When the encoding quality evaluation index is lower than the preset quality threshold, apply the dynamic range compression coefficient to the dominant unit of the corresponding bitrate sharing group;
[0046] Recalculate the bitrate allocation ratio of other encoding units based on the compressed dominant unit;
[0047] Recode the high-dynamic range video frame based on the updated bitrate allocation ratio.
[0048] By adopting the above technical solution, through analyzing the luminance distribution of each bitrate sharing group in the decoded video frame, combining the encoding quality evaluation index and the dynamic range compression coefficient, an adaptive bitrate optimization feedback mechanism is established. When it is detected that the encoding quality does not meet the standard, the system can perform targeted dynamic range compression on the dominant unit of the bitrate sharing group according to the peak-valley ratio calculated from the luminance distribution curve, and recalculate the bitrate allocation ratio of other encoding units accordingly. This adaptive adjustment method based on the feedback of the encoding result enables the system to timely discover and correct unreasonable bitrate allocation situations. By applying dynamic range compression to the dominant unit, more bitrate space can be released while maintaining the visual quality, and the updated bitrate allocation ratio ensures that the released bitrate can be reasonably reallocated to other encoding units, thereby improving the overall encoding effect.
[0049] Combined with some embodiments of the first aspect, in some embodiments, calculating the peak-valley ratio of each bitrate sharing group based on the luminance distribution curve, and determining the dynamic range compression coefficient of the bitrate sharing group specifically includes:
[0050] Calculate the standard deviation and mean of the luminance values within each bitrate sharing group on the luminance distribution curve;
[0051] Determine the peak and valley positions of the luminance distribution according to the standard deviation, and calculate the luminance difference between adjacent peaks and valleys;
[0052] Divide the luminance difference by the mean to obtain the normalized peak-valley ratio;
[0053] Construct a piecewise function based on the peak-valley ratio. When the peak-valley ratio is greater than the first dynamic threshold, calculate the dynamic range compression coefficient according to the preset compression curve; when the peak-valley ratio is less than the second dynamic threshold, set the dynamic range compression coefficient to 1; when the peak-valley ratio is between the first dynamic threshold and the second dynamic threshold, linearly interpolate to obtain the dynamic range compression coefficient.
[0054] By adopting the above technical solution, by calculating the standard deviation and mean on the luminance distribution curve and constructing a piecewise function based on the peak and valley positions to determine the dynamic range compression coefficient, precise control of dynamic range compression is achieved. Dividing the luminance difference by the mean to obtain the normalized peak-valley ratio reduces the influence of different luminance levels on the calculation of the compression coefficient. The design of the piecewise function takes into account the characteristics of different peak-valley ratio intervals. When the peak-valley ratio is large, compression is performed using the preset compression curve; when the peak-valley ratio is small, the original dynamic range is maintained; when the peak-valley ratio is in the middle interval, smooth transition is achieved through linear interpolation. This adaptive compression scheme based on luminance distribution characteristics can apply an appropriate degree of dynamic range compression according to the specific characteristics of the image content, reducing the detail loss caused by over-compression and also reducing the bitrate waste caused by under-compression, enabling dynamic range compression to achieve the purpose of saving bitrate while well maintaining the visual quality of the image.
[0055] In a second aspect, an embodiment of the present application provides a dynamic bitrate allocation system based on AI multi-token prediction. The dynamic bitrate allocation system based on AI multi-token prediction includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the system to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0056] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, which when running on the system, cause the system to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0057] In a fourth aspect, an embodiment of the present application provides a computer program product, which when running on the system, causes the system to execute the method described in any possible implementation manner in the first aspect.
[0058] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0059] 1. The present application provides a dynamic bitrate allocation method based on AI multi-token prediction. By analyzing the visual saliency of coding units in the non-linear visual response space and constructing an attention propagation map, regions with strong visual relevance are accurately identified and clustered into bitrate sharing groups, enabling regions with similar visual characteristics to share bitrate resources. By establishing a bitrate sharing mechanism based on visual difference within the bitrate sharing groups, high-quality coding of the dominant units is ensured, and reasonable bitrates are allocated to other coding units according to the degree of visual difference. Perceptual redundancy analysis is performed on the boundary regions between bitrate sharing groups, and the redundant bitrates in the identified transition regions are transferred to visually critical regions, which not only ensures smooth transitions in the boundary regions, improves the coding quality of visually important regions, realizes the perceptual optimization allocation of bitrate resources, enhances the accuracy of allocating bitrates that adapt to the characteristics of human eye perception, and thus improves the subjective quality of the encoded video.
[0060] 2. The present application provides a dynamic bitrate allocation method based on AI multi-token prediction. By analyzing the luminance distribution of each bitrate sharing group in the decoded video frame and combining the coding quality evaluation index and the dynamic range compression coefficient, an adaptive bitrate optimization feedback mechanism is established. When it is detected that the coding quality does not meet the standard, the system can perform targeted dynamic range compression on the dominant units of the bitrate sharing group according to the peak-valley ratio calculated from the luminance distribution curve, and recalculate the bitrate allocation ratio of other coding units accordingly. This adaptive adjustment method based on the feedback of the coding result enables the system to timely detect and correct unreasonable bitrate allocation situations. By applying dynamic range compression to the dominant units, more bitrate space can be released while maintaining visual quality, and the updated bitrate allocation ratio ensures that the released bitrates can be reasonably reallocated to other coding units, thereby improving the overall coding effect.
[0061] 3. The present application provides a dynamic bitrate allocation method based on AI multi-token prediction. By adopting the above technical solutions, the standard deviation and mean are calculated on the luminance distribution curve, and a piecewise function is constructed based on the peak and valley positions to determine the dynamic range compression coefficient, realizing precise control of dynamic range compression. Dividing the luminance difference by the mean to obtain the normalized peak-valley ratio reduces the influence of different luminance levels on the calculation of the compression coefficient. The design of the piecewise function takes into account the characteristics of different peak-valley ratio intervals. When the peak-valley ratio is large, a preset compression curve is used for compression; when the peak-valley ratio is small, the original dynamic range is maintained; when the peak-valley ratio is in the middle interval, linear interpolation is used for smooth transition. This adaptive compression scheme based on luminance distribution characteristics can apply an appropriate degree of dynamic range compression according to the specific characteristics of the image content, reducing the detail loss caused by over-compression and the bitrate waste caused by under-compression, enabling dynamic range compression to not only achieve the purpose of saving bitrates but also well maintain the visual quality of the image. Brief Description of the Drawings
[0062] Figure 1 is a schematic flowchart of a dynamic bitrate allocation method based on AI multi-token prediction in an embodiment of the present application.
[0063] Figure 2 is a schematic flowchart of an adaptive optimization method for encoding quality in an embodiment of the present application.
[0064] Figure 3 is a schematic structural diagram of an entity device of a dynamic bitrate allocation system based on AI multi-token prediction provided in an embodiment of the present application. Detailed Embodiments
[0065] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above", "said", "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations including one or more of the listed items.
[0066] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0067] Next, an embodiment is used in combination with Figure 1 to describe a dynamic bitrate allocation method based on AI multi-token prediction in an embodiment of the present application:
[0068] Please refer to Figure 1 , which is a schematic flowchart of a dynamic bitrate allocation method based on AI multi-token prediction in an embodiment of the present application.
[0069] S101. Determine the light and dark alternating area in the high-dynamic range video frame to be encoded;
[0070] The system determines the light and dark alternating regions in the high-dynamic range video frame to be encoded, specifically including: calculating the luminance difference between adjacent pixel points in the high-dynamic range video frame to be encoded to obtain a luminance difference matrix; performing multi-scale decomposition on the luminance difference matrix to obtain luminance gradient information at different scales; constructing a luminance change frequency map based on the luminance gradient information at different scales, and marking the regions with a luminance change frequency greater than the first preset threshold in the luminance change frequency map as candidate light and dark alternating regions; calculating the spatial connectivity of the candidate light and dark alternating regions, and determining the candidate light and dark alternating regions with a spatial connectivity greater than the second preset threshold as the light and dark alternating regions. In this step, the system first needs to analyze the high-dynamic range video frame to be encoded to determine the regions with light and dark alternation. The light and dark alternating regions are usually caused by factors such as the reflection characteristics of objects in the scene and the change of lighting conditions, and are visually manifested as drastic changes in luminance. The purpose of determining the light and dark alternating regions is to perform targeted encoding processing on these regions subsequently to improve the encoding efficiency and visual quality.
[0071] To determine the light and dark alternating regions, the luminance difference between adjacent pixel points in the video frame can be calculated to generate a luminance difference matrix. Then, multi-scale decomposition is performed on this matrix to obtain luminance gradient information at different scales. Based on the gradient information, the system can construct a luminance change frequency map and mark the regions with a luminance change frequency greater than a certain threshold as candidate light and dark alternating regions. Finally, by calculating the spatial connectivity of the candidate regions, the candidate regions with a connectivity greater than a certain threshold are determined as the final light and dark alternating regions.
[0072] S102: Divide the light and dark alternating regions into multiple coding units, and map the luminance values of the coding units to the non-linear visual response space based on the human eye visual perception characteristics;
[0073] After determining the light and dark alternating regions, the system needs to divide them into multiple coding units for subsequent encoding processing. The size and shape of the coding units can be determined according to specific application scenarios and performance requirements. Common coding units include macroblocks, sub-blocks, coding tree units, etc.
[0074] To better adapt to the human eye visual perception characteristics, the system can map the luminance values of the coding units from the linear space to the non-linear visual response space. The human eye's perception of luminance is non-linear, being more sensitive to luminance changes in low-luminance regions and less sensitive in high-luminance regions. Mapping the luminance values to the non-linear space can make the luminance distribution of the coding units more conform to the human eye's perception characteristics, thereby improving the subjective visual quality. Commonly used non-linear visual response models include gamma correction, logarithmic transformation, sigmoid function, etc. The system can select an appropriate non-linear mapping function according to specific application scenarios and visual effect requirements.
[0075] S103. In the non - linear visual response space, calculate the visual saliency of each coding unit, and construct the attention propagation map of the coding unit according to the visual saliency;
[0076] In the non - linear visual response space, the system calculates the visual saliency of each coding unit, specifically including: extracting the local contrast feature, local direction feature and local color feature of the coding unit; calculating the luminance saliency map according to the local contrast feature, calculating the texture saliency map according to the local direction feature, and calculating the chromaticity saliency map according to the local color feature; determining the weight coefficients of the luminance saliency map, texture saliency map and chromaticity saliency map based on the human eye visual perception characteristic curve; and performing weighted fusion on the luminance saliency map, texture saliency map and chromaticity saliency map according to the corresponding weight coefficients to obtain the visual saliency of the coding unit. And construct the attention propagation map of the coding unit according to the visual saliency.
[0077] In the non - linear visual response space, the system needs to calculate the visual saliency of each coding unit, that is, the degree of attraction of the unit to the human eye visual attention. The visual saliency reflects the importance of the coding unit in the entire video frame, and is of great significance for subsequent bit - rate allocation and quality control.
[0078] There are various methods for calculating visual saliency. Common methods include those based on features such as local contrast, local direction, and local color. For example, the system can extract the local contrast feature of the coding unit, calculate the contrast difference between it and the surrounding units, and generate a luminance saliency map. Similarly, local direction features and local color features can also be extracted to generate a texture saliency map and a chromaticity saliency map respectively. Finally, these saliency maps are weighted and fused according to certain weight coefficients to obtain the comprehensive visual saliency of the coding unit.
[0079] After obtaining the visual saliency of the coding unit, the system can further construct the attention propagation map. The attention propagation map describes the transfer and diffusion relationship of visual attention between coding units, and can be used to guide subsequent bit - rate allocation. One method for constructing the attention propagation map is to regard the coding unit as a node of the graph, establish a directed edge according to their visual saliency and spatial adjacency relationship, and the weight of the edge reflects the intensity of attention propagation from one unit to another.
[0080] S104. Determine the visual association strength between coding units based on the attention propagation map;
[0081] The system determines the visual association strength between coding units based on the attention propagation map, specifically including: constructing the adjacency matrix of the attention propagation map, where the matrix elements of the adjacency matrix represent the visual saliency differences between adjacent coding units; performing eigenvalue decomposition on the adjacency matrix to obtain eigenvectors and eigenvalues; calculating the visual association propagation path between coding units according to the eigenvectors and eigenvalues; and calculating the visual association strength between coding units based on the length of the visual association propagation path and the visual saliency difference.
[0082] After constructing the attention propagation map, the system needs to further analyze the visual association strength between coding units, that is, their visual correlation and dependence. The visual association strength reflects the redundancy between coding units and is of great significance for subsequent bitrate allocation and quality control.
[0083] To determine the visual association strength, the system can perform graph-theoretic analysis on the attention propagation map. A feasible method is to represent the attention propagation map as an adjacency matrix, where the matrix elements represent the visual saliency differences between adjacent coding units. Then, perform eigenvalue decomposition on the adjacency matrix to obtain eigenvectors and eigenvalues. The eigenvectors represent the main directions and patterns of visual association of coding units, and the eigenvalues reflect the importance of these patterns. Based on the eigenvectors and eigenvalues, the system can calculate the visual association propagation path between coding units and quantify the visual association strength according to the length of the path and the saliency difference.
[0084] S105. Cluster the coding units with visual association strength greater than a preset threshold into a bitrate sharing group;
[0085] According to the visual association strength between coding units, the system can cluster them into multiple bitrate sharing groups. The coding units within a bitrate sharing group have strong visual correlation and can share the same quantization parameter or coding mode, thereby reducing the bitrate overhead and improving the compression efficiency.
[0086] To achieve the division of bitrate sharing groups, the system can set a threshold for visual association strength. The coding unit pairs with association strength greater than this threshold are grouped into the same bitrate sharing group. The size of the threshold determines the granularity and number of bitrate sharing groups, and needs to be balanced according to the specific application scenario and performance requirements. Generally speaking, the larger the threshold, the fewer the bitrate sharing groups, the stronger the correlation of the coding units within the group, and the greater the potential for bitrate savings; conversely, the smaller the threshold, the more the bitrate sharing groups, the weaker the correlation of the coding units within the group, and the smaller the potential for bitrate savings.
[0087] S106. Within each bitrate sharing group, select the coding unit with the highest visual saliency as the dominant unit, and determine the visual difference degree between other coding units except the dominant unit and the dominant unit;
[0088] After dividing the bitrate sharing groups, the system needs to select a dominant unit in each group as a reference for bitrate allocation and quality control. The dominant unit should have the highest visual saliency and be able to attract the viewer's attention to the greatest extent. Selecting the coding unit with the highest visual saliency as the dominant unit can ensure that important visual information is preferentially protected and transmitted under limited bitrates.
[0089] After selecting the dominant unit, the system also needs to calculate the visual difference degree between other coding units and the dominant unit. The visual difference degree reflects the secondary degree of non-dominant units relative to the dominant unit and is of great significance for bitrate allocation and quality control. One method of calculating the visual difference degree is to normalize the visual saliency of the dominant unit to 1 and then calculate the difference between the visual saliencies of other units and it. The larger the difference, the greater the visual difference between the unit and the dominant unit, and more consideration should be given to it during bitrate allocation.
[0090] S107. Establish a bitrate sharing coefficient based on the visual difference degree, and allocate the bitrate of the dominant unit to other coding units through the bitrate sharing coefficient;
[0091] The system establishes a bitrate sharing coefficient according to the visual difference degree, specifically including: calculating the Euclidean distances between the dominant unit and other coding units in terms of luminance, texture, and chromaticity features; constructing a feature similarity matrix based on the Euclidean distances and performing normalization processing on the feature similarity matrix; calculating the bitrate sharing weights of each other coding unit and the dominant unit according to the normalized feature similarity matrix; taking the bitrate sharing weights as the bitrate sharing coefficient. And allocate the bitrate of the dominant unit to other coding units through the bitrate sharing coefficient.
[0092] After determining the dominant unit and the visual difference degree, the system needs to further establish a bitrate sharing coefficient and allocate the bitrate of the dominant unit to other coding units accordingly. The bitrate sharing coefficient reflects the weight of non-dominant units in bitrate allocation, and its magnitude is inversely proportional to the visual difference degree. The larger the visual difference degree of a unit, the smaller its bitrate sharing coefficient and the less bitrate it is allocated; conversely, the smaller the visual difference degree of a unit, the larger its bitrate sharing coefficient and the more bitrate it is allocated.
[0093] To establish the bitrate sharing coefficient, the system can first calculate the Euclidean distances between the dominant unit and other coding units in each visual feature dimension, such as luminance, texture, chromaticity, etc. Then, combine these distances into a feature similarity matrix and perform normalization processing to make the value range of matrix elements between [0, 1]. The normalized feature similarity can be directly used as the bitrate sharing coefficient, or further adjusted through non-linear mappings such as exponential transformation or power function to adjust the sensitivity and dynamic range of bitrate allocation.
[0094] After obtaining the bitrate sharing coefficient, the system can distribute the bitrate of the dominant unit to other coding units according to the ratio of the coefficient. For example, assume the bitrate of the dominant unit is R, and there are N coding units in the bitrate sharing group, with their bitrate sharing coefficients being {w_1, w_2,..., w_N}, then the bitrate r_i assigned to the i-th coding unit is r_i = R * w_i / sum(w_1, w_2,..., w_N). This distribution method can balance the bitrate and distortion while ensuring the quality of the dominant unit and taking into account the visual importance of other coding units.
[0095] S108. Perform perceptual redundancy analysis on the boundary regions between bitrate sharing groups to identify transition regions;
[0096] When the system performs perceptual redundancy analysis on the boundary regions between bitrate sharing groups, it can utilize various image analysis techniques. The boundary regions usually contain pixels from two adjacent bitrate sharing groups, and their characteristics may lie between the two groups. By analyzing the visual characteristics of these pixels, such as luminance gradient, texture complexity, color change, etc., the system can evaluate the visual redundancy degree of this region.
[0097] Specifically, the system can extract the local contrast, local direction, and local color features of the pixels in the boundary region and compare them with the corresponding features of the adjacent bitrate sharing groups. If the features of the boundary region are closer to one of the groups, it can be classified into that group; if its features are significantly different from both groups, it can be marked as a transition region. In addition, the system can also analyze the continuity of the boundary region in the time dimension, that is, examine the feature changes of the corresponding boundary regions in adjacent video frames to further optimize the identification of transition regions.
[0098] S109. Transfer the spatial redundancy bitrate of the transition region to the visually significant regions whose visual saliency is greater than the preset saliency to obtain an optimized bitrate allocation scheme;
[0099] After identifying the transition regions, the system needs to reallocate their spatial redundancy bitrates to improve the overall quality of video coding. Since the transition regions usually have high visual redundancy, reducing their bitrate allocation will not significantly affect the perceptual quality. At the same time, the visually significant regions are more critical to the viewing experience, so they are worthy of being allocated more coding resources.
[0100] The system can calculate the visual saliency of each coding unit and set a preset saliency threshold. For regions whose saliency exceeds this threshold, the system can mark them as visually critical regions. Then, the system can transfer some or even all of the bitrates originally allocated to the transitional regions to these critical regions. The transfer ratio of the bitrate can be determined according to the saliency level of the critical regions, and the regions with higher saliency can obtain more bitrate increase. In this way, the system obtains an optimized bitrate allocation scheme, achieving the optimal visual quality under limited coding resources.
[0101] S110. Encode the high-dynamic-range video frames according to the optimized bitrate allocation scheme.
[0102] After obtaining the optimized bitrate allocation scheme, the system needs to perform actual encoding processing on the high-dynamic-range video frames accordingly. The encoding process usually includes multiple steps such as predictive coding, transform quantization, and entropy coding.
[0103] In the predictive coding stage, the system can make full use of the visual saliency information provided by the bitrate allocation scheme. For example, in intra-frame prediction, the system can preferentially select visually critical regions as reference samples to improve the prediction accuracy; in inter-frame prediction, the system can perform more fine-grained motion estimation and compensation for critical regions to achieve more accurate temporal redundancy elimination.
[0104] The transform quantization stage is crucial for bitrate control. The system can adopt different quantization parameters for different regions according to the bitrate allocation scheme. Smaller quantization steps can be used for visually critical regions to retain more details, while larger steps can be used for non-critical regions to save bitrate. This adaptive quantization strategy can maximize the perceptual quality while meeting the bitrate limit.
[0105] Finally, in the entropy coding stage, the system can adopt different coding modes and parameters for different regions to further improve the coding efficiency. For example, for critical regions with complex textures, the system can select a higher-order entropy coding model; for non-critical regions with smooth textures, the system can select a more concise model to save computational overhead.
[0106] In the above embodiments, by analyzing the visual saliency of coding units in the non-linear visual response space and constructing an attention propagation map, regions with strong visual relevance are accurately identified and clustered into rate-sharing groups, enabling regions with similar visual features to share rate resources. By establishing a rate-sharing mechanism based on visual difference degrees within the rate-sharing groups, high-quality coding of the dominant units is ensured, and reasonable rates are allocated to other coding units according to the visual difference degrees. Perceptual redundancy analysis is performed on the boundary regions between the rate-sharing groups, and the redundant rates in the identified transition regions are transferred to the visually critical regions, which not only ensures smooth transitions in the boundary regions, improves the coding quality of visually important regions, realizes the perceptual optimal allocation of rate resources, and enhances the accuracy of allocating rates that adapt to the human eye perception characteristics, thereby improving the subjective quality of the encoded video.
[0107] Based on the above dynamic rate allocation method based on AI multi-token prediction, the present application also provides an adaptive optimization method for coding quality. This method can dynamically adjust and optimize the rate allocation scheme when it detects that the coding quality does not meet the standard by establishing a feedback regulation mechanism for coding quality evaluation and dynamic range compression. The following combines Figure 2 to describe an adaptive optimization method for coding quality in the embodiments of the present application:
[0108] Please refer to Figure 2 which is a schematic flowchart of an adaptive optimization method for coding quality in the embodiments of the present application.
[0109] S201. Obtain the coding quality evaluation index of the high-dynamic range video frame and the decoded video frame;
[0110] In the process of adaptively optimizing the coding quality, the system first needs to obtain two key input data: the coding quality evaluation index and the decoded video frame. The coding quality evaluation index reflects the performance of the current coding scheme and can be generated based on subjective evaluation or objective metrics. For example, the system can adopt common objective quality evaluation indexes such as PSNR and SSIM, or collect user feedback through subjective evaluation experiments and calculate the subjective quality score (such as MOS score) accordingly. The decoded video frame provides an intuitive presentation of the coding output, facilitating subsequent visual analysis by the system.
[0111] Specifically, the system can connect a quality assessment module to the output of the encoder to calculate the quality assessment metrics for each frame in real time. The quality assessment can be carried out based on the comparison between the original video and the reconstructed video, or can be estimated only by using the statistical characteristics (such as distortion, blur, etc.) of the reconstructed video itself. For subjective quality assessment, the system can provide an interactive interface that allows users to rate the video quality and collect feedback data for metric calculation. In addition, the system also needs to decode the bitstream output by the encoder to obtain a complete sequence of decoded video frames, providing input for subsequent processing.
[0112] S202. Analyze the luminance distribution of each rate-sharing group in the decoded video frames to obtain a luminance distribution curve;
[0113] After obtaining the decoded video frames, the system needs to analyze their luminance distribution to gain insights into the luminance change characteristics of different rate-sharing groups. Luminance is an important component of visual signals, and its distribution has a decisive impact on the dynamic range and contrast of the image. By analyzing the luminance distribution of each group, the system can evaluate the rationality of the current rate allocation scheme and provide a basis for subsequent dynamic adjustment.
[0114] In specific implementation, the system can traverse each rate-sharing group, extract the luminance values of the pixels it contains, and count the frequencies of different luminance levels. Then, the system can draw a luminance distribution histogram based on the frequency information, with the horizontal axis representing the luminance range (such as 0 - 255) and the vertical axis representing the frequency or relative frequency. To better characterize the distribution characteristics, the system can also smooth the histogram to obtain a continuous luminance distribution curve. Features such as the shape, peak position, and dynamic range of the distribution curve contain rich luminance distribution information.
[0115] S203. Calculate the peak-valley ratio of each rate-sharing group based on the luminance distribution curve, and determine the dynamic range compression coefficient of the rate-sharing group according to the peak-valley ratio;
[0116] The system calculates the peak-valley ratio of each rate-sharing group based on the luminance distribution curve and determines the dynamic range compression coefficient of the rate-sharing group according to the peak-valley ratio, specifically including: calculating the standard deviation and mean of the luminance values within each rate-sharing group on the luminance distribution curve; determining the peak and valley positions of the luminance distribution according to the standard deviation, and calculating the luminance difference between adjacent peaks and valleys; dividing the luminance difference by the mean to obtain a normalized peak-valley ratio; constructing a piecewise function based on the peak-valley ratio, when the peak-valley ratio is greater than the first dynamic threshold, calculating the dynamic range compression coefficient according to a preset compression curve; when the peak-valley ratio is less than the second dynamic threshold, setting the dynamic range compression coefficient to 1; when the peak-valley ratio is between the first dynamic threshold and the second dynamic threshold, linearly interpolating to obtain the dynamic range compression coefficient.
[0117] After obtaining the luminance distribution curves of each bitrate sharing group, the system needs to further quantify its dynamic range characteristics. The peak-to-valley ratio is a commonly used dynamic range description index, which reflects the contrast difference between the highest peak and the lowest valley on the luminance distribution curve. A larger peak-to-valley ratio means a large span of luminance distribution within the bitrate sharing group, and more coding resources may be required to ensure the picture quality; conversely, a smaller peak-to-valley ratio means that the luminance distribution is relatively concentrated and the coding pressure is relatively small.
[0118] To calculate the peak-to-valley ratio, the system first needs to locate the peak and valley points on the luminance distribution curve. This can be achieved by finding the first derivative of the curve and detecting the points where the gradient changes sign. It should be noted that there may be multiple local extreme points on the curve, so the system needs to set certain screening conditions (such as peak-valley height difference threshold, adjacent peak-valley distance threshold, etc.) to filter out false peak-valley points. After finding the main peak and valley positions, the system can calculate the corresponding luminance difference and divide it by the global average luminance to obtain the normalized peak-to-valley ratio. The normalization process makes the peak-to-valley ratio independent of the specific luminance range and more general.
[0119] The level of the peak-to-valley ratio provides important clues for dynamic range compression. The system can design a segmented compression strategy based on the peak-to-valley ratio, that is, adopt different intensities of dynamic range compression for different ratio intervals. For example, when the peak-to-valley ratio is very high, the system can adopt a stronger compression curve to reduce the excessive luminance dynamic range; when the peak-to-valley ratio is low, the system can maintain the original dynamic range without compression; when the peak-to-valley ratio is at a medium level, the system can determine the compression coefficient between the two extremes by interpolation according to its proximity to the high and low thresholds. This adaptive compression strategy can take into account different luminance distribution characteristics, avoiding over-compression while suppressing excessive waste of bitrate.
[0120] S204. When the coding quality evaluation index is lower than the preset quality threshold, apply the dynamic range compression coefficient to the dominant unit of the corresponding bitrate sharing group;
[0121] The coding quality evaluation index provides timely feedback to the system to judge whether the current coding scheme has achieved the expected quality level. When the evaluation index is lower than the preset threshold, it means that the existing bitrate allocation scheme is not yet optimized and urgent dynamic adjustment is needed. For the bitrate sharing groups that do not meet the standard, the system can choose to apply the dynamic range compression coefficient obtained in the previous step to their dominant units to reduce the luminance dynamic range and thus save bitrate resources.
[0122] In specific implementation, the system first needs to set a reasonable quality threshold. This threshold can be determined according to factors such as the type of video content, user preferences, bandwidth limitations, etc. For example, for videos with rich texture details, the quality threshold can be set relatively high; while for videos with limited dynamic range, the quality threshold can be appropriately relaxed. The setting of the threshold can adopt heuristic rules or be optimized through machine learning methods. When the evaluation metric fails to reach the threshold, the system can identify the dominant unit (usually the coding unit with the highest visual saliency) within the bitrate sharing group and apply the dynamic range compression coefficient to the luminance value of this unit. The compression can be achieved through look-up table mapping or mathematical transformation. After compression, the luminance dynamic range of the dominant unit is reduced, and the encoder can correspondingly reduce the allocated bitrate, thereby saving resources.
[0123] S205. Recalculate the bitrate allocation ratio of other coding units based on the compressed dominant unit;
[0124] Once the luminance dynamic range of the dominant unit is compressed, the bitrate required for its encoding will also decrease accordingly. To make full use of the saved bitrate, the system needs to recalculate the bitrate allocation ratio of other coding units within the bitrate sharing group to achieve rebalancing of bitrate resources. The new bitrate allocation should, on the premise of ensuring the quality of the dominant unit, improve the encoding quality of non-dominant units as much as possible, thereby improving the overall performance of the bitrate sharing group.
[0125] In specific implementation, the system can adopt a bitrate reallocation strategy based on visual weights. That is, first calculate the visual weight of each coding unit within the bitrate sharing group. The calculation of the weight can comprehensively consider various visual saliency factors such as luminance, contrast, texture complexity, and temporal continuity. Then, the system can use the bitrate released after compression of the dominant unit and reallocate it to other coding units according to the proportion of visual weights. The unit with a larger weight will receive a greater increase in bitrate. This weighted allocation method can ensure that visually more important regions receive higher-priority bitrate guarantees. The specific bitrate allocation can be solved by optimizing the objective function, such as minimizing the weighted distortion or maximizing the subjective visual quality score, etc.
[0126] S206. Re-encode the high dynamic range video frame based on the updated bitrate allocation ratio.
[0127] When the bitrate allocation ratio within the bitrate sharing group is updated, the system needs to re-encode the current high dynamic range video frames to generate the final encoded bitstream. The re-encoding process will make full use of the new bitrate allocation ratio, dynamically adjust the quantization parameters and encoding modes of each coding unit, and strive to optimize the visual quality under the limited bitrate. At the same time, re-encoding also provides an opportunity for the encoder to perform adaptive optimization, enabling it to dynamically fine-tune and feedback correct the encoding decisions based on the previous encoding information.
[0128] When specifically implementing the re-encoding, the system first needs to convert the updated bitrate allocation ratio into actual encoding parameters. Usually, there is a definite functional relationship between the bitrate and the quantization parameter. For example, the higher the bitrate, the smaller the quantization step size and the lower the quantization distortion. Therefore, the system can use the rate-distortion model to find the optimal quantization parameter according to the target bitrate. After determining the quantization parameter, the system also needs to decide the encoding mode of the image block, including the intra prediction mode and the inter prediction mode. Generally, intra encoding is suitable for areas with complex textures and difficult to predict; while inter encoding is suitable for areas with smooth motion and high temporal redundancy. The system can use the previous encoding mode and correlation information to optimize the current encoding mode selection.
[0129] In the above embodiment, by analyzing the luminance distribution of each bitrate sharing group in the decoded video frame, combining the encoding quality evaluation index and the dynamic range compression coefficient, an adaptive bitrate optimization feedback mechanism is established. When it is detected that the encoding quality does not meet the standard, the system can perform targeted dynamic range compression on the dominant unit of the bitrate sharing group according to the peak-valley ratio calculated from the luminance distribution curve, and accordingly recalculate the bitrate allocation ratio of other coding units. This adaptive adjustment method based on the feedback of the encoding result enables the system to timely detect and correct unreasonable bitrate allocation situations. By applying dynamic range compression to the dominant unit, more bitrate space can be released while maintaining the visual quality, and the updated bitrate allocation ratio ensures that the released bitrate can be reasonably reallocated to other coding units, thereby improving the overall encoding effect.
[0130] The following describes the system in the embodiments of the present invention application from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of a dynamic bitrate allocation system based on AI multi-token prediction provided by the embodiments of the present application.
[0131] It should be noted that Figure 3 The structure of the system shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0132] As Figure 3As shown, the system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 302 or the program loaded from the storage section 308 into the Random Access Memory (RAM) 303, such as executing the method in the above embodiment. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.
[0133] The following components are connected to the I / O interface 305: an input section 306 including a camera, an infrared sensor, etc.; an output section 307 including a Liquid Crystal Display (LCD), a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.
[0134] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the Central Processing Unit (CPU) 301, various functions defined in the present invention are executed.
[0135] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0137] As another aspect, the present invention also provides a computer-readable storage medium, which may be included in the system described in the above embodiments; or it may exist alone without being assembled into the system. The above storage medium carries one or more computer programs, and when the above one or more computer programs are executed by a processor of a system, the system implements the method provided in the above embodiments.
[0138] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.
[0139] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted as "if determining...", "in response to determining...", "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".
[0140] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0141] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.
Claims
1. A dynamic bitrate allocation method based on AI multi-token prediction, characterized in that Including: Determine the bright-dark alternating regions in the high-dynamic range video frame to be encoded; Divide the bright-dark alternating regions into multiple coding units, and map the luminance values of the coding units to a non-linear visual response space based on the human visual perception characteristics; In the non-linear visual response space, calculate the visual saliency of each coding unit, and construct an attention propagation map for the coding units according to the visual saliency; Determine the visual association strength between the coding units based on the attention propagation map; Cluster the coding units with the visual association strength greater than a preset threshold into a bitrate sharing group; Within each bitrate sharing group, select the coding unit with the highest visual saliency as the dominant unit, and determine the visual difference degree between the other coding units except the dominant unit and the dominant unit; Establish a bitrate sharing coefficient according to the visual difference degree, and allocate the bitrate of the dominant unit to the other coding units through the bitrate sharing coefficient; Perform perceptual redundancy analysis on the boundary regions between the bitrate sharing groups to identify the transition regions; Transfer the spatial redundant bitrate of the transition regions to the visual key regions with visual saliency greater than a preset saliency to obtain an optimized bitrate allocation scheme; Encode the high-dynamic range video frame according to the optimized bitrate allocation scheme.
2. The method according to claim 1, wherein The determination of the bright-dark alternating regions in the high-dynamic range video frame to be encoded specifically includes: Calculate the luminance difference between adjacent pixel points in the high-dynamic range video frame to be encoded to obtain a luminance difference matrix; Perform multi-scale decomposition on the luminance difference matrix to obtain luminance gradient information at different scales; Construct a luminance change frequency map according to the luminance gradient information at different scales, and mark the regions with luminance change frequency greater than a first preset threshold in the luminance change frequency map as candidate bright-dark alternating regions; Calculate the spatial connectivity of the candidate bright-dark alternating regions, and determine the candidate bright-dark alternating regions with spatial connectivity greater than a second preset threshold as the bright-dark alternating regions.
3. The method according to claim 1, wherein In the non-linear visual response space, the calculation of the visual saliency of each coding unit specifically includes: Extract the local contrast feature, local direction feature, and local color feature of the coding unit; Calculate a luminance saliency map according to the local contrast feature, calculate a texture saliency map according to the local direction feature, and calculate a chromaticity saliency map according to the local color feature; Determine the weight coefficients of the luminance saliency map, the texture saliency map, and the chromaticity saliency map based on the human visual perception characteristic curve; Perform weighted fusion on the luminance saliency map, the texture saliency map, and the chromaticity saliency map according to the corresponding weight coefficients to obtain the visual saliency of the coding unit.
4. The method according to claim 1, characterized in that The determination of the visual association strength between the coding units based on the attention propagation map specifically includes: Construct an adjacency matrix of the attention propagation map, and the matrix elements of the adjacency matrix represent the visual saliency difference between adjacent coding units; Perform eigen-decomposition on the adjacency matrix to obtain eigenvectors and eigenvalues; Calculate the visual association propagation path between the coding units according to the feature vectors and the eigenvalues; Calculate the visual association strength between the coding units based on the length of the visual association propagation path and the visual saliency difference.
5. The method according to claim 1, wherein The establishment of the bitrate sharing coefficient according to the visual difference degree specifically includes: Calculate the Euclidean distances between the dominant unit and the other coding units in terms of luminance, texture, and chrominance features; Construct a feature similarity matrix based on the Euclidean distances and perform normalization processing on the feature similarity matrix; Calculate the bitrate sharing weight of each of the other coding units with respect to the dominant unit according to the normalized feature similarity matrix; Use the bitrate sharing weight as the bitrate sharing coefficient.
6. The method according to claim 1, characterized in that, After encoding the high-dynamic range video frame according to the optimized bitrate allocation scheme, the method further includes: Obtain the coding quality evaluation index of the high-dynamic range video frame and the decoded video frame; Analyze the luminance distribution of each bitrate sharing group in the decoded video frame to obtain a luminance distribution curve; Calculate the peak-valley ratio of each bitrate sharing group based on the luminance distribution curve, and determine the dynamic range compression coefficient of the bitrate sharing group according to the peak-valley ratio; When the coding quality evaluation index is lower than a preset quality threshold, apply the dynamic range compression coefficient to the dominant unit of the corresponding bitrate sharing group; Recalculate the bitrate allocation ratio of the other coding units according to the compressed dominant unit; Recode the high-dynamic range video frame based on the updated bitrate allocation ratio.
7. The method according to claim 6, wherein The calculation of the peak-valley ratio of each bitrate sharing group based on the luminance distribution curve, and the determination of the dynamic range compression coefficient of the bitrate sharing group according to the peak-valley ratio specifically include: Calculate the standard deviation and mean of the luminance values within each bitrate sharing group on the luminance distribution curve; Determine the peak and valley positions of the luminance distribution according to the standard deviation, and calculate the luminance difference between adjacent peaks and valleys; Divide the luminance difference by the mean to obtain a normalized peak-valley ratio; Construct a piecewise function based on the peak-valley ratio. When the peak-valley ratio is greater than a first dynamic threshold, calculate the dynamic range compression coefficient according to a preset compression curve; when the peak-valley ratio is less than a second dynamic threshold, set the dynamic range compression coefficient to 1; when the peak-valley ratio is between the first dynamic threshold and the second dynamic threshold, linearly interpolate to obtain the dynamic range compression coefficient.
8. A dynamic bitrate allocation system based on AI multi-token prediction, characterized in that, The system includes: One or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the system to execute the method according to any one of claims 1-7.
9. A computer-readable storage medium, comprising instructions, characterized in that, When the instructions run on the system, cause the system to execute the method according to any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product runs on the system, cause the system to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Face region detection based light field video compression
CN111492657A
Underground video coding method and device
CN119653091A