Video coding method and device
By semantic segmentation of video frames and label map-guided encoding units divided, predicted and quantified, the problem of insufficient utilization of semantic information in the HEVC encoding method is solved, and more efficient video encoding and compression performance is achieved.
Patent Information
- Application Number
- CN202510199061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-22
AI Technical Summary
The traditional HEVC video encoding method lacks the use of semantic information, and the encoding efficiency needs to be further improved, making it difficult to meet the needs of efficient video encoding.
By performing semantic segmentation processing on video frames, a semantic label map is generated, and the encoding unit is divided, predicted, transformed and quantified by the encoding unit based on the semantic label map. Combined with the optimization criteria of rate distortion, adaptively adjust the size and prediction mode of the encoding unit to optimize the encoding process.
Improves the efficiency and compression performance of video encoding, reduces encoding artifacts, and improves reconstruction quality and encoding efficiency.
Smart Images

Figure CN120358352A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video encoding and decoding technologies, and particularly to a video encoding method and apparatus. Background Art
[0002] Widespread application of the High Efficiency Video Coding (HEVC) standard: As a new generation of video coding standard, compared with the previous H.264 / AVC standard, the HEVC has significantly improved in both coding efficiency and compression performance. The HEVC has been widely applied in fields such as ultra-high definition video coding, low-latency video communication, and video streaming transmission. However, with the diversification of video content and the continuous improvement of application requirements, traditional HEVC coding methods still have some limitations, such as insufficient utilization of semantic information and room for further improvement in coding efficiency. Summary of the Invention
[0003] This application provides a video coding method and apparatus, which can improve the efficiency and compression performance of video coding.
[0004] To achieve the above object, this application provides a video coding method, which includes:
[0005] Performing semantic segmentation processing on the current frame in the video to obtain the semantic label map of the current frame;
[0006] Determining the prediction information of the coding units in the current frame based on the semantic label map;
[0007] Encoding the current frame based on the prediction information.
[0008] To achieve the above object, this application also provides an electronic device, which includes a processor; the processor is used to execute instructions to implement the steps of the above method.
[0009] To achieve the above object, this application also provides a computer-readable storage medium, which is used to store instructions / program data, and the instructions / program data can be executed to implement the above method.
[0010] The video coding method of this application performs semantic segmentation processing on the current frame in the video to obtain the semantic label map of the current frame; encodes the current frame based on the semantic label map to obtain the encoding result of the current frame, so as to introduce the semantic label map of the image to optimize the encoding of the image, thereby improving the efficiency and compression performance of video coding. Description of the Drawings
[0011] The drawings described herein are used to provide a further understanding of this application, and constitute a part of this application. The illustrative embodiments and descriptions of this application are used to explain this application, and do not constitute an improper limitation to this application. In the drawings:
[0012] Figure 1 It is a schematic flowchart of an embodiment of the video encoding method of the present application;
[0013] Figure 2 It is a schematic structural diagram of an embodiment of the electronic device of the present application;
[0014] Figure 3 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Specific Embodiments
[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application. Additionally, unless otherwise specified (e.g., "or alternatively" or "or in an alternative"), the term "or" as used herein refers to a non-exclusive "or" (i.e., "and / or"). And, the various embodiments described herein are not necessarily mutually exclusive, because some embodiments can be combined with one or more other embodiments to form new embodiments.
[0016] The present application proposes a video encoding method. This video encoding method performs semantic segmentation processing on the current frame in the video to obtain the semantic label map of the current frame; encodes the current frame based on the semantic label map to obtain the encoding result of the current frame, so as to introduce the semantic label map of the image to optimize the encoding of the image, thereby improving the efficiency and compression performance of video encoding.
[0017] Specifically, as Figure 1 shown, a video encoding method of an embodiment proposed by the present application specifically includes the following steps. It should be noted that the following step numbers are only used for simplified description and are not intended to limit the execution order of the steps. The steps of this embodiment can be arbitrarily changed in the execution order on the basis of not violating the technical idea of the present application.
[0018] S101: Perform semantic segmentation processing on the current frame in the video to obtain the semantic label map of the current frame.
[0019] The semantic segmentation processing can be performed on the current frame in the video to obtain the semantic label map of the current frame; so as to subsequently encode the current frame based on the semantic label map of the current frame, introduce the semantic label map of the image to optimize the encoding of the image, thereby improving the efficiency and compression performance of video encoding.
[0020] In one implementation, the semantic segmentation model can be used to perform semantic segmentation processing on the current frame to obtain the semantic label map.
[0021] Among them, a pre-trained deep learning semantic segmentation model (such as FCN, U-Net) can be used to perform semantic segmentation processing on the current frame to obtain a pixel-level semantic label map, where each pixel in the current frame can be identified as a specific semantic category, such as people, cars, buildings, green vegetation, etc.
[0022] In another embodiment, the semantic segmentation model may include a backbone network, a semantic segmentation branch, and a contrast learning branch. That is, semantic segmentation and contrast learning can share the same backbone network (such as ResNet, EfficientNet, etc.) and perform semantic segmentation and contrast learning tasks on different branches respectively. Backbone network: Use a deep convolutional neural network (such as DCNN) as the backbone network to extract low-level features of the input image. This backbone network can be a pre-trained model or can be trained from scratch. Semantic segmentation branch: Add a segmentation head (such as FCN, U-Net, etc.) on the output feature map of the backbone network to generate a pixel-level semantic label map. The task of the segmentation head is to classify each pixel and label its belonging semantic category (such as people, cars, background, etc.). Contrast learning branch: Add a contrast learning module on the output feature map of the backbone network to extract high-level semantic features. The contrast learning module learns discriminative feature representations by maximizing the similarity between positive sample pairs and the difference between negative sample pairs. During the training process, the loss functions of the semantic segmentation task and the contrast learning task can be optimized simultaneously. For example, the loss functions of the two tasks are weighted and summed to form a multi-task loss function. During actual semantic segmentation processing, the current frame can be input into the semantic segmentation model including the backbone network, the semantic segmentation branch, and the contrast learning branch, and then the outputs of the semantic segmentation branch and the contrast learning branch are fused. For example, the outputs of the semantic segmentation branch and the contrast learning branch are fused by means of channel-level fusion, weighted fusion, recursive fusion, etc. to obtain the semantic label map of the current frame.
[0023] In another implementation, the semantic segmentation results of the current frame obtained by using the semantic segmentation model for semantic segmentation processing can be used for contrast learning of different semantic regions. Specifically, the image can be divided into different semantic regions (such as foreground, background, people, vehicles, etc.) through the label map generated by the semantic segmentation model, and then contrast learning is performed within these regions to ensure that the contrast learning module focuses on specific semantic categories and learns more discriminative feature representations to obtain the semantic label map. In this way, the semantic features can be enhanced through contrast learning to generate a more robust and discriminative semantic label map.
[0024] In another implementation, the current frame can be semantically segmented by a semantic segmentation model to obtain a first feature map; the current frame can be processed by a contrastive learning model for higher-level semantic feature extraction to obtain a second feature map; the first feature map and the second feature map are fused to obtain a semantic label map of the current frame.
[0025] In yet another implementation, an adaptive contrastive learning mechanism can be designed to dynamically adjust the contrastive learning strategy according to the semantic segmentation result, so that the semantic segmentation model after contrastive learning can generate a more robust and discriminative semantic label map. For example, positive and negative sample pairs can be adaptively selected according to the class distribution in the semantic label map, or the weight of the contrastive learning loss function can be adjusted.
[0026] In the above implementation, a semantic segmentation model and a contrastive learning module can be introduced to perform semantic analysis on the input video frames. The semantic segmentation model divides the video frames into different semantic regions, such as foreground, background, people, cars, etc., and generates a pixel-level semantic label map. The contrastive learning module extracts high-level semantic features through self-supervised learning to form a semantic feature map. The outputs of these two modules provide rich semantic prior information for the subsequent encoding process, enabling the encoder to perform adaptive encoding optimization according to the semantic structure of the video content. The semantic segmentation model can adopt a deep convolutional neural network and generate accurate semantic labels through end-to-end training. The contrastive learning module can learn a discriminative semantic feature representation by maximizing the similarity between positive sample pairs and the difference between negative sample pairs.
[0027] Optionally, the semantically segmented map generated by the semantic segmentation process can be a low-resolution semantic label map, that is, the resolution of the semantic label map obtained by the semantic segmentation process is lower than the resolution of the original image frame; in this case, the semantic label map obtained by the semantic segmentation process can be upsampled to obtain a semantic label map with the same resolution as the original image frame (i.e., the current frame), so as to reduce the computational overhead of the semantic segmentation process and enable the use of the final semantic label map to guide the encoding of the current frame through the upsampling restoration method. In other embodiments, the resolution of the semantic label map generated by the semantic segmentation process can also be the same as that of the original image frame. In this way, the semantic label map generated by the semantic segmentation process can be directly used for guided encoding in the subsequent process. Of course, in some scenarios, the semantic label map can also be further processed before being used to guide the encoding of the current frame.
[0028] S102: Encode the current frame based on the semantic label map to obtain the encoding result of the current frame.
[0029] Optionally, after performing semantic segmentation processing on the current frame in the video to obtain the semantic label map of the current frame, the current frame can be encoded based on the semantic label map of the current frame to introduce the semantic label map of the image to optimize the encoding of the image, thereby improving the efficiency and compression performance of video encoding.
[0030] Among them, the current frame can be divided into multiple coding units, and then the multiple coding units are compressed by sequentially performing processing such as prediction, transformation, quantization, and / or entropy coding on the multiple coding units to obtain the coding result of the current frame.
[0031] In this way, when encoding the current frame, the current frame can be first divided into multiple coding units.
[0032] Optionally, the current frame can be divided into multiple coding units using the division method of modern coding standards (such as HEVC / H.265, etc.).
[0033] Alternatively, based on the semantic label map of the current frame, the current frame can be divided to obtain multiple coding units.
[0034] In one implementation, each semantic region can be divided separately, that is, each semantic region is separately divided into coding units to determine at least one coding unit in each semantic region.
[0035] In another implementation, the video frame can be initially divided into coding units of the same size, and these coding units will serve as the basis for further adaptive division; then, in combination with the semantic label map, the initially divided coding units are adjusted. The size of the initially divided coding units can be set according to the actual situation and is not limited here. For example, it can be 64*64 or 32*32.
[0036] Optionally, based on the semantic label map, the semantic importance of each coding unit can be determined, and then whether to merge and / or further divide the coding unit can be determined based on the semantic importance to adjust the initially divided coding units based on the semantic label map.
[0037] Among them, for the initially divided coding units, the semantic category distribution and importance score inside them can be calculated. Among them, calculate the pixel ratio of different semantic categories inside each coding unit, and combine these ratios and predefined semantic weights to evaluate the semantic importance score of each coding unit. For regions with important semantics (such as faces, texts, etc.), mark them as high-importance coding units; for regions with secondary semantics (such as backgrounds), mark them as low-importance coding units.
[0038] Subsequently, the size of the coding unit can be adaptively adjusted according to the importance score of the coding unit.
[0039] For high-importance coding units, the coding unit can be further divided, i.e., the final coding unit adopts a smaller size (such as 32x32 or 16x16) to obtain finer prediction and transformation. That is, the coding unit splitting decision can be made according to the information in the semantic label map. For semantically important regions, it is inclined to be split into smaller coding units to obtain finer prediction and transformation. During the splitting decision process, the rate-distortion optimization criterion can also be combined to dynamically adjust the splitting strategy to ensure the best balance between the coding efficiency and the reconstruction quality of the split coding units.
[0040] For low-importance coding units, the coding unit may not be further divided, or two adjacent coding units with the same semantic category and the same semantic importance can be merged. In this way, the low-importance coding unit can adopt a larger size (such as 64x64) to reduce the coding overhead. That is to say, the coding unit merging decision can be made according to the information in the semantic label map. For coding units within the same semantic region, they are inclined to be merged into larger coding units to reduce the coding overhead. During the merging decision process, the rate-distortion optimization criterion is combined to ensure the best balance between the coding efficiency and the reconstruction quality of the merged coding units.
[0041] During the above coding unit merging and splitting process, the decision-making strategy can be adaptively adjusted. Specifically, according to the semantic importance score of the current coding unit, the merging and splitting strategies can be dynamically adjusted to improve the overall coding efficiency and reconstruction quality.
[0042] During the process of adjusting the coding unit size, the boundaries between different-sized coding units can also be aligned to avoid coding artifacts at the boundaries.
[0043] In addition, when dividing the coding unit, try to align the coding unit boundary with the semantic boundary. This can be achieved by analyzing the boundary information in the semantic label map and referring to these boundaries during the coding unit division process to ensure that the coding unit boundary does not cross the semantic boundary as much as possible. By aligning with the semantic boundary, the coding error and artifacts at the semantic boundary can be reduced, and the reconstruction quality can be improved.
[0044] The method of recursive splitting can be adopted. For each initial coding unit, the sub-region division is recursively performed until the optimal semantic adaptive division effect is achieved, that is, the division strategy of the coding unit is adaptively adjusted according to the semantic importance score and boundary information of the coding unit.
[0045] In some embodiments, the rate-distortion optimization (RDO) criterion can also be combined to comprehensively consider the trade-off between coding efficiency and reconstruction quality, and the division strategy is dynamically adjusted during the semantic adaptive division process to achieve overall optimization.
[0046] Among them, rate-distortion optimization can be performed on the basis of the preliminary semantic adaptive coding unit division. By calculating the coding cost and reconstruction distortion of each coding unit, the advantages and disadvantages of different division strategies are evaluated. The dynamically adjusted RDO criterion is adopted to adjust the weight between the coding cost and the reconstruction distortion according to the semantic importance score in different regions.
[0047] Its implementation method can be:
[0048] (1) Refine the coding unit division: On the basis of the preliminary RDO, that is, after using the rate-distortion optimized coding unit preliminary division result, further refine the coding unit division strategy. For high-importance coding units, further divide them into smaller sub-coding units to achieve more refined coding; for low-importance coding units, merge them into larger coding units to reduce the coding complexity. Combine the refined semantic information to dynamically adjust the coding unit division strategy to ensure the best balance between coding efficiency and reconstruction quality.
[0049] (2) Multi-level RDO optimization: During the coding unit division process, adopt a multi-level RDO optimization strategy. At each level, according to the division result of the current coding unit, dynamically adjust the RDO criterion, evaluate the effects of different division strategies, and perform refined optimization. Through the multi-level optimization process, gradually refine the coding unit division strategy to ensure that the optimal coding effect can be achieved at each level.
[0050] After dividing the current frame into multiple coding units by the above method or the conventional division method, each coding unit in the current frame can be processed such as prediction, transformation, quantization, and entropy coding to obtain the coding results of each coding unit in the current block.
[0051] When predicting the coding unit, the information of the semantic label map can also be referred to.
[0052] In one implementation, the coding unit in the current frame can be predicted based on the semantic label map to determine the prediction information of the coding unit, where the prediction information can include information such as the prediction mode, reference frame, and / or motion vector.
[0053] Optionally, if a coding unit uses the intra-frame prediction method to generate a prediction block, the intra-frame prediction mode of the coding unit can be determined based on the semantic label map (such as the specific intra-frame prediction direction in the ordinary intra-frame prediction mode, the block copy prediction mode, etc.).
[0054] In the first embodiment, the information of the semantic label map can be utilized to select the same or similar prediction modes for coding units within the same semantic region, so as to improve the prediction consistency of the same semantic region. For coding units at the semantic boundary, prediction modes that can better adapt to the edge direction can be selected to reduce coding artifacts. For example, for coding units in the background region, the horizontal intra prediction mode or the vertical intra prediction mode can be selected as candidate prediction modes, and for coding units in the foreground region, more or more complex candidate prediction modes can be selected to capture detailed features.
[0055] In the second embodiment, the mapping relationship between each semantic category and the preferred prediction mode can be predefined according to the semantic label, so that the preferred prediction mode corresponding to the semantic category of the coding unit is used as the candidate prediction mode of the coding unit; and then the final prediction mode of the coding unit is selected from the candidate prediction modes of the coding unit through methods such as cost comparison. This embodiment not only reduces the number of candidate modes but also reduces the computational complexity of prediction; and compared with the traditional method of mainly determining the intra prediction mode of the coding unit based on the correlation of the already reconstructed pixels in the neighborhood and the rate-distortion optimization criterion, this embodiment combines the semantic information of the coding unit and the directional features of the already reconstructed pixels around the coding unit during the prediction of the coding unit, improving the overall prediction effect. Among them, the cost comparison method can be the rate-distortion optimization comparison method. For example, for coding units belonging to the "person" category, their optimal prediction modes may mainly be the vertical and diagonal intra prediction modes; for coding units belonging to the "sky" category, their optimal prediction modes may mainly be the horizontal intra prediction mode.
[0056] In an example, the preferred intra prediction mode corresponding to the texture property of the coding unit is used as the candidate intra prediction mode of the coding unit, and the texture property of the coding unit is determined by the semantic category of the coding unit. In this way, the preferred prediction modes corresponding to the smooth region and the texture region can be preset; then, based on the discrimination result of which of the smooth region and the texture region the semantic category of the coding unit corresponds to, the preferred prediction mode corresponding to the judgment result is used as the candidate prediction mode of the coding unit; when the number of candidate prediction modes is multiple, the final prediction mode of the coding unit is selected from the candidate prediction modes of the coding unit through methods such as cost comparison. Exemplarily, for the texture region, more angular prediction modes are preferably selected as the preferred prediction modes; for the smooth region, the DC mode is preferably selected as the preferred prediction mode; if a coding unit is classified as "sky", then it may be automatically assigned to the processing of the smooth region; if it is "grassland" or "trees", it may be assigned to the processing of the texture region.
[0057] In another example, the preferred intra prediction mode corresponding to the importance of the coding unit is used as the candidate intra prediction mode of the coding unit, and the importance of the coding unit is determined by the semantic category of the coding unit. In this way, the preferred prediction mode corresponding to each importance can be preset; the importance of the coding unit determined by the semantic category of the coding unit; the preferred prediction mode corresponding to the importance of the coding unit is used as the candidate prediction mode of the coding unit; when the number of candidate prediction modes is multiple, the final prediction mode of the coding unit is selected from the candidate prediction modes of the coding unit by means of cost comparison, etc. For example, for high-importance coding units, multiple intra prediction modes are used for comprehensive evaluation to select the optimal prediction mode; for low-importance coding units, simplified prediction modes are used to reduce the computational complexity. That is, the number of preferred intra prediction modes corresponding to the coding unit is positively correlated with the importance of the coding unit.
[0058] In a specific example, the steps of performing intra prediction on a coding unit based on a semantic label map may include: prediction mode initialization, semantics-based consistency prediction, and prediction direction optimization.
[0059] 1. Prediction mode initialization
[0060] 1.1 Initialize the prediction mode according to the semantic label map
[0061] During the coding process, first, the prediction mode of each coding unit (Coding Unit) can be initialized according to the semantic segmentation result (i.e., the semantic label map). The semantic label map divides the image into different semantic regions (such as sky, road, building, person, etc.), and these regions have similar visual features and structural characteristics. Therefore, for coding units within the same semantic region, similar prediction modes tend to be selected to maintain semantic consistency.
[0062] Initial prediction mode selection: For each coding unit, an initial prediction mode can be preset according to its semantic category. For example:
[0063] Smoothing regions (such as sky, wall): Initially select the DC mode or the planar mode because the changes in these regions are small, and using simple prediction modes can effectively reduce the computational complexity.
[0064] Texture regions (such as grassland, trees, building surface): Initially select the angular prediction mode, especially those modes that can capture the texture direction (such as horizontal, vertical, diagonal, etc.) to better adapt to the complex texture structure.
[0065] Edge regions (such as object boundaries): Initially select the angular prediction mode that can capture the edge direction to reduce the prediction error at the edge.
[0066] Based on historical data or pre-trained models: To ensure the rationality of the initial prediction pattern, historical data or pre-trained models can be utilized to guide the selection of the prediction pattern. For example, by analyzing the relationship between semantic regions and prediction patterns in a large number of encoded images, a mapping table can be constructed for quickly looking up the optimal prediction pattern corresponding to each semantic category. Additionally, deep learning models (such as convolutional neural networks, CNNs) can be used to automatically learn the method of inferring the best prediction pattern from the semantic label map.
[0067] 1.2 Maintaining semantic consistency
[0068] Within the same semantic region, adjacent coding units usually have similar visual features and structural characteristics, so their prediction patterns should also be kept as consistent as possible. This helps reduce prediction errors and improve the overall efficiency of encoding. Specific practices include:
[0069] Region division: According to the semantic label map, the image is divided into multiple semantic regions. For each region, the same prediction pattern can be selected during initialization, or a set of similar prediction patterns can be chosen within the region.
[0070] Neighborhood constraint: During the initialization process, consider the semantic information between adjacent coding units to ensure the selection of the same or similar prediction patterns within the same semantic region. For example, if the left and upper neighbors of a certain coding unit have both selected the vertical prediction pattern, then the current coding unit can also preferentially select the vertical prediction pattern to maintain spatial consistency.
[0071] 2. Semantic-based consistent prediction
[0072] 2.1 Using the semantic label map to guide the selection of prediction patterns
[0073] During the actual encoding process, intra-frame prediction is not only based on the initial prediction pattern but can also be dynamically adjusted according to the specific content of the current block. At this time, the semantic label map can serve as important reference information to help the encoder select prediction patterns that can maintain semantic consistency.
[0074] Semantic consistency constraint: For coding units within the same semantic region, try to select the same or similar prediction patterns to maintain spatial consistency. For example, when processing the "building" region, all coding units can preferentially select directional prediction patterns (such as vertical, horizontal, or diagonal patterns) that can capture the texture of the building. This not only reduces prediction errors but also avoids inconsistent prediction results between different coding units, thereby improving the overall encoding quality.
[0075] Neighborhood analysis: When selecting a prediction mode, not only the semantic category of the current coding unit should be considered, but also the semantic information of its adjacent coding units should be analyzed. If the adjacent coding units belong to the same semantic category, the same prediction mode can be preferentially selected; if the adjacent coding units belong to different semantic categories, an appropriate prediction mode can be selected according to the features at the boundary (such as the edge direction) to reduce the cross-category prediction error.
[0076] Adaptive adjustment: Although the initial prediction mode is set based on the semantic label map, during the actual coding process, the encoder can make adaptive adjustments according to the content of the current block. For example, if the actual content of a coding unit does not match its initial prediction mode (such as a large prediction error), other prediction modes can be reselected to obtain a better coding effect.
[0077] 2.2 Improving the spatial consistency and accuracy of prediction
[0078] Through semantic-based consistency prediction, the spatial consistency of the prediction mode can be maintained within the same semantic region, thereby reducing the prediction error and improving the coding accuracy. Specifically:
[0079] Reducing block artifacts: Since the prediction modes between adjacent coding units are kept consistent, there will be no obvious prediction differences at the boundaries between blocks, thus reducing block artifacts and improving the visual quality of the reconstructed image.
[0080] Enhancing detail retention: For regions with rich textures, selecting an appropriate angular prediction mode can better capture details and reduce the loss of high-frequency components, thereby retaining more image details.
[0081] Simplifying the coding process: By maintaining the consistency of the prediction mode, the encoder can reduce unnecessary mode switches, lower the coding complexity, and improve the coding efficiency.
[0082] 3. Prediction direction optimization
[0083] 3.1 Adaptive selection of the optimal prediction direction according to the semantic feature map
[0084] When selecting a prediction direction, the optimal prediction direction can be adaptively selected according to the information in the semantic feature map (such as the gradient map, LBP map, etc.). The semantic feature map not only provides pixel-level semantic category information but also contains local visual features (such as gradients, textures, etc.), and this information can help the encoder more accurately select the prediction direction.
[0085] Texture region: For the parts marked as texture regions, more angular prediction modes tend to be selected. This is because texture regions usually contain rich details and variations. Using multi-angle prediction modes can better capture these details and reduce prediction errors. For example, for natural textures such as trees and grasslands, angular prediction modes such as 45° and 63° can be selected to better adapt to the complex texture structure.
[0086] Smoothing region: For the parts marked as smoothing regions, the DC mode or low-frequency transform basis functions tend to be selected. The variations in smoothing regions are small. Using simple DC mode or planar mode can effectively reduce the computational complexity while maintaining high prediction accuracy. For example, for large-area smoothing regions such as the sky and walls, selecting the DC mode can reduce prediction errors and at the same time reduce the coding complexity.
[0087] 3.2 Combining Rate-Distortion Optimization Criterion
[0088] When selecting the prediction direction, in addition to considering semantic features, the Rate-Distortion Optimization (RDO) criterion can also be combined to ensure the optimization of the coding effect. RDO is a commonly used coding optimization method that selects the optimal coding parameters by weighing the coding bit rate (rate) and the quality of the reconstructed image (distortion).
[0089] Dynamically adjusting the prediction direction: When selecting the prediction direction, the encoder can dynamically adjust the prediction direction according to the RDO criterion to achieve the best coding effect. Specifically, the encoder calculates the rate-distortion cost of each candidate prediction direction and selects the direction with the minimum cost as the final prediction direction. This can ensure that the coding bit rate is minimized while ensuring the coding quality.
[0090] Balancing prediction accuracy and coding complexity: By combining the RDO criterion, the encoder can find the best balance between prediction accuracy and coding complexity. For example, for regions with rich textures, although using more angular prediction modes can improve prediction accuracy, it will also increase the coding complexity. Therefore, the encoder can select the prediction direction that can minimize the coding bit rate while ensuring prediction accuracy through the RDO criterion.
[0091] Optionally, if an encoding unit uses the inter-frame prediction method to generate prediction blocks, the inter-frame prediction information (such as reference frames, motion vectors, etc.) of the encoding unit can be determined based on the semantic label map.
[0092] Among them, in inter-frame prediction, the reference frame of the current encoding unit can be selected according to the information of the semantic label map.
[0093] Based on the semantic label maps of the current frame and the candidate reference frames, an image frame containing a region with a high semantic similarity degree (i.e., the similarity degree meets the first preset condition) to the coding unit can be determined from the candidate reference frames; based on the image frame containing the region with a high semantic similarity degree to the coding unit, the reference frame of the current frame can be determined.
[0094] Optionally, the above candidate reference frames can be at least some of the previously encoded frames before the current frame. More preferably, the above candidate reference frames can be at least some of the image frames in the reference frame list. In other embodiments, the above candidate reference frames can be image frames selected by measurement conditions other than the semantic similarity degree measurement condition. For example, they can be image frames selected by pixel similarity and / or distance difference.
[0095] The semantic similarity degree reflects the consistency and coincidence degree of the semantic content of two regions. The semantic similarity degree can be calculated from the data in the semantic label map of the coding unit in the current frame and the data in the semantic label map of the region corresponding to the coding unit in the reference frame. For example, the semantic IoU (Intersection over Union) or semantic cosine similarity of the data in the semantic label map of the coding unit in the current frame and the data in the semantic label map of the region corresponding to the coding unit in the reference frame can be calculated to obtain the semantic similarity degree between the coding unit and the corresponding region of the reference frame.
[0096] Among them, a high semantic similarity degree can be understood as the semantic similarity degree being higher than the first preset value. The first preset value can be a default value or set dynamically according to the actual situation. For example, the selection condition is to select about 10% of the candidate reference frames from high to low based on the semantic similarity degree. In this case, the first preset value can be set to the highest value of the semantic similarity degree among the 90% of the candidate reference frames that are not selected. Among them, a high semantic similarity degree between two regions can also be understood as the semantic categories of the two regions being the same.
[0097] If there are multiple image frames containing regions with a high semantic similarity degree to the coding unit, the image frames containing regions with a high semantic similarity degree to the coding unit can be further selected by measurement conditions such as pixel similarity and / or distance difference to obtain the reference frame of the current frame.
[0098] In one embodiment, based on the semantic label maps of the current frame and the candidate reference frames, an image frame containing a region with a high degree of semantic similarity to the coding unit is determined from the encoded frames and used as the rough selection reference frame for the coding unit in the current frame. Then, based on the distance and / or similarity between the rough selection reference frame and the current frame, the reference frame for the coding unit is determined from the rough selection reference frames of the coding unit. For example, the rough selection reference frame with the shortest distance to the current frame is used as the reference frame for the coding unit, or the rough selection reference frame with the highest similarity to the current frame is used as the reference frame for the coding unit.
[0099] In another embodiment, based on the distance and / or pixel similarity between the encoded frames and the current frame, a rough selection reference frame for the current frame can be determined from the encoded frames before the current frame. Then, an image frame containing a region with the same semantic category as the coding unit can be determined from the rough selection reference frames of the current frame to further screen the rough selection reference frames of the current frame and determine the rough selection reference frame for the coding unit. Furthermore, the reference frame for the coding unit can be determined from the screened rough selection reference frames.
[0100] In yet another embodiment, an image frame containing a region with the same semantic category as the coding unit can be determined from the reference frame list to further screen the candidate reference frames for the coding unit. Furthermore, the reference frame for the coding unit can be determined from the screened rough selection reference frames.
[0101] In still another embodiment, the candidate reference frames for the coding unit can be determined based on information such as the spatial domain motion information and / or historical motion information of the coding unit. Then, an image frame containing a region with the same semantic category as the coding unit can be determined from the candidate reference frames to further screen the candidate reference frames for the coding unit. Furthermore, the reference frame for the coding unit can be determined from the screened rough selection reference frames.
[0102] In other embodiments, the encoded frames can be prioritized based on whether there are regions with the same / similar semantic category as the coding unit in the encoded frames. Preferably, the candidate frames in the reference frame list can be prioritized so that reference frames with the same / similar semantic category can be preferentially selected for predictive reference, thereby improving the accuracy of motion compensation and solving the problem of insufficiently fine reference frame selection in traditional methods. More preferably, when prioritizing the encoded frames, whether there are regions with the same / similar semantic category as the coding unit in the encoded frames can be used as the main sorting condition, and the distance and / or pixel similarity degree between the encoded frames and the current frame can be used as the secondary sorting condition. In this way, traditional metric conditions such as distance and / or similarity can be comprehensively considered to sort the encoded frames, so as to preferentially select encoded frames with the same / similar semantics and relatively short distance or high similarity as reference frames.
[0103] In other implementation manners, it is also possible to weight the semantic similarity degree between the coding unit and the corresponding area of the reference frame and the values of other metric conditions such as the pixel similarity and / or distance difference to obtain the similarity; determine an image frame with a high similarity (i.e., the similarity meets the second preset condition) from the candidate reference frames; and determine the reference frame of the current frame based on the image frame with a high similarity.
[0104] Among them, a high similarity can be understood as a similarity higher than a second preset value. The second preset value can be a default value or dynamically set according to the actual situation. Alternatively, if a preset number of image frames are to be selected, the similarities of the preset number of image frames ranked at the head of the candidate reference frames in the order of similarity from high to low are all high, and the similarities of the remaining image frames are low.
[0105] In a specific example, it is possible to weight the pixel similarity (i.e., pixel difference) and the semantic similarity degree between the coding unit and the corresponding area of the reference frame to obtain the similarity; then determine an image frame with a high similarity from the candidate reference frames; and determine the reference frame of the current frame based on the image frame with a high similarity. In this way, according to the comprehensive similarity index, a reference frame with a high semantic similarity degree and a small pixel difference can be preferentially selected. Such a reference frame can provide a more accurate match during motion compensation, reduce the prediction residual, and improve the coding efficiency. By introducing a weight parameter, the weights of the semantic similarity degree and the pixel similarity degree can be dynamically adjusted, and optimized selection can be performed according to different scenarios and semantic information. By combining semantic information for reference frame selection, the temporal correlation of the video content can be more accurately reflected, and the accuracy and efficiency of inter-frame prediction can be improved.
[0106] Through the above embodiments, for coding units in the same semantic region, it is possible to preferentially select reference frames containing the same semantic information to improve the semantic consistency of reference frame selection. For example, for a foreground moving object, a reference frame containing the same object is selected. Exemplarily, for a coding block belonging to the "car" category, a reference frame also containing "car" is preferentially selected. For a static background, a reference frame with a similar background can be selected, or a reference frame with a relatively close time distance can be selected. And by combining the distance and similarity measurement methods to optimize the reference frame selection. In this way, compared with the traditional method of selecting reference frames based on distance and similarity, by introducing semantic information, the reference frame selection strategy can be further optimized. Through the reference frame selection based on the semantic similarity degree, the accuracy of motion compensation can be improved, and the prediction residual can be reduced.
[0107] In addition, semantic segmentation can be performed on each reference frame to generate its semantic label map, and the semantic segmentation map can be stored in a memory for subsequent use in intra-frame prediction frames.
[0108] In inter-frame prediction, the motion vector of the current coding unit can also be determined according to the information of the semantic label map.
[0109] In one implementation manner, the candidate motion vectors of the current coding unit may be determined based on the coded blocks in the current frame or the reference frame that have a high semantic similarity degree with the current coding unit (i.e., the similarity degree meets the third preset condition); and the best motion vector of the current coding unit may be determined based on the candidate motion vectors of the current coding unit. In this way, the optimal motion vector prediction mode can be selected by analyzing the semantic consistency between adjacent coding units. Among them, according to the information of the semantic label map, when predicting the motion vector, the adjacent blocks with the same semantic category as the current coding unit are preferentially selected as candidate predictors and given a higher prediction weight to determine the best motion vector of the current coding unit, so as to improve the accuracy of motion vector prediction and reduce the residual that can be coded.
[0110] In one embodiment, the motion vectors of the coded blocks in the current frame or the reference frame that have a high semantic similarity degree with the current coding unit may be added to the motion information candidate list as candidate motion vectors. When generating the motion vector prediction value, the adjacent blocks with the same or similar semantic category as the current coded block are preferentially selected as candidate blocks. For example, for a foreground moving object, a reference frame containing the same object is selected and a consistent motion vector is used; for the coded blocks belonging to the "person" category, the motion vectors of the adjacent blocks also belonging to the "person" category are preferentially selected as the prediction values. For a static background, a reference frame with a relatively short temporal distance is selected and a zero motion vector is used.
[0111] During the prediction process, the motion vector may be corrected according to the information of the semantic label map to ensure the rationality and accuracy of the motion vector in different regions. Optionally, the motion information such as the motion information of the spatial block and / or the temporal block and / or the historical motion vector (HMVP) of the current coding unit may also be added to the motion information candidate list, that is, the traditional inter-frame prediction method (such as the AWVP method) can be jointly used to optimize the motion vector prediction. Based on the semantically similar candidate blocks, the spatial and temporal adjacent blocks are added to form a joint prediction, so that through the semantically guided motion vector prediction, the motion correlation can be better utilized and the coding efficiency can be improved.
[0112] In another embodiment, the position difference between the coded blocks in the current frame or the reference frame that have a high semantic similarity degree with the current coding unit and the current coding unit may be calculated and used as the candidate motion vector of the current coding unit.
[0113] Optionally, the best motion vector of the current coding unit may be selected from the candidate motion vectors of the current coding unit by means of prediction cost comparison or the like.
[0114] During the motion vector prediction process, the prediction model is adaptively adjusted by combining the motion prior information in the semantic label map to improve the prediction accuracy and robustness. The process may include: generating a motion prior map, generating a motion prior map based on the semantic label map to represent the motion trend of each block; adaptively adjusting the prediction model, adjusting the motion vector prediction model according to the motion prior map.
[0115] In another embodiment, the search range of the encoded region of the current coding unit in the reference frame may be determined based on the semantic category of the current coding unit; a matching block of the current coding unit is searched in the determined search range; the motion information of the current block is calculated based on the positions of the current coding unit and the matching block.
[0116] Optionally, during the motion estimation process, according to the information in the semantic label map, an initial motion estimation search range is set for different semantic regions. Compared with the traditional method of using a fixed-size search window, by comprehensively considering semantic information and dynamically adjusting the size and position of the search window, it can adapt to the motion characteristics of different semantic regions to reduce unnecessary computational overhead, thereby improving the efficiency and accuracy of motion estimation. For example, for a static background region, the search range can be smaller; for a moving foreground object, the search range can be larger; then, according to the predefined mapping relationship between the semantic category and the search range, the size of the search window suitable for the current coding block is determined. For example, for a coding block belonging to the "background" category, its motion is usually small, and a smaller search range can be used; for a coding block belonging to the "moving object" category, its motion is usually large, and a larger search range can be used.
[0117] In addition, a potential motion area can be determined according to the semantic tag of the current coding unit and the semantic tag map of the reference frame. For example, the search range of the current coding unit is set to an area in the reference frame where the semantic similarity degree with the current coding unit meets a fourth preset condition. Here, the fourth preset condition may refer to the area with the highest semantic similarity degree between the reference frame and the current coding unit. For example, for areas belonging to the same semantic category, the motion displacement may be relatively consistent, so the search range can be narrowed. That is, if there is an encoded block in the coding units belonging to the same semantic area as the current coding unit, the search range of the current block can be determined based on the motion vector of the encoded block belonging to the same semantic area as the current coding unit. For example, the search range of the current coding unit is set to a smaller area around the position pointed by the motion vector of the encoded block belonging to the same semantic area as the current coding unit. For areas near the semantic boundary, the motion may be more complex, and the search range can be expanded. In this way, by combining the motion prior information provided by the semantic tag map, such as the motion direction and speed of the semantic area, the search range and direction are further optimized. By analyzing the motion features in the semantic feature map, the size and shape of the search window are adaptively adjusted to improve the efficiency and accuracy of motion estimation.
[0118] Furthermore, the selection of the search starting point can be guided by combining the semantic tag map to further improve the efficiency of motion estimation. In this way, through the adaptive determination of the search range, unnecessary computational overhead can be reduced while ensuring the accuracy of motion estimation.
[0119] In addition, a multi-resolution search strategy can also be used. Based on a rough search, refinement is carried out to gradually narrow the search range and improve the efficiency and accuracy of motion estimation. Through the semantic-guided determination of the search range, unnecessary computations can be significantly reduced, and the accuracy and efficiency of motion estimation can be improved.
[0120] Optionally, the matching cost between multiple candidate blocks and the current coding unit can be compared to search for the matching block of the current coding unit within the determined search range.
[0121] Among them, in the cost comparison method in the embodiments of the present application and when calculating the above matching cost, the cost can be calculated to evaluate the advantages and disadvantages of corresponding prediction information such as candidate blocks, motion vectors, and prediction modes. The steps may include: determining the candidate prediction information of the coding unit, where the prediction information includes a motion vector and / or a prediction mode; combining the semantic tag map to calculate the semantic cost of the candidate prediction information of the coding unit; and determining the best prediction information of the coding unit from the candidate prediction information based on the semantic cost.
[0122] In one embodiment, traditional methods can be used to calculate the corresponding cost of the current coding unit, such as SAD, SATD or SSD (Sum of Squared Differences), and these metrics reflect the degree of difference at the pixel level. For example, the matching cost between the current block and the candidate block is calculated by methods such as SAD or SSD.
[0123] In another embodiment, the information of the semantic label map can be combined to calculate the corresponding cost of the current coding unit. For example, by combining the information of the semantic label map, the semantic matching cost between the current coding unit and the candidate block in the reference frame is calculated. The semantic matching cost (i.e., the above-mentioned semantic cost) can be based on metrics such as semantic IoU and semantic cross-entropy to measure the degree of overlap and similarity between two blocks in terms of semantic categories.
[0124] In addition, the pixel cost of the candidate prediction information of the coding unit can also be calculated; based on the total cost, the best prediction information of the coding unit is determined from the candidate prediction information, and the total cost is obtained by weighting the semantic cost and the pixel cost. In this way, the semantic matching cost and the pixel matching cost (i.e., the above-mentioned pixel cost) are weighted and combined to form a comprehensive motion estimation cost function. By introducing semantic factors, it is possible to better distinguish between semantically consistent motion vectors and semantically inconsistent motion vectors, and improve the accuracy of motion estimation. When calculating the comprehensive matching cost, the weights of the semantic matching cost and the pixel matching cost are dynamically adjusted and optimized according to different semantic regions and scenarios. Optionally, the weight of the semantic cost of the coding unit is positively correlated with the semantic importance of the coding unit, and the semantic importance of the coding unit is determined based on the data of the coding unit in the semantic label map. In this way, for semantically important regions, the weight of the semantic matching cost is increased; for secondary regions, the weight of the pixel matching cost is increased. Using the multi-objective optimization strategy can reduce the computational complexity while ensuring the motion estimation accuracy, and improve the overall coding efficiency. Through the semantic-guided matching cost calculation, it is possible to more accurately evaluate and select motion vectors, and improve the accuracy and efficiency of motion estimation.
[0125] In the above prediction scheme, intra-frame prediction can be guided by semantics, which can make full use of the spatial correlation of semantics, improve the prediction efficiency, reduce the energy of the coding residuals, and improve the compression performance. Moreover, through semantic-guided inter-frame prediction and motion estimation, the temporal correlation of video content can be better captured, the temporal prediction effect can be improved, and the compression efficiency can be enhanced.
[0126] In addition, semantic information can also assist in the merging and refinement of motion vectors, avoid redundant motion information, and reduce the coding overhead.
[0127] Among them, after obtaining the residual data of the coding unit through intra-frame prediction / inter-frame prediction of the coding unit, the residual data of the coding unit can be subjected to transformation processing. Among them, the residual data can be calculated based on the original block and the prediction block of the coding unit, and the prediction block of the coding unit is obtained by performing intra-frame prediction / inter-frame prediction on the coding unit. That is to say, the residual data of the coding unit can be determined based on the prediction information of the coding unit.
[0128] Conventional methods can be used to perform transformation processing on the residual data.
[0129] More preferably, semantic label map-guided transformation processing can also be combined. Among them, the semantic label map can be combined to adjust the transformation parameters of the coding unit. In this way, through semantic-guided transformation, the spatial correlation of semantics can be fully utilized, the energy of coding residuals can be reduced, and the compression performance can be improved.
[0130] In one implementation, the transformation mode of each coding unit of the current frame can be initialized according to the semantic label map of the current frame. For example, the transformation mode of each coding unit of the current frame is initialized to the transformation mode corresponding to the semantic category of each coding unit. In this way, for the coding units within the same semantic region, similar transformation modes tend to be selected to maintain semantic consistency. Optionally, the initial transformation mode can be set based on historical data or a pre-trained model to ensure the rationality of the initial selection.
[0131] In another implementation, the transformation parameters of the current coding unit can be determined based on the already encoded blocks in the current frame or the reference frame that have a high degree of semantic similarity with the current coding unit (that is, the similarity degree meets the fifth preset condition). For example, the transformation mode of the already encoded block in the current frame or the reference frame that has a high degree of semantic similarity with the current coding unit is used as the transformation mode of the current coding unit, so that for the coding units within the same semantic region, a transformation mode that can maintain semantic consistency is selected. More preferably, the transformation parameters of the current coding unit are determined based on the already encoded blocks in the surrounding area of the coding unit whose semantic similarity with the coding unit meets the fifth preset condition, that is, the transformation parameters of the current coding unit can be determined based on the already encoded blocks with a high degree of semantic similarity among the surrounding coding units of the current coding unit. For example, the transformation mode of the already encoded block with the highest degree of semantic similarity among the surrounding coding units of the current coding unit is used as the transformation mode of the current coding unit. In this way, by analyzing the semantic information between adjacent coding units, it is ensured that the same or similar transformation modes are selected within the same semantic region, improving the spatial consistency and accuracy of the transformation.
[0132] In addition, for coding units at semantic boundaries, it is preferable to select transformation modes that can better adapt to the edge direction, such as selecting the direction adaptive transformation mode or the direction lifting transformation mode, etc., to reduce coding artifacts. Further, the prediction mode and the transformation mode can be adaptively adjusted by analyzing the semantic information of the boundary region to ensure the coding quality at the boundary. It can be understood that in this solution, semantic boundary detection can be performed first, where the semantic boundaries within the coding unit can be detected based on the information of the semantic label map. During the coding unit partitioning and intra-frame prediction processes, special attention should be paid to processing the semantic boundary region. The semantic label map is processed using an edge detection algorithm (such as the Canny edge detection) to accurately locate the semantic boundaries. And during the semantic boundary processing, coding artifacts are suppressed by combining semantic information. Specifically, artifact removal processing can be performed in the semantic boundary region by introducing an edge protection filter. Methods such as Gaussian filtering and bilateral filtering are used to smooth the boundary region, reduce artifacts, and improve the visual quality.
[0133] Optionally, the optimal transformation mode and / or transformation kernel can be adaptively selected according to the information of the semantic label map. For example, the transformation mode and / or transformation kernel of the coding unit can be determined based on the texture property corresponding to the semantic category of the coding unit. For example, if it is confirmed according to the semantic category of the coding unit that the coding unit belongs to the texture region, a high-frequency transformation basis function can be selected for the coding unit, that is, the texture region tends to select a high-frequency transformation basis function; if it is confirmed according to the semantic category of the coding unit that the coding unit belongs to the smooth region, a low-frequency transformation basis function can be selected for the coding unit, that is, the smooth region tends to select a low-frequency transformation basis function. During the transformation kernel selection process, the rate-distortion optimization criterion can also be combined to dynamically adjust the transformation kernel to achieve the best coding effect. In addition, the transformation mode and / or transformation kernel of the coding unit can also be determined based on the importance of the coding unit.
[0134] Optionally, after the transformation, the transformed data can be quantized. Optionally, the frequency coefficients after the transformation can be quantized. The frequency coefficients after the transformation can be converted from the residual data in the spatial domain by a discrete cosine transform (DCT) or other similar transformation methods (such as wavelet transform, direction adaptive transformation, etc.).
[0135] In one implementation, traditional quantization methods can be used to quantize the transformed data.
[0136] In another implementation, the quantization process of the transformed data of the current frame can be guided by combining the semantic label map of the current frame.
[0137] Among them, the quantization parameter of the coding unit can be adaptively adjusted according to the semantic information of the coding unit in the semantic label map, so as to guide the quantization processing of the transform data of the current frame. Through the semantic label map, different quantization strategies can be adopted for different semantic regions to achieve more refined and flexible control.
[0138] Before performing the quantization of the transform data, the importance of different regions can be evaluated according to the semantic information and mapped to the quantization parameter. Optionally, the quantization step size of the coding unit is negatively correlated with the importance and / or visual sensitivity of the coding unit, and the importance and / or visual sensitivity of the coding unit is determined by the data of the coding unit in the semantic label map.
[0139] Among them, the semantic importance of the coding unit can be determined according to the semantic category of the coding unit, and the quantization step size of the coding unit can be adjusted according to the semantic importance of the coding unit. Among them, the quantization step size of the coding unit is negatively correlated with the semantic importance of the coding unit. The semantic category of the coding unit is determined by the semantic information of the coding unit in the semantic label map. Among them, in step S101, the current coding unit (coding unit) is divided into different semantic categories, such as foreground, background, face, text, etc. by using a semantic segmentation algorithm; in step S102, corresponding weight coefficients can be assigned to each semantic category according to a predefined semantic importance weight table, and the weight coefficients reflect the degree of influence of the semantic category on the visual quality. Usually, the weight coefficients of semantic important regions such as foreground and face are higher, while the weight coefficients of background regions are lower; next, the semantic importance weight is mapped to the adjustment factor of the quantization parameter. For semantic regions with higher weights, the quantization parameter adjustment factor is smaller to retain more high-frequency details; for semantic regions with lower weights, the quantization parameter adjustment factor is larger to reduce the coding overhead. Further, the quantization step size of each coding unit can be adaptively adjusted according to the mapped quantization parameter adjustment factor. In video coding and decoding, a quantization parameter QP (Quantization Parameter) can be used to control the quantization step size. The larger the QP, the larger the quantization step size, the sparser the quantized coefficients, the lower the coding rate, but the greater the distortion. Traditional quantization methods usually set a fixed QP value at the frame level or slice level and dynamically adjust it according to the rate-distortion optimization process. In semantic adaptive quantization, the QP value can further consider semantic importance. Specifically, for semantic important regions, the quantization parameter adjustment factor can be subtracted from the original QP value to obtain a smaller QP value, so as to adopt a smaller quantization step size, which can retain more high-frequency details and improve the reconstruction quality. For semantic secondary regions, the quantization parameter adjustment factor can be added to the original QP value to obtain a larger QP value, so as to adopt a larger quantization step size, which can reduce the coding overhead and save the bit rate. In this way, through adaptive quantization step size adjustment, the coding quality and bit rate of different semantic regions can be better balanced. Thus, for semantic important regions (such as face, text), a smaller quantization step size is adopted to retain more high-frequency details; for semantic secondary regions (such as background), a larger quantization step size is adopted to reduce the coding overhead. In addition, in the semantic-level coding control of this implementation method, a more personalized and intelligent coding scheme can be customized according to the actual application requirements and the preferences of the target audience.
[0140] Optionally, the quantization parameter can also consider the visual sensitivity of the semantic region. For regions with high visual sensitivity, such as texture, edge, etc., a smaller quantization step size is adopted.
[0141] Optionally, during the process of adjusting quantization parameters, in combination with the rate-distortion optimization criterion, the quantization strategy is dynamically adjusted to ensure an optimal balance between coding efficiency and reconstruction quality.
[0142] In addition, there is a concept of Dead Zone in the quantization process of video coding and decoding. The Dead Zone refers to setting a relatively large quantization interval near zero to achieve more zero coefficients. The size of the Dead Zone is generally controlled by a parameter β. The larger β is, the larger the Dead Zone is, and the more zero coefficients after quantization. In traditional solutions, the setting of the Dead Zone is usually fixed, lacking flexibility, resulting in limited video compression. Based on this, in order to further improve the performance of adaptive quantization, the Dead Zone parameter can be optimized according to semantic sensitivity. Among them, the semantic sensitivity of each region in the current frame can be determined according to the semantic category of each region, and then the Dead Zone parameter is adjusted according to the semantic sensitivity of the coding unit. Optionally, the Dead Zone parameter of the coding unit is negatively correlated with the semantic sensitivity of the coding unit, and the semantic sensitivity of the coding unit is determined by the data of the coding unit in the semantic label map. That is, for semantically sensitive regions, such as textures and edges, the Dead Zone parameter β is reduced to retain more non-zero coefficients and improve the ability to represent details. For semantically insensitive regions, such as smooth regions, the Dead Zone parameter β is increased to generate more zero coefficients and reduce the coding overhead. Through semantically sensitive Dead Zone optimization, the human visual characteristics can be better considered during quantization, improving the subjective quality. At the same time, Dead Zone optimization and adaptive quantization step adjustment can work together to further enhance the performance of semantic adaptive quantization.
[0143] After quantization processing, entropy coding can also be performed on the quantized transform coefficients. Among them, the transform coefficients can be reordered and scanned; then entropy coding is performed on the reordered and scanned coefficients. Exemplarily, the quantized coefficient matrix can be rearranged to reduce redundant information; then the rearranged coefficient matrix is scanned according to a specific scanning order (such as Zig-Zag scanning) to generate a one-dimensional sequence; then entropy coding is performed on the scanned coefficient sequence.
[0144] In one implementation, the coding order can be optimized according to the information of the semantic label map, that is, the quantized coefficient matrix is rearranged according to the information of the semantic label map. Optionally, the important foreground object regions can be encoded first, and then the background regions can be encoded to ensure the overall coding efficiency.
[0145] In another implementation, it can be rearranged row-first or column-first.
[0146] In addition, reordering and scanning in the entropy encoding process can be used in combination. That is, in some cases, the reordering of transform coefficients can be achieved by changing the scanning order. For example, in an adaptive scanning order, the optimal scanning order can be dynamically selected according to the statistical characteristics of the actual data, so as to achieve the effect of optimizing the encoding order. In this way, the entropy encoding process can include: scanning the quantized transform coefficients; performing entropy encoding on the scanned coefficient sequence.
[0147] In the above embodiments, the coefficients can be scanned in a conventional scanning manner. Among them, the conventional scanning order is usually fixed, such as zigzag scanning or diagonal scanning.
[0148] More preferably, a semantic-driven coefficient scanning strategy can be further introduced. Among them, for different semantic regions, different scanning orders are adopted. That is, the quantized transform coefficients can be scanned using the scanning order corresponding to the texture of the coding unit. For example, for the texture region, a scanning order that can better capture the texture direction is adopted, such as direction-adaptive scanning; for the edge region, a scanning order that can better capture the edge structure is adopted, such as edge-aware scanning. This can make the transform coefficients more conform to the semantic characteristics after scanning, which is beneficial to subsequent entropy encoding.
[0149] Furthermore, in some cases, the scanned coefficients can be reordered again, and then entropy encoding is performed on the transform coefficients after reordering again, which can better utilize the statistical characteristics of the data, thereby improving the compression effect. Optionally, semantic-driven reordering can be performed on the coefficients after scanning. Among them, the coefficients can be sorted according to semantic importance, with the coefficients in the semantically important region placed in front and the coefficients in the semantically less important region placed behind. This can enable the important coefficients to be encoded first, improve the reconstruction quality, while the less important coefficients can be appropriately discarded to save the bit rate to ensure the overall encoding efficiency. Optionally, by introducing a priority parameter, the semantically important region can be given priority consideration during encoding to improve the overall encoding quality.
[0150] When performing entropy encoding on the quantized transform coefficients, CABAC (Context-Adaptive Binary Arithmetic Coding) can be used as the entropy encoding method to adaptively select the probability model and context to improve the encoding efficiency.
[0151] More preferably, when performing entropy coding on the quantized transform coefficients, semantic information can be used to adaptively adjust the coding model and / or context to better match the statistical characteristics of different semantic regions. Among them, corresponding probability models and contexts can be assigned to each semantic region according to the semantic category and semantic importance of each region to better match its statistical characteristics and improve the coding efficiency. That is, entropy coding can be performed on the quantized transform coefficients based on the probability model corresponding to the importance and / or texture of the coding unit. Optionally, for semantically important regions, more complex and refined probability models [such as Context-Adaptive Variable-Length Coding (CAVLC) or Context-Adaptive Binary Arithmetic Coding] can be used to capture their complex statistical characteristics; for semantically secondary regions, simpler and coarser probability models (such as Huffman coding or run-length coding) can be used to reduce the coding overhead. That is, the complexity of the probability model used by the region can be positively correlated with the semantic importance of the region. Among them, the semantic importance of the region can be determined by the semantic category of the region. In addition, the complexity of the probability model used by the region can be positively correlated with the texture complexity of the region. For example, for complex texture regions, more complex probability models are used; for smooth regions, simpler probability models are used.
[0152] Meanwhile, when performing entropy coding on a region, if context modeling is involved, for example, when using CAVLC or CABAC to perform entropy coding on a region, semantic information can also be considered during context modeling. Among them, for adjacent coefficients of the same semantic category, their contexts should be more similar and the context model can be shared; for coefficients of different semantic categories, their contexts should be distinguished and different context models can be used. In this way, through semantic-based adaptive entropy coding, the statistical characteristics of different semantic regions can be better adapted and the coding efficiency can be improved.
[0153] Among them, when performing entropy coding using some probability models (such as Huffman coding), which involves allocating codewords, the allocation of codewords can be guided by semantic information when using these probability models. Among them, codewords can be adaptively allocated to the region according to the information of the semantic label map. Optionally, codewords can be allocated to the region according to the semantic importance and / or texture complexity of the region. Among them, the allocated codewords can be positively correlated with the semantic importance and / or texture complexity of the corresponding region. For example, for semantically important regions (such as foreground objects), more codewords are allocated to ensure their coding accuracy; for semantically secondary regions (such as the background), fewer codewords are allocated to save the bit rate.
[0154] In addition, during the entropy coding process, the bitrate allocation strategy can be adaptively adjusted according to the information in the semantic label map to optimize the compression effect and bitrate. Specifically, an adaptive algorithm (such as a reinforcement learning algorithm) can be introduced to dynamically optimize the bitrate allocation model so that it can adapt to the changes in video content in real time. Among them, the bitrate allocation model can be used to determine the bit representation length or coding method of each symbol based on the occurrence probability of symbols or symbol sequences determined by a probability model to achieve the best compression effect. That is, the probability model provides the basic data for the bitrate allocation model, and the bitrate allocation model optimizes the encoding process based on this data.
[0155] During the entropy coding process, the encoding parameters can also be adaptively adjusted according to the information in the semantic label map. Specifically, by introducing semantic weight parameters, the encoding parameters can be dynamically adjusted to better meet the actual needs of different regions. Among them, the weight value of each region or object can be calculated based on the results of semantic analysis; then, the entropy coding parameters can be adjusted according to the weight value. For example, the probability model can be adjusted in arithmetic coding or the symbol frequency can be adjusted in Huffman coding. The weight value can be determined according to the following factors: the importance of the object, for example, key objects (such as faces, text) are given higher weights; the importance of the region, foreground regions are given higher weights, and background regions are given lower weights; the motion intensity, in video coding, regions with stronger motion are given higher weights.
[0156] Optionally, during the entropy coding process, enhanced coding of high-frequency information can also be performed according to the information in the semantic label map. Specifically, by introducing a high-frequency enhancement algorithm, the coding accuracy of high-frequency information can be improved to ensure that the high-frequency details after reconstruction are not lost. Among them, a semantic weight map can be generated based on the semantic label map, and each weight in the semantic weight map represents the importance of each pixel or block; according to the semantic weight map, the quantization step of high-frequency coefficients can be reduced or their bitrate allocation can be increased to perform enhanced coding of high-frequency information.
[0157] Using semantic information to guide the entropy coding process, the coding efficiency and coding quality can be further improved by optimizing the coding order and parameters.
[0158] Semantic-guided inter-frame prediction and motion estimation involve multiple steps, including reference frame selection, search range determination, matching cost calculation, and / or motion vector coding, etc. To obtain the best coding performance, these steps can be jointly optimized. Specifically, an iterative approach can be adopted to alternately optimize the parameters and decisions of each step. For example, in each iteration, first fix the reference frame and search range, and optimize the matching cost function and motion vector coding; then, based on the optimized motion vectors, update the reference frame selection and search range. Through multiple iterations, continuously refine and improve the results of inter-frame prediction and motion estimation. At the same time, during the optimization process, the semantic information can also be updated. Since motion compensation will change the content of the reconstructed frame, the semantic label map and semantic feature map of the reconstructed frame can be recalculated after each iteration. The updated semantic information can be used to guide the optimization process of the next iteration, forming a closed-loop feedback. Through joint optimization and update, the role of semantic information in inter-frame prediction and motion estimation can be fully utilized to improve the coding efficiency and video quality.
[0159] In addition, since quantization and entropy coding will change the content of the reconstructed image, the semantic label map and semantic weight map can be recalculated after each iteration. The updated semantic information can be used to guide the optimization process of the next round of iteration to form a closed-loop feedback. The process can include: initial coding, performing initial transformation, quantization, and entropy coding on the input data; reconstructed image generation, decoding the encoded data to generate a reconstructed image; semantic analysis, recalculating the semantic label map and semantic importance map based on the reconstructed image; feedback adjustment, adjusting the quantization step size, entropy coding parameters, etc. according to the new semantic information for a new round of coding optimization; iterative optimization, repeating the above steps until a predetermined termination condition is reached (such as reaching the maximum number of iterations or meeting the quality requirements).
[0160] In joint optimization, a rate-distortion optimization framework can be adopted to determine the optimal quantization and entropy coding parameters by minimizing the weighted sum of distortion and bit rate. During the optimization process, factors such as semantic information, visual sensitivity, and coding efficiency need to be comprehensively considered to balance the quality and bit rate of different semantic regions. At the same time, the semantic information also needs to be dynamically updated during the optimization process.
[0161] In the process of semantic-guided intra-frame prediction and transformation, an overall optimization strategy can be adopted. Specifically, by combining the information of the semantic label map and the feature map, dynamically adjust the prediction and transformation parameters to achieve overall optimization. By introducing a multi-level optimization strategy, adjust the parameters at different levels to ensure the optimality of the overall coding effect.
[0162] In addition, during the training process, the model can be regularly verified and evaluated. By introducing visual quality assessment metrics (such as PSNR, SSIM), the encoding effect of the model is evaluated to ensure the actual encoding performance of the model. Combining user feedback and actual application scenarios, the model is further optimized to improve the overall encoding effect and user experience.
[0163] Multiple coding modes can also be combined to evaluate the effects of different coding strategies, dynamically adjust the coding mode selection strategy, and improve the overall coding efficiency and reconstruction quality. Among them, an appropriate coding mode can be selected according to the data type and application scenario; the effects of different coding strategies can be evaluated, for example, the effects of coding strategies can be evaluated from multiple dimensions such as compression ratio, reconstruction quality, and computational complexity; the coding mode selection strategy can be dynamically adjusted, for example, the most suitable coding mode can be selected according to the content characteristics of different regions, and another example is to dynamically adjust the coding parameters and modes according to the actual needs of users; another example is to dynamically adjust the coding strategy according to changes in the external environment (such as network bandwidth, device performance, etc.).
[0164] In addition, during the video encoding and decoding process, image frame reconstruction may also be involved. For example, during the video encoding process, after a prediction, residual calculation, transformation, quantization, and entropy encoding are sequentially performed on an image frame to obtain the bitstream of the image frame, the reconstructed frame of the image frame can be reconstructed based on the bitstream of the image frame, and the reconstructed frame can be used as reference image data for subsequent frames.
[0165] During the reconstruction process of the image frame, the information of the semantic label map can also be combined to adopt different reconstruction strategies for different regions. Specifically, for high-importance regions, a high-precision reconstruction strategy is adopted to retain more details; for low-importance regions, a simplified reconstruction strategy is adopted to reduce the computational complexity. By introducing visual enhancement algorithms (such as image enhancement algorithms, super-resolution algorithms, etc.), important semantic regions are enhanced to improve the overall visual quality. Optionally, during the reconstruction process, compression artifact suppression can be performed by introducing the information of the semantic label map. Specifically, compression artifacts in the reconstructed frame can be suppressed by introducing artifact suppression algorithms (such as deblocking effect algorithms, de-ringing effect algorithms, etc.) to improve the overall visual effect.
[0166] After the reconstructed frame is obtained, the reconstructed frame can be adaptively filtered and optimized according to the information of the semantic label map. Specifically, for high-importance regions, a weaker filtering intensity is adopted to retain more details; for low-importance regions, a stronger filtering intensity is adopted to effectively remove noise and artifacts. Combining a temporal consistency filter, based on the information of the semantic label map, the pixel values between adjacent frames are smoothed to improve the temporal coherence of the video.
[0167] Among them, after the reconstructed frame is obtained, loop filtering can be performed on the reconstructed frame. Optionally, the loop filtering operation of the reconstructed frame of the image frame can be guided by combining with the semantic label map of the image frame to achieve semantic-aware quality enhancement.
[0168] Optionally, the loop filtering operation may include deblocking filtering and sample adaptive offset (SAO) filtering operations.
[0169] During the adaptive deblocking filtering process, the deblocking filtering strength and filtering template can be adjusted according to the semantic label map. For semantically important regions, weaker smoothing filtering is adopted to retain details; for semantically less important regions, stronger smoothing filtering is adopted to suppress noise. In deblocking filtering, the filtering strength and filtering template can be adaptively adjusted according to the semantic label map and semantic feature map. For semantically important regions, such as faces, weaker smoothing filtering is adopted to retain more details; for semantically less important regions, stronger smoothing filtering is adopted to suppress coding noise.
[0170] Compared with using a fixed filtering strength, in this embodiment, corresponding filtering strength parameters are set for different semantic regions according to the semantic importance map. Specifically, for regions with high semantic importance (such as faces, texts, etc.), weaker smoothing filtering is adopted to retain more detailed information; for regions with low semantic importance (such as backgrounds), stronger smoothing filtering is adopted to suppress coding artifacts and noise.
[0171] Specifically, according to the semantic importance map, the image is divided into multiple semantic regions, and corresponding filtering strength parameters (such as filtering coefficients, thresholds, etc.) are set for each semantic region. In this way, during the DBF process, according to the semantic region to which each pixel belongs, the corresponding filtering strength parameter is selected for adaptive filtering. In this way, by adaptively adjusting the DBF strength, while retaining the details of semantically important regions, the coding artifacts in semantically less important regions can be suppressed, and the overall visual quality can be improved. Through semantic adaptive filtering strength adjustment, a better balance between detail retention and artifact suppression can be achieved.
[0172] In addition, in deblocking filtering, traditional DBF usually uses fixed filtering templates (such as 4x4 or 8x8 square templates), which cannot be optimized for different semantic regions and may result in poor filtering effects. Compared with the prior art, the present application selects the most suitable filtering template according to the structural characteristics of different semantic regions. For example, for the face region, an elliptical or polygonal template adapted to the facial structure is used; for the text region, a directional template adapted to the character strokes is used; for the texture region, an anisotropic template adapted to the texture direction is used. Among them, the structural characteristics (such as edges, textures, etc.) of each semantic region are extracted; according to the structural characteristics, the most matching filtering template is selected from the predefined template library; during the DBF process, the corresponding filtering template is applied to different semantic regions for smoothing. Through semantic-driven filtering template selection, the present application can better adapt to the structural characteristics of different semantic regions, improve the effects of detail preservation and artifact suppression. Through semantic-driven filtering template selection, the filtering effect can be significantly improved, and the visual quality of the reconstructed image can be improved.
[0173] In sample adaptive offset (SAO) filtering, the filtering type and offset value can be adaptively selected according to semantic information.
[0174] Among them, the texture complexity characteristics (such as gradients, variances, etc.) of each semantic region can be extracted; according to the texture complexity characteristics, the dominant filtering mode of each semantic region is determined; during the SAO filtering process, the corresponding filtering mode is applied to different semantic regions for adaptive offset adjustment. In this way, through semantic-guided filtering mode selection, SAO filtering can better adapt to the image content characteristics of different semantic regions and improve the filtering effect. Through semantic-guided filtering mode selection, the image content can be more accurately matched, and the filtering effect and visual quality can be improved.
[0175] SAO filtering supports multiple filtering modes, such as edge offset (EO) mode and band offset (BO) mode, to adapt to different types of image content. In semantic-aware loop filtering, different filtering modes can be adopted according to different semantic regions. Specifically, for semantic regions containing rich edges and textures (such as buildings, vegetation, etc.), the EO mode or edge-preserving filtering is preferably selected to better preserve edge and texture details; for semantic regions containing smooth areas (such as the sky, roads, etc.), the BO mode or brightness compensation filtering is preferably selected to better eliminate noise and artifacts, or to reduce brightness distortion.
[0176] Balance detail retention and noise suppression through adaptive offset adjustment. The filtering type and offset value can be adaptively selected according to semantic information. That is, for different semantic regions, different offset adjustment strategies are adopted. Among them, the offset value of SAO filtering can also be adjusted according to semantic importance. For semantically important regions (such as faces, texts, etc.), a smaller offset adjustment step size is adopted to avoid detail loss caused by excessive smoothing. For semantically secondary regions (such as backgrounds), a larger offset adjustment step size is adopted to better suppress noise and artifacts. Through semantic-aware loop filtering, the subjective visual quality of the reconstructed image can be effectively improved, making it more in line with the perceptual characteristics of the human eye. Implementation of offset adjustment: Divide the reconstructed image into multiple semantic regions according to the semantic importance map; Set corresponding offset adjustment parameters (such as adjustment step size, threshold, etc.) for each semantic region; During the SAO filtering process, select the corresponding offset adjustment parameters for adaptive filtering according to the semantic region to which each sample belongs. Adaptive offset adjustment can better balance detail retention in semantically important regions and noise suppression in semantically secondary regions, improving the subjective visual quality of the reconstructed image. SAO filtering with adaptive offset adjustment can effectively reduce compression artifacts and noise and improve the quality of the reconstructed image.
[0177] In the sample adaptive offset (SAO) filter of HEVC, the filtering type and intensity can be adaptively adjusted according to the semantic label map. By introducing the semantic label map, the filtering type and intensity can be adaptively adjusted according to the semantic information of different regions, so as to remove artifacts while retaining details. The specific process is as follows:
[0178] (1) Filter parameter adjustment:
[0179] ① Semantic information-guided filtering: Adjust the parameters of the SAO filter adaptively according to the information of the semantic label map. For semantically important regions (such as foreground objects), edge-preserving filtering is used to retain more details; for semantically secondary regions (such as backgrounds), smoothing filtering is used to effectively remove noise and artifacts.
[0180] ② Region-adaptive filtering: During the filtering process, different filtering types and intensities are adaptively selected according to different regions of the semantic label map. For example, for the foreground object region, a lower filtering intensity can be selected to retain details; for the background region, a higher filtering intensity can be selected to remove noise.
[0181] (2) Filtering strategy optimization:
[0182] ①Semantics-based filtering strategy: In the SAO filtering process, a semantics-based filtering strategy is introduced. Specifically, according to the information of the semantic label map, the most suitable filtering strategy for the current region can be selected. For example, for complex foreground objects, a multi-pass filtering strategy can be chosen to retain more detailed information; for simple background regions, a single-pass filtering strategy can be selected to remove noise and artifacts.
[0183] ②Filtering order optimization: During the filtering process, the filtering order is optimized according to the information of the semantic label map. Specifically, the important foreground object regions can be filtered first, and then the background regions can be filtered to ensure the overall visual effect.
[0184] (3) Filtering result evaluation:
[0185] ①Visual quality evaluation: During the filtering process, the filtering result is evaluated according to the information of the semantic label map. Specifically, visual quality evaluation metrics (such as PSNR, SSIM) can be introduced to evaluate the reconstructed frames after filtering to ensure that the filtering effect meets the expectations.
[0186] ②Semantic information verification: During the filtering process, the reconstructed frames after filtering can be verified according to the information of the semantic label map. Specifically, by comparing the semantic information consistency before and after filtering, it can be verified whether the filtering result retains sufficient semantic information.
[0187] Among them, the semantic-guided encoding of video frames using the semantic label map can also be applied to intelligent encoding. Among them, the implementation process of an intelligent semantic HEVC encoding method includes the following steps:
[0188] 1. Generation of semantic feature maps: Before performing joint semantic encoding, semantic feature maps can be generated. This step can be based on the results of semantic segmentation and contrast learning.
[0189] (1) Generation of semantic label maps: For each video frame, a pixel-level semantic label map is obtained using a semantic segmentation model.
[0190] (2) Generation of semantic feature maps: The label map is multiplied element-wise with the original video frame to obtain a semantic feature map. Each position in the semantic feature map represents the semantic category to which the corresponding pixel belongs, and different semantic categories are represented by different grayscale values or colors.
[0191] (3) Enhancement of semantic feature maps: The semantic feature maps are enhanced using the semantic embedding vectors obtained from contrast learning. Specifically, the semantic embedding vectors are fused with the semantic feature maps in the spatial dimension to obtain enhanced semantic feature maps. This fusion can be achieved through simple concatenation operations or by using attention mechanisms to adaptively adjust the weights at different positions.
[0192] 2. Extraction of Content Features: Joint semantic coding can also extract traditional content features.
[0193] (1) Residual Feature Extraction: Calculate the difference information (residual) between the current frame and the reference frame, and then transform and quantize the residual to obtain the compressed residual coefficients. Residual features can effectively capture the temporal redundancy between video frames and reduce the coding bitrate.
[0194] (2) Texture Feature Extraction: Use a convolutional neural network to extract the texture information in the video frame. By introducing a texture feature extraction module in the encoder, a more compact and discriminative content representation can be learned.
[0195] 3. Feature Fusion and Coding: After obtaining the semantic feature map and content features, they can be fused and jointly encoded.
[0196] (1) Feature Concatenation: Concatenate the semantic feature map and content features in the channel dimension to form a joint feature map.
[0197] (2) Attention Mechanism: Introduce an attention mechanism to adaptively adjust the weights of the semantic feature map and content features at different positions. The attention weights can be calculated based on the statistical information of the feature map (such as mean, variance) or context information (such as features in the surrounding area).
[0198] (3) Joint Coding: On the fused joint feature map, use a traditional video encoder (such as H.264, H.265) or a deep learning-based encoder (such as variational autoencoder, generative adversarial network) to achieve compression coding. The encoder maps the joint feature map to a low-dimensional latent representation, and performs quantization and entropy coding to obtain the compressed bitstream.
[0199] 4. Rate-Distortion Optimization: An important goal of joint semantic coding is to minimize the reconstructed distortion after compression under a given bitrate constraint.
[0200] (1) Rate-Distortion Loss: Consists of rate loss and distortion loss. Rate loss measures the length of the encoded bitstream and encourages the encoder to generate a more compact representation. Distortion loss measures the difference between the reconstructed frame and the original frame and encourages the encoder to retain more video content information.
[0201] (2) Lagrange Multiplier: Used to balance the rate loss and distortion loss in the rate-distortion loss to adapt to different rate-distortion preferences.
[0202] (3) Gradient Descent Optimization: Use the gradient descent algorithm to minimize the rate-distortion loss while updating the parameters of the encoder and decoder.
[0203] (4) Multi-scale coding: Generate multiple bitstream versions at different bitrates. During decoding, select an appropriate version for reconstruction according to the bandwidth or quality requirements. This adaptive coding strategy can flexibly adapt to different transmission and storage scenarios and provide a better user experience.
[0204] 5. Reconstruction and post-processing: At the receiver, the decoder recovers the joint feature map from the compressed bitstream and divides it into a semantic feature map and content features.
[0205] (1) Symmetric decoding: Use a decoding module symmetric to the encoder to map the semantic feature map and content features back to the pixel domain respectively to obtain the reconstructed video frame.
[0206] (2) Intra-frame filtering based on semantic labels: According to the semantic label map of the reconstructed frame, adopt different filtering strategies for different semantic regions. For example, for smooth regions (such as the background), use a denoising filter with a stronger intensity, while for texture regions (such as foreground objects), use a sharpening filter with a weaker intensity.
[0207] (3) Inter-frame filtering based on temporal consistency: By analyzing the semantic correspondence between adjacent reconstructed frames, perform smooth transitions on temporally discontinuous regions, reduce the flicker and jitter introduced by compression, and improve the temporal coherence of the video.
[0208] In summary, joint semantic coding is a key step in video compression optimization. By fusing the semantic feature map and content features, it considers both semantic information and detail information during the encoding stage, improving the reconstruction quality after compression. Reasonable feature extraction, fusion strategies, and rate-distortion optimization can significantly improve the efficiency and performance of video compression. At the same time, post-processing techniques after reconstruction, such as intra-frame filtering and inter-frame filtering, can further improve the visual quality of the compressed video and bring a better viewing experience to users. Joint semantic coding is compatible with traditional video coding frameworks and can be easily integrated into existing codecs, having broad application prospects.
[0209] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of an embodiment of the electronic device of the present application. This electronic device 20 includes a processor 22, and the processor 22 is used to execute instructions to implement the above method. For the specific implementation process, please refer to the description of the above embodiment, which will not be elaborated here.
[0210] The processor 22 may also be referred to as a CPU (Central Processing Unit). The processor 22 may be an integrated circuit chip with the ability to process signals. The processor 22 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 22 may also be any conventional processor, etc.
[0211] The electronic device 20 may further include a memory 21 for storing instructions and data required for the operation of the processor 22.
[0212] The processor 22 is used to execute instructions to implement the method provided by any embodiment of the method of the present application and any non-conflicting combination.
[0213] Among them, the electronic device of the present application may be a confusion control system.
[0214] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present application. The computer-readable storage medium 30 of the embodiment of the present application stores instruction / program data 31, and when the instruction / program data 31 is executed, it implements the method provided by any embodiment of the video encoding method of the present application and any non-conflicting combination. In one embodiment, the instruction / program data 31 may form a program file and be stored in the storage medium 30 in the form of a software product, so that a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium 30 includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.
[0215] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of components is only a logical function division. In actual implementation, there may be other division methods. For example, multiple components or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or components, and can be in electrical, mechanical, or other forms.
[0216] In addition, in each embodiment of the present application, the various functional components can be integrated in a processing component, or each component can exist physically alone, or two or more components can be integrated in one component. The above-mentioned integrated components can be implemented in the form of hardware or in the form of software functional components.
[0217] It can also be stated that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity, or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the element.
[0218] The above is only the implementation manner of the present application, and does not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A video encoding method, characterized in that, The method includes: Performing semantic segmentation processing on a current frame in a video to obtain a semantic label map of the current frame; Determining prediction information of coding units in the current frame based on the semantic label map; Encoding the current frame based on the prediction information.
2. The video encoding method according to claim 1, wherein The determining the prediction information of coding units in the current frame based on the semantic label map includes: If a coding unit uses an intra prediction mode to generate a prediction block, determining the intra prediction mode of the coding unit based on the semantic label map; or, If a coding unit uses an inter prediction mode to generate a prediction block, determining a reference frame and / or a motion vector of the coding unit based on the semantic label map.
3. The video encoding method according to claim 2, wherein The determining the intra prediction mode of the coding unit based on the semantic label map includes: Taking the preferred intra prediction mode corresponding to the semantic category of the coding unit as the candidate intra prediction mode of the coding unit; Determining the final intra prediction mode of the coding unit from the candidate intra prediction modes of the coding unit.
4. The video encoding method according to claim 3, wherein The taking the preferred intra prediction mode corresponding to the semantic category of the coding unit as the candidate intra prediction mode of the coding unit includes: Taking the preferred intra prediction mode corresponding to the texture property of the coding unit as the candidate intra prediction mode of the coding unit, where the texture property of the coding unit is determined by data of the coding unit in the semantic label map; or, Taking the preferred intra prediction mode corresponding to the importance of the coding unit as the candidate intra prediction mode of the coding unit, where the importance of the coding unit is determined by data of the coding unit in the semantic label map.
5. The video encoding method according to claim 3, wherein The taking the preferred intra prediction mode corresponding to the semantic category of the coding unit as the candidate intra prediction mode of the coding unit includes: Taking the preferred intra prediction mode corresponding to the importance of the coding unit as the candidate intra prediction mode of the coding unit, where the importance of the coding unit is determined by the semantic category of the coding unit; The number of preferred intra prediction modes corresponding to the coding unit is positively correlated with the importance of the coding unit.
6. The video encoding method according to claim 2, wherein The if a coding unit uses an inter prediction mode to generate a prediction block, determining a reference frame and / or a motion vector of the coding unit based on the semantic label map includes: Based on the semantic label maps of the current frame and candidate reference frames of the current frame, determining an image frame containing a preset region from the candidate reference frames, where the preset region is a region whose similarity with the coding unit meets a first preset condition, or the preset region is a region whose similarity with the coding unit meets a second preset condition, and the similarity between a region in the reference frame and the coding unit is obtained by weighting the semantic similarity degree and pixel similarity between a region in the reference frame and the coding unit; Determining the reference frame of the current frame based on the image frame containing the preset region.
7. The video encoding method according to claim 6, wherein The determining an image frame containing a preset region from the candidate reference frames based on the semantic label maps of the current frame and candidate reference frames of the current frame includes: Calculate the semantic intersection over union or semantic cosine similarity between the semantic label map data of the coding unit and the semantic label map data of a region in the candidate reference frame, to obtain the semantic similarity degree between the coding unit and a region in the reference frame.
8. The video encoding method according to claim 2, wherein If a coding unit uses an inter prediction mode to generate prediction blocks, determining the reference frame and / or motion vector of the coding unit based on the semantic label map includes: Determining the candidate motion vectors of the coding unit based on the coded blocks in the current frame or the reference frame that have a semantic similarity degree with the coding unit satisfying a third preset condition; determining the best motion vector of the coding unit based on the candidate motion vectors of the coding unit; or, Determining the search range of the coded region of the coding unit in the reference frame based on the semantic category of the coding unit; searching for the matching block of the coding unit within the determined search range; calculating the motion vector of the current coding unit based on the positions of the coding unit and the matching block.
9. The video encoding method according to claim 8, wherein Determining the search range of the coded region of the coding unit in the reference frame based on the semantic category of the coding unit includes: The size of the search range of the coding unit is determined based on the semantic category of the coding unit; or, If there are coded blocks among the coding units belonging to the same semantic region as the current coding unit, setting the search range of the current coding unit as the region around the position pointed to by the motion vector of the coded blocks belonging to the same semantic region as the current coding unit; or, Setting the search range of the current coding unit as the region in the reference frame that has a semantic similarity degree with the current coding unit satisfying a fourth preset condition.
10. The video encoding method according to claim 1, wherein Determining the prediction information of the coding unit in the current frame based on the semantic label map includes: Determining the candidate prediction information of the coding unit, where the prediction information includes a motion vector and / or a prediction mode; Combining the semantic label map, calculating the semantic cost of the candidate prediction information of the coding unit; Based on the semantic cost, determining the best prediction information of the coding unit from the candidate prediction information.
11. The video encoding method according to claim 10, wherein The method further includes: Calculating the pixel cost of the candidate prediction information of the coding unit; Based on the total cost, determining the best prediction information of the coding unit from the candidate prediction information, where the total cost is obtained by weighting the semantic cost and the pixel cost; The weight of the semantic cost of the coding unit is positively correlated with the semantic importance of the coding unit, and the semantic importance of the coding unit is determined based on the data of the coding unit in the semantic label map.
12. The video encoding method according to claim 1, wherein Encoding the current frame based on the prediction information includes: Determining the residual data of the coding unit based on the prediction information; Performing a transformation process on the residual data of the coding unit based on the semantic label map.
13. The video encoding method according to claim 12, wherein Performing a transformation process on the residual data of the coding unit based on the semantic label map includes: Determine the transform mode and / or transform kernel of the coding unit based on the texture and / or importance of the coding unit, where the texture and / or importance of the coding unit is determined by the data of the coding unit in the semantic label map; or, Determine the transform parameters of the current coding unit based on the coded blocks in the peripheral area of the coding unit whose semantic similarity with the coding unit satisfies a fifth preset condition.
14. The video encoding method according to claim 1, wherein The encoding of the current frame based on the prediction information includes: Determine the residual data of the coding unit based on the prediction information; Perform a transform process on the residual data of the coding unit to obtain transform coefficients; Perform a quantization process on the transform coefficients; The quantization step size of the coding unit is negatively correlated with the importance and / or visual sensitivity of the coding unit, where the importance and / or visual sensitivity of the coding unit is determined by the data of the coding unit in the semantic label map.
15. The video encoding method according to claim 14, characterized in that, The dead zone parameter of the coding unit is negatively correlated with the semantic sensitivity of the coding unit, where the semantic sensitivity of the coding unit is determined by the data of the coding unit in the semantic label map.
16. The video encoding method according to claim 1, characterized in that, The encoding of the current frame based on the prediction information includes: Determine the residual data of the coding unit based on the prediction information; Perform a transform process on the residual data of the coding unit to obtain transform coefficients; Perform a quantization process on the transform coefficients to obtain quantized transform coefficients; Perform entropy coding on the quantized transform coefficients based on the semantic label map.
17. The video encoding method according to claim 16, wherein The entropy coding of the quantized transform coefficients based on the semantic label map includes: Scan the quantized transform coefficients in the scanning order corresponding to the texture of the coding unit, where the texture of the coding unit is determined by the data of the coding unit in the semantic label map; Perform entropy coding on the scanned coefficient sequence.
18. The video encoding method according to claim 16, wherein The entropy coding of the quantized transform coefficients based on the semantic label map includes: Reorder the quantized transform coefficients in descending order of semantic importance; Perform entropy coding on the reordered coefficient sequence.
19. The video encoding method according to claim 16, wherein The entropy coding of the quantized transform coefficients based on the semantic label map includes: Perform entropy coding on the quantized transform coefficients based on the probability model corresponding to the importance and / or texture of the coding unit, where the importance and texture of the coding unit are determined by the data of the coding unit in the semantic label map; or, For several adjacent quantized transform coefficients of the same semantic class, perform entropy coding on the several adjacent quantized transform coefficients based on a shared context model; or, Allocate codewords for a region according to the importance and / or texture of the region, where the codewords allocated for the region are positively correlated with the importance and / or texture of the region, and the importance and texture of the region are determined by the data of the region in the semantic label map; or, Adaptive adjust the code rate allocation strategy and / or encoding parameters according to the information of the semantic label map; or, Perform enhanced coding on the high-frequency information in the quantized transform coefficients according to the information of the semantic label map.
20. The video encoding method according to claim 1, wherein The method further includes: A reconstructed frame of the current frame is reconstructed based on the encoding result of the current frame, wherein the reconstruction strategy used for each region in the current frame is determined based on the data of the semantic label map of each region.
21. The video encoding method according to claim 1, wherein The method further includes: Reconstructing a reconstructed frame of the current frame based on the encoding result of the current frame; Performing adaptive filtering on the reconstructed frame based on the semantic label map.
22. The video encoding method according to claim 21, wherein The performing adaptive filtering on the reconstructed frame based on the semantic label map includes: Determining the filtering strength and / or filtering template of deblocking filtering for each region in the current frame based on the semantic label map; and / or, Adaptive selection of the filtering type and / or offset value of sample adaptive offset filtering for each region in the current frame based on the semantic label map.
23. The video encoding method according to claim 22, wherein The determining the filtering strength and / or filtering template of deblocking filtering for each region in the current frame based on the semantic label map includes: Determining the filtering strength of each region based on the importance of each region, the filtering strength of the region is negatively correlated with the importance of the region, and the importance of the region is determined by the data of the region in the semantic label map; and / or, Determining the filtering template of each region based on the semantic structure feature of each region, and the semantic structure feature of the region is determined by the data of the region in the semantic label map.
24. The video encoding method according to claim 22, wherein The adaptive selection of the filtering type and / or offset value of sample adaptive offset filtering for each region in the current frame based on the semantic label map includes: Determining the filtering mode of each region based on the texture of each region, and the texture of the region is determined by the data of the region in the semantic label map; and / or, Determining the offset value during sample adaptive offset filtering for each region based on the importance of each region, the offset value of the region is negatively correlated with the importance of the region, and the importance of the region is determined by the data of the region in the semantic label map.
25. The video encoding method according to claim 1, wherein Before the determining the prediction information of the coding unit in the current frame based on the semantic label map, it includes: Dividing the current frame into multiple coding units with the same size; Determining the semantic importance of each coding unit based on the semantic label map; Performing merging or splitting processing on the coding units based on the semantic importance of the coding units to adjust the size of the coding units, and the size of the adjusted coding unit is negatively correlated with its semantic importance; The determining the prediction information of the coding unit in the current frame based on the semantic label map includes: Performing prediction on the adjusted coding unit based on the semantic label map to determine the prediction information of the adjusted coding unit.
26. An electronic device, characterized in that, The electronic device includes a processor; the processor is configured to execute instructions to implement the steps of the method according to any one of claims 1-25.
27. A computer-readable storage medium having a program and / or instructions stored thereon, characterized in that, When the program and / or instructions are executed, the steps of the method according to any one of claims 1-25 are implemented.
Citation Information
Cited By
Video stream processing method and device and video stream generation method
CN122137953A