Video coding method and device, electronic equipment and storage medium

By reusing motion compensation time domain filtering motion vectors and feature point topological relationships to optimize ROI encoding, the problems of insufficient real-time performance and high resource consumption in existing technologies are solved, and high-definition and low-bitrate real-time live broadcast effects are achieved.

CN120676151APending Publication Date: 2025-09-19BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511102517.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing ROI encoding technology has the disadvantages of insufficient real-time performance, high resource consumption and rough bit rate allocation in live streaming, and cannot effectively meet the high efficiency and high quality requirements in live streaming scenarios.

Method used

By reusing the motion vectors in the motion compensation time domain filtering process, the ROI position prediction value is determined, and the position prediction is corrected based on the spatial topological relationship of the feature points. The ROI encoding process is optimized by combining differentiated quantization parameters and dynamic bit rate allocation.

Benefits of technology

It improves the real-time performance and quality of encoding, reduces latency, enhances bandwidth adaptability, reduces the risk of live broadcast freezes, and realizes high-definition, low-bitrate real-time live streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676151A_ABST
    Figure CN120676151A_ABST
Patent Text Reader

Abstract

The invention provides a video coding method and device, electronic equipment and a storage medium, relates to the technical field of image processing, in particular to the field of video coding and the like, and can be applied to application scenes such as video live broadcast and the like. The specific implementation scheme is as follows: acquiring a reference frame which is adjacent to a current frame and has established initial ROI hierarchical division; multiplexing a motion vector generated by the reference frame in a motion compensation time domain filtering process, and determining an ROI position predicted value of the current frame; extracting feature points of the five sense organs of the current frame, and correcting an ROI position predicted value based on a feature point spatial topological relation in the initial ROI; respectively configuring differential quantization parameters for the corrected main face region, the corrected secondary face region and the corrected background region; dynamically allocating a three-layer area code rate according to a real-time network bandwidth; and outputting the coded frame of the current frame and the associated ROI level metadata. According to the scheme, the coding efficiency and quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, in particular to the field of video coding, and can be used in application scenarios such as live video broadcasting, and specifically relates to video coding methods, devices, electronic devices and storage media. Background Art

[0002] In the field of live streaming, to enhance the viewing experience, it's common to optimize the encoding of regions of interest (ROIs) within live video streams to ensure high-quality presentation of key content within a limited bitrate. However, existing ROI encoding technologies suffer from issues such as insufficient real-time performance, high resource consumption, and extensive bitrate allocation, making them unable to effectively meet the high efficiency and quality requirements of live streaming scenarios. Summary of the Invention

[0003] The present disclosure provides a video encoding method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of the present disclosure, a video encoding method is provided, comprising: obtaining a reference frame adjacent to a current frame and for which an initial ROI hierarchical division has been established; reusing a motion vector generated by the reference frame during motion-compensated temporal filtering to determine an ROI position prediction value for the current frame; extracting facial feature points of the current frame, and correcting the ROI position prediction value based on the spatial topological relationship of the feature points in the initial ROI; configuring differentiated quantization parameters for the corrected primary facial region, secondary facial region, and background region respectively; dynamically allocating three-layer regional bit rates based on real-time network bandwidth; and outputting a coded frame of the current frame and associated ROI-level metadata; wherein the ROI-level metadata comprises: boundary coordinates of the primary facial region, secondary facial region, and background region; quantization parameter offset values ​​corresponding to each region; and bit rate allocation weight values ​​for each region.

[0005] According to a second aspect of the present disclosure, a video encoding device is provided, comprising: a reference frame acquisition module for acquiring a reference frame adjacent to a current frame and for which an initial ROI hierarchical division has been established; a position prediction module for reusing a motion vector generated by the reference frame during motion compensation temporal filtering to determine an ROI position prediction value for the current frame; a position correction module for extracting facial feature points of the current frame and correcting the ROI position prediction value based on the spatial topological relationship of the feature points in the initial ROI; a parameter configuration module for respectively configuring differentiated quantization parameters for the corrected primary facial region, secondary facial region, and background region; a bit rate allocation module for dynamically allocating bit rates for three layers of regions based on real-time network bandwidth; and a data output module for outputting a coded frame of the current frame and associated ROI-level metadata; wherein the ROI-level metadata includes: boundary coordinates of the primary facial region, secondary facial region, and background region; quantization parameter offset values ​​corresponding to each region; and bit rate allocation weight values ​​for each region.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0007] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.

[0009] The disclosed solution eliminates the overhead of independent motion estimation by reusing the motion vectors generated by the motion compensation time-domain filtering process; implements cross-frame tracking by predicting the ROI position based on the motion vector; extracts feature points in the first frame and establishes a spatial topological relationship as a correction benchmark; and uses the topological relationship to correct the predicted ROI position in subsequent frames to ensure positioning accuracy; and improves coding efficiency and quality by combining differentiated quantization parameters and dynamically adapting to the network environment.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0012] Figure 1 is a flowchart of a video encoding method according to an embodiment of the present disclosure;

[0013] Figure 2 is another flowchart of a video encoding method according to an embodiment of the present disclosure;

[0014] Figure 3 is a structural diagram of a video encoding device according to an embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of a scenario of a video encoding method according to an embodiment of the present disclosure;

[0016] Figure 5 3 is a structural diagram of an electronic device used to implement the video encoding method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0019] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0020] Before introducing the technical solutions of the embodiments of the present disclosure, the following technical terms that may be used in the present disclosure are further explained:

[0021] Video encoding refers to the process of compressing and encoding video data. The goal is to reduce the size of video data while minimizing video quality. This technology achieves efficient storage and transmission by removing redundant information, such as temporal, spatial, and statistical redundancy.

[0022] In online live streaming technology, in order to enhance the audience's viewing experience, ROI encoding optimization technology is usually used to focus on encoding the host's face or key operation areas. However, existing technologies have significant flaws: insufficient real-time performance, complex deep learning algorithm detection takes a long time, resulting in increased encoding delays and live broadcast freezes; high resource consumption, mobile devices are difficult to support, and need to rely on cloud servers, further increasing network transmission delays; coarse bit rate distribution, simply distinguishing between ROI and background, without fine-tuning key parts within the ROI (such as facial features). In addition, the frame extraction detection strategy ignores the dynamic characteristics of the scene, and is prone to coding resource mismatch problems when switching scenes, resulting in bit rate waste and reduced quality in key areas.

[0023] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, the present disclosure proposes a video encoding method, which eliminates redundant calculations by reusing motion compensation time-domain filtering motion vectors, significantly reducing the processing delay of the motion tracking module; avoiding independent motion estimation operations, achieving nearly zero-overhead ROI cross-frame tracking, and thus improving the real-time performance of encoding. Through the cross-frame multiplexing mechanism of the first-frame topological relationship, the spatial consistency of feature points is maintained; the quality fluctuations of the feature points reconstructed frame by frame are eliminated, and the clarity of details in key areas is guaranteed, thereby optimizing the encoding quality; combining layered coding with dynamic bit rate allocation, intelligently balancing quality and smoothness when bandwidth fluctuates, and enhancing bandwidth adaptability; through the coordinated optimization of motion tracking and bit rate control, the risk of live broadcast freezes is reduced. Under the premise of ensuring the visual quality of the main facial area, the overall bit rate efficiency is optimized, and a low-latency high-definition encoding framework of "tracking-layering-flow control" is established.

[0024] The present disclosure provides a video encoding method. Figure 1This is a flow chart of a video encoding method according to an embodiment of the present disclosure, which can be applied to a video encoding device. The video encoding device is located in an electronic device. The electronic device includes but is not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, and servers can be cloud servers or ordinary servers. For example, mobile devices include but are not limited to video live broadcast devices, and video live broadcast devices can be mobile phones, tablet computers, vehicle-mounted terminals, etc. In some possible implementations, the video encoding method can also be implemented by a processor calling computer-readable instructions stored in a memory. Figure 1 As shown, the video encoding method includes:

[0025] S101, obtaining a reference frame adjacent to the current frame and for which initial ROI hierarchical division has been established;

[0026] S102, multiplexing the motion vector generated by the reference frame in the motion compensation temporal filtering process to determine the ROI position prediction value of the current frame;

[0027] S103, extracting facial feature points of the current frame, and correcting the ROI position prediction value based on the spatial topological relationship of the feature points in the initial ROI;

[0028] S104, configuring differential quantization parameters for the corrected primary facial region, secondary facial region, and background region respectively;

[0029] S105, dynamically allocating the three-layer regional bit rate according to the real-time network bandwidth;

[0030] S106 : Output the coded frame of the current frame and the associated ROI-level metadata.

[0031] Here, the current frame refers to the frame in the video sequence that currently needs to be encoded. In the embodiment of the present disclosure, the video to be encoded is composed of a series of continuous image frames. In particular, when processing the current frame, information of the previous and next frames can be referenced.

[0032] Here, ROI refers to a region or object of particular interest in image processing and computer vision. In the disclosed embodiment, the ROI can be an area in a video frame that requires focused encoding. For example, the ROI can be a face area.

[0033] In an embodiment of the present disclosure, a reference frame adjacent to the current frame can be extracted from an acquired video sequence. Exemplarily, the reference frame can be the chronologically previous frame or the chronologically subsequent frame of the current frame; at the same time, the reference frame can also be the chronologically N-th frame before the current frame, where N≥1, and N is a positive integer; similarly, the reference frame can also be the chronologically M-th frame after the current frame, where M≥1, and M is a positive integer; similarly, the reference frame can also be a key frame selected through inter-frame prediction; in particular, the reference frame can also be determined by other methods in the prior art, which is not limited by the present disclosure. Furthermore, based on the determined reference frame, the levels of the main facial area, secondary facial area, and background area that have been divided can be obtained, and then the boundary information of the above areas can be recorded. The above is only an exemplary explanation and is not intended to limit all possible situations for obtaining reference frames, but it is not exhaustive here.

[0034] Motion compensated temporal filtering (MCTF) is a method that uses temporal information to predict and compensate for inter-frame errors. In this disclosed embodiment, MCTF analyzes the motion trends of pixels between previous and next frames to predict the pixel values ​​of the current frame, thereby reducing temporal redundancy in the video.

[0035] Here, a motion vector (MV) can describe the motion trajectory of a pixel block between previous and next frames. In the disclosed embodiment, the MV can be used to indicate the displacement of an object from a reference frame to a current frame.

[0036] Here, the ROI position prediction value refers to the estimation of the specific position of the ROI in the current frame based on the MV.

[0037] In the disclosed embodiment, the motion vector generated by the MCTF process in the reference frame can be first extracted. Then, using the MV, the ROI position in the reference frame is mapped to the current frame. The predicted ROI position is determined through operations such as translation and scaling. The above is only an example and does not limit all possible situations for determining the predicted ROI position value of the current frame. This is just not an exhaustive list.

[0038] Here, facial feature points refer to key points in the face region, such as the outline points of the eyes, nose, and mouth. In the disclosed embodiment, facial feature points can be automatically extracted by a face detection algorithm.

[0039] Here, the spatial topological relationship of the feature points can describe the positional relationship between the feature points. In the embodiment of the present disclosure, the spatial topological relationship of the feature points can describe the relative positions between the feature points of the facial features (such as the relative positions of the eyes and nose).

[0040] In the disclosed embodiments, a deep learning model or a specific face detector can be used to first extract the facial feature points of the current frame. Subsequently, the predicted ROI value can be adjusted based on the locations of the feature points and the position of the ROI in the reference frame. The above is merely an example and does not limit all possible scenarios for modifying the predicted ROI position. This is not intended to be an exhaustive list.

[0041] Here, the main facial area refers to the area within the ROI that contains the key feature points among the facial features. For example, the key feature points may include the eyes, nose, and mouth. In particular, the key feature points can be divided according to actual circumstances, and this disclosure does not limit this.

[0042] Here, the secondary facial region refers to the area within the ROI, excluding the primary facial region, that contains edge feature points among the facial features. For example, edge feature points may include cheeks and chin. In particular, edge feature points can be divided based on actual circumstances, and this disclosure does not limit this.

[0043] Here, the background region refers to other regions in the ROI except the primary face region and the secondary face region.

[0044] Here, the quantization parameter refers to a parameter used to control encoding quality. In the embodiment of the present disclosure, the smaller the quantization parameter, the higher the bit rate, the higher the quality, and the larger the file size.

[0045] In the disclosed embodiment, different quantization parameters can be assigned to the primary facial region, secondary facial region, and background based on the ROI segmentation results, thereby reducing resource usage in the background region. Specifically, the primary facial region can be assigned the smallest quantization parameter, the secondary facial region can be assigned a medium quantization parameter, and the background region can be assigned the largest quantization parameter. The above is merely an example and does not limit all possible scenarios for configuring differentiated quantization parameters. This is simply not an exhaustive list.

[0046] For example, the spatial topological relationship of feature points can be established simultaneously when dividing the three-layer area. Specifically, the process of establishing the spatial topological relationship of feature points can first establish a local rectangular coordinate system with the upper left corner vertex of the initial ROI area as the spatial origin; then, the absolute image coordinates of each feature point can be converted into normalized coordinates relative to the initial ROI; finally, a key point spatial topological structure record can be generated based on the normalized coordinate values, wherein the topological structure record contains the normalized coordinate values ​​of each feature point and a description of the relative position relationship between any two feature points.

[0047] For example, the process of converting the absolute image coordinates of each feature point into normalized coordinates relative to the initial ROI can be expressed by the following formula:

[0048]

[0049]

[0050] Where x Normalization Represents the horizontal normalized coordinate; y Normalization Represents the vertical normalized coordinate; x j Indicates the horizontal coordinate of the jth feature point; y j Indicates the ordinate of the jth feature point; L ROI,L Indicates the left boundary of ROI; L ROI,U Indicates the upper boundary of ROI; W ROI Indicates ROI width; H ROI Indicates the ROI height.

[0051] Here, network bandwidth refers to the maximum amount of data that a communication channel or network can transmit within a specific time period, and can be measured in bits per second. In the embodiments of the present disclosure, network bandwidth can represent the network's ability to transmit data. The greater the bandwidth, the more data can be transmitted per unit time, and the higher the network's transmission efficiency.

[0052] Here, bitrate refers to the amount of data transmitted per unit time, typically measured in kilobits per second. In the disclosed embodiments, bitrate can represent the amount of data transmitted per unit time. Higher bitrates increase the amount of data transmitted and improve video quality. Specifically, bitrate affects both video quality and transmission speed.

[0053] In the disclosed embodiment, the bitrate allocation for the primary face, secondary face, and background regions can be dynamically adjusted based on the current network bandwidth to ensure a better user experience. The above is merely an example and does not limit all possible scenarios for allocating bitrates for the three regions. This is not an exhaustive list.

[0054] Here, the coded frame is a compressed image frame generated during the video encoding process. Its core function is to convert the original video data into a smaller form that is easier to store and transmit, while trying to maintain the visual quality of the original video.

[0055] Here, metadata is additional information related to the coded frame, describing the structure, characteristics or content of the coded frame, and providing guidance information for the decoder or subsequent processing.

[0056] In the embodiment of the present disclosure, the ROI information can be added to the metadata after encoding the current frame to facilitate use during decoding. The above is only an example and does not limit all possible situations of outputting encoded frames and metadata. This is just not an exhaustive list.

[0057] The technical solution of the embodiment of the present disclosure avoids independent motion estimation calculations by reusing the motion vectors generated by the MCTF process; predicts the current frame ROI based on the physical ROI position of the reference frame to ensure real-time tracking; extracts facial feature points in the first frame and establishes a spatial topological relationship, and reuses the topological relationship in subsequent frames to correct the predicted ROI position and maintain consistency across frames. Differentiated quantization parameters are configured for different levels of regions; combined with a dynamic bit rate allocation strategy, the quality of key areas is prioritized; the encoding end outputs ROI level metadata; by outputting encoded frames and metadata, it can support the decoding end to optimize display, making it easier for the decoder to implement regional post-processing (such as main face sharpening) according to metadata to restore the display quality of key areas. The technical solution of the embodiment of the present disclosure can improve the clarity of the main face while reducing encoding delay, and significantly reduce the freeze rate when bandwidth fluctuates, achieving high-definition and low-bitrate real-time live streaming.

[0058] In some embodiments, the establishment of the initial ROI includes: using a lightweight face detection algorithm to locate the face area in the first frame of the video; dividing the face area into three layers: primary face area, secondary face area and background area according to the key points of the facial features; wherein the primary face area includes the smallest rectangular area covering the eyebrows, eyes, nose and mouth; the secondary face area includes the annular area covering the cheeks and chin; the background area includes the remaining area outside the face.

[0059] Here, a lightweight face detection algorithm is an efficient, low-computing resource-consuming algorithm designed for face detection tasks. For example, the lightweight face detection algorithm can be implemented using existing technologies such as a hierarchical fast detection framework or region extraction technology combined with a specific pre-trained model. The selection can be based on actual circumstances and is not limited in this disclosure.

[0060] In the disclosed embodiment, the first frame of the video stream can be extracted as the first frame of the video. Subsequently, a lightweight model can be used to scan the image, detect the location of the face, and output the rectangular bounding box of the face as the positioning result. In particular, if there are multiple detected faces in the first frame, the main face can be filtered by rules such as area size and position, and the results can be verified or filtered to remove false detections. The above is only an example and does not limit all possible situations for outputting coded frames and metadata. It is just not an exhaustive list here.

[0061] In the disclosed embodiment, a facial key point detection algorithm can first be used to locate key points of the facial features, including eyebrows, eyes, nose, mouth, cheeks, and chin, and the coordinates of these key points can be output as the basis for facial region division. Subsequently, based on the key points of the eyebrows, eyes, nose, and mouth, the minimum rectangular boundary covering these feature points can be calculated as the primary facial region. Next, the primary facial region can be expanded outward to cover the cheeks and chin, forming a ring-shaped area as the secondary facial region. In particular, the expansion range can be determined based on the proportion of the primary facial region. Then, the primary facial region and the secondary facial region are subtracted from the ROI, and the remaining image area is extracted as the background region. Finally, the boundary coordinates of the primary facial region, the secondary facial region, and the background region can be saved, and the region division results can be stored as metadata and associated with the first frame encoding information. The above is only an example and does not limit all possible situations for dividing the three-layer region, but it is not intended to be exhaustive.

[0062] This lightweight algorithm can quickly locate the face region in the first frame of a video, meeting the requirements of real-time video encoding. By extracting key facial features, the face can be precisely divided into primary and secondary facial regions, focusing encoding resources on key areas. Using key facial features to delineate the primary, secondary, and background regions improves the encoding quality of key content while reducing the encoding quality of less important content, reducing resource usage in irrelevant areas and improving the user viewing experience. By outputting the region segmentation results, they can be used as metadata, suitable for dynamically adjusting bitrate allocation and quantization parameters.

[0063] In some embodiments, the video encoding method also includes: for multiple face areas detected in the first frame, calculating the central area weight and area weight of each face respectively; the central area weight reflects the proximity between the center position of the face and the center of the image; the area weight reflects the proportion of the face area to the total area of ​​the image; according to the preset central weight factor and area weight factor, calculating the comprehensive score of each face; selecting the face with the highest comprehensive score as the main tracking target, and determining the face area of ​​the main tracking target as the main face area; merging the remaining face areas that are not selected as the main tracking targets into the secondary face area.

[0064] Here, the central region weight is an evaluation metric used to measure the proximity of a face region to the image center. In the disclosed embodiment, the central region weight reflects whether a face is located at the visual midpoint of the image. In particular, faces closer to the image center are more likely to be the user's focus and are therefore assigned a higher weight.

[0065] Here, the area weight reflects the ratio of the face area to the total image area and is used to assess the importance of the face area. In the disclosed embodiment, faces with larger areas are more prominent and require more emphasis on encoding. In particular, the area weight is used to determine the importance of faces that occupy a larger visual proportion.

[0066] In the embodiment of the present disclosure, a lightweight face detection algorithm can be used to locate all face regions in the first frame and record the bounding box of each face. Subsequently, the weight of the central region can be calculated. Specifically, the coordinates of the center point of the image (x c ,y c ), and then determine the center point coordinates of each face area (x i ,y i ), then the distance between each face center and the image center can be calculated. Finally, the calculated distance can be normalized and mapped to the [0, 1] interval, and then inverted to obtain the central area weight C i , so that the smaller the distance, the greater the weight. Furthermore, the bounding box area of ​​each face can be obtained, and then the ratio of the face area to the total image area can be calculated, and the ratio can be used as the area weight A i The above is only an example and is not intended to limit all possible situations for calculating the central region weight and area weight. This is just not an exhaustive list.

[0067] For example, the process of calculating the distance between the center point of each face and the center point of the image can be expressed by the following formula:

[0068]

[0069] Where, L i Represents the distance between the center point of the i-th face and the center point of the image.

[0070] For example, the ratio of the face area to the total image area is calculated and used as the area weight A i The process can be expressed by the following formula:

[0071]

[0072] Where S i represents the face area of ​​the i-th person; S total Represents the total area of ​​the image.

[0073] Here, the central weight factor is a preset weight coefficient used to emphasize the importance of the central area weight in the comprehensive score. In the embodiment of the present disclosure, the central weight factor can be determined according to actual scenario requirements.

[0074] Here, the area weight factor is a preset weight coefficient used to adjust the influence of the area weight in the comprehensive score. In the embodiment of the present disclosure, the importance of the area weight can be highlighted or weakened by adjusting the size of the area weight factor.

[0075] Here, the comprehensive score is the overall score of the face region calculated by combining the central region weight and the area weight using the weight factor. In the disclosed embodiment, the comprehensive score can be used to compare the importance of multiple face regions.

[0076] In the disclosed embodiments, the center weighting factor and area weighting factor can be set based on actual scenario requirements or preset parameters, and then a comprehensive score can be calculated for each face region. The above is only an example and does not limit all possible scenarios for calculating the comprehensive score for each face. This is not an exhaustive list.

[0077] For example, the process of calculating the comprehensive score can be expressed by the following formula:

[0078] P i =β·C i +γ·A i

[0079] Where, P i represents the comprehensive score of the face area; β represents the central weight factor; γ represents the area weight factor.

[0080] Here, the primary tracking target is the selected face region with the highest comprehensive score, representing the most important face in the image. In the embodiment of the present disclosure, the primary tracking target will be coded as the main face region.

[0081] In the disclosed embodiment, all detected facial regions can be first traversed, the comprehensive scores of each region compared, and the facial region with the highest comprehensive score can be selected and marked as the primary tracking target. Next, the facial features of the primary tracking target can be further extracted to further delineate the primary facial region. The above description is merely illustrative and does not limit all possible scenarios for determining the primary facial region; it is simply not intended to be exhaustive.

[0082] In the embodiment of the present disclosure, for the remaining face regions, their bounding boxes can be merged into a whole and marked as a secondary face region. The above is only an example and does not limit all possible situations for determining the secondary face region. It is just not an exhaustive list here.

[0083] In this way, by calculating the central region weight and area weight, the importance of each face region in the image can be quantified. This also supports multi-target detection and can handle situations where multiple faces are present in the first frame, ensuring that no face regions are missed. By adjusting the central and area weight factors, importance evaluation can be optimized for different scenarios. By calculating a comprehensive score, the bias of a single metric is avoided, making the results more consistent with the needs of the actual scenario. The face with the highest comprehensive score is selected as the primary tracking target, ensuring accurate positioning of the key area. This allows encoding resources to be focused on the most important face regions, improving encoding efficiency and video quality. After the secondary facial regions are merged, the remaining unselected face regions are still encoded, preventing the neglect of secondary targets. At the same time, they can be processed uniformly, saving encoding resources.

[0084] In some embodiments, the video encoding method further includes: obtaining an intra-frame prediction cost and an inter-frame prediction cost of the current frame; in response to the intra-frame prediction cost not exceeding the product of a preset coefficient and the inter-frame prediction cost, determining that a scene switch has occurred; in response to the occurrence of a scene switch, interrupting the current ROI tracking process, using the current frame as a reference frame for the new scene, and re-executing the ROI establishment operation on the reference frame.

[0085] Here, the intra-frame prediction cost is the cost required for predictive coding using the current frame's own content when measuring the compression efficiency of the current frame. In the disclosed embodiment, the intra-frame prediction cost can reflect the resource usage required for intra-frame coding of the current frame.

[0086] Here, the inter-frame prediction cost is the cost required for predictive coding using a reference frame when measuring the compression efficiency of the current frame. In the disclosed embodiment, the inter-frame prediction cost can reflect the resource requirements of the current frame when performing inter-frame coding.

[0087] In the disclosed embodiment, intra-frame prediction coding can be first performed on the current frame, the optimal intra-frame prediction mode can be selected, and the size of the prediction residual, the quantization coefficient, and the coding complexity can be measured to obtain the intra-frame prediction cost. Subsequently, inter-frame prediction coding can be performed using a reference frame, and the complexity of the motion vector, the reference frame matching accuracy, and the processing cost of the coding residual can be calculated to obtain the inter-frame prediction cost. The above is only an example and does not limit all possible situations for obtaining intra-frame prediction costs and inter-frame prediction costs. It is just not an exhaustive list here.

[0088] Here, the preset coefficient is a set weight factor used to adjust the comparison condition between the intra-frame prediction cost and the inter-frame prediction cost. In the embodiment of the present disclosure, the preset coefficient can reflect the relative importance of intra-frame coding and inter-frame coding.

[0089] Here, a scene change refers to a significant change in video content. For example, a scene change can be a transition from one scene to another, resulting in the content of the reference frame and the current frame being no longer relevant. In the disclosed embodiments, the scene change process is accompanied by a sharp increase in the inter-frame prediction cost, while the intra-frame prediction cost is lower.

[0090] In the embodiment of the present disclosure, the relationship between the intra-frame prediction cost and the inter-frame prediction cost can be compared first. When the intra-frame prediction cost does not exceed the product of the preset coefficient and the inter-frame prediction cost, it is considered that a scene switch has occurred. Specifically, since the prediction efficiency of the reference frame for the current frame drops sharply when the scene switch occurs, the inter-frame prediction cost is significantly higher than the intra-frame prediction cost, which indicates that the content of the reference frame can no longer effectively predict the current frame. Exemplarily, the value range of the preset coefficient can be [0.3, 0.6]. In particular, the preset coefficient can also select other value ranges according to actual conditions, and the present disclosure does not limit this. The above is only an exemplary explanation and is not intended to limit all possible situations for determining the occurrence of scene switching. It is just that this is not an exhaustive list.

[0091] In the embodiment of the present disclosure, when a scene switch occurs, since the reference frame and the current frame are no longer correlated, the ROI tracking process based on the reference frame, such as motion vector prediction and feature point correction, can be stopped. Furthermore, the current frame can be marked as a scene switch point, the relevant processes can be reinitialized in the encoding, and the current frame can be set as the new reference frame for inter-frame prediction of subsequent frames. In particular, the data related to the previous reference frame, such as motion vectors and ROI information, can be cleared. Then, a lightweight face detection algorithm can be used on the new reference frame to relocate the face area and establish a new ROI hierarchical division, thereby dividing the main face area, secondary face area and background area based on feature point detection. Finally, the new ROI can be associated with the encoding process to provide an accurate prediction basis for subsequent frames. The above is only an exemplary description and does not limit all possible situations for re-executing the ROI establishment operation, but it is not exhaustive here.

[0092] For example, in the process of re-executing the ROI establishment operation, the current frame where the scene switch occurs can be regarded as the first frame of the new scene, and the initial ROI detection process of face detection and hierarchical division can be re-executed on this frame, and this frame can be used as the reference frame of the next frame in the sequence, and the spatial topological relationship of the feature points established by the ROI of the current frame can be used as the spatial topological relationship of the feature points of the subsequent frames.

[0093] In this way, the calculation of intra-frame and inter-frame prediction costs can dynamically evaluate the coding efficiency of the current frame, providing an accurate basis for scene change determination. By comparing intra-frame and inter-frame prediction costs, significant changes in video scenes can be effectively detected, ensuring that the encoding process adapts to the new scene. When a scene change occurs, the content of the reference frame differs significantly from the current frame, and inter-frame prediction can lead to reduced coding efficiency. Scene change determination can avoid the use of inefficient reference frames. By interrupting the current ROI tracking process, the use of incorrect reference frames for predicting ROI positions is avoided, improving ROI positioning accuracy in subsequent frames. Initializing a new reference frame allows timely adaptation to scene changes, avoiding degradation in encoding performance due to reference frame errors. Reestablishing the ROI ensures high-quality encoding of key areas in the new scene, improving video quality.

[0094] In some embodiments, the motion vector generated by the reference frame during the motion compensation time domain filtering process is reused to determine the ROI position prediction value of the current frame, including: dividing the physical ROI of the reference frame into a set of macroblocks according to the standard macroblock size; obtaining the motion vector calculated for each macroblock during the motion compensation time domain filtering process; and using a motion vector aggregation algorithm to generate the ROI position prediction value of the current frame.

[0095] Here, the physical ROI of the reference frame refers to the actual location area of ​​the ROI extracted from the reference frame during the video encoding process and determined after precise detection and correction. In the embodiment of the present disclosure, this area is the part of the reference frame that needs to be encoded most.

[0096] Here, a standard macroblock is a fixed-size rectangular block obtained by dividing an image during video encoding and used for subsequent compression processing. In the disclosed embodiments, a macroblock is the basic unit of video encoding. For example, in the H.264 standard, a typical macroblock size is 16×16 pixels; similarly, in the H.265 standard, a typical macroblock size ranges from 8×8 pixels to 64×64 pixels.

[0097] Here, the macroblock set refers to a set formed by dividing the physical ROI of the reference frame into a number of macroblocks. In the embodiment of the present disclosure, the macroblocks in the macroblock set collectively cover the entire ROI.

[0098] In the disclosed embodiment, a detected ROI can first be extracted from a reference frame and then divided into several rectangular blocks according to a preset standard macroblock size. The resulting rectangular blocks are used as macroblocks. In particular, if the ROI boundary does not completely match the macroblock size, padding or cropping can be used to complete the division. Finally, all the divided macroblocks can be recorded as a set, where each macroblock contains its position coordinates and pixel or feature data. The above is only an example and does not limit all possible situations for dividing macroblock sets. However, this is not an exhaustive list.

[0099] In the disclosed embodiment, MV is used to estimate the motion relationship between the reference frame and the current frame. Typical methods include block matching and optical flow. In the MCTF process, motion estimation is performed on each macroblock, the best matching block in the current frame is found, and then an MV is generated. In particular, the motion vector of each macroblock contains a horizontal component and a vertical component; wherein the horizontal component reflects the displacement of the macroblock in the horizontal direction; similarly, the vertical component reflects the displacement of the macroblock in the vertical direction. Finally, the motion vector of each macroblock can be bound to its spatial position to obtain the MV of each macroblock. The above is only an exemplary explanation and is not intended to limit all possible situations for obtaining MV, but it is not exhaustive here.

[0100] Here, the aggregation algorithm refers to generating an overall result by merging, statistically analyzing, or optimizing the MVs of multiple macroblocks. In the embodiment of the present disclosure, the aggregation algorithm is used to integrate the motion information of a set of macroblocks into a predicted ROI position value for the current frame.

[0101] Here, the ROI position prediction value is the position estimation result of the ROI in the current frame obtained by motion vector analysis. In the embodiment of the present disclosure, the ROI position prediction value is the predicted position of the ROI position of the reference frame after motion compensation adjustment.

[0102] In the embodiment of the present disclosure, the horizontal component and the vertical component can be extracted from the motion vectors of all macroblocks respectively, and the median of the horizontal component and the median of the vertical component can be calculated respectively to eliminate the influence of abnormal motion vectors. Then, the center point coordinates of the reference frame ROI can be obtained, and then the median of the horizontal component and the median of the vertical component can be superimposed on the center point coordinates of the reference frame ROI to obtain the predicted center point of the current frame ROI. Finally, the predicted center point can be used as the center point of the current frame ROI, keeping the width and height of the area consistent with the reference frame ROI, and reconstructing the ROI position of the current frame. In particular, the ROI position prediction value includes the center point coordinates and the region boundary of the current frame ROI, where the region boundary can be calculated from the center point and size. The above is only an exemplary explanation and is not intended to limit all possible situations for generating ROI position prediction values, but it is not exhaustive here.

[0103] This ensures that each prediction is based on the most recently corrected physical ROI, ensuring that the predicted position does not deviate from the actual position due to error propagation. Aggregation using the median of motion vectors effectively filters out abnormal or noisy motion vectors, improving prediction accuracy. By dividing the ROI into a set of macroblocks, the computational complexity of motion compensation is reduced. Aggregation of motion vectors directly generates predictions without re-detecting the ROI, significantly improving real-time processing efficiency. If the ROI in the video moves over time, the motion compensation and aggregation algorithms can dynamically adjust the ROI position to ensure the validity of the prediction results.

[0104] In some embodiments, the video encoding method further includes: calculating the consistency of the motion vectors of each macroblock within the physical ROI of the reference frame as a confidence score; in response to the confidence score exceeding a preset occlusion determination threshold, determining that the corresponding macroblock area is partially occluded; for the macroblock area that is partially occluded, using a feature point matching method to update its position information in the current frame.

[0105] Here, the confidence score is a metric that measures the consistency of motion vectors across macroblocks within a physical ROI of a reference frame. In the disclosed embodiments, the confidence score can reflect whether the motion vectors of macroblocks within the region tend to move consistently. Specifically, the confidence score is used to determine whether there is occlusion or irregular motion within the ROI. Occlusion can lead to inconsistent motion vectors for macroblocks, thereby reducing the confidence score.

[0106] In the embodiment of the present disclosure, the statistical features of all macroblock motion vectors within the physical ROI of the reference frame can be first calculated, and the statistical features can be used as consistency indicators. Exemplarily, the statistical features can be selected as mean or variance; wherein, when the mean is used as the statistical feature, the average direction and amplitude of all macroblock motion vectors can be calculated; similarly, when the variance is used as the statistical feature, the degree of discreteness of the motion vector can be evaluated. Then, the confidence score can be calculated based on the consistency indicator. In particular, the lower the confidence score value, the more dispersed the motion vector and the stronger the inconsistency. The above is only an exemplary explanation and is not intended to limit all possible situations for calculating the confidence score, but it is not exhaustive here.

[0107] For example, when the variance is selected as the consistency index, the process of calculating the confidence score can be expressed by the following formula:

[0108]

[0109] Where Co represents the confidence score; D represents the variance of all macroblock MVs.

[0110] Specifically, if the MVs are highly consistent, the variance is small and the confidence score is close to 1; if the motion vectors are significantly different, the variance increases and the confidence score decreases, approaching 0.

[0111] Here, the occlusion determination threshold is a preset value used to compare with the confidence score to determine whether occlusion may occur in certain macroblocks within the ROI. In the embodiment of the present disclosure, the occlusion determination threshold can be used as a sensitivity adjustment parameter for occlusion detection. In particular, the lower the value of the occlusion determination threshold, the more sensitive the determination of occlusion is; conversely, the higher the value, the greater the tolerance to occlusion. Exemplarily, the value range of the occlusion determination threshold can be [0.2, 0.5]. In particular, the occlusion determination threshold can also select other value ranges according to actual conditions, and the present disclosure does not limit this.

[0112] In the disclosed embodiment, a confidence score can be calculated for each ROI macroblock area and compared with the occlusion determination threshold. If the confidence score is lower than the preset occlusion determination threshold (such as 0.3), it is determined that the macroblock area has been partially occluded. Specifically, occlusion can cause significant differences or irregular distributions in the motion vectors of the reference frame and the current frame in the local area, which is manifested as increased inconsistency in the motion vectors, thereby reducing the confidence score. The above is only an exemplary explanation and is not intended to limit all possible situations for determining local occlusion, but it is not exhaustive here.

[0113] Here, the feature point matching method is an algorithm for detecting and tracking feature points in the ROI. In the embodiment of the present disclosure, the position information of the local area can be updated by matching feature points in the current frame and the reference frame.

[0114] In the disclosed embodiment, a feature point detection algorithm can be used in the occluded macroblock area of ​​the reference frame to extract key points. Subsequently, feature point detection can be run again in the corresponding position area of ​​the current frame to obtain a feature point candidate set. Next, the feature points of the reference frame and the current frame can be compared to find matching pairs, and then the displacement of the local area can be estimated based on the matched feature point pairs. Exemplarily, the matching method can use Euclidean distance or Hamming distance. Finally, the feature point matching results can be mapped to the macroblock area to update the position of the occluded macroblock in the current frame. The above is only an exemplary explanation and is not intended to limit all possible situations for updating position information, but it is not exhaustive here.

[0115] Exemplarily, when a feature point is blocked, the actual straight-line distance of each pair of feature points visible in the current frame can be measured first; when the relative deviation between the measured distance value of a pair of feature points and the corresponding distance value stored in the relative distance matrix exceeds a preset error threshold (such as the preset error threshold has a value range of 20% to 40%), the positioning of the feature point is determined to be invalid; all feature points that are determined to be invalid are eliminated; and only the remaining valid feature points are used to recalculate the ROI position correction amount.

[0116] By calculating the consistency of motion vectors, we can effectively detect abnormal motion within macroblocks and accurately identify local occlusions. By setting different ranges for the occlusion threshold, we can adapt to the occlusion detection needs of different scenarios. When occlusion causes motion vectors to fail, feature point matching methods can use local features to match, ensuring the updated position information of the occluded area and preserving the integrity of the ROI. By detecting and correcting motion vectors in occluded areas in real time, we reduce error propagation.

[0117] In some embodiments, the ROI position prediction value is corrected based on the spatial topological relationship of the feature points in the initial ROI, including: calculating the expected position coordinates of each feature point in the current frame based on the spatial topological relationship of the feature points recorded in the initial ROI and combining the ROI position prediction value; obtaining the actual position coordinates of the corresponding feature points in the current frame through the facial feature point detection algorithm; calculating the spatial transformation parameters based on the correspondence between the expected position coordinates and the actual position coordinates; and applying the spatial transformation parameters to correct the ROI position prediction value.

[0118] Here, the expected position coordinates are the theoretical positions of the feature points in the current frame calculated through the spatial topological relationship of the feature points recorded by the initial ROI and the ROI position prediction value.

[0119] In the embodiment of the present disclosure, the spatial topological relationship of the feature points of the initial ROI can be obtained first, and the initial topological relationship can be mapped to the current frame using the predicted ROI center point, size and shape. For example, the relative offset of the feature point can be added to the center point of the ROI position prediction value to obtain the expected position coordinates. Subsequently, the theoretical position of each feature point in the facial feature points in the current frame can be calculated based on the predicted ROI position to form an expected position set of feature points. The above is only an exemplary explanation and is not intended to limit all possible situations for calculating the expected position coordinates of feature points, but it is not exhaustive here.

[0120] Here, the facial feature point detection algorithm is used to detect the key feature points in the face and extract the actual position coordinates of the feature points.

[0121] Here, the actual position coordinates are the real positions of the feature points detected in the current frame by the facial feature point detection algorithm.

[0122] In the embodiment of the present disclosure, the facial feature point detection algorithm can be first run in the current frame to accurately extract the actual positions of feature points such as eyes, nose, and mouth. For example, a deep learning model can be used to locate key points or image processing technology can be used to extract edges or corners. Subsequently, the detected feature point coordinates can be recorded to form the actual position coordinates of the corresponding feature points. The above is only an exemplary explanation and is not intended to limit all possible situations for obtaining the actual position coordinates of feature points. It is just that this is not an exhaustive list.

[0123] Here, the spatial transformation parameters are used to describe the conversion information of the geometric relationship between the desired position and the actual position. In the embodiment of the present disclosure, the spatial transformation parameters can be represented by a mathematical model and can include translation parameters, rotation parameters, and scaling parameters.

[0124] In the embodiment of the present disclosure, the coordinates of the expected position and the coordinates of the actual position can be matched one by one, and then the deviations between each corresponding group can be compared, and then the geometric transformation parameters can be calculated based on the deviation information. For example, the offset between the center of the expected position and the center of the actual position can be calculated to obtain the translation parameters; the angular change of the feature point set can be analyzed to obtain the rotation parameters; the change ratio of the distance between the feature points can be calculated to obtain the scaling parameters. In particular, a linear transformation or an affine transformation model can also be used to fit the transformation relationship from the expected position to the actual position. The above is only an exemplary explanation and is not intended to limit all possible situations for calculating spatial transformation parameters. It is just that this is not an exhaustive list.

[0125] In the disclosed embodiment, the spatial transformation parameters can be applied to the predicted ROI position, and the corrected ROI position is then used as the final ROI position of the current frame. The above is only an example and does not limit all possible situations in which the ROI position prediction value can be corrected. This is just not an exhaustive list.

[0126] For example, the process of correcting the predicted ROI position can first calculate the expected position coordinates of each feature point in the current frame based on the current frame ROI position predicted by the motion vector and the normalized coordinates. Subsequently, the actual position coordinates of each feature point in the current frame can be detected. Next, the ROI position correction value can be calculated by minimizing the total distance deviation between the actual position coordinates of all feature points and the expected position coordinates. Finally, the correction value can be applied to adjust the position of the predicted ROI area.

[0127] In this way, the predicted ROI position is corrected using spatial transformation parameters, reducing deviations caused by scene changes or prediction errors. The facial feature point detection algorithm can adapt to various facial expressions, head movement, and occlusions, ensuring accurate detection of the actual position. Furthermore, spatial transformation parameters can dynamically adjust the ROI position to adapt to dynamically changing scenes. The corrected ROI position is closer to the actual area, avoiding wasted resources caused by misjudgments.

[0128] In some embodiments, the bitrate of the three-layer region is dynamically allocated according to the real-time network bandwidth, including: in response to detecting that the real-time network bandwidth is not lower than a preset bandwidth threshold, controlling the proportion of the bitrate allocated to the main facial region to the total bitrate to be not lower than a preset ratio threshold; in response to detecting that the real-time network bandwidth is lower than the preset bandwidth threshold, ensuring that the bitrate of the main facial region is not lower than a preset minimum guaranteed bitrate, and compressing the bitrate of the background region to within a first compression coefficient range of its original baseline value.

[0129] Here, the bandwidth threshold is a preset reference value. In the disclosed embodiments, the bandwidth threshold can be used to determine whether the network bandwidth is sufficient. Specifically, if the network bandwidth is greater than or equal to the bandwidth threshold, the network status is considered good and a higher bitrate can be allocated; otherwise, the bitrate needs to be reduced to accommodate the bandwidth limitation.

[0130] Here, the ratio threshold is the minimum ratio of the bitrate allocated to the main facial region to the total bitrate when sufficient bandwidth is available. In the disclosed embodiments, the ratio threshold ensures that the main facial region receives sufficient bitrate resources to maintain high-quality encoding when network bandwidth is sufficient. For example, the ratio threshold can range from [0.5 to 0.7]. Specifically, the ratio threshold can also be selected from other ranges based on practical circumstances, and this disclosure does not impose any limitations thereon.

[0131] In the embodiment of the present disclosure, the current network bandwidth value can be obtained in real time through the network status monitoring mechanism, and then the current network bandwidth value can be compared with the bandwidth threshold. For example, the process of obtaining the current network bandwidth value can be carried out through the Real-time Transport Control Protocol (RTCP) and other methods. Then, when it is detected that the real-time network bandwidth is not lower than the preset bandwidth threshold, that is, when it is considered that the bandwidth is sufficient, the bit rate of the main facial area can be allocated according to the proportional threshold, that is, the product of the total bit rate value corresponding to the real-time network bandwidth and the proportional threshold is used as the bit rate of the main facial area. The above is only an exemplary explanation and is not intended to limit all possible situations for controlling bit rate allocation, but it is not exhaustive here.

[0132] Here, the minimum guaranteed bitrate is the minimum encoding bitrate that must be allocated to the primary facial region even when bandwidth is insufficient, to maintain basic visual quality for that region. In the disclosed embodiments, this minimum guaranteed bitrate prevents significant degradation in the quality of the primary facial region, ensuring clarity in key areas even with limited network bandwidth.

[0133] Here, the original reference value is the initial bitrate assigned to the background region when bandwidth is sufficient. In the disclosed embodiment, the original reference value represents the normal encoding state of the background region. In particular, when network bandwidth is insufficient, bitrate compression can be performed on the background region based on the original reference value.

[0134] Here, the first compression coefficient is the compression ratio range of the background region when bandwidth is insufficient, which is used to reduce the bit rate allocation of the background region to save bandwidth. In the embodiment of the present disclosure, the first compression coefficient can control the bit rate compression amplitude of the background region to maintain efficient utilization of network bandwidth. Exemplarily, the value range of the first compression coefficient can be [0.5, 0.8]. In particular, the first compression coefficient can also be selected from other value ranges based on actual conditions, which is not limited by the present disclosure.

[0135] In this disclosed embodiment, when network bandwidth is insufficient, the bitrate of the primary facial region is first ensured to be no less than the minimum guaranteed bitrate. This is done by comparing the minimum guaranteed bitrate with the product of the total bitrate value and the ratio threshold, and taking the maximum value between the two as the bitrate of the primary facial region. Subsequently, the bitrate of the background region can be compressed to within a first compression factor of the original baseline value. The product of the original baseline value and the first compression factor is used as the bitrate of the background region. The above description is merely illustrative and does not limit all possible scenarios for compressing the background region bitrate; this is not intended to be exhaustive.

[0136] This prioritizes bitrate allocation for the primary facial area, ensuring encoding quality for the primary facial area, regardless of network bandwidth availability, and enhancing the user experience. A minimum guarantee mechanism ensures that the primary facial area maintains the lowest bitrate, even when network bandwidth is insufficient, preventing significant distortion in key video areas. By monitoring network bandwidth in real time, the system dynamically adjusts the bitrate allocation strategy to ensure video stream stability under varying bandwidth conditions. By controlling the range of the first compression coefficient, flexible bitrate compression is applied to the background area, improving bandwidth utilization efficiency.

[0137] In some embodiments, the video encoding method also includes: when insufficient bandwidth is triggered for the first time, compressing the background area bit rate to a specific value within a first compression coefficient range; if the bandwidth continues to decrease and the decrease per unit time exceeds a preset rate threshold, compressing the secondary facial area bit rate to within a second compression coefficient range; during the compression process, ensuring that the main facial area bit rate is not lower than a protection redundancy coefficient times the preset minimum guaranteed bit rate.

[0138] In the disclosed embodiments, network bandwidth can be monitored in real time. When the bandwidth is detected to fall below a preset bandwidth threshold for the first time, background region bitrate compression logic is triggered. Specifically, a specific compression value can be selected within a first compression coefficient range, and the background region bitrate can be adjusted to the product of the compression value and the total bitrate, freeing up bandwidth resources for the primary and secondary facial regions. The above description is merely illustrative and does not limit all possible scenarios for background region bitrate compression; this is not intended to be exhaustive.

[0139] Here, the rate threshold is a preset standard for the bandwidth drop rate. In the embodiment of the present disclosure, the rate threshold can be used to determine whether the network bandwidth drops rapidly within a unit time.

[0140] Here, the second compression coefficient represents the compression range for the secondary facial region's bitrate when bandwidth drops dramatically, reducing the bitrate allocation for the secondary facial region. In the disclosed embodiments, the second compression coefficient dynamically adjusts the bitrate allocation for the secondary facial region, reserving more bandwidth resources for the primary facial region. For example, the second compression coefficient can range from [0.8, 0.95]. Other ranges of values ​​for the second compression coefficient can also be selected based on practical circumstances, and this disclosure does not limit this.

[0141] In this disclosed embodiment, the bandwidth reduction rate can be calculated in real time. If the reduction exceeds a rate threshold, the secondary facial region bitrate compression logic is triggered. Specifically, a specific compression value can be selected within the second compression coefficient range, and the bitrate of the secondary facial region can be adjusted to the product of the compression value and the total bitrate, freeing up bandwidth resources for the primary facial region. The above is merely an example and does not limit all possible scenarios for compressing the bitrate of the secondary facial region. This is simply not an exhaustive list.

[0142] Here, the protection redundancy factor is a safeguard multiplier for the bitrate of the primary facial region when the bandwidth drops sharply, ensuring minimum visual quality in that region. In the disclosed embodiments, even if the bandwidth drops rapidly, the protection redundancy factor ensures that the primary facial region will not experience severe distortion. For example, the protection redundancy factor can range from [1.1, 1.3]. In particular, the protection redundancy factor can also be selected from other ranges based on practical circumstances, and this disclosure does not impose any limitations on this.

[0143] In the disclosed embodiment, the minimum bitrate for the main area can be calculated based on the minimum guaranteed bitrate and the protection redundancy factor. For example, the product of the minimum guaranteed bitrate and the protection redundancy factor can be used as the minimum bitrate for the main area. Subsequently, even if the bandwidth decreases rapidly, the bitrate allocated to the main facial area is ensured to be no less than the minimum bitrate for the main area, thereby safeguarding the visual quality of the key area. The above is merely an example and does not limit all possible scenarios for bitrate protection in the main facial area; it is simply not an exhaustive list.

[0144] In this way, the main facial area always maintains a high bit rate, and even if the bandwidth drops rapidly, it can meet the minimum visual quality requirements and avoid distortion or blurring. During the compression process, the quality of the main facial area is prioritized, and the allocation priority of encoding resources is clarified. The bit rate of the background area is compressed when insufficient bandwidth is first detected to avoid affecting the main area and secondary areas. Bandwidth changes are dynamically monitored according to the rate threshold, and secondary area bit rate compression is triggered in time to ensure the stability of the video stream. When the bandwidth drops further, the bit rate of the secondary facial area is compressed, further saving bandwidth resources. The main facial area always remains clear, and the quality of key areas of the video stream will not be affected even if the network bandwidth drops.

[0145] In some embodiments, differentiated quantization parameters are configured for the corrected primary facial area, secondary facial area, and background area, respectively, including: setting a minimum quantization parameter offset value for the primary facial area so that its quantization parameter is lower than the encoder base quantization parameter; enabling a first prediction mode combination and allocating the highest priority encoding resource; setting a medium quantization parameter offset value for the secondary facial area so that its quantization parameter is equal to or slightly higher than the encoder base quantization parameter; enabling a second prediction mode combination and allocating normal priority encoding resources; setting a maximum quantization parameter offset value for the background area, enabling a third prediction mode combination, and allocating the lowest priority encoding resource.

[0146] Here, the first prediction mode combination is a prediction mode set assigned by the encoder to the primary facial region. In the embodiment of the present disclosure, the first prediction mode combination may include prediction modes divided into 4×4 and 8×8 block sizes.

[0147] In the disclosed embodiment, since the main facial area is the key area that the user pays attention to, the lowest quantization parameter is required to reduce the quantization error. Exemplarily, the quantization parameter of the main facial area can be the difference between the base quantization parameter and the lowest quantization parameter offset value. Subsequently, the first prediction mode combination can be enabled, and the main facial area can be divided using 4×4 and 8×8 block sizes to perform high-grained encoding. In particular, the 4×4 size division is suitable for processing areas with complex textures and details such as the eyes and mouth, and the 8×8 size division is suitable for processing slightly larger uniform areas. Furthermore, the main facial area obtains more computing resources and bandwidth to ensure that the details of the area are retained with the highest quality. The above is only an exemplary explanation and is not intended to limit all possible situations for enabling the first prediction mode combination, but it is not exhaustive here.

[0148] Here, the second prediction mode combination is a prediction mode set assigned by the encoder to the secondary facial region. In the embodiment of the present disclosure, the second prediction mode combination may include prediction modes divided into 8×8 and 16×16 block sizes.

[0149] In the disclosed embodiment, since the secondary facial area is not the main focus of the user, but still requires good visual quality, a medium quantization parameter is required to reduce the quantization error. Exemplarily, the quantization parameter of the secondary facial area can be the difference between the base quantization parameter and the medium quantization parameter offset value. Subsequently, the second prediction mode combination can be enabled, and the secondary facial area can be divided using block sizes of 8×8 and 16×16 to perform medium-fine-grained encoding. Furthermore, the computing resources and bandwidth allocated to the secondary facial area are lower than those of the main facial area, but higher than those of the background area, ensuring a balance between visual quality and resource allocation. The above is only an exemplary explanation and is not intended to limit all possible situations for enabling the second prediction mode combination, but it is not exhaustive here.

[0150] Here, the third prediction mode combination is a prediction mode set assigned by the encoder to the background area. In the embodiment of the present disclosure, the third prediction mode combination may include prediction modes divided by a block size of 16×16 or larger.

[0151] In the disclosed embodiment, the background area is the part that the user pays the least attention to, so it can accept the highest quantization parameter, reducing the detail retention of the background to save resources. Exemplarily, the quantization parameter of the background area can be the difference between the base quantization parameter and the highest quantization parameter offset value. Subsequently, the third prediction mode combination can be enabled, and the background area can be divided using a block size of 16×16 or larger to perform low-grained encoding. Furthermore, the background area obtains the least computing resources and bandwidth, giving priority to ensuring the encoding quality of the main facial area. The above is only an exemplary explanation and is not intended to limit all possible situations for enabling the third prediction mode combination, but it is not exhaustive here.

[0152] In this way, the primary facial area uses the lowest quantization parameters and small block division, ensuring the highest encoding quality for the key areas of the video that users are interested in. At the same time, the first prediction mode combination allows for fine encoding of complex details, reducing visual distortion and enhancing the user experience. Resources are concentrated in the primary facial area, balanced with secondary facial areas, and reduced in the background area, achieving efficient use of encoding resources. Quantization parameters are dynamically adjusted based on the importance of the video content, ensuring clear details in the key areas and reducing resources in the secondary and background areas. Different block division methods are used in different areas to adapt to the texture complexity of the area, achieving the optimal balance between quality and efficiency.

[0153] In some embodiments, the first frame ROI detection step can be performed first. Specifically, a lightweight face detection algorithm can be used to detect the face area in the first frame of the live video, determine the initial ROI, and record its position, size, shape, and other information. At the same time, based on the key feature points of the face (such as the eyes, nose, mouth, etc.), the face area is further divided into three levels: the primary face area (including the core part of the facial features), the secondary face area (the facial outline and the surrounding area), and the background area (the rest of the face).

[0154] Furthermore, an ROI tracking step based on MCTF motion estimation can be performed. Specifically, the MV obtained in the MCTF process can be reused to continuously track the face ROI area. For example, the MV of each macroblock between the current frame and the adjacent reference frame can be calculated in the motion estimation stage of MCTF; in particular, this process is usually a necessary module of the encoder, so it does not take up additional time; then, for the face ROI determined in the first frame, it can be divided into multiple macroblocks, and the motion vectors of the adjacent macroblocks are aggregated and analyzed to predict the position of the ROI in the current frame; finally, the predicted position can be corrected in combination with the relative position relationship of the key feature points of the face to obtain an accurate ROI position, and the area ranges of the three levels can be updated synchronously.

[0155] Furthermore, a scene cut determination step can be performed. Specifically, whether a scene cut has occurred can be determined by calculating the difference in image features between the current frame and adjacent frames. Exemplarily, whether a scene cut has occurred can be determined by calculating the difference in cost between the intra-frame prediction mode and the inter-frame prediction mode for the current frame. That is, if the cost of the intra-frame prediction mode is significantly lower than that of the inter-frame mode, a scene cut is determined to have occurred; otherwise, no scene cut is determined to have occurred.

[0156] In particular, if it is determined that a scene switch occurs, the ROI detection step is re-executed to update the face ROI and its three-level division; if no scene switch occurs, the ROI and level division obtained based on MCTF motion estimation tracking continue to be used.

[0157] Furthermore, a rate-layered encoding step can be performed. Specifically, the encoding parameters of the encoder can be adjusted according to the determined ROI and its three levels for differential encoding. For the main face area, a lower quantization parameter, a higher encoding quality level, and a finer prediction mode are used to ensure clear facial features and details; for the secondary face area, the quantization parameter is set to a medium level and a conventional prediction mode is used; for the background area, the quantization parameter is appropriately increased, a coarser prediction mode is used, the encoding quality is reduced, and the bit rate allocation is reduced. At the same time, the bit rate ratio of each level is adjusted in real time according to the network bandwidth. When the bandwidth is sufficient, the bit rate share of the main face area is increased; when the bandwidth is insufficient, the basic bit rate requirement of the main face area is prioritized.

[0158] Finally, the encoding output step can be performed. Specifically, the video frames after the rate-layered encoding can be encapsulated to generate a live video stream output.

[0159] Figure 2 Another flow chart of the video encoding method according to the embodiment of the present disclosure is shown as follows: Figure 2 As shown, the video encoding method includes:

[0160] S201, obtaining live video frames;

[0161] S202. Perform ROI detection using a lightweight face detection algorithm:

[0162] S203, multiplexing the MV in the MTCF process and performing ROI tracking;

[0163] S204, determine whether a scene switch occurs, if a scene switch occurs, go to S202; if no scene switch occurs, go to S205;

[0164] S205, rate-layered coding;

[0165] S206: Encapsulate the video frame and output it.

[0166] This solution provides a low-latency high-definition encoding framework with the trinity of "tracking-layering-flow control". By reusing MCTF motion vectors to eliminate redundant calculations, the processing delay of the motion tracking module is significantly reduced. Independent motion estimation operations are avoided to achieve near-zero-overhead ROI cross-frame tracking, thereby improving the real-time performance of encoding. The spatial consistency of feature points is maintained through the cross-frame multiplexing mechanism of the first-frame topological relationship. The quality fluctuations of feature points reconstructed frame by frame are eliminated to ensure the clarity of details in key areas, thereby optimizing the encoding quality. By combining layered encoding with dynamic bitrate allocation, the quality and smoothness are intelligently balanced when bandwidth fluctuates, enhancing bandwidth adaptability. The risk of live broadcast freezes is reduced through the coordinated optimization of motion tracking and bitrate control. The overall bitrate efficiency is optimized while ensuring the visual quality of the main facial area.

[0167] It should be understood that Figure 2 The schematic diagram shown is only exemplary and not restrictive, and it is scalable, and those skilled in the art can Figure 2 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0168] The present disclosure provides a video encoding device. Figure 3As shown, the apparatus may include: a reference frame acquisition module 301 for acquiring a reference frame adjacent to a current frame and for which an initial ROI hierarchical division has been established; a position prediction module 302 for reusing a motion vector generated by the reference frame during motion-compensated temporal filtering to determine a predicted ROI position value for the current frame; a position correction module 303 for extracting facial feature points of the current frame and correcting the predicted ROI position value based on the spatial topological relationship of the feature points in the initial ROI; a parameter configuration module 304 for respectively configuring differentiated quantization parameters for the corrected primary facial region, secondary facial region, and background region; a bit rate allocation module 305 for dynamically allocating bit rates for the three layers of regions based on real-time network bandwidth; and a data output module 306 for outputting a coded frame of the current frame and associated ROI-level metadata; wherein the ROI-level metadata includes: boundary coordinates of the primary facial region, secondary facial region, and background region; quantization parameter offset values ​​corresponding to each region; and bit rate allocation weight values ​​for each region.

[0169] In some embodiments, the video encoding apparatus further includes: a face positioning module 307 ( Figure 3 ), for locating the face area using a lightweight face detection algorithm for the first frame of the video; a hierarchical division module 308 ( Figure 3 (not shown in the figure), which is used to divide the face area into three layers: the main face area, the secondary face area and the background area according to the key points of the facial features; the main face area includes the smallest rectangular area covering the eyebrows, eyes, nose and mouth; the secondary face area includes the annular area covering the cheeks and chin; the background area includes the remaining area outside the face.

[0170] In some embodiments, the video encoding apparatus further includes: a weight calculation module 309 ( Figure 3 ), for calculating the central region weight and area weight of each face for the multiple face regions detected in the first frame; the central region weight reflects the proximity between the center position of the face and the center of the image; the area weight reflects the proportion of the face region to the total area of ​​the image; the comprehensive score calculation module 310 ( Figure 3 ), which is used to calculate the comprehensive score of each face according to the preset central weight factor and area weight factor; the main face determination module 311 ( Figure 3 ), for selecting the face with the highest comprehensive score as the primary tracking target, and determining the face area of ​​the primary tracking target as the primary facial area; the secondary facial determination module 312 ( Figure 3 (not shown in the figure), which is used to merge the remaining face areas that are not selected as the main tracking target into the secondary face area.

[0171] In some embodiments, the video encoding apparatus further includes: a prediction cost acquisition module 313 ( Figure 3), used to obtain the intra-frame prediction cost and inter-frame prediction cost of the current frame; scene switching determination module 314 ( Figure 3 ), for determining that a scene switch occurs in response to the intra-frame prediction cost not exceeding the product of a preset coefficient and an inter-frame prediction cost; a region update module 315 ( Figure 3 (not shown) is used to interrupt the current ROI tracking process in response to a scene switch, use the current frame as a reference frame of the new scene, and re-execute the ROI establishment operation on the reference frame.

[0172] In some embodiments, the position prediction module 302 includes: a macroblock division submodule, which is used to divide the physical ROI of the reference frame into a set of macroblocks according to the standard macroblock size; a motion vector acquisition submodule, which is used to obtain the motion vector calculated by each macroblock during the motion compensation time domain filtering process; and a prediction generation submodule, which is used to generate the ROI position prediction value of the current frame using a motion vector aggregation algorithm.

[0173] In some embodiments, the video encoding apparatus further includes: a confidence calculation module 316 ( Figure 3 ), used to calculate the consistency of the motion vectors of each macroblock in the physical ROI of the reference frame as a confidence score; local occlusion determination module 317 ( Figure 3 ), for determining that a local occlusion occurs in the corresponding macroblock area in response to a confidence score exceeding a preset occlusion determination threshold; a position information updating module 318 ( Figure 3 (not shown) is used to update the position information of the macroblock area where partial occlusion occurs in the current frame using the feature point matching method.

[0174] In some embodiments, the position correction module 303 includes: an expected position calculation submodule, which is used to calculate the expected position coordinates of each feature point in the current frame based on the spatial topological relationship of the feature points recorded in the initial ROI and the ROI position prediction value; an actual position acquisition submodule, which is used to obtain the actual position coordinates of the corresponding feature points in the current frame through the facial feature point detection algorithm; a transformation parameter calculation submodule, which is used to calculate the spatial transformation parameters based on the correspondence between the expected position coordinates and the actual position coordinates; and a position prediction correction submodule, which is used to apply the spatial transformation parameters to correct the ROI position prediction value.

[0175] In some embodiments, the bitrate allocation module 305 includes: a main face bitrate allocation submodule, configured to, in response to detecting that the real-time network bandwidth is not less than a preset bandwidth threshold, control the ratio of the bitrate allocated to the main face region to the total bitrate to be not less than a preset ratio threshold; and a background bitrate allocation submodule, configured to, in response to detecting that the real-time network bandwidth is less than a preset bandwidth threshold, ensure that the bitrate of the main face region is not less than a preset minimum guaranteed bitrate, and compress the bitrate of the background region to within a first compression coefficient range of its original reference value.

[0176] In some embodiments, the video encoding device further includes: a background bit rate dynamic adjustment module 319 ( Figure 3 ), which is used to compress the bit rate of the background area to a specific value within the first compression coefficient range when the bandwidth is insufficient for the first time; if the bandwidth continues to decrease and the decrease per unit time exceeds the preset rate threshold, the bit rate of the secondary facial area is compressed to a range of the second compression coefficient; the main facial bit rate dynamic adjustment module 321 ( Figure 3 (not shown in the figure), which is used to ensure that the code rate of the main facial area is not lower than the protection redundancy coefficient times the preset minimum guaranteed code rate during the compression process.

[0177] In some embodiments, the parameter configuration module 304 includes: a first combining submodule, used to set a minimum quantization parameter offset value for the primary facial area so that its quantization parameter is lower than the encoder base quantization parameter: enable the first prediction mode combination, and allocate the highest priority encoding resources; a second combining submodule, used to set a medium quantization parameter offset value for the secondary facial area so that its quantization parameter is equal to or slightly higher than the encoder base quantization parameter; enable the second prediction mode combination, and allocate normal priority encoding resources; and a third combining submodule, used to set a maximum quantization parameter offset value for the background area, enable the third prediction mode combination, and allocate the lowest priority encoding resources.

[0178] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0179] The video encoding device in the disclosed embodiment avoids independent motion estimation calculations by reusing the motion vectors generated by the MCTF process; predicts the current frame ROI based on the physical ROI position of the reference frame to ensure real-time tracking; extracts facial feature points in the first frame and establishes a spatial topological relationship, and reuses the topological relationship in subsequent frames to correct the predicted ROI position and maintain consistency across frames. Differentiated quantization parameters are configured for different levels of regions; combined with a dynamic bit rate allocation strategy, the quality of key areas is prioritized; the encoding end outputs ROI level metadata; by outputting encoded frames and metadata, it can support the decoding end to optimize display, making it easier for the decoder to implement regional post-processing (such as main face sharpening) based on metadata to restore the display quality of key areas. This can improve the clarity of the main face while reducing encoding delay, significantly reduce the freeze rate when bandwidth fluctuates, and achieve high-definition, low-bitrate real-time live streaming.

[0180] The embodiment of the present disclosure provides a scenario diagram of a video encoding method, such as Figure 4 shown.

[0181] As mentioned above, the video encoding method provided by the embodiments of the present disclosure is applied to electronic devices. The electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0182] Specifically, the electronic device can perform the following operations:

[0183] Acquire a reference frame that is adjacent to the current frame and has established an initial ROI hierarchical division; reuse the motion vector generated by the reference frame during the motion compensation time domain filtering process to determine the ROI position prediction value of the current frame; extract the facial feature points of the current frame and correct the ROI position prediction value based on the spatial topological relationship of the feature points in the initial ROI; configure differentiated quantization parameters for the corrected main facial area, secondary facial area and background area respectively; dynamically allocate three-layer regional bit rates based on the real-time network bandwidth; output the coded frame of the current frame and the associated ROI level metadata; wherein the ROI level metadata includes: the boundary coordinates of the main facial area, secondary facial area and background area; the quantization parameter offset value corresponding to each area; and the bit rate allocation weight value for each area.

[0184] It should be understood that Figure 4 The scene diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 4 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0185] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0186] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0187] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0188] like Figure 5As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0189] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0190] The computing unit 501 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the video encoding method. For example, in some embodiments, the video encoding method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the video encoding method described above can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the video encoding method in any other appropriate manner (eg, by means of firmware).

[0191] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0192] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0193] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0194] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0195] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0196] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0197] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0198] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A video encoding method, comprising: Obtain a reference frame adjacent to the current frame and having established an initial region of interest (ROI) hierarchical division; Multiplexing the motion vector generated by the reference frame in the motion compensation temporal filtering process to determine the ROI position prediction value of the current frame; Extracting facial feature points of the current frame, and correcting the ROI position prediction value based on the spatial topological relationship of the feature points in the initial ROI; configuring differential quantization parameters for the corrected primary facial region, secondary facial region, and background region respectively; Dynamically allocate three-layer regional bit rates based on real-time network bandwidth; Outputting the coded frame of the current frame and associated ROI-level metadata; wherein the ROI-level metadata includes: boundary coordinates of the primary facial region, the secondary facial region, and the background region; quantization parameter offset values ​​corresponding to each region; and bit rate allocation weight values ​​for each region.

2. The method according to claim 1, wherein The establishment of the initial ROI includes: A lightweight face detection algorithm is used to locate the face area in the first frame of the video; The facial area is divided into three layers according to the key points of the facial features: the primary facial area, the secondary facial area and the background area; wherein the primary facial area includes the smallest rectangular area covering the eyebrows, eyes, nose and mouth; the secondary facial area includes the annular area covering the cheeks and chin; and the background area includes the remaining area outside the face.

3. The method according to claim 2, wherein: The method further comprises: For the multiple face regions detected in the first frame, respectively calculate the central region weight and area weight of each face; the central region weight reflects the proximity between the center position of the face and the center of the image; the area weight reflects the proportion of the face region to the total area of ​​the image; Calculate the comprehensive score of each face based on the preset central weight factor and area weight factor; Selecting the face with the highest comprehensive score as the main tracking target, and determining the face area of ​​the main tracking target as the main facial area; The remaining face regions that are not selected as the primary tracking targets are merged into the secondary face region.

4. The method according to claim 1, wherein The method further comprises: Obtaining an intra-frame prediction cost and an inter-frame prediction cost of the current frame; In response to the intra-frame prediction cost not exceeding the product of a preset coefficient and the inter-frame prediction cost, determining that a scene cut occurs; In response to the scene switching, the current ROI tracking process is interrupted, the current frame is used as a reference frame of the new scene, and the ROI establishment operation is re-executed on the reference frame.

5. The method according to claim 1, wherein The step of multiplexing the motion vector generated by the reference frame in the motion compensation time domain filtering process to determine the ROI position prediction value of the current frame includes: Dividing the physical ROI of the reference frame into a set of macroblocks according to a standard macroblock size; Obtaining the motion vector of each macroblock calculated during the motion compensation temporal filtering process; A motion vector aggregation algorithm is used to generate a predicted value of the ROI position of the current frame.

6. The method according to claim 5, wherein: The method further comprises: Calculating the consistency of the motion vectors of each macroblock in the physical ROI of the reference frame as a confidence score; In response to the confidence score exceeding a preset occlusion determination threshold, determining that local occlusion occurs in the corresponding macroblock area; For the macroblock area where partial occlusion occurs, the feature point matching method is used to update its position information in the current frame.

7. The method according to claim 1, wherein The correcting the ROI position prediction value based on the spatial topological relationship of the feature points in the initial ROI includes: Calculating the expected position coordinates of each feature point in the current frame based on the spatial topological relationship of the feature points recorded in the initial ROI and the ROI position prediction value; Obtaining the actual position coordinates of the corresponding feature points in the current frame through a facial feature point detection algorithm; Calculate spatial transformation parameters based on the correspondence between the expected position coordinates and the actual position coordinates; The spatial transformation parameters are applied to modify the ROI position prediction value.

8. The method according to claim 1, wherein The method of dynamically allocating the three-layer regional bit rate according to the real-time network bandwidth includes: In response to detecting that the real-time network bandwidth is not less than a preset bandwidth threshold, controlling the proportion of the bit rate allocated to the main facial region to the total bit rate to be not less than a preset ratio threshold; In response to detecting that the real-time network bandwidth is lower than a preset bandwidth threshold, the bit rate of the main facial area is guaranteed not to be lower than a preset minimum guaranteed bit rate, and the bit rate of the background area is compressed to a first compression coefficient range of its original reference value.

9. The method according to claim 8, wherein The method further comprises: When insufficient bandwidth is triggered for the first time, compressing the bit rate of the background area to a specific value within the first compression coefficient range; If the bandwidth continues to decrease and the decrease per unit time exceeds a preset rate threshold, compressing the bit rate of the secondary facial region to within a second compression coefficient range; During the compression process, the bit rate of the main facial area is guaranteed to be no less than the protection redundancy coefficient times the preset minimum guaranteed bit rate.

10. The method according to claim 1, wherein Differentiated quantization parameters are configured for the corrected main facial area, secondary facial area, and background area, including: Setting a minimum quantization parameter offset value for the main facial area so that its quantization parameter is lower than the encoder base quantization parameter; enabling the first prediction mode combination and allocating the highest priority encoding resources; Set a medium quantization parameter offset value for the secondary facial area so that its quantization parameter is equal to or slightly higher than the encoder base quantization parameter; enable the second prediction mode combination and allocate normal priority coding resources; The highest quantization parameter offset value is set for the background area, the third prediction mode combination is enabled, and the lowest priority coding resources are allocated.

11. A video encoding apparatus, comprising: A reference frame acquisition module is used to acquire a reference frame adjacent to the current frame and for which an initial ROI hierarchical division has been established; A position prediction module, configured to multiplex the motion vector generated by the reference frame during the motion compensation temporal filtering process to determine the ROI position prediction value of the current frame; A position correction module, configured to extract facial feature points of the current frame and correct the ROI position prediction value based on the spatial topological relationship of the feature points in the initial ROI; A parameter configuration module, configured to configure differential quantization parameters for the corrected primary facial region, secondary facial region, and background region respectively; The bitrate allocation module is used to dynamically allocate the bitrate of three-layer regions according to the real-time network bandwidth; A data output module is configured to output the coded frame of the current frame and associated ROI-level metadata; wherein the ROI-level metadata includes: boundary coordinates of the primary facial region, the secondary facial region, and the background region; quantization parameter offset values ​​corresponding to each region; and bit rate allocation weight values ​​for each region.

12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are for causing a computer to perform a method according to any one of claims 1-10.

14. A computer program product comprising a computer program stored on a storage medium, the computer program implementing the method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Cited By

  • Unmanned aerial vehicle edge video processing method

    CN121442095A

  • Semantic hierarchical multi-path video return method and system for vehicle insurance evidence collection and claim settlement

    CN121585825A

  • Image storage method, image storage device and computer storage medium

    CN121705445A