Video communication data efficient compression method and system based on artificial intelligence

Through the video compression method based on artificial intelligence, the head posture depth map is generated and the deformation-sensitive area is identified, and the number of encoded bits is dynamically allocated, which solves the problems of face distortion and unreasonable bit allocation in video compression, and improves the quality and efficiency of video communication.

CN120547348APending Publication Date: 2025-08-26SHENZHEN BANGLIAN TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510773753.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When existing video compression technology faces large head movements and rapid expression changes, it cannot accurately identify facial texture stretching, resulting in distortion of key face areas and unreasonable bit allocation, affecting the video quality and character recognizability.

Method used

Using an artificial intelligence-based method, a head attitude depth map is generated through monocular depth estimation, deformation-sensitive areas are identified, and the encoded bit count is dynamically allocated based on the deformation entropy value, and a deformation-sensitive area mask is constructed to achieve local deformation-driven compression.

Benefits of technology

It improves the visual quality and compression efficiency of the face area in video communication, reduces the risk of distortion in key areas, and improves character recognizability and expression accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547348A_ABST
    Figure CN120547348A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video compression, in particular to a video communication data efficient compression method and system based on artificial intelligence, and the method comprises the following steps: extracting a head posture depth map of a video frame through a monocular depth estimation network, the head posture depth map comprising three-dimensional space coordinates of face key points; calculating a face curvature change gradient according to the head posture depth map, and outputting a deformation sensitive area mask; generating a local deformation entropy value based on the deformation sensitive area mask, dynamically allocating a coding bit number according to the local deformation entropy value, and generating a compressed video stream; wherein the higher the deformation entropy value is, the more the bit number is allocated to the region. According to the method, the visual quality and the anti-artifact capability of the key region are effectively improved, and the distortion risk of the traditional constant compression on the face region is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video compression technology, and in particular to an artificial intelligence-based efficient video communication data compression method and system. Background Art

[0002] With the prevalence of remote work, online education, and video social networking, the frequency of video communication usage continues to grow across various terminal devices, especially on mobile and embedded devices, which place higher demands on efficient compression and low-bandwidth transmission. Existing video compression technologies (such as H.264 / AVC and H.265 / HEVC) primarily rely on unified macroblock partitioning and fixed or semi-adaptive quantization parameter (QP) control strategies. When faced with large head movements and rapid changes in facial expressions, the following problems are prone to occur:

[0003] Severe distortion in moving areas: Traditional methods fail to accurately identify facial texture stretching and compression areas caused by changes in head posture. This results in distortion such as blurring and mosaics in key facial areas under constant compression parameters, seriously affecting video quality and character recognizability.

[0004] Unreasonable bit allocation: Most existing compression frameworks use intra-frame or inter-frame overall complexity assessment to control bit rate, which is unable to dynamically tilt bit resources towards important local areas (such as faces, eyes, and mouths), resulting in over-compression of key areas and waste of background area resources.

[0005] Lack of ability to model three-dimensional information in monocular videos: Although some compression optimization methods have introduced perceptual attention mechanisms based on deep learning, most methods are unable to restore the true three-dimensional structure of the face without a depth camera, resulting in inaccurate judgment of facial area deformation and affecting the effectiveness of the compression optimization strategy. Summary of the Invention

[0006] The present invention provides an efficient compression method and system for video communication data based on artificial intelligence. The efficient compression method combines artificial intelligence facial posture estimation, curvature perception and content-adaptive compression strategy. It can not only identify texture deformation caused by head movement at a low cost, but also realize differentiated coding control, thereby improving the visual quality and compression efficiency of the facial area in video communication.

[0007] An efficient compression method for video communication data based on artificial intelligence, comprising the following steps:

[0008] S1. Head pose depth map generation: extracting a head pose depth map of a video frame through a monocular depth estimation network, wherein the head pose depth map includes the three-dimensional spatial coordinates of facial key points;

[0009] S2. Identifying deformation-sensitive areas: Calculating the facial curvature gradient based on the head posture depth map and outputting a deformation-sensitive area mask, wherein the deformation-sensitive area mask marks pixel areas where texture stretching / compression due to head rotation occurs;

[0010] S3. Deformation entropy value driven compression: Generate a local deformation entropy value based on the deformation sensitive area mask, dynamically allocate coding bits according to the local deformation entropy value, and generate a compressed video stream; wherein the area with a higher deformation entropy value has more allocated bits.

[0011] Optionally, the S1 specifically includes:

[0012] S11, spatial feature encoding: uses lightweight MobileNetV3 as the encoder to downsample the input RGB video frame and output a multi-channel feature tensor;

[0013] S12, three-dimensional coordinate regression: The feature tensor output by the encoder is input into the deconvolution decoder, and a heat map including multiple facial key points is generated through the deconvolution layer and jump connection in the deconvolution decoder. Each key point is associated with a three-dimensional spatial coordinate (x i ,y i ,z i ); where (x i ,y i ) represents the pixel coordinates of the i-th key point in the image plane, z i Represents the depth distance of the i-th key point relative to the camera;

[0014] S13, posture parameter fusion: input the three-dimensional space coordinates into a fully connected layer, and output multi-degree-of-freedom head posture parameters, including three Euler angles and a three-dimensional translation vector;

[0015] S14, depth map synthesis: construct a head posture rotation matrix based on the head posture parameters, map the three-dimensional key points to the two-dimensional image plane through perspective projection transformation, and use bilinear interpolation filling to generate a head posture depth map, where each pixel value of the posture depth map represents the normalized depth of the point in the local coordinate system of the head.

[0016] Optionally, the facial key points include eyebrow area key points, eye area key points, nose area key points, mouth area key points, face contour area key points and eye area key points.

[0017] Optionally, in the depth map synthesis: the three-dimensional space coordinates (x i ,y i ,z i ) The coordinates after posture transformation R·p i +t is projected onto the two-dimensional image plane and expressed as follows using the perspective projection relationship:

[0018]

[0019] Among them, (x′ i ,y′ i ,z′ i )=R·p i +t;(u i ,v i ) represents the pixel coordinates of the key point on the synthetic image plane, f x ,f y is the focal length of the camera intrinsic parameter, c x ,c y represents the camera principal point position, p i represents the coordinate vector of the i-th three-dimensional key point, and t represents the three-dimensional translation vector of the head center point.

[0020] Optionally, the S2 specifically includes:

[0021] S21, curvature field construction: performing Gaussian curvature calculation on the head posture depth map to generate a facial curvature distribution map;

[0022] S22, gradient sensitivity analysis: performing Sobel gradient detection on the facial curvature distribution map to calculate the curvature change gradient amplitude;

[0023] S23, deformation area marking: marking the connected areas whose gradient amplitude is greater than the deformation judgment threshold as deformation sensitive areas, and generating a binary mask; the deformation judgment threshold is adaptively adjusted according to the head deflection angle θ.

[0024] Optionally, the curvature change gradient amplitude calculation in S22 includes executing a two-dimensional Sobel operator on the curvature distribution map to calculate the curvature gradient components G in the x-axis and y-axis directions respectively. x ,G y , and synthesize the overall curvature change gradient amplitude, expressed as:

[0025] in, By calculating the Sobel convolution kernel in the x direction, By calculating the Sobel convolution kernel in the y direction, G represents the intensity of the curvature change in a certain area of ​​the face, which is used to measure the severity of local deformation.

[0026] Optionally, the deformation judgment threshold in S23 is calculated as: τ d =τ0·(1+γ|θ|); where τ0 is the static reference threshold, γ is the angle sensitivity factor used to control the magnification of the angle deflection to the threshold, θ is the head yaw angle, τ dIt is an adaptive deformation judgment threshold. The larger the yaw angle, the higher the value, which improves the detection coverage of sensitive areas.

[0027] Optionally, the S3 specifically includes:

[0028] S31, mask-guided segmentation: dividing the video frame into rectangular macroblocks, and performing subsequent entropy calculation only on the macroblocks with the mask mark of the deformation sensitive area as 1;

[0029] S32, local deformation entropy value calculation: for each macroblock covered by the mask, extract the brightness component Y of the YUV color space and calculate the local deformation entropy value;

[0030] S33, dynamic bit allocation: allocating actual coding bits to each macroblock according to the local deformation entropy value of the macroblock;

[0031] S34, uses an H.264 encoder to perform a compression process to generate a compressed video stream.

[0032] Optionally, the S34 specifically includes:

[0033] For macroblocks within the mask area, the quantization parameters are adjusted according to the number of allocated bits to achieve variable bit rate coding;

[0034] For non-masked areas, constant compression is performed using a fixed quantization parameter;

[0035] After all macroblocks are encoded, the final compressed video stream is output.

[0036] An artificial intelligence-based efficient video communication data compression system, used to implement the above-mentioned efficient video communication data compression method, includes the following:

[0037] Video frame acquisition module: used to obtain the input video stream and extract the continuous RGB frame sequence as the basic data source for subsequent encoding processing;

[0038] The posture depth estimation module is used to perform three-dimensional head posture recognition and depth map generation on the video frame, specifically including:

[0039] Spatial feature encoding unit: uses a lightweight MobileNetV3 network structure to extract features and spatially downsample the input RGB frame, and outputs a multi-channel intermediate feature tensor;

[0040] Keypoint regression unit: Based on the deconvolution decoding structure and skip connection mechanism, it outputs a two-dimensional heat map of facial key points and their corresponding three-dimensional spatial coordinates;

[0041] Posture parameter extraction unit: Build a fully connected neural network to regress the six-degree-of-freedom posture parameters of the head, including yaw angle, pitch angle, roll angle and three-dimensional translation vector;

[0042] Depth map generation unit: Constructs a rotation matrix based on the posture parameters, maps the three-dimensional key points to the image plane through perspective projection and bilinear interpolation, and generates a head posture depth map.

[0043] Deformation recognition module: used to identify facial texture stretching or compression areas caused by head movement and generate deformation-sensitive mask images;

[0044] Entropy calculation and bit allocation module: used to analyze the complexity of deformation areas by block and dynamically adjust the allocation of coding resources. Different numbers of coding bits are allocated to each macroblock based on the deviation between the local entropy value and the average entropy value.

[0045] Compression encoding module: used to perform the actual video compression encoding process and output the final compressed video stream.

[0046] Beneficial effects of the present invention:

[0047] This paper constructs a lightweight MobileNetV3 feature encoder and deconvolution decoding structure and introduces a fusion mechanism for the head's six-degree-of-freedom posture parameters. It achieves the generation of high-quality head posture depth maps from ordinary RGB video frames without the need for a depth camera. This depth map has real spatial structure information, providing a data basis for subsequent deformation recognition and region-aware compression, and solving the problem of insufficient representation of facial motion areas in monocular video compression schemes in existing technologies.

[0048] This invention proposes a "deformation-sensitive area mask" constructed based on Gaussian curvature and gradient amplitude, combined with the adaptive adjustment threshold of the head yaw angle to identify facial deformation areas caused by head movement. On this basis, the local texture complexity is quantified by the brightness entropy value, driving the dynamic allocation of bit resources on demand, effectively improving the visual quality and anti-artifact ability of key areas, and avoiding the distortion risk of traditional constant QP compression on the face area.

[0049] The present invention uses mask guidance to perform entropy analysis and dynamic bit allocation only on the significantly deformed areas, and adopts a fixed quantization compression strategy for non-critical areas, forming a hierarchical compression model of significant area enhancement + background simplification. Under the same bit rate control, the present invention improves the PSNR of the key head area, significantly improving the face recognition and expression accuracy of video communications in scenarios such as remote conferencing, online teaching, and telemedicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 Schematic diagram of the compression method according to an embodiment of the present invention;

[0052] Figure 2 Schematic diagram of the functional modules of the compression system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art may also adopt other alternative methods to implement some known technologies. Moreover, the accompanying drawings are only for describing the embodiments in more detail and are not intended to specifically limit the present invention.

[0054] like Figure 1 As shown, an efficient compression method for video communication data based on artificial intelligence includes the following steps:

[0055] S1. Head pose depth map generation: The head pose depth map of the video frame is extracted through a monocular depth estimation network. The head pose depth map includes the three-dimensional spatial coordinates of facial key points.

[0056] S2, deformation-sensitive area identification: Calculate the facial curvature change gradient based on the head posture depth map and output the deformation-sensitive area mask. The deformation-sensitive area mask marks the pixel area where the texture is stretched / compressed due to head rotation;

[0057] S3, deformation entropy driven compression: Generate local deformation entropy values ​​based on deformation-sensitive area masks, dynamically allocate coding bits according to the local deformation entropy values, and generate compressed video streams; the higher the deformation entropy value, the more bits are allocated to the area.

[0058] S1 specifically includes:

[0059] S11, spatial feature encoding: Use lightweight MobileNetV3 as the encoder to perform spatial downsampling on the input RGB video frame and output an intermediate feature tensor with 128 channels for subsequent key point positioning and depth reasoning.

[0060] S12, 3D coordinate regression: The feature tensor output by the encoder is input into the deconvolution decoder, and the ability to retain structural information is improved through multi-scale skip connections. A heat map set containing 68 facial key points is generated; the 3D coordinates associated with each key point are expressed as:

[0061] (x i ,y i ,z i ),i=1,2,…,68; where (x i ,y i ) represents the pixel coordinates of the i-th key point in the image plane, z i Indicates the depth distance of the i-th key point relative to the camera.

[0062] Facial key points refer to typical anatomical or geometric feature points on the face. The 68 common facial key points include:

[0063] Eyebrow area (10 points): 5 points on each eyebrow;

[0064] Eye area (12 o'clock): 6 o'clock for each eye;

[0065] Nose area (9 o'clock): 4 o'clock on the bridge of the nose + 5 o'clock on the nose wings;

[0066] Mouth area (20 points): outer lip 12 points + inner lip 8 points;

[0067] Facial contour area (17 o'clock): the area from the chin to the ear;

[0068] Eye point (1 point): the center point at the root of the nose;

[0069] These points refer to face annotation standards such as dlib or OpenFace and are widely used in expression analysis, pose estimation and face recognition.

[0070] S13, posture parameter fusion: The three-dimensional coordinates of the above 68 key points are input into the fully connected neural network, and the six-degree-of-freedom posture parameters of the head are output, including three Euler angles and a three-dimensional translation vector:

[0071] Yaw angle: θ y ;

[0072] Pitch angle: θ p ;

[0073] Roll angle: θ r ;

[0074] Translation vector:

[0075] Based on this, the head posture rotation matrix is ​​constructed: R = R y (θ y )·R p (θ p )·R r (θ r ); where R y (θ y) represents the rotation matrix around the y-axis, R p (θ p ) represents the rotation matrix around the x-axis, R r (θ r ) represents the rotation matrix around the z-axis.

[0076] S14, depth map synthesis: the three-dimensional coordinates of the key points (x i ,y i ,z i ) The coordinates after posture transformation R·p i +t is projected onto the two-dimensional image plane using the perspective projection relationship:

[0077]

[0078] Among them, (x′ i ,y′ i ,z′ i )=R·p i +t;(u i ,v i ) represents the pixel coordinates of the key point on the synthetic image plane, f x ,f y is the focal length of the camera internal parameter (unit pixel), c x ,c y Indicates the camera principal point position (unit pixel), p i Represents the coordinate vector of the i-th three-dimensional key point.

[0079] The depth value z′ of the sparse key point is converted into i Interpolation diffusion is performed in the image plane to generate a dense head posture depth map D(u,v), that is, the head depth value of the pixel point (u,v);

[0080] Finally, the synthesized depth map is normalized and the normalized depth value is calculated:

[0081] z′∈[0,1]; where: z represents the original depth value of the pixel in the depth map, z min ,z max Indicates the minimum and maximum depth values ​​of all key points in the current frame, used for normalization.

[0082] 1. Structure and feature extraction process of lightweight MobileNetV3 encoder:

[0083] 1. Network Structure: MobileNetV3 is a convolutional neural network built based on automatic structure search and lightweight module stacking. It has high computational efficiency and is suitable for mobile deployment. Its core components include:

[0084] Depthwise separable convolution: significantly reduces the number of parameters;

[0085] Squeeze-and-Excitation module (SE module): used for channel attention modeling;

[0086] h-swish activation function: improves nonlinear expression capabilities;

[0087] Residual connection and bottleneck structure: enhancing feature transfer stability.

[0088] 2. Feature extraction process: The input RGB video frame size is H×W×3. The main processing steps after the MobileNetV3 encoder are as follows:

[0089] Initial convolution and activation, output H / 2×W / 2×16;

[0090] Stack 5 to 7 Bottleneck modules for step-by-step downsampling and channel expansion, outputting H / 16×W / 16×128;

[0091] The final output is a spatial feature tensor with 128 channels for facial key points and depth information regression.

[0092] This output retains the local structure and edge information of the original video frame while compressing the computational complexity, and is the basis for subsequent key point positioning and posture estimation.

[0093] 2. Structure of deconvolution decoder and skip connection and key point heat map generation:

[0094] 1. Decoder structure design: The deconvolution decoder uses a symmetric upsampling structure to gradually restore the spatial features output by the encoder to a size similar to the original image (H / 2×W / 2 or higher):

[0095] It includes 34 transposed convolutional layers, and the number of channels in each layer is gradually reduced from 128 to 321;

[0096] Each deconvolution module is followed by batch normalization (BatchNorm) and ReLU nonlinear activation;

[0097] The final output tensor size is H / 2×W / 2×68, which is a 68-channel heat map set, with each channel corresponding to a facial key point.

[0098] 2. Multi-scale skip connection mechanism: To prevent high-level semantic features from losing their original spatial location information, a U-Net-style skip connection mechanism is adopted. Specifically:

[0099] Concatenate the feature maps output by the encoder's intermediate layers (such as the 2nd and 4th Bottleneck modules) with the deconvolution output of the corresponding decoding layers;

[0100] After splicing, 1×1 convolution is used to reduce the dimension and retain the local structure information;

[0101] Strengthen the response of the texture area around the key points and improve the accuracy of heat map prediction.

[0102] The decoder structure finally outputs a heat map for key point positioning and can further extract the three-dimensional spatial position of each key point.

[0103] 3. Fully connected neural network outputs head posture parameters:

[0104] 1. Network input design: The input is the three-dimensional spatial coordinate vector of 68 facial key points obtained in the decoding stage: This vector is flattened into a feature vector of length 204 and input into the fully connected network.

[0105] 2. Network structure diagram:

[0106] FC1 layer: input 204 dimensions → output 128 dimensions, fully connected layer + ReLU;

[0107] FC2 layer: 128-dimensional → output 64-dimensional, fully connected layer + BatchNorm + ReLU;

[0108] FC3 layer: 64-dimensional → output 6-dimensional, fully connected layer + Tanh;

[0109] 3. Output content: The output 6-dimensional vectors correspond to the six-degree-of-freedom posture parameters of the head:

[0110] (θ y ,θ p ,θ r ,t x ,t y ,t z ), where θ y ,θ p ,θ r Denote the yaw angle, pitch angle and roll angle respectively, t x ,t y ,t z Represents the three-dimensional translation vector of the head center point in the camera coordinate system. The output is subsequently used to generate the rotation matrix R and projection operation to achieve posture correction and depth map synthesis of the original key points.

[0111] S2 specifically includes:

[0112] S21, curvature field construction: Perform Gaussian curvature calculation on the head posture depth map to construct the facial curvature distribution map. The curvature value K of each pixel is defined as follows:

[0113] in, They represent the second-order partial derivatives, which are used to characterize the curvature of the surface at that point. The curvature value K is obtained by fitting a quadratic surface in the local window of the image and calculating its Hessian matrix.

[0114] S22, Gradient Sensitivity Analysis: Execute a two-dimensional Sobel operator on the curvature distribution map to calculate the curvature gradient components G in the x-axis and y-axis directions respectively. x ,G y , and synthesize the overall gradient amplitude:

[0115] in, By calculating the Sobel convolution kernel in the x direction, By calculating the Sobel convolution kernel in the y direction, G represents the intensity of the curvature change in a certain area of ​​the face, which is used to measure the severity of local deformation.

[0116] S23, deformation area marking: mark all deformation areas in the gradient amplitude map G that are greater than the deformation judgment threshold τ d The connected areas of are extracted and marked as deformation-sensitive areas to form a binary mask map:

[0117] If G(x,y)>τ d , then the pixel is marked as a sensitive area (the value is set to 1);

[0118] Otherwise, it is marked as a non-sensitive area (the value is set to 0).

[0119] Among them, the deformation judgment threshold τ d Calculated as: τ d =τ0·(1+γ|θ|); τ0 is the static reference threshold (the preset constant is 0.15), γ is the angle sensitivity factor, which is used to control the magnification of the angle deflection to the threshold, with a value of 0.01 to 0.05, and θ is the head yaw angle (θ output by S1). y )τ d This is an adaptive deformation judgment threshold. The larger the yaw angle, the higher the value, improving the detection coverage of sensitive areas. The output binary mask image is used as the deformation sensitive area mask to drive the subsequent dynamic allocation of compression bits.

[0120] The Gaussian curvature K is obtained by fitting a quadratic surface to the normalized depth map z′(x,y) in a local window of the image and calculating its Hessian matrix, as follows:

[0121] 1. Fitting the quadratic surface: In the neighborhood window (e.g. 5×5) of each pixel (x0, y0), the local depth function is fitted using the least squares method: z′(x, y)≈ax 2 +bxy+cy 2 +dx+ey+f; where a, b, c, d, e, and f are the coefficients to be fitted, representing the shape information of the surface in this local area.

[0122] 2. Construct the Hessian matrix. From the fitted surface, the Hessian matrix composed of the second-order partial derivatives can be directly calculated:

[0123]

[0124] 3. Calculate the Gaussian curvature. The Gaussian curvature K is equal to the determinant of the Hessian matrix, that is:

[0125] K=det(H)=(2a)(2c)-b 2 =4ac-b 2 ; This value reflects the superposition of the bending characteristics of the local surface in two orthogonal directions and is a stable measurement indicator of facial deformation.

[0126] To calculate and That is, the gradient components of the Gaussian curvature map K(x,y) in the x and y directions, and a two-dimensional convolution operation is performed on K using a standard Sobel convolution kernel (3×3). The details are as follows:

[0127] 1. Definition of Sobel convolution kernel:

[0128] Sobel convolution kernel in the x direction (to find G x ):

[0129]

[0130] Sobel convolution kernel in the y direction (to find G y ):

[0131]

[0132] 2. Convolution operation: Let K(x,y) represent the Gaussian curvature image and perform the following two-dimensional convolution:

[0133] Horizontal gradient: G x (x,y)=S x *K(x,y);

[0134] Vertical gradient: G y (x,y)=S y *K(x,y); where * represents a two-dimensional discrete convolution operation (equivalent to a weighted sliding window in image processing).

[0135] 3. Gradient amplitude synthesis: Calculate the gradient amplitude of curvature change: The gradient amplitude map is the basis for sensitive area detection and is used for subsequent dynamic threshold τ d Compare and generate deformation mask.

[0136] S3 specifically includes:

[0137] S31, mask-guided segmentation:

[0138] Each frame of video image is divided into macroblocks of 16×16 pixels. Subsequent processing operations are performed only on the macroblocks marked as 1 in the deformation-sensitive area mask in the previous step. The areas not covered by the mask are not involved in the entropy calculation.

[0139] S32, local deformation entropy calculation: For each macroblock k covered by the mask, extract the corresponding brightness component Y in the YUV color space, and calculate its local deformation entropy according to the following formula:

[0140] Among them, p i represents the probability that a pixel with brightness value i∈[0,255] appears in macroblock k, It represents the local deformation entropy value of the kth macroblock, reflecting the complexity of its brightness texture and indirectly representing the deformation intensity caused by the change of head posture.

[0141] S33, dynamic bit allocation: according to the local deformation entropy value of each macroblock Allocate the actual number of coding bits B for this macroblock k , the calculation rules are as follows:

[0142]

[0143] Among them, B base is the standard benchmark bit number, the initial coding bit set for all macroblocks, E avg is the average value of the local deformation entropy of all mask macroblocks in the current frame, β m is a high entropy enhancement adjustment factor used to amplify the bit allocation in high texture complexity areas; m A low entropy attenuation adjustment factor used to compress the number of bits in simple texture areas.

[0144] S34, compressed stream generation: The H.264 encoder is used to perform the compression process. The specific strategy is as follows:

[0145] For macroblocks within the mask area, the number of bits allocated is B k Adjust the quantization parameter (QP) to achieve variable bit rate coding;

[0146] For non-masked areas, constant compression is performed using a fixed quantization parameter QP = 32;

[0147] After all macroblocks are encoded, the final compressed video stream is output.

[0148] like Figure 2 As shown, an artificial intelligence-based efficient video communication data compression system is used to implement the above compression method, including the following modules:

[0149] Video frame acquisition module: used to obtain the input video stream and extract the continuous RGB frame sequence as the basic data source for subsequent encoding processing;

[0150] Pose and Depth Estimation Module: This module is used to perform 3D head pose recognition and depth map generation on video frames. It includes:

[0151] Spatial feature encoding unit: uses a lightweight MobileNetV3 network structure to extract features and spatially downsample the input RGB frame, and outputs a multi-channel intermediate feature tensor;

[0152] Keypoint regression unit: Based on the deconvolution decoding structure and skip connection mechanism, it outputs a two-dimensional heat map of facial key points and their corresponding three-dimensional spatial coordinates;

[0153] Posture parameter extraction unit: Build a fully connected neural network to regress the six-degree-of-freedom posture parameters of the head, including yaw angle, pitch angle, roll angle and three-dimensional translation vector;

[0154] Depth map generation unit: Constructs a rotation matrix based on the posture parameters, maps the three-dimensional key points to the image plane through perspective projection and bilinear interpolation, and generates a head posture depth map.

[0155] Deformation recognition module: used to identify facial texture stretching or compression areas caused by head movement and generate deformation-sensitive mask images;

[0156] Entropy calculation and bit allocation module: used to analyze the complexity of deformation areas by block and dynamically adjust the allocation of coding resources. Different numbers of coding bits are allocated to each macroblock based on the deviation between the local entropy value and the average entropy value.

[0157] Compression encoding module: used to perform the actual video compression encoding process and output the final compressed video stream.

[0158] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.

[0159] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. An efficient compression method for video communication data based on artificial intelligence, characterized in that: The following steps are involved: S1. Head pose depth map generation: extracting a head pose depth map of a video frame through a monocular depth estimation network, wherein the head pose depth map includes the three-dimensional spatial coordinates of facial key points; S2. Identifying deformation-sensitive areas: Calculating the facial curvature gradient based on the head posture depth map and outputting a deformation-sensitive area mask, wherein the deformation-sensitive area mask marks pixel areas where texture stretching / compression due to head rotation occurs; S3, deformation entropy value driven compression: generating a local deformation entropy value based on the deformation sensitive area mask, dynamically allocating coding bits according to the local deformation entropy value, and generating a compressed video stream; The region with higher deformation entropy value is allocated more bits.

2. The method for efficiently compressing video communication data based on artificial intelligence according to claim 1, characterized in that: Said S1 specifically includes: S11, spatial feature encoding: uses lightweight MobileNetV3 as the encoder to downsample the input RGB video frame and output a multi-channel feature tensor; S12, three-dimensional coordinate regression: The feature tensor output by the encoder is input into the deconvolution decoder, and a heat map including multiple facial key points is generated through the deconvolution layer and jump connection in the deconvolution decoder. Each key point is associated with a three-dimensional spatial coordinate (x i ,y i ,z i ); where (x i ,y i ) represents the pixel coordinates of the i-th key point in the image plane, z i Represents the depth distance of the i-th key point relative to the camera; S13, posture parameter fusion: input the three-dimensional space coordinates into a fully connected layer, and output multi-degree-of-freedom head posture parameters, including three Euler angles and a three-dimensional translation vector; S14, depth map synthesis: construct a head posture rotation matrix based on the head posture parameters, map the three-dimensional key points to the two-dimensional image plane through perspective projection transformation, and use bilinear interpolation filling to generate a head posture depth map, where each pixel value of the posture depth map represents the normalized depth of the point in the local coordinate system of the head.

3. The method for efficiently compressing video communication data based on artificial intelligence according to claim 1, characterized in that: The facial key points include eyebrow area key points, eye area key points, nose area key points, mouth area key points, face contour area key points and eye area key points.

4. The method for efficiently compressing video communication data based on artificial intelligence according to claim 2, characterized in that: In the depth map synthesis: the three-dimensional space coordinates (x i ,y i ,z i ) The coordinates after posture transformation R·p i +t is projected onto the two-dimensional image plane and expressed as follows using the perspective projection relationship: Among them, (x′ i ,y′ i ,z′ i )=R·p i +t;(u i ,v i ) represents the pixel coordinates of the key point on the synthetic image plane, f x ,f y is the focal length of the camera internal parameter, c x ,c y Indicates the position of the camera's principal point, p i represents the coordinate vector of the i-th three-dimensional key point, and t represents the three-dimensional translation vector of the head center point.

5. The method for efficiently compressing video communication data based on artificial intelligence according to claim 1, characterized in that: The S2 specifically includes: S21, curvature field construction: performing Gaussian curvature calculation on the head posture depth map to generate a facial curvature distribution map; S22, gradient sensitivity analysis: performing Sobel gradient detection on the facial curvature distribution map to calculate the curvature change gradient amplitude; S23, deformation area marking: marking the connected areas whose gradient amplitude is greater than the deformation judgment threshold as deformation sensitive areas, and generating a binary mask; the deformation judgment threshold is adaptively adjusted according to the head deflection angle θ.

6. The method for efficiently compressing video communication data based on artificial intelligence according to claim 5, characterized in that: The curvature change gradient amplitude calculation in S22 includes executing a two-dimensional Sobel operator on the curvature distribution map to calculate the curvature gradient components G in the x-axis and y-axis directions respectively. x ,G y , and synthesize the overall curvature change gradient amplitude, expressed as: in, By calculating the Sobel convolution kernel in the x direction, By calculating the Sobel convolution kernel in the y direction, G represents the intensity of the curvature change in a certain area of ​​the face, which is used to measure the severity of local deformation.

7. The method for efficiently compressing video communication data based on artificial intelligence according to claim 5, characterized in that: The deformation judgment threshold in S23 is calculated as: d =τ0·(1+γ|θ|); where τ0 is the static reference threshold, γ is the angle sensitivity factor used to control the magnification of the angle deflection to the threshold, θ is the head yaw angle, τ d It is an adaptive deformation judgment threshold. The larger the yaw angle, the higher the value, which improves the detection coverage of sensitive areas.

8. The method for efficiently compressing video communication data based on artificial intelligence according to claim 1, characterized in that: The S3 specifically includes: S31, mask-guided segmentation: dividing the video frame into rectangular macroblocks, and performing subsequent entropy calculation only on the macroblocks with the mask mark of the deformation sensitive area as 1; S32, local deformation entropy value calculation: for each macroblock covered by the mask, extract the brightness component Y of the YUV color space and calculate the local deformation entropy value; S33, dynamic bit allocation: allocating actual coding bits to each macroblock according to the local deformation entropy value of the macroblock; S34, uses an H.264 encoder to perform a compression process to generate a compressed video stream.

9. The method for efficiently compressing video communication data based on artificial intelligence according to claim 8, characterized in that: The S34 specifically includes: For macroblocks within the mask area, the quantization parameters are adjusted according to the number of allocated bits to achieve variable bit rate coding; For non-masked areas, constant compression is performed using a fixed quantization parameter; After all macroblocks are encoded, the final compressed video stream is output.

10. An artificial intelligence-based video communication data efficient compression system, used to implement the artificial intelligence-based video communication data efficient compression method according to any one of claims 1 to 9, characterized in that: Includes the following modules: Video frame acquisition module: used to obtain the input video stream and extract the continuous RGB frame sequence as the basic data source for subsequent encoding processing; The posture depth estimation module is used to perform three-dimensional head posture recognition and depth map generation on the video frame, specifically including: Spatial feature encoding unit: uses a lightweight MobileNetV3 network structure to extract features and spatially downsample the input RGB frame, and outputs a multi-channel intermediate feature tensor; Keypoint regression unit: Based on the deconvolution decoding structure and skip connection mechanism, it outputs a two-dimensional heat map of facial key points and their corresponding three-dimensional spatial coordinates; Posture parameter extraction unit: Build a fully connected neural network to regress the six-degree-of-freedom posture parameters of the head, including yaw angle, pitch angle, roll angle and three-dimensional translation vector; Depth map generation unit: Constructs a rotation matrix based on the posture parameters, maps the 3D key points to the image plane through perspective projection and bilinear interpolation, and generates a head posture depth map; Deformation recognition module: used to identify facial texture stretching or compression areas caused by head movement and generate deformation-sensitive mask images; Entropy calculation and bit allocation module: used to analyze the complexity of deformation areas by block and dynamically adjust the allocation of coding resources. Different numbers of coding bits are allocated to each macroblock based on the deviation between the local entropy value and the average entropy value. Compression encoding module: used to perform the actual video compression encoding process and output the final compressed video stream.

Citation Information

Cited By

  • Intelligent monitoring image recognition processing system for community emergency disposal

    CN121982656A

  • Intelligent monitoring image recognition processing system for community emergency disposal

    CN121982656B