Intelligent conference video frame dynamic coding method based on multi-mode semantic understanding
By employing a dynamic encoding method for intelligent video frames based on multimodal semantic understanding, and combining semantic association analysis of video and audio, this method solves the problem that single-modal visual features cannot accurately represent semantic activity in intelligent video conferencing. It enables precise speaker identification and optimized allocation of encoding resources, thereby improving the efficiency and accuracy of video streams.
Patent Information
- Application Number
- CN202512035945.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-02-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing intelligent video conferencing, analysis methods based on single-modal visual features cannot accurately represent the true semantic activity, resulting in the inability to accurately locate the core content of the meeting. The generated visual attention model contains a large number of false high-response areas, which cannot provide a reliable semantic localization basis for subsequent processing.
A dynamic encoding method for intelligent conferencing video frames based on multimodal semantic understanding is adopted. By constructing a multimodal semantic understanding model, the video frame content is subjected to deep semantic analysis. Combined with semantic correlation analysis of video stream and audio stream, the pronunciation region that is highly related to the current speech content is identified and located. The quantization bias value is calculated based on the semantic relevance for dynamic encoding.
It achieves accurate identification of the real speaker under multi-target interference, avoids misjudgment of background and non-speaker areas, optimizes the allocation of encoding resources, and ensures the efficiency of video stream and semantic understanding capabilities.
Smart Images

Figure CN121585864A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, in particular to an intelligent conference video frame dynamic coding method based on multi-modal semantic understanding. BACKGROUND
[0002] In the visual analysis task of intelligent video conference and human-computer interaction scene, constructing an accurate visual attention model or saliency map is a key prerequisite for realizing scene understanding and subsequent data processing. Existing scene analysis technologies usually use recognition methods based on single-modal visual features, such as using a face detection algorithm to locate the positions of participants, or using an optical flow method to capture the motion area in the picture, and determining the high attention area in the image by calculating the brightness contrast, edge density or motion vector amplitude of the pixels; However, this analysis method relying only on shallow visual features has the defect of being unable to accurately represent the real semantic activity in a complex conference scene; the existing technology is mainly based on the assumption that visual saliency is equivalent to semantic importance, that is, the area in the picture where the face is clearer and the motion amplitude is larger is considered more important. However, in actual application, this assumption is often no longer valid, for example, when there are multiple participants in the conference, the face of a non-speaker may be determined by the algorithm as a high saliency area because of strong light or close distance; or irrelevant objects in the background may be misidentified as active areas because they produce large optical flow modulus; on the contrary, the real speaker may be ignored because of small motion amplitude; This mapping deviation between shallow visual features and deep semantic attributes makes it difficult for existing visual analysis technologies to accurately lock the real conference core content under multi-target interference, resulting in the generated visual attention model containing a large number of false high response areas, which cannot provide reliable semantic positioning basis for subsequent accurate image processing; Therefore, an intelligent conference video frame dynamic coding method based on multi-modal semantic understanding is proposed. SUMMARY
[0003] The purpose of the present application is to provide an intelligent conference video frame dynamic coding method based on multi-modal semantic understanding, which realizes conference video frame dynamic coding by constructing a multi-modal semantic understanding model to perform deep semantic analysis on the content of the video frame.
[0004] To achieve the above purpose, the present application provides the following technical solutions: The intelligent conference video frame dynamic coding method based on multi-modal semantic understanding comprises: real-time acquisition of video stream sequences and synchronous audio streams in a conference scene; performing content-based semantic analysis and decoupling processing on the video stream sequence, extracting key frames and subsequent frames, the key frames containing static facial structure and texture features; performing facial motion recognition on the subsequent frames, extracting sparse motion fields of facial key point geometric displacement and local affine transformation parameters, and performing semantic segmentation processing on the video frames, dividing into N candidate visual regions containing faces, mouth shapes and backgrounds; performing semantic content analysis on the synchronous audio stream, extracting audio semantic features; performing cross-modal semantic correlation analysis, calculating the semantic correlation degree of the distribution features of the sparse motion fields in the candidate visual regions and the audio semantic features, and identifying and positioning the pronunciation regions highly related to the current voice content based on the semantic correlation degree; According to the semantic correlation degree, the quantization bias value of each candidate visual region is calculated; the pronunciation region is applied with a negative bias, and the background region is applied with a positive bias, and is packaged into a variable rate video code stream.
[0005] Preferably, the process of extracting key frames and subsequent frames of the video stream sequence comprises: performing spatio-temporal topology analysis based on graph neural network on the video stream sequence, and constructing a semantic state topology graph containing participant identity nodes and facial pose nodes; monitoring the semantic state topology graph, when a quadrant switching event of facial pose category is detected, determining that a semantic state transition occurs at the current time, defining the current video frame as a key frame, the key frame containing static facial structure and texture features; in the time domain interval where the semantic state transition does not occur, the current video frame is defined as a subsequent frame.
[0006] Preferably, the sparse motion field extraction process is to use a feature recognition network to perform facial key point detection on the key frame and the subsequent frame respectively, and identify semantic key points; establishing the spatial correspondence relationship of the semantic key points between the key frame and the subsequent frame, generating facial key point geometric displacement by comparing the position difference of the same semantic key points in the two frames of images; extracting the local neighborhood features of each semantic key point, and identifying the local affine transformation parameters by comparing the feature distribution changes in the local neighborhood of the key frame and the subsequent frame; aggregate the facial key point geometric displacement and the local affine transformation parameters to construct the sparse motion field.
[0007] Preferably, the process of performing semantic segmentation processing on the video frame comprises: performing pixel-level semantic classification on the video frame to generate a semantic label map containing pixel category labels; Based on the semantic label map, perform region extraction and hierarchy definition: aggregate the set of pixels with background labels to construct a background region; identify the facial contour boundary in the foreground to define the facial region; and further locate the lip texture boundary within the facial region to define the mouth shape region; output N candidate visual regions containing the background, face, and mouth shape and their corresponding pixel indices.
[0008] Preferably, the specific process of generating the semantic label map containing pixel category labels is as follows: Pixel-level feature extraction is performed on the video frames to obtain a high-dimensional semantic feature vector corresponding to each pixel. The high-dimensional semantic feature vector is mapped to the semantic category space, and the classification confidence of each pixel belonging to the background, face, and mouth shape categories is calculated respectively. Perform the highest probability decision, select the category with the highest confidence as the unique label for that pixel, and traverse all pixels to generate the semantic label map.
[0009] Preferably, the cross-modal semantic association analysis is performed as follows: For each candidate visual region, the geometric displacement of facial key points in the sparse motion field within the region is aggregated using vector magnitudes to generate a visual signal; the audio semantic features are analyzed to extract the audio temporal envelope; the visual motion distribution sequence and the audio rhythm envelope sequence are subjected to waveform fitting analysis on the time axis to calculate the semantic correlation between the changes in the two sequences; By comparing the scores of different candidate visual regions, the region with the highest score and significant waveform synchronization characteristics is selected and identified as the pronunciation region that is highly related to the current speech content.
[0010] Preferably, the process of calculating the quantization bias value of each candidate visual region is as follows: Set the numerical mapping range of the global baseline quantization parameters and quantization offset values for the current video frame; The semantic relevance values are mapped to the numerical mapping interval, such that high semantic relevance corresponds to negative values within the interval, and low semantic relevance corresponds to positive values within the interval. For the pronunciation region, a negative vectorized bias value is obtained by matching within the numerical mapping interval based on semantic relevance; for the background region, a positive vectorized bias value is obtained by matching within the numerical mapping interval based on its semantic relevance.
[0011] Preferably, the variable bitrate video stream encapsulation process is as follows: The layered code stream structure including a static texture layer and a semantic driven layer is constructed, the key frame is encoded based on the global reference quantization parameter and written into the static texture layer, the global reference quantization parameter and the negative quantization bias value are added to obtain a refined quantization step, and the sparse motion field parameters in the pronunciation area are encoded, the global reference quantization parameter and the positive quantization bias value are added to obtain a rough quantization step, and the sparse motion field parameters in the background area are encoded, the encoded sparse motion field parameters of each area are written into the semantic driven layer, and the static texture layer, the semantic driven layer and the synchronous audio stream are multiplexed and packaged to generate the variable bit rate video code stream.
[0012] Compared with the prior art, the present application has the following beneficial effects: 1. By performing cross-modal semantic correlation analysis, the sparse motion field distribution in the visual area and the audio semantic features are waveform fitted on the time axis, so that the pronunciation area highly related to the current speech content is accurately located by using semantic correlation. The limitation that the existing method only relies on visual saliency for judgment in the background technology is effectively solved, the problem of misjudging non-speakers as high importance areas due to strong light, short distance or background object motion interference is avoided, accurate locking of the real speaker in a multi-participant scene is realized, and the deviation between visual saliency and semantic importance is corrected.
[0013] 2. By calculating the quantization bias value of each candidate visual area according to the semantic correlation, a negative bias is applied to the pronunciation area with high semantic correlation, and a positive bias is applied to the background area with low correlation. This dynamic encoding strategy solves the problem of unreasonable allocation of encoding resources caused by a large number of false high response areas. By reducing the bit rate proportion of irrelevant background and non-speaker areas while ensuring the quality details of the core speaker area, the present application can generate an efficient variable bit rate video code stream according to the real semantic activity, overcoming the defect of traditional technology that wastes data in non-core areas.
[0014] 3. The spatio-temporal topology analysis based on graph neural network is used to monitor semantic state transition, the video stream is decoupled into key frames containing static features and subsequent frames containing sparse motion fields, and the key frame is updated only when the facial posture switches quadrant. This mechanism solves the problem that the real speaker may be ignored due to small motion amplitude. By finely extracting the geometric displacement of facial key points and local affine transformation parameters, the present application can sensitively capture small facial semantic changes in complex conference scenes, and no longer simply rely on large magnitude of optical flow modulus, thereby improving the semantic understanding ability of dynamic conference scenes while ensuring spatio-temporal consistency. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1A flowchart of a method for dynamic coding of intelligent conference video frames based on multi-modal semantic understanding; Figure 2 A main flowchart of a method for dynamic coding of intelligent conference video frames based on multi-modal semantic understanding; Figure 3 A specific implementation flowchart of cross-modal semantic correlation analysis and quantitative bias calculation of the application. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the application.
[0017] Embodiment one: Please refer to Figure 1 Figure 2 and Figure 3 The application provides a method for dynamic coding of intelligent conference video frames based on multi-modal semantic understanding, and the technical solutions are as follows: Real-time acquisition of a video stream sequence and a synchronous audio stream in a conference scene; Content-based semantic analysis and decoupling processing are performed on the video stream sequence, key frames and subsequent frames are extracted, the key frames contain static facial structure and texture features, facial motion recognition is performed on the subsequent frames, a sparse motion field of facial key point geometric displacement and local affine transformation parameters is extracted, and semantic segmentation processing is performed on the video frames, which are divided into N candidate visual regions containing faces, mouth shapes and backgrounds; Semantic content analysis is performed on the synchronous audio stream, audio semantic features are extracted, cross-modal semantic correlation analysis is performed, distribution features of the sparse motion field in the candidate visual regions and semantic correlation degrees of the audio semantic features are calculated, and a pronunciation region highly related to current voice content is recognized and located based on the semantic correlation degrees; Quantitative bias values of each candidate visual region are calculated according to the semantic correlation degrees, a negative bias is applied to the pronunciation region, a positive bias is applied to the background region, and the video stream is encapsulated as a variable bit rate video stream.
[0018] The specific process of extracting the key frames and the subsequent frames of the video stream sequence includes: performing spatio-temporal topology analysis based on a graph neural network on the video stream sequence to construct a semantic state topology graph containing participant identity nodes and face posture nodes; the constructed graph structure has a node set containing identity nodes and face posture nodes of all participants, and an edge set including three types of connections: (1) identity-posture edges between identity nodes of the same participant and face posture nodes of the participant in the current frame, with a weight initialized as 1; (2) time continuity edges between identity nodes of homologous participants in adjacent two frames, with a weight initialized as 0.5 to reflect the time proximity; (3) spatial context edges between different participants in the same frame, with a weight of 0.2; the input features: the input feature of the identity node is a 512-dimensional face feature vector (extracted by a face recognition network such as ArcFace); the input feature of the face posture node is a 3-dimensional Euler angle vector. The backbone of the graph neural network contains two GAT layers, each containing 8 attention heads, with a hidden layer dimension of 64, and is supervised trained on a labeled dataset containing 10,000 multi-person conference video clips, with the labeled quadrant switching moments as the supervision signal. In the inference stage, the node features of the current and historical frames are input, and the key class vectors of the output nodes are calculated through forward propagation, wherein the output vector of the face posture node is mapped to the quadrant prediction probability distribution through a softmax classification layer, and the quadrant index with the maximum probability is the current predicted quadrant; by comparing the quadrant indexes of adjacent frames, if the index changes, it is determined as a quadrant switching event; monitoring the semantic state topology graph, when a quadrant switching event of the face posture class is detected, determining that a semantic state transition occurs at the current time, defining the current video frame as a key frame, the key frame containing static face structure and texture features; within the time domain interval where the semantic state transition does not occur, the current video frame is defined as a subsequent frame.
[0019] The real-time acquisition of the video stream sequence and the synchronous audio stream in the conference scene is realized by deploying an acquisition device in the conference scene, using a high-definition camera (4K resolution, 60fps high-frame-rate camera is used in this embodiment) to capture the visual picture of the conference scene, and generating a continuous video stream sequence; at the same time, using an omnidirectional microphone array to collect multiple voice signals on the spot, and generating a synchronous audio stream. In order to ensure the time sequence alignment of the multi-modal data, a uniform high-precision timestamp is stamped on each frame of the acquired video image and the corresponding audio sampling segment, ensuring the strict correspondence of the video stream sequence and the synchronous audio stream in the time dimension; For each captured video frame, a face detection algorithm (RetinaFace in this embodiment) is first used to locate all participants in the frame. The facial feature vector of each participant is extracted as an identity node in the topology graph, representing the uniqueness of the subject. Simultaneously, a head pose estimation network (HopeNet in this embodiment) is used to calculate the current Euler angles of each participant's head. These Euler angles include pitch, yaw, and roll angles, which are mapped to facial pose nodes in the topology graph. By mapping identity features to head Euler angles as topology nodes, this invention achieves precise dual-tracking of the identities and visual attention directions of multiple participants. This not only effectively avoids target confusion in multi-person meeting scenarios but also keenly captures the shift in participants' attention through quantified pose parameters. Based on this, a graph neural network is constructed, connecting identity nodes to corresponding facial pose nodes and connecting nodes from the same source in adjacent frames through temporal edges to form a semantic state topology graph. This semantic state topology graph not only reflects the static features at the current moment but also implies temporal motion trends through the connection weights between nodes. A semantic space coordinate system for facial poses is pre-established, and the Euler angle space is divided into several discrete semantic quadrants. In this embodiment, the yaw angle is divided into the left view area [-90, -30°], the front view area [-30°, +30°], and the right view area [+30°, +90°]. The graph neural network monitors the position of facial pose nodes in the semantic space in real time and calculates the quadrant index of the pose node between the previous and next frames. When the pose coordinate associated with a certain identity node is detected to cross the boundary from one preset quadrant to another preset quadrant, it is determined that a quadrant switching event has occurred, which indicates that the visual focus or interactive object of the participant has undergone an essential change. Therefore, it is determined that a semantic state transition has occurred at the current moment. Once a semantic state transition is determined, the encoder locks the current video frame as a keyframe and performs full texture encoding on the keyframe to fully preserve the static facial structure and high-frequency texture features, serving as a visual benchmark for a subsequent period. In the time domain interval where no quadrant switching occurs, the visual semantic state is determined to be stable, and these frames are defined as subsequent frames. For subsequent frames, the complete static texture is no longer encoded repeatedly; instead, only the minute changes relative to the keyframe are considered to prepare for the subsequent extraction of sparse motion field parameters, thereby achieving semantic-level decoupling and compression of video data. By monitoring quadrant switching events in the semantic state topology map, keyframes are extracted only at semantic transition moments when the visual focus of participants changes substantially. This effectively eliminates redundant coding caused by head micro-movements, solving the problem that traditional technologies cannot distinguish between invalid motion and valid semantic actions. It significantly reduces the transmission bitrate while ensuring image clarity at key interaction moments.
[0020] The sparse motion field extraction process is to perform face key point detection on the key frame and the subsequent frame respectively by using a feature recognition network, and identify semantic key points; The spatial correspondence relationship of the semantic key points between the key frame and the subsequent frame is established, the position difference of the same semantic key points in the two frames of images is compared, and the face key point geometric displacement is generated; the local neighborhood features of each semantic key point are extracted, the feature distribution changes in the local neighborhood of the key frame and the subsequent frame are compared, and the local affine transformation parameters are identified; The face key point geometric displacement and the local affine transformation parameters are aggregated to construct the sparse motion field.
[0021] A pre-trained face key point detection model is used, and PFLD is selected as the feature recognition network in this embodiment. The key frame and the current subsequent frame are input into the face key point detection model, and a plurality of semantic key point coordinates (the number of key points is set to 106 in this embodiment) including a face contour, eyebrows, eyes, a nose bridge and a lip region are output. These key points can accurately cover the high semantic active area of the face, and serve as anchor points for subsequent motion estimation; Based on the index order of the key points in the set, a point-to-point spatial one-to-one mapping relationship is established. For each index corresponding key point, the horizontal and vertical coordinate values thereof in the subsequent frame and the horizontal and vertical coordinate values thereof in the key frame are read respectively, the numerical difference between the two is calculated, and the difference value constitutes a two-dimensional translation vector, representing the absolute movement distance and direction of the face key point on the image plane, which is used for subsequent coarse-grained position change of the reconstructed face; Taking each semantic key point in the key frame as the center, a local image block with a preset pixel size, 16x16 pixels in this embodiment, is intercepted as a reference local neighborhood feature. Similarly, an image block with the same size is intercepted as a target local neighborhood feature with the key point at the corresponding position in the subsequent frame as the center. By minimizing the pixel gray difference or feature descriptor distance between the two local neighborhood feature blocks, the affine transformation parameters describing the local texture deformation are solved by applying the local solving strategy of the optical flow method. The affine transformation parameters specifically include a scaling factor, a rotation angle and a shear parameter, which are used to describe the micro deformation and stretching of the skin texture around the key point, so as to capture the micro expression details of the face; The two-dimensional geometric displacement vector calculated for each key point is spliced with the local affine transformation parameters to generate a complete motion descriptor. By traversing all the identified semantic key points, the motion descriptor set of all key points is defined as a sparse motion field. The sparse motion field only stores motion information at the semantic key point positions, and the data amount is exponentially reduced compared with the full-pixel optical flow field. At the same time, the texture deformation information required for reconstructing the face details is retained through the affine parameters, realizing efficient semantic-level compression.
[0022] The semantic segmentation process for the video frame is specifically as follows: Performing pixel-level semantic classification on the video frame to generate a semantic label map containing pixel class labels; Based on the semantic label map, region extraction and hierarchy definition are performed: aggregating pixel sets with background labels to construct a background region; identifying facial contour boundaries in the foreground to define a face region; and further locating lip texture boundaries within the face region to define a mouth region; outputting N candidate visual regions containing the background, face and mouth and their corresponding pixel indexes; A semantic segmentation neural network based on an encoder-decoder architecture (BiSeNetV2 network model is used in this embodiment) is used as the processing core. First, the current video frame is scaled to the model input size (512x512 pixels in this embodiment) and normalized; multi-scale deep features of the image are extracted through the backbone network (ResNet-50 is used); the decoder gradually up-samples the deep features to restore the original image resolution, and the probability distribution of each pixel belonging to the three preset categories of background, face skin and mouth is calculated through the Softmax classification layer; the maximum probability decision is made for each pixel position, and the class index with the maximum probability is assigned to the pixel to generate a semantic label map consistent with the resolution of the original video frame, where each pixel carries only one unique class label, i.e., 0 represents background, 1 represents face, and 2 represents mouth; Iterate through the semantic label map to retrieve the pixel coordinate index of all pixels with a label value of background, i.e., a label of 0, and use binary mask technology to generate a background mask image, in which the background pixels are set to 1 and the remaining pixels are set to 0. In order to eliminate noise, morphological opening operation is performed on the background mask to smooth the boundaries and remove isolated misjudgment pixels, and the pixel set of the background region is determined; The specific implementation process of identifying facial contour boundaries in the foreground to define a face region and further locating lip texture boundaries within the face region to define a mouth region is as follows: using a hierarchical progressive extraction strategy, first combine pixels with label values of face skin and mouth as the full foreground region, use connected component analysis algorithm to extract the largest connected region as the head range of the participant, remove the fine spots caused by background misjudgment, and determine the overall contour boundary of the face; within the determined head range, specifically retrieve pixels with a label value of mouth, i.e., a label of 2, use the difference characteristics of lip color texture and facial skin color in the color space to correct the segmentation result, accurately locate the texture boundary of the lip peak and lip corner, and thus define the mouth region, and the remaining head range pixels are defined as pure face region; According to the processing result, a region index map including three independent channels is constructed, the first channel stores a pixel coordinate list of the background region, the second channel stores a pixel coordinate list of the pure face region, and the third channel stores a pixel coordinate list of the mouth shape region. The mask matrix of the N (N=3 in this embodiment) candidate visual regions is packed with the original video frame data; The hierarchical region extraction strategy based on deep learning is combined with morphological post-processing to realize pixel-level accurate decoupling of the conference scene. Compared with the traditional coarse-grained detection method, through the combined application of connected domain analysis and chroma space correction, isolated noise points in the background are effectively filtered out and the lip texture boundary is finely corrected, ensuring that the high-sensitivity mouth shape region can be accurately separated from the face skin and avoiding the mismatch of coding resources caused by fuzzy segmentation edges.
[0023] The specific process of generating the semantic label map including pixel class labels is as follows: Pixel-level feature extraction is performed on the video frame to obtain a high-dimensional semantic feature vector corresponding to each pixel point; The high-dimensional semantic feature vector is mapped to a semantic category space, and the classification confidence of each pixel point belonging to the background, face, and mouth shape categories is calculated respectively; The maximum probability decision is executed, the class with the highest confidence is selected as the unique label of the pixel, and the semantic label map is generated by traversing all pixel points.
[0024] The specific implementation process of obtaining a high-dimensional semantic feature vector corresponding to each pixel point is as follows: a semantic segmentation neural network based on an encoder-decoder architecture is used as the processing core, and BiSeNetV2 network model is specifically selected in this embodiment. First, the current video frame collected in real time is preprocessed, scaled to the input size preset by the model, which is set to 512x512 pixels in this embodiment, and normalized to unify the data distribution. The preprocessed image is input into the backbone extraction module of the network, and ResNet-50 model is used as the backbone network in this embodiment to extract multi-scale deep features of the image using its multi-layer convolution structure. These deep features not only contain the texture details of the image, but also aggregate the context information under different receptive fields, thereby forming a high-dimensional semantic feature vector corresponding to each pixel point; The decoder module of the network is used to perform step-by-step upsampling operation on the deep features output by the backbone network to restore the spatial resolution of the feature map to the same resolution as the original video frame. After the classification layer, the feature channels are processed using the normalized exponential function to calculate the probability distribution values of each pixel point belonging to the three preset categories, i.e., background, face skin, and mouth shape. The probability distribution values accurately reflect the possibility of the current pixel belonging to each semantic category, i.e., the classification confidence; For each pixel position, compare the probability value in the background, face and mouth shape three categories, execute the maximum probability decision; select the probability value of the maximum category index assignment to the pixel, as its only class label; traversal all pixel points, generate a two-dimensional matrix consistent with the original video frame resolution, namely the semantic label map containing pixel class label; Using ResNet-50 to extract multi-scale deep features can effectively overcome the interference of complex light and background, and realize accurate pixel-level classification of face, mouth shape and background, providing a high-reliability spatial reference for subsequent differential coding of pronunciation regions.
[0025] The execution of cross-modal semantic correlation analysis is: For each candidate visual region, the vector length of the geometric displacement of the face key points in the sparse motion field is aggregated to generate a visual signal; the audio semantic feature is analyzed, and the audio time domain envelope is extracted; waveform fitting analysis is performed on the visual motion distribution sequence and the audio rhythm envelope sequence on the time axis, and the semantic correlation of the two sequences is calculated; By comparing the scores of different candidate visual regions, the region with the highest score and significant waveform synchronization characteristics is selected as the pronunciation region highly related to the current speech content.
[0026] The extraction of audio semantic features is to preprocess the collected original audio stream, and remove the environmental noise by using spectral subtraction; the audio data is processed by frame based on the video frame rate, and the video in this embodiment is 60fps, so the audio data is divided into 60 time windows per second; for the audio signal in each time window, the root mean square energy value is calculated, which reflects the volume intensity at the current time; the energy values of the continuous time windows are arranged in time sequence to construct an audio time domain envelope sequence reflecting the rhythm change and pause law of the speech, which is used as the semantic feature input on the audio side; The mouth shape region, pure face region and background region are traversed respectively; for each region, all semantic key points falling within the region are searched; the corresponding geometric displacement vectors of these key points in the sparse motion field are read, the length of each displacement vector is calculated using the Euclidean distance principle, and the displacement vector lengths of all key points in the region are accumulated to obtain the overall motion amplitude of each region in the current frame; the overall motion amplitudes calculated from the continuous video frames are connected in time sequence to generate three independent visual motion distribution sequences for the mouth shape, face and background respectively, which directly quantify the activity level of each region over time; A short-time sliding window is set, and the embodiment is a data frame containing the past half second to one second. The audio time-domain envelope sequence and the visual motion distribution sequence of each candidate region are synchronously intercepted on the time axis. A normalized cross-correlation function is used as an evaluation index to calculate the linear correlation degree between the audio envelope waveform and the mouth motion waveform, the facial motion waveform and the background motion waveform. The calculation process is essentially to evaluate the fluctuation of visual action, such as the opening and closing of the lips, and the synchronization in time of the sound intensity, such as the syllable pronunciation. The calculation result is three floating-point values between-1 and 1. The closer the value is to 1, the higher the synchronization in time between the two. The three correlation degree scores obtained by the above calculation are compared in real time. The minimum synchronization threshold is set to 0.6. If the score of a certain region is lower than the threshold, it is directly determined as a non-correlated region. In the region higher than the threshold, the maximum value optimization strategy is executed. If the correlation degree score of the mouth region is significantly higher than that of the pure facial region and the background region, which is set to be higher than 0.1 in the embodiment, it is determined that the person in the current frame is making sound through the mouth. The mouth region is confirmed as the pronunciation region. On the contrary, if all the scores are very low, which is set to be lower than 0.3 in the embodiment, it is determined that the current state is silent or voice-over state, and there is no visual pronunciation region. The mouth movement when speaking and the nodding and shaking movement when not speaking are effectively distinguished, ensuring the accuracy of semantic positioning. By performing cross-modal semantic correlation analysis, the visual motion distribution sequence and the audio rhythm envelope sequence are waveform-fitted on the time axis, realizing high-robustness semantic positioning of the pronunciation region. The introduction of audio loudness change as a strong semantic constraint and the use of normalized cross-correlation function to calculate the time sequence synchronization of audio-visual signals greatly reduce the false detection rate of false pronunciation regions, ensuring that the dynamic coding resources can accurately lock the visual region that truly carries the speech information at the current time, thereby maintaining excellent semantic understanding accuracy in a complex environment with multiple people or noisy background.
[0027] The process of calculating the quantization bias value of each candidate visual region is as follows: The numerical mapping interval of the global reference quantization parameter and the quantization bias value of the current video frame is set. The semantic correlation degree value is mapped to the numerical mapping interval, so that the negative value in the interval corresponding to the high semantic correlation degree and the positive value in the interval corresponding to the low semantic correlation degree. For the pronunciation region, a negative quantization bias value is obtained by matching the semantic correlation degree in the numerical mapping interval. For the background region, a positive quantization bias value is obtained by matching the semantic correlation degree in the numerical mapping interval.
[0028] The numerical mapping interval of the global reference quantization parameter and the quantization bias value of the current video frame is set, and the specific implementation process is that the encoding controller first calculates the basic quantization step suitable for the current frame according to the real-time available bandwidth of the current transmission channel and the water level state of the encoding buffer, which is defined as the global reference quantization parameter; at the same time, the numerical mapping interval of the quantization bias value is defined in advance, and the interval is composed of a negative integer lower limit representing the quality enhancement limit and a positive integer upper limit representing the compression limit, and the span of the interval determines the dynamic range of dynamic encoding adjustment, and the interval is set to a closed interval from -12 to 12 in this embodiment, so as to ensure that the final quantization parameter after adjustment will not exceed the effective value range allowed by the video encoding standard; The semantic correlation value is mapped to the numerical mapping interval, so that high semantic correlation corresponds to a negative value in the interval, and low semantic correlation corresponds to a positive value in the interval, and the specific implementation process is to construct a linear inverse mapping function, taking the semantic correlation as an input variable, and the logic of the linear inverse mapping function is that when the semantic correlation value tends to the maximum value one, the mapping function output approaches the negative value of the lower limit of the interval, for example, negative ten, which represents that the semantic of the region is extremely important and needs to reduce the quantization parameter to reduce the distortion; on the contrary, when the semantic correlation value tends to the minimum value zero, the mapping function output approaches the positive value of the upper limit of the interval, for example, positive ten, which represents that the semantic of the region is redundant and can increase the quantization parameter to compress data; Since the pronunciation region shows extremely high waveform synchronization in cross-modal analysis, its semantic correlation score is usually high (set to zero point eight or more in this embodiment), and the specific negative integer is calculated by substituting the mapping function, the negative integer is a negative bias value, and the negative bias value is marked as the special adjustment parameter of the pronunciation region, which is forced to reduce the quantization step in subsequent encoding, so as to completely retain the high-frequency visual information such as lip muscle texture and tooth details, and ensure the intelligibility of lip language in video conference; The background region lacks synchronous changes with the current voice rhythm, and its semantic correlation score is extremely low (i.e. less than zero point three), and a specific positive integer is matched by the mapping function according to this, the positive integer is a positive bias value, and the positive bias value is marked as the adjustment parameter of the background region, which is forced to expand the quantization step in subsequent encoding, and actively discards the high-frequency noise and texture details in the background, thereby greatly reducing the number of encoding bits occupied by the region, so as to divert the limited transmission bandwidth resources to the core pronunciation region; By establishing a reverse mapping mechanism between semantic correlation and quantitative bias value, the video coding resource is realized in the semantic level on-demand allocation, which is different from the traditional method of allocating code rate only according to human visual saliency, the application directly hooks the audio-video synchronization rate, ensures that only the pronunciation area carrying the real voice interaction information can obtain a negative bias, and the irrelevant background is applied with a positive bias, effectively solving the problem of code rate waste caused by background clutter or non-speaker motion interference in the conference video; The variable bit rate video stream encapsulation process is: A layered stream structure including a static texture layer and a semantic driven layer is constructed, the key frame is encoded based on the global reference quantization parameter and written into the static texture layer, the global reference quantization parameter and the negative quantization bias value are added to obtain a refined quantization step, the sparse motion field parameters in the pronunciation area are encoded, the global reference quantization parameter and the positive quantization bias value are added to obtain a rough quantization step, the sparse motion field parameters in the background area are encoded, and the encoded sparse motion field parameters of each area are written into the semantic driven layer, and the static texture layer, the semantic driven layer and the synchronous audio stream are multiplexed and encapsulated to generate the variable bit rate video stream. Two independent data logic channels are constructed at the output end of the encoder, the first channel is defined as a static texture layer and is specially used for storing high-fidelity key frame image data, the data update frequency is low and only when the semantic state jumps, the data is written; the second channel is defined as a semantic driven layer and is specially used for storing inter-frame continuous changing motion parameters, the data update frequency is consistent with the video frame rate, and through the layered design, the redundant background pixel information and the core semantic action information are decoupled in physical storage. The current determined global reference quantization parameter is read, and the key frame is encoded by using an intra-frame image compression algorithm; during the encoding process, the global reference quantization parameter is used to generate a quantization matrix to quantize the frequency domain coefficients of the key frame. Since the key frame is the basis for subsequent video reconstruction, a lower compression rate is usually used to retain complete facial structure and texture details, and the binary data stream generated by encoding is directly written into the data packet of the static texture layer. For the candidate visual area determined as the pronunciation area, the corresponding negative quantization bias value is read. An addition operation is performed to add the global reference quantization parameter and the negative bias value, since adding a negative number is equal to subtracting the absolute value, the operation result is a quantization parameter with a smaller value, which is defined as a refined quantization step. The refined quantization step is used for high-precision quantization of the sparse motion field parameters in the pronunciation area. The refined quantization step means retaining more decimal places or using more intensive quantization intervals, so as to losslessly record the slight opening and closing of the lips and the slight twitching of the facial muscles, and ensure the visual intelligibility of voice interaction. For the visual area determined as the background area, a corresponding positive vector quantization bias value is read. An addition operation is performed to add the global reference quantization parameter and the positive value, and the operation result is a quantization parameter with a larger value, which is defined as a rough quantization step. The motion parameters in the background area are quantized with low precision using the rough quantization step. The rough quantization step means that high-frequency jitter details in the data are discarded, and only large-amplitude position movement information is retained, thereby greatly reducing the data bit number occupied by background noise; The pronunciation area parameters and the background area parameters after the differential quantization processing are packaged and written in the semantic driving layer in time sequence. A multiplexer is started, and a synchronous time axis is introduced to interleave and encapsulate the key frame data packet of the static texture layer, the motion data packet of the semantic driving layer, and the synchronous audio stream data packet. In the encapsulation process, an aligned timestamp label is added to each data packet to ensure that the audio, facial texture reference, and dynamic expression driving parameters at a certain time can be accurately synchronized and reorganized at the decoding end, and finally a single file of variable bit rate video code stream is generated. The intra-frame image compression algorithm adopts the intra-frame encoding mode in the H.265 video encoding standard, recursively divides the key frame using the coding tree unit structure, quantizes using a specified quantization parameter, and uses arithmetic coding to entropy encode the quantization coefficients to generate a code stream conforming to the H.265 standard. The geometric displacement vector and affine transformation parameter in the sparse motion field parameter are encoded using context adaptive binary arithmetic coding, the background area parameter is quantized using a rough quantization step, and the pronunciation area parameter is quantized using a fine quantization step. Finally, the static texture layer, the semantic driving layer, and the synchronous audio stream are encapsulated according to the multiplexing specification of MPEG-4 Part 12.
[0029] By constructing a layered code stream structure separating static texture and semantic driving, and combining a differential quantization strategy based on semantic bias, semantic-level intelligent compression of video data is achieved, breaking the limitation of traditional video encoding using a uniform quantization parameter for the entire image. The motion capture accuracy of the pronunciation area is forcibly improved through negative bias, and the motion data of irrelevant background is deeply compressed through positive bias. Under extremely low bandwidth conditions, conference videos with high-fidelity lip details can still be smoothly transmitted, significantly solving the problem of video conference experience lag and blur in a network congestion environment.
[0030] The application solves the problem of waste of coding resources or loss of key information caused by misjudgment of the core area in a complex conference scene by performing cross-modal semantic correlation analysis, jointly solving the sparse motion field distribution characteristics in the visual area and the semantic characteristics of the synchronous audio stream, using the high synchronization of audio-visual signals in the time sequence to determine the core area, effectively overcoming the limitations of traditional single-mode methods relying on visual saliency for judgment, accurately excluding high visual response interference caused by strong light, close distance or background object motion, and accurately locking the pronunciation area that is truly expressing the speech content; and then applying negative bias to the pronunciation area and positive bias to the background area according to the semantic correlation, realizing accurate allocation of coding resources based on real semantic activity, and solving the problem of waste of coding resources or loss of key information caused by misjudgment of the core area in a complex conference scene.
[0031] Embodiment two: This embodiment applies the intelligent conference video frame dynamic coding method based on multi-modal semantic understanding to a certain round table conference, including one main speaker, three conference participants and a projection screen background in the conference room; Real-time acquisition of video stream sequence and synchronous audio stream in the conference scene, a panoramic camera device is deployed at the center position of the round table conference room, a 4K panoramic camera with a resolution of 3840x2160 and a frame rate of 30fps is used in this embodiment to capture a 360-degree visual picture covering the whole scene, and a video stream sequence is generated; at the same time, a six-array omnidirectional microphone deployed on the desktop is used to collect the surround sound signals on the spot, and a synchronous audio stream with a sampling rate of 48kHz is generated; based on the audio sampling clock, the video frame is aligned in hardware level time stamping, to ensure that the lip movement of the main speaker and the audio signal are strictly synchronized in time dimension, with an error controlled within 10 milliseconds; Performing spatio-temporal topology analysis based on graph neural network on the video stream sequence to extract key frames and subsequent frames; using a multi-task cascaded convolutional neural network to locate the faces of the four conference participants in the picture to construct four identity nodes; using a fine-grained structure aggregation network to calculate the head Euler angle of each conference participant to construct a face posture node; establishing a semantic state topology graph, defining the yaw angle in the interval [-30°, 30°] as the attention face area and the range beyond the range as the attention screen area; monitoring the topology graph, when the main speaker turns his head from the attention face area to the attention screen area to watch the projection content, determining that a quadrant switching event and a semantic state transition occur, defining the current frame as a key frame, and performing static texture coding on the side face of the main speaker and the projection screen content; during the duration that the main speaker keeps facing the screen to explain, the video frame is defined as a subsequent frame; The face motion recognition is performed on the subsequent frame, a sparse motion field is extracted, a lightweight convolutional neural network is used as a feature recognition network, and 106-point high-precision face key point detection is performed on the key frame and the subsequent frame; a point-to-point mapping relationship is established, and a geometric displacement vector of the face key point of the main speaker is calculated; meanwhile, a local neighborhood of 32 by 32 pixels around the key point is intercepted, and local affine transformation parameters are calculated by comparing feature distributions; the geometric displacement and the affine parameters are aggregated to construct a sparse motion field that only describes the opening and closing of the mouth shape and the facial micro-expression of the main speaker during explanation, and the data amount is only one percent of the original pixel data; The semantic segmentation processing is performed on the video frame to generate a semantic label map containing pixel class labels, a DeepLabV3 Plus deep semantic segmentation network is used as the processing core, the video frame is scaled to 1024 by 1024 pixels to input the network, and a 101-layer deep residual network is used as the backbone network to extract deep features; the high-dimensional semantic feature vector is mapped to the class space, and the classification confidence of each pixel belonging to the conference room background, the participant face, and the main speaker mouth shape is calculated; the maximum probability decision is performed to generate a semantic label map with the same proportion as the original picture, and the active speaking area and the static conference room background area are distinguished; Based on the semantic label map, the region extraction and level definition are performed, which are divided into N candidate visual regions, the pixels of all labels that are the conference room background and non-active participants are aggregated to construct a background region; the contour boundary of the main speaker's head is recognized to define the face region; the lip color texture feature is used to correct the boundary in the main speaker's face region to locate the mouth shape region; finally, the pixel index mask of the three candidate visual regions is output; The semantic content analysis is performed on the synchronous audio stream, and the cross-modal semantic correlation analysis is performed, specifically, the Mel frequency cepstral coefficient is extracted from the collected audio stream, and the root mean square energy is calculated to generate an audio time domain envelope sequence; at the same time, the motion vector module length of the key points in the main speaker mouth shape region, the main speaker face region and the background region is aggregated respectively to generate three groups of visual motion distribution sequences; a sliding window of 1.0 second is set, and the waveform fitting degree of the audio envelope and each visual sequence is calculated by using the normalized cross-correlation function; when the correlation coefficient of the motion waveform of the main speaker mouth shape region and the speech energy envelope is 0.85, and the correlation coefficient of the background region is 0.15, the main speaker's mouth shape region is recognized and positioned as the pronunciation area highly related to the current speech content; The quantitative bias value of each candidate visual region is calculated according to the semantic correlation degree, the global reference quantitative parameter under the current network environment is set to 28, and the value mapping interval is set to negative fifteen to positive fifteen; a linear mapping function is constructed, the semantic correlation degree 0.85 of the lip shape region of the main speaker is mapped to a negative bias value negative twelve, aiming to greatly reduce the quantization step to retain the details of the lip language; the semantic correlation degree 0.15 of the conference room background region is mapped to a positive bias value positive fifteen, aiming to maximize the quantization step to compress redundant data; The package is a variable bit rate video stream, a hierarchical stream structure is constructed, the key frame is encoded with the reference quantitative parameter 28 and written into the static texture layer; the fine quantization step is calculated, that is, 28 plus negative 12 equals 16, the sparse motion parameters of the pronunciation region are encoded with the step to ensure that the remote participants can clearly identify the lip shape of the main speaker; the coarse quantization step is calculated, that is, 28 plus 15 equals 43, the motion parameters of the background region are compressed with the step; finally, the encoded parameters of each region are written into the semantic driving layer, multiplexed with the static texture layer and the synchronous audio stream, and a variable bit rate video stream suitable for the low delay demand of business meetings is generated.
[0032] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for dynamic encoding of intelligent conference video frames based on multimodal semantic understanding, characterized in that, include: Real-time acquisition of video stream sequences and synchronous audio streams in conference scenarios; Content-based semantic analysis and decoupling processing are performed on the video stream sequence to extract key frames and subsequent frames, wherein the key frames contain static facial structure and texture features; Facial motion recognition is performed on the subsequent frames to extract the sparse motion field of the geometric displacement of facial key points and local affine transformation parameters. At the same time, semantic segmentation processing is performed on the video frames to divide them into N candidate visual regions containing face, lip shape and background. Semantic content analysis is performed on the synchronous audio stream to extract audio semantic features; cross-modal semantic association analysis is performed to calculate the semantic correlation between the distribution features of the sparse motion field in the candidate visual region and the audio semantic features, and the pronunciation region is identified and located based on the semantic correlation. Quantization bias values for each candidate visual region are calculated based on the semantic relevance. A negative bias is applied to the sound region, and a positive bias is applied to the background region, which is then encapsulated into a variable bitrate video stream.
2. The intelligent conferencing video frame dynamic encoding method based on multimodal semantic understanding according to claim 1, characterized in that, The specific process of extracting keyframes and subsequent frames from the video stream sequence includes: A spatiotemporal topology analysis based on graph neural networks is performed on the video stream sequence to construct a semantic state topology graph containing participant identity nodes and facial pose nodes; The semantic state topology is monitored. When a quadrant switching event of facial pose category is detected, it is determined that a semantic state transition has occurred at the current moment. The current video frame is defined as a key frame, which contains static facial structure and texture features. In the time domain interval where the semantic state transition has not occurred, the current video frame is defined as a subsequent frame.
3. The intelligent conferencing video frame dynamic encoding method based on multimodal semantic understanding according to claim 1, characterized in that, The sparse motion field extraction process involves using a feature recognition network to perform facial key point detection on the key frame and the subsequent frame respectively, and identifying semantic key points. Establish the spatial correspondence between the semantic key points between the key frame and the subsequent frame, generate the geometric displacement of the facial key points by comparing the positional differences of the same semantic key points in the two frames; extract the local neighborhood features of each semantic key point, and identify the local affine transformation parameters by comparing the feature distribution changes in the local neighborhood between the key frame and the subsequent frame. The geometric displacements of the facial key points are aggregated with the local affine transformation parameters to construct the sparse motion field.
4. The intelligent conferencing video frame dynamic encoding method based on multimodal semantic understanding according to claim 1, characterized in that, The specific process of performing semantic segmentation on video frames is as follows: Perform pixel-level semantic classification on the video frames to generate a semantic label map containing pixel category labels; Based on the semantic label map, perform region extraction and hierarchy definition: aggregate the set of pixels with background labels to construct a background region; identify the facial contour boundary in the foreground to define the facial region; and further locate the lip texture boundary within the facial region to define the mouth shape region; output N candidate visual regions containing the background, face, and mouth shape and their corresponding pixel indices.
5. The intelligent conferencing video frame dynamic coding method based on multimodal semantic understanding according to claim 4, characterized in that, The specific process for generating a semantic label map containing pixel category labels is as follows: Pixel-level feature extraction is performed on the video frames to obtain a high-dimensional semantic feature vector corresponding to each pixel. The high-dimensional semantic feature vector is mapped to the semantic category space, and the classification confidence of each pixel belonging to the background, face, and mouth shape categories is calculated respectively. Perform the highest probability decision, select the category with the highest confidence as the unique label for that pixel, and traverse all pixels to generate the semantic label map.
6. The intelligent conferencing video frame dynamic encoding method based on multimodal semantic understanding according to claim 1, characterized in that, The cross-modal semantic association analysis is performed as follows: For each candidate visual region, the geometric displacement of facial key points in the sparse motion field within the region is aggregated using vector magnitudes to generate a visual signal; the audio semantic features are analyzed to extract the audio temporal envelope; the visual motion distribution sequence and the audio rhythm envelope sequence are subjected to waveform fitting analysis on the time axis to calculate the semantic correlation between the changes in the two sequences; By comparing the scores of different candidate visual regions, the region with the highest score and significant waveform synchronization characteristics is selected and identified as the pronunciation region that is highly related to the current speech content.
7. The intelligent conferencing video frame dynamic coding method based on multimodal semantic understanding according to claim 1, characterized in that, The process of calculating the quantization bias value of each candidate visual region is as follows: Set the numerical mapping range of the global baseline quantization parameters and quantization offset values for the current video frame; The semantic relevance values are mapped to the numerical mapping interval, such that high semantic relevance corresponds to negative values within the interval, and low semantic relevance corresponds to positive values within the interval. For the pronunciation region, a negative vectorized bias value is obtained by matching within the numerical mapping interval based on semantic relevance; for the background region, a positive vectorized bias value is obtained by matching within the numerical mapping interval based on its semantic relevance.
8. The intelligent conferencing video frame dynamic encoding method based on multimodal semantic understanding according to claim 1, characterized in that, The variable bitrate video stream encapsulation process is as follows: A hierarchical bitstream structure comprising a static texture layer and a semantic driving layer is constructed. The keyframes are encoded based on global reference quantization parameters and written into the static texture layer. The global reference quantization parameters are added to the negative vectorization bias value to obtain a fine quantization step size, which is then used to encode the sparse motion field parameters in the pronunciation region. The global baseline quantization parameter is added to the positive vectorization bias value to obtain the coarsened quantization step size, and the sparse motion field parameters in the background region are encoded. The encoded sparse motion field parameters of each region are written into the semantic driving layer, and the static texture layer, the semantic driving layer and the synchronous audio stream are multiplexed and encapsulated to generate the variable bitrate video stream.
Citation Information
Cited By
Virtual anchor real-time driving system based on facial motion capture
CN121842342A
Narrowband transmission method and system based on semantic compression and audio and video joint perception
CN121963737A
A multi-modal content intelligent extraction and automatic picture matching method for news media
CN122290025A
A full-link-aware conference device failure prediction method and system
CN122290316A