Multimedia conference room sound system based on artificial intelligence

By using an AI-based multimedia conference room audio system, the main speaking channel can be identified and optimized in real time, solving the problems of cumbersome operation and noise interference in multi-person interactive scenarios of traditional systems, and achieving high-quality voice output and natural conference interaction.

CN120812477BActive Publication Date: 2026-03-24GUANGDONG RUIZHAO AUDIO EQUIPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traditional multimedia conference room audio systems are cumbersome to operate in real-time interactive scenarios such as multiple people speaking alternately, free discussion, and interruption responses. This can easily lead to problems such as speech not being picked up or environmental noise, echoes, and overlap with the conversation, affecting the clarity of speech recognition and the accuracy of meeting records.

Method used

An AI-based multimedia conference room audio system is adopted. The system uses a personnel positioning module to acquire the image position and head posture of the participants in real time. Combined with the microphone layout, it automatically constructs the spatial mapping relationship between the personnel and the microphone, determines the main speaking channel in real time and performs gain maintenance and suppression processing. Combined with behavioral feature extraction and speech recognition, it predicts the next speaker and optimizes the strategy model.

Benefits of technology

It achieves automated, high-quality voice channel recognition and retention, reduces the need for manual intervention with microphones and speakers, and improves voice clarity and naturalness of interaction during meetings, making it suitable for various medium to large-scale intelligent meeting scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812477B_ABST
    Figure CN120812477B_ABST
Patent Text Reader

Abstract

The present application relates to conference room sound control technical field, especially in kind based on artificial intelligence's multimedia conference room sound system, through personnel positioning module real-time acquisition of the image position and head posture of the participant, the automatic establishment of the space mapping relationship of personnel and channel is combined with microphone layout, the dynamic binding of microphone channel is realized. Voice activity detection based on audio data automatically identifies the main speaking channel, and differentiates the control of the channel gain through the sound output module, effectively suppresses the background noise of the non-speaking microphone. By extracting the behavior characteristics and interaction intention of the main speaker, the behavior characteristics and interaction intention are jointly modeled based on the reinforcement learning model, the prediction of the next speaker and the dynamic update of the main channel are realized, and the strategy parameters are continuously optimized based on the speech feedback. Reduce manual operation, improve the clarity of voice output and the natural fluency of conference interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of conference room audio control technology, and more particularly to a multimedia conference room audio system based on artificial intelligence. Background Technology

[0002] Multimedia conference rooms are important venues for communication, decision-making reporting, and collaborative discussions among departments in enterprises and research institutions. To ensure clear, high-quality audio output, current multimedia conference rooms generally adopt a layout of multiple microphone arrays combined with speaker systems to cover the speaking areas of all participants.

[0003] However, with the increase in the number of participants and the diversification of speaking methods, traditional conference audio systems have revealed a series of problems in practical applications. Existing systems mostly rely on preset fixed microphone channels, requiring manual adjustments to channel volume and microphone activation / deactivation. In highly interactive scenarios with real-time interaction, such as multiple participants speaking alternately, free discussion, and interruptions, participants often need to manually turn on their microphones before each speech and turn them off afterward. This results in frequent and cumbersome operations, severely impacting the naturalness and continuity of the meeting flow. Furthermore, oversights in switching microphones on and off can easily lead to unrecorded speech or overlapping of ambient noise, echoes, and other audio, interfering with other participants' speeches and reducing the clarity of speech recognition and the accuracy of meeting recordings. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides a multimedia conference room audio system based on artificial intelligence.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] An artificial intelligence-based multimedia conference room audio system includes:

[0007] The personnel positioning module is used to perform face recognition and head pose estimation on the collected image data to generate the identity identifier and corresponding spatial location of each participant.

[0008] The microphone association module is used to construct a temporary relationship table between participants and microphone channels based on the spatial location and the microphone layout in the conference room;

[0009] The main speaker initial determination module is used to detect the voice activity frequency of the audio data collected by each microphone channel and determine the main speaker channel and the corresponding personnel identity based on the temporary relationship table.

[0010] The audio output module is used to perform channel gain maintenance processing on the main speaking channel and gain suppression processing on the other channels to obtain the current effective audio signal set, and to output the effective audio signal set as the audio signal output of the conference room speaker.

[0011] The interaction feature extraction module is used to extract behavioral features such as gaze direction, head orientation, and hand movements from the personnel area of ​​the main speaking channel in the image data based on the main speaking channel and the corresponding personnel identification, and generate a behavioral feature set; and to perform speech recognition and semantic parsing processing on the effective speech audio signal set to extract the interaction intent of the current speech content.

[0012] The main speaking channel update module is used to predict the next speaker based on the behavioral feature set and interaction intent using a reinforcement learning model, update the main speaking channel based on a temporary relationship table, calculate a reward value based on the actual speaking feedback, and update the policy parameters of the reinforcement learning model based on the reward value.

[0013] Furthermore, the personnel positioning module is used to perform the following steps:

[0014] Based on the collected image data, face detection processing is performed on the image data to obtain a set of face regions of all participants in the image frame;

[0015] Based on the set of face regions, time-series tracking processing is performed, and a unique temporary identifier is assigned to each face region according to the frame sequence to obtain the participant's identity identifier and image location pair;

[0016] Based on the image location, head pose estimation is performed on each face region to obtain the head pose parameters corresponding to each identifier. Invalid faces with viewpoint deviations outside the threshold are filtered out by combining the image location to obtain the face spatial projection parameter set after pose correction.

[0017] Based on the set of facial spatial projection parameters and camera calibration parameters, the facial center point and orientation information of each identifier are mapped in three dimensions to obtain the identity identifier and corresponding spatial location of each participant.

[0018] Furthermore, the microphone association module is used to perform the following steps:

[0019] Based on the preset spatial location of the conference room microphones, the Euclidean distance calculation is performed on the participant identification and corresponding spatial location output by the personnel positioning module to obtain the set of distances between each participant and all microphones;

[0020] Based on the aforementioned distance set, spatial distance threshold filtering and minimum distance priority matching are performed to determine the temporary binding relationship between each participant and the microphone channel;

[0021] Based on the temporary binding relationship, construct a temporary relationship table that is updated in real time.

[0022] Furthermore, the main speaker initialization module is used to perform the following steps:

[0023] Based on the audio data collected from each microphone channel, the audio data is processed by frame-by-frame windowing, and the short-time energy is calculated for each frame of audio data to obtain the short-time energy sequence corresponding to each microphone channel.

[0024] For the short-time energy sequence, perform energy threshold comparison and activity frame ratio calculation processing to obtain the real-time voice activity frequency of each microphone channel;

[0025] Based on the real-time voice activity frequency, the microphone channel with the highest real-time voice activity frequency is determined as the main speaking channel, and the personnel identity identifier corresponding to the main speaking channel is determined based on the temporary relationship table.

[0026] Furthermore, the audio output module is used to perform the following steps:

[0027] Calculate the main channel gain coefficient of the current frame based on the short-time energy of the main speaking channel audio signal;

[0028] For the audio signal of the non-main speaking channel, short-time energy analysis is performed frame by frame to calculate the short-time energy ratio between the audio signal of the non-main speaking channel and the audio signal of the main speaking channel, and the real-time attenuation coefficient corresponding to the non-main speaking channel is determined based on the short-time energy ratio.

[0029] Based on the gain coefficient of the main channel and the real-time attenuation coefficient of the non-main channel, the audio signals of each channel are weighted and superimposed frame by frame to obtain the fused effective audio signal set, which is then output to the speaker for sound output.

[0030] Furthermore, the step of extracting behavioral features such as gaze direction, head orientation, and hand gestures from the personnel area of ​​the main speaking channel in the image data based on the main speaking channel and the corresponding personnel identification, and generating a behavioral feature set, includes the following steps:

[0031] Based on the personnel identification corresponding to the main speaking channel, the image region of the personnel in the current frame is extracted from the image data, and the facial key points and hand key points are located in the image region to obtain a standard key point coordinate set;

[0032] Based on the key point coordinate set, a two-dimensional behavioral feature matrix is ​​constructed that integrates head posture, gaze direction and hand movements. Each row of the two-dimensional behavioral feature matrix represents the spatial position and relative relationship of various key points in a frame, and each column corresponds to different types of action dimension information.

[0033] The two-dimensional behavior feature matrix is ​​input into a multi-layer residual neural network containing a residual connection structure, and local perceptual encoding and time step residual information fusion are performed sequentially to output an intermediate behavior embedding representation containing action trend and interaction pattern representation.

[0034] The intermediate behavior embedding representation is subjected to fully connected mapping and dimensionality compression to generate a behavior feature vector with a unified structure as the output behavior feature set.

[0035] Furthermore, the step of performing speech recognition and semantic parsing processing on the effective speech audio signal set to extract the interactive intent of the current speech content includes:

[0036] Perform speech preprocessing and feature encoding on the audio signal corresponding to the main speaking channel, and input it into the Transformer speech recognition model to output the text transcription result of the current speech;

[0037] Named entity recognition and role dereference are performed on the text transcription result to extract the names or titles of the participants appearing in the speech, and match them with the identity identifier set generated by the personnel positioning module to obtain the target identity of the interaction.

[0038] Based on the text content, a word vector sequence is constructed and input into a semantic parsing network to extract the interaction type label and core intent phrase in the current speech, thereby obtaining the interaction purpose corresponding to the current speech.

[0039] The target identity and interaction purpose are integrated to construct the interaction intent vector.

[0040] Furthermore, the speech preprocessing includes noise suppression, silence removal, and spectral normalization.

[0041] Furthermore, the reinforcement learning model is constructed through the following steps:

[0042] Based on behavioral feature sets and mutual intention vectors, a state vector is constructed to describe the current meeting state.

[0043] Based on the identity identifier set generated by the personnel positioning module, an action space is established with each participant's identity as a discrete action.

[0044] The reward value is calculated using a reward function based on the matching between the reinforcement learning action output and the actual identity of the next speaker;

[0045] The state vector is input into the policy network to perform forward computation, outputting the probability distribution corresponding to each candidate action in the action space. Based on the reward value, the gradient value of the network parameters is calculated using the policy gradient algorithm, and the weight parameters of the policy network are updated according to the gradient value.

[0046] Furthermore, the formula for the reward function is as follows:

[0047] ;

[0048] in, The immediate reward for the i-th decision; The identifier of the speaker who actually spoke after the i-th decision; The identifier of the highest probability candidate speaker output by the reinforcement learning model in the i-th decision; For matching functions; In the state The following model identifies the speakers in actual events. The given probability distribution values ​​for the policy network; and These are the weight parameters.

[0049] The beneficial effects of this invention are as follows: This invention acquires the image position and head posture of each participant in real time through a personnel positioning module. Combined with the preset layout of microphones in the conference room, it automatically constructs a spatial mapping relationship between personnel and microphones, realizing dynamic binding and updating of channels. Furthermore, based on voice activity detection of audio data, it determines the main speaking channel in real time and performs gain maintenance of the main channel and gain suppression of non-main channels through the audio output module, thereby automatically completing the identification and preservation of high-quality voice channels and avoiding background noise from non-speaking microphones from mixing into the output audio. Through a behavioral feature extraction module, it analyzes the main speaker's image area for gaze, head orientation, and gesture movements. Combined with speech recognition and semantic parsing models, it extracts interactive intent information in the current speech, constructing an interactive feature vector containing directional relationships and speaking purposes, improving the modeling ability for potential speaking transition trends in multi-round interactions. Through a reinforcement learning model, it jointly models and predicts behavioral features and interactive intent, realizing intelligent prediction of the next speaker and dynamic updating of the main channel, and continuously optimizing the strategy model based on actual feedback. This invention reduces the need for manual intervention with microphones and audio equipment in meetings, improves voice clarity and natural interaction during meetings, and is suitable for various medium-to-large-scale intelligent meeting scenarios. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the structure of a multimedia conference room audio system based on artificial intelligence, as described in this invention.

[0051] Figure 2 This is a flowchart of the steps involved in constructing the reinforcement learning model in this invention. Detailed Implementation

[0052] Please see Figures 1-2As shown, the present invention relates to an artificial intelligence-based multimedia conference room audio system, comprising:

[0053] The personnel positioning module is used to perform face recognition and head pose estimation on the collected image data to generate the identity identifier and corresponding spatial location of each participant.

[0054] The microphone association module is used to construct a temporary relationship table between participants and microphone channels based on the spatial location and the microphone layout in the conference room;

[0055] The main speaker initial determination module is used to detect the voice activity frequency of the audio data collected by each microphone channel and determine the main speaker channel and the corresponding personnel identity based on the temporary relationship table.

[0056] The audio output module is used to perform channel gain maintenance processing on the main speaking channel and gain suppression processing on the other channels to obtain the current effective audio signal set, and to output the effective audio signal set as the audio signal output of the conference room speaker.

[0057] The interaction feature extraction module is used to extract behavioral features such as gaze direction, head orientation, and hand movements from the personnel area of ​​the main speaking channel in the image data based on the main speaking channel and the corresponding personnel identification, and generate a behavioral feature set; and to perform speech recognition and semantic parsing processing on the effective speech audio signal set to extract the interaction intent of the current speech content.

[0058] The main speaking channel update module is used to predict the next speaker based on the behavioral feature set and interaction intent using a reinforcement learning model, update the main speaking channel based on a temporary relationship table, calculate a reward value based on the actual speaking feedback, and update the policy parameters of the reinforcement learning model based on the reward value.

[0059] In some embodiments, this system is deployed inside a conference room, equipped with a multi-channel directional microphone array, a high-definition camera, and a speaker system. First, the personnel positioning module receives the real-time image stream, extracts the facial region of each participant based on a face recognition algorithm (MTCNN), and obtains the three-axis rotation angle of the head using a pose estimation network (HopeNet). Combining the image coordinates with camera calibration parameters, the system maps the relative position and identification number of each participant in three-dimensional space. The system does not require fixed seating and has dynamic adaptability. Subsequently, the microphone association module uses known microphone layout coordinates and calculates Euclidean distance to spatially match the real-time position of each participant with the microphone channel. A temporary mapping table between participants and channels is constructed using a minimum distance priority and exclusive binding rule, providing accurate person-microphone correspondence for audio channel management. During the meeting, the main speaker initialization module continuously receives audio streams from each channel, performs energy detection-based voice activity recognition (VAD), determines the current main speaker channel by statistically analyzing the frequency of voice activity within a sliding window, and identifies the speaker based on the mapping table. Building upon this foundation, the audio output module performs gain preservation processing on the main speaking channel while applying dynamic suppression coefficients to other channels. Based on a gain adjustment strategy calculated using short-time energy ratios, it achieves real-time sound source focusing, reducing background noise from non-speaker channels from superimposing into the main channel. This results in a higher signal-to-noise ratio and improved speech intelligibility for the audio signal output to the speakers. To enhance the system's understanding and prediction of speaking behavior, the interaction feature extraction module extracts facial key points and hand contours within the main speaker's image region, constructing triples for head posture, gaze direction, and gesture. Simultaneously, it preprocesses the speech, combining this with a BERT semantic encoder to identify target titles, keywords, and interactive intentions. Finally, this module outputs a fused feature vector representing the dynamics of behavior and semantic orientation. The main speaking channel update module uses this feature vector as state input to construct a policy gradient-based reinforcement learning model. Using candidate identities as the action space and actual speaking success as reward feedback, it continuously optimizes the main channel switching strategy. This mechanism achieves intelligent predictive channel scheduling, distinct from traditional rule-based judgment and manual operation, adapting to complex speech scenarios such as multi-person interaction and interruptions, enhancing the system's responsiveness and prediction accuracy.

[0060] Furthermore, the personnel positioning module is used to perform the following steps:

[0061] Based on the collected image data, face detection processing is performed on the image data to obtain a set of face regions of all participants in the image frame;

[0062] Based on the set of face regions, time-series tracking processing is performed, and a unique temporary identifier is assigned to each face region according to the frame sequence to obtain the participant's identity identifier and image location pair;

[0063] Based on the image location, head pose estimation is performed on each face region to obtain the head pose parameters corresponding to each identifier. Invalid faces with viewpoint deviations outside the threshold are filtered out by combining the image location to obtain the face spatial projection parameter set after pose correction.

[0064] Based on the set of facial spatial projection parameters and camera calibration parameters, the facial center point and orientation information of each identifier are mapped in three dimensions to obtain the identity identifier and corresponding spatial location of each participant.

[0065] In some embodiments, real-time image frame sequences are acquired using multiple high-definition RGB cameras deployed at the front and sides of the conference room. First, the system invokes a multi-scale face detection model based on a deep convolutional neural network to perform face detection on each frame, outputting a set of candidate face regions including face bounding boxes, confidence scores, and coordinates of facial key points. Then, a temporal association algorithm based on optical flow guidance and IoU matching is used to consistently label the same face region in consecutive frames. Combined with the multi-object tracking network DeepSort, a unique temporary identifier is assigned to each face trajectory, forming an identifier-image spatial location pair to ensure continuous and stable tracking under conditions of free movement, partial occlusion, or changes in lighting. For each confirmed face region, a 3D head pose estimation method based on a pose regression network is further invoked, such as HopeNet or a 6DoF head pose regression model, to analyze its pitch, yaw, and roll pose parameters. The system calculates the angle between the estimated result and the viewing direction. If the angle deviates from the camera's main viewing angle by more than a set threshold (e.g., ±45°), the face information in that frame is deemed unreliable and is removed to improve positioning accuracy. For face regions with valid poses, PnP (Programmable Noise Proof) is performed using the face center pixel coordinates and known distortion correction parameters within the camera. This projects the two-dimensional image coordinates into three-dimensional space, restoring the position and orientation vector relative to the camera coordinate system, forming a complete set of face spatial projection parameters. Finally, based on the three-dimensional coordinates and orientation vector of the face center, the system maps the identity of each participant to their physical spatial position in the conference room scenario, providing an accurate spatial basis for subsequent dynamic binding of microphone channels. Compared to traditional methods that rely solely on seat numbers or manual binding, this implementation significantly improves the adaptability and real-time accuracy of personnel positioning in dynamic scenes by introducing structured pose estimation and multi-target temporal association algorithms.

[0066] Furthermore, the microphone association module is used to perform the following steps:

[0067] Based on the preset spatial location of the conference room microphones, the Euclidean distance calculation is performed on the participant identification and corresponding spatial location output by the personnel positioning module to obtain the set of distances between each participant and all microphones;

[0068] Based on the aforementioned distance set, spatial distance threshold filtering and minimum distance priority matching are performed to determine the temporary binding relationship between each participant and the microphone channel;

[0069] Based on the temporary binding relationship, construct a temporary relationship table that is updated in real time.

[0070] In some embodiments, the 3D spatial coordinates of all microphones are obtained through pre-configured parameters, typically calibrated based on conference room modeling or layout diagrams. Simultaneously, the personnel positioning module outputs the identity identifier of each participant and their corresponding 3D spatial position in the current frame (e.g., world coordinates of the tip of the nose or the center of the face). Subsequently, the system performs Euclidean distance calculations on the position vector of each participant and the coordinates of each microphone. To avoid misbinding distant microphones, the system sets a spatial distance threshold ε, retaining only microphone channels with a distance less than ε as valid candidate channels. If this set is not empty, the microphone channel with the smallest distance value is selected as the temporary binding channel for that participant. If it is empty, it is considered that the microphone was not effectively associated in the current frame, and the system retryes in the next frame. By traversing all participants and completing the temporary binding of their respective microphones, the system constructs a mapping table of personnel and microphone channels for the current frame and caches it as a temporary relationship table. This table is updated in real time as the personnel's positions change dynamically, serving as a key reference in subsequent modules such as main speaker recognition and gain adjustment.

[0071] Furthermore, the main speaker initialization module is used to perform the following steps:

[0072] Based on the audio data collected from each microphone channel, the audio data is processed by frame-by-frame windowing, and the short-time energy is calculated for each frame of audio data to obtain the short-time energy sequence corresponding to each microphone channel.

[0073] For the short-time energy sequence, perform energy threshold comparison and activity frame ratio calculation processing to obtain the real-time voice activity frequency of each microphone channel;

[0074] Based on the real-time voice activity frequency, the microphone channel with the highest real-time voice activity frequency is determined as the main speaking channel, and the personnel identity identifier corresponding to the main speaking channel is determined based on the temporary relationship table.

[0075] Specifically, firstly, the system continuously receives audio data from each microphone channel and divides each audio segment into short time frames. A windowing method is used to locally enhance the audio signal in each frame, extracting its short-time energy features. By calculating the energy level of the audio signal frame by frame, the system can construct a short-time energy sequence for each channel within a certain time range to reflect the current intensity of speech activity. Subsequently, based on a preset energy threshold, the energy sequence of each channel is judged, and the proportion of active frames with energy exceeding the threshold within the time window is counted, thereby quantifying the speech activity frequency of each channel. The higher the speech activity frequency, the greater the likelihood of continuous speaking behavior in that channel. This indicator can effectively eliminate the interference of brief noise or non-speaking behavior in the judgment in meeting scenarios where multiple people speak concurrently or take turns. Finally, the main speaker initialization module determines the microphone channel with the highest speech activity frequency as the main speaker channel based on the comparison of the speech activity frequencies of each channel, and accurately locates the participant's identity corresponding to that channel by combining the temporary relationship table provided by the microphone association module.

[0076] Furthermore, the audio output module is used to perform the following steps:

[0077] Calculate the main channel gain coefficient of the current frame based on the short-time energy of the main speaking channel audio signal;

[0078] For the audio signal of the non-main speaking channel, short-time energy analysis is performed frame by frame to calculate the short-time energy ratio between the audio signal of the non-main speaking channel and the audio signal of the main speaking channel, and the real-time attenuation coefficient corresponding to the non-main speaking channel is determined based on the short-time energy ratio.

[0079] Based on the gain coefficient of the main channel and the real-time attenuation coefficient of the non-main channel, the audio signals of each channel are weighted and superimposed frame by frame to obtain the fused effective audio signal set, which is then output to the speaker for sound output.

[0080] Specifically, firstly, for the currently identified main speaking channel, its audio signal is extracted and divided into frames. Short-time energy calculations are used to obtain the energy value corresponding to each frame, and the gain coefficient of the main channel is further set based on this energy level. This coefficient is used to maintain the clarity and volume proportion of the main speaking signal during subsequent weighted processing, preventing the speech content from being weakened by overall dynamic suppression. Next, the same short-time energy analysis is performed on the audio signals of all non-main speaking channels, and the ratio is calculated with the energy value of the corresponding frame of the main channel. This ratio reflects the relative intensity of the non-main speaking signal relative to the main channel. The system maps the comparison value based on a preset nonlinear attenuation function or proportional function (reciprocal function or sigmoid mapping) to generate the real-time attenuation coefficient of the non-main channel in the current frame. This processing ensures timely dynamic suppression even when there is slight environmental interference or low-amplitude noise, thereby improving the overall speech output quality. Finally, the audio output module performs weighted superposition processing on the audio frame signals of all channels based on the gain coefficient of the main channel and the attenuation coefficient of the non-main channels to obtain a fused effective audio signal set. This audio set maintains the dominance of the main speaker's content in both the time and frequency domains, while significantly reducing background noise or non-target content that may be present in other channels. The processed audio signal is output to the conference room speakers in real time, achieving clear and focused sound playback, significantly improving the auditory experience of participants and the efficiency of meeting communication.

[0081] Furthermore, the step of extracting behavioral features such as gaze direction, head orientation, and hand gestures from the personnel area of ​​the main speaking channel in the image data based on the main speaking channel and the corresponding personnel identification, and generating a behavioral feature set, includes the following steps:

[0082] Based on the personnel identification corresponding to the main speaking channel, the image region of the personnel in the current frame is extracted from the image data, and the facial key points and hand key points are located in the image region to obtain a standard key point coordinate set;

[0083] Based on the key point coordinate set, a two-dimensional behavioral feature matrix is ​​constructed that integrates head posture, gaze direction and hand movements. Each row of the two-dimensional behavioral feature matrix represents the spatial position and relative relationship of various key points in a frame, and each column corresponds to different types of action dimension information.

[0084] The two-dimensional behavior feature matrix is ​​input into a multi-layer residual neural network containing a residual connection structure, and local perceptual encoding and time step residual information fusion are performed sequentially to output an intermediate behavior embedding representation containing action trend and interaction pattern representation.

[0085] The intermediate behavior embedding representation is subjected to fully connected mapping and dimensionality compression to generate a behavior feature vector with a unified structure as the output behavior feature set.

[0086] In some embodiments, firstly, based on the person's identity identifier bound to the main speaking channel, the image region of that person is accurately extracted in the current image frame. The image region is typically the upper body, with particular attention paid to key areas such as the face and hands. The system calls a facial landmark detection algorithm (e.g., a deep learning-based MTCNN or HRNet model) to locate 68 standard facial landmarks, and simultaneously combines a hand landmark recognition model (MediaPipe Hands) to obtain positional data containing 21 hand nodes. This results in a set of landmark coordinates that integrates the face and hands, with each landmark represented by two-dimensional or three-dimensional coordinates. Subsequently, a two-dimensional behavioral feature matrix is ​​constructed based on this landmark set. Specifically, each row of the matrix represents the coordinate features of a certain type of landmark in the image frame, such as the position of the corner of the eye, the vector from the top of the head to the chin, the angle between the fingers, etc.; each column corresponds to the dimensional information of different time steps or landmark categories, such as the first column being the head pose angle, the second column being the cosine value of the gaze direction, and the third column being the degree of left finger spread, etc. This matrix is ​​used to simultaneously model spatial structural relationships (such as left-right symmetry and angular changes) and temporal dynamic features (such as the start and end of actions and duration). After construction, the two-dimensional behavioral feature matrix is ​​input into a deep residual neural network containing residual connection structures. Each residual block in the network consists of two convolutional layers and an identity mapping path; the former extracts local spatial change patterns, while the latter preserves the original temporal feature information. To enhance the modeling ability of the temporal dimension, the network inserts temporal convolutional layers or a Transformer-based temporal attention mechanism between every two residual blocks, enabling the model to learn the mutual influence patterns between key points in different time periods. For example, behaviors such as the head turning from left to right or the fingers changing from closed to open can be identified as complete trends within multiple frames. The network output is an intermediate behavioral embedding representation, which has a more compact clustering effect for similar behavioral patterns in the feature space. This embedding vector is then non-linearly mapped through a multilayer perceptron (MLP) and subjected to dimensionality compression using feature selection mechanisms (such as attention weighting or information entropy filtering), ultimately generating a fixed-length, semantically clear behavioral feature vector. This vector can characterize the intention of the current speaker, such as whether they are pointing at someone or making a request (such as raising their hand or indicating that they want to speak), providing structured input for interactive intent recognition and subsequent speech prediction.

[0087] Furthermore, the step of performing speech recognition and semantic parsing processing on the effective speech audio signal set to extract the interactive intent of the current speech content includes:

[0088] Perform speech preprocessing and feature encoding on the audio signal corresponding to the main speaking channel, and input it into the Transformer speech recognition model to output the text transcription result of the current speech;

[0089] Named entity recognition and role dereference are performed on the text transcription result to extract the names or titles of the participants appearing in the speech, and match them with the identity identifier set generated by the personnel positioning module to obtain the target identity of the interaction.

[0090] Based on the text content, a word vector sequence is constructed and input into a semantic parsing network to extract the interaction type label and core intent phrase in the current speech, thereby obtaining the interaction purpose corresponding to the current speech.

[0091] The target identity and interaction purpose are integrated to construct the interaction intent vector.

[0092] In some embodiments, the audio signal corresponding to the main speaking channel is first subjected to speech preprocessing, including silence removal, noise filtering, and audio normalization, to ensure the stability and clarity of the input signal in the time and frequency domains. Subsequently, the processed audio signal is input to an end-to-end speech recognition model built using a Transformer architecture. This model models the time-series features of the input audio based on an attention mechanism and outputs the corresponding text transcription result in conjunction with a phoneme-level decoder module. This text transcription has high semantic restoration capabilities, accurately restoring the key words, sentence structure, and pauses of the speaker's statements. After obtaining the text transcription result, the system uses a Named Entity Recognition (NER) model in Natural Language Processing to label and identify names, titles, and other terms appearing in the text. For example, "What does Mr. Li think of this project?" will identify "Mr. Li" as a name / title entity. Subsequently, the system performs role referencing resolution processing to determine the specific identity attribution of pronouns such as "he," "that person," and "this colleague" in the context. By matching the identified names or titles with the identity identifier set generated by the personnel location module, the target personnel identifier semantically pointed to in the text can be obtained, achieving a closed loop from natural language to structured identity mapping. Further, to clarify the pragmatic intent expressed by the speaker, the system segments the text transcription results and constructs a word vector sequence using a word embedding model (BERT or Word2Vec), which is then input into a custom semantic parsing network. This network typically employs a BiGRU+Attention structure or a lightweight Transformer sub-network, classifying common speaking intent types in meeting scenarios (such as requesting to speak, issuing instructions, soliciting opinions, and topic shifting), and extracting phrase combinations containing semantic cores as the core of intent expression. For example, "Please ask Teacher Zhang for his opinion on this period's budget" will be parsed as "pointing to Teacher Zhang, the interaction purpose is to solicit opinions." Finally, the system fuses and encodes the extracted target identity (such as Teacher Zhang's identity identifier) ​​with the interaction purpose (such as "soliciting opinions") to generate a unified format interaction intent vector. This vector serves as one of the important state inputs in the reinforcement learning model, participating in the policy judgment for predicting the next speaker.

[0093] Furthermore, the speech preprocessing includes noise suppression, silence removal, and spectral normalization.

[0094] Specifically, firstly, to address potential background noise in the meeting environment (such as air conditioning noise, page-turning noise, and whispering), the system employs a combination of spectral subtraction and neural network enhancement models (such as DCCRN or Wavesplit) for noise suppression. Spectral subtraction estimates the background noise spectrum of non-speech frames and subtracts it from the speech spectrum, thus reducing static noise. The DCCRN model, based on a trained time-frequency enhancement network, effectively restores the speech characteristics of the speech signal while preserving natural sound quality. Secondly, to improve the processing efficiency of the speech recognition system, the system performs Voice Activity Detection (VAD) processing on the original speech signal. This processing uses a dual-channel VAD network based on gated recurrent units (GRUs) to jointly judge the short-time energy of the speech and features such as Mel-frequency cepstral coefficients (MFCCs), accurately identifying and removing non-speech frames. Through this step, the system effectively reduces meaningless data input, lowers the computational load, and improves the response speed and accuracy of the speech transcription model in multi-turn speeches. Finally, to ensure consistency in amplitude and spectral distribution of audio signals from different microphone channels, the system performs spectral normalization on the audio signals. Specifically, the audio signals are first converted into a spectral representation using a Short-Time Fourier Transform (STFT), and then the spectrum is subjected to logarithmic amplitude compression and Z-score normalization for each channel. This ensures that the feature spaces of different input channels have a uniform numerical distribution, thereby guaranteeing stable recognition performance of the Transformer speech recognition model under multi-source input.

[0095] Furthermore, the reinforcement learning model is constructed through the following steps:

[0096] Based on behavioral feature sets and mutual intention vectors, a state vector is constructed to describe the current meeting state.

[0097] Based on the identity identifier set generated by the personnel positioning module, an action space is established with each participant's identity as a discrete action.

[0098] The reward value is calculated using a reward function based on the matching between the reinforcement learning action output and the actual identity of the next speaker;

[0099] The state vector is input into the policy network to perform forward computation, outputting the probability distribution corresponding to each candidate action in the action space. Based on the reward value, the gradient value of the network parameters is calculated using the policy gradient algorithm, and the weight parameters of the policy network are updated according to the gradient value.

[0100] In some embodiments, firstly, during the state vector construction phase, the system receives a two-dimensional behavioral feature matrix output by the behavioral feature extraction module. This matrix integrates action elements such as the speaker's gaze direction, head posture, and gestures over a time series dimension. Simultaneously, the interaction feature extraction module outputs an interaction intent vector of the current speech content, containing target identity and semantic core phrase information. The system performs feature-level concatenation or attention fusion on these two features (e.g., using a multi-head attention mechanism to jointly encode the behavioral features and semantic vectors) to construct a state vector with contextual semantic awareness, used to represent the current meeting context and behavioral state. During the action space construction phase, the system sets up an action space based on the identity set identified by the personnel positioning module, where each action represents switching speaking control to a specific participant. Since the action space is finite and enumerable, the policy network outputs a probability distribution vector for all candidate actions. In the policy network design phase, this embodiment uses a multi-layer feedforward neural network as the policy network. The network input is the state vector, and the output is the predicted probability distribution π(a|s) for all candidate actions. This network transforms the linearly transformed logits into action selection probabilities through a softmax output layer. To enhance the model's ability to model long-term behavioral sequences, a Transformer encoder structure can be used to pre-encode the state vector, and residual connections and layer normalization can be introduced to improve training stability. During the reward function design phase, a reward value is defined to measure the effectiveness of action selection. A positive reward is set when the system's predicted next speaker matches the actual speaker; otherwise, a penalty is imposed. The reward value can be further constructed by incorporating soft factors such as the degree of interaction semantic matching, action transition confidence, and delayed response time, forming a composite reward structure. The state vector from each round of the meeting is input into the policy network. The policy network generates a probability distribution of candidate speakers based on the current parameters and samples the current predicted action from it. Subsequently, the reward value is calculated based on the matching between the action and the actual speaker, and this reward value serves as a feedback signal for optimizing the objective function. To improve the expected total reward of the policy, a REINFORCE-based policy gradient method is adopted. This method estimates the gradient of the policy network parameters by maximizing the expected reward function under the current policy, using the sampled actions and their corresponding rewards. Then, using the backpropagation algorithm, the network parameters are adjusted based on this gradient information, making the policy network more inclined to output actions that bring high rewards in subsequent states. This policy optimization mechanism combines feedback information from actual meeting behavior with probabilistic modeling characteristics, which not only improves the model's ability to adapt to speaking patterns, but also has the ability to continuously learn and generalize inference, thus maintaining high-efficiency speaking channel prediction performance in complex meeting scenarios such as multi-round free interaction and frequent switching among multiple people.

[0101] Furthermore, the formula for the reward function is as follows:

[0102] ;

[0103] in, The immediate reward for the i-th decision; The identifier of the speaker who actually spoke after the i-th decision; The identifier of the highest probability candidate speaker output by the reinforcement learning model in the i-th decision; For matching functions; In the state The following model identifies the speakers in actual events. The given probability distribution values ​​for the policy network; and These are the weight parameters.

[0104] It's important to note that the reward function compares the model's predicted results with the actual speaker's identity identifier to determine if they match perfectly. A high base reward is given if they match, and zero if they don't. Simultaneously, the system further examines the probability value the model assigns to the actual speaker, serving as a soft score for the prediction's reasonableness. This probability value comes from the probability distribution of all candidate speakers output by the policy network in the current meeting state, with the probability corresponding to the actual speaker specifically extracted as a measure of policy confidence. Ultimately, the reward value consists of two parts: a hard match based on accurate prediction and a probability score based on the model's confidence in the actual speaker. These two parts are weighted and fused using preset weighting factors, allowing the model to both reinforce accurate predictions during training and receive appropriate incentives when approaching the correct answer, thus improving learning stability and convergence speed.

[0105] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A multimedia conference room audio system based on artificial intelligence, characterized in that, include: The personnel positioning module is used to perform face recognition and head pose estimation on the collected image data to generate the identity identifier and corresponding spatial location of each participant. The microphone association module is used to construct a temporary relationship table between participants and microphone channels based on the spatial location and the microphone layout in the conference room; The main speaker initial determination module is used to detect the voice activity frequency of the audio data collected by each microphone channel and determine the main speaker channel and the corresponding personnel identity based on the temporary relationship table. The audio output module is used to perform channel gain maintenance processing on the main speaking channel and gain suppression processing on the other channels to obtain the current effective audio signal set, and to output the effective audio signal set as the audio signal output of the conference room speaker. The interaction feature extraction module is used to extract behavioral features such as gaze direction, head orientation, and hand movements from the personnel area of ​​the main speaking channel in the image data based on the main speaking channel and the corresponding personnel identification, and generate a behavioral feature set; and to perform speech recognition and semantic parsing processing on the effective speech audio signal set to extract the interaction intent of the current speech content. The main speaking channel update module is used to predict the next speaker based on the behavioral feature set and interaction intent through a reinforcement learning model and update the main speaking channel according to a temporary relationship table. It also calculates a reward value based on the actual speaking feedback and updates the policy parameters of the reinforcement learning model based on the reward value. The personnel positioning module is used to perform the following steps: Based on the collected image data, face detection processing is performed on the image data to obtain a set of face regions of all participants in the image frame; Based on the set of face regions, time-series tracking processing is performed, and a unique temporary identifier is assigned to each face region according to the frame sequence to obtain the participant's identity identifier and image location pair; Based on the image location, head pose estimation is performed on each face region to obtain the head pose parameters corresponding to each identifier. Invalid faces with viewpoint deviations outside the threshold are filtered out by combining the image location to obtain the face spatial projection parameter set after pose correction. Based on the set of facial spatial projection parameters and camera calibration parameters, the facial center point and orientation information of each identifier are mapped in three dimensions to obtain the identity identifier and corresponding spatial location of each participant.

2. The multimedia conference room audio system based on artificial intelligence according to claim 1, characterized in that, The microphone association module is used to perform the following steps: Based on the preset spatial location of the conference room microphones, the Euclidean distance calculation is performed on the participant identification and corresponding spatial location output by the personnel positioning module to obtain the set of distances between each participant and all microphones; Based on the aforementioned distance set, spatial distance threshold filtering and minimum distance priority matching are performed to determine the temporary binding relationship between each participant and the microphone channel; Based on the temporary binding relationship, construct a temporary relationship table that is updated in real time.

3. The multimedia conference room audio system based on artificial intelligence according to claim 1, characterized in that, The main speaker initialization module is used to perform the following steps: Based on the audio data collected from each microphone channel, the audio data is processed by frame-by-frame windowing, and the short-time energy is calculated for each frame of audio data to obtain the short-time energy sequence corresponding to each microphone channel. For the short-time energy sequence, perform energy threshold comparison and activity frame ratio calculation processing to obtain the real-time voice activity frequency of each microphone channel; Based on the real-time voice activity frequency, the microphone channel with the highest real-time voice activity frequency is determined as the main speaking channel, and the personnel identity identifier corresponding to the main speaking channel is determined based on the temporary relationship table.

4. The multimedia conference room audio system based on artificial intelligence according to claim 3, characterized in that, The audio output module is used to perform the following steps: Calculate the main channel gain coefficient of the current frame based on the short-time energy of the main speaking channel audio signal; For the audio signal of the non-main speaking channel, short-time energy analysis is performed frame by frame to calculate the short-time energy ratio between the audio signal of the non-main speaking channel and the audio signal of the main speaking channel, and the real-time attenuation coefficient corresponding to the non-main speaking channel is determined based on the short-time energy ratio. Based on the gain coefficient of the main channel and the real-time attenuation coefficient of the non-main channel, the audio signals of each channel are weighted and superimposed frame by frame to obtain the fused effective audio signal set, which is then output to the speaker for sound output.

5. The multimedia conference room audio system based on artificial intelligence according to claim 1, characterized in that, The step of extracting behavioral features such as gaze direction, head orientation, and hand gestures from the personnel area in the main speaking channel of the image data based on the main speaking channel and the corresponding personnel identification, and generating a behavioral feature set, includes the following steps: Based on the personnel identification corresponding to the main speaking channel, the image region of the personnel in the current frame is extracted from the image data, and the facial key points and hand key points are located in the image region to obtain a standard key point coordinate set; Based on the key point coordinate set, a two-dimensional behavioral feature matrix is ​​constructed that integrates head posture, gaze direction and hand movements. Each row of the two-dimensional behavioral feature matrix represents the spatial position and relative relationship of various key points in a frame, and each column corresponds to different types of action dimension information. The two-dimensional behavior feature matrix is ​​input into a multi-layer residual neural network containing a residual connection structure, and local perceptual encoding and time step residual information fusion are performed sequentially to output an intermediate behavior embedding representation containing action trend and interaction pattern representation. The intermediate behavior embedding representation is subjected to fully connected mapping and dimensionality compression to generate a behavior feature vector with a unified structure as the output behavior feature set.

6. The multimedia conference room audio system based on artificial intelligence according to claim 5, characterized in that, The step of performing speech recognition and semantic parsing processing on the valid audio signal set to extract the interactive intent of the current speech content includes: Perform speech preprocessing and feature encoding on the audio signal corresponding to the main speaking channel, and input it into the Transformer speech recognition model to output the text transcription result of the current speech; Named entity recognition and role dereference are performed on the text transcription result to extract the names or titles of the participants appearing in the speech, and match them with the identity identifier set generated by the personnel positioning module to obtain the target identity of the interaction. Based on the text content, a word vector sequence is constructed and input into a semantic parsing network to extract the interaction type label and core intent phrase in the current speech, thereby obtaining the interaction purpose corresponding to the current speech. The target identity and interaction purpose are integrated to construct the interaction intent vector.

7. A multimedia conference room audio system based on artificial intelligence according to claim 6, characterized in that, The speech preprocessing includes noise suppression, silence removal, and spectrum normalization.

8. A multimedia conference room audio system based on artificial intelligence according to claim 6, characterized in that, The reinforcement learning model is constructed through the following steps: Based on behavioral feature sets and mutual intention vectors, a state vector is constructed to describe the current meeting state. Based on the identity identifier set generated by the personnel positioning module, an action space is established with each participant's identity as a discrete action. The reward value is calculated using a reward function based on the matching between the reinforcement learning action output and the actual identity of the next speaker; The state vector is input into the policy network to perform forward computation, outputting the probability distribution corresponding to each candidate action in the action space. Based on the reward value, the gradient value of the network parameters is calculated using the policy gradient algorithm, and the weight parameters of the policy network are updated according to the gradient value.

9. A multimedia conference room audio system based on artificial intelligence according to claim 8, characterized in that, The formula for the reward function is as follows: ; in, The immediate reward for the i-th decision; The identifier of the speaker who actually spoke after the i-th decision; The identifier of the highest probability candidate speaker output by the reinforcement learning model in the i-th decision; For matching functions; In the state The following model identifies the speakers in actual events. The given probability distribution values ​​for the policy network; and These are the weight parameters.

Citation Information

Patent Citations

  • Intelligent conference content real-time translation method based on voice recognition

    CN120164479A

  • Intelligent conference management method and system

    WO2019148583A1