Multimedia conference room sound system based on artificial intelligence
Through the AI-based multimedia conference room sound system, the location and behavioral characteristics of participants are identified and processed in real time, and the microphone channel gain is automatically adjusted, which solves the problems of cumbersome operation and noise interference of traditional systems and improves the voice clarity and naturalness of interaction in meetings.
Patent Information
- Application Number
- CN202511005194.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Traditional multimedia conference room audio systems are cumbersome to operate in scenarios where multiple people take turns speaking, free discussions, and highly real-time interactions. This can easily result in speeches not being picked up or being interfered with by ambient noise, affecting the clarity of speech recognition and the accuracy of meeting records.
An artificial intelligence-based multimedia conference room sound system is used. The personnel positioning module obtains the image position and head posture of participants in real time, builds a temporary relationship table based on the microphone layout, determines the main speaking channel in real time and performs gain maintenance and suppression processing. Combined with behavioral feature extraction and speech recognition, the reinforcement learning model is used to predict the next speaker and update the main speaking channel.
It achieves automated high-quality voice channel recognition and retention, reduces the need for manual intervention in microphones and speakers, and improves voice clarity and natural interaction during meetings.
Smart Images

Figure CN120812477A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of conference room sound control, and in particular to a multimedia conference room sound system based on artificial intelligence. BACKGROUND
[0002] A multimedia conference room is an important place for work departments such as enterprises and scientific research institutions to organize communication, decision-making reporting, and collaborative discussion. To ensure clear and high-quality audio output, current multimedia conference rooms generally use a layout method of combining multiple microphone arrays with a loudspeaker system to cover the speaking areas of all participants.
[0003] However, with the increase in the number of participants and the diversification of speaking methods, traditional conference sound systems have exposed a series of problems in actual application. Existing systems rely on pre-set fixed microphone channels and perform operations such as channel volume adjustment, microphone start-stop control, and the like through manual methods. In real-time interactive scenarios such as multiple people speaking alternately, free discussion, interrupting responses, and the like, conference participants often need to manually turn on the microphone before each speech and turn it off after the speech, resulting in frequent operations and cumbersome use, which seriously affects the naturalness and continuity of the conference process. At the same time, if there is a switching operation oversight, it can cause environmental noise, echo, and overlapping problems of the speech not being picked up or the microphone not being turned off, which further interferes with the speech of other personnel, reducing the clarity of voice recognition and the accuracy of conference records. SUMMARY
[0004] To solve the above problems, the present application provides a multimedia conference room sound system based on artificial intelligence.
[0005] To achieve the above purpose, the technical solution adopted by the present application is: A multimedia conference room sound system based on artificial intelligence, comprising: A personnel positioning module for performing face recognition and head pose estimation on image data based on collected image data, generating the identity of each participant and the corresponding spatial position; A microphone association module for constructing a temporary relationship table between participants and microphone channels based on the spatial position and the relationship between the conference room microphone layout; A primary speaker initial determination module for performing voice activity frequency detection on audio data collected by each microphone channel and determining the primary speaker channel and the corresponding participant identity according to the temporary relationship table; A sound output module for performing channel gain maintenance processing on the primary speaker channel and gain suppression processing on the remaining channels to obtain a current effective audio signal set, and outputting the effective audio signal set as a conference room loudspeaker output audio signal for sound output; An interaction feature extraction module is configured to extract behavior features of a line-of-sight direction, a head orientation and a hand movement of a person in a main speaking channel in the image data based on the person identity identifier of the main speaking channel, to generate a behavior feature set, and to perform speech recognition and semantic analysis processing on the set of valid speaking audio signals to extract an interaction intention of the current speaking content. A main speaking channel updating module is configured to predict a next speaking person by a reinforcement learning model according to the behavior feature set and the interaction intention, to update the main speaking channel according to a temporary relationship table, to calculate a reward value according to an actual speaking feedback result, and to update a policy parameter of the reinforcement learning model by the reward value.
[0006] Further, the person positioning module is configured to perform the following steps: Based on the collected image data, face detection processing is performed on the image data to obtain a face region set of all participants in the image frame. Based on the face region set, time sequence tracking processing is performed to assign a unique temporary identifier to each face region according to a frame sequence to obtain a person identity identifier and an image position pair. According to the image position, head posture estimation processing is performed on each face region to obtain head posture parameters corresponding to each identifier, and invalid faces outside a view angle deviation threshold are filtered in combination with the image position to obtain a set of posture-corrected face space projection parameters. According to the set of face space projection parameters and camera calibration parameters, three-dimensional mapping is performed on the face center point and orientation information of each identifier to obtain a person identity identifier and a corresponding spatial position of each participant.
[0007] Further, the microphone association module is configured to perform the following steps: Based on the preset spatial position of the conference room microphone, a Euclidean distance calculation processing is performed on the person identity identifier and the corresponding spatial position of each participant output by the person positioning module to obtain a distance set between each participant and all microphones. Based on the distance set, spatial distance threshold screening and minimum distance priority matching are performed to determine a temporary binding relationship between each participant and a microphone channel. According to the temporary binding relationship, a real-time updated temporary relationship table is constructed.
[0008] Further, the main speaking initial determination module is configured to perform the following steps: Based on the audio data collected by each microphone channel, frame windowing processing is performed on the audio data, and short-time energy is calculated for each frame of audio data to obtain a short-time energy sequence corresponding to each microphone channel. Energy threshold comparison and active frame proportion calculation processing are performed on the short-time energy sequence to obtain a real-time speech activity frequency of each microphone channel. determine the microphone channel with the highest real-time voice activity frequency as the main speaking channel, and determine the personnel identity corresponding to the main speaking channel based on the temporary relationship table.
[0009] Further, the audio output module is configured to perform the following steps: calculate a main channel gain coefficient for the current frame based on the short-time energy of the main speaking channel audio signal; perform short-time energy analysis on the audio signal of each non-main speaking channel, calculate the short-time energy ratio between the audio signal of the non-main speaking channel and the audio signal of the main speaking channel, and determine a real-time attenuation coefficient of the non-main speaking channel based on the short-time energy ratio; perform frame-by-frame weighted superposition of the audio signals of each channel based on the gain coefficient of the main channel and the real-time attenuation coefficient of the non-main channel, obtain a fused effective audio signal set, and output to the loudspeaker for audio output.
[0010] Further, the behavior feature set is generated by performing line-of-sight direction, head orientation, and hand movement behavior feature extraction on the personnel region of the main speaking channel in the image data based on the personnel identity corresponding to the main speaking channel, including the following steps: extract the image region of the personnel in the current frame from the image data based on the personnel identity corresponding to the main speaking channel, perform face key point and hand key point positioning on the image region to obtain a standard key point coordinate set; based on the key point coordinate set, construct a two-dimensional behavior feature matrix that fuses head posture, line-of-sight direction, and hand movement, each row of the two-dimensional behavior feature matrix represents the spatial position and relative relationship of each type of key point in a frame, and each column corresponds to different categories of movement dimension information; input the two-dimensional behavior feature matrix into a multi-layer residual neural network containing a residual connection structure, sequentially perform local perception coding and time step residual information fusion, and output an intermediate behavior embedding representation containing movement trend and interaction mode representation; perform full connection mapping and dimension compression processing on the intermediate behavior embedding representation to generate a uniform structure behavior feature vector as the behavior feature set output.
[0011] Further, the speech recognition and semantic analysis processing on the effective speaking audio signal set to extract the interactive intent of the current speaking content includes: perform speech preprocessing and feature encoding on the audio signal corresponding to the main speaking channel, and input into a Transformer speech recognition model to output a text transcription result of the current speaking; Perform named entity recognition and role anaphora resolution on the text transcription result, extract the names or titles of participants appearing in the speech, and match them with the identity set generated by the personnel positioning module to obtain the target identity of the interaction; Based on the text content, construct a word vector sequence and input it into a semantic parsing network to extract the interaction type label and intent core phrase in the current speech, and obtain the interaction purpose corresponding to the current speech; Fuse the target identity and interaction purpose to construct an interaction intent vector.
[0012] Further, the speech preprocessing includes noise suppression, silence segment rejection, and spectral normalization.
[0013] Further, the reinforcement learning model is constructed by the following steps: Based on the behavior feature set and the interaction intent vector, a state vector is constructed to describe the current conference state; Based on the identity set generated by the personnel positioning module, an action space is established with each participant's identity as a discrete action; According to the matching between the reinforcement learning action output and the actual next speaker identity, the reward value is calculated by a reward function; Input the state vector into the policy network for forward calculation, output the probability distribution corresponding to each candidate action in the action space, and based on the reward value, calculate the gradient value of the network parameters by the policy gradient algorithm, and update the weight parameters of the policy network according to the gradient value.
[0014] Further, the reward function is as follows: ; Where, is the immediate reward of the i-th decision; is the actual speaker identity after the i-th decision; is the highest probability candidate speaker identity output by the reinforcement learning model in the i-th decision; is the matching function; is the policy network probability distribution value given by the model to the actual speaker identity in state ; and are weight parameters.
[0015] The beneficial effects of the present application are that the present application acquires the image position and head posture of each participant in real time through the personnel positioning module, automatically constructs the spatial mapping relationship between the personnel and the microphone based on the preset layout of the microphone in the conference room, realizes the dynamic binding and updating of the channel, and further judges the main speaking channel in real time based on the voice activity detection of the audio data, and executes the gain maintenance of the main channel and the gain suppression processing of the non-main channel through the sound output module, thereby automatically completing the identification and reservation of the high-quality voice channel, avoiding the mixing of background noise of the non-speaking microphone into the output audio. Through the line of sight, head orientation and gesture action analysis of the image area of the main speaker by the behavior feature extraction module, the interactive intention information in the current speech is extracted by cooperating with the speech recognition and semantic analysis model, the interactive feature vector containing the pointing relationship and the speaking purpose is constructed, and the modeling ability of the potential speaking transfer trend in the multi-round interaction is improved. The behavior features and interactive intentions are jointly modeled and predicted through the reinforcement learning model, the intelligent prediction of the next speaker and the dynamic updating of the main channel are realized, and the strategy model is continuously optimized based on the actual feedback. The present application reduces the artificial intervention demand of the microphone and the sound in the conference, improves the voice clarity and interaction naturalness in the conference process, and is suitable for various large-scale intelligent conference scenes. BRIEF DESCRIPTION OF DRAWINGS
[0016] Fig. 1 Fig. 1 is a structural schematic diagram of a multimedia conference room sound system based on artificial intelligence in the present application.
[0017] Fig. 2 Fig. 5 is a construction step flow chart of a reinforcement learning model in the present application. DETAILED DESCRIPTION
[0018] Referring to Figs. 1-2 The present application relates to a multimedia conference room sound system based on artificial intelligence, which comprises: A personnel positioning module is used for performing face recognition and head posture estimation on the image data based on the collected image data, and generating the identity of each participant and the corresponding spatial position; A microphone association module is used for constructing a temporary relationship table between the participants and the microphone channels based on the spatial position and the layout relationship of the conference room microphone; A main speaking preliminary determination module is used for performing voice activity frequency detection on the audio data collected by each microphone channel, and determining the main speaking channel and the corresponding personnel identity according to the temporary relationship table; A sound output module is used for performing channel gain maintenance processing on the main speaking channel, performing gain suppression processing on the remaining channels, obtaining a current effective audio signal set, and outputting the effective audio signal set as the conference room loudspeaker output audio signal for sound output; An interaction feature extraction module is configured to extract the behavior features of the line-of-sight direction, head orientation and hand movement of the personnel region of the main speaking channel in the image data based on the main speaking channel and the corresponding personnel identity, and generate a behavior feature set; and perform speech recognition and semantic analysis processing on the valid speaking audio signal set to extract the interaction intent of the current speaking content. A main speaking channel updating module is configured to predict the next speaking personnel by a reinforcement learning model according to the behavior feature set and the interaction intent, update the main speaking channel according to a temporary relationship table, calculate a reward value according to an actual speaking feedback result, and update the policy parameters of the reinforcement learning model through the reward value.
[0019] In some embodiments, the system is deployed in a conference room, and is configured with a multi-path directional microphone array, a high-definition camera, and a speaker system. First, the personnel positioning module accesses real-time image streams, extracts the face area of each participant based on a face recognition algorithm (MTCNN), and obtains the three-axis rotation angle of the head by combining a pose estimation network (HopeNet). By combining the image coordinates and the camera calibration parameters, the relative position and identity number of the participants in the three-dimensional space are mapped. The system does not require fixed seats and has dynamic adaptability. Subsequently, the microphone association module uses the known microphone layout coordinates to perform spatial matching between the real-time positions of each participant and the microphone channels by calculating the Euclidean distance. By using the minimum distance priority and exclusive binding rules, a temporary mapping table between the participants and the channels is constructed, providing an accurate person-microphone correspondence for audio channel management. During the meeting, the main speaker determination module continuously receives the audio streams collected by each channel, performs voice activity detection (VAD) based on energy detection, determines the current main speaker channel by counting the voice activity frequency in the sliding window, and locks the speaker identity according to the mapping table. Based on this, the sound output module performs gain preservation processing on the main speaker channel, and applies a dynamic suppression coefficient to other channels. Based on the gain adjustment strategy calculated by the short-time energy ratio, real-time sound source focusing is achieved, the background noise of non-speaker channels is reduced, and the audio signal output to the speaker has better signal-to-noise ratio and language clarity. To enhance the system's understanding and prediction of speaking behavior, the interactive feature extraction module extracts facial key points and hand contours in the image area of the main speaker, constructs a head pose, gaze direction, and gesture action triplet, and simultaneously performs preprocessing on the speaker's voice, identifies the target title, keywords, and interactive intent in the speech by combining the BERT semantic encoder. Finally, the module outputs a fusion feature vector representing the behavior dynamics and semantic direction. The main speaker channel update module uses the feature vector as the state input, constructs a strategy gradient type reinforcement learning model, uses the candidate personnel identity as the action space, and uses the actual speaking hit situation as the reward feedback to continuously optimize the main channel switching strategy. This mechanism realizes intelligent prediction channel scheduling that is different from traditional rule judgment and manual operation, adapts to complex speech scenarios such as multi-person interaction and interruption, and enhances the responsiveness and prediction accuracy of the system.
[0020] Further, the personnel positioning module is configured to perform the following steps: Based on the acquired image data, perform face detection processing on the image data to obtain a set of face areas of all participants in the image frame; Based on the set of face areas, perform time sequence tracking processing to assign a unique temporary identifier to each face area according to the frame sequence, and obtain a participant identity identifier and image position pair; According to the image position, head pose estimation processing is performed on each face region to obtain head pose parameters corresponding to each identifier, and invalid faces deviating from a view angle threshold are filtered based on the image position to obtain a set of face space projection parameters after pose correction; According to the face space projection parameter set and the camera calibration parameter, the face center point and the orientation information of each identifier are mapped in three dimensions to obtain the identity identifier of each participant and the corresponding spatial position.
[0021] In some embodiments, a real-time image frame sequence is obtained by deploying multiple high-definition RGB cameras in front of and on the side of the conference room. First, the system calls a multi-scale face detection model based on a deep convolutional neural network to perform face detection on each frame of image, outputting a set of face candidate regions including face bounding boxes, confidence scores, and five feature point coordinates. Subsequently, a time sequence correlation algorithm based on optical flow guidance and IoU matching is used to consistently mark the same face region in consecutive frames, and a unique temporary identity identifier is assigned to each face track by combining a multi-target tracking network DeepSort, forming an identity identifier and image spatial position pair, ensuring continuous and stable tracking under conditions of free movement, partial occlusion, or changes in lighting. For each confirmed face region, a three-dimensional head pose estimation method based on a pose regression network is further called, which can be HopeNet or a 6DoF head pose regression model, to analyze the pitch, yaw, and roll pose parameters. The angle between the estimation result and the view direction is calculated, and if the angle deviates from the main view angle of the camera by more than a set threshold (such as ±45°), the face information in that frame is deemed to be unreliable and is discarded to improve positioning accuracy. For the face region with valid pose, the face center point pixel coordinates, and the known camera internal and distortion correction parameters are combined to perform PnP solving, projecting the two-dimensional image coordinates into three-dimensional space to restore their position and orientation vector relative to the camera coordinate system, forming a complete set of face space projection parameters. Finally, the system maps the identity identifier of each participant and its physical spatial position relationship in the conference room scene based on the face center three-dimensional coordinates and the orientation vector, providing an accurate spatial basis for the dynamic binding of the subsequent microphone channels. Compared with the traditional method of relying only on seat numbers or manual binding, this implementation significantly improves the adaptability and real-time accuracy of personnel positioning in dynamic scenarios by introducing structured pose estimation and multi-target time sequence correlation algorithms.
[0022] Further, the microphone association module is configured to perform the following steps: Based on the preset spatial positions of the conference room microphones, the Euclidean distance calculation processing is performed on the participant identity identifier and the corresponding spatial position output by the personnel positioning module to obtain a set of distances between each participant and all microphones; Based on the distance set, a spatial distance threshold screening and a minimum distance priority matching are performed to determine a temporary binding relationship between each participant and a microphone channel; According to the temporary binding relationship, a real-time updated temporary relationship table is constructed.
[0023] In some embodiments, the three-dimensional spatial coordinates of all microphones are obtained through pre-configured parameters, and the calibration is usually completed based on conference room modeling or layout drawings. At the same time, the personnel positioning module outputs the identity of each participant and the three-dimensional spatial position (such as the world coordinates of the nose tip or the face center point) corresponding to the current frame. Then, the system performs Euclidean distance calculation on the position vector of each participant and the coordinates of each microphone in turn. In order to avoid misbinding of long-distance microphones, the system sets a spatial distance threshold ε, and only keeps the microphone channels that satisfy the distance less than ε as the effective candidate set. If the set is not empty, the microphone channel with the minimum distance value is selected as the temporary binding channel of the participant. If it is empty, it is considered that the current frame cannot effectively associate the microphone, and the next frame is retried. By traversing all participants and completing the temporary binding of each microphone, the system constructs a personnel and microphone channel mapping table for the current frame, and caches it as a temporary relationship table. The table is updated in real time with the dynamic changes of the personnel position, and is used as a key reference basis in subsequent main speech recognition, gain adjustment and other modules.
[0024] Further, the main speech preliminary determination module is configured to perform the following steps: Based on the audio data collected by each microphone channel, the audio data is subjected to frame windowing processing, and the short-time energy of each frame of audio data is calculated to obtain the short-time energy sequence corresponding to each microphone channel; The short-time energy sequence is subjected to energy threshold comparison and active frame proportion calculation processing to obtain the real-time speech activity frequency of each microphone channel. Based on the real-time speech activity frequency, the microphone channel with the highest real-time speech activity frequency is determined as the main speech channel, and the identity of the participant corresponding to the main speech channel is determined based on the temporary relationship table.
[0025] Specifically, first, the system continuously receives audio data from each microphone channel, and divides each piece of audio into short time frames, locally enhances each frame of audio signal in a windowed manner, and extracts its short-time energy features. By calculating the energy level of the audio signal frame by frame, the system can construct a short-time energy sequence of each channel within a certain time range to reflect the current voice activity intensity. Subsequently, based on the preset energy threshold, the energy sequence of each channel is judged, and the proportion of active frames whose energy exceeds the threshold within the time window is counted, so as to quantify the voice activity frequency of each channel. The higher the voice activity frequency, the greater the possibility of continuous speaking behavior in the channel. This indicator can effectively eliminate the interference of short-term noise or non-speaking behavior on the judgment in the multi-person concurrent or turn-by-turn speaking conference scene. Finally, the main speaking initial determination module determines the microphone channel with the highest voice activity frequency as the main speaking channel based on the comparison results of the voice activity frequencies of each channel, and accurately locates the identity of the participant corresponding to the channel in combination with the temporary relationship table provided by the microphone association module.
[0026] Further, the sound output module is configured to perform the following steps: calculating a main channel gain coefficient of the current frame according to the short-time energy of the main speaking channel audio signal; performing short-time energy analysis on the audio signal of the non-main speaking channel frame by frame, calculating the short-time energy ratio between the non-main speaking channel audio signal and the main speaking channel audio signal, and determining the real-time attenuation coefficient corresponding to the non-main speaking channel according to the short-time energy ratio; based on the gain coefficient of the main channel and the real-time attenuation coefficient of the non-main channel, performing frame-by-frame weighted superposition of the audio signals of each channel to obtain a fused effective audio signal set, and outputting the set to the loudspeaker for sound output.
[0027] Specifically, first, for the currently identified main speaking channel, the audio signal thereof is extracted and divided by frames, the energy value corresponding to each frame is obtained through short-time energy calculation, and the gain coefficient of the main channel is further set according to the energy level. The coefficient is used to maintain the intelligibility and volume ratio of the main speaking signal during subsequent weighting processing, preventing the voice content from being weakened due to overall dynamic suppression. Then, the same short-time energy analysis is performed on the audio signals of all non-main speaking channels, and the energy values of the corresponding frames of the main channel are calculated by ratio. The ratio reflects the relative strength of the non-main speaking signal relative to the main channel. The system maps the ratio based on a preset nonlinear decay function or a proportional function (inverse function or sigmoid mapping), thereby generating a real-time attenuation coefficient of the non-main channel in the current frame. This processing ensures that dynamic suppression can be performed in time when there is slight environmental interference or low-amplitude noise, thereby improving the overall voice output quality. Finally, the sound output module performs weighted superposition processing on the audio frame signals of all channels according to the main channel gain coefficient and the attenuation coefficient of the non-main channel, to obtain the fused effective audio signal set. The audio set maintains the dominance of the main speaking content in the time domain or frequency domain, while significantly reducing the background noise or non-target speaking content that may exist in other channels. The processed audio signal is output in real time to the conference room loudspeaker, realizing clear and focused sound playback, and significantly improving the auditory experience of the participants and the communication efficiency of the meeting.
[0028] Further, based on the main speaking channel and the corresponding personnel identity identifier, the personnel region of the main speaking channel in the image data is subjected to behavior feature extraction of the line of sight direction, head orientation and hand movement, and a behavior feature set is generated, including the following steps: Based on the personnel identity identifier corresponding to the main speaking channel, the image region of the personnel in the current frame is extracted from the image data, and face key points and hand key points are positioned on the image region to obtain a standard key point coordinate set; Based on the key point coordinate set, a two-dimensional behavior feature matrix integrating head posture, line of sight direction and hand movement is constructed, each row of the two-dimensional behavior feature matrix representing the spatial position and relative relationship of each type of key point in a frame, and each column corresponding to different categories of action dimension information; The two-dimensional behavior feature matrix is input into a multi-layer residual neural network containing a residual connection structure, and local perception coding and time step residual information fusion are sequentially performed, and an intermediate behavior embedding representation containing action trend and interaction mode representation is output; The intermediate behavior embedding representation is subjected to full connection mapping and dimension compression processing, and a behavior feature vector of a unified structure is generated as a behavior feature set output.
[0029] In some embodiments, first, according to the identity of the person bound to the main speaking channel, the image area of the person is accurately extracted in the current image frame. The image area is usually the upper body area, and the face and the key parts of the hands are particularly concerned. The system calls a face key point detection algorithm (such as the deep learning-based MTCNN or HRNet model) to locate 68 standard key points of the face, and at the same time, a hand key point recognition model (MediaPipe Hands) is used to obtain the position data of 21 hand nodes. Finally, a key point coordinate set integrating the face and the hands is formed, and each key point is represented by two-dimensional or three-dimensional coordinates. Subsequently, a two-dimensional behavior feature matrix is constructed based on the key point set. Specifically, each row of the matrix represents the coordinate features of a certain type of key point in the image frame, such as the position of the eye corner, the vector from the top of the head to the chin, the angle between the fingers, etc.; each column corresponds to the dimension information of different time steps or key point categories, such as the head posture angle in the first column, the line of sight direction cosine value in the second column, and the left finger opening degree in the third column. The matrix is used to model both spatial structural relationships (such as left-right symmetry and angle changes) and temporal dynamic features (such as the start and end of actions and the duration). After construction, the two-dimensional behavior feature matrix is input into a deep residual neural network containing residual connection structures. Each residual block in the network is composed of two convolution layers and an identity mapping path, the former extracts local spatial variation patterns, and the latter preserves the original temporal feature information. To enhance the modeling capability of the time dimension, the network inserts a temporal convolution layer (Temporal Convolution) or a time attention mechanism based on the Transformer structure between every two residual blocks, so that the model can learn the mutual influence patterns between key points in different time periods. For example, the behavior of turning the head from left to right and opening the fingers can be identified in multiple frames to recognize the complete trend. The output of the network is an intermediate behavior embedding representation, which has a more compact clustering effect on similar behavior patterns in the feature space. The embedding vector is then nonlinearly mapped through a multi-layer perceptron (MLP) and dimensionally compressed through a feature selection mechanism (such as attention weighting or information entropy filtering), and finally a fixed-length behavior feature vector with clear semantics is generated. This vector can represent the intention of the current speaker, such as whether he is pointing at someone or performing a request action (such as raising his hand or signaling to speak), providing a structured input for interactive intention recognition and subsequent speech prediction.
[0030] Further, the speech recognition and semantic analysis processing on the set of valid speech audio signals to extract the interactive intention of the current speech content includes: Performing speech preprocessing and feature encoding on the audio signal corresponding to the main speaking channel and inputting it into a Transformer speech recognition model to output the text transcription result of the current speech; Perform named entity recognition and role anaphora resolution on the text transcription result, extract the names or titles of the participants appearing in the speech, and match them with the identity set generated by the personnel positioning module to obtain the target identity of the interaction; Based on the text content, construct a word vector sequence and input it into a semantic parsing network to extract the interaction type label and intent core phrase in the current speech, and obtain the interaction purpose corresponding to the current speech; Fuse the target identity and the interaction purpose to construct an interaction intent vector.
[0031] In some embodiments, first, perform speech preprocessing on the audio signal corresponding to the main speech channel, including mute segment rejection, noise filtering and audio normalization operation, to ensure the stability and clarity of the input signal in time domain and frequency domain. Subsequently, the processed audio signal is input into an end-to-end speech recognition model constructed using the Transformer structure. This model models the time sequence features of the input audio based on the attention mechanism, and outputs the corresponding text transcription result in combination with the phoneme-level decoder module. This text transcription has high semantic restoration ability and can accurately restore the key words and sentences, sentence structure and pause rhythm of the speaker's statement. After obtaining the text transcription result, the system uses the named entity recognition (NER) model in natural language processing to label and identify the names, positions, titles, etc. appearing in the text. For example, "General Li, how do you view this project" will identify "General Li" as a type of name / position entity. Subsequently, the system performs role anaphora resolution processing to solve the specific identity attribution of pronouns such as "he", "that", "this colleague" in the context. By matching the identified names or titles with the identity set generated by the personnel positioning module, the target personnel identity pointed to by the text semantics can be obtained, realizing the closed loop from natural language to structured identity mapping. Further, to clarify the pragmatic intent expressed by the speaker, the system tokenizes the text transcription result and constructs a word vector sequence through a word embedding model (BERT or Word2Vec) and inputs it into a custom semantic parsing network. This network usually uses a BiGRU+Attention structure, or a lightweight Transformer subnetwork, to perform classification on common speech intent types in a meeting scenario (such as requesting to speak, issuing instructions, soliciting opinions, topic shifting), and extract short phrase combinations containing semantic cores as intent expression cores. For example, "Please ask Teacher Zhang to talk about his opinion on this quarter's budget", will be parsed as "pointing to Teacher Zhang, the interaction purpose is to solicit opinions". Finally, the system fuses the above extracted target identity (such as the identity of Teacher Zhang) and the interaction purpose (such as "soliciting opinions") to generate an interaction intent vector in a unified format. This vector is one of the important state inputs in the reinforcement learning model, participating in the strategy judgment of the next speaker prediction.
[0032] Further, the voice preprocessing includes noise suppression, silence segment elimination and spectrum normalization.
[0033] Specifically, first, for the background noise that may exist in the conference environment (such as air conditioning sound, page turning sound, whispering, etc.), the system uses a combination of spectral subtraction and a neural network enhancement model (such as DCCRN or Wavesplit) to suppress noise. Among them, the spectral subtraction estimates the background noise spectrum of the non-speech frame and subtracts it from the speech spectrum to reduce static noise; while the DCCRN model is based on a trained time-frequency enhancement network, which can effectively restore the speech features of the speech signal and retain natural sound quality. Second, to improve the processing efficiency of the speech recognition system, the system performs voice activity detection (VAD) on the original speech signal. This processing uses a two-channel VAD network based on a gated recurrent unit (GRU) to jointly judge features such as short-time energy and mel-frequency cepstral coefficients (MFCC) of the speech, accurately identify non-speech frames and eliminate them. Through this step, the system can effectively reduce meaningless data input, reduce computational load, and improve the response speed and accuracy of the speech transcription model in multiple rounds of speaking. Finally, to ensure the consistency of audio signals from different microphones in amplitude and spectral distribution, the system performs spectrum normalization on the audio signal. The specific approach is: first, convert the audio signal to a spectral representation through short-time Fourier transform (STFT), then perform log amplitude compression and Z-score normalization on the spectrum by channel, so that the feature space input by different channels has uniform numerical distribution characteristics, thereby ensuring that the Transformer speech recognition model maintains stable recognition performance under multi-source input.
[0034] Further, the reinforcement learning model is constructed by the following steps: Based on the behavior feature set and the mutual intention vector, a state vector is constructed to describe the current conference state; Based on the identity set generated by the personnel positioning module, an action space is established with the identity of each participant as a discrete action; According to the matching between the reinforcement learning action output and the actual next speaker identity, the reward value is calculated through the reward function; The state vector is input into the policy network for forward calculation, outputting the probability distribution corresponding to each candidate action in the action space, and based on the reward value, the gradient value of the network parameters is calculated through the policy gradient algorithm, and the weight parameters of the policy network are updated according to the gradient value.
[0035] In some embodiments, first, in the state vector construction stage, the system receives a two-dimensional behavior feature matrix output by the behavior feature extraction module, which integrates the line-of-sight direction, head posture, and gesture action of the main speaker in the time sequence dimension. At the same time, the interaction feature extraction module outputs the interaction intention vector of the current speech content, which contains the target identity and semantic core phrase information. The system performs feature-level splicing or attention fusion (such as using a multi-head attention mechanism to jointly encode the behavior features and semantic vectors) on the above two features to construct a state vector with context semantic awareness capability, which is used to represent the context and behavior state of the current meeting. In the action space construction stage, the system sets the action space according to the identity set identified by the personnel positioning module, where each action represents the switching of the speaking right to a certain participant. Since the action space is limited and enumerable, the strategy network output is a probability distribution vector for all candidate actions. In the strategy network design stage, this embodiment uses a multi-layer feedforward neural network as the strategy network, with the state vector as the input and the predicted probability distribution π(a|s) for all candidate actions as the output. The network converts the logits after linear transformation into action selection probabilities through the softmax output layer. To enhance the model's ability to model long-time behavior sequences, a Transformer encoder structure can be used to pre-encode the state vector, and residual connections and layer normalization can be introduced to improve training stability. In the reward function design stage, the reward value is defined to measure the effectiveness of action selection. When the predicted next speaker identity matches the actual speaker identity, a positive reward is set; if not, a penalty is given. The construction of the reward value can further introduce soft factors such as interaction semantic matching degree, action transfer confidence, and delay response time to form a composite reward structure. The state vector in each meeting round is input into the strategy network, which generates a probability distribution for the candidate speakers based on the current parameters, and the current predicted action is sampled from the probability distribution. Then, the reward value is calculated based on the matching between the action and the actual speaker, which serves as the feedback signal for the optimization objective function. To improve the expected total reward of the strategy, the REINFORCE-based policy gradient method is used. This method estimates the gradient of the policy network parameters by maximizing the expected reward function under the current policy using the sampled actions and their corresponding rewards. Then, using the backpropagation algorithm, the network parameters are adjusted based on the gradient information, so that the policy network is more likely to output actions that bring high rewards in subsequent states. This strategy optimization mechanism combines the feedback information of actual meeting behavior and the probability modeling characteristics, not only improving the model's adaptive ability to speaking patterns, but also having the ability of continuous learning and generalization reasoning, thereby maintaining high efficiency in predicting the speaking channel in complex meeting scenarios such as multiple rounds of free interaction and frequent switching among multiple people.
[0036] Further, the formula of the reward function is as follows: ; wherein, is the immediate reward of the ith decision; is the actual speaker identity after the ith decision; is the highest probability candidate speaker identity output by the reinforcement learning model in the ith decision; is the matching function; is the state of the model; is the policy network probability distribution value given by the model to the actual speaker identity and are weight parameters.
[0037] It should be noted that the reward function compares the predicted result output by the model with the actual speaker identity, and judges whether the two are completely matched. If matched, a higher basic reward is given, and if not matched, the part of the reward is zero. At the same time, the system further investigates the probability value given by the model to the real speaker as a soft score of the prediction rationality. The probability value comes from the probability distribution output by the policy network to all candidate speakers under the current conference state, and the probability corresponding to the actual speaker is extracted as a measure of the policy confidence. Finally, the reward value is composed of two parts: one part is based on the hard matching result of whether the prediction is accurate, and the other part is based on the probability score of the model's confidence in the real speaker. The two parts are weighted and fused by a preset weight factor, so that the model can not only strengthen the accurate prediction behavior in the training process, but also obtain moderate incentive when close to the correct answer, and improve the learning stability and convergence speed.
[0038] The above embodiments only describe the preferred embodiments of the present application, and do not limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements of the technical solutions of the present application made by ordinary engineering technicians in the art shall fall within the protection scope determined by the claims of the present application.
Claims
1. A multimedia conference room sound system based on artificial intelligence, characterized in that: include: The personnel positioning module is used to perform face recognition and head posture estimation on the collected image data to generate the identity of each participant and the corresponding spatial location; A microphone association module is used to construct a temporary relationship table between participants and microphone channels based on the relationship between the spatial position and the microphone layout of the conference room; A main speaker initial determination module is used to detect the voice activity frequency of the audio data collected by each microphone channel and determine the main speaker channel and the corresponding person identity according to the temporary relationship table; An audio output module is configured to perform channel gain maintenance processing on the main speaking channel and gain suppression processing on the remaining channels to obtain a current valid audio signal set, and output the valid audio signal set as an audio signal output by a conference room speaker for audio output; An interaction feature extraction module is used to extract behavioral features such as gaze direction, head orientation, and hand movements from the person area in the main speaking channel in the image data based on the main speaking channel and the corresponding person identity, thereby generating a behavioral feature set; and to perform speech recognition and semantic parsing on the valid speaking audio signal set to extract the interactive intent of the current speech content; The main speaking channel update module is used to predict the next speaker through the reinforcement learning model based on the behavioral feature set and interaction intention, and update the main speaking channel according to the temporary relationship table, calculate the reward value based on the actual speech feedback results, and update the strategy parameters of the reinforcement learning model through the reward value.
2. The artificial intelligence-based multimedia conference room sound system according to claim 1, characterized in that: The personnel positioning module is used to perform the following steps: Based on the collected image data, perform face detection processing on the image data to obtain a set of face regions of all participants in the image frame; Based on the face region set, a time-series tracking process is performed to assign a unique temporary identifier to each face region according to the frame sequence, thereby obtaining a participant identity identifier and image position pair; According to the image position, head posture estimation processing is performed on each face area to obtain head posture parameters corresponding to each identifier, and invalid faces with a viewing angle deviation outside a threshold are filtered out in combination with the image position to obtain a face space projection parameter set after posture correction; According to the face space projection parameter set and the camera calibration parameters, the face center point and orientation information of each identifier are three-dimensionally mapped to obtain the identity identifier of each participant and the corresponding spatial position.
3. The multimedia conference room sound system based on artificial intelligence according to claim 1 is characterized in that: The microphone association module is used to perform the following steps: Based on the preset spatial positions of the conference room microphones, perform Euclidean distance calculation on the participant identification output by the personnel positioning module and the corresponding spatial position to obtain the set of distances between each participant and all microphones; Based on the distance set, perform spatial distance threshold screening and minimum distance priority matching to determine a temporary binding relationship between each participant and the microphone channel; According to the temporary binding relationship, a temporary relationship table that is updated in real time is constructed.
4. The artificial intelligence-based multimedia conference room sound system according to claim 1, characterized in that: The main speech initialization module is used to perform the following steps: Based on the audio data collected by each microphone channel, the audio data is framed and windowed, and the short-time energy of each frame of audio data is calculated to obtain the short-time energy sequence corresponding to each microphone channel; Performing energy threshold comparison and activity frame ratio calculation on the short-time energy sequence to obtain the real-time voice activity frequency of each microphone channel; Based on the real-time voice activity frequency, the microphone channel with the highest real-time voice activity frequency is determined as the main speaking channel, and the person identity corresponding to the main speaking channel is determined based on the temporary relationship table.
5. The artificial intelligence-based multimedia conference room sound system according to claim 4, characterized in that: The audio output module is used to perform the following steps: Calculate the main channel gain coefficient of the current frame according to the short-time energy of the main speech channel audio signal; Performing short-time energy analysis on the audio signal of the non-main speaking channel frame by frame, calculating the short-time energy ratio between the audio signal of the non-main speaking channel and the audio signal of the main speaking channel, and determining the real-time attenuation coefficient corresponding to the non-main speaking channel based on the short-time energy ratio; Based on the gain coefficient of the main channel and the real-time attenuation coefficient of the non-main channel, frame-by-frame weighted superposition of the audio signals of each channel is performed to obtain a fused effective audio signal set, which is output to the speaker for sound output.
6. The artificial intelligence-based multimedia conference room sound system according to claim 1, characterized in that: The extracting of behavioral features of the gaze direction, head orientation, and hand movements of the person region in the main speaking channel in the image data based on the main speaking channel and the corresponding person identity identifier to generate a behavioral feature set includes the following steps: Based on the person's identity corresponding to the main speaking channel, the image area of the person in the current frame is extracted from the image data, and the facial key points and hand key points are located in the image area to obtain a standard key point coordinate set; Based on the key point coordinate set, a two-dimensional behavior feature matrix is constructed that integrates head posture, gaze direction, and hand movements. Each row of the two-dimensional behavior feature matrix represents the spatial position and relative relationship of various key points in a frame, and each column corresponds to different categories of action dimension information. Inputting the two-dimensional behavior feature matrix into a multi-layer residual neural network with a residual connection structure, sequentially performing local perception encoding and time step residual information fusion, and outputting an intermediate behavior embedding representation containing action trend and interaction pattern representation; Fully connected mapping and dimension compression processing are performed on the intermediate behavior embedding representation to generate a behavior feature vector with a unified structure as an output of the behavior feature set.
7. The artificial intelligence-based multimedia conference room sound system according to claim 6, characterized in that: The performing of speech recognition and semantic analysis on the valid speech audio signal set to extract the interaction intent of the current speech content includes: Perform speech preprocessing and feature encoding on the audio signal corresponding to the main speaking channel, input it into the Transformer speech recognition model, and output the text transcription result of the current speech; Perform named entity recognition and role reference resolution on the text transcription results to extract the names or titles of the participants appearing in the speech, and match them with the identity identification set generated by the person location module to obtain the target identity of the interaction; Based on the text content, a word vector sequence is constructed and input into the semantic parsing network to extract the interaction type label and the core phrase of the intent in the current speech to obtain the interaction purpose corresponding to the current speech; The target identity and interaction purpose are integrated to construct the interaction intention vector.
8. The artificial intelligence-based multimedia conference room sound system according to claim 7, characterized in that: The speech preprocessing includes noise suppression, silent segment removal and spectrum normalization.
9. The artificial intelligence-based multimedia conference room sound system according to claim 7, characterized in that: The reinforcement learning model is constructed by the following steps: Based on the behavioral feature set and the mutual intention vector, a state vector is constructed to describe the current meeting status; Based on the identity set generated by the personnel positioning module, an action space is established with each participant's identity as a discrete action; The reward function calculates the reward value based on the match between the reinforcement learning action output and the actual next speaker identity; The state vector is input into the policy network to perform forward calculation, and the probability distribution corresponding to each candidate action in the action space is output. Based on the reward value, the gradient value of the network parameter is calculated by the policy gradient algorithm, and the weight parameters of the policy network are updated according to the gradient value.
10. The artificial intelligence-based multimedia conference room sound system according to claim 9, characterized in that: The formula of the reward function is as follows: ; in, is the immediate reward for the i-th decision; is the identity of the person who actually spoke after the i-th decision; is the identity of the candidate speaker with the highest probability output by the reinforcement learning model in the i-th decision; is the matching function; In state The following model identifies the actual speaker The probability distribution value of the given policy network; and is the weight parameter.
Citation Information
Patent Citations
Conference spokesman identity recognition method based on artificial intelligence technology
CN115100701A
Attention analysis method and system based on video conference system, and storage medium
CN116665111A
Conference data management system and method based on Internet of Things and microphone switching technology
CN116996337A
Multifunctional conference management system suitable for construction site
CN118741037A
Speaker separation method in intelligent conference system based on voiceprint recognition
CN119724217A
Cited By
Deep learning-based medical self-service check-in terminal interactor identity binding method and system
CN121389094A