Conference state control method, system and equipment for AI digital people
By acquiring meeting environment data and conducting multimodal analysis to generate meeting status assessment results, the problem of AI digital humans lacking proactive prediction in meeting control is solved, the natural connection and dynamic optimization of the meeting process are achieved, and meeting efficiency and user experience are improved.
Patent Information
- Application Number
- CN202510850980.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-17
AI Technical Summary
Existing AI digital humans lack an active prediction mechanism in meeting control, resulting in rigid control of meeting processes and an inability to dynamically adjust the pace or agenda according to the meeting environment.
By acquiring meeting environment data and performing multimodal analysis, the system generates meeting status assessment results, and generates meeting control instructions based on the assessment results, including process control operations such as adjusting the meeting rhythm, rescheduling the agenda, or switching roles.
It achieves comprehensive perception of the meeting scene, can proactively predict meeting needs, ensure the natural connection and dynamic optimization of the meeting process, improve the flexibility, naturalness and intelligence of meeting control, and enhance meeting efficiency and the interactive experience of participants.
Smart Images

Figure CN120812205A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video conferencing, in particular to a conference state control method for AI digital person, a conference state control system for AI digital person, an electronic device and a computer readable storage medium. BACKGROUND
[0002] With the rapid development of artificial intelligence, computer vision and speech recognition technologies, virtual characters are increasingly widely used in conference scenarios, such as serving as virtual hosts for conference guidance, process control and content recording. These technologies significantly improve conference management and collaboration efficiency, providing users with a more convenient interactive experience.
[0003] However, existing virtual characters, such as AI digital persons, have significant shortcomings in actual conference control applications, mainly manifested in the lack of proactive prediction mechanisms. Current AI digital persons mostly adopt a passive response mode, only performing corresponding operations after receiving explicit instructions from users, resulting in a rigid conference process control that lacks natural fluency and the ability to dynamically adjust pace or agenda according to the conference environment. SUMMARY
[0004] In view of the above problems, the present application embodiments are proposed to provide a conference state control method for AI digital person, a conference state control system for AI digital person, an electronic device and a computer readable storage medium that overcome the above problems or at least partially solve the above problems.
[0005] To solve the above problems, the present application embodiments disclose a conference state control method, which comprises:
[0006] acquiring conference environment data;
[0007] performing multi-modal analysis on the conference environment data to obtain a conference state evaluation result;
[0008] generating a conference control instruction according to the conference state evaluation result;
[0009] performing a conference process control operation according to the conference control instruction.
[0010] Optionally, the multi-modal analysis on the conference environment data to obtain a conference state evaluation result comprises:
[0011] extracting multi-class feature data from the conference environment data, the multi-class feature data comprising at least one of facial expression features identified based on a deep learning model, speech emotion features extracted based on an end-to-end speech recognition model, and gesture dynamic features identified based on a dynamic time warping algorithm;
[0012] inputting the multi-class feature data into a state evaluation model, and outputting the conference state evaluation result by the state evaluation model.
[0013] Optionally, the inputting the multi-class feature data into the state evaluation model and outputting the conference state evaluation result by the state evaluation model comprises:
[0014] inputting the multi-class feature data into the state evaluation model, identifying emotion types of the facial expression features by the state evaluation model, and / or performing semantic sentence segmentation and speech rhythm analysis on the speech emotion features, and / or performing interactive activity score on the gesture dynamic features to obtain the conference state evaluation result.
[0015] Optionally, the generating a conference control instruction according to the conference state evaluation result comprises:
[0016] mapping at least one of an emotion category identification result, a semantic speech analysis result, and an interactive activity score result to a conference state level based on a preset multi-dimensional state evaluation model;
[0017] generating the conference control instruction when the conference state level reaches a preset target level, the conference control instruction being used to adjust a conference rhythm, rearrange a conference agenda, or switch a conference role;
[0018] verifying the authenticity of the conference control instruction by a multi-modal consistency detection model.
[0019] Optionally, the generating the conference control instruction when the conference state level reaches a preset target level comprises:
[0020] continuously judging whether the conference state level reaches the target level within a preset control trigger evaluation period;
[0021] generating the conference control instruction when the conference state level reaches the target level for multiple times continuously within the control trigger evaluation period and a time interval of each time of judgment is within a preset time window.
[0022] Optionally, the performing a conference process control operation according to the conference control instruction comprises:
[0023] judging whether a current conference process node allows the conference process control operation to be inserted;
[0024] If the current conference process node allows the insertion of the conference process control operation, adjusting the conference rhythm, rearranging the conference agenda or switching the conference role through voice prompts or subtitle prompts, and aligning the audio-video time reference by using an optical manifold algorithm and a deep learning model;
[0025] If the current conference process node does not allow the insertion of the conference process control operation, the conference control instruction is cached for execution at the next conference process node that allows the insertion.
[0026] Optionally, after the conference process control operation is performed according to the conference control instruction, the method further comprises:
[0027] In the subsequent conference process, new conference environment data is continuously acquired and multi-modal analysis processing is performed to obtain a subsequent evaluation result;
[0028] When the subsequent evaluation result does not reach the target level continuously for multiple times within a preset control feedback evaluation period, the generation frequency of the conference control instruction is reduced, or the generation trigger condition of the conference control instruction is improved;
[0029] When the subsequent evaluation result continuously reaches the target level, the generation frequency of the conference control instruction is maintained, and the generation trigger condition of the conference control instruction is recorded;
[0030] In the subsequent conference process, if conference environment data matching the recorded generation trigger condition is acquired, a conference control instruction is directly generated based on the recorded generation trigger condition.
[0031] Embodiments of the present application also disclose a conference state control system, which comprises:
[0032] A data acquisition module is configured to acquire conference environment data;
[0033] A data analysis module is configured to perform multi-modal analysis on the conference environment data to obtain a conference state evaluation result;
[0034] An instruction generation module is configured to generate a conference control instruction according to the conference state evaluation result;
[0035] A conference control module is configured to perform a conference process control operation according to the conference control instruction.
[0036] Optionally, the data analysis module comprises:
[0037] The feature data extraction module is configured to extract multi-type feature data from the conference environment data, the multi-type feature data including at least one of facial expression features identified based on a deep learning model, speech emotion features extracted based on an end-to-end speech recognition model, and gesture dynamic features identified based on a dynamic time warping algorithm.
[0038] The state evaluation module is configured to input the multi-type feature data into a state evaluation model, and output the conference state evaluation result through the state evaluation model.
[0039] Optionally, the state evaluation module is configured to input the multi-type feature data into the state evaluation model, and perform emotion type identification on the facial expression features, and / or perform semantic sentence segmentation and speech rhythm analysis on the speech emotion features, and / or perform interactive activity level scoring on the gesture dynamic features through the state evaluation model to obtain the conference state evaluation result.
[0040] Optionally, the instruction generation module includes:
[0041] The state level mapping module is configured to map at least one of the emotion category identification result, the semantic speech analysis result, and the interactive activity level scoring result to a conference state level based on a preset multi-dimensional state evaluation model.
[0042] The control instruction generation module is configured to generate the conference control instruction when the conference state level reaches a preset target level, the conference control instruction being used to adjust a conference rhythm, rearrange a conference agenda, or switch a conference role.
[0043] The authenticity verification module is configured to verify authenticity of the conference control instruction through a multi-modal consistency detection model.
[0044] Optionally, the control instruction generation module includes:
[0045] The state level judgment module is configured to continuously judge whether the conference state level reaches the target level within a preset control trigger evaluation period.
[0046] The conference control instruction generation module is configured to generate the conference control instruction when the conference state level reaches the target level for multiple times continuously and a time interval of each time of judgment is within a preset time window within the control trigger evaluation period.
[0047] Optionally, the conference control module includes:
[0048] The control operation judgment module is configured to judge whether a current conference process node allows insertion of the conference flow control operation.
[0049] The control operation insertion module is configured to adjust conference rhythm, rearrange conference agenda or switch conference roles through voice prompt or subtitle prompt, and align audio and video time base by using an optical manifold algorithm and a deep learning model if the current conference process node allows insertion of the conference process control operation.
[0050] The control instruction cache module is configured to cache the conference control instruction to a next conference process node that allows insertion of the conference process control operation if the current conference process node does not allow insertion of the conference process control operation.
[0051] Optionally, the system further comprises:
[0052] The re-evaluation module is configured to continue to acquire new conference environment data and perform multi-modal analysis processing to obtain a subsequent evaluation result in a subsequent conference process after the conference control module performs the conference process control operation according to the conference control instruction.
[0053] The frequency condition adjustment module is configured to reduce a generation frequency of the conference control instruction or improve a generation trigger condition of the conference control instruction when the subsequent evaluation result does not reach the target level continuously for multiple times within a preset control feedback evaluation period.
[0054] The frequency condition maintenance module is configured to maintain the generation frequency of the conference control instruction and record the generation trigger condition of the conference control instruction when the subsequent evaluation result continuously reaches the target level.
[0055] The instruction direct generation module is configured to directly generate the conference control instruction based on the recorded generation trigger condition if conference environment data matching the recorded generation trigger condition is acquired in the subsequent conference process.
[0056] The embodiment of the application further discloses an electronic device, comprising: one or more processors; and one or more machine-readable media having instructions stored thereon that, when executed by the one or more processors, cause the electronic device to perform the conference state control method.
[0057] The embodiment of the application further discloses a computer-readable storage medium storing a computer program that causes a processor to perform the conference state control method.
[0058] The embodiment of the application has the following advantages:
[0059] The conference state control scheme for an AI digital person provided by the embodiment of the application acquires conference environment data, performs multi-modal analysis on the conference environment data to obtain a conference state evaluation result, generates a conference control instruction according to the conference state evaluation result, and performs a conference process control operation according to the conference control instruction.
[0060] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0061] The embodiments of the present application realize comprehensive perception of the conference scene by acquiring conference environment data. The conference state evaluation result is generated through multi-modal analysis, which can capture the dynamic information of the participants in real time. The conference control instruction is generated based on the evaluation result, which can actively predict the conference demand and generate the corresponding instruction. The conference process control operation is performed according to the conference control instruction, which ensures the natural connection and dynamic optimization of the conference process. The embodiments of the present application overcome the process rigidity problem caused by passive response in the prior art, significantly improve the flexibility, naturalness and intelligent level of conference control, thereby effectively improving the conference efficiency and enhancing the interactive experience of the participants, and adapting to the dynamic change demand in the complex conference scene. BRIEF DESCRIPTION OF DRAWINGS
[0062] Figure 1 is a step flow chart of a conference state control method for an AI digital person according to an embodiment of the present application;
[0063] Figure 2 is a step flow chart of an AI digital person conference control method according to an embodiment of the present application;
[0064] Figure 3 is a structural block diagram of a conference state control system for an AI digital person according to an embodiment of the present application. DETAILED DESCRIPTION
[0065] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0066] The embodiments of the present application provide a conference state control scheme for an AI digital person. By acquiring conference environment data including audio and video data, text data and network transmission data, etc., multi-modal analysis technology is used to extract features such as facial expressions, speech emotions and gesture dynamics in real time, generate conference state evaluation results, and generate conference control instructions based on the conference state evaluation results, and then execute conference control instructions to realize process control operations such as conference rhythm adjustment, agenda rearrangement or role switching. The embodiments of the present application realize audio and video quality optimization, content authenticity verification and multi-language support through deep learning models, speech recognition technology and optical manifold algorithm, significantly improve the natural smoothness, intelligent level and user experience of the conference process, and effectively solve the deficiencies of passive response and lack of prediction mechanism of traditional AI digital person conference control system.
[0067] Reference Figure 1, a step flow chart of a conference state control method for an AI digital person (referred to as a conference state control method) is shown. The conference state control method for an AI digital person can be applied to a conference control system, a conference management system, and the like (hereinafter referred to as a system). The conference state control method for an AI digital person can specifically include the following steps:
[0068] Step 101, conference environment data is acquired.
[0069] First, the system configures conference parameters, including participant information (such as name, position, company), conference theme (such as conference name, topic classification), and time setting (such as conference time, break time). These parameters are completed through user input or preset templates to ensure accurate conference environment initialization. Second, the system deploys conference equipment, including a 1080P high-definition camera for collecting video data, an omnidirectional microphone array for collecting audio data, and a real-time caption recording system for generating text data. The video data uses a high-resolution format to capture the facial expressions and hand gestures of the participants, the audio data is collected through the omnidirectional microphone array to ensure multi-directional sound collection, and the text data records the conference content and caption information. In addition, the system establishes a distributed data storage system that supports local storage and cloud storage for storing audio and video data, text data, and metadata (such as conference timestamps, device status). The distributed storage architecture uses high availability and high security design to ensure the reliability and accessibility of data during the conference process. The system also acquires network transmission data through real-time bandwidth testing and delay measurement to monitor network quality to support subsequent optimization.
[0070] Step 102, multi-modal analysis is performed on the conference environment data to obtain conference state evaluation results.
[0071] Firstly, the system extracts multi-class feature data from the meeting environment data, including facial expression features, speech emotion features, and gesture dynamic features. Facial expression features are extracted through a deep learning-based facial expression recognition model (FERM), which recognizes six basic expressions (happy, joy, delight, ease, sadness, and fear). Speech emotion features are extracted through an end-to-end speech recognition model, which analyzes features such as fundamental frequency, amplitude, and duration, and combines semantic segmentation and speech rhythm analysis. Gesture dynamic features are identified through a dynamic time warping algorithm (DTW), which recognizes ten common gestures such as nodding and waving, and is used to assess interaction activity. Secondly, the system inputs the extracted multi-class feature data into a state assessment model, which combines convolutional neural networks (CNN) and recurrent neural networks (RNN) to analyze facial expressions, speech emotions, and gesture dynamics, and outputs meeting state assessment results. In addition, the system uses adaptive noise cancellation technology (ANC) and a deep learning-based speech enhancement model to remove background noise (such as keyboard tapping and air conditioner noise), improving speech clarity and signal-to-noise ratio. Network transmission quality is evaluated through real-time bandwidth testing and delay measurement, recording key moment information (such as speaker switching and emotional peak moments).
[0072] Step 103, generating meeting control instructions according to the meeting state assessment results.
[0073] The system adopts a preset multi-dimensional state evaluation model to map the generated facial expression recognition results, speech emotion analysis results, and gesture interaction activity scores into conference state levels. This model can be based on Bidirectional Encoder Representations from Transformers (BERT) or Generative Pretrained Transformer (GPT) to analyze high-frequency discussion topics and emotional peaks. For example, when detecting that the participants frequently exhibit happy expressions and active gestures, the system may map to a "high activity" state level. When the conference state level reaches the preset target level (such as "adjust the pace" or "switch the agenda"), the system generates corresponding conference control instructions, including adjusting the conference pace (such as speeding up or slowing down the discussion), rearranging the conference agenda (such as prioritizing high-frequency topics), or switching conference roles (such as switching from the moderator to the speaker). To ensure the reliability of the instructions, the system verifies the authenticity of the instructions through a multi-modal consistency detection model that combines speech features and facial expression features using a Cycle-Consistent Loss Function (CycleGAN) to identify potential fake content. In addition, the system continuously determines whether the conference state level continuously reaches the target level within the preset control trigger evaluation period, and the time interval for each determination must be within the preset time window to avoid false triggering.
[0074] Step 104, perform conference process control operations according to conference control instructions.
[0075] First, the system determines whether the current conference progress node allows the insertion of control operations, such as checking whether it is in the topic discussion, break section, or key speech stage. If the current node allows insertion, the system performs control operations through voice prompts or subtitle prompts, such as playing prompt voice to adjust the conference pace, rearranging the agenda through subtitle prompts, or switching conference roles (such as switching from an AI moderator to a designated participant to speak). To ensure audio-video synchronization, the system uses an Optical Flow Algorithm (OFA) and a deep learning model to align the speech and video time references, avoiding the disconnection between the prompts and the conference content. If the current node does not allow insertion, the system caches the control instructions to the next node that allows insertion for execution, such as caching the "switch agenda" instruction after the current speech ends.
[0076] For example, in a video conference, the system detects that network delays cause interruptions in speaking, and inserts a caption "Please wait, network optimization", and adjusts the agenda sequence after the network is restored. After the control operation is completed, the system generates the optimized conference highlights content, encodes the video with an efficient encoding standard, adds metadata such as titles, descriptions, and tags, and ensures the high-quality output of the content. This step significantly improves the natural smoothness of the conference process and the user experience through intelligent execution.
[0077] The embodiment of the present application realizes comprehensive perception of the conference scene by acquiring conference environment data. The conference state evaluation result is generated through multi-modal analysis, which can capture the dynamic information of the participants in real time. The conference control instruction is generated based on the evaluation result, which can actively predict the conference demand and generate the corresponding instruction. The conference process control operation is executed according to the conference control instruction, which ensures the natural connection and dynamic optimization of the conference process. The embodiment of the present application overcomes the process rigidity problem caused by passive response in the prior art, significantly improves the flexibility, naturalness and intelligent level of conference control, thereby effectively improves the conference efficiency, enhances the interactive experience of the participants, and adapts to the dynamic change demand in complex conference scene.
[0078] In an exemplary embodiment of the present application, the multi-modal analysis is performed on the conference environment data to obtain the conference state evaluation result. One implementation of the multi-modal analysis is as follows: multi-class feature data is extracted from the conference environment data, the multi-class feature data includes at least one of the following: facial expression features identified based on a deep learning model, speech emotion features extracted based on an end-to-end speech recognition model, and gesture dynamic features identified based on a dynamic time warping algorithm; the multi-class feature data is input into a state evaluation model, and the conference state evaluation result is output by the state evaluation model.
[0079] Firstly, the system extracts multi-class feature data from the conference environment data, including at least one of the following: facial expression features, which are extracted by FERM to identify six basic expressions (happy, joy, delight, relaxed, sad, and fear) to capture the emotions of the participants; for example, in a product launch conference, the system captures the smiles or frowns of the participants through a 1080P high-definition camera to judge the reaction to the new product. Speech emotion features are extracted by E2E-SRM to analyze features such as fundamental frequency, amplitude, and duration, and combined with semantic sentence segmentation and rhythm analysis to reflect the speaker's emotion; for example, fast speech speed may indicate intense argument. Gesture dynamic features are identified by DTW to evaluate the interaction activity, such as nodding and waving hands; for example, frequent gestures indicate high participation. The extraction process uses ANC to remove background noise (such as air conditioner noise), and SEM to improve speech clarity to ensure accurate feature data.
[0080] Secondly, the system inputs multiple types of feature data into the state assessment model. The model combines CNN and RNN to comprehensively analyze facial expressions, voice emotions and gesture dynamics to generate meeting state assessment results, including meeting atmosphere (positive or dull), interactive activity (high or low) and network quality score.
[0081] For example, during a video conference, the system might output a result of "high activity, slight network latency," providing a basis for generating subsequent instructions. Network quality is assessed through real-time bandwidth testing and latency measurements, recording key moments (such as speaker switching).
[0082] This implementation utilizes multimodal feature extraction and combined CNN and RNN analysis to systematically perceive the dynamics of meeting scenes, overcoming the limitations of traditional AI digital humans, which rely solely on a single data source (such as voice commands). Comprehensive analysis of multiple features (such as facial expressions, voice, and gestures) enhances assessment accuracy. The processing power of CNN and RNN ensures real-time and in-depth analysis, laying the foundation for generating precise control commands and significantly improving the intelligence, naturalness, and user experience of meeting processes.
[0083] In an exemplary embodiment of the present invention, multiple categories of feature data are input into a state assessment model, and an implementation method for outputting a conference state assessment result through the state assessment model is: multiple categories of feature data are input into the state assessment model, and the state assessment model is used to identify the emotion type of facial expression features, and / or, to perform semantic segmentation and speech rhythm analysis on the voice emotion features, and / or, to perform interactive activity scoring on the gesture dynamic features to obtain the conference state assessment result.
[0084] The system extracts multiple types of feature data from the meeting environment data, including facial expression features, voice emotion features, and gesture dynamic features, and inputs them into the state assessment model, which is processed by combining convolutional neural networks and recurrent neural networks. The specific operations include at least one of the following: (1) Emotion type recognition of facial expression features, using the facial expression recognition model to analyze six basic expressions (happiness, joy, happiness, comfort, sadness, and fear); for example, in a project review meeting, if the system detects that most participants show "sad" expressions, it may indicate that the discussion has reached an impasse.
[0085] (2) Perform semantic punctuation and speech rhythm analysis on the speech emotion features. Extract features such as fundamental frequency, amplitude, and duration through an end-to-end speech recognition model, and perform semantic punctuation and speech rhythm analysis using natural language processing (NLP) technology. For example, in a technical seminar, fast speech and frequent punctuation may reflect heated debate. (3) Perform interactive activity score on gesture dynamic features. Identify ten common gestures (such as nodding and waving) through dynamic time warping algorithm, and calculate activity score based on gesture frequency and amplitude. For example, frequent waving may indicate high participation. To ensure data quality, the system uses adaptive noise suppression technology to remove background noise (such as air conditioner noise) before feature extraction, and uses a speech enhancement model to improve speech clarity. The state evaluation model analyzes the above features comprehensively and outputs the conference state evaluation results, including conference atmosphere (such as positive or dull), interactive activity (such as high or low), and network quality score. For example, in a video conference, the system may output the result of "positive atmosphere, high activity, and slight network delay", which provides the basis for subsequent generation of conference control instructions.
[0086] This embodiment accurately captures the dynamic state of the conference through the state evaluation model's targeted analysis of multiple features (expression, speech, gesture), overcoming the one-sidedness of traditional AI digital people relying solely on a single feature (such as speech instructions). The joint processing of CNN and RNN enhances the depth and timing of feature analysis, the semantic analysis of NLP improves the accuracy of speech understanding, and the scoring mechanism of DTW quantifies the interactive activity, thereby generating a comprehensive and reliable conference state evaluation result, providing a solid foundation for actively optimizing the conference process, and significantly improving the intelligence and natural flow of the conference.
[0087] In an exemplary embodiment of the present application, one embodiment of generating conference control instructions based on conference state evaluation results is as follows: based on a pre-set multi-dimensional state evaluation model, at least one of the emotion category recognition result, the semantic speech analysis result, and the interactive activity score result is mapped to a conference state level; when the conference state level reaches a pre-set target level, a conference control instruction is generated, the conference control instruction is used to adjust the conference rhythm, rearrange the conference agenda, or switch the conference role; the authenticity of the conference control instruction is verified through a multi-modal consistency detection model.
[0088] Firstly, the system adopts a preset multi-dimensional state evaluation model to map at least one of the analysis results of multi-class feature data (including emotion category recognition results, semantic speech rate analysis results, and interactive activity level score results) into a conference state level. This model is based on a bidirectional encoder representation transformation model or a generative pre-training transformation model, which generates a conference state level by analyzing emotion categories (such as happy, sad), semantic speech rate (such as fast or slow), and interactive activity level (such as high or low). For example, in a product planning conference, the system detects that the participants' emotions are "sad" and the speech rate is slow, which may be mapped to a "low activity" state level. Secondly, when the conference state level reaches a preset target level (such as "adjust the pace" or "switch the agenda"), the system generates a conference control instruction, and the instruction content includes adjusting the conference pace (such as speeding up the discussion), rearranging the conference agenda (such as prioritizing high-frequency topics), or switching the conference role (such as switching from an AI host to a specific participant). To ensure the reliability of the instruction, the system verifies the authenticity of the instruction through a multi-modal consistency detection model, which combines facial expression features and speech emotion features, uses a cyclic consistency loss function to identify potential fake content, and ensures that the instruction is generated based on real data. In addition, the system continuously determines whether the state level continuously reaches the target level within a preset control trigger evaluation period, and the time interval of each determination needs to be within a preset time window to avoid false triggering. For example, in a video conference, if the system continuously detects a "high activity" state, it may generate the instruction "extend the current topic discussion time".
[0089] This embodiment overcomes the passive response limitations of traditional AI digital humans that rely only on explicit user instructions by using the mapping mechanism of the multi-dimensional state evaluation model and the authenticity verification. The semantic analysis capabilities of BERT and GPT ensure accurate mapping of state levels, and the verification mechanism improves the reliability of instructions, thereby achieving active optimization of conference processes and significantly improving the intelligence, natural flow, and user experience of conferences.
[0090] In an exemplary embodiment of the present application, when the conference state level reaches a preset target level, one embodiment of generating a conference control instruction is: continuously determining whether the conference state level reaches the target level within a preset control trigger evaluation period; when the conference state level continuously reaches the target level multiple times within the control trigger evaluation period, and the time interval of each determination is within a preset time window, a conference control instruction is generated.
[0091] The system continuously determines whether the conference state level generated by the multi-dimensional state evaluation model reaches the preset target level, such as "high activity" (reflecting intense discussion) or "low activity" (reflecting stagnant discussion), within a preset control trigger evaluation period (e.g., evaluated once every 10 seconds). The conference state level is based on the conference state evaluation results obtained through multi-modal analysis, taking into account the emotion category recognition results, semantic speech rate analysis results, and interactive activity score. To avoid false triggering due to transient fluctuations, the system requires the conference state level to reach the target level continuously for multiple times (e.g., 3 times) within the control trigger evaluation period, and the time interval for each determination must be within a preset time window (e.g., 5 seconds to 15 seconds) to ensure the stability of the state. For example, in a video conference, if the system continuously detects "low activity" state for 3 times (due to the silence or slow speech of the participants), and the interval between each detection is 10 seconds, an instruction "insert prompt voice to activate discussion" is generated. The generated conference control instruction can be used to adjust the conference pace (such as speeding up the discussion), rearrange the conference agenda (such as switching to a more attractive topic), or switch the conference role (such as transferring the hosting right to an active participant). After the instruction is generated, the system verifies the instruction authenticity through a multi-modal consistency detection model to ensure that it is generated based on real data. For example, the system can avoid false operations caused by temporary network delays through continuous determination.
[0092] This embodiment uses a continuous judgment mechanism of control trigger evaluation period and time window to effectively filter transient state fluctuations, improving the generation accuracy and stability of conference control instructions. The periodic evaluation combined with the analysis capabilities of BERT and GPT ensures the reliability of the state level, thereby achieving precise dynamic adjustment of the conference process and significantly improving the intelligence, natural flow, and user participation experience of the conference.
[0093] In an exemplary embodiment of the present application, one implementation of performing conference process control operations according to conference control instructions is: determining whether the current conference process node allows the insertion of conference process control operations; if the current conference process node allows the insertion of conference process control operations, adjusting the conference pace, rearranging the conference agenda, or switching the conference role through voice or subtitle prompts, and aligning the audio and video time references using optical manifold algorithms and deep learning models; if the current conference process node does not allow the insertion of conference process control operations, the conference control instructions are cached and executed at the next conference process node that allows insertion.
[0094] Firstly, the system determines whether the current conference process node allows the insertion of conference flow control operations. The conference process node includes stages such as topic discussion, break, and key speech. The feasibility of insertion is determined by preset rules (such as conference agenda timetable or priority setting). For example, if the current node is "main speaker speech", the system may determine not to insert the prompt to avoid interruption. If the current node allows insertion, the system executes the conference control instruction through voice prompts (such as playing "please speed up the discussion pace") or subtitle prompts (such as displaying "switch to the next topic"). The specific operations include adjusting the conference pace (such as speeding up or slowing down the discussion speed), rearranging the conference agenda (such as prioritizing high-frequency topics), or switching conference roles (such as switching from an AI host to a designated participant). To ensure that the prompt is synchronized with the conference content, the system uses an optical manifold algorithm and a deep learning model to align the audio and video time base, avoiding the disconnection of voice or subtitles from video content. In addition, the system provides multilingual support through a neural network-based machine translation model, generating multilingual voice or subtitles to meet the needs of international conferences; through adaptive noise suppression technology, it filters background noise (such as environmental noise), and uses a dynamic volume adjustment algorithm (Dynamic Volume Adjustment Algorithm, DVAA) to dynamically adjust the volume of background music according to the conference sentiment and scene. If the current node does not allow insertion, the system caches the conference control instruction to the next node that allows insertion, for example, caching the "rearrange agenda" instruction to execute after the current speech ends. For example, the system may cache the "optimize network prompt" instruction after detecting network delay, and insert the prompt after the network stabilizes.
[0095] This embodiment ensures the timely execution of conference control instructions by judging the insertion feasibility of conference process nodes and the audio and video synchronization mechanism of OFA and CNN. The node judgment and caching mechanism improves the accuracy of instruction execution, and the use of NMT and DVAA enhances the multilingual adaptability and sound optimization, thereby significantly improving the natural flow, intelligence level and user experience of the conference process.
[0096] In an exemplary embodiment of the present application, after performing the conference process control operation according to the conference control instruction, one implementation is to continue to obtain new conference environment data and perform multi-modal analysis processing to obtain subsequent evaluation results in the subsequent conference process; when the subsequent evaluation results do not reach the target level for continuous multiple times within the preset control feedback evaluation period, the generation frequency of the conference control instruction is reduced, or the generation trigger condition of the conference control instruction is improved; when the subsequent evaluation results continuously reach the target level, the generation frequency of the current conference control instruction is maintained, and the generation trigger condition of the current conference control instruction is recorded; in the subsequent conference process, if the conference environment data matching the recorded generation trigger condition is obtained, the conference control instruction is directly generated based on the recorded generation trigger condition.
[0097] Firstly, the system continues to obtain new conference environment data, including audio and video data, text data and network transmission data, in the subsequent conference process, and performs multi-modal analysis to generate subsequent conference state evaluation results. The multi-modal analysis utilizes a facial expression recognition model to recognize six basic expressions (such as happy, sad), an end-to-end speech recognition model combined with natural language processing technology to analyze speech speed and semantics, a dynamic time warping algorithm to evaluate gesture activity, and a convolutional neural network and a recurrent neural network to generate evaluation results. Secondly, the system analyzes whether the subsequent conference state evaluation results continuously reach the preset target level (such as "high activity") for multiple times (such as 3 times) within the preset control feedback evaluation period (such as every 15 seconds). If not, the system reduces the generation frequency of the conference control instruction (such as from once every minute to once every two minutes) or improves the generation trigger condition threshold (such as requiring higher emotional intensity); if the target level is continuously reached, the system maintains the current generation frequency and records the generation trigger condition (such as a specific emotion and gesture combination). Finally, if the subsequent conference environment data matches the recorded generation trigger condition, the system directly generates the conference control instruction (such as "speed up the discussion pace") based on the recorded condition, without the need for reanalysis. For example, if the system detects that the participants are continuously silent (not reaching the target level), the instruction frequency can be reduced to reduce interference; if it detects intense discussion (continuously reaching the target), the trigger condition is recorded to quickly respond to similar scenarios. After the instruction is generated, the authenticity is verified by a multi-modal consistency detection model.
[0098] This implementation realizes adaptive conference control by continuously monitoring and dynamically adjusting the instruction generation strategy, overcoming the excessive or insufficient intervention caused by fixed instruction generation in traditional AI digital human systems. Continuous evaluation and threshold adjustment improve the adaptability of the instruction, trigger condition recording and reuse improve the response efficiency, thereby significantly enhancing the intelligence, flexibility of the conference process and continuous optimization of user experience.
[0099] Based on the above related description of the conference state control method for AI digital person, the AI digital person conference control method is introduced below, which is applied to a conference control system and used for intelligent and active conference process management. The implementation process of the method is described in detail below in combination with specific steps.
[0100] Referring to Figure 2 , a step flow diagram of an AI digital person conference control method according to an embodiment of the application is shown.
[0101] Step 201, initializing the conference environment and collecting conference environment data
[0102] The system first initializes the conference environment and collects conference environment data by configuring conference parameters and deploying hardware devices. The conference parameters include participant information (such as name, department, role), conference theme (such as project review, technical discussion), and time setting (such as conference duration, stage division). The hardware devices include high-resolution video acquisition devices (1080P high-definition cameras), audio acquisition devices (omnidirectional microphone arrays), and real-time subtitle acquisition devices (supporting text transcription). The video acquisition device captures the facial expressions and actions of the participants, the audio acquisition device records the speech content, and the subtitle acquisition device generates real-time conference records. The system establishes a distributed storage architecture, including local storage and cloud storage modules, for storing the collected audio and video data, text data, and metadata (such as timestamps, device status). In addition, the system collects network transmission data in real time through a network monitoring module, including bandwidth, delay, and packet loss rate. For example, in a multi-party remote conference, the system can capture the dynamics of the participants through the camera, record the discussion content through the microphone array, and monitor the network fluctuations to ensure data integrity.
[0103] Step 202, real-time analysis of conference environment data and generation of quality evaluation report
[0104] The system performs real-time analysis on the collected conference environment data to generate a quality assessment report to evaluate the conference operation status. The analysis process includes: (1) Video analysis: using deep learning-based video analysis algorithms to detect changes in facial expressions (such as smiling, frowning) and body movements (such as nodding) of participants, generating expression and movement feature reports. (2) Audio analysis: through voice analysis technology, extracting the pitch, volume and speed features of the speaker, and identifying the background noise level; the system uses noise suppression technology to remove interference (such as fan noise). (3) Network analysis: through real-time network monitoring, evaluating bandwidth utilization and transmission delay, generating network quality scores. The analysis results are summarized into a quality assessment report, including expression activity, audio clarity and network stability indicators. For example, in a technical seminar, the system may generate a report showing "high expression activity, good audio clarity, and slight network delay", providing a basis for subsequent optimization. During the analysis process, the system uses data encryption technology to protect data security.
[0105] Step 203, optimizing conference parameters based on quality assessment report
[0106] The system dynamically optimizes conference parameters based on the quality assessment report to improve conference effectiveness. If the report shows that optimization is needed (such as low audio clarity or high network delay), the system adjusts the relevant parameters: (1) Video optimization: through adaptive resolution adjustment algorithm, adjusting video resolution according to network conditions (such as from 1080P to 720P). (2) Audio optimization: through dynamic audio enhancement technology, improving voice clarity and suppressing background noise. (3) Agenda optimization: according to expression activity and discussion heat, adjusting agenda priority (such as extending the time of popular topics). After optimization, the system returns to step 202 for continuous monitoring; if no optimization is needed, it goes to step 204.
[0107] Step 204, generating and outputting optimized conference content
[0108] The system generates and outputs conference content based on conference parameters, including conference records and highlights videos. The generation process includes: (1) Content organization: through text analysis technology, semantically organizing real-time subtitles and speech content, extracting key discussion points. (2) Video editing: using intelligent editing algorithms, automatically editing highlight segments according to discussion heat and expression activity. (3) Output processing: through efficient video encoding technology, encoding highlight videos and adding metadata (such as conference title, time). The system supports multi-language output, generating multi-language subtitles through neural network-based machine translation models. The final content is distributed to participants through cloud storage.
[0109] It should be noted that for the method embodiments, the methods can be described as a series of acts combined to achieve the intended purpose, however, the person skilled in the art should understand that the present application is not limited to the order of the acts described, because depending on the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, the person skilled in the art should understand that the embodiments described in the specification are all preferred embodiments, and the acts involved are not necessarily essential to the embodiments of the present application.
[0110] Referring to Figure 3 , a structural block diagram of a conference state control system for an AI digital person (referred to as a conference state control system) is shown. The conference state control system for the AI digital person can specifically include the following modules.
[0111] The data acquisition module 31 is configured to acquire conference environment data.
[0112] The data analysis module 32 is configured to perform multi-modal analysis on the conference environment data to obtain a conference state evaluation result.
[0113] The instruction generation module 33 is configured to generate a conference control instruction according to the conference state evaluation result.
[0114] The conference control module 34 is configured to perform conference process control operations according to the conference control instruction.
[0115] In an exemplary embodiment of the present application, the data analysis module 32 includes:
[0116] The feature data extraction module is configured to extract multi-class feature data from the conference environment data, the multi-class feature data including at least one of facial expression features identified based on a deep learning model, speech emotion features extracted based on an end-to-end speech recognition model, and gesture dynamic features identified based on a dynamic time warping algorithm.
[0117] The state evaluation module is configured to input the multi-class feature data into a state evaluation model, and output the conference state evaluation result through the state evaluation model.
[0118] In an exemplary embodiment of the present application, the state evaluation module is configured to input the multi-class feature data into the state evaluation model, identify the emotion type of the facial expression features through the state evaluation model, and / or perform semantic sentence segmentation and speech rhythm analysis on the speech emotion features, and / or score the interactive activity level of the gesture dynamic features to obtain the conference state evaluation result.
[0119] In an exemplary embodiment of the present application, the instruction generation module 33 includes:
[0120] a state level mapping module, configured to map at least one of the emotion category recognition result, the semantic speech rate analysis result and the interaction activity score result into a conference state level based on a preset multi-dimensional state evaluation model;
[0121] a control instruction generation module, configured to generate the conference control instruction when the conference state level reaches a preset target level, the conference control instruction being used to adjust a conference rhythm, rearrange a conference agenda or switch a conference role;
[0122] a reality verification module, configured to verify the reality of the conference control instruction through a multi-modal consistency detection model.
[0123] In an exemplary embodiment of the present application, the control instruction generation module comprises:
[0124] a state level judgment module, configured to continuously judge whether the conference state level reaches the target level within a preset control trigger evaluation period;
[0125] a conference control instruction generation module, configured to generate the conference control instruction when the conference state level reaches the target level for multiple times continuously and the time interval of each time is within a preset time window within the control trigger evaluation period.
[0126] In an exemplary embodiment of the present application, the conference control module 34 comprises:
[0127] a control operation judgment module, configured to judge whether a current conference process node allows insertion of the conference process control operation;
[0128] a control operation insertion module, configured to adjust the conference rhythm, rearrange the conference agenda or switch the conference role through voice prompt or subtitle prompt if the current conference process node allows insertion of the conference process control operation, and align the audio and video time benchmarks by using an optical manifold algorithm and a deep learning model;
[0129] a control instruction cache module, configured to cache the conference control instruction to a next conference process node allowing insertion of the conference process control operation for execution if the current conference process node does not allow insertion of the conference process control operation.
[0130] In an exemplary embodiment of the present application, the system further comprises:
[0131] a reevaluation module, configured to continue to acquire new conference environment data and perform multi-modal analysis processing to obtain a subsequent evaluation result in a subsequent conference process after the conference control module 34 performs the conference process control operation according to the conference control instruction.
[0132] a frequency condition adjusting module, configured to reduce a generation frequency of the conference control instruction, or improve a generation trigger condition of the conference control instruction, when the subsequent evaluation result fails to reach the target level continuously for a plurality of times within a preset control feedback evaluation period;
[0133] a frequency condition maintaining module, configured to maintain the current generation frequency of the conference control instruction, and record the generation trigger condition of the conference control instruction, when the subsequent evaluation result continuously reaches the target level;
[0134] an instruction directly generating module, configured to generate the conference control instruction directly based on the recorded generation trigger condition, if conference environment data matching the recorded generation trigger condition is acquired in a subsequent conference process.
[0135] For the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts are referred to the part of the method embodiment.
[0136] Each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts between the embodiments are referred to each other.
[0137] Those skilled in the art should understand that the embodiments of the embodiments of the present application can be provided as a method, device or computer program product. Therefore, the embodiments of the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0138] The embodiments of the present application are described with reference to flowcharts and / or block diagrams according to the method, terminal device (system) and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of the flows and / or blocks in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal device produce a machine that implements the functions specified in the flowchart and / or block diagram. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The device for realizing the functions specified in one flow or multiple flows and / or blocks
[0139] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0141] While preferred embodiments of the application have been described, those skilled in the art will recognize that additional modifications and changes can be made thereto without departing from the scope of the application. Accordingly, the appended claims are intended to cover all such modifications and changes as fall within the scope of the application.
[0142] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are not intended to denote a particular order, quantity, combination, or importance of, or between, the elements so designated. Also, the use of the terms "including", "containing", or "comprising" and variations thereof, is meant to encompass the inclusion of zero or more elements, steps, or components, and is not meant to exclude the addition of other elements, steps, or components, or the performance of further additions, whether optional or required. Without further limitation, an element, step, or component preceded by "comprising" does not exclude the addition of more elements, steps, or components, whether optional or required.
[0143] The above describes in detail the conference state control method for an AI digital person and the conference state control system for an AI digital person provided by the present application, and the principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, the specific implementation manners and application ranges can be changed according to the idea of the present application. In conclusion, the content of the present description should not be understood as a limitation of the present application.
Claims
1. A conference status control method, characterized in that: The method comprises: Obtain meeting environment data; Performing multimodal analysis on the conference environment data to obtain a conference status assessment result; generating a conference control instruction according to the conference status evaluation result; Execute conference process control operations according to the conference control instructions.
2. The method according to claim 1, characterized in that The performing multimodal analysis on the conference environment data to obtain a conference status evaluation result includes: Extracting multiple types of feature data from the conference environment data, the multiple types of feature data comprising at least one of the following: facial expression features recognized based on a deep learning model, speech emotion features extracted based on an end-to-end speech recognition model, and gesture dynamic features recognized based on a dynamic time warping algorithm; The multiple types of feature data are input into a state evaluation model, and the conference state evaluation result is output through the state evaluation model.
3. The method according to claim 2, characterized in that The step of inputting the multiple types of feature data into a state assessment model and outputting the conference state assessment result through the state assessment model includes: The multiple categories of feature data are input into the state assessment model, and the state assessment model is used to identify the emotion type of the facial expression features, and / or to perform semantic segmentation and speech rhythm analysis on the voice emotion features, and / or to perform interaction activity scoring on the gesture dynamic features to obtain the meeting state assessment result.
4. The method according to claim 3, characterized in that Generating a conference control instruction according to the conference status evaluation result includes: Based on a preset multi-dimensional state evaluation model, at least one of the emotion category recognition result, the semantic speech speed analysis result, and the interaction activity score result is mapped into a meeting state level; When the conference status level reaches a preset target level, the conference control instruction is generated, and the conference control instruction is used to adjust the conference rhythm, reschedule the conference agenda, or switch conference roles; The authenticity of the conference control instruction is verified through a multimodal consistency detection model.
5. The method according to claim 4, characterized in that When the conference status level reaches a preset target level, generating the conference control instruction includes: Within a preset control trigger evaluation period, continuously determining whether the conference status level reaches the target level; During the control trigger evaluation period, when the conference status level reaches the target level multiple times in succession and the time interval between each judgment is within a preset time window, the conference control instruction is generated.
6. The method according to claim 1, characterized in that The performing of the conference process control operation according to the conference control instruction includes: Determine whether the current conference process node allows the insertion of the conference process control operation; If the current conference process node allows the insertion of the conference process control operation, the conference rhythm is adjusted, the meeting agenda is rearranged, or the meeting roles are switched through voice prompts or subtitle prompts, and the audio and video time bases are aligned using the optical manifold algorithm and deep learning model; If the current conference process node does not allow the conference process control operation to be inserted, the conference control instruction is cached and executed at the next conference process node that allows insertion.
7. The method according to claim 4, characterized in that After performing the conference process control operation according to the conference control instruction, the method further includes: During subsequent meetings, new meeting environment data will continue to be acquired and multimodal analysis and processing will be performed to obtain subsequent evaluation results; When the subsequent evaluation results fail to reach the target level for multiple consecutive times within a preset control feedback evaluation period, the frequency of generating conference control instructions is reduced, or the triggering conditions for generating conference control instructions are increased; When the subsequent evaluation results continue to reach the target level, maintaining the current frequency of generating conference control instructions and recording the generation triggering conditions of the current conference control instructions; During the subsequent conference, if conference environment data matching the recorded generation trigger condition is obtained, a conference control instruction is directly generated based on the recorded generation trigger condition.
8. A conference status control system, characterized in that: The system comprises: A data acquisition module is used to obtain conference environment data; A data analysis module, configured to perform multimodal analysis on the conference environment data to obtain a conference status assessment result; An instruction generation module, configured to generate a conference control instruction according to the conference status evaluation result; The conference control module is used to perform conference process control operations according to the conference control instructions.
9. An electronic device, characterized in that: include: one or more processors; and One or more machine-readable media having instructions stored thereon, when executed by the one or more processors, enable the electronic device to execute the conference status control method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer program stored therein enables the processor to execute the conference status control method according to any one of claims 1 to 7.
Citation Information
Cited By
Intelligent activity control method and system based on feature matching
CN121052612A
An active intelligent control method and system based on feature matching
CN121052612B