Intelligent agent collaborative barrier-free conference management method and system

By using an intelligent agent-based approach to manage accessible meetings, audio and non-verbal visual information are collected and processed in real time to generate a structured data stream of meeting content. Based on the accessibility needs of participants, multimodal information is output synchronously, solving the problems of single modality and rigid strategies in traditional accessible meeting solutions and achieving efficient accessibility information transmission.

CN121788093APending Publication Date: 2026-04-03广东公信智能会议股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional accessible meeting solutions suffer from problems such as single modality, lack of collaboration, and rigid strategies. They are unable to fully reproduce the interactive scenarios and emotional intentions of participants in the meeting, and cannot dynamically adjust the output strategy according to the personalized accessibility needs of different participants.

Method used

By employing an agent collaboration approach, a voice agent, a visual agent, a semantic parsing agent, and an accessibility output agent work together to collect and process conference audio and non-verbal visual information in real time, generating a structured conference content data stream. Based on the accessibility needs of the participants, multimodal accessibility information, such as speech synthesis broadcasting, real-time subtitle display, and sign language animation, is output synchronously.

Benefits of technology

It achieves deep integration and precise adaptation of multimodal information, improves the completeness, real-time nature and relevance of meeting information transmission, and enhances the inclusiveness and practical effectiveness of barrier-free meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121788093A_ABST
    Figure CN121788093A_ABST
Patent Text Reader

Abstract

The invention provides an agent collaborative barrier-free conference management method and system, and relates to the technical field of conference management, and the method comprises the steps that a voice agent collects a conference audio stream in real time, carries out the recognition and transcription of the conference audio stream, and generates a first text stream; the visual agent collects non-lingual visual information of the participants in real time and converts the non-lingual visual information into a second text stream; the semantic analysis agent receives and fuses the first text stream and the second text stream, extracts conference key information through a context semantic understanding model, and generates a structured conference content data stream; and the barrier-free output agent is used for receiving the conference content data stream and synchronously generating at least two different modes of barrier-free output information from the conference content data stream according to the barrier-free requirements of the participants. The technical problems of missing multi-modal information fusion, insufficient collaboration and output strategy solidification existing in a barrier-free conference scheme in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of meeting management technology, and in particular to a barrier-free meeting management method and system based on intelligent agent collaboration. Background Technology

[0002] With the diversification of meeting scenarios and the personalization of participants' needs, barrier-free meeting management technology has become key to improving meeting inclusivity and efficiency.

[0003] However, traditional accessible meeting solutions have many technical pain points: on the one hand, traditional accessible meeting solutions mostly focus on single-modal output, making it difficult to fully reproduce the interactive scenarios and emotional intentions of participants in the meeting; on the other hand, each functional module is mostly independent and isolated, without building an efficient collaborative mechanism, which not only leads to low efficiency in meeting content processing and fragmented and incomplete information transmission, but also makes it impossible to dynamically adjust the output strategy according to the personalized accessibility needs of different participants, resulting in a lack of targeted output content and insufficient effectiveness in practical application.

[0004] Therefore, there is an urgent need for an accessible meeting management method based on intelligent agent collaboration to address the pain points of existing technologies, such as single modality, lack of collaboration, and rigid strategies. Summary of the Invention

[0005] This invention addresses the technical problems of lack of multimodal information fusion, insufficient collaboration, and fixed output strategies in existing accessible meeting solutions by providing an accessible meeting management method and system based on intelligent agent collaboration.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides an accessible meeting management method for intelligent agent collaboration, comprising a voice intelligent agent, a visual intelligent agent, a semantic parsing intelligent agent, and an accessible output intelligent agent deployed and working collaboratively, the method comprising: The voice agent collects the conference audio stream in real time, and identifies and transcribes the conference audio stream to generate a first text stream; The visual agent collects non-verbal visual information from participants in real time and converts the non-verbal visual information into a second text stream; The semantic parsing agent receives and merges the first text stream and the second text stream, extracts key information of the meeting through the context semantic understanding model, and generates a structured meeting content data stream; The accessibility output agent receives the meeting content data stream and, based on the accessibility needs of the participants, synchronously generates at least two different modalities of accessibility output information from the meeting content data stream. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.

[0007] Secondly, the present invention provides an intelligent agent collaborative barrier-free meeting management system, comprising: The audio acquisition and processing module is used by the voice agent to acquire the conference audio stream in real time, and to identify and transcribe the conference audio stream to generate a first text stream; A visual information acquisition module is used by the visual intelligent agent to acquire non-verbal visual information of the participants in real time and convert the non-verbal visual information into a second text stream. The information fusion module is used by the semantic parsing agent to receive and fuse the first text stream and the second text stream, extract key information of the meeting through the context semantic understanding model, and generate a structured meeting content data stream. An accessibility information output module is used by the accessibility output agent to receive the meeting content data stream and, based on the accessibility needs of the participants, synchronously generate at least two different modalities of accessibility output information from the meeting content data stream. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.

[0008] The beneficial effects of this invention are: Compared to existing technologies, this application firstly, uses a voice intelligence agent to collect the conference audio stream in real time, and then identifies and transcribes it to generate a first text stream. This improves the real-time performance and accuracy of conference audio processing and provides a reliable voice data foundation for the cross-modal information fusion of the semantic parsing intelligence agent. Secondly, a visual intelligence agent collects the non-verbal visual information of the participants in real time and converts it into a second text stream, achieving temporal alignment and correlation with the first text stream. This provides complete and reliable visual semantic data support for the cross-modal information fusion of the semantic parsing intelligence agent. Thirdly, the semantic parsing intelligence agent receives and merges the first and second text streams, extracts key conference information through a contextual semantic understanding model, and generates a structured conference content data stream. This solves the fragmentation and incompleteness problems caused by the independent processing of traditional conference information, and achieves deep integration of cross-modal information. Finally, the accessibility output agent receives the meeting content data stream and, based on the accessibility needs of the participants, synchronously generates at least two different modalities of accessibility output information from the meeting content data stream. This overcomes the limitations of the single modality in traditional solutions, achieves precise adaptation to different types of accessibility needs, and ensures the integrity, real-time nature, and relevance of meeting information transmission, thereby improving the inclusiveness and practical effectiveness of accessibility meetings.

[0009] Through the aforementioned technical solution, this application achieves collaborative work among multiple intelligent agents by deploying a voice intelligent agent, a visual intelligent agent, a semantic parsing intelligent agent, and an accessibility output intelligent agent. By comprehensively collecting meeting audio and non-verbal visual information, a structured meeting content data stream is generated through fusion. Based on the accessibility needs of participants, at least two different modalities of accessibility output information are simultaneously generated. This overcomes the technical pain points of traditional accessibility meeting solutions, such as single modality, lack of collaboration, and rigid strategies. It achieves comprehensive collection, deep cross-modal fusion, and structured integration of meeting audio and non-verbal visual information, enabling precise adaptation and dynamic response to different types of accessibility needs. This improves the completeness, real-time nature, and relevance of meeting information transmission, effectively ensuring fair information transmission and enhancing the inclusiveness, adaptability, and practical effectiveness of accessibility meetings. Attached Figure Description

[0010] Figure 1 A flowchart illustrating an intelligent agent collaborative barrier-free meeting management method provided by the present invention; Figure 2 This is a schematic diagram of the structure of an intelligent agent collaborative barrier-free meeting management system provided by the present invention.

[0011] In the attached diagram, the components represented by each number are as follows: Audio acquisition and processing module 11, visual information acquisition module 12, information fusion module 13, and accessibility information output module 14. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0014] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0015] Example 1, as Figure 1 As shown, this embodiment of the invention provides an accessible meeting management method based on intelligent agent collaboration, including a voice intelligent agent, a visual intelligent agent, a semantic parsing intelligent agent, and an accessible output intelligent agent deployed and working collaboratively, specifically including: This application deploys a voice agent, a visual agent, a semantic parsing agent, and an accessibility output agent. By synchronizing and associating these agents with participant identifiers via timestamps, collaborative work among them is achieved. Timestamp synchronization means that each agent timestamps the collected and processed information in milliseconds, ensuring consistency in the output of information from different agents. Participant identifier association involves assigning a unique identifier, such as an ID code, to each participant. Each agent uses this identifier to associate the collected and processed information with the corresponding participant, avoiding information confusion. Through this deployment and collaboration mechanism, information sharing and collaborative work among multiple agents can be achieved, ensuring comprehensive and accurate processing of meeting content.

[0016] Optionally, the voice intelligent agent refers to an intelligent module used for the acquisition, processing and transcription of conference audio information. Its core components may include an audio acquisition module, an audio preprocessing module, and a real-time speech recognition and transcription module. Its core function is to generate a first text stream with timestamps and participant identifiers.

[0017] Optionally, the visual intelligent agent refers to an intelligent module used for the acquisition, parsing, and semantic conversion of non-verbal visual information of participants. Its core components may include a visual acquisition module, a target detection and key point recognition module, and a visual semantic parsing module. Its core function is to generate a second text stream with timestamps and participant identifiers.

[0018] Optionally, the semantic parsing agent refers to an intelligent module used for cross-modal text stream fusion, contextual semantic understanding, and extraction of key meeting information. Its core components may include a text stream alignment module, a cross-modal fusion module, a contextual semantic understanding model, and a key meeting information extraction module. Its core function is to fuse the first text stream and the second text stream to generate a structured meeting content data stream.

[0019] Optionally, the accessibility output agent refers to an intelligent module used to convert structured meeting content data streams into multimodal accessibility output information according to the personalized needs of participants. Its core components include a participant needs matching module, a dynamic strategy generation module, and a multimodal output module. Its core function is to simultaneously output at least two types of multimodal information, such as speech synthesis broadcasting, real-time subtitle display, and sign language animation.

[0020] S10: The voice agent collects the conference audio stream in real time, and identifies and transcribes the conference audio stream to generate a first text stream.

[0021] Traditional conference audio processing suffers from drawbacks such as incomplete acquisition coverage, severe environmental noise and echo interference, high latency in speech recognition and transcription, and output of only single text content, making it difficult to meet the needs of accessible conferences for real-time performance, accuracy, and multi-agent collaboration.

[0022] To address the aforementioned issues, this application utilizes a voice intelligence agent to collect conference audio streams in real time, and then identifies and transcribes the conference audio streams to generate a first text stream.

[0023] Specifically, step S10 in the method includes: The voice intelligence agent collects multiple audio streams in the conference scene in real time through distributed audio acquisition devices, and performs noise reduction, echo cancellation and sound source localization preprocessing on the multiple audio streams to obtain a clean audio stream; A real-time speech recognition model based on the Transformer architecture is used to recognize and transcribe the clean audio stream sentence by sentence, and the corresponding participant identifiers and speaking timestamps are simultaneously annotated to obtain the speech-transcribed text; The speech-to-text is processed for grammatical error correction, sentence segmentation optimization, and terminology standardization to generate a first text stream with a unified format and complete semantics, and the first text stream is pushed to the semantic parsing agent in real time.

[0024] In this embodiment, the voice agent first acquires multiple audio streams in the conference scene in real time through a distributed audio acquisition device, and then performs noise reduction, echo cancellation, and sound source localization preprocessing on the multiple audio streams to obtain a clean audio stream. Specifically, the distributed audio acquisition device includes multiple microphones arranged in different locations in the conference scene, enabling 360-degree audio acquisition without blind spots; the noise reduction processing can use a spectral subtraction-based noise reduction algorithm to remove environmental noise; the echo cancellation can use an adaptive filtering algorithm to eliminate echo interference in the conference scene; and the sound source localization preprocessing can use a localization algorithm based on time difference of arrival to accurately determine the speaker's location.

[0025] For example, when five participants speak at the same time in a meeting, the distributed audio acquisition device can acquire five audio streams. After preprocessing, five clean audio streams are obtained, each corresponding to the speech of one of the five participants.

[0026] Secondly, a real-time speech recognition model based on the Transformer architecture is used to recognize and transcribe the clean audio stream sentence by sentence, and the corresponding participant identifiers and speaking timestamps are simultaneously annotated to obtain the speech-transcribed text. Here, the real-time speech recognition model refers to a speech recognition model that uses Transformer as its core network structure, which has the advantages of strong parallel computing capabilities and high recognition accuracy; the participant identifier is a unique identifier assigned to each participant, such as participant A, participant B, etc.; and the speaking timestamp is a time stamp of the speaking content in milliseconds.

[0027] For example, a real-time speech recognition model can be constructed through the following technical path: Collect audio data under different acoustic environments, number of speakers, and interference conditions. After preprocessing using the same process as a speech agent, label clean audio samples, transcribed text, participant identifiers, and millisecond-level speech timestamps as the training dataset. Divide the training dataset into training, validation, and test sets in a 7:2:1 ratio. Then, design a lightweight Conformer backbone network, integrating Transformer self-attention and convolutional local feature extraction capabilities. The encoder employs a sparse multi-head attention mechanism to reduce computation, and the decoder uses CTC and attention... The hybrid decoding structure of the force mechanism adds a multi-task joint learning branch to the output layer to synchronously output the transcribed text, participant identification, and timestamp-related information, resulting in a real-time speech recognition model. Subsequently, the real-time speech recognition model is pre-trained using a training set, employing a weighted sum of CTC loss and attention loss combined with association loss. Then, real-time performance is optimized through knowledge distillation, model pruning, and quantization. The hyperparameters are adjusted until the recognition accuracy is ≥95% and the end-to-end latency is ≤200 milliseconds, at which point the model is considered to have converged. Finally, the trained real-time speech recognition model is deployed to the speech agent to achieve real-time, high-precision recognition and transcription of clean audio streams and synchronous output of related information.

[0028] For example, when participant A starts speaking at 10:00:00.000, and the content is: Today we will discuss the upgrade plan for the intelligent mine safety monitoring system, after recognition and transcription, the resulting speech-to-text is: Participant A [10:00:00.000-10:00:05.000]: Today we will discuss the upgrade plan for the intelligent mine safety monitoring system.

[0029] Finally, the speech-to-text is processed for grammatical correction, sentence segmentation optimization, and terminology standardization to generate a first text stream with a unified format and complete semantics. This first text stream is then pushed to the semantic parsing agent in real time. Specifically, grammatical correction can employ grammatical correction algorithms to correct grammatical errors in the speech-to-text; sentence segmentation optimization can use semantic-based segmentation algorithms to make the segmentation of the transcribed text more consistent with language habits; and terminology standardization refers to unifying the technical terms in the speech-to-text into preset standard terms, such as standardizing AI to artificial intelligence.

[0030] For example, if the speech-to-text is: Participant A [10:00:00.000-10:00:05.000]: Today we discuss the upgrade plan for the intelligent mine safety monitoring system, after terminology standardization, it becomes: Participant A [10:00:00.000-10:00:05.000]: Today we discuss the upgrade plan for the artificial intelligence mine safety monitoring system. Finally, the first text stream is generated and pushed to the semantic parsing agent.

[0031] In summary, compared to existing technologies, this application utilizes the aforementioned speech agent to acquire conference audio streams in real time, and then identifies and transcribes these audio streams to generate a first text stream. This distributed acquisition achieves comprehensive coverage of the conference audio, preprocessing eliminates noise, echo, and other interference, and through real-time speech recognition model transcription, grammatical error correction, sentence segmentation optimization, and terminology standardization, a structured first text stream with participant identifiers and timestamps is generated. This not only improves the real-time performance and accuracy of conference audio processing but also provides a reliable speech data foundation for cross-modal information fusion of the semantic parsing agent.

[0032] S20: The visual agent collects non-verbal visual information of the participants in real time and converts the non-verbal visual information into a second text stream.

[0033] Information processing in traditional meeting scenarios often focuses on voice signals, neglecting non-verbal visual information such as participants' facial expressions, body movements, and posture changes. However, these are important carriers of participants' emotions and intentions. Their absence not only fails to fully recreate the meeting interaction scenario but also makes it difficult to meet the demand for multimodal information in barrier-free meetings.

[0034] Meanwhile, with the rapid development of deep learning technologies such as object detection, key point recognition, and visual semantic generation in the field of computer vision, visual agents are being able to collect non-verbal visual information of attendees in real time and convert it into a structured second text stream that can collaborate with the first text stream output by the speech agent, providing complete visual semantic data support for the cross-modal fusion of subsequent semantic parsing agents.

[0035] To address the aforementioned issues, this application utilizes a visual intelligent agent to collect non-verbal visual information from participants in real time and converts this non-verbal visual information into a second text stream.

[0036] Specifically, step S20 in the method includes: The visual intelligent agent collects non-verbal visual information of the participants in real time through a high-definition camera array, wherein the non-verbal visual information includes at least facial expressions, body movements and posture changes. The target detection and key point recognition model is used to detect and track the non-verbal visual information, and extract the facial expression feature points, limb joint points and movement trajectory of the participants to form non-verbal visual recognition results. The non-verbal visual recognition results are converted into semantic text descriptions by a pre-trained visual semantic parsing model, wherein the semantic text descriptions include at least emotion type, action meaning, and posture state. Associate the semantic text description with the corresponding participant identifier and timestamp, generate a second text stream, and send the second text stream to the semantic parsing agent.

[0037] In this embodiment, the visual agent first collects nonverbal visual information of participants in real time through a high-definition camera array. This nonverbal visual information includes at least facial expressions, body movements, and posture changes. Specifically, the high-definition camera array includes multiple high-definition cameras positioned at different locations within the meeting room, enabling omnidirectional imaging of each participant. Facial expressions include basic emotional expressions such as joy, anger, sorrow, and happiness; body movements include gestures and sitting postures; and posture changes include leaning forward and backward.

[0038] For example, when participant B frowns, crosses his arms, and leans forward while participant A is speaking, the high-definition camera array can capture participant B's facial expressions, body movements, and posture changes in real time.

[0039] Secondly, a target detection and keypoint recognition model is used to detect and track targets in non-verbal visual information, extracting facial expression feature points, limb joints, and movement trajectories of attendees to form non-verbal visual recognition results. The target detection and keypoint recognition model can employ the YOLO target detection model and the OpenPose keypoint recognition model. The YOLO target detection model can detect the attendee's body regions, while the OpenPose keypoint recognition model can extract facial expression feature points and limb joints. Facial expression feature points are feature points of important facial areas, and limb joints include key body joints. The number of facial expression feature points and limb joints can be dynamically set according to computing power and accuracy requirements.

[0040] For example, after collecting the facial expressions, body movements, and posture changes of participant B, 68 facial expression feature points and 18 body joint points are extracted through the target detection and key point recognition model, and the movement change trajectory is generated to form a non-verbal visual recognition result.

[0041] It should be noted that both the YOLO object detection model and the OpenPose key point recognition model are open source models, which can be obtained through public technical communities or code repositories. Those skilled in the art can make targeted fine-tuning of the above models based on the meeting scenario requirements of this invention, such as optimizing the extraction accuracy of facial expression feature points and limb joint points of participants and optimizing the speed of real-time tracking of multiple participants, so as to meet the technical requirements of visual intelligent agents for efficient object detection and key point recognition of non-verbal visual information.

[0042] Secondly, a pre-trained visual semantic parsing model is used to convert non-verbal visual recognition results into semantic text descriptions. These semantic text descriptions include at least emotion type, action meaning, and posture. Specifically, the pre-trained visual semantic parsing model refers to a deep learning model trained with a large amount of visual information and corresponding semantic text descriptions, capable of converting visual information into semantic text; emotion types include questioning, agreement, and neutrality; action meanings include liking, disagreeing, and thinking; and postures include sitting upright and leaning forward.

[0043] For example, the visual semantic parsing model can be constructed using the following technical approach: First, collect non-verbal visual information data such as facial expressions, body movements, and posture changes of participants under different meeting room acoustic environments, number of participants, and interactive scenarios. After target detection and key point recognition, extract facial expression feature points, body joints, and movement trajectory changes. Simultaneously, manually annotate the corresponding semantic text descriptions and associate them with participant identifiers and millisecond-level timestamps to form a visual feature dataset and a semantic text description set. Divide the visual feature dataset and the semantic text description set into training, validation, and test sets according to a 7:1.5:1.5 ratio. Second, design a deep learning architecture that integrates visual feature extraction and semantic generation, using visual... The Transformer extracts global semantic features from visual feature data and combines them with pre-trained language models, such as BERT, to construct a cross-modal semantic mapping module, realizing the conversion of visual features to natural language semantics. At the same time, an association branch between participant identifiers and timestamps is added to the model output layer to support multi-task joint learning. Finally, the model is trained using a dataset, and the model parameters are optimized using the cross-entropy loss function. Lightweight techniques such as knowledge distillation and model pruning are used to improve inference speed. Hyperparameters are adjusted on the validation set until the matching degree between the semantic text description and non-linguistic visual information output by the model is ≥90% and the end-to-end latency is ≤150 milliseconds. The model is considered to have converged, and the trained visual semantic parsing model is obtained.

[0044] For example, the visual semantic parsing model converts non-verbal visual recognition results into semantic text descriptions: Participant B exhibits a questioning emotional state and crosses his arms to indicate opposition.

[0045] Finally, a second text stream is generated by associating the semantic text description with the corresponding participant identifier and timestamp, and then sent to the semantic parsing agent. Specifically, the participant identifier is consistent with the participant identifier used in the voice agent, and the timestamp is synchronized with the timestamp used in the voice agent.

[0046] For example, when the semantic text describes participant B as exhibiting a questioning emotional state and crossing their arms to indicate opposition, the corresponding participant identifier is associated with participant B and the timestamp is 10:00:05.000-10:00:10.000, a second text stream is generated and sent to the semantic parsing agent.

[0047] In summary, compared to existing technologies, this application utilizes a visual intelligent agent to collect non-verbal visual information of participants in real time and converts this non-verbal visual information into a second text stream. Thus, by collecting and generating the second text stream in real time, the emotions and intentions of participants during meeting interactions are reconstructed, achieving temporal alignment and correlation with the first text stream. This provides complete and reliable visual semantic data support for cross-modal information fusion of the semantic parsing intelligent agent.

[0048] S30: The semantic parsing agent receives and merges the first text stream and the second text stream, extracts key information of the meeting through the context semantic understanding model, and generates a structured meeting content data stream.

[0049] In traditional meeting information processing, speech-to-text and visual semantic description text are often processed independently, lacking effective cross-modal fusion and contextual analysis mechanisms. This results in fragmented and incomplete extracted meeting information, making it difficult to meet the needs of multi-agent collaboration and barrier-free meetings for structured and accurate meeting content.

[0050] To address the aforementioned issues, this application utilizes a semantic parsing agent to receive and fuse the first text stream and the second text stream, extracts key meeting information through a contextual semantic understanding model, and generates a structured meeting content data stream.

[0051] Specifically, step S30 in the method includes: The semantic parsing agent timestamps the first text stream and the second text stream, and establishes an association mapping between the speech-to-text and semantic text description of the same participant based on the participant identifier, thereby obtaining fused text information; The fused text information is analyzed by a contextual semantic understanding model to extract key meeting information, which includes at least the meeting topics, agenda nodes, key points of speeches, decision results, to-do items, interaction relationships among participants, and emotional tendencies. According to a preset structured data format, the key meeting information, the first text stream, and the second text stream are integrated to generate a meeting content data stream containing hierarchical relationships.

[0052] In this embodiment, the semantic parsing agent first performs timestamp alignment on the first and second text streams, and establishes an association mapping between the speech-transcribed text and semantic text description of the same participant based on the participant identifier, thereby obtaining fused text information. Specifically, timestamp alignment refers to matching information with the same timestamp in the first and second text streams; association mapping refers to associating the speech-transcribed text and semantic text description of the same participant based on the participant identifier.

[0053] For example, when the first text stream contains participant A [10:00:00.000-10:00:05.000]: Today we discuss the upgrade plan for the artificial intelligence mine safety monitoring system, and the second text stream contains participant B [10:00:05.000-10:00:10.000]: exhibiting a questioning emotional state, with crossed arms indicating opposition, the association between the speech-transcribed text and semantic text description of the same participant is established based on the participant identifier, and then after timestamp alignment, the fused text information is obtained.

[0054] Secondly, a contextual semantic understanding model is used to perform contextual association analysis on the fused text information to extract key meeting information. This key information includes at least the meeting topics, agenda nodes, key points of speeches, decision results, to-do items, participant interactions, and emotional tendencies. The contextual semantic understanding model refers to a deep learning model that uses a pre-trained language model as its core network structure, possessing powerful contextual understanding capabilities. The meeting topics refer to the core themes and discussion content of the meeting; agenda nodes refer to different stages of the meeting; key points of speeches refer to the core content of participants' speeches; decision results refer to the consensus reached at the meeting; to-do items refer to tasks determined by the meeting that need to be completed subsequently; participant interactions refer to the dialogue relationships between participants; and emotional tendencies refer to the participants' emotional attitudes towards the meeting content.

[0055] For example, by performing contextual association analysis on the fused text information through a contextual semantic understanding model, the meeting topic was extracted as follows: the agenda node was the proposal to upgrade the AI ​​mine safety monitoring system; the key points of the speech were: participant A proposed the AI ​​mine safety monitoring system upgrade plan; the decision result was that no consensus had been reached; the to-do items were: the technical department would submit a detailed design document next week; the interaction relationship between the participants was a dialogue between participant A and participant B; and the emotional tendency was that participant B was questioning.

[0056] Finally, following a preset structured data format, the key meeting information, the first text stream, and the second text stream are integrated to generate a hierarchical meeting content data stream. The preset structured data format refers to using standardized and extensible data formats such as JSON and XML, integrating the key meeting information, the first text stream, and the second text stream according to a preset hierarchical order. For example, JSON format can be used to construct structured data according to the hierarchical order of key meeting information, the first text stream, and the second text stream.

[0057] Furthermore, the phrase "performing contextual association analysis on the fused text information through a contextual semantic understanding model to extract key meeting information" includes: Entity recognition and dependency parsing are performed on the speech-to-text in the fused text information to generate a text semantic graph. The semantic text descriptions in the fused text information are subjected to emotion classification and action intent recognition to generate visual semantic tags; Based on the text semantic graph and the visual semantic tags, cross-modal semantic fusion and disambiguation are performed through the attention mechanism in the context semantic understanding model to identify and label meeting topics, agenda nodes, key points of speeches, decision results, to-do items, participant interaction relationships and emotional tendencies as key meeting information.

[0058] In this embodiment, entity recognition and dependency parsing are first performed on the speech-to-text transcribed from the fused text information to generate a text semantic graph. Entity recognition refers to identifying entities in the speech-to-text, such as attendees, meeting topics, and program names; dependency parsing refers to analyzing the dependency relationships between words in the speech-to-text, such as subject-verb and verb-object relationships; and the text semantic graph is a graph structure with entities as nodes and dependency relationships as edges.

[0059] For example, when the speech-to-text is: Participant A proposes an upgrade plan for an AI-powered mine safety monitoring system, after entity recognition, the entities are: Participant A, AI-powered mine safety monitoring system upgrade plan. After dependency parsing, the dependency relation is: propose. The generated text semantic graph is the following triple: Participant A - propose - AI-powered mine safety monitoring system upgrade plan.

[0060] Secondly, the semantic text descriptions in the fused text information are subjected to emotion classification and action intent recognition to generate visual semantic labels. Emotion classification refers to classifying the emotions in the semantic text descriptions into types such as questioning, agreeing, and neutral; action intent recognition refers to classifying the actions in the semantic text descriptions into intentions such as liking, disagreeing, and thinking; visual semantic labels are labels that contain both emotion type and action intent.

[0061] For example, when the semantic text is described as: Participant B is in a questioning emotional state and his crossed arms indicate opposition, after emotion classification and action intention recognition, a visual semantic label is generated: the emotion is questioning and the action intention is opposition.

[0062] Finally, based on the text semantic graph and visual semantic labels, cross-modal semantic fusion and disambiguation are performed using the attention mechanism in the context semantic understanding model. This identifies and labels meeting topics, agenda nodes, key points of speeches, decision results, to-do items, participant interactions, and emotional tendencies as key meeting information. The attention mechanism refers to the multi-head attention mechanism in the context semantic understanding model, which automatically focuses on important parts of the fused text information; cross-modal semantic fusion refers to fusing the text semantic graph and visual semantic labels to form a unified semantic representation; and disambiguation refers to eliminating semantic ambiguities that arise during the fusion process.

[0063] For example, a contextual semantic understanding model can be constructed through the following technical path: First, collect the first text stream, second text stream, and fused text information from historical multi-scenario meetings, manually annotate key information such as meeting topics, agenda nodes, and key points of speeches, and associate them with participant identifiers and millisecond-level timestamps to form a plural dataset and an annotated dataset. The plural dataset and the annotated dataset are then divided into training, validation, and test sets in a 7:1.5:1.5 ratio. Second, use a Transformer with cross-modal attention mechanism as the base network, extracting the first text stream through text branches. The semantic features and visual branches are used to extract visual semantic features from the second text stream. A context association layer is introduced to capture long-distance dependencies in the meeting content. A key information extraction branch is set in the output layer to obtain the context semantic understanding model. Subsequently, the context semantic understanding model is trained using a dataset. The weighted sum of cross-modal fusion loss and key information extraction loss is used as the loss function. Lightweight methods such as knowledge distillation, model pruning, and quantization are used to optimize the inference speed. Hyperparameters are adjusted on the validation set until the accuracy of extracting key information from the meeting is ≥92%, which is considered convergence, and the trained context semantic understanding model is obtained.

[0064] For example, by using the multi-head attention mechanism of the contextual semantic understanding model, attention is paid to the upgrade plan of the artificial intelligence mine safety monitoring system in the text semantic graph and the questioning in the visual semantic label. After fusion, the meeting topic is identified as the upgrade plan of the artificial intelligence mine safety monitoring system, and the emotional tendency of participant B is questioning, which is marked as key information of the meeting.

[0065] In summary, compared to existing technologies, this application uses a semantic parsing agent to receive and fuse the first and second text streams, extracts key meeting information through a contextual semantic understanding model, and generates a structured meeting content data stream. This solves the fragmentation and incompleteness problems caused by the independent processing of traditional meeting information, achieves deep integration of cross-modal information, and provides unified and reliable structured data support for the multimodal accurate output of the barrier-free output agent.

[0066] S40: The accessibility output agent receives the meeting content data stream and, based on the accessibility needs of the participants, synchronously generates at least two different modalities of accessibility output information from the meeting content data stream. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.

[0067] As the accessibility needs of participants with different types of disabilities, such as visual impairment, hearing impairment, and cognitive impairment, become increasingly diverse, traditional accessibility meeting output solutions suffer from the technical pain point of fixed output strategies. That is, relying solely on single voice or text output cannot simultaneously cover various needs, and lacks the ability to dynamically adjust the output format according to the personalized needs of participants, making it difficult for some participants to efficiently obtain meeting information.

[0068] To address the aforementioned issues, this application utilizes an accessibility output agent to receive the meeting content data stream and, based on the accessibility needs of the participants, synchronously generate at least two different modalities of accessibility output information from the meeting content data stream. These different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.

[0069] Specifically, step S40 in the method includes: The current meeting stage is determined based on the agenda nodes of key meeting information in the meeting content data stream; The accessibility needs of the participants and the current meeting stage are input into a dynamic strategy generation model built on an ensemble learning algorithm to generate accessibility content generation strategies for different participants. Based on the aforementioned accessibility content generation strategy, the importance of key meeting information in the meeting content data stream is analyzed and ranked. The sorted key meeting information is linked and integrated with the speech-to-text and semantic text descriptions in the meeting content data stream according to timestamps to form an integrated information stream, and content tags are added to the integrated information stream. Based on the integrated information flow with the aforementioned content tags, enhanced meeting content data is generated to meet the accessibility needs of different participants. The accessibility output agent generates at least two different modalities of accessibility output information in time alignment based on the enhanced meeting content data. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation. The construction process of the dynamic strategy generation model includes: Multiple sets of historical meeting scenario data are acquired to form a historical feature dataset, wherein each set of historical meeting scenario data includes the accessibility needs of historical participants and the corresponding historical meeting stage; Obtain the historical content generation strategy that corresponds to each group of data in the historical feature dataset and has been validated for effectiveness, and form a strategy tag set; Based on the ensemble learning algorithm, the historical feature dataset and the policy label set are trained in multiple rounds of iterative training. In each round of iterative training, a training subset is generated by sampling from the historical feature dataset according to the current sample weight distribution, and a base learner is trained based on the training subset; Calculate the policy generation error of the base learner on the training subset, and update the sample weight distribution in the next round of training based on the policy generation error, wherein updating the sample weight distribution in the next round of training is to give higher weights to samples with larger policy generation errors. After completing all iterations of training, the multiple base learners obtained from the training are integrated to generate the dynamic policy generation model.

[0070] In this embodiment, the current meeting stage is first determined based on the agenda nodes of key meeting information in the meeting content data stream. These agenda nodes refer to stages such as the proposal discussion stage, the proposal voting stage, and the summary stage, with each current meeting stage corresponding to a specific agenda node.

[0071] For example, when the agenda node in the key information of the meeting is the solution discussion stage, the current meeting stage is determined to be the solution discussion stage.

[0072] Secondly, the accessibility needs of the participants, along with the current meeting stage, are input into a dynamic strategy generation model built on an ensemble learning algorithm to generate accessibility content generation strategies tailored to different participants. Here, the participants' accessibility needs refer to needs related to visual impairment, hearing impairment, cognitive impairment, etc.; the accessibility content generation strategy refers to the content generation strategy generated based on the different accessibility needs of the participants and the current meeting stage.

[0073] For example, when participant C has visual impairment needs and the current meeting stage is the solution discussion stage, after inputting into the dynamic strategy generation model, the generated accessibility content generation strategy for participant C is: prioritize the generation of speech synthesis broadcast, highlight the key points of the speech and the decision results.

[0074] Secondly, based on the accessibility content generation strategy, the importance of key meeting information in the meeting content data stream is analyzed and ranked. Specifically, importance analysis refers to analyzing the degree of importance of key meeting information to different participants according to the accessibility content generation strategy; ranking refers to sorting the key meeting information from highest to lowest importance.

[0075] For example, the accessibility content generation strategy is as follows: prioritize the generation of speech synthesis broadcasts, highlight key points of speeches and decision results, analyze and rank the importance of key information in the meeting, and obtain the following ranking results: key points of speeches, decision results, to-do items, meeting topics, agenda nodes, participant interaction relationships, and emotional tendencies.

[0076] Furthermore, the sorted key meeting information is linked and integrated with the speech-to-text and semantic text descriptions in the meeting content data stream according to timestamps to form an integrated information stream, and content tags are added to the integrated information stream. Content tags refer to the tags added to different types of information in the integrated information stream, such as [Key Points of Speech], [Decision Results], [To-Do Items], etc.

[0077] For example, the key information of the meeting after sorting is as follows: the key points of the speech are: participant A proposed an upgrade plan for the artificial intelligence mine safety monitoring system; the decision result is that no consensus has been reached yet; the to-do items are: the technical department will submit a detailed design document next week; the meeting topic is the upgrade plan for the artificial intelligence mine safety monitoring system; the agenda node is the discussion stage of the plan; the interaction relationship between participants is the dialogue between participant A and participant B; the emotional tendency is that participant B raises questions. These are then linked and integrated according to timestamps to form an integrated information flow, and content tags such as [key points of the speech], [decision result], and [to-do items] are added.

[0078] Furthermore, based on the integrated information flow with content tags, enhanced meeting content data is generated to adapt to the accessibility needs of different participants. Specifically, enhanced meeting content data refers to data obtained by adding relevant information based on the accessibility needs of different participants, on top of the integrated information flow.

[0079] For example, for participant C with visual impairment needs, speech feature information required for speech synthesis is added to the integrated information flow to generate meeting content enhancement data; for participant D with hearing impairment needs, visual style information required for subtitle display is added to the integrated information flow to generate meeting content enhancement data.

[0080] Finally, the accessibility output agent synchronously generates at least two different modalities of accessibility output information with time alignment based on the enhanced meeting content data. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.

[0081] The construction process of the dynamic strategy generation model includes: First, multiple sets of historical meeting scenario data are acquired to form a historical feature dataset. Each set of historical meeting scenario data includes the accessibility needs of historical participants and the corresponding historical meeting stage. For example, 1000 sets of historical meeting scenario data are acquired, each set including the accessibility needs of historical participants, such as visual impairment needs, hearing impairment needs, etc., and the corresponding historical meeting stage, such as the solution discussion stage, solution voting stage, etc., thus forming a historical feature dataset.

[0082] Secondly, obtain the historical content generation strategies corresponding to each set of data in the historical feature dataset, which have been validated for effectiveness, and form a strategy tag set. For example, obtain the historical content generation strategies corresponding to 1000 sets of historical meeting scenario data. For instance, for a set of data on visually impaired needs and solution discussion stages, the historical content generation strategy is to prioritize generating speech synthesis broadcasts. After validating the historical content generation strategies, a strategy tag set is formed.

[0083] Furthermore, based on ensemble learning algorithms, multiple rounds of iterative training are performed on historical feature datasets and policy label sets. Specifically, by leveraging the characteristic of integrating strong learners with weak learners in ensemble learning algorithms, compared to a single model, ensemble learning can effectively reduce the bias and variance of a single model through the collaborative decision-making of multiple base learners, and better adapt to complex combinations of scenarios with different accessibility requirements and different meeting stages.

[0084] For example, the ensemble learning algorithm can adopt the AdaBoost algorithm, and the number of rounds of multi-round iterative training can be set to 10 rounds. The main logic of multi-round iterative training is to gradually focus on difficult examples, that is, to continuously optimize the model's ability to generate policies for complex and special scenarios through multiple rounds of training, so as to avoid the model overfitting in simple scenarios and undergeneralizing in complex scenarios.

[0085] Furthermore, in each round of iterative training, a training subset is generated by sampling from the historical feature dataset based on the current sample weight distribution, and a base learner is trained based on this training subset. Here, the sample weight distribution refers to the weight of each sample in the historical feature dataset, and the initial sample weight distribution is uniform. Specifically, the sample weight distribution in the initial round is uniform, meaning that each group of samples in the historical feature dataset has an equal probability of being selected. The first training subset is generated by randomly sampling from the historical feature dataset; subsequently, the first base learner is trained based on this first training subset.

[0086] For example, in the initial round, 500 sets are uniformly sampled from 1000 sets of historical feature data in the historical feature dataset as a training subset to train the first decision tree base learner, enabling it to initially learn the basic policy mapping relationship. In subsequent iterations, the sample weight distribution will be dynamically adjusted according to the training effect of the base learner in the previous round, and will no longer remain uniform.

[0087] For example, a decision tree can be chosen as the base learner because it has a simple structure, fast training speed, and can clearly capture the non-linear relationship between features and labels.

[0088] Furthermore, the policy generation error of the base learner on the training subset is calculated, and the sample weight distribution in the next iteration of training is updated based on the policy generation error. This update aims to give higher weights to samples with larger policy generation errors. The policy generation error refers to the error between the content generation policy generated by the base learner and the historical content generation policies in the policy label set. Updating the sample weight distribution means adjusting the weight of each sample based on the policy generation error, giving higher weights to samples with larger policy generation errors so they can be prioritized for learning in subsequent iterations.

[0089] Specifically, the sample weights are dynamically adjusted based on the calculated policy generation error, allowing subsequent training to focus more on samples that the model struggles to predict accurately. The policy generation error refers to the deviation between the content generation policy output by the current base learner and the corresponding historical content generation policies in the policy label set. It can be calculated using a policy matching score system: if the output content generation policy matches the core requirements of the historical content generation policy (e.g., voice broadcasting for visually impaired needs), but there are differences in details (e.g., failure to insert timely periodic summaries), the error is calculated as 0.2; if the output policy does not match the core requirements of the standard policy (e.g., outputting sign language animation for visually impaired needs), the error is calculated as 1.0. The sample weights are then updated based on the error value: for samples with large policy generation errors, their weights are increased in the next iteration, making them easier to sample into the training subset; for simple samples with small errors, their weights are decreased to reduce redundancy in repeated learning.

[0090] For example, for a set of samples in the cognitive impairment demand + free communication stage, if the policy generation error of the base learner in the previous round is 0.8, then its weight is increased from the initial 0.001 to 0.005, and the probability of it being sampled into the training subset in the next round increases by 5 times, ensuring that the subsequent base learner focuses on optimizing such difficult case scenarios.

[0091] Finally, after completing all iterative training, the trained base learners are integrated to generate a dynamic policy generation model. Specifically, after completing a preset number of iterative training rounds, several base learners with varying performance will be obtained. For example, some base learners excel at handling simple accessibility requirements—meeting scenarios—while others are better suited for handling complex and challenging scenarios. The integration process can employ a weighted voting mechanism, assigning weights based on the policy generation error of each base learner. Base learners with smaller policy generation errors receive higher weights, and their decisions have a greater impact on the final output policy.

[0092] For example, after 10 iterations, 10 base learners are obtained. Among them, the policy generation error of the 3 base learners that handle simple scenarios is 0.1, and the weight is assigned to each of them 0.2. Among the 7 base learners that handle complex scenarios, the policy generation error of 2 is 0.3, and the weight is assigned to each of them 0.15. The policy generation error of 5 is 0.5, and the weight is assigned to each of them 0.04. The sum of the weights of all base learners is 1.

[0093] For example, after inputting the accessibility needs of the participants and the current stage of the meeting into the dynamic strategy generation model, the 10 base learners output strategies respectively, and the final accessibility content generation strategy is determined by weighted voting.

[0094] Specifically, generating the speech synthesis broadcast includes: Extract key points of speeches, emotional tendencies, and associated participant identifiers from the enhanced meeting content data; The corresponding synthetic speech features are matched according to the participant identifier, and the key points of the speech are converted into synthetic speech containing identity information using the synthetic speech features. When the enhanced meeting content data contains semantic text descriptions that are time-synchronized with the key points of the speech, a prompt sound is inserted at the corresponding time position of the synthesized speech, and the non-verbal visual information corresponding to the semantic text description is verbally described. When the agenda node in the meeting content data stream switches, a meeting phase summary audio is generated and output based on the to-do items and decision results in the key meeting information.

[0095] In this embodiment, the key points of the speeches, emotional tendencies, and associated participant identifiers are first extracted from the enhanced meeting content data. For example, the key points of the speeches extracted from the enhanced meeting content data are: participant A proposes an upgrade plan for an AI-powered mine safety monitoring system; the emotional tendency is: participant B raises questions; and the associated participant identifiers are participant A and participant B.

[0096] Secondly, corresponding synthesized speech features are matched based on the participant's identifier, and these features are used to convert the key points of the speech into synthesized speech containing the participant's identity information. These synthesized speech features include timbre, speech rate, and intonation, with different synthesized speech features corresponding to different participant identifiers.

[0097] For example, the synthesized speech features of participant A are matched as follows: male timbre, moderate speaking speed, and calm tone. The main points of the speech, "Participant A proposes an upgrade plan for an artificial intelligence mine safety monitoring system," are converted into synthesized speech containing identity information: "Participant A proposes an upgrade plan for an artificial intelligence mine safety monitoring system."

[0098] Furthermore, when the enhanced data of the meeting content includes semantic text descriptions that are time-synchronized with the key points of the speech, a prompt tone is inserted at the corresponding time position of the synthesized speech, and the non-verbal visual information corresponding to the semantic text description is verbally described. The prompt tone is a preset prompt tone, such as a ding-dong sound; the verbal description refers to converting the semantic text description into a natural language description.

[0099] For example, when the enhanced data of the meeting content contains semantic text descriptions that are synchronized with the key points of the speech in time: when participant B shows a questioning emotional state and crosses his arms to indicate opposition, a prompt sound "ding-dong" is inserted at the corresponding time position of the synthesized speech, and a voice-over description is given: participant B is frowning and crossing his arms at this time, showing a questioning emotional state.

[0100] Finally, when the agenda node in the meeting content data stream switches, a meeting summary audio is generated and output based on the to-do items and decision results in the key meeting information. For example, when the agenda node switches from the solution discussion stage to the solution voting stage, based on the to-do item: "The technical department will submit the detailed design document next week," and the decision result: "No consensus has been reached yet," a meeting summary audio is generated: "The solution discussion stage has ended, no consensus has been reached yet, and the technical department will submit the detailed design document next week."

[0101] Specifically, generating the real-time caption display includes: Extract speech-to-text, semantic text description, and associated participant identifiers from the enhanced meeting content data; The transcribed text is segmented according to semantic integrity to form continuous subtitle text for display, and different visual styles are used to distinguish the subtitle text corresponding to different participants. The semantic text descriptions in the enhanced meeting content data that are synchronized with the speech-to-text in time are converted into icons or short auxiliary texts and embedded into the subtitle text at the corresponding time point for display. For content in the subtitle text that belongs to meeting topics, decision results, or to-do items, dynamic visual highlighting is performed based on the content markers.

[0102] In this embodiment, the speech-to-text, semantic text description, and associated participant identifiers are first extracted from the enhanced meeting content data. For example, the speech-to-text extracted from the enhanced meeting content data is: Participant A [10:00:00.000-10:00:05.000]: Today we discuss the upgrade plan for the artificial intelligence mine safety monitoring system. The extracted semantic text description is: Participant B [10:00:05.000-10:00:10.000]: Presenting a questioning emotional state, with crossed arms indicating opposition. The extracted associated participant identifiers are Participant A and Participant B.

[0103] Secondly, the speech-to-text is segmented according to semantic integrity to form continuous subtitle text for display, and different visual styles are used to distinguish the subtitle text corresponding to different participant identifiers. These visual styles can include color, font, font size, etc., with different visual styles corresponding to different participant identifiers.

[0104] For example, the speech-to-text is: Participant A [10:00:00.000-10:00:05.000]: Today we will discuss the upgrade plan for the artificial intelligence mine safety monitoring system. The speech is segmented according to semantic integrity and displayed in blue font, representing Participant A. The speech-to-text of Participant B is displayed in red font, thereby achieving accurate differentiation of the speech information of different participants.

[0105] Next, the semantic text descriptions in the enhanced meeting content data that are synchronized with the speech-to-text in time are converted into icons or brief auxiliary text and embedded into the caption text at the corresponding time point for display. The icons can include question icons, approval icons, and thinking icons, while the auxiliary text can include [Question], [Agree], [Think], etc.

[0106] For example, the semantic text description in the meeting content enhancement data that is synchronized with the speech-to-text in time: Participant B shows a questioning emotional state and crosses his arms to indicate opposition, which is converted into auxiliary text: [Questioning], and embedded in the subtitle text at the corresponding time point, displayed as: Today we discuss the upgrade plan of the artificial intelligence mine safety monitoring system [Questioning].

[0107] Finally, for content in the subtitle text that pertains to meeting topics, decision results, or to-do items, dynamic visual highlighting is applied based on content markers. This dynamic visual highlighting can include techniques such as bolding, color changing, and flashing.

[0108] For example, for content in the subtitle text that belongs to the meeting agenda: "Upgrade plan for artificial intelligence mine safety monitoring system", it can be bolded and colored according to the content marker "[meeting agenda]" to highlight it.

[0109] Specifically, generating the sign language animation includes: Extract key information from the meeting content enhancement data, including key points of speeches, meeting topics, emotional tendencies, and participant interactions. Based on the context of the key points of the speech and the meeting agenda, determine the corresponding sequence of sign language actions; Drive the virtual character to execute the sign language action sequence, and control the virtual character's facial expressions, movement rhythm and match the emotional tendency; When the interaction relationship between the participants indicates that the dialogue is taking place between specific participants, the virtual character's gaze is directed towards the participant associated with the dialogue.

[0110] In this embodiment, the key information of the meeting, including key points of speech, meeting topics, emotional tendencies, and participant interactions, is first extracted from the enhanced meeting content data. For example, the extracted key points of speech are: Participant A proposes an upgrade plan for an AI-powered mine safety monitoring system; the meeting topic is: "Upgrade Plan for an AI-powered Mine Safety Monitoring System"; the emotional tendency is: Participant B raises questions; and the participant interaction is a dialogue between Participant A and Participant B.

[0111] Secondly, determine the corresponding sign language sequence based on the context of the key points of the speech and the meeting topic. The sign language sequence refers to the sequence of actions determined according to the national standard sign language.

[0112] For example, based on the key points of the speech: Participant A proposes an upgrade plan for an artificial intelligence mine safety monitoring system, and the meeting topic: upgrade plan for an artificial intelligence mine safety monitoring system, the corresponding sign language sequence is determined as: Participant A, propose, artificial intelligence, mine, safety, monitoring, system, upgrade, plan.

[0113] Next, the virtual character is driven to perform a sequence of sign language actions, and the facial expressions, movement rhythm, and emotional inclinations of the virtual character are controlled to match. Here, the virtual character refers to a preset 3D virtual character; facial expressions can include basic emotional expressions such as joy, anger, sorrow, and happiness, and movement rhythm can include speed, intensity, etc.

[0114] For example, a virtual character is driven to perform a sequence of sign language actions, and the virtual character's facial expression is controlled to be questioning, and the movement rhythm is slow, matching the emotional tendency of participant B.

[0115] Finally, when participant interaction relationships indicate that the dialogue is taking place between specific participants, the virtual character's gaze is directed towards the participant associated with the dialogue. For example, when participant interaction relationships indicate that the dialogue is taking place between participant A and participant B, the virtual character of participant A is directed towards the virtual character of participant B.

[0116] In summary, compared to existing technologies, this application utilizes an accessible output agent to receive the meeting content data stream and, based on the accessibility needs of participants, synchronously generate at least two different modalities of accessible output information. These different modalities include speech synthesis and playback, real-time subtitle display, and sign language animation. This overcomes the limitations of traditional single-modality solutions, achieves precise adaptation to different types of accessibility needs, ensures the integrity, real-time nature, and relevance of meeting information transmission, and enhances the inclusivity and practical effectiveness of accessible meetings.

[0117] In summary, the embodiments of this application have at least the following technical effects: Compared to existing technologies, this application first acquires the conference audio stream in real time through the speech agent, and then identifies and transcribes the audio stream to generate a first text stream. In this way, comprehensive coverage of the conference audio is achieved through distributed acquisition. Preprocessing eliminates noise, echo, and other interference, and real-time speech recognition model transcription, grammar correction, sentence segmentation optimization, and terminology standardization generate a structured first text stream with participant identifiers and timestamps. This improves the real-time performance and accuracy of conference audio processing and provides a reliable voice data foundation for cross-modal information fusion of the semantic parsing agent.

[0118] Secondly, this application uses a visual intelligent agent to collect non-verbal visual information of participants in real time and converts this non-verbal visual information into a second text stream. In this way, by collecting and generating the second text stream in real time, the emotions and intentions of participants during the meeting interaction are reconstructed, achieving temporal alignment and correlation with the first text stream. This provides complete and reliable visual semantic data support for cross-modal information fusion of the semantic parsing intelligent agent. Furthermore, this application uses a semantic parsing agent to receive and fuse the first and second text streams, extracts key meeting information through a contextual semantic understanding model, and generates a structured meeting content data stream. This solves the fragmentation and incompleteness problems caused by the independent processing of traditional meeting information, achieves deep integration of cross-modal information, and provides unified and reliable structured data support for the multimodal accurate output of the barrier-free output agent.

[0119] Finally, this application utilizes an accessibility output agent to receive the meeting content data stream and, based on the accessibility needs of the participants, synchronously generate at least two different modalities of accessibility output information from the meeting content data stream. These different modalities include speech synthesis and playback, real-time subtitle display, and sign language animation. This overcomes the limitations of traditional single-modality solutions, achieves precise adaptation to different types of accessibility needs, and ensures the integrity, real-time nature, and relevance of meeting information transmission, thereby enhancing the inclusivity and practical effectiveness of accessible meetings.

[0120] Through the aforementioned technical solution, this application achieves collaborative work among multiple intelligent agents by deploying a voice intelligent agent, a visual intelligent agent, a semantic parsing intelligent agent, and an accessibility output intelligent agent. By comprehensively collecting meeting audio and non-verbal visual information, a structured meeting content data stream is generated through fusion. Based on the accessibility needs of participants, at least two different modalities of accessibility output information are simultaneously generated. This overcomes the technical pain points of traditional accessibility meeting solutions, such as single modality, lack of collaboration, and rigid strategies. It achieves comprehensive collection, deep cross-modal fusion, and structured integration of meeting audio and non-verbal visual information, enabling precise adaptation and dynamic response to different types of accessibility needs. This improves the completeness, real-time nature, and relevance of meeting information transmission, effectively ensuring fair information transmission and enhancing the inclusiveness, adaptability, and practical effectiveness of accessibility meetings.

[0121] Example 2, as Figure 2 As shown, based on the same inventive concept as the intelligent agent collaborative barrier-free meeting management method provided in Embodiment 1, this embodiment of the invention also provides an intelligent agent collaborative barrier-free meeting management system, including: The audio acquisition and processing module 11 is used by the voice agent to acquire the conference audio stream in real time, and to identify and transcribe the conference audio stream to generate a first text stream; The visual information acquisition module 12 is used by the visual intelligent agent to acquire non-verbal visual information of the participants in real time and convert the non-verbal visual information into a second text stream; Information fusion module 13 is used by the semantic parsing agent to receive and fuse the first text stream and the second text stream, extract key meeting information through the context semantic understanding model, and generate a structured meeting content data stream; Accessibility information output module 14 is used by the accessibility output agent to receive the meeting content data stream and, according to the accessibility needs of the participants, synchronously generate at least two different modalities of accessibility output information from the meeting content data stream. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.

[0122] The audio acquisition and processing module 11 is specifically used for: The voice intelligence agent collects multiple audio streams in the conference scene in real time through distributed audio acquisition devices, and performs noise reduction, echo cancellation and sound source localization preprocessing on the multiple audio streams to obtain a clean audio stream; A real-time speech recognition model based on the Transformer architecture is used to recognize and transcribe the clean audio stream sentence by sentence, and the corresponding participant identifiers and speaking timestamps are simultaneously annotated to obtain the speech-transcribed text; The speech-to-text is processed for grammatical error correction, sentence segmentation optimization, and terminology standardization to generate a first text stream with a unified format and complete semantics, and the first text stream is pushed to the semantic parsing agent in real time.

[0123] The visual information acquisition module 12 is specifically used for: The visual intelligent agent collects non-verbal visual information of the participants in real time through a high-definition camera array, wherein the non-verbal visual information includes at least facial expressions, body movements and posture changes. The target detection and key point recognition model is used to detect and track the non-verbal visual information, and extract the facial expression feature points, limb joint points and movement trajectory of the participants to form non-verbal visual recognition results. The non-verbal visual recognition results are converted into semantic text descriptions by a pre-trained visual semantic parsing model, wherein the semantic text descriptions include at least emotion type, action meaning, and posture state. Associate the semantic text description with the corresponding participant identifier and timestamp, generate a second text stream, and send the second text stream to the semantic parsing agent.

[0124] The information fusion module 13 is specifically used for: The semantic parsing agent timestamps the first text stream and the second text stream, and establishes an association mapping between the speech-to-text and semantic text description of the same participant based on the participant identifier, thereby obtaining fused text information; The fused text information is analyzed by a contextual semantic understanding model to extract key meeting information, which includes at least the meeting topics, agenda nodes, key points of speeches, decision results, to-do items, interaction relationships among participants, and emotional tendencies. According to a preset structured data format, the key meeting information, the first text stream, and the second text stream are integrated to generate a meeting content data stream containing hierarchical relationships.

[0125] Furthermore, the fused text information is analyzed using a contextual semantic understanding model to extract key meeting information, including: Entity recognition and dependency parsing are performed on the speech-to-text in the fused text information to generate a text semantic graph. The semantic text descriptions in the fused text information are subjected to emotion classification and action intent recognition to generate visual semantic tags; Based on the text semantic graph and the visual semantic tags, cross-modal semantic fusion and disambiguation are performed through the attention mechanism in the context semantic understanding model to identify and label meeting topics, agenda nodes, key points of speeches, decision results, to-do items, participant interaction relationships and emotional tendencies as key meeting information.

[0126] The accessibility information output module 14 is specifically used for: The current meeting stage is determined based on the agenda nodes of key meeting information in the meeting content data stream; The accessibility needs of the participants and the current meeting stage are input into a dynamic strategy generation model built on an ensemble learning algorithm to generate accessibility content generation strategies for different participants. Based on the aforementioned accessibility content generation strategy, the importance of key meeting information in the meeting content data stream is analyzed and ranked. The sorted key meeting information is linked and integrated with the speech-to-text and semantic text descriptions in the meeting content data stream according to timestamps to form an integrated information stream, and content tags are added to the integrated information stream. Based on the integrated information flow with the aforementioned content tags, enhanced meeting content data is generated to meet the accessibility needs of different participants. The accessibility output agent generates at least two different modalities of accessibility output information in time alignment based on the enhanced meeting content data. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation. The construction process of the dynamic strategy generation model includes: Multiple sets of historical meeting scenario data are acquired to form a historical feature dataset, wherein each set of historical meeting scenario data includes the accessibility needs of historical participants and the corresponding historical meeting stage; Obtain the historical content generation strategy that corresponds to each group of data in the historical feature dataset and has been validated for effectiveness, and form a strategy tag set; Based on the ensemble learning algorithm, the historical feature dataset and the policy label set are trained in multiple rounds of iterative training. In each round of iterative training, a training subset is generated by sampling from the historical feature dataset according to the current sample weight distribution, and a base learner is trained based on the training subset; Calculate the policy generation error of the base learner on the training subset, and update the sample weight distribution in the next round of training based on the policy generation error, wherein updating the sample weight distribution in the next round of training is to give higher weights to samples with larger policy generation errors. After completing all iterations of training, the multiple base learners obtained from the training are integrated to generate the dynamic policy generation model.

[0127] Specifically, generating the speech synthesis broadcast includes: Extract key points of speeches, emotional tendencies, and associated participant identifiers from the enhanced meeting content data; The corresponding synthetic speech features are matched according to the participant identifier, and the key points of the speech are converted into synthetic speech containing identity information using the synthetic speech features. When the enhanced meeting content data contains semantic text descriptions that are time-synchronized with the key points of the speech, a prompt sound is inserted at the corresponding time position of the synthesized speech, and the non-verbal visual information corresponding to the semantic text description is verbally described. When the agenda node in the meeting content data stream switches, a meeting phase summary audio is generated and output based on the to-do items and decision results in the key meeting information.

[0128] Specifically, generating the real-time caption display includes: Extract speech-to-text, semantic text description, and associated participant identifiers from the enhanced meeting content data; The transcribed text is segmented according to semantic integrity to form continuous subtitle text for display, and different visual styles are used to distinguish the subtitle text corresponding to different participants. The semantic text descriptions in the enhanced meeting content data that are synchronized with the speech-to-text in time are converted into icons or short auxiliary texts and embedded into the subtitle text at the corresponding time point for display. For content in the subtitle text that belongs to meeting topics, decision results, or to-do items, dynamic visual highlighting is performed based on the content markers.

[0129] Specifically, generating the sign language animation includes: Extract key information from the meeting content enhancement data, including key points of speeches, meeting topics, emotional tendencies, and participant interactions. Based on the context of the key points of the speech and the meeting agenda, determine the corresponding sequence of sign language actions; Drive the virtual character to execute the sign language action sequence, and control the virtual character's facial expressions, movement rhythm and match the emotional tendency; When the interaction relationship between the participants indicates that the dialogue is taking place between specific participants, the virtual character's gaze is directed towards the participant associated with the dialogue.

[0130] In summary, the embodiments of this application have at least the following technical effects: Compared to existing technologies, this application firstly utilizes an audio acquisition and processing module. The voice agent acquires the conference audio stream in real time, identifies and transcribes it to generate a first text stream, improving the real-time performance and accuracy of conference audio processing and providing a reliable voice data foundation for the cross-modal information fusion of the semantic parsing agent. Secondly, through a visual information acquisition module, a visual agent acquires non-verbal visual information from participants in real time and converts it into a second text stream, achieving temporal alignment and correlation with the first text stream. This provides complete and reliable visual semantic data support for the cross-modal information fusion of the semantic parsing agent. Thirdly, through an information fusion module, the semantic parsing agent receives and fuses the first and second text streams, extracts key conference information through a contextual semantic understanding model, and generates a structured conference content data stream. This solves the fragmentation and incompleteness problems caused by the independent processing of traditional conference information, achieving deep integration of cross-modal information. Finally, through the accessibility information output module, the accessibility output agent receives the meeting content data stream and, based on the participants' accessibility needs, simultaneously generates at least two different modalities of accessibility output information. This overcomes the limitations of traditional single-modality solutions, achieving precise adaptation to different types of accessibility needs while ensuring the integrity, real-time nature, and relevance of meeting information delivery, thus enhancing the inclusivity and practical effectiveness of accessible meetings. In this way, it addresses the technical pain points of traditional accessible meeting solutions, such as single modality, lack of collaboration, and rigid strategies, improving the integrity, real-time nature, and relevance of meeting information delivery.

[0131] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0132] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0133] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0134] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0135] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0136] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0137] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. A barrier-free meeting management method for intelligent agent collaboration, characterized in that, The method includes deploying and cooperating voice agents, visual agents, semantic parsing agents, and accessibility output agents, and comprises: The voice agent collects the conference audio stream in real time, and identifies and transcribes the conference audio stream to generate a first text stream; The visual agent collects non-verbal visual information from participants in real time and converts the non-verbal visual information into a second text stream; The semantic parsing agent receives and merges the first text stream and the second text stream, extracts key information of the meeting through the context semantic understanding model, and generates a structured meeting content data stream; The accessibility output agent receives the meeting content data stream and, based on the accessibility needs of the participants, synchronously generates at least two different modalities of accessibility output information from the meeting content data stream. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.

2. The barrier-free meeting management method for intelligent agent collaboration according to claim 1, characterized in that, The voice agent acquires the conference audio stream in real time, and recognizes and transcribes the conference audio stream to generate a first text stream, including: The voice intelligence agent collects multiple audio streams in the conference scene in real time through distributed audio acquisition devices, and performs noise reduction, echo cancellation and sound source localization preprocessing on the multiple audio streams to obtain a clean audio stream; A real-time speech recognition model based on the Transformer architecture is used to recognize and transcribe the clean audio stream sentence by sentence, and the corresponding participant identifiers and speaking timestamps are simultaneously annotated to obtain the speech-transcribed text; The speech-to-text is processed for grammatical error correction, sentence segmentation optimization, and terminology standardization to generate a first text stream with a unified format and complete semantics, and the first text stream is pushed to the semantic parsing agent in real time.

3. The barrier-free meeting management method for intelligent agent collaboration according to claim 1, characterized in that, The visual agent collects nonverbal visual information from participants in real time and converts this nonverbal visual information into a second text stream, including: The visual intelligent agent collects non-verbal visual information of the participants in real time through a high-definition camera array, wherein the non-verbal visual information includes at least facial expressions, body movements and posture changes. The target detection and key point recognition model is used to detect and track the non-verbal visual information, and extract the facial expression feature points, limb joint points and movement trajectory of the participants to form non-verbal visual recognition results. The non-verbal visual recognition results are converted into semantic text descriptions by a pre-trained visual semantic parsing model, wherein the semantic text descriptions include at least emotion type, action meaning, and posture state. Associate the semantic text description with the corresponding participant identifier and timestamp, generate a second text stream, and send the second text stream to the semantic parsing agent.

4. The barrier-free meeting management method for intelligent agent collaboration according to claim 1, characterized in that, The semantic parsing agent receives and merges the first text stream and the second text stream, extracts key meeting information through a contextual semantic understanding model, and generates a structured meeting content data stream, including: The semantic parsing agent timestamps the first text stream and the second text stream, and establishes an association mapping between the speech-to-text and semantic text description of the same participant based on the participant identifier, thereby obtaining fused text information; The fused text information is analyzed by a contextual semantic understanding model to extract key meeting information, which includes at least the meeting topics, agenda nodes, key points of speeches, decision results, to-do items, interaction relationships among participants, and emotional tendencies. According to a preset structured data format, the key meeting information, the first text stream, and the second text stream are integrated to generate a meeting content data stream containing hierarchical relationships.

5. The barrier-free meeting management method for intelligent agent collaboration according to claim 4, characterized in that, By performing contextual association analysis on the fused text information using a contextual semantic understanding model, key meeting information is extracted, including: Entity recognition and dependency parsing are performed on the speech-to-text in the fused text information to generate a text semantic graph. The semantic text descriptions in the fused text information are subjected to emotion classification and action intent recognition to generate visual semantic tags; Based on the text semantic graph and the visual semantic tags, cross-modal semantic fusion and disambiguation are performed through the attention mechanism in the context semantic understanding model to identify and label meeting topics, agenda nodes, key points of speeches, decision results, to-do items, participant interaction relationships and emotional tendencies as key meeting information.

6. The barrier-free meeting management method for intelligent agent collaboration according to claim 1, characterized in that, The accessibility output agent receives the meeting content data stream and, based on the accessibility needs of the participants, synchronously generates at least two different modalities of accessibility output information from the meeting content data stream, including: The current meeting stage is determined based on the agenda nodes of key meeting information in the meeting content data stream; The accessibility needs of the participants and the current meeting stage are input into a dynamic strategy generation model built on an ensemble learning algorithm to generate accessibility content generation strategies for different participants. Based on the aforementioned accessibility content generation strategy, the importance of key meeting information in the meeting content data stream is analyzed and ranked. The sorted key meeting information is linked and integrated with the speech-to-text and semantic text descriptions in the meeting content data stream according to timestamps to form an integrated information stream, and content tags are added to the integrated information stream. Based on the integrated information flow with the aforementioned content tags, enhanced meeting content data is generated to meet the accessibility needs of different participants. The accessibility output agent generates at least two different modalities of accessibility output information in time alignment based on the enhanced meeting content data. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation. The construction process of the dynamic strategy generation model includes: Multiple sets of historical meeting scenario data are acquired to form a historical feature dataset, wherein each set of historical meeting scenario data includes the accessibility needs of historical participants and the corresponding historical meeting stage; Obtain the historical content generation strategy that corresponds to each group of data in the historical feature dataset and has been validated for effectiveness, and form a strategy tag set; Based on the ensemble learning algorithm, the historical feature dataset and the policy label set are trained in multiple rounds of iterative training. In each round of iterative training, a training subset is generated by sampling from the historical feature dataset according to the current sample weight distribution, and a base learner is trained based on the training subset; Calculate the policy generation error of the base learner on the training subset, and update the sample weight distribution in the next round of training based on the policy generation error, wherein updating the sample weight distribution in the next round of training is to give higher weights to samples with larger policy generation errors. After completing all iterations of training, the multiple base learners obtained from the training are integrated to generate the dynamic policy generation model.

7. The barrier-free meeting management method for intelligent agent collaboration according to claim 6, characterized in that, Generating the speech synthesis broadcast includes: Extract key points of speeches, emotional tendencies, and associated participant identifiers from the enhanced meeting content data; The corresponding synthetic speech features are matched according to the participant identifier, and the key points of the speech are converted into synthetic speech containing identity information using the synthetic speech features. When the enhanced meeting content data contains semantic text descriptions that are time-synchronized with the key points of the speech, a prompt sound is inserted at the corresponding time position of the synthesized speech, and the non-verbal visual information corresponding to the semantic text description is verbally described. When the agenda node in the meeting content data stream switches, a meeting phase summary audio is generated and output based on the to-do items and decision results in the key meeting information.

8. The barrier-free meeting management method for intelligent agent collaboration according to claim 6, characterized in that, Generating the real-time caption display includes: Extract speech-to-text, semantic text description, and associated participant identifiers from the enhanced meeting content data; The transcribed text is segmented according to semantic integrity to form continuous subtitle text for display, and different visual styles are used to distinguish the subtitle text corresponding to different participants. The semantic text descriptions in the enhanced meeting content data that are synchronized with the speech-to-text in time are converted into icons or short auxiliary texts and embedded into the subtitle text at the corresponding time point for display. For content in the subtitle text that belongs to meeting topics, decision results, or to-do items, dynamic visual highlighting is performed based on the content markers.

9. The barrier-free meeting management method for intelligent agent collaboration according to claim 6, characterized in that, Generating the sign language animation includes: Extract key information from the meeting content enhancement data, including key points of speeches, meeting topics, emotional tendencies, and participant interactions. Based on the context of the key points of the speech and the meeting agenda, determine the corresponding sequence of sign language actions; Drive the virtual character to execute the sign language action sequence, and control the virtual character's facial expressions, movement rhythm and match the emotional tendency; When the interaction relationship between the participants indicates that the dialogue is taking place between specific participants, the virtual character's gaze is directed towards the participant associated with the dialogue.

10. A barrier-free meeting management system for intelligent agent collaboration, characterized in that, A method for managing barrier-free meetings that enables intelligent agent collaboration as described in any one of claims 1-9, comprising: The audio acquisition and processing module is used by the voice agent to acquire the conference audio stream in real time, and to identify and transcribe the conference audio stream to generate a first text stream; A visual information acquisition module is used by the visual intelligent agent to acquire non-verbal visual information of the participants in real time and convert the non-verbal visual information into a second text stream. The information fusion module is used by the semantic parsing agent to receive and fuse the first text stream and the second text stream, extract key information of the meeting through the context semantic understanding model, and generate a structured meeting content data stream. An accessibility information output module is used by the accessibility output agent to receive the meeting content data stream and, based on the accessibility needs of the participants, synchronously generate at least two different modalities of accessibility output information from the meeting content data stream. The different modalities of accessibility output information include speech synthesis broadcasting, real-time subtitle display, and sign language animation.