Intelligent conference assistant system based on human-computer interaction and AI digital human

CN122472704BActive Publication Date: 2026-09-29TIANJIN WEASTED TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610837782.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-29
Estimated Expiration
2046-06-11

AI Technical Summary

Technical Problem

传统会议助手主要依赖语音转写、关键词提取或会后纪要整理方式,对会议过程中的议题推进、参会人发言归属、任务事项形成以及会议事项确认关系缺乏连续建模能力,容易出现会议内容记录完整但事项归属不清、纪要生成滞后、任务跟踪断裂等问题

Benefits of technology

本发明通过交互采集、话轮绑定和长序列构建,将参会人交互输入、数字人交互指令、议题锚定结果和交互标注结果统一形成会议长序列输入对象,使语音、文字、屏幕交互和数字人交互过程能够在同一会议语义链条中被处理,减少普通会议助手只记录文本而忽略交互来源、话轮边界和议题归属的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472704B_ABST
    Figure CN122472704B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent conference assistant system based on human-computer interaction and AI digital people, and relates to the technical field of intelligent conferences, and comprises the following: an interaction collecting module, which obtains conference original interaction data; a speech turn binding module, which obtains an identity binding speech turn sequence; a long sequence constructing module, which generates a conference long sequence input object and conference interaction features; a semantic understanding module, which inputs an improved Longformer model to obtain conference semantic understanding results; an interaction decision module, which inputs an improved LinUCB algorithm to obtain a digital person interaction strategy; a digital person response module, which generates a digital person conference interaction response; a matter confirmation module, a minutes tracking module and a feedback optimization module, which generate conference matter records and optimization results. The application introduces topic global attention in the Longformer model and introduces interaction constraint correction in the LinUCB algorithm, so that conference understanding and digital person decision-making are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent meeting technology, and in particular to an intelligent meeting assistant system based on human-computer interaction and AI digital human. Background Technology

[0002] With the increasing prevalence of online meetings, hybrid work environments, and enterprise digital collaboration scenarios, interactive content during meetings, such as voice communication, text discussion, screen sharing, and task confirmation, is showing continuous growth and multi-source parallelism. Traditional meeting assistants mainly rely on speech transcription, keyword extraction, or post-meeting minutes compilation. They lack the ability to continuously model the progression of meeting topics, the attribution of participants' speeches, the formation of tasks, and the confirmation of meeting items. This can easily lead to problems such as complete meeting content records but unclear attribution of tasks, delayed minutes generation, and broken task tracking.

[0003] While existing intelligent meeting systems can generate meeting summaries through natural language processing models or provide broadcast, reminder, and Q&A services through digital human avatars, most solutions still treat meeting text as ordinary long text, lacking joint processing of the relationship between speaking turns, topic anchors, digital human interaction commands, and participant confirmation feedback.

[0004] Meanwhile, the intervention of AI digital humans usually relies on fixed rules or manual triggering, making it difficult to dynamically select interaction strategies such as prompts, follow-up questions, restatements, or confirmations based on the meeting stage, gaps in the agenda, and turn-taking boundaries. This can lead to problems such as inappropriate intervention timing, responses that are out of touch with the meeting agenda, and difficulty in reverse optimization of post-meeting feedback.

[0005] Therefore, how to provide an intelligent meeting assistant system based on human-computer interaction and AI digital human is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an intelligent meeting assistant system based on human-computer interaction and AI digital human. This invention acquires participant interaction inputs and digital human interaction commands during the meeting through an interaction acquisition module, forming raw meeting interaction data; and establishes identity-bound turn sequences, meeting long sequence input objects, and meeting interaction features through a turn binding module and a long sequence construction module, enabling meeting voice, text, screen interaction, and digital human interaction commands to enter a unified processing chain according to the relationship between turn and topic.

[0007] The semantic understanding module inputs the long sequence of conference input objects into the improved Longformer model, and combines a cross-turn topic preservation mechanism in the long sequence embedding layer, turn attention layer, topic global attention layer, and item output layer to generate the conference semantic understanding results. The interaction decision module inputs the conference semantic understanding results and conference interaction features into the improved LinUCB algorithm, and obtains the digital human interaction strategy through interaction constraint correction. The digital human response module, item confirmation module, minutes tracking module, and feedback optimization module generate the digital human conference interaction response, candidate item confirmation results, conference item records, and conference assistant optimization results in sequence.

[0008] This invention can improve the accuracy of semantic understanding in long meetings, enabling AI digital humans to participate in reminders, follow-up questions, reiterations, and confirmations based on meeting stages and interaction states, reducing omissions in meeting matters and breaks in task tracking, and improving the continuity and executability of intelligent meeting assistance processes.

[0009] An intelligent meeting assistant system based on human-computer interaction and AI digital human according to an embodiment of the present invention includes: The interactive data acquisition module is used to collect the interactive inputs of participants and the interactive commands of the digital human during the meeting, and to obtain the raw interactive data of the meeting. The turn-binding module is used to bind the source of participants' interactions and speaking time periods in the original conference interaction data to obtain an identity-bound turn sequence. The long sequence construction module is used to anchor topics and annotate interactions in identity-bound turn sequences to obtain the long sequence input objects and meeting interaction features of the meeting; The semantic understanding module is used to input the long sequence of conference input objects into the improved Longformer model to obtain the conference semantic understanding results. The improved Longformer model includes a long sequence embedding layer, a turn attention layer, an issue global attention layer, and an item output layer. The issue global attention layer embeds a cross-turn issue retention mechanism. The interactive decision module is used to input the semantic understanding results of the meeting and the interactive features of the meeting into the improved LinUCB algorithm to obtain the digital human interaction strategy; The digital human response module is used to generate digital human meeting interaction responses based on the digital human interaction strategy; The item confirmation module is used to collect participants' interaction confirmations based on the digital human conference interaction responses and obtain the candidate item confirmation results; The minutes tracking module is used to generate meeting minutes based on the confirmation results of candidate items; The feedback optimization module is used to collect feedback from participants on the meeting minutes and update the digital human interaction strategy based on the feedback to obtain the meeting assistant optimization results.

[0010] Optionally, the interactive acquisition module specifically comprises: In response to the meeting start command, a unified meeting clock is established, and an interaction source identifier is generated based on the meeting access record; An environmental noise baseline is extracted from the conference audio input channel. A voice endpoint threshold is generated based on the environmental noise baseline. The voice stream of the participants is subjected to frame-by-frame energy detection. Voice frames that continuously reach the voice endpoint threshold are identified as valid voice frames. A voice time identifier is generated based on the start and end frame times of the valid voice frames. The system reads the trigger times of text input by participants, screen interaction content, and digital human interaction commands, generates event time identifiers based on the trigger times, and generates interaction category identifiers based on the data source. The participant's voice segments, text content, screen interaction content, and digital human interaction commands corresponding to the valid voice frames are bound to the interaction source identifier, interaction category identifier, and time identifier, respectively, and sorted in ascending order according to the time identifier to obtain the original meeting interaction data.

[0011] Optionally, the talk-turn binding module specifically comprises: Read the original meeting interaction data that has been sorted in ascending order by time identifier, and divide the original meeting interaction data into speech interaction data and event interaction data according to the interaction category identifier; Perform speech transcription on the speech segments of participants in the speech-type interactive data to obtain the speech text of the participants, and bind the speech text of the participants with the corresponding interaction source identifier and speech time identifier; The text content in the speech-type interaction data is used as the text speech text, and the text speech text is bound to the corresponding interaction source identifier and event time identifier; Using the participants' spoken text and text spoken text as the text to be bound, the time interval between adjacent texts to be bound is calculated sequentially according to the time identifier; when the interaction source identifier of adjacent texts to be bound is the same and the time interval is less than or equal to the turn interval threshold, the adjacent texts to be bound are grouped into the same turn. When the interaction source identifier changes or the time interval exceeds the turn interval threshold, the current speaking turn ends and a new speaking turn is established. The speaking turn time period is generated based on the time identifier of the first and last unbound speaking text in each speaking turn. Event-type interactive data is bound to speech turns with overlapping time intervals according to the event time identifier. If the event time identifier does not fall into any speech turn time interval, the event-type interactive data is bound to the speech turn with the smallest time interval, thus obtaining the identity-bound speech turn sequence.

[0012] Optionally, the long sequence construction module specifically comprises: Read the identity-bound turn sequence and the meeting agenda, and extract the turn text, interaction source identifier, turn time period, and event-type interaction data for each speaking turn from the identity-bound turn sequence; The topic text in the meeting agenda is converted into topic anchors. Semantic encoding is performed on the dialogue text. The semantic encoding results of the dialogue text are matched with the topic anchors for similarity. The topic anchors that meet the anchoring conditions are determined as the topic anchoring results of the corresponding speaking dialogue. Based on the data source and trigger time of the event-type interaction data, screen interaction events and digital human command events that fall within the speaking turn time period are bound to the corresponding speaking turns, and event-type interaction data that do not fall within the speaking turn time period are bound to the speaking turn with the closest time distance, thus obtaining the interaction annotation results; Based on the order of the turn-by-turn time periods, the turn-by-turn text, interaction source identifier, topic anchoring result, and interaction annotation result are written into the corresponding long sequence position to generate a long sequence input object for the meeting; Based on the digital human command triggering status, participant confirmation status, and speaking turn time boundaries in the interaction annotation results, conference interaction features are generated.

[0013] Optionally, the semantic understanding module specifically comprises: The long sequence embedding layer reads the turn-around text, interaction source identifier, topic anchoring result, and interaction annotation result from the conference long sequence input object, calls the trained text encoding parameters, source encoding parameters, topic encoding parameters, and interaction state encoding parameters, and maps the turn-around text, interaction source identifier, topic anchoring result, and interaction annotation result into a conference turn-around sequence representation of a unified dimension; The turn attention layer is based on the Longformer local sparse self-attention structure. It establishes a turn attention window with the speaking turn as the encoding unit, and writes adjacent speaking turns and speaking turns on the same topic into the turn attention mask. The turn context representation is obtained through multi-head sparse self-attention encoding. The global attention layer sets the topic anchor points in the meeting agenda and the speech turns with digital human interaction annotations as global attention nodes, so that the topic anchor points aggregate the semantics of speech turns under the same topic, and the speech turns with digital human interaction annotations aggregate the semantics of interaction confirmation, thus obtaining a global topic interaction representation. The cross-turn topic retention mechanism generates a topic consistency attention bias based on the topic anchoring result and a digital human interaction attention bias based on the interaction annotation result. The attention weight of the topic global attention layer is corrected by the topic consistency attention bias and the digital human interaction attention bias to obtain the topic-enhanced conference representation. The output layer includes a meeting stage classification header, a minutes item boundary labeling header, and a tracking item boundary labeling header. The meeting stage classification header enhances the meeting representation based on the agenda and outputs the meeting status identification result. The minutes item boundary labeling header enhances the meeting representation based on the agenda and outputs candidate minutes items. The tracking item boundary labeling header enhances the meeting representation based on the agenda and outputs candidate tracking items. The meeting status identification result, candidate minutes items, and candidate tracking items are combined into the meeting semantic understanding result.

[0014] Optionally, the interactive decision-making module specifically comprises: Read the semantic understanding results and interaction features of the meeting, and convert the meeting stage information and candidate item information in the semantic understanding results and the digital human trigger information and turn-taking boundary information in the interaction features into a digital human decision context vector. The following are candidate digital human actions: remaining silent, providing interface prompts, asking follow-up questions via voice, summarizing the phases, reiterating the tasks, and providing confirmation guidance. For each candidate digital human action, the improved LinUCB algorithm is invoked to calculate the action gain estimate and action confidence compensation value based on the digital human decision context vector, and the action gain estimate and action confidence compensation value are combined into an action confidence score; A set of action constraint coefficients is generated based on the semantic understanding results and interaction features of the meeting. The set of action constraint coefficients includes turn-by-turn boundary constraint coefficients, item confirmation constraint coefficients, and active trigger constraint coefficients. When the turn boundary information indicates that the current speech has not ended, the turn boundary constraint coefficient is configured as the constraint coefficient for voice follow-up questions, phase summary, and confirmation guidance; when the candidate item information indicates that there are unconfirmed items, the item confirmation constraint coefficient is configured as the constraint coefficient for item restatement and confirmation guidance. When the digital human trigger message indicates that the participant actively issues a digital human interaction command, the active trigger constraint coefficient is configured as the constraint coefficient for interface prompts, voice follow-up questions, phase summaries, task reiterations, and confirmation guidance; a comprehensive action constraint coefficient is generated based on the constraint coefficient of each candidate digital human action hit, and the action confidence score is multiplied by the comprehensive action constraint coefficient to obtain the corrected action confidence score; Candidate digital human actions are ranked according to the revised action confidence score. The highest-ranked candidate digital human action is taken as the digital human action type. Based on the digital human action type, strategy-related objects are selected from the current speaking turn, candidate minutes, candidate tracking items, and current topic anchoring results. The highest-ranked candidate digital human action, digital human action type, and strategy-related objects are combined to form the digital human interaction strategy.

[0015] Optionally, the digital human response module specifically comprises: Read the digital human interaction strategy and parse the digital human action type and strategy-related objects; When the digital human's action type is to remain silent, generate a digital human meeting interaction response that does not contain content visible to the participants. When the digital human's action type is interface prompt, voice questioning, phase summary, item reiteration or confirmation guidance, the related topic content, related speech content, related minutes or related tracking items are retrieved from the meeting semantic understanding results according to the strategy associated object. Based on the type of digital human action and the retrieved related content, digital human response text is generated. Specifically, interface prompts generate topic reminder text, voice follow-up questions generate item completion text, stage summaries generate topic summary text, item repetitions generate item repetition text, and confirmation guidance generates confirmation guidance text. Determine the response carrier based on the type of digital human action, write the digital human response text corresponding to the interface prompt action into the meeting interface, and convert the digital human response text corresponding to voice follow-up questions, phase summaries, item repetitions or confirmation guidance actions into digital human voice and synchronized subtitles. When the digital human's action type is a summary of an event or a confirmation guide, a confirmation entry is generated based on the associated minutes or tracked events. The digital human response text, response carrier, and strategy-related objects are combined, and a confirmation entry is written when one exists to generate a digital human conference interactive response.

[0016] Optionally, the item confirmation module specifically comprises: Read the confirmation entry and strategy-related objects in the digital human conference interaction response, retrieve related minutes or related tracking items from the conference semantic understanding results based on the strategy-related objects, generate items to be confirmed carrying topic anchoring results and turn time periods, determine the item source type based on related minutes or tracking items, and bind the items to be confirmed to the confirmation entry. The system displays the items to be confirmed to participants through the confirmation entry point and collects the confirmation, modification, or rejection instructions entered by the participants. When a confirmation command is received, the confirmation status of the item to be confirmed will be updated to confirmed. When a modification instruction is received, the modification content entered by the participant is read, the corresponding content in the item to be confirmed is replaced according to the modification content, and the confirmation status of the modified item to be confirmed is updated to confirmed. When a rejection instruction is received, the confirmation status of the item to be confirmed will be updated to "not approved". The process involves binding the items to be confirmed, the type of item source, the type of instruction entered by the participant, the confirmation status, the topic anchoring result, and the turn-around time period to generate candidate item confirmation results.

[0017] Optionally, the minutes tracking module specifically comprises: Read the pending items, item source type, participant input instruction type, and confirmation status from the candidate item confirmation results. Items with a confirmation status of "confirmed" are considered valid items, while pending items with a confirmation status of "not approved" are excluded from the meeting item record. When the instruction type is a modification instruction, the updated pending item in the candidate item confirmation result is used as the content to be written for the valid item; when the instruction type is a confirmation instruction, the original pending item is used as the content to be written for the valid item. Read the topic anchoring results and turn-by-turn time periods carried by valid matters, and classify valid matters into minutes matters and follow-up matters according to the source type of matters; Using the topic anchoring result as the basis for recording partitions, minutes belonging to the same topic anchoring result are written into the corresponding topic minutes area in ascending order of turn-by-turn time; when there are minutes with the same content in the same topic minutes area, the minutes with the later turn-by-turn time are retained. Based on the topic anchoring result, the tracking items belonging to the same topic anchoring result are written into the corresponding topic tracking area in ascending order according to the turn-by-turn time period; the content, responsible party, completion deadline and confirmation status of the tracking items are aligned, and the tracking items with missing fields are written into the pending supplementation status; The minutes and tracking areas for each topic are merged according to the order of the meeting agenda to generate a record of meeting items.

[0018] Optionally, the feedback optimization module specifically comprises: Read the meeting event record, the digital human interaction strategy used when generating the meeting event record, and the digital human decision context vector saved when generating the digital human interaction strategy. Read the strategy-related objects and candidate digital human actions from the digital human interaction strategy. Establish feedback association records based on the event content, strategy-related objects, and candidate digital human actions in the meeting event record. The meeting agenda display interface collects feedback instructions from participants regarding the agenda content and determines the feedback type based on the feedback instructions. The feedback types include acceptance feedback, modification feedback, and rejection feedback. Based on the feedback type, the action reward mapping is performed as follows: Positive feedback is rewarded with a positive reward value, and candidate digital human actions associated with the task content are marked as valid actions; Correction feedback is rewarded with a correction reward value, the modified task content is compared with the original task content, the correction reward value is determined based on the comparison results, and the modified task content is written back to the meeting minutes; Rejection feedback is rewarded with a negative reward value, and candidate digital human actions associated with the task content are marked as invalid actions. Write the digital human decision context vector, candidate digital human actions, and corresponding reward values ​​into the action feedback samples of the improved LinUCB algorithm; Update the parameter matrix and reward vector of the corresponding candidate digital human action based on the action feedback sample, recalculate the action parameters of the candidate digital human action, and obtain the updated digital human interaction strategy. Bind the updated digital human interaction strategy to the written-back meeting event records to generate meeting assistant optimization results.

[0019] The beneficial effects of this invention are: This invention unifies participant interaction input, digital human interaction commands, topic anchoring results, and interaction annotation results into a long sequence input object for the meeting through interactive acquisition, turn binding, and long sequence construction. This enables voice, text, screen interaction, and digital human interaction processes to be processed in the same semantic chain of the meeting, reducing the problem that ordinary meeting assistants only record text and ignore the source of interaction, turn boundaries, and topic attribution.

[0020] The improved Longformer model introduces a global attention layer for topics and a cross-turn topic preservation mechanism on the basis of the original local sparse self-attention for long texts. This makes topic anchors and digital human interaction annotations global attention nodes, enhances the ability to aggregate semantics across turns under the same topic, and reduces the risk of topic confusion, missed identification of minutes, and mismatch of tracking items in long meetings.

[0021] The improved LinUCB algorithm introduces meeting interaction features and interaction constraint corrections on the basis of the original context action selection. This enables the digital human's interaction strategy to be adjusted according to the meeting stage, the status of the matter, and the turn boundary. This reduces the situation where the AI ​​digital human inappropriately intervenes before the speech is finished, and improves the matching of matter restation, confirmation guidance, and voice follow-up questions.

[0022] Through continuous processing of digital human response, task confirmation, minutes tracking, and feedback optimization, meeting minutes can be continuously revised based on participant confirmation and feedback, improving the semantic understanding accuracy, digital human interaction adaptability, and closed-loop management capabilities of the intelligent meeting assistant. Attached Figure Description

[0023] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of the intelligent meeting assistant system based on human-computer interaction and AI digital human proposed in this invention; Figure 2 This is a schematic diagram illustrating the working principle of the improved Longformer model proposed in this invention; Figure 3 This is a flowchart illustrating the improved LinUCB algorithm proposed in this invention for generating digital human interaction strategies. Detailed Implementation

[0024] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0025] refer to Figures 1-3 Intelligent meeting assistant systems based on human-computer interaction and AI digital humans include: The interactive data acquisition module is used to collect the interactive inputs of participants and the interactive commands of the digital human during the meeting, and to obtain the raw interactive data of the meeting. The turn-binding module is used to bind the source of participants' interactions and speaking time periods in the original conference interaction data to obtain an identity-bound turn sequence. The long sequence construction module is used to anchor topics and annotate interactions in identity-bound turn sequences to obtain the long sequence input objects and meeting interaction features of the meeting; The semantic understanding module is used to input the long sequence of conference input objects into the improved Longformer model to obtain the conference semantic understanding results. The improved Longformer model includes a long sequence embedding layer, a turn attention layer, an issue global attention layer, and an item output layer. The issue global attention layer embeds a cross-turn issue retention mechanism. The interactive decision module is used to input the semantic understanding results of the meeting and the interactive features of the meeting into the improved LinUCB algorithm to obtain the digital human interaction strategy; The digital human response module is used to generate digital human meeting interaction responses based on the digital human interaction strategy; The item confirmation module is used to collect participants' interaction confirmations based on the digital human conference interaction responses and obtain the candidate item confirmation results; The minutes tracking module is used to generate meeting minutes based on the confirmation results of candidate items; The feedback optimization module is used to collect feedback from participants on the meeting minutes and update the digital human interaction strategy based on the feedback to obtain the meeting assistant optimization results.

[0026] In this embodiment, the interactive data acquisition module specifically comprises: Upon responding to the meeting start command, the unified meeting clock is activated. The unified clock uses the meeting server time as the reference, and the clock synchronization error is controlled within 50ms. The meeting access record includes the participant account, terminal access sequence number, and AI digital human access sequence number. The interaction acquisition module generates an interaction source identifier according to the meeting number, access sequence number, and access type. The access type includes participant source and digital human source. In the conference audio input channel, the audio sampling rate is set to 16kHz and the sampling bit depth is set to 16bit. The no-speaking audio interval detected before or after the conference starts is used as the environmental noise sampling interval. The interactive acquisition module calculates the short-time energy mean and standard deviation of the audio frames within the environmental noise sampling interval, and generates the voice endpoint threshold by adding 3 times the standard deviation to the short-time energy mean. The participant's voice stream is divided into frames for energy detection based on a 25ms frame length and a 10ms frame shift. When the short-time energy of 5 consecutive frames reaches the voice endpoint threshold, the time of the first qualified voice frame is taken as the voice start time. When the short-time energy of 80 consecutive frames is lower than the voice endpoint threshold, the time of the last valid voice frame is taken as the voice end time. The voice time identifier is generated from the voice start time and the voice end time, and the participant's voice segment is captured. For participants' text input, screen interaction content, and digital human interaction commands, the interaction acquisition module reads the corresponding trigger time to generate an event time identifier and generates an interaction category identifier based on the data source; screen interaction content includes shared page switching, document selection, and whiteboard annotation content; digital human interaction commands include digital human wake-up commands, phase summary commands, task reiteration commands, and confirmation entry trigger commands. The interactive acquisition module binds participants' voice segments, text content, screen interaction content, and digital human interaction commands to interaction source identifiers, interaction category identifiers, and time identifiers, respectively. It writes these data into the meeting interaction cache in ascending order of time identifiers, forming the original meeting interaction data. This provides a data foundation with continuous time and clear source for the turn binding module to identify the source of participants' interactions and speaking periods.

[0027] In this embodiment, the call-turn binding module specifically comprises: After reading the original meeting interaction data, the voice and text categories are classified into speech interaction data according to the interaction category identifier, and the screen interaction and digital human interaction command categories are classified into event interaction data. The voice category data is input into the meeting speech-to-text model to obtain the participant's speech text. The text category data is directly read to obtain the text speech text. Both the participant's speech text and the text speech text inherit the corresponding interaction source identifier and time identifier to form the speech text to be bound. The turn-by-turn interval threshold is set to 1.5s; the text to be bound is read sequentially according to the time identifier, and the time interval between adjacent texts to be bound is calculated; when the interaction source identifiers of adjacent texts to be bound are the same and the time interval is less than or equal to 1.5s, the adjacent texts to be bound are grouped into the same turn-by-turn. When the interaction source identifier changes or the time interval is greater than 1.5s, the current speaking turn ends and a new speaking turn is established; the start time of each speaking turn is taken as the starting point of the time identifier of the first speech text to be bound in the speaking turn, and the end time is taken as the ending point of the time identifier of the last speech text to be bound, thus generating a speaking turn time period; For event-based interactive data, determine whether it falls within the speaking turn's time period based on the event time identifier; when the event time identifier falls within the speaking turn's time period, bind the event-based interactive data to the corresponding speaking turn; when the event time identifier does not fall within any speaking turn's time period, bind the event-based interactive data to the speaking turn with the smallest time difference. After binding is complete, the speech turn text, interaction source identifier, speech turn time period, and event-type interaction data are written into the speech turn record in chronological order to obtain the identity-bound speech turn sequence, which provides a clear source and continuous time basis for the subsequent generation of long sequence input objects in the conference.

[0028] In this embodiment, the long sequence construction module specifically includes: Read the identity-bound turn sequence and the meeting agenda, extract the turn text, interaction source identifier, turn time period, and event-type interaction data from each turn; concatenate the topic sequence number, topic title, and topic description from the meeting agenda into topic text, and use a text encoder to generate a 768-dimensional topic anchor vector; perform sentence segmentation and semantic encoding on the turn text to generate a 768-dimensional turn semantic vector, and calculate the cosine similarity between the turn semantic vector and each topic anchor vector; When the maximum cosine similarity is greater than or equal to 0.62, the corresponding topic anchor point is determined as the topic anchoring result; when the maximum cosine similarity is less than 0.62 and there is a valid topic anchoring result in the previous speaking turn, the topic anchoring result of the previous speaking turn is used; when the screen interaction content associated with the speaking turn corresponds to the topic text in the meeting agenda, the topic anchoring result is corrected with the corresponding topic anchor point; when the maximum cosine similarity of the first speaking turn is less than 0.62, the first topic anchor point in the meeting agenda is used as the topic anchoring result. Generate interaction labels for event-type interactive data according to the data source and trigger time; when the trigger time of a screen interaction event or digital human command event falls within the speech turn time period, write the corresponding interaction label into the speech turn; When the trigger time does not fall within the time period of any speaking turn, calculate the absolute time difference between the trigger time and the time boundary of the adjacent speaking turn, and write the event-type interaction data into the speaking turn with the smallest absolute time difference to obtain the interaction annotation result; According to the order of the turn-taking time periods, the turn-taking text, interaction source identifier, topic anchoring result, and interaction annotation result are written into the long sequence position; the maximum length of a single long sequence segment is set to 4096 text tags. When the maximum length is exceeded, it is divided into continuous segments according to the topic anchoring result, and the last two speaking turns of the previous segment are retained as contextual overlap content between adjacent segments to generate the conference long sequence input object. Based on the interaction annotation results, the trigger status of digital human commands and the confirmation status of participants are statistically analyzed, and the current speaking turn is judged to determine whether the current speaking turn has ended based on the speaking turn time period, thus generating meeting interaction features. The long sequence input objects of the meeting are used to improve the Longformer model for semantic understanding, and the meeting interaction features are used to improve the LinUCB algorithm to generate digital human interaction strategies.

[0029] In this embodiment, the semantic understanding module specifically comprises: The original Longformer model is used to process long text sequences. It mainly reduces the computational cost of long sequence attention by using a local sparse self-attention structure and aggregates key semantics in long texts by using global attention nodes. The improved Longformer model retains the original Longformer model's long text encoding capability, local sparse self-attention structure and global attention computation method, without changing the original model's basic ability to perform long-distance semantic modeling of long meeting texts. The difference in this invention is that the improved Longformer model no longer uses only ordinary text tags as input, but simultaneously inputs the turn text, interaction source identifier, topic anchoring result and interaction annotation result from the long sequence input object of the meeting into the model, and introduces a topic global attention layer at the original Longformer global attention position, embedding a cross-turn topic retention mechanism in the topic global attention layer; The hidden dimension of the long sequence embedding layer is set to 768, and the upper limit of the input length is set to 4096 text tags; the turn-by-turn text calls the text encoding parameters obtained during training to generate the text representation, the interaction source identifier calls the source encoding parameters to generate the source representation, the topic anchoring result calls the topic encoding parameters to generate the topic representation, and the interaction annotation result calls the interaction state encoding parameters to generate the interaction representation. Four types of representations are written into a unified dimension of the conference turn sequence according to the same long sequence position. Among them, the participant source is used to retain the speaking attribution relationship, the topic representation is used to retain the topic attribution relationship, and the interaction representation is used to retain the relationship between digital human instructions, digital human responses, and participant confirmations. The turn attention layer adopts the local sparse self-attention structure of the original Longformer model, and changes the local window that slides according to text tags in the original model to a turn attention window organized according to speech turns. The turn attention window covers the current speaking turn, adjacent speaking turns, and speaking turns on the same topic. Sparse self-attention encoding is performed using 12 attention heads within the window to obtain the turn context representation. The global attention layer sets the topic anchors in the meeting agenda and the speech turns with digital human interaction annotations as global attention nodes. The topic anchors are used to aggregate the semantics of speech turns under the same topic, and the speech turns with digital human interaction annotations are used to aggregate the interactive semantics related to reminders, follow-up questions, restatements of matters, and confirmation guidance. The cross-turn topic retention mechanism generates a topic consistency attention bias based on the topic anchoring results, generates a digital human interaction attention bias based on the interaction annotation results, and uses the two types of attention biases to correct the attention weight of the topic global attention layer, thus obtaining a topic-enhanced conference representation. The item output layer has three output heads. The meeting stage classification head outputs the meeting status identification result based on the enhanced meeting representation of the agenda items. The meeting status includes agenda item development, solution discussion, item confirmation, and task assignment. The minutes item boundary labeling head performs boundary labeling on the resolution expression, conclusion expression, and outstanding issue expression in the agenda item enhanced meeting representation, and generates candidate minutes items. The tracking item boundary annotation header performs boundary annotation on expressions related to task actions, responsible parties, completion deadlines, and confirmation status, generating candidate tracking items; the meeting status identification results, candidate minutes items, and candidate tracking items are combined into the meeting semantic understanding results; Through the above improvements, the Longformer model is enhanced on the basis of the original Longformer model's long text processing capabilities, and further gains the ability to perceive turn ownership, maintain agenda, and perceive digital human interaction semantics for intelligent meeting scenarios. The global attention layer for topics enhances the aggregation effect of cross-turn content under the same topic, and the cross-turn topic retention mechanism reduces semantic drift caused by topic switching, interruption and screen interaction. The item output layer directly transforms the semantics of long meetings into meeting status, candidate minutes items and candidate tracking items, providing stable, continuous and usable semantic input for improving the digital human interaction strategy generated by the LinUCB algorithm.

[0030] In this embodiment, the interactive decision-making module specifically comprises: The original LinUCB algorithm is a context-based multi-arm decision-making algorithm. It calculates the revenue estimates and confidence compensation values ​​of different candidate actions based on the current context vector and selects the candidate action with the highest confidence upper bound score. The improved LinUCB algorithm retains the basic framework of the original LinUCB algorithm, which selects actions based on context vectors and updates action parameters based on action feedback. It does not change the original algorithm's fundamental ability to balance action utilization and action exploration through "revenue estimation + confidence compensation". The difference in this invention is that the improved LinUCB algorithm no longer selects from ordinary recommended actions, but instead uses the AI ​​digital human's meeting interaction actions as candidate digital human actions, and converts the meeting semantic understanding results and meeting interaction features into a digital human decision context vector, so that the action selection process is jointly constrained by the meeting stage, candidate items, turn boundaries and active triggering state. After reading the semantic understanding results and interaction features of the meeting, extract the meeting status recognition results, candidate minutes items and candidate tracking items from the semantic understanding results, and extract the digital human trigger status and turn-by-turn boundary status from the interaction features. The meeting status identification results are converted into meeting stage codes according to topic development, solution discussion, item confirmation and task assignment. Candidate minutes items and candidate tracking items are converted into candidate item codes according to whether there are unconfirmed items. Digital human trigger state and turn-by-turn boundary state are converted into interaction state codes. The meeting stage codes, candidate item codes and interaction state codes are concatenated and then subjected to max-min normalization to obtain the digital human decision context vector. The candidate digital human actions include remaining silent, interface prompts, voice follow-up questions, phase summary, task reiteration, and confirmation guidance; each candidate digital human action corresponds to a set of action parameter matrices and reward vectors. The action parameter matrix is ​​initially set as an identity matrix, the reward vector is initially set as a zero vector, and the confidence adjustment parameter is set to 0.4. For each candidate digital human action, the improved LinUCB algorithm reads the action parameter matrix and reward vector of the corresponding candidate digital human action, and uses the digital human decision context vector as the input for the current action selection. First, the inverse matrix is ​​solved for the action parameter matrix. When the action parameter matrix is ​​not invertible, a numerically stable term of 10^-5 is superimposed on the diagonal position of the action parameter matrix before the inverse matrix is ​​solved. Multiply the inverse of the action parameter matrix by the reward vector to obtain the action parameter vector of the candidate digital human action; perform an inner product operation on the digital human decision context vector and the action parameter vector to obtain the action reward estimate. The square root of the product of the digital human decision context vector, the inverse matrix of the action parameter matrix, and the transpose of the digital human decision context vector is obtained by multiplying them sequentially and then multiplying by the confidence adjustment parameter of 0.4. The action confidence compensation value is obtained by adding the action gain estimate and the action confidence compensation value. A set of action constraint coefficients is generated based on the semantic understanding results and interaction features of the meeting. The turn-taking boundary state in the interaction features is used as the turn-taking boundary information, the candidate item encoding is used as the candidate item information, and the digital human triggering state is used as the digital human triggering information. The set of action constraint coefficients includes turn-taking boundary constraint coefficients, item confirmation constraint coefficients, and active triggering constraint coefficients. The turn-taking boundary constraint coefficient is set to 0.6, the item confirmation constraint coefficient is set to 1.25, and the active triggering constraint coefficient is set to 1.15. For each candidate digital human action, a comprehensive action constraint coefficient is established, with an initial value of 1. When the turn boundary information indicates that the current speech has not ended and the candidate digital human action belongs to voice follow-up questioning, phase summary, or confirmation guidance, the comprehensive action constraint coefficient is multiplied by the turn boundary constraint coefficient. When the candidate item information indicates that there is an unconfirmed item and the candidate digital human action belongs to item restatement or confirmation guidance, the comprehensive action constraint coefficient is multiplied by the item confirmation constraint coefficient. When the digital human trigger information indicates that the participant actively issues a digital human interaction command and the candidate digital human action belongs to interface prompt, voice follow-up questioning, phase summary, item restatement, or confirmation guidance, the comprehensive action constraint coefficient is multiplied by the active trigger constraint coefficient. The action confidence score corresponding to the candidate digital human action is multiplied by the comprehensive action constraint coefficient to obtain the corrected action confidence score. Candidate digital human actions are sorted in descending order according to the corrected action confidence score, and the candidate digital human action with the highest ranking is selected as the digital human action type; when two candidate digital human actions have the same corrected action confidence score, the priority action is determined in the following order: confirmation guidance, matter restation, voice follow-up questioning, stage summary, interface prompts and keeping silent. The strategy-related objects are selected from the current speaking turn, candidate minutes, candidate follow-up items, and current topic anchoring results based on the digital human's action type. When the digital human's action type is to remain silent, the current speaking turn is used as the strategy-related object. When the digital human's action type is an interface prompt or a phase summary, the current topic anchoring result is used as the strategy-related object. When the digital human's action type is a voice follow-up question, the current speaking turn and candidate minutes or candidate follow-up items in the meeting semantic understanding results that are in an unconfirmed state are used as the strategy-related objects. When the digital human's action type is a matter reiteration or confirmation guidance, candidate minutes or candidate follow-up items are used as the strategy-related objects. The highest-ranked candidate digital human action, digital human action type, policy associated object, digital human decision context vector, and action confidence score are combined into a digital human interaction strategy. The digital human decision context vector, candidate digital human action, policy associated object, and action confidence score are saved simultaneously for generating action feedback samples during feedback optimization. Through the above improvements, the LinUCB algorithm is enhanced to further develop dynamic decision-making capabilities for AI digital human conference interaction, building upon the original LinUCB algorithm's context action selection capabilities. Conference stage encoding, candidate item encoding, turn boundary information, and digital human trigger information jointly participate in the generation of action constraint coefficient sets, enabling digital human interaction strategies to be selected around the conference process, item confirmation requirements, speaking boundaries, and active triggering states. This provides the digital human response module with digital human interaction strategies that are adapted to the conference semantics and interaction states.

[0031] In this embodiment, the digital human response module specifically comprises: Read the digital human interaction strategy and parse the digital human action type and strategy associated object. The digital human action type includes keeping silent, interface prompt, voice follow-up question, phase summary, item reiteration and confirmation guidance. The strategy associated object is used to point to the corresponding speaking turn, current topic anchoring result, candidate minutes item or candidate tracking item. When the digital human action type is to keep silent, an internal response record is generated and no content visible to the participants is generated. When the digital human's action type is interface prompt, voice questioning, phase summary, item reiteration or confirmation guidance, the related topic content, related speech content, related minutes or related tracking items are retrieved from the meeting semantic understanding results according to the strategy associated object. The interface prompts call the topic reminder template, the voice follow-up calls the item completion template, the phase summary calls the topic summary template, the item reiteration calls the item reiteration template, and the confirmation guidance calls the confirmation guidance template; the retrieved related content is written into the corresponding template to generate the digital human response text; The response carrier is determined according to the type of digital human action; the digital human response text corresponding to the interface prompt is written into the meeting interface prompt area; the digital human response text corresponding to voice follow-up questions, phase summary, matter reiteration and confirmation guidance is input into the digital human speech synthesis unit to generate digital human speech, and subtitles are generated synchronously according to the response text; the digital human speech sampling rate is set to 24kHz, and the length of a single subtitle is controlled within 32 Chinese characters. If the length is exceeded, it is segmented according to punctuation. When the digital human's action type is a matter reiteration or confirmation guidance, a confirmation entry is generated based on the associated minutes or tracked matters; the confirmation entry is bound to the policy-related object and the temporary number of the associated matter, and the interaction options include confirm, modify, and reject; the digital human's response text, response carrier, policy-related object, and confirmation entry are combined into a digital human meeting interaction response; Through the above processing, the digital human meeting interaction response can transform the digital human interaction strategy into meeting interface prompts, digital human voice, synchronized subtitles, and confirmation entry points, enabling the AI ​​digital human to perform reminders, follow-up questions, summaries, reiterations, and confirmation guidance around meeting matters, and provide an interaction entry point for subsequent matter confirmation.

[0032] In this embodiment, the item confirmation module specifically includes: Read the confirmation entry and strategy-related objects from the digital human conference interaction response. Based on the strategy-related objects, retrieve related minutes or related tracking items from the conference semantic understanding results to generate items to be confirmed. The items to be confirmed retain the item content, topic anchoring result, turn-taking time period, related topics, responsible party, completion deadline, and item source type. The item source type is used to distinguish between minutes and tracking items. Establish a one-to-one binding relationship between the confirmation entry and the items to be confirmed. The binding relationship records the temporary number of the related item and the strategy-related object. The confirmation portal displays the items to be confirmed to participants. The displayed content includes the item details, related topics, and fields that need to be confirmed. The confirmation portal provides three types of instructions: confirm, modify, and reject. When a participant selects the confirmation instruction, the confirmation status of the item to be confirmed will be updated to confirmed. When a participant selects a modification instruction, the system reads the modification content entered by the participant and replaces the corresponding content in the item to be confirmed according to the position of the item field, forming the modified item to be confirmed, and then updates the confirmation status to confirmed; when a participant selects a rejection instruction, the confirmation status of the item to be confirmed is updated to rejected. If no participant input is received within the set confirmation time limit, the item to be confirmed will remain in the pending confirmation state, and a reconfirmation entry will be retained in the digital human conference interaction response. When multiple participants input the same pending confirmation item, the input from the meeting host or the person responsible for the item will be given priority, and the input from other participants will be written into the confirmation remarks. After the confirmation process is completed, the pending confirmation item, item source type, participant input instruction type, confirmation status, topic anchoring result, and turn-around time period will be bound together to generate candidate item confirmation results. Through the above processing, candidate minutes and candidate tracking items can be confirmed, modified, or rejected by the participants before being written into the meeting minutes, reducing the errors in writing items, mismatch of responsible parties, and missing tracking content caused by the direct generation of records by AI digital humans, and providing clear confirmation results of candidate items for minutes and tracking generation.

[0033] In this embodiment, the minutes tracking module specifically includes: Read the pending items, item source type, instruction type, and confirmation status from the candidate item confirmation results; when the confirmation status is confirmed, the corresponding pending item is treated as a valid item; when the confirmation status is not approved, the corresponding pending item is excluded from the meeting item record; when the instruction type is a modification instruction, the modified pending item is written as a valid item; when the instruction type is a confirmation instruction, the original pending item is written as a valid item. Based on the source type of the matter, valid matters are divided into minutes matters and follow-up matters, and the topic anchoring results and turn-around time periods in the valid matters are read; when a valid matter lacks topic anchoring results, the nearest anchored topic is matched according to the turn-around time period. Using the topic anchoring result as the basis for recording partitioning, the minutes are written into the corresponding topic minutes area in ascending order of the turn time period; when there are minutes with the same content in the same topic minutes area, the minutes with the later turn time period are retained. Based on the topic anchoring results, the tracked items are written into the corresponding topic tracking area in ascending order according to the turn-by-turn time period, and the item content, responsible party, completion deadline and confirmation status are written into the tracking field; when the responsible party or completion deadline is not obtained, the corresponding tracking field is marked as pending supplementation. Merge the minutes and tracking areas of each agenda item according to the order of the meeting agenda to generate meeting event records; the meeting event records provide a basis for feedback and optimization by organizing meetings by agenda item, sorting them by time, and having a confirmed status.

[0034] In this embodiment, the feedback optimization module specifically comprises: Read meeting event records, the digital human interaction strategy used when generating meeting event records, and the digital human decision context vector saved when generating digital human interaction strategies; establish feedback association records based on the event content in the meeting event records, the strategy-related objects in the digital human interaction strategies, and candidate digital human actions, so that each meeting event can be associated with the digital human action that triggered the event; The meeting agenda display interface collects feedback instructions from participants regarding the agenda content. Feedback instructions include acceptance, modification, and rejection. The reward value for acceptance is set to 1, and the reward value for rejection is set to -1. For modification, the modified agenda content is read, the number of modified fields is divided by the total number of agenda fields to obtain the difference ratio, and 1 minus 2 times the difference ratio is used to obtain the correction reward value. When the correction reward value is less than -1, the correction reward value is truncated to -1; when the correction reward value is greater than 1, the correction reward value is truncated to 1. At the same time, the modified agenda content is written back to the meeting agenda. Write the digital human decision context vector, candidate digital human actions, and corresponding reward values ​​into the action feedback samples of the improved LinUCB algorithm; for candidate digital human actions that generate feedback, multiply the digital human decision context vector by its own transpose and add it to the parameter matrix of the corresponding action, multiply the reward value by the digital human decision context vector and add it to the reward vector of the corresponding action, and then recalculate the action parameters based on the updated parameter matrix and reward vector. The digital human interaction strategy is updated based on the recalculated action parameters, and the updated digital human interaction strategy is bound to the written meeting agenda record to generate meeting assistant optimization results. Through the above processing, the participants' adoption, modification and rejection of the meeting agenda record can have a reverse effect on improving the LinUCB algorithm, enabling the AI ​​digital human to more accurately select interactive actions such as interface prompts, voice follow-up questions, phase summaries, agenda restatements and confirmation guidance in subsequent meetings.

[0035] Example 1: To verify the feasibility of the present invention in practice, it was applied to a project review meeting scenario in a software development company. The meeting types included requirement review meetings, iteration planning meetings, defect review meetings, and delivery acceptance meetings. These meetings typically revolve around the scope of requirements, development schedule, defect responsibility, delivery risks, and acceptance criteria. During the meetings, there are problems such as multiple people speaking continuously, frequent switching of screen sharing, verbal confirmation of task responsibility, and secondary revision of meeting minutes after the meeting. Traditional meeting assistants are prone to issues such as confusion of topic attribution, omission of items, and breakage of task tracking.

[0036] During application, the conference terminal connects to a microphone, conference chat window, screen sharing component, and AI digital human interactive interface. Participants interact during the conference by speaking, supplementing with text, switching shared pages, and using digital human wake-up commands. This invention uniformly collects and binds the above-mentioned interactive content to form an identity-bound turn sequence. It then combines the conference agenda to generate a long sequence of input objects and conference interaction features. By improving the Longformer model, it identifies the semantic understanding results of the conference, and by improving the LinUCB algorithm, it controls the AI ​​digital human to perform prompts, follow-up questions, repetitions, and confirmation guidance, ultimately generating a record of conference matters.

[0037] Twenty-eight meetings held within a consecutive month were selected as samples. The average meeting duration was 48 minutes, with 5 to 10 participants. Each meeting included audio presentations, text supplements, and screen sharing. Standard minutes and standard tracking items were jointly confirmed by two project managers and one meeting recorder, resulting in a comparative annotation. Comparison methods included manual transcription, standard speech-to-text summarization, standard Longformer semantic recognition with rule-triggered digital human method, and the method of this invention. The meeting minutes of the 28 meetings were manually reviewed, and the output results of each method were averaged. The results are shown in Table 1 below.

[0038] Table 1. Comparison of the processing effects of different meeting assistant methods

[0039] As shown in Table 1, the meeting semantic recognition accuracy of the method of the present invention reaches 92.5%, which is 4.9 percentage points higher than the ordinary Longformer semantic recognition plus rule-triggered digital human method; the item recall rate reaches 89.6%, and the responsibility object matching accuracy reaches 87.9%, indicating that the global attention layer of the topic and the cross-turn topic retention mechanism in the improved Longformer model can reduce topic mismatch and omission of task responsibility objects in long meetings.

[0040] The adoption rate of digital human intervention reached 80.8%, higher than the 63.9% of the rule-triggered digital human method; the minutes correction rate dropped to 8.6%, indicating that after the improved LinUCB algorithm selects digital human interaction strategies based on the meeting semantic understanding results and meeting interaction characteristics, voice follow-up questions, matter repetition and confirmation guidance are closer to the current stage of the meeting and the gaps in the matters, reducing invalid interruptions and error prompts.

[0041] Among them, the meeting semantic recognition accuracy rate represents the proportion of meeting status, topic attribution, and matter boundary identification that are consistent with the results of manual annotation; the matter recall rate represents the proportion of minutes and tracking matters identified by the system that cover matters confirmed by humans; the responsibility object matching accuracy rate represents the proportion of task responsibility objects that are consistent with the responsibility objects confirmed by humans; the digital human intervention adoption rate represents the proportion of participants who accept the prompts, follow-up questions, or confirmation guidance from the AI ​​digital human; the minutes correction rate represents the proportion of matters that are modified by participants after the meeting matter records are generated; and the record generation latency represents the average time from the end of the meeting to the completion of the generation of available meeting matter records.

[0042] The above results demonstrate that this invention preserves the relationships between turn-taking, agenda items, and interaction annotations through long sequence input objects in meetings, improves the semantic understanding capability of long meetings by improving the Longformer model, enhances the adaptability of AI digital human intervention strategies by improving the LinUCB algorithm, and forms a closed loop of meeting items through item confirmation and feedback optimization. This system can reduce the burden of manual minutes preparation in enterprise project review meetings and improve the accuracy and feasibility of meeting item recording.

[0043] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An intelligent meeting assistant system based on human-computer interaction and AI digital human, characterized in that: Includes the following modules: The interactive data acquisition module is used to collect the interactive inputs of participants and the interactive commands of the digital human during the meeting, and to obtain the raw interactive data of the meeting. The turn-binding module is used to bind the source of participants' interactions and speaking time periods in the original conference interaction data to obtain an identity-bound turn sequence. The long sequence construction module is used to anchor topics and annotate interactions in identity-bound turn sequences to obtain the long sequence input objects and meeting interaction features of the meeting; The semantic understanding module is used to input the long sequence of conference input objects into the improved Longformer model to obtain the conference semantic understanding results. The improved Longformer model includes a long sequence embedding layer, a turn attention layer, an issue global attention layer, and an item output layer. The issue global attention layer embeds a cross-turn issue retention mechanism. The semantic understanding module specifically includes: The long sequence embedding layer reads the turn-around text, interaction source identifier, topic anchoring result, and interaction annotation result from the conference long sequence input object, calls the trained text encoding parameters, source encoding parameters, topic encoding parameters, and interaction state encoding parameters, and maps the turn-around text, interaction source identifier, topic anchoring result, and interaction annotation result into a conference turn-around sequence representation of a unified dimension; The turn attention layer is based on the Longformer local sparse self-attention structure. It establishes a turn attention window with the speaking turn as the encoding unit, and writes adjacent speaking turns and speaking turns on the same topic into the turn attention mask. The turn context representation is obtained through multi-head sparse self-attention encoding. The global attention layer sets the topic anchor points in the meeting agenda and the speech turns with digital human interaction annotations as global attention nodes, so that the topic anchor points aggregate the semantics of speech turns under the same topic, and the speech turns with digital human interaction annotations aggregate the semantics of interaction confirmation, thus obtaining a global topic interaction representation. The cross-turn topic retention mechanism generates a topic consistency attention bias based on the topic anchoring result and a digital human interaction attention bias based on the interaction annotation result. The attention weight of the topic global attention layer is corrected by the topic consistency attention bias and the digital human interaction attention bias to obtain the topic-enhanced conference representation. The output layer includes a meeting stage classification header, a minutes item boundary labeling header, and a tracking item boundary labeling header. The meeting stage classification header outputs the meeting status identification result based on the enhanced meeting representation of the agenda items. The minutes item boundary labeling header outputs candidate minutes items based on the enhanced meeting representation of the agenda items. The tracking item boundary labeling header outputs candidate tracking items based on the enhanced meeting representation of the agenda items. The meeting status identification result, candidate minutes items, and candidate tracking items are combined into the meeting semantic understanding result. The interactive decision module is used to input the semantic understanding results of the meeting and the interactive features of the meeting into the improved LinUCB algorithm to obtain the digital human interaction strategy; The interactive decision-making module is specifically as follows: Read the semantic understanding results and interaction features of the meeting, and convert the meeting stage information and candidate item information in the semantic understanding results and the digital human trigger information and turn-taking boundary information in the interaction features into a digital human decision context vector. The following are candidate digital human actions: remaining silent, providing interface prompts, asking follow-up questions via voice, summarizing the phases, reiterating the tasks, and providing confirmation guidance. For each candidate digital human action, the improved LinUCB algorithm is invoked to calculate the action gain estimate and action confidence compensation value based on the digital human decision context vector, and the action gain estimate and action confidence compensation value are combined into an action confidence score; A set of action constraint coefficients is generated based on the semantic understanding results and interaction features of the meeting. The set of action constraint coefficients includes turn-by-turn boundary constraint coefficients, item confirmation constraint coefficients, and active trigger constraint coefficients. When the turn boundary information indicates that the current speech has not ended, the turn boundary constraint coefficient is configured as the constraint coefficient for voice follow-up questions, phase summary, and confirmation guidance; when the candidate item information indicates that there are unconfirmed items, the item confirmation constraint coefficient is configured as the constraint coefficient for item restatement and confirmation guidance. When the digital human trigger message indicates that the participant actively issues a digital human interaction command, the active trigger constraint coefficient is configured as the constraint coefficient for interface prompts, voice follow-up questions, phase summaries, task reiterations, and confirmation guidance; a comprehensive action constraint coefficient is generated based on the constraint coefficient of each candidate digital human action hit, and the action confidence score is multiplied by the comprehensive action constraint coefficient to obtain the corrected action confidence score; The candidate digital human actions are ranked according to the revised action confidence score. The highest-ranked candidate digital human action is taken as the digital human action type. Based on the digital human action type, strategy-related objects are selected from the current speaking turn, candidate minutes, candidate tracking items, and current topic anchoring results. The highest-ranked candidate digital human action, digital human action type, and strategy-related objects are combined into a digital human interaction strategy. The digital human response module is used to generate digital human meeting interaction responses based on the digital human interaction strategy; The item confirmation module is used to collect participants' interaction confirmations based on the digital human conference interaction responses and obtain the candidate item confirmation results. The minutes tracking module is used to generate meeting minutes based on the confirmation results of candidate items; The feedback optimization module is used to collect feedback from participants on the meeting minutes and update the digital human interaction strategy based on the feedback to obtain the meeting assistant optimization results.

2. The intelligent meeting assistant system based on human-computer interaction and AI digital human as described in claim 1, characterized in that, The interactive data acquisition module is specifically: In response to the meeting start command, a unified meeting clock is established, and an interaction source identifier is generated based on the meeting access record; An environmental noise baseline is extracted from the conference audio input channel. A voice endpoint threshold is generated based on the environmental noise baseline. The voice stream of the participants is subjected to frame-by-frame energy detection. Voice frames that continuously reach the voice endpoint threshold are identified as valid voice frames. A voice time identifier is generated based on the start and end frame times of the valid voice frames. The system reads the trigger times of text input by participants, screen interaction content, and digital human interaction commands, generates event time identifiers based on the trigger times, and generates interaction category identifiers based on the data source. The participant's voice segments, text content, screen interaction content, and digital human interaction commands corresponding to the valid voice frames are bound to the interaction source identifier, interaction category identifier, and time identifier, respectively, and sorted in ascending order according to the time identifier to obtain the original meeting interaction data.

3. The intelligent meeting assistant system based on human-computer interaction and AI digital human as described in claim 1, characterized in that, The specific features of the talk-turn binding module are: Read the original meeting interaction data that has been sorted in ascending order by time identifier, and divide the original meeting interaction data into speech interaction data and event interaction data according to the interaction category identifier; Perform speech transcription on the speech segments of participants in the speech-type interactive data to obtain the speech text of the participants, and bind the speech text of the participants with the corresponding interaction source identifier and speech time identifier; The text content in the speech-type interaction data is used as the text speech text, and the text speech text is bound to the corresponding interaction source identifier and event time identifier; Using the participants' spoken text and text spoken text as the text to be bound, the time interval between adjacent texts to be bound is calculated sequentially according to the time identifier; when the interaction source identifier of adjacent texts to be bound is the same and the time interval is less than or equal to the turn interval threshold, the adjacent texts to be bound are grouped into the same turn. When the interaction source identifier changes or the time interval exceeds the turn interval threshold, the current speaking turn ends and a new speaking turn is established. The speaking turn time period is generated based on the time identifier of the first and last unbound speaking text in each speaking turn. Event-type interactive data is bound to speech turns with overlapping time intervals according to the event time identifier. If the event time identifier does not fall into any speech turn time interval, the event-type interactive data is bound to the speech turn with the smallest time interval, thus obtaining the identity-bound speech turn sequence.

4. The intelligent meeting assistant system based on human-computer interaction and AI digital human as described in claim 1, characterized in that, The long sequence construction module is specifically as follows: Read the identity-bound turn sequence and the meeting agenda, and extract the turn text, interaction source identifier, turn time period, and event-type interaction data for each speaking turn from the identity-bound turn sequence; The topic text in the meeting agenda is converted into topic anchors. Semantic encoding is performed on the dialogue text. The semantic encoding results of the dialogue text are matched with the topic anchors for similarity. The topic anchors that meet the anchoring conditions are determined as the topic anchoring results of the corresponding speaking dialogue. Based on the data source and trigger time of the event-type interaction data, screen interaction events and digital human command events that fall within the speaking turn time period are bound to the corresponding speaking turns, and event-type interaction data that do not fall within the speaking turn time period are bound to the speaking turn with the closest time distance, thus obtaining the interaction annotation results; Based on the order of the turn-by-turn time periods, the turn-by-turn text, interaction source identifier, topic anchoring result, and interaction annotation result are written into the corresponding long sequence position to generate a long sequence input object for the meeting; Based on the digital human command triggering status, participant confirmation status, and speaking turn time boundaries in the interaction annotation results, conference interaction features are generated.

5. The intelligent meeting assistant system based on human-computer interaction and AI digital human as described in claim 1, characterized in that, The digital human response module is specifically as follows: Read the digital human interaction strategy and parse the digital human action type and strategy-related objects; When the digital human's action type is to remain silent, generate a digital human meeting interaction response that does not contain content visible to the participants. When the digital human's action type is interface prompt, voice questioning, phase summary, item reiteration or confirmation guidance, the related topic content, related speech content, related minutes or related tracking items are retrieved from the meeting semantic understanding results according to the strategy associated object. Based on the type of digital human action and the retrieved related content, digital human response text is generated. Specifically, interface prompts generate topic reminder text, voice follow-up questions generate item completion text, stage summaries generate topic summary text, item repetitions generate item repetition text, and confirmation guidance generates confirmation guidance text. Determine the response carrier based on the type of digital human action, write the digital human response text corresponding to the interface prompt action into the meeting interface, and convert the digital human response text corresponding to voice follow-up questions, phase summaries, item repetitions or confirmation guidance actions into digital human voice and synchronized subtitles. When the digital human's action type is a summary of an event or a confirmation guide, a confirmation entry is generated based on the associated minutes or tracked events. The digital human response text, response carrier, and strategy-related objects are combined, and a confirmation entry is written when one exists to generate a digital human conference interactive response.

6. The intelligent meeting assistant system based on human-computer interaction and AI digital human as described in claim 1, characterized in that, The specific content confirmation module is as follows: Read the confirmation entry and strategy-related objects in the digital human conference interaction response, retrieve related minutes or related tracking items from the conference semantic understanding results based on the strategy-related objects, generate items to be confirmed carrying topic anchoring results and turn time periods, determine the item source type based on related minutes or tracking items, and bind the items to be confirmed to the confirmation entry. The system displays the items to be confirmed to participants through the confirmation entry point and collects the confirmation, modification, or rejection instructions entered by the participants. When a confirmation command is received, the confirmation status of the item to be confirmed will be updated to confirmed. When a modification instruction is received, the modification content entered by the participant is read, the corresponding content in the item to be confirmed is replaced according to the modification content, and the confirmation status of the modified item to be confirmed is updated to confirmed. When a rejection instruction is received, the confirmation status of the item to be confirmed will be updated to "not approved". The process involves binding the items to be confirmed, the type of item source, the type of instruction entered by the participant, the confirmation status, the topic anchoring result, and the turn-around time period to generate candidate item confirmation results.

7. The intelligent meeting assistant system based on human-computer interaction and AI digital human as described in claim 1, characterized in that, The minutes tracking module is specifically as follows: Read the pending items, item source type, participant input instruction type, and confirmation status from the candidate item confirmation results. Items with a confirmation status of "confirmed" are considered valid items, while pending items with a confirmation status of "not approved" are excluded from the meeting item record. When the instruction type is a modification instruction, the updated pending item in the candidate item confirmation result is used as the content to be written for the valid item; when the instruction type is a confirmation instruction, the original pending item is used as the content to be written for the valid item. Read the topic anchoring results and turn-by-turn time periods carried by valid matters, and classify valid matters into minutes matters and follow-up matters according to the source type of matters; Using the topic anchoring result as the basis for recording partitions, minutes belonging to the same topic anchoring result are written into the corresponding topic minutes area in ascending order of turn-by-turn time; when there are minutes with the same content in the same topic minutes area, the minutes with the later turn-by-turn time are retained. Based on the topic anchoring results, the tracking items belonging to the same topic anchoring results are written into the corresponding topic tracking area in ascending order according to the turn-by-turn time period; Align the fields of item content, responsible party, completion deadline and confirmation status in the tracked items, and write the missing field status to be supplemented for the tracked items; The minutes and tracking areas for each topic are merged according to the order of the meeting agenda to generate a record of meeting items.

8. The intelligent meeting assistant system based on human-computer interaction and AI digital human as described in claim 1, characterized in that, The feedback optimization module specifically includes: Read the meeting event record, the digital human interaction strategy used when generating the meeting event record, and the digital human decision context vector saved when generating the digital human interaction strategy. Read the strategy-related objects and candidate digital human actions from the digital human interaction strategy. Establish feedback association records based on the event content, strategy-related objects, and candidate digital human actions in the meeting event record. The meeting agenda display interface collects feedback instructions from participants regarding the agenda content and determines the feedback type based on the feedback instructions. The feedback types include acceptance feedback, modification feedback, and rejection feedback. Based on the feedback type, the action reward mapping is performed as follows: Positive feedback is rewarded with a positive reward value, and candidate digital human actions associated with the task content are marked as valid actions; Correction feedback is rewarded with a correction reward value, the modified task content is compared with the original task content, the correction reward value is determined based on the comparison results, and the modified task content is written back to the meeting minutes; Rejection feedback is rewarded with a negative reward value, and candidate digital human actions associated with the task content are marked as invalid actions. Write the digital human decision context vector, candidate digital human actions, and corresponding reward values ​​into the action feedback samples of the improved LinUCB algorithm; Update the parameter matrix and reward vector of the corresponding candidate digital human action based on the action feedback sample, recalculate the action parameters of the candidate digital human action, and obtain the updated digital human interaction strategy. Bind the updated digital human interaction strategy to the written-back meeting event records to generate meeting assistant optimization results.

Citation Information

Patent Citations

  • Conference record generation method and device, computer equipment and readable storage medium

    CN120431934A

  • Intelligent conference summary generation method and device, equipment and storage medium

    CN121884818A