Conference interaction method and device, electronic equipment and storage medium

By collecting and recognizing conference audio and video streams in real time, dynamically building a knowledge base and providing multi-dimensional information display, the problem of insufficient conference information integration and operability is solved, thereby improving conference efficiency and intelligence.

CN121750618APending Publication Date: 2026-03-27GUANGZHOU KINGSOFT MOBILE TECH +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies for meeting recording and management suffer from insufficient real-time performance, integration, and operability, leading to low meeting efficiency and difficulties in knowledge retention.

Method used

By capturing audio and video streams in real time during the meeting, performing streaming content recognition to generate real-time transcribed content streams, and updating the meeting-related knowledge base based on this, the system responds to user interactions to provide accurate information services, and achieves multi-dimensional information display and linkage.

Benefits of technology

It enables real-time structuring and continuous accumulation of meeting information, improves the efficiency of information flow and utilization, provides accurate, timely and actionable information feedback, and significantly enhances the level of intelligence in the meeting process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750618A_ABST
    Figure CN121750618A_ABST
Patent Text Reader

Abstract

The invention relates to a conference interaction method and device, electronic equipment and a storage medium, and the method comprises the steps: in a conference process, in response to the update of conference content, adjusting a knowledge base associated with a conference; and in response to an interaction triggering operation of a user, providing corresponding conference information based on the knowledge base. Therefore, the utilization efficiency of the conference information and the intelligent level of cooperative work can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computers, and in particular to a conference interaction method and device, an electronic device, and a storage medium. BACKGROUND

[0002] In the daily operation of various organizations and enterprises, conferences are a key form of collaboration for information synchronization, scheme discussion, and decision-making.

[0003] Currently, conference recording and management mainly rely on a series of separate tools and manual processes. For example, conference content is saved through independent audio or video recording devices. After the conference, a dedicated person needs to listen to and organize the audio and video files to form a written summary. During the conference, if a participant needs to refer to related documents (such as reports, resumes, or schemes), they must switch back and forth between the conference window and the file browser, office software, and other different applications. The task assignments and to-do lists generated during the conference rely on the participants to record them and manually organize them after the conference. In addition, if a specific discussion or decision in a past conference needs to be found, a large amount of written summary or a long playback of the audio and video recording is often required.

[0004] As can be seen, the prior art has obvious deficiencies in the real-time, integration, and operability of conference information, resulting in low conference efficiency, difficulty in knowledge sedimentation, and other problems. SUMMARY

[0005] The present application provides a conference interaction method, device, electronic device, and storage medium to solve the problem of the prior art, which has obvious deficiencies in the real-time, integration, and operability of conference information, resulting in low conference efficiency, difficulty in knowledge sedimentation, and other problems.

[0006] In a first aspect, the present application provides a conference interaction method, which comprises: In the conference process, in response to the update of the conference content, adjusting the knowledge base associated with the conference; In response to the user's interaction trigger operation, providing corresponding conference information based on the knowledge base.

[0007] In a possible implementation, the adjusting of the knowledge base associated with the conference in the conference process in response to the update of the conference content comprises: In the conference process, collecting conference audio and video streams; Performing stream content recognition on the conference audio and video streams to generate a real-time transcription content stream arranged in chronological order; Taking the real-time transcription content stream as new conference content, and incrementally updating the knowledge base associated with the conference based on the new conference content.

[0008] In one possible implementation, the incremental update of the knowledge base associated with the meeting based on the new meeting content includes: Based on the real-time transcribed content stream, perform at least one of the following processes: Based on the real-time transcribed content stream, update the meeting minutes content in the knowledge base; Update the metadata in the knowledge base to associate the timestamps and speaker identifiers with the content in the real-time transcribed content stream; Update the content structure information in the knowledge base to record the content units divided by semantic analysis of the real-time transcribed content stream; Update the association information in the knowledge base to record the association between specified information items extracted from the real-time transcribed content stream and the corresponding content.

[0009] In one possible implementation, the step of providing corresponding meeting information based on the knowledge base in response to a user's interactive trigger operation includes: In response to a user's interactive trigger action targeting a specific speaker, the meeting information associated with that specific speaker is retrieved from the knowledge base; Based on the meeting information associated with the specific speaker, perform at least one of the following operations: The timeline view graphically displays the distribution of speaking time slots for the specific speaker throughout the meeting. Display a continuous audio stream containing segments of the speaker's speech.

[0010] In one possible implementation, the step of providing corresponding meeting information based on the knowledge base in response to a user's interactive trigger operation includes: In response to a user's interactive trigger operation, response content is generated based on the knowledge base for the interactive trigger operation; At least one interactive reference tag is provided for the response content association, and the interactive reference tag is associated with a specific information fragment in the knowledge base that serves as the information source of the response content.

[0011] In one possible implementation, the method further includes: In response to the triggering operation of the interactive reference mark, a positioning and display operation corresponding to the specific information fragment is performed.

[0012] In one possible implementation, the step of providing corresponding meeting information based on the knowledge base in response to a user's interactive trigger operation includes: In response to user interaction, one or more recommended meeting discussion points are generated based on the knowledge base; In response to the selection operation of any of the recommended meeting discussion points, the selected recommended meeting discussion point is imported into the meeting process.

[0013] In one possible implementation, the step of responding to a user instruction by providing meeting information associated with the user instruction based on the knowledge base includes: In response to user interaction triggers, the corresponding target information fragment is determined from the knowledge base; In the corresponding meeting information display view, locate and highlight the target information segment.

[0014] In one possible implementation, the method further includes: During the meeting process, multiple meeting information display views are shown, each of which is used to display different dimensions of information in the meeting content.

[0015] In one possible implementation, the display of multiple meeting information viewpoints includes: The target interface layout mode is determined based on the attributes of the meeting; Arrange multiple meeting information display views according to the target interface layout pattern.

[0016] In one possible implementation, the method further includes: In response to an interactive event, the display state of the plurality of meeting information display views is adjusted, including at least one of: dragging the separator line to adjust the view size, expanding or collapsing a specified view, and entering full-screen focus mode.

[0017] In one possible implementation, the method further includes; In response to a user action in the first meeting information display view, the content associated with the user action is located and displayed in the second meeting information display view, thus creating a linkage between the first and second meeting information display views.

[0018] In one possible implementation, the method further includes: In response to an interactive command initiated on the currently located display content in the second meeting information display view, an intelligent auxiliary response is provided for the currently located display content in the third meeting information display view.

[0019] In one possible implementation, the first meeting information display view is a meeting minutes view, the second meeting information display view is an associated document view, and the third meeting information display view is a smart question and answer view; Alternatively, the first meeting information display view can be an associated document view, and the second meeting information display view can be an intelligent question and answer view.

[0020] In one possible implementation, the method further includes: Extract the tasks to be done from the meeting content; The extracted to-do items are associated with the content fragments in the knowledge base that generated the to-do items; The system displays the to-do items and, in response to an operation on the to-do items, navigates to the meeting information display view where the content segment is located and positions and displays the content segment.

[0021] In one possible implementation, the method further includes: Based on the analysis of the content fragments associated with the to-do items, the scope of responsible persons for the to-do items is determined from the participants of the meeting.

[0022] Secondly, this application provides a conference interaction device, the device comprising: The knowledge base adjustment module is used to adjust the knowledge base associated with the meeting in response to updates to the meeting content during the meeting process. The intelligent interaction module is used to respond to user interaction triggers and provide corresponding meeting information based on the knowledge base.

[0023] In one possible implementation, the knowledge base adjustment module includes: The acquisition unit is used to acquire audio and video streams during the meeting process; A streaming recognition unit is used to perform streaming content recognition on the conference audio and video streams and generate a real-time transcribed content stream arranged in chronological order. The knowledge base update unit is used to take the real-time transcribed content stream as new meeting content and update the knowledge base associated with the meeting incrementally based on the new meeting content.

[0024] In one possible implementation, the knowledge base updating unit is specifically used for: Based on the real-time transcribed content stream, perform at least one of the following processes: Based on the real-time transcribed content stream, update the meeting minutes content in the knowledge base; Update the metadata in the knowledge base to associate the timestamps and speaker identifiers with the content in the real-time transcribed content stream; Update the content structure information in the knowledge base to record the content units divided by semantic analysis of the real-time transcribed content stream; Update the association information in the knowledge base to record the association between specified information items extracted from the real-time transcribed content stream and the corresponding content.

[0025] In one possible implementation, the intelligent interaction module includes: The first interaction unit is used to respond to a user's interactive trigger operation for a specific speaker and retrieve meeting information associated with the specific speaker from the knowledge base; Based on the meeting information associated with the specific speaker, perform at least one of the following operations: The timeline view graphically displays the distribution of speaking time slots for the specific speaker throughout the meeting. Display a continuous audio stream containing segments of the speaker's speech.

[0026] In one possible implementation, the intelligent interaction module includes: The second interaction unit is used to respond to the user's interaction trigger operation and generate response content for the interaction trigger operation based on the knowledge base; At least one interactive reference tag is provided for the response content association, and the interactive reference tag is associated with a specific information fragment in the knowledge base that serves as the information source of the response content.

[0027] In one possible implementation, the intelligent interaction module is further used for: In response to the triggering operation of the interactive reference mark, a positioning and display operation corresponding to the specific information fragment is performed.

[0028] In one possible implementation, the intelligent interaction module includes: The third interaction unit is used to respond to the user's interactive trigger operation and generate one or more recommended meeting discussion points based on the knowledge base; In response to the selection operation of any of the recommended meeting discussion points, the selected recommended meeting discussion point is imported into the meeting process.

[0029] In one possible implementation, the intelligent interaction module includes: The fourth interaction unit is used to determine the corresponding target information fragment from the knowledge base in response to the user's interactive trigger operation. In the corresponding meeting information display view, locate and highlight the target information segment.

[0030] In one possible implementation, the device further includes: The multi-dimensional display module is used to display multiple meeting information display views during the meeting process, and the multiple meeting information display views are used to display different dimensions of information in the meeting content.

[0031] In one possible implementation, the multi-dimensional display module is specifically used for: The target interface layout mode is determined based on the attributes of the meeting; Arrange multiple meeting information display views according to the target interface layout pattern.

[0032] In one possible implementation, the multidimensional display module is further used for: In response to an interactive event, the display state of the plurality of meeting information display views is adjusted, including at least one of: dragging the separator line to adjust the view size, expanding or collapsing a specified view, and entering full-screen focus mode.

[0033] In one possible implementation, the device further includes; The view linkage module is used to respond to user operations in the first meeting information display view, locate and display content associated with the user operations in the second meeting information display view, thereby forming a linkage between the first meeting information display view and the second meeting information display view.

[0034] In one possible implementation, the device further includes: The intelligent assistance module is used to respond to interactive commands initiated on the currently located display content in the second meeting information display view, and to provide intelligent assistance responses for the currently located display content in the third meeting information display view.

[0035] In one possible implementation, the first meeting information display view is a meeting minutes view, the second meeting information display view is an associated document view, and the third meeting information display view is a smart question and answer view; Alternatively, the first meeting information display view can be an associated document view, and the second meeting information display view can be an intelligent question and answer view.

[0036] In one possible implementation, the device further includes: The to-do item extraction module is used to extract to-do items from the meeting content; and associate the extracted to-do items with the content fragments in the knowledge base that generated the to-do items; The to-do item processing module is used to display the to-do items and, in response to operations on the to-do items, jump to the meeting information display view where the content segment is located and position and display the content segment.

[0037] In one possible implementation, the device further includes: The responsible party determination module is used to determine the scope of responsible parties for the to-do items from the participants of the meeting based on the analysis of the content fragments associated with the to-do items.

[0038] Thirdly, this application provides an electronic device, including: a processor and a memory, wherein the processor is configured to execute a conference interaction program stored in the memory to implement the conference interaction method described in any one of the first aspects.

[0039] Fourthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the conference interaction method described in any one aspect.

[0040] Compared with the prior art, the technical solution provided in this application has the following advantages: Firstly, by dynamically constructing and maintaining a meeting knowledge base in real-time response to meeting content updates, the method provides real-time structuring and continuous accumulation of meeting information. This transforms previously fragmented and outdated audio, video, and document content into a real-time, internally interconnected structured data system, fundamentally solving the problems of information processing delays and information silos. Secondly, by responding to user interactions based on this real-time updated knowledge base to provide precise information services, it achieves a leap from static data storage to dynamic intelligent services. This allows users to obtain accurate, timely, and actionable information feedback based on a complete knowledge context that is always synchronized with the meeting's progress. These two aspects together construct a complete intelligent meeting interaction system, significantly improving the efficiency and intelligence level of information flow and utilization during the meeting process. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0042] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0044] Figure 1 A flowchart illustrating an embodiment of a meeting interaction method provided in this application; Figure 2 A schematic diagram of a meeting interface provided for an embodiment of this application; Figure 3 A schematic diagram of another meeting interface provided in an embodiment of this application; Figure 4 A block diagram illustrating an embodiment of a conference interaction device provided in this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0046] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0047] To address the significant shortcomings of existing technologies in terms of the real-time nature, integration, and operability of meeting information, which lead to low meeting efficiency and difficulties in knowledge accumulation, this application provides a meeting interaction method, device, electronic device, and storage medium that can significantly improve the utilization efficiency of meeting information and the level of intelligence in collaborative work.

[0048] Figure 1This is a flowchart illustrating an embodiment of a meeting interaction method provided in this application. In one embodiment, the executing entity of this application embodiment can be a device with data processing and network communication capabilities. The device can be a hardware entity or a software functional module deployed on a hardware entity. When the device is a hardware entity, it can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, desktop computers, servers, etc. When the device is a software functional module, it can be installed and run on the aforementioned electronic devices, providing corresponding network services through client applications or server applications. Specifically, if the device installs and runs a client application, it can act as a user device in network communication; if the device installs and runs a server application, it can act as a server in network communication, providing corresponding services. In the application scenarios involved in this application, the services include, but are not limited to, online meeting services. Figure 1 As shown, the process includes the following steps: Step 101: During the meeting, adjust the knowledge base associated with the meeting in response to updates to the meeting content.

[0049] "Updates to meeting content" refers to any form of information that is continuously generated or added during the meeting process, including but not limited to: content streams generated in real time through streaming content recognition of the meeting audio and video streams, newly uploaded or associated meeting documents (such as presentations and reports), user communication points regarding documents, user screenshots of documents, notes or annotations manually added by users in meeting documents, and key information extracted in real time from the dialogue (such as decision points and to-do items).

[0050] In step 101, in response to updates to such meeting content, adjustments are triggered to the knowledge base associated with the meeting. This includes adding new meeting content to the knowledge base, supplementing or correcting existing content in the knowledge base, establishing or updating relationships between content (such as associating speech segments with corresponding document paragraphs), and structuring the content (such as dividing topics and labeling speakers). Through this step, the knowledge base associated with the meeting evolves synchronously with the progress of the meeting, forming a real-time, complete, and structured meeting information system, providing a data foundation for subsequent intelligent interaction.

[0051] In one embodiment, during the meeting process, in response to updates to the meeting content, the knowledge base associated with the meeting is adjusted, including: during the meeting process, acquiring meeting audio and video streams; performing streaming content recognition on the meeting audio and video streams to generate a real-time transcribed content stream arranged in chronological order; using the real-time transcribed content stream as new meeting content, and incrementally updating the knowledge base associated with the meeting based on the new meeting content.

[0052] Streaming content recognition is a comprehensive real-time processing technology that processes continuously acquired audio and video streams from the conference acquisition end. This processing includes not only speech recognition of the audio stream, converting it into a continuous text stream, but also content recognition of visual information in the video stream. For example, when the video stream contains shared screens, presentation documents, or physical whiteboards, optical character recognition can be performed in real time to extract text from the screen. It can also detect and perform structured analysis on non-textual visual elements such as graphics, charts, and illustrations in the screen, generating corresponding descriptive tags or metadata.

[0053] The resulting "real-time transcription content stream" is not a single text sequence, but a multimodal content sequence aligned to time stamps. Each content unit in this sequence can contain the text content identified at that specific moment, as well as a summary or reference to the visual information associated with that moment. For example, when a speaker discusses a chart on the fifth page of a presentation, the content stream may contain not only a text transcription of their words, but also the chart and its visual recognition results (such as "chart type is bar chart," "title is 'quarterly revenue comparison'," etc.).

[0054] Furthermore, an incremental, structured knowledge building process is initiated by using "real-time transcription of the content stream" as "new meeting content." In one embodiment, this process not only appends text chronologically but also dynamically detects topic boundaries through real-time semantic analysis, intelligently dividing the content stream into logically structured topic paragraphs. Simultaneously, key information items such as keywords, decision conclusions, and to-do items are continuously extracted from the text and annotated with structured tags. All processed meeting content is not loosely stored but organized into a multi-level tree structure, systematically integrated into the knowledge base. For example, this tree structure, from top to bottom, includes at least: Conference level: Stores conference metadata (title, time, participants) and a dynamically updated overall summary; Topic level: For each semantic paragraph, record the topic, start and end time, participating speakers, and AI (Artificial Intelligence) generated summary; Speech level: Saves the original text of each speech, its precise timestamp, the speaker's identifier, and associates it with the documents or actions mentioned. Element level: Attached to the corresponding speech or topic, managing and extracting atomic information such as keywords, to-do items, and decision points.

[0055] In one embodiment, updating the knowledge base associated with the meeting based on the new meeting content increment includes: performing at least one of the following processes based on the real-time transcription content stream: updating the meeting minutes content in the knowledge base based on the real-time transcription content stream; updating the metadata in the knowledge base to associate timestamps and speaker identifiers with the content in the real-time transcription content stream; updating the content structure information in the knowledge base to record the content units divided by semantic analysis of the real-time transcription content stream; and updating the association information in the knowledge base to record the association relationship between specified information items extracted from the real-time transcription content stream and the corresponding content.

[0056] Among them, updating the meeting minutes content in the knowledge base based on real-time transcription of content streams includes continuously adding the identified content (such as text, charts, etc.) to the "Meeting Minutes" main body of the knowledge base in chronological order.

[0057] Updating the metadata in the knowledge base to link timestamps and speaker identifiers with the content in the real-time transcribed content stream includes: building a precise association between content (what was said) and context (who said it and when). This not only preserves the original record of the discussion, but also automatically annotates each sentence with precise timestamps and speaker identifiers, laying the foundation for subsequent time- and person-based retrieval and location.

[0058] The content structure information in the knowledge base is updated to record the content units segmented by semantic analysis of the real-time transcribed content stream. This includes: using AI to perform real-time semantic analysis on the continuous transcribed content stream, automatically detecting semantic boundaries such as topic shifts and changes in discussion focus, thereby intelligently segmenting the content stream into logically meaningful content units (such as paragraphs, corresponding to a complete topic or sub-topic of discussion). This structural information is recorded in the knowledge base, enabling meeting minutes to be upgraded from "a sequence of sentences arranged chronologically" to "a structured document organized by topic."

[0059] Updating the knowledge base with related information to record the relationships between specified information items extracted from the real-time transcription content stream and their corresponding content includes: real-time text analysis to extract key specified information items, including but not limited to keywords (core concepts of the discussion), to-do items (specific tasks generated), and decision points (conclusions reached). These information items are extracted and linked to the original content in the knowledge base that generated them, using structured tags (such as [To-Do], [Decision]). Through this association, users can not only see the original discussion but also quickly overview the tasks and conclusions produced by the meeting, and can navigate back to the detailed discussion context by clicking on tags.

[0060] In summary, the four processing methods described above can work together to transcribe the raw, linear audio and video streams into a multi-level, multi-dimensional, and internally interconnected tree-like knowledge network. This dynamically constructed knowledge network is the data foundation supporting all subsequent intelligent question answering, precise positioning, and cross-dimensional linkage in this embodiment of the application.

[0061] Furthermore, the "meeting-related knowledge base" is physically manifested as a structured data storage format, used to centrally store, index, and manage all meeting-related information. This knowledge base is flexible in deployment; it can be stored on cloud servers for cross-device access and collaboration, or deployed on local servers or terminal devices to meet the needs of data privatization or low-latency processing. From a logical organization perspective, an independent meeting process (such as a single interview or a project review) can correspond to an independent knowledge base instance, ensuring the isolation and security of meeting data. Simultaneously, in practical applications, it also supports multiple related meetings (such as a series of project meetings or multiple rounds of interviews with the same candidate) sharing or being associated with the same knowledge base, enabling cross-meeting content retrieval, knowledge accumulation, and information linkage. This design allows the knowledge base to serve as both a real-time information hub for a single meeting and a repository of accumulated knowledge assets for long-term projects or topics.

[0062] Step 102: In response to user interaction triggers, provide relevant meeting information based on the knowledge base.

[0063] The user's "interaction-triggered actions" refer to any action initiated by the user to obtain meeting information, including but not limited to asking questions in natural language, clicking on interactive elements in the interface (such as links, buttons, timestamps), filtering or searching, etc.

[0064] Upon receiving such an operation, the executing entity in this embodiment first parses its intent and objective, and then performs retrieval, analysis, and reasoning within the knowledge base constructed and continuously updated in step 101. Finally, it provides the corresponding meeting information.

[0065] In one embodiment, "providing relevant meeting information" refers not only to returning text answers, but also includes, but is not limited to: locating and highlighting relevant minutes paragraphs in the interface, jumping to a specific location in the related document, playing audio clips at corresponding time points, visually displaying the speaker's timeline, listing relevant to-do items or recommended discussion points, etc.

[0066] The specific implementation methods and typical interaction processes for "interactive triggering operations" and "providing corresponding meeting information" will be described in detail in different embodiments below, and will not be elaborated here.

[0067] The technical solution provided in this application, on the one hand, dynamically constructs and maintains a meeting knowledge base by responding to real-time updates of meeting content, achieving instant structuring and continuous accumulation of meeting information. This transforms previously fragmented and outdated audio, video, and document content into a real-time, internally interconnected structured data system, fundamentally solving the problems of information processing delays and information silos. On the other hand, based on this real-time updated knowledge base, responding to user interactions to provide precise information services achieves a leap from static data storage to dynamic intelligent services. This allows users to obtain accurate, timely, and actionable information feedback based on a complete knowledge context that is always synchronized with the meeting's progress. These two aspects together construct a complete intelligent meeting interaction system, significantly improving the efficiency and intelligence level of information flow and utilization during the meeting process.

[0068] In one embodiment, the method of this application further includes: displaying multiple meeting information display views during the meeting process, wherein the multiple meeting information display views are respectively used to display different dimensions of information in the meeting content.

[0069] This embodiment aims to overcome the limitations of traditional meetings' single, linear information presentation. By displaying multiple meeting information display views dedicated to different dimensions of information in parallel during the meeting process, complex meeting content is structurally deconstructed and presented synchronously. This constructs a multi-dimensional, three-dimensional information space, allowing key dimensions such as content presentation, participant interaction, real-time communication, and meeting metadata to be extracted and visualized, thereby significantly improving the information density and operational efficiency of the meeting. For example, the multiple meeting information display views include, but are not limited to: a meeting minutes view for displaying structured text, a document view for displaying related document content, a speaker view for visualizing the speaking sequence, an intelligent Q&A view for displaying interactive Q&A, and so on. In the case of multi-person meetings within the same meeting space, different meeting display content and / or different meeting functions can be configured for different participants' meeting interfaces based on permission settings.

[0070] Furthermore, these meeting information display views are not displayed in isolation, but are interconnected and dynamically updated around the same meeting process, providing users (especially the host) with an integrated decision-making and control interface. This design fundamentally changes the way participants obtain and process meeting information, shifting from passively receiving a single information stream to actively planning and integrating multi-dimensional information streams. Ultimately, it optimizes meeting pacing, enhances participant immersion and collaboration depth, and greatly improves the user experience. For example, in response to operations on specific timestamps or spoken content in the meeting minutes view, the document content discussed at that moment is located and displayed in the associated document view, and / or the audio player is controlled to jump to the corresponding time point; in response to operations on specific document paragraphs or pages in the document view, the minutes segment discussing that paragraph or page is located and displayed in the meeting minutes view; in response to operations on citation marks within answers in the smart question-and-answer view, the user jumps to the corresponding information location in the meeting minutes view, document view, or audio player, and so on.

[0071] Therefore, the technical solution provided in this application effectively solves the problems of single information dimension and fragmented operation interface in traditional meeting tools through the association and collaborative display mechanism between multiple views. It organically integrates core information such as documents, minutes, and Q&A that were originally scattered in different windows or applications into a unified interactive interface and realizes real-time linkage between views. Users do not need to perform cumbersome manual switching and searching between multiple independent windows. They can smoothly complete a series of operations such as viewing, locating, tracing, and interacting in an integrated information space, thereby significantly reducing cognitive load and operational costs, and effectively improving the overall efficiency and collaborative experience of meeting information processing.

[0072] The following provides a detailed explanation of the two layouts and two types of scenarios: Layout Scheme 1 Taking an interview scenario as an example, such as Figure 2 As shown, the meeting interface uses a three-column layout (left, middle, and right), organically integrating the core information and tools needed for the interview. Specifically: The left panel displays interview information and documents, centrally presenting static input information. At the top is a candidate's basic information card, displaying the interview stage, interviewee's name, date, etc., for quick identification. The upper half of the panel displays the candidate's resume in a structured format, while the lower half presents the key job requirements side-by-side. This side-by-side comparison provides interviewers with an intuitive framework for matching and analyzing candidate abilities. In practical applications, the information in this area can be automatically retrieved from recruitment platforms or obtained through file parsing.

[0073] The central panel is the interview record and evaluation editing area; this is the interviewer's core workspace. For example... Figure 2As shown, this area allows for the pre-setting of structured interview assessment templates, where interviewers can record key points of candidates' answers, provide sub-scores, or write free evaluations in real time. All records are synchronized to the cloud in real time, ensuring no information is lost and ultimately forming a complete and traceable interview file. Furthermore, this area supports flexible content visibility control. In multi-participant interview meetings, interviewers can set visibility permissions for specific records or assessments for other participants (such as candidates). This means that within the same meeting space, the information accessible to participants with different roles can be presented differently. For example, interviewer-recorded assessments, internal discussion annotations, or certain reference documents can be set to be visible only to the interviewer, and inaccessible to candidates. This mechanism ensures the integrity of the interview process records and the security of internal information, allowing public information and internal assessments to be managed in parallel within the same collaborative environment.

[0074] The right-hand panel serves as the intelligent Q&A and recommendation area, acting as the system's real-time intelligent support hub. It can include two main functional modules: "Intelligent Q&A" and "Question Recommendation." Interviewers can input natural language questions (such as "What are this candidate's strengths?") in the Q&A area, and the system will generate answers with source annotations based on the resume, job description, and previously discussed content. Users can click on the source marker in the answer (such as "Resume - Project Experience"), and the resume panel on the left will automatically scroll to the corresponding content. The question recommendation module dynamically recommends follow-up questions based on resume analysis and real-time dialogue context, assisting interviewers in conducting in-depth interviews.

[0075] As can be seen, this layout, through the clear division of labor and intelligent linkage of multiple views, seamlessly integrates the information and operations that are scattered across multiple browser tabs, documents and notes in traditional interviews into a unified intelligent workspace. Essentially, it reconstructs the interviewer's workflow, enabling them to complete assessments and decisions more focused and efficiently.

[0076] Layout Scheme 2 Taking a project meeting scenario as an example, such as Figure 3 As shown, the meeting interface also uses a three-column layout (left, middle, and right) to create a complete knowledge center encompassing the entire meeting lifecycle. Specifically: The left panel is the meeting resources and document management area, used to collect various resources related to the meeting, such as meeting recordings, AI-generated minutes, documents shared during the meeting (PPT, Word, etc.), pre-meeting preparation materials, and post-meeting action plans. These resources can be intelligently aggregated using rules such as time matching and keyword association to form a complete meeting digital asset package. The bottom of the panel also provides management entry points such as "Upload Files," "Share Space," and "Export & Package" (not shown in the image), facilitating user maintenance and distribution of the knowledge base.

[0077] The central panel is the real-time content display area. The upper half displays the meeting video stream, while the lower half supports switching between multiple modes, including "Speech-to-Text View," "Smart Chapter View," and "Speaker View." In Speech-to-Text View, the transcribed text of the meeting audio is displayed in real-time, with timestamps and speaker identifiers. In Smart Chapter View, AI automatically identifies turning points in the agenda, dividing the meeting content into logical chapters for easy structured browsing. In Speaker View, the speaking time, document operation records, and key conclusion markers for each speaker are visualized in a timeline swimlane chart. Clicking on any event point automatically jumps to the corresponding recording and locates the relevant context, enabling precise playback.

[0078] The right panel is a multi-mode information organization area, supporting switching between "Minutes View," "Document View," and "Mind Map" modes. In Minutes View, structured records and spoken texts can be filtered by speaker and keywords; in Document View, shared files can be previewed, and when the page is hovered over, meeting segments discussing this page can be displayed on the side; in Mind Map View, the meeting discussion is graphically organized, presenting the logical relationships between topics, sub-topics, and conclusions.

[0079] In addition, the panel can be opened or closed via buttons for the intelligent Q&A and to-do list management areas. The intelligent Q&A module allows users to ask questions about the entire meeting content in natural language (e.g., "Regarding the risk of project delays, what countermeasures were proposed at the meeting?"). The system will generate answers by combining the minutes, recordings, and document content, and can locate the original source. The to-do list management module can automatically identify and extract tasks from meeting discussions (e.g., "Zhang San is responsible for submitting the proposal by Friday"), and supports editing, assignment, and one-click synchronization to external task tools to ensure that decisions are implemented.

[0080] As can be seen, this layout transforms a meeting event with a beginning and an end into a living knowledge base and collaboration hub that can be accessed and traced at any time by organically integrating and intelligently linking scattered meeting materials, discussion records and follow-up tasks. This fundamentally reconstructs the user's knowledge management and action process based on meetings, and greatly improves the collaboration efficiency based on meeting information.

[0081] based on Figure 2 and Figure 3 The following are exemplary embodiments of the interface layout provided in this application: In one embodiment, displaying multiple meeting information display views includes: determining a target interface layout mode based on the attributes of the meeting; and organizing the arrangement of multiple meeting information display views according to the target interface layout mode.

[0082] This embodiment aims to address the diverse needs for information organization and presentation across different meeting scenarios. Instead of a fixed, single interface, it intelligently determines the target interface layout based on the meeting's attributes (e.g., the meeting type is labeled "recruitment interview," "project review," or the scenario is automatically identified through analysis of related documents). Each layout mode predefines the arrangement logic and size proportions of multiple functional panels (such as minutes panel, document panel, and Q&A panel). These views are then arranged according to the selected mode.

[0083] For example, a three-column layout of "document on the left, records in the middle, and questions and answers on the right" might be automatically adopted for interview scenarios, while a layout that focuses on resource management and timeline visualization might be adopted for internal discussion scenarios.

[0084] Furthermore, the interface layout in this embodiment can not only adapt to the meeting type but also dynamically adjust based on the user's role (such as host or regular participant). The system can identify the current user's role and present differentiated function panels or information content within the same meeting interface accordingly. For example, for the meeting host, the interface may include more management controls, internal evaluation panels, or settings options; while for regular participants, the interface may focus on core content display and personal interaction functions, with some management views or sensitive information automatically hidden or disabled. This role-driven differentiated interface presentation, while ensuring information security and operational permissions, further optimizes the user experience and interaction efficiency for users with different roles.

[0085] This embodiment achieves a high degree of alignment between the interface and the meeting objectives, thereby enhancing the user experience.

[0086] In one embodiment, in response to an interactive event, the display state of multiple meeting information display views is adjusted, including at least one of: dragging a separator line to adjust the view size, expanding or collapsing a specified view, and entering a full-screen focus mode.

[0087] This embodiment aims to provide a highly personalized user experience to adapt to different screen spaces, attention focuses, and individual operating habits.

[0088] Specifically, it allows users to actively adjust the display status of various information display views through a variety of interactive events. This includes: adjusting the width or height ratio of adjacent views in real time by dragging the separator line; expanding or collapsing secondary or auxiliary views by clicking the button to temporarily expand the display space of the main content area; and hiding all auxiliary panels by switching to full-screen focus mode, allowing users to focus entirely on the content of a specific view (such as a long document or detailed minutes).

[0089] This embodiment ensures that the system provides an optimal visual workspace under different devices and different stages of use.

[0090] In one embodiment, in response to a user operation in the first conference information display view, content associated with the user operation is located and displayed in the second conference information display view, thus forming a linkage between the first conference information display view and the second conference information display view.

[0091] The core of this embodiment lies in responding to user actions in one view and providing direct contextual feedback in another. When a user performs an action (such as clicking a timestamp, a speaker's name, or a keyword) in the first meeting information display view (e.g., meeting minutes), the system understands the user's intention to view related details and automatically locates and displays content directly related to that action in the second meeting information display view (e.g., a related document view or a speaker timeline view). For example, clicking on "the chart on page five" mentioned in the minutes will automatically flip the page and highlight the chart in the document view on the right. This linkage breaks down the barriers between views, forming an intuitive information network navigation.

[0092] In one embodiment, in response to an interactive command initiated on the currently located display content in the second conference information display view, an intelligent auxiliary response is provided for the currently located display content in the third conference information display view.

[0093] This embodiment further enhances the system's proactive service capabilities based on cross-view linkage. Once a user has located specific content in the second view (e.g., viewing a project description in a resume) based on the aforementioned linkage, and if further questions or needs arise based on this content, the user can initiate a new interactive command (e.g., selecting text and clicking "Ask AI"). At this point, the third meeting information display view (e.g., an intelligent Q&A view) is activated, providing intelligent auxiliary responses to the currently located content (e.g., generating in-depth questions about the project experience, performing skills matching analysis, etc.). This achieves a seamless progression from "finding information" to "deeply understanding information."

[0094] The interface layout and linkage mechanism provided in the above embodiments jointly construct a flexible and collaborative multi-dimensional information presentation space. This space is not only a "display platform" for information, but also an "operating console" and "response interface" for realizing various intelligent interactive functions. Based on this, the embodiments of this application can also support the following series of core interactive functions that are deeply suited to meeting scenarios: In one embodiment, in response to a user's interactive triggering operation, providing corresponding meeting information based on a knowledge base includes: in response to a user's interactive triggering operation for a specific speaker, retrieving meeting information associated with the specific speaker from the knowledge base; and based on the meeting information associated with the specific speaker, performing at least one of the following operations: graphically displaying the distribution of the specific speaker's speaking time segments in the meeting process in a timeline view; and displaying a continuous audio stream containing speaking segments of the specific speaker.

[0095] This embodiment specifically provides a method for monitoring and analyzing the speeches of specific participants and providing intelligent interactive feedback. When a user initiates an interactive trigger operation targeting a specific speaker by clicking on the speaker's avatar, filtering the list, or entering a natural language command containing the speaker's name, the meeting information associated with that specific speaker is retrieved from the knowledge base constructed in step 101. This meeting information can be a structured collection, including not only the original text of each sentence spoken by the speaker, but also the precise timestamp of each sentence, the topic it belongs to, and the document content that may have been associated with the speech.

[0096] Based on meeting information associated with a specific speaker, at least one highly visual and convenient contextualized response action can be performed: 1. In the speaker view, the distribution of a specific speaker's speaking time throughout the meeting is displayed graphically: This transforms the abstract speaking record into an intuitive timeline graph. For example, on a timeline representing the total meeting duration, color blocks or line segments of different lengths and colors clearly indicate the start and end times and duration of each speaker's remarks. This visualization allows users to quickly grasp the speaker's activity pattern (e.g., whether they speak intensively at the beginning of the meeting or participate throughout), speaking frequency, and the rhythm of their interaction with other speakers.

[0097] 2. Display a continuous audio stream containing segments of a specific speaker's speech: This operation provides an efficient, targeted review experience. Instead of playing the entire meeting recording and having users manually skip irrelevant parts, the system automatically extracts and splices audio clips of all the speaker's statements from the original recording based on the timestamp information associated with the speaker in the knowledge base, generating a continuous audio stream containing only their statements for playback. This is equivalent to creating a "personal set of statements" for the user, which is particularly suitable for quickly reviewing their points, evaluating their expression, or preparing feedback.

[0098] In summary, this embodiment transforms the structured information in the knowledge base, which is associated by speaker dimension, into two forms that best suit users' cognitive habits: visual graphics and continuously playable audio. This greatly optimizes the experience of reviewing, analyzing, and understanding the contributions of specific individuals in a meeting, and solves the pain points of "insufficient utilization of speaker information" and "low review efficiency" in the prior art.

[0099] In one embodiment, in response to a user's interactive trigger operation, corresponding meeting information is provided based on a knowledge base, including: generating response content for the interactive trigger operation based on the knowledge base; providing at least one interactive reference tag for associating the response content, the interactive reference tag being associated with a specific information fragment in the knowledge base that serves as the information source of the response content. Further, in response to a trigger operation on the interactive reference tag, performing a positioning and display operation corresponding to the specific information fragment.

[0100] In this embodiment, when a user initiates interaction through natural language questions or triggering operations, the execution entity of this application embodiment will retrieve and understand the knowledge base associated with the meeting to form corresponding "response content." For example, a user might ask questions like: "What were the discussions about project progress during the meeting?", "What tasks was Zhang San responsible for?", or "What were the conclusions of the discussion on the fifth slide of the PPT?". More importantly, while generating the response content, the system automatically identifies and records the knowledge base information source upon which each key assertion or fact in the response content is based. For example, when answering "What are Zhang San's views on project risks?", not only is a summary provided, but it also links this summary to a specific speech segment by Zhang San at a certain point in time, found in the knowledge base.

[0101] Furthermore, to make the aforementioned source information visible and usable to users, interactive citation tags are provided to associate the response content. These interactive citation tags can be visual elements embedded in the answer text (such as highlighted text, underlined citations, icons, etc.), associated with a specific information fragment in the knowledge base that serves as its source. This fragment could be a speech paragraph, a document page, or a point in time. For example, the end of the answer might display "[Source: Minutes 15:30]" or "[See the Resume Project Experience section]", which are examples of interactive citation tags.

[0102] When a user clicks or hovers over an interactive reference marker, the system automatically performs the location and display operations corresponding to the specific information segment associated with that marker. For example, if the marker is associated with a speech segment in meeting minutes, the system can automatically jump to the minutes view, scroll and highlight the text, and simultaneously control the audio player to jump to the corresponding moment for playback, achieving "audio-visual synchronization." If the marker is associated with a section of a document (such as a resume), the system can automatically open the document in document view and scroll to the relevant paragraph for highlighting. If the marker is associated with a to-do item, the system can jump to the to-do panel and locate the task, while also allowing a reverse association back to the context in which the minutes were generated.

[0103] The interactive method proposed in this embodiment implements a response mechanism that integrates information tracing and rapid location. By associating interactive reference tags with the intelligently generated response content and allowing users to directly jump to the original context of the information source by triggering the tags, this scheme technically achieves a bidirectional traceable link between the answer and the source text. This not only significantly improves the efficiency of users in verifying and exploring information after obtaining it, but also optimizes the navigation experience in multi-source, structured meeting information by reducing interface switching and manual search steps, thereby enhancing the practicality and reliability of the system's intelligent interaction.

[0104] In one embodiment, in response to a user's interactive triggering operation, providing corresponding meeting information based on a knowledge base includes: in response to a user's interactive triggering operation, generating one or more recommended meeting discussion points based on the knowledge base; and in response to a selection operation of any recommended meeting discussion point, importing the selected recommended meeting discussion point into the meeting process.

[0105] This embodiment aims to provide proactive and intelligent assistance for meeting processes. It offers a dynamic guidance mechanism to address potential issues in meetings (such as interviews and reviews with clear agendas), including discussion interruptions, omissions of key points, or insufficient preparation for questions.

[0106] The process of generating recommended meeting discussion points is not random; instead, it involves real-time analysis of multi-source information in the knowledge base to generate meaningful "recommended meeting discussion points." These information sources can include: static document content: for example, in an interview scenario, analyzing a candidate's resume and job description to automatically identify skill matching points, project experience highlights, or potential questions (such as gaps in work experience), thereby generating questions such as "Please describe in detail the technical challenges you faced in the XX project" or "Please explain this period of career gap."

[0107] In one embodiment, the generated recommended points are not merely a static list for browsing, but are directly actionable. When a user (such as an interviewer) clicks to select a recommended discussion point, the system will perform an "import meeting process" operation. This means that, in terms of interface interaction, the question can be inserted or copied to the currently used meeting minutes document, question list, or chat box with a single click, eliminating the hassle of manual input. In terms of process recording, the system can record when the recommended point is used and associate it with subsequent discussion responses, ensuring the integrity of the meeting minutes.

[0108] This embodiment proactively generates valuable discussion guidance suggestions through deep understanding of meeting materials and real-time dialogue, and integrates them seamlessly into the meeting flow through a user-friendly interactive design. This not only reduces the preparation burden on the meeting facilitator but also ensures that the meeting can cover key issues more systematically and comprehensively, directly improving the quality and productivity of the meeting itself.

[0109] In one embodiment, in response to a user instruction, providing meeting information associated with the user instruction based on a knowledge base includes: in response to a user's interactive triggering operation, determining a corresponding target information fragment from the knowledge base; and locating and highlighting the target information fragment in a corresponding meeting information display view.

[0110] This embodiment aims to understand the user's discussion intent in real time during the meeting process, and based on a structured knowledge base, accurately locate and highlight the content that the user is currently concerned with in the correct information view.

[0111] For example, when a user mentions content related to a relevant document (such as a resume or report) during a meeting discussion (e.g., an interviewer asks "What was your experience at XX company?"), the transcribed text stream is analyzed in real time. Keywords such as "XX company" are detected, and the corresponding section is located in the relevant document (e.g., "Work Experience Paragraph 2"). Then, in the corresponding document view (e.g., the left-hand resume panel), scrolling and highlighting are automatically performed, directly presenting the relevant document paragraph to the user. This process achieves "see what you say," greatly improving the efficiency of information comparison and verification.

[0112] This embodiment achieves intelligent real-time linkage between meeting discussions and the knowledge base. It can automatically and accurately map the focus of users' oral discussions to the specific location of relevant documents and present it to users intuitively, effectively avoiding the tedious operations of manual searching, page turning, and comparison required in traditional methods. This not only significantly improves the immediacy and accuracy of information acquisition and verification but also makes the meeting process smoother, reducing interruptions caused by information switching, thereby improving overall meeting efficiency and the convenience of collaborative work.

[0113] In one embodiment, to-do items are extracted from the meeting content; the extracted to-do items are associated with content fragments in the knowledge base that generate to-do items; the to-do items are displayed, and in response to operations on the to-do items, the system jumps to the meeting information display view where the content fragments are located and positions and displays the content fragments. Further, based on the analysis of the content fragments associated with the to-do items, the scope of responsible parties for the to-do items is determined from the meeting participants.

[0114] This embodiment provides an intelligent meeting task management method, which aims to solve the problems of scattered records, unclear sources, and unclear allocation of to-do items in traditional meetings, and realize closed-loop management of the entire process from identification and association to allocation.

[0115] Specifically, firstly, through real-time semantic analysis, to-do items are automatically extracted from continuously updated meeting content (such as transcribed text). This differs from simple note-taking; it refers to identifying statements with clear action directions (such as "Next, we need to complete the solution design" or "Please have Zhang San provide the test report").

[0116] Next, each extracted to-do item is linked to the original discussion content fragment that generated that item in the knowledge base. This ensures that each task carries the complete context in which it was generated, forming a traceable "task-source" link.

[0117] Based on this association, a list of to-do items is displayed in a dedicated panel. When a user performs an action on any item in the list (such as clicking its source link), the system can use the above association to immediately jump to the corresponding meeting information display view (such as the minutes view), and automatically locate and highlight the specific discussion segment that generated the task, thereby achieving one-click tracing from abstract tasks to specific discussion contexts.

[0118] Furthermore, this embodiment includes an intelligent responsibility allocation mechanism: it analyzes the original content fragments associated with the to-do item, and intelligently determines whether the item is an individual task or a group task based on the semantics of the discussion (e.g., whether the task is explicitly assigned to an individual or requires teamwork). Based on this, it determines the scope of responsible persons from the list of participants (e.g., associating with a single person or multiple relevant participants). This automates and rationalizes task allocation, reduces the workload of manual assignment, and improves accuracy.

[0119] In summary, this embodiment transforms loosely structured tasks generated during meetings into actionable items with clear structure, traceable origins, and well-defined allocation, significantly improving the efficiency and traceability of meeting decisions.

[0120] Based on the above embodiments, the technical solution provided in this application constructs an intelligent interaction system that spans the entire lifecycle of a meeting. This system uses a structured meeting knowledge base built in real-time and incrementally as its core data hub. It provides users with an integrated working interface through adaptive and interconnected multi-dimensional information display views. Furthermore, it supports a series of interactive functions, including speaker analysis, traceable intelligent question answering, intent-driven positioning, and intelligent task management. The entire solution achieves a complete closed loop from real-time perception and structured accumulation of meeting content to contextualized, on-demand retrieval and in-depth utilization of information. This fundamentally changes the limitations of traditional meeting tools, such as information lag, single-dimensionality, and passive interaction, significantly improving the efficiency of meeting information processing, decision-making quality, and knowledge collaboration effectiveness.

[0121] See Figure 4 This is a block diagram illustrating an embodiment of a conference interaction device provided in this application. Figure 4 As shown, the device includes: The knowledge base adjustment module 41 is used to adjust the knowledge base associated with the meeting in response to updates to the meeting content during the meeting process. The intelligent interaction module 42 is used to respond to user interaction triggers and provide corresponding meeting information based on the knowledge base.

[0122] In one possible implementation, the knowledge base adjustment module 41 includes: The acquisition unit is used to acquire audio and video streams during the meeting process; A streaming recognition unit is used to perform streaming content recognition on the conference audio and video streams and generate a real-time transcribed content stream arranged in chronological order. The knowledge base update unit is used to take the real-time transcribed content stream as new meeting content and update the knowledge base associated with the meeting incrementally based on the new meeting content.

[0123] In one possible implementation, the knowledge base updating unit is specifically used for: Based on the real-time transcribed content stream, perform at least one of the following processes: Based on the real-time transcribed content stream, update the meeting minutes content in the knowledge base; Update the metadata in the knowledge base to associate the timestamps and speaker identifiers with the content in the real-time transcribed content stream; Update the content structure information in the knowledge base to record the content units divided by semantic analysis of the real-time transcribed content stream; Update the association information in the knowledge base to record the association between specified information items extracted from the real-time transcribed content stream and the corresponding content.

[0124] In one possible implementation, the intelligent interaction module 42 includes: The first interaction unit is used to respond to a user's interactive trigger operation for a specific speaker and retrieve meeting information associated with the specific speaker from the knowledge base; Based on the meeting information associated with the specific speaker, perform at least one of the following operations: The timeline view graphically displays the distribution of speaking time slots for the specific speaker throughout the meeting. Display a continuous audio stream containing segments of the speaker's speech.

[0125] In one possible implementation, the intelligent interaction module 42 includes: The second interaction unit is used to respond to the user's interaction trigger operation and generate response content for the interaction trigger operation based on the knowledge base; At least one interactive reference tag is provided for the response content association, and the interactive reference tag is associated with a specific information fragment in the knowledge base that serves as the information source of the response content.

[0126] In one possible implementation, the intelligent interaction module 42 is further configured to: In response to the triggering operation of the interactive reference mark, a positioning and display operation corresponding to the specific information fragment is performed.

[0127] In one possible implementation, the intelligent interaction module 42 includes: The third interaction unit is used to respond to the user's interactive trigger operation and generate one or more recommended meeting discussion points based on the knowledge base; In response to the selection operation of any of the recommended meeting discussion points, the selected recommended meeting discussion point is imported into the meeting process.

[0128] In one possible implementation, the intelligent interaction module 42 includes: The fourth interaction unit is used to determine the corresponding target information fragment from the knowledge base in response to the user's interactive trigger operation. In the corresponding meeting information display view, locate and highlight the target information segment.

[0129] In one possible implementation, the device further includes: The multi-dimensional display module is used to display multiple meeting information display views during the meeting process, and the multiple meeting information display views are used to display different dimensions of information in the meeting content.

[0130] In one possible implementation, the multi-dimensional display module is specifically used for: The target interface layout mode is determined based on the attributes of the meeting; Arrange multiple meeting information display views according to the target interface layout pattern.

[0131] In one possible implementation, the multidimensional display module is further used for: In response to an interactive event, the display state of the plurality of meeting information display views is adjusted, including at least one of: dragging the separator line to adjust the view size, expanding or collapsing a specified view, and entering full-screen focus mode.

[0132] In one possible implementation, the device further includes; The view linkage module is used to respond to user operations in the first meeting information display view, locate and display content associated with the user operations in the second meeting information display view, thereby forming a linkage between the first meeting information display view and the second meeting information display view.

[0133] In one possible implementation, the device further includes: The intelligent assistance module is used to respond to interactive commands initiated on the currently located display content in the second meeting information display view, and to provide intelligent assistance responses for the currently located display content in the third meeting information display view.

[0134] In one possible implementation, the first meeting information display view is a meeting minutes view, the second meeting information display view is an associated document view, and the third meeting information display view is a smart question and answer view; Alternatively, the first meeting information display view can be an associated document view, and the second meeting information display view can be an intelligent question and answer view.

[0135] In one possible implementation, the device further includes: The to-do item extraction module is used to extract to-do items from the meeting content; and associate the extracted to-do items with the content fragments in the knowledge base that generated the to-do items; The to-do item processing module is used to display the to-do items and, in response to operations on the to-do items, jump to the meeting information display view where the content segment is located and position and display the content segment.

[0136] In one possible implementation, the device further includes: The responsible party determination module is used to determine the scope of responsible parties for the to-do items from the participants of the meeting based on the analysis of the content fragments associated with the to-do items.

[0137] like Figure 5 As shown in the figure, this application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113 is used to store computer programs; In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the conference interaction method provided in any of the foregoing method embodiments, including: During the meeting, the knowledge base associated with the meeting is adjusted in response to updates to the meeting content; In response to user interaction, relevant meeting information is provided based on the knowledge base.

[0138] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the conference interaction method provided in any of the foregoing method embodiments.

[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0141] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0142] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A meeting interaction method, characterized in that, The method includes: During the meeting, the knowledge base associated with the meeting is adjusted in response to updates to the meeting content; In response to a user's interactive triggering action, the system provides meeting information associated with the interactive triggering action based on the knowledge base.

2. The method according to claim 1, characterized in that, During the meeting process, in response to updates to the meeting content, the knowledge base associated with the meeting is adjusted, including: During the meeting, capture the audio and video streams of the meeting; Perform streaming content recognition on the conference audio and video streams to generate a real-time transcribed content stream arranged in chronological order; The real-time transcribed content stream is used as new meeting content, and the knowledge base associated with the meeting is updated incrementally based on the new meeting content.

3. The method according to claim 2, characterized in that, The incremental update of the knowledge base associated with the meeting based on the new meeting content includes: Based on the real-time transcribed content stream, perform at least one of the following processes: Based on the real-time transcribed content stream, update the meeting minutes content in the knowledge base; Update the metadata in the knowledge base to associate the timestamps and speaker identifiers with the content in the real-time transcribed content stream; Update the content structure information in the knowledge base to record the content units divided by semantic analysis of the real-time transcribed content stream; Update the association information in the knowledge base to record the association between specified information items extracted from the real-time transcribed content stream and the corresponding content.

4. The method according to claim 1, characterized in that, The response to user interaction triggers an operation, providing corresponding meeting information based on the knowledge base, including: In response to a user's interactive trigger action targeting a specific speaker, the meeting information associated with that specific speaker is retrieved from the knowledge base; Based on the meeting information associated with the specific speaker, perform at least one of the following operations: The timeline view graphically displays the distribution of speaking time slots for the specific speaker throughout the meeting. Display a continuous audio stream containing segments of the speaker's speech.

5. The method according to claim 1, characterized in that, The response to user interaction triggers an operation, providing corresponding meeting information based on the knowledge base, including: In response to a user's interactive trigger operation, response content is generated based on the knowledge base for the interactive trigger operation; At least one interactive reference tag is provided for the response content association, and the interactive reference tag is associated with a specific information fragment in the knowledge base that serves as the information source of the response content.

6. The method according to claim 5, characterized in that, The method further includes: In response to the triggering operation of the interactive reference mark, a positioning and display operation corresponding to the specific information fragment is performed.

7. The method according to claim 1, characterized in that, The response to user interaction triggers an operation, providing corresponding meeting information based on the knowledge base, including: In response to user interaction, one or more recommended meeting discussion points are generated based on the knowledge base; In response to the selection operation of any of the recommended meeting discussion points, the selected recommended meeting discussion point is imported into the meeting process.

8. The method according to claim 1, characterized in that, The response to a user instruction, providing meeting information associated with the user instruction based on the knowledge base, includes: In response to user interaction triggers, the corresponding target information fragment is determined from the knowledge base; In the corresponding meeting information display view, locate and highlight the target information segment.

9. The method according to claim 1, characterized in that, The method further includes: During the meeting process, multiple meeting information display views are shown, each of which is used to display different dimensions of information in the meeting content.

10. The method according to claim 9, characterized in that, The presentation of multiple meeting information views includes: The target interface layout mode is determined based on the attributes of the meeting; Arrange multiple meeting information display views according to the target interface layout pattern.

11. The method according to claim 9, characterized in that, The method further includes: In response to an interactive event, the display state of the plurality of meeting information display views is adjusted, including at least one of: dragging the separator line to adjust the view size, expanding or collapsing a specified view, and entering full-screen focus mode.

12. The method according to claim 9, characterized in that, The method further includes; In response to a user action in the first meeting information display view, the content associated with the user action is located and displayed in the second meeting information display view, thus creating a linkage between the first and second meeting information display views.

13. The method according to claim 12, characterized in that, The method further includes: In response to an interactive command initiated on the currently located display content in the second meeting information display view, an intelligent auxiliary response is provided for the currently located display content in the third meeting information display view.

14. The method according to claim 12 or 13, characterized in that, The first meeting information display view is a meeting minutes view, the second meeting information display view is a related document view, and the third meeting information display view is a smart question and answer view; Alternatively, the first meeting information display view can be an associated document view, and the second meeting information display view can be an intelligent question and answer view.

15. The method according to claim 1, characterized in that, The method further includes: Extract the tasks to be done from the meeting content; The extracted to-do items are associated with the content fragments in the knowledge base that generated the to-do items; The system displays the to-do items and, in response to an operation on the to-do items, navigates to the meeting information display view where the content segment is located and positions and displays the content segment.

16. The method according to claim 15, characterized in that, The method further includes: Based on the analysis of the content fragments associated with the to-do items, the scope of responsible persons for the to-do items is determined from the participants of the meeting.

17. A conference interaction device, characterized in that, The device includes: The knowledge base adjustment module is used to adjust the knowledge base associated with the meeting in response to updates to the meeting content during the meeting process. The intelligent interaction module is used to respond to user interaction triggers and provide corresponding meeting information based on the knowledge base.

18. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute a conference interaction program stored in the memory to implement the conference interaction method according to any one of claims 1-16.

19. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the conference interaction method according to any one of claims 1-16.