Conference summary automatic generation method and device, equipment and medium
By using cross-modal semantic alignment and graph construction, the problems of fragmented meeting information and inaccurate summaries were solved, enabling efficient and structured meeting content summarization and knowledge management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MERCHANTS FINANCIAL LEASING CO LTD
- Filing Date
- 2025-12-10
- Publication Date
- 2026-05-08
AI Technical Summary
Existing meeting processing tools lack cross-modal correlation, resulting in fragmented information, errors in speech-to-text conversion, insufficient speaker identification, and a lack of structured and differentiated meeting summaries, making it difficult to support subsequent reuse and traceability.
By acquiring meeting images, text, and audio, cross-modal semantic alignment is performed to construct a semantic network, extract structured information, generate meeting discussion flowcharts and knowledge graphs, and form hierarchical summaries.
It achieves highly accurate summarization of meeting content, effective integration of cross-modal information, and supports subsequent reuse and traceability.
Smart Images

Figure CN121996786A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for automatically generating meeting summaries. Background Technology
[0002] Meetings are a core setting for organizational communication, decision-making, and work advancement. The accurate recording and efficient transmission of their content (such as discussion viewpoints, decision results, and task assignments) are crucial for team collaboration. With the increasing prevalence of remote work, meeting formats are becoming more diverse, encompassing multimodal information such as video feeds, shared screens (PPT / whiteboard), real-time chat, and voice conversations, placing higher demands on the comprehensiveness and integration capabilities of information processing.
[0003] Existing meeting processing tools have significant limitations: On the one hand, they often handle a single modality separately (e.g., speech-to-text tools only output text, and screen-sharing tools only record images), lacking cross-modal connections (e.g., the "Solution B" mentioned in the speech does not correspond to the PPT slides presented at the same time), resulting in fragmented information; on the other hand, speech-to-text often contains errors due to technical terms and accents, and lacks speaker identification labels, affecting the judgment of content attribution; in addition, meeting summaries are mostly simple text piles, not structured according to the logic of "issue-decision-task", and cannot generate differentiated information based on roles, resulting in low efficiency in extracting core content; at the same time, key entities in the meeting (such as solutions and tasks) and their relationships have not formed a knowledge system, making it difficult to support subsequent reuse and traceability. Summary of the Invention
[0004] This invention provides a method, apparatus, computer equipment, and medium for automatically generating meeting summaries, in order to solve the problem of low accuracy in existing meeting summary methods on the market.
[0005] Firstly, a method for automatically generating meeting summaries is provided, including: Acquire the meeting image, meeting text, and meeting audio of the target meeting; perform text conversion on the meeting audio to obtain the converted text; Cross-modal semantic alignment is performed on the meeting images, the meeting text, and the transformed text sequence to obtain a semantic network; Extract the structured information of the semantic network, and construct a meeting discussion flowchart based on the transformed text and the timestamp of the structured information. Generate a summary of the discussion content based on the meeting discussion flowchart. A meeting knowledge graph is constructed based on the semantic network, and a hierarchical summary of the meeting is generated based on the meeting knowledge graph. The summary of the discussion content and the hierarchical summary of the meeting are then combined to obtain a summary of the meeting content.
[0006] Secondly, an automatic meeting summary generation device is provided, comprising: The data conversion module is used to acquire the meeting images, meeting text, and meeting audio of the target meeting, and to convert the meeting audio into text to obtain converted text; The semantic alignment module is used to perform cross-modal semantic alignment on the conference images, the conference text, and the transformed text sequence to obtain a semantic network. The graph construction module is used to extract the structured information of the semantic network, construct a meeting discussion flowchart based on the transformed text and the timestamp of the structured information, and generate a summary of the discussion content based on the meeting discussion flowchart. The summary generation module is used to construct a conference knowledge graph based on the semantic network, generate a hierarchical summary of the conference based on the conference knowledge graph, and summarize the discussion content summary and the hierarchical summary of the conference to obtain a conference content summary.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for automatically generating conference summaries.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described method for automatically generating meeting summaries.
[0009] The above-mentioned scheme for automatically generating meeting summaries, including methods, devices, computer equipment, and storage media, can improve the accuracy of meeting summaries by acquiring meeting images, text, and audio from the target meeting; converting the audio into text to obtain converted text; performing cross-modal semantic alignment on the meeting images, text, and converted text sequences to obtain a semantic network; extracting the structured information from the semantic network; constructing a meeting discussion flowchart based on the timestamps of the converted text and the structured information; generating a summary of discussion content based on the meeting discussion flowchart; constructing a meeting knowledge graph based on the semantic network; generating a hierarchical summary of the meeting based on the meeting knowledge graph; and summarizing the summary of discussion content and the hierarchical summary of the meeting to obtain a meeting content summary. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for the automatic generation method of meeting summaries in one embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for automatically generating meeting summaries according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a meeting summary automatic generation device in one embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The automatic meeting summary generation method provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain meeting images, text, and audio from the target meeting through the client. It performs text conversion on the audio to obtain converted text, performs cross-modal semantic alignment on the meeting images, text, and converted text sequences to obtain a semantic network, extracts the structured information from the semantic network, and constructs a meeting discussion flowchart based on the timestamps of the converted text and the structured information. It then generates a summary of the discussion content based on the flowchart, constructs a meeting knowledge graph based on the semantic network, generates a hierarchical summary of the meeting based on the knowledge graph, and summarizes the discussion content summary and the hierarchical summary to obtain a meeting summary, thus improving the accuracy of the meeting summary. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the automatic meeting summary generation method provided in this embodiment of the invention includes the following steps: S1. Acquire the meeting image, meeting text, and meeting audio of the target meeting, and perform text conversion on the meeting audio to obtain the converted text.
[0015] In this embodiment of the invention, the meeting images may include digital content displayed on a shared screen (such as PPT slides, data reports, and design drawings), images obtained by photographing or scanning handwritten content on a physical whiteboard (such as flowcharts, formula derivations, and meeting notes), and real-time images captured by the participants' cameras during the video conference.
[0016] In this embodiment of the invention, the meeting text refers to written materials generated before, during, or in connection with the meeting. It may include pre-uploaded electronic documents (such as structured agenda PDFs, requirements specifications in Word, and background information in Markdown), real-time text discussions and voting records generated in chat tools (such as Teams / Slack) during the meeting, and pre-set auxiliary information (such as participant name-role mapping tables and project-specific terminology databases). Additionally, it may also include structured summary data from past meetings.
[0017] In this embodiment of the invention, the conference audio is the core real-time audio information stream of the conference, typically originating from conference recordings or real-time audio transmission from online conferencing systems. Its core value lies in carrying the spoken expressions of the participants, including speaker identification markers and precise timestamp information.
[0018] In this embodiment of the invention, the step of converting the conference audio into text to obtain converted text includes: Based on speech activity detection, non-silent segments in the conference speech are identified to obtain segmented audio block sequences; Voiceprint recognition is performed on the segmented audio block sequence to obtain voiceprint recognition data; Based on a preset voiceprint database, the speaker's identity is labeled in the segmented audio block sequence according to the voiceprint recognition data to obtain the identity-labeled audio block; A pre-trained speech recognition model is used to generate timestamped recognition text based on the identified audio blocks; The identified text is then corrected using context and a pre-defined domain terminology list to obtain the converted text.
[0019] In detail, the step of identifying non-silent segments in the conference audio based on speech activity detection to obtain a segmented audio block sequence involves scanning the audio stream using a speech activity detection algorithm to identify valid speech segments (such as human voices and non-background noise) with energy exceeding a silence threshold, and then segmenting them into independent segments based on silence intervals (>1.5 seconds). This step filters out invalid audio such as coughs and keystrokes, ensuring that subsequent processing focuses on valid content.
[0020] In detail, the speech activity detection algorithm can determine speech activity by calculating the short-time energy of the speech signal. When the short-time energy exceeds a certain threshold, speech activity is considered to exist.
[0021] In detail, the voiceprint recognition of the segmented audio block sequence can be performed by using a voiceprint feature extraction model (such as x-vector / ECAPA-TDNN) to extract the speaker's biometric vector (such as spectral envelope, fundamental frequency mode) from each audio block and converting it into a mathematically represented voiceprint fingerprint.
[0022] In detail, the pre-trained speech recognition model is a model that converts speech signals into text information. For example, a hidden Markov model can model the temporal characteristics of speech signals through state transition probabilities and emission probabilities.
[0023] In detail, the terminology correction of the identified text by combining the context and the preset domain terminology list can be achieved by using the semantics of adjacent text to correct ambiguities (e.g., "testing department" misidentified as "testing department" → corrected according to the topic "quality testing"), and by replacing abbreviations / proper terms according to the terminology list.
[0024] S2. Perform cross-modal semantic alignment on the conference image, the conference text, and the transformed text sequence to obtain a semantic network.
[0025] In this embodiment of the invention, the step of performing cross-modal semantic alignment on the meeting image, the meeting text, and the transformed text sequence to obtain a semantic network includes: OCR text extraction and target detection are performed on the meeting images to obtain image semantic units; The meeting text is parsed in a structured manner to obtain structured text; The image semantic units, the structured text, and the transformed text sequence are mapped to a unified time axis to obtain a time-synchronized multimodal stream; Cross-modal entity connections are performed on the time-synchronized multimodal stream to obtain an entity connection graph; Joint semantic encoding is performed on the entity connectivity graph and the time-synchronized multimodal stream to obtain a cross-modal semantic vector sequence; The semantic network is constructed based on the cross-modal semantic vector sequence and the entity connectivity graph.
[0026] Specifically, the OCR text extraction and object detection of the meeting image are to extract the text content in the image (such as PPT titles, handwritten formulas on whiteboards) through an optical character recognition (OCR) engine, and at the same time use an object detection model (such as YOLO) to identify graphic elements (chart / flowchart nodes). The detected elements are annotated with spatial positions and timestamps.
[0027] Specifically, the structured parsing of the meeting text is to split the agenda document into topic nodes with time estimates by chapter; the chat records identify task-based statements (such as "@Li Si submit the report before Wednesday") through an intent classification model; and a pre-set glossary is used to complete entity standardization.
[0028] Specifically, mapping the image semantic units, the structured text, and the transformed text sequence to a unified timeline to obtain a time-synchronized multimodal stream is based on the speech timestamp, and the dynamic time warping (DTW) algorithm is used to calculate the temporal offset of other modalities from the speech, aligning image page-turning events and agenda nodes to the speech timeline. The clock offset is corrected through cross-correlation analysis (such as a 0.5-second delay in the video stream).
[0029] Specifically, cross-modal entity connection of the time-synchronized multimodal stream to obtain an entity connection graph is to extract entities (person names / organizations / technical terms) in each modality, and link multimodal expressions of the same entity through a coreference resolution algorithm. For example, "@Engineer Wang" in the chat record + "General Manager Wang" in the speech → entity#Wang Wu, "This solution" in the text + "Solution A" marked on the PPT in the image → entity#Scheme_A.
[0030] Specifically, the joint semantic encoding of the entity connection graph and the time-synchronized multimodal stream is to perform feature fusion using a multimodal Transformer (such as VL-BERT).
[0031] In detail, the construction of the semantic network based on the cross-modal semantic vector sequence and the entity connection graph uses entity nodes (such as decision schemes, participants, and task items) in the entity connection graph as the basic nodes of the semantic network. Then, through dynamic contextual relationship analysis of the cross-modal semantic vector sequence, the logical relationship type between entities is automatically inferred—for example, based on the spatial distribution and temporal continuity of semantic vectors, relationships such as "support", "oppose", and "dependency" are identified, generating directional edges such as [Zhang San]-[Proposal]->[Solution A]. At the same time, discrete event nodes (issue proposal, point of contention, conclusion) are linked into a temporal causal chain based on timestamps, forming a complete event flow of issue discussion → solution debate → voting decision. The multimodal evidence (voice segment ID, PPT coordinates, text position) stored in the entity connection graph is bound as the traceable attributes of the nodes. Finally, a machine-parseable graph structure semantic network that integrates entity topological relationships, event temporal logic, and multimodal evidence chains is output.
[0032] S3. Extract the structured information of the semantic network, and construct a meeting discussion flowchart based on the converted text and the timestamp of the structured information. Generate a summary of the discussion content based on the meeting discussion flowchart.
[0033] In this embodiment of the invention, extracting the structured information of the semantic network includes: The semantic network is used to perform type recognition using a pre-trained graph neural network to obtain a set of classification marker points; Based on the relation edges of the semantic network, relation paths are extracted from the set of classification marker points to obtain structured relation paths; Based on the timestamps of the semantic network, the structured relation paths are bound by temporal constraints to obtain structured units; The multimodal evidence contained in the semantic network is connected to the structured unit to obtain structured information.
[0034] In detail, the step of using a pre-trained graph neural network to perform type recognition on the semantic network and obtain a set of classified labeled points involves using the graph neural network to analyze the semantic features and topological connection patterns of nodes, identifying and labeling key node types (such as decision points, action items, and controversial issues). The model classifies nodes based on their attributes (textual description, relation density) and adjacency structure (such as decision nodes often being associated with multiple "support / oppose" edges), outputting a set of nodes with type labels.
[0035] In detail, the process of extracting relational paths from the set of classification markers based on the relational edges of the semantic network to obtain structured relational paths involves scanning the semantic network based on predefined relational path templates (such as proposal → discussion → decision) to extract node sequences that satisfy specific logical chains. A directed path matching algorithm is then used to capture key processes (such as the action item allocation-responsible person-deadline chain), transforming unstructured graph relationships into machine-parseable semantic chains.
[0036] In detail, the process of binding the structured relationship paths with time-series constraints based on the timestamps of the semantic network to obtain structured units involves extracting time metadata (such as decision timestamps and task deadline texts) from semantic network nodes and converting it into standardized time constraints using time parsing rules (such as "this Friday" → specific date). Time-series attributes (such as decision effective time and task deadline) are then bound to each structured path to form executable time-sensitive instructions.
[0037] In detail, connecting the multimodal evidence contained in the semantic network to the structured unit to obtain structured information involves injecting the multimodal evidence indexes (speech segment IDs, image coordinates, text positions) stored in the semantic network into the structured unit. Through cross-modal reference mapping (such as associating decision points with corresponding PPT slides for a given time period), a verifiable original data source is attached to each piece of structured information, thus constructing an audit trail chain.
[0038] In this embodiment of the invention, constructing a meeting discussion flowchart based on the timestamps of the converted text and the structured information includes: The converted text and the structured information are subjected to time-series time node identification to obtain a time-series time node set; Perform time topology sorting on the set of time nodes to obtain the event sequence; The converted text is semantically recognized to obtain the semantics of the converted text; Identify the logical relationships of the transformed text semantics based on the event sequence; A meeting discussion flow graph is generated based on the logical relationships and the event sequence.
[0039] In detail, the step of identifying time nodes in the converted text and the structured information involves identifying discussion topics (such as "budget proposal debate") from the converted text, extracting key nodes (such as "decision: pass proposal B") from the structured information, merging events with the same theme (such as consecutive discussions of "technical risks" as a single node) based on timestamps, and generating a comprehensive node set covering discussion events and decision-making actions.
[0040] In detail, the process of performing time topological sorting on the set of time nodes to obtain the event sequence involves first arranging the nodes in ascending order by timestamp; for events with overlapping times (such as multiple people speaking at the same time), priority is calculated based on the speaking density weight (number of speakers × duration × keyword importance) to ensure that high-profile discussions are prioritized; and finally, a conflict-free linear event sequence is generated.
[0041] In detail, the semantic recognition of the transformed text to obtain the semantics of the transformed text is achieved by using the Semantic Role Labeling (SRL) model and dependency parsing to parse the semantic structure of each text segment, extracting the action subject (such as "Zhang San"), behavior ("oppose"), object ("Solution A"), and logical connectives ("because" / "but"), and constructing a machine-understandable predicate-argument framework.
[0042] In detail, the step of identifying the logical relationship of the transformed text semantics based on the event sequence is to inject logical markers into the transformed text semantics based on the temporal position of the event sequence, such as causal inference, opposition detection, and task association.
[0043] In detail, the process of generating a meeting discussion flow graph based on the logical relationships and the event sequence involves constructing a directed graph with the time-ordered event sequence as the main path and the set of logical relationships as the branch edges.
[0044] In this embodiment of the invention, the step of generating a summary of discussion content based on the meeting discussion flowchart involves extracting the main chronological path in the flowchart (such as topic proposal → solution debate → voting decision → task allocation) as the narrative skeleton, and then injecting logical relational semantics—when a branch node (such as a point of contention) is detected, a transitional statement is inserted (such as "but some members question..."); at the same time, multimodal evidence summaries are bound to key nodes (such as "based on data analysis on page 5 of PPT"), and differentiated templates are applied according to the node type: decision points highlight the conclusion ("finally approved solution B"), and action items specify the responsible person and time ("Li Si needs to submit a report before Friday"). Finally, the chronological context, logical connections, and evidence tracing are integrated to generate a coherent natural language summary.
[0045] S4. Construct a meeting knowledge graph based on the semantic network, generate a hierarchical summary of the meeting based on the meeting knowledge graph, and summarize the discussion content and the hierarchical summary of the meeting to obtain a summary of the meeting content.
[0046] In this embodiment of the invention, constructing a conference knowledge graph based on the semantic network includes: Extract the core entities and original relation edges of the semantic network; The core entities are then normalized to obtain a set of normalized entities; The original relation edges are labeled with their types to obtain a set of valid relation edges; The attribute information of the normalized entity set and the effective relation edge set is bound to obtain the entity relation structure; The entity relationship structure is subjected to hierarchical constraints to obtain a hierarchical entity relationship structure; A meeting knowledge graph is generated based on the hierarchical entity relationship structure.
[0047] In detail, the extraction of the core entities and original relation edges of the semantic network is achieved by identifying and extracting key entities and their associated relationships through a preset whitelist of entity types (such as people, products, and decisions) and relation patterns (such as subject-verb-object structures).
[0048] In detail, the normalization process for the core entities, resulting in a normalized entity set, is achieved through an entity alignment algorithm. Combining a predefined entity dictionary (such as a list of attendees) and contextual semantics (such as "General Manager Wang" referring to himself as "Wang Wei" in speech), synonymous entities are mapped to a unified identifier (such as "Wang Wei (ID:001)"). Simultaneously, entity names are standardized (e.g., expanding "AI Project" to "Artificial Intelligence Project"). In detail, the step of annotating the original relation edges to obtain a set of valid relation edges is to analyze the text description of the relation edges (such as "Zhang San suggests adopting scheme B") using a natural language processing model and annotate their relation types ("suggestion").
[0049] In detail, binding attribute information to the normalized entity set and the effective relation edge set refers to adding static attributes (such as "Zhang San - Role - Product Manager") and dynamic attributes (such as "Solution A - Creation Time - 20250701") to entities, and adding timestamps (such as "Zhang San - Proposal - Solution A - Time - 09:15:23") and evidence sources (such as "Voice Segment ID: V003") to relation edges. In detail, the hierarchical constraint on the entity relationship structure is based on entity type and relationship logic, establishing parent-child hierarchical relationships (such as "Meeting" → "Agenda" → "Decision") and inclusion relationships (such as "Project Budget" including "Hardware Procurement" and "Personnel Costs"). Through hierarchical constraints, the flat entity relationship structure is transformed into a tree network.
[0050] In detail, the generation of the meeting knowledge graph based on the hierarchical entity relationship structure involves constructing a complete meeting knowledge graph using hierarchical entities as nodes, typified relationship edges as connections, and attribute information as supplementary information. This graph supports complex queries (such as "find all solutions proposed by the product department and supported by the technical team") and intelligent analysis (such as identifying decision conflict points through relationship paths).
[0051] In this embodiment of the invention, generating a hierarchical summary of a meeting based on the meeting knowledge graph includes: The core path node sequence of the meeting knowledge graph is identified based on node importance calculation; Based on the preset user role configuration, a hierarchical summary template is injected into the core path node sequence to obtain a hierarchical summary skeleton. Obtain the node description of each node in the meeting knowledge graph, and perform semantic condensation and evidence binding based on the hierarchical summary skeleton and the node description to obtain a summary paragraph with evidence tags; A hierarchical summary of the meeting is generated based on the cross-level association of the summary paragraphs with evidence markers.
[0052] In detail, the identification of the core path node sequence of the meeting knowledge graph based on node importance calculation involves using the PageRank algorithm to identify highly central nodes (such as decision points), then connecting related events (such as discussion nodes before decision-making) through time-series backtracking, and finally generating a linear node sequence that runs through the core issues, decisions, and action items.
[0053] In detail, the hierarchical summary template injection of the core path node sequence based on the preset user role configuration is to dynamically select the summary framework from the preset template library according to the role requirements.
[0054] In detail, the step of obtaining the node description of each node in the conference knowledge graph, performing semantic condensation and evidence binding based on the hierarchical summary skeleton and the node description involves compressing the node description using a text summarization model (e.g., "Solution B can reduce costs by 20%" → "Solution B significantly reduces costs"), attaching multimodal source tracing tags to key conclusions, and converting evidence IDs into accessible links through a coordinate mapping algorithm.
[0055] In detail, the step of generating a hierarchical summary of the meeting based on the cross-level association of the summary paragraphs with evidence tags is to analyze the logical dependency of the hierarchical summary and add bidirectional interactive links.
[0056] As can be seen, the above scheme first acquires images (including PPT slides, whiteboard images, participant images, etc.), text (including agenda documents, chat logs, terminology databases, etc.), and audio from the target meeting. For the audio, non-silent segments are first identified, and the speaker's identity is marked using voiceprint recognition. A timestamped text is generated using a speech recognition model, and then combined with context and domain terminology correction to obtain the converted text. Subsequently, cross-modal semantic alignment is performed on the multimodal data: images are extracted using OCR to obtain text and target detection to obtain semantic units. After structured parsing of the text, the three are synchronized into a multimodal stream along the timeline. Cross-modal entities are connected to form a graph, and semantic vectors are obtained through joint semantic encoding. Finally, a semantic network containing entity nodes, logical relationships, and multimodal evidence is constructed. Next, the structured information of the semantic network is extracted: graph neural networks are used to identify node types, extract relationship paths, and bind timestamps and associated evidence; then, time-series nodes are identified and sorted into event sequences. The semantic logic of the text is analyzed to construct a meeting discussion flowchart, from which a summary of the discussion content containing the timeline and evidence is generated. Finally, a knowledge graph of the meeting is constructed based on the semantic network: entities are standardized, relationship types are labeled and attributes are bound, and hierarchical constraints are added; then the core path of the graph is identified, templates are injected according to user roles, semantics are condensed and evidence is bound to generate hierarchical summaries, and the discussion content is summarized and hierarchical summaries are compiled to obtain the final summary of the meeting content.
[0057] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0058] In one embodiment, an automatic meeting summary generation device is provided, which corresponds one-to-one with the automatic meeting summary generation method described in the above embodiments. For example... Figure 3 As shown, the automatic conference summary generation device includes a data conversion module 101, a semantic alignment module 102, a graph construction module 103, and a summary generation module 104. Detailed descriptions of each functional module are as follows: Data conversion module 101 is used to acquire meeting images, meeting text, and meeting audio of the target meeting, and to convert the meeting audio into text to obtain converted text; Semantic alignment module 102 is used to perform cross-modal semantic alignment on the conference image, the conference text, and the transformed text sequence to obtain a semantic network; The graph construction module 103 is used to extract the structured information of the semantic network, construct a meeting discussion flowchart based on the converted text and the timestamp of the structured information, and generate a summary of the discussion content based on the meeting discussion flowchart. The summary generation module 104 is used to construct a conference knowledge graph based on the semantic network, generate a hierarchical summary of the conference based on the conference knowledge graph, and summarize the discussion content summary and the hierarchical summary of the conference to obtain a conference content summary.
[0059] In one embodiment, when performing text conversion on the conference audio to obtain converted text, the data conversion module 101 is specifically used for: Based on speech activity detection, non-silent segments in the conference speech are identified to obtain segmented audio block sequences; Voiceprint recognition is performed on the segmented audio block sequence to obtain voiceprint recognition data; Based on a preset voiceprint database, the speaker's identity is labeled in the segmented audio block sequence according to the voiceprint recognition data to obtain the identity-labeled audio block; A pre-trained speech recognition model is used to generate timestamped recognition text based on the identified audio blocks; The identified text is then corrected using context and a pre-defined domain terminology list to obtain the converted text.
[0060] In one embodiment, the semantic alignment module 102, when performing cross-modal semantic alignment of the conference image, the conference text, and the transformed text sequence to obtain a semantic network, is specifically used for: OCR text extraction and target detection are performed on the meeting images to obtain image semantic units; The meeting text is parsed in a structured manner to obtain structured text; The image semantic units, the structured text, and the transformed text sequence are mapped to a unified time axis to obtain a time-synchronized multimodal stream; Cross-modal entity connections are performed on the time-synchronized multimodal stream to obtain an entity connection graph; Joint semantic encoding is performed on the entity connectivity graph and the time-synchronized multimodal stream to obtain a cross-modal semantic vector sequence; The semantic network is constructed based on the cross-modal semantic vector sequence and the entity connectivity graph.
[0061] In one embodiment, the graph construction module 103, when performing the extraction of structured information from the semantic network, is specifically used for: The semantic network is used to perform type recognition using a pre-trained graph neural network to obtain a set of classification marker points; Based on the relation edges of the semantic network, relation paths are extracted from the set of classification marker points to obtain structured relation paths; Based on the timestamps of the semantic network, the structured relation paths are bound by temporal constraints to obtain structured units; The multimodal evidence contained in the semantic network is connected to the structured unit to obtain structured information.
[0062] In one embodiment, the graph construction module 103, when performing the construction of the meeting discussion flowchart based on the timestamps of the converted text and the structured information, is specifically used for: The converted text and the structured information are subjected to time-series time node identification to obtain a time-series time node set; Perform time topology sorting on the set of time nodes to obtain the event sequence; The converted text is semantically recognized to obtain the semantics of the converted text; Identify the logical relationships of the transformed text semantics based on the event sequence; A meeting discussion flow graph is generated based on the logical relationships and the event sequence.
[0063] In one embodiment, the summary generation module 104, when performing the step of constructing a conference knowledge graph based on the semantic network, is specifically used for: Extract the core entities and original relation edges of the semantic network; The core entities are then normalized to obtain a set of normalized entities; The original relation edges are labeled with their types to obtain a set of valid relation edges; The attribute information of the normalized entity set and the effective relation edge set is bound to obtain the entity relation structure; The entity relationship structure is subjected to hierarchical constraints to obtain a hierarchical entity relationship structure; A meeting knowledge graph is generated based on the hierarchical entity relationship structure.
[0064] In one embodiment, the summary generation module 104, when performing the step of generating a hierarchical summary of the conference based on the conference knowledge graph, is specifically used for: The core path node sequence of the meeting knowledge graph is identified based on node importance calculation; Based on the preset user role configuration, a hierarchical summary template is injected into the core path node sequence to obtain a hierarchical summary skeleton. Obtain the node description of each node in the meeting knowledge graph, and perform semantic condensation and evidence binding based on the hierarchical summary skeleton and the node description to obtain a summary paragraph with evidence tags; A hierarchical summary of the meeting is generated based on the cross-level association of the summary paragraphs with evidence markers.
[0065] This invention provides an automatic meeting summary generation device. First, it acquires images (including PPT slides, whiteboard images, participant images, etc.), text (including agenda documents, chat logs, terminology databases, etc.), and audio from the target meeting. For the audio, non-silent segments are identified, and speaker identities are marked using voiceprint recognition. A timestamped text is generated using a speech recognition model, and then combined with context and domain terminology correction to obtain the converted text. Next, cross-modal semantic alignment is performed on the multimodal data: images are extracted using OCR to obtain text and target detection to obtain semantic units. After structured parsing of the text, the three are synchronized into a multimodal stream along the timeline. Cross-modal entities are connected to form a graph, and semantic vectors are obtained through joint semantic encoding. Finally, a semantic network containing entity nodes, logical relationships, and multimodal evidence is constructed. Then, the structured information of the semantic network is extracted: a graph neural network is used to identify node types, extract relationship paths, and bind timestamps and associated evidence; then, time-series nodes are identified and sorted into event sequences. The semantic logic of the text is analyzed to construct a meeting discussion flowchart, based on which a summary of the discussion content containing the timeline and evidence is generated. Finally, a knowledge graph of the meeting is constructed based on the semantic network: entities are standardized, relationship types are labeled and attributes are bound, and hierarchical constraints are added; then the core path of the graph is identified, templates are injected according to user roles, semantics are condensed and evidence is bound to generate hierarchical summaries, and the discussion content is summarized and hierarchical summaries are compiled to obtain the final summary of the meeting content.
[0066] Specific limitations regarding the automatic meeting summary generation device can be found in the limitations of the automatic meeting summary generation method described above, and will not be repeated here. Each module in the aforementioned automatic meeting summary generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0067] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a server-side method for automatically generating meeting summaries.
[0068] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the client-side functions or steps of a method for automatically generating meeting summaries. In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire the meeting image, meeting text, and meeting audio of the target meeting; perform text conversion on the meeting audio to obtain the converted text; Cross-modal semantic alignment is performed on the meeting images, the meeting text, and the transformed text sequence to obtain a semantic network; Extract the structured information of the semantic network, and construct a meeting discussion flowchart based on the transformed text and the timestamp of the structured information. Generate a summary of the discussion content based on the meeting discussion flowchart. A meeting knowledge graph is constructed based on the semantic network, and a hierarchical summary of the meeting is generated based on the meeting knowledge graph. The summary of the discussion content and the hierarchical summary of the meeting are then combined to obtain a summary of the meeting content.
[0069] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire the meeting image, meeting text, and meeting audio of the target meeting; perform text conversion on the meeting audio to obtain the converted text; Cross-modal semantic alignment is performed on the meeting images, the meeting text, and the transformed text sequence to obtain a semantic network; Extract the structured information of the semantic network, and construct a meeting discussion flowchart based on the transformed text and the timestamp of the structured information. Generate a summary of the discussion content based on the meeting discussion flowchart. A meeting knowledge graph is constructed based on the semantic network, and a hierarchical summary of the meeting is generated based on the meeting knowledge graph. The summary of the discussion content and the hierarchical summary of the meeting are then combined to obtain a summary of the meeting content.
[0070] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0071] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0072] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0073] Finally, it should be noted that if any software tools or components not belonging to this company appear in the embodiments of the application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for automatically generating meeting summaries, characterized in that, include: Acquire the meeting image, meeting text, and meeting audio of the target meeting; perform text conversion on the meeting audio to obtain the converted text; Cross-modal semantic alignment is performed on the meeting images, the meeting text, and the transformed text sequence to obtain a semantic network; Extract the structured information of the semantic network, and construct a meeting discussion flowchart based on the transformed text and the timestamp of the structured information. Generate a summary of the discussion content based on the meeting discussion flowchart. A meeting knowledge graph is constructed based on the semantic network, and a hierarchical summary of the meeting is generated based on the meeting knowledge graph. The summary of the discussion content and the hierarchical summary of the meeting are then combined to obtain a summary of the meeting content.
2. The method for automatically generating meeting summaries as described in claim 1, characterized in that, The step of converting the conference audio into text to obtain the converted text includes: Based on speech activity detection, non-silent segments in the conference speech are identified to obtain segmented audio block sequences; Voiceprint recognition is performed on the segmented audio block sequence to obtain voiceprint recognition data; Based on a preset voiceprint database, the speaker's identity is labeled in the segmented audio block sequence according to the voiceprint recognition data to obtain the identity-labeled audio block; A pre-trained speech recognition model is used to generate timestamped recognition text based on the identified audio blocks; The identified text is then corrected using context and a pre-defined domain terminology list to obtain the converted text.
3. The method for automatically generating meeting summaries as described in claim 1, characterized in that, The step of performing cross-modal semantic alignment on the conference image, the conference text, and the transformed text sequence to obtain a semantic network includes: OCR text extraction and target detection are performed on the meeting images to obtain image semantic units; The meeting text is parsed in a structured manner to obtain structured text; The image semantic units, the structured text, and the transformed text sequence are mapped to a unified time axis to obtain a time-synchronized multimodal stream; Cross-modal entity connections are performed on the time-synchronized multimodal stream to obtain an entity connection graph; Joint semantic encoding is performed on the entity connectivity graph and the time-synchronized multimodal stream to obtain a cross-modal semantic vector sequence; The semantic network is constructed based on the cross-modal semantic vector sequence and the entity connectivity graph.
4. The method for automatically generating meeting summaries as described in claim 1, characterized in that, The extraction of structured information from the semantic network includes: The semantic network is used to perform type recognition using a pre-trained graph neural network to obtain a set of classification marker points; Based on the relation edges of the semantic network, relation paths are extracted from the set of classification marker points to obtain structured relation paths; Based on the timestamps of the semantic network, the structured relation paths are bound by temporal constraints to obtain structured units; The multimodal evidence contained in the semantic network is connected to the structured unit to obtain structured information.
5. The method for automatically generating meeting summaries as described in claim 1, characterized in that, The construction of the meeting discussion flowchart based on the timestamps of the converted text and the structured information includes: The converted text and the structured information are subjected to time-series time node identification to obtain a time-series time node set; Perform time topology sorting on the set of time nodes to obtain the event sequence; The converted text is semantically recognized to obtain the semantics of the converted text; Identify the logical relationships of the transformed text semantics based on the event sequence; A meeting discussion flow graph is generated based on the logical relationships and the event sequence.
6. The method for automatically generating meeting summaries as described in claim 1, characterized in that, The step of constructing a conference knowledge graph based on the semantic network includes: Extract the core entities and original relation edges of the semantic network; The core entities are normalized to obtain a normalized entity set; The original relation edges are labeled with their types to obtain a set of valid relation edges; The attribute information is bound to the normalized entity set and the effective relation edge set to obtain the entity relation structure; The entity relationship structure is subjected to hierarchical constraints to obtain a hierarchical entity relationship structure; A meeting knowledge graph is generated based on the hierarchical entity relationship structure.
7. The method for automatically generating meeting summaries as described in claim 1, characterized in that, The step of generating a hierarchical summary of the meeting based on the meeting knowledge graph includes: The core path node sequence of the meeting knowledge graph is identified based on node importance calculation; Based on the preset user role configuration, a hierarchical summary template is injected into the core path node sequence to obtain a hierarchical summary skeleton. Obtain the node description of each node in the meeting knowledge graph, and perform semantic condensation and evidence binding based on the hierarchical summary skeleton and the node description to obtain a summary paragraph with evidence tags; A hierarchical summary of the meeting is generated based on the cross-level association of the summary paragraphs with evidence markers.
8. A device for automatically generating meeting summaries, characterized in that, include: The data conversion module is used to acquire the meeting images, meeting text, and meeting audio of the target meeting, and to convert the meeting audio into text to obtain converted text; The semantic alignment module is used to perform cross-modal semantic alignment on the conference images, the conference text, and the transformed text sequence to obtain a semantic network; The graph construction module is used to extract the structured information of the semantic network, construct a meeting discussion flowchart based on the transformed text and the timestamp of the structured information, and generate a summary of the discussion content based on the meeting discussion flowchart. The summary generation module is used to construct a conference knowledge graph based on the semantic network, generate a hierarchical summary of the conference based on the conference knowledge graph, and summarize the discussion content summary and the hierarchical summary of the conference to obtain a conference content summary.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the automatic generation method for meeting summaries as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the automatic generation method for meeting summaries as described in any one of claims 1 to 7.