Intelligent conference automated meeting record and summary generation method
Through directional microphone array and vocalprint clustering technology, combined with acoustic and semantic intent models, the problems of speaker identity identification and semantic paragraph division in multiple spokesperson meetings were solved, and structured meeting summary was generated, which improved the traceability and utilization efficiency of meeting content.
Patent Information
- Application Number
- CN202510796057.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-06-16
AI Technical Summary
It is difficult to accurately identify the spokesperson's identity in the multi-speaker scenarios, the semantic paragraph division is unreasonable, and the abstract content structure is unclear, resulting in information breakage or redundancy, and it is impossible to effectively support subsequent traceability and utilization.
Directed microphone arrays and K-means++ voiceprint clustering are used to separate speech streams in real time, and dynamic paragraph division is divided by combining acoustic and semantic meaning intensity models to generate a structured summary of agenda perception.
It realizes accurate identification of spokesperson identity and dynamic boundary judgment of semantic paragraphs, improves the traceability and structured utilization of conference content, and supports decision-making traceability and information retrieval.
Smart Images

Figure CN120340497B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method for generating automatic meeting records and summaries for intelligent conferences. Background Art
[0002] With the widespread use of remote work, remote collaboration, and intelligent conferencing systems, digital recording and structured management of meeting content have become key to improving organizational decision-making efficiency and knowledge asset utilization. Traditional meeting recording methods rely heavily on manual note-taking or one-way text transcription based on voice recognition. While these methods can retain some content, they present the following significant challenges:
[0003] First, it is difficult to distinguish the voice streams of multiple speakers. In scenarios where multiple people are speaking continuously or in free discussions, the system often cannot accurately identify the speakers, resulting in a lack of semantic attribution in the transcribed text, which affects subsequent review and accountability.
[0004] Second, semantic segmentation is crude. Existing methods are mostly based on fixed time windows or fixed grammatical rules. This makes it difficult to cope with nonlinear semantic structures such as jumps in intent and significant fluctuations in tone in actual meetings, resulting in information fragmentation or redundant repetition.
[0005] Third, the summary content structure is unclear and lacks agenda alignment. Traditional compressed summaries ignore the context of the meeting process and role differences, and are unable to distinguish key elements such as task allocation, focus of controversy, and decision-making results. This results in low summary value density, poor readability, and difficulty in subsequent use. Summary of the Invention
[0006] The present invention provides an intelligent conference automated record and summary generation method. Aiming at real multi-speaker conference scenarios, the method integrates identity separation, semantic recognition, summary extraction and structure mapping capabilities to improve the organization, traceability and structured utilization of conference content.
[0007] The method for automatically generating meeting minutes and summaries for intelligent meetings includes the following steps:
[0008] S1: Capture multi-speaker speech streams in real time through a directional microphone array, generate a raw text stream with time series tags based on real-time voiceprint clustering, and simultaneously extract intent intensity parameters from the speech stream;
[0009] S2: performing intent-driven dynamic segmentation processing on the original text stream, including:
[0010] Semantic paragraphs are divided according to the mutation points of the intention intensity parameter;
[0011] Merge paragraphs and calibrate boundaries based on the semantic conflict between adjacent paragraphs;
[0012] S3: Perform agenda-aware summary block generation on the segmented text, including extracting decision statements that meet the semantic density threshold as summary core blocks;
[0013] S4: Dynamically map the summary core blocks to the preset agenda template, and assemble the summary core blocks into a structured summary document according to the hierarchical structure of the agenda template.
[0014] Optionally, the S1 specifically includes:
[0015] S11, controls the directional parameters of the microphone array through the beamforming algorithm to lock the voice signal source of the current speaker;
[0016] S12, performs real-time voiceprint clustering on the continuous speech stream, uses the K-means++ algorithm to divide the voiceprint feature vector into speaker subsets, and binds each speaker subset to a unique identity code;
[0017] S13, embedding a time stamp engine in the speech transcription process, inserting a timestamp delimiter when a speaker switch is detected, and generating an original text stream with an identification code and time stamps;
[0018] S14, synchronously extract the acoustic intent features and semantic intent features of the speech stream, input the acoustic intent features and semantic intent features into the pre-trained intention strength calculation model, and output the intention strength parameter curve that changes with time.
[0019] Optionally, the acoustic intent features include fundamental frequency change rate and energy mutation gradient; the semantic intent features include logical connective density and sentiment polarity intensity.
[0020] Optionally, the intention intensity calculation model includes a two-dimensional feature coding layer, a feature fusion layer, a time series modeling layer and an output layer. The two-dimensional feature coding layer includes an acoustic coding module and a semantic coding module. The feature fusion layer uses a bimodal attention mechanism to calculate attention weights for acoustic intent features and semantic intent features respectively, and weightedly fuses the two parts of attention to obtain a fused feature representation. The time series modeling layer inputs a bidirectional LSTM model based on the fused feature representation to capture the time context features. The output layer outputs the intention intensity value at each moment, and then draws the intention intensity parameter curve.
[0021] Optionally, the S2 specifically includes:
[0022] S21, semantic paragraph segmentation: perform differential calculation on the intent intensity parameter curve. When the intensity change rate of adjacent time windows exceeds a first threshold, it is marked as a mutation boundary point. Extract text segments within a preset time length before and after the mutation boundary point, calculate the semantic integrity score of the segment through a bidirectional LSTM model, and select segments with scores above the second threshold as candidate paragraph boundaries.
[0023] S22, calculating the semantic conflict degree of adjacent candidate paragraphs and generating a comprehensive conflict value;
[0024] S23, paragraph merging and calibration: When the comprehensive conflict value is lower than the third threshold, adjacent paragraphs are merged and the boundaries are recalculated, including expanding to both sides based on the mutation boundary point until a logical connective or responsible subject change event is detected;
[0025] Insert enhanced time markers at the beginning of the merged segments to record the original segmentation points and the merge decision chain.
[0026] Optionally, the generation of the comprehensive conflict value in S22 includes:
[0027] S221, extracting the key decision word set and sentiment polarity vector of each paragraph;
[0028] S222, integrating the keyword overlap rate and the sentiment polarity vector angle to generate a comprehensive conflict value.
[0029] Optionally, after the segment merging and calibration in S23, the step further includes inserting a reinforced time marker at the beginning of the merged segment to record the original segmentation point and the merging decision chain.
[0030] Optionally, the S3 includes constructing a decision pattern map and defining candidate text block extraction rules; calculating the semantic density value of the candidate text block: using a multi-head attention mechanism to calculate the relevance weight of the sentence and the agenda topic, integrating the sentence length penalty factor and the decision hierarchy coefficient to generate a semantic density score, selecting candidate text blocks whose semantic density scores exceed the dynamic semantic density threshold and marking them as summary core blocks, and injecting the timestamp chain of the source paragraph and the speaker identity label.
[0031] Optionally, the candidate text block extraction rules include:
[0032] Decision statement block: extract triples of action verbs, responsible parties, and time nodes;
[0033] Controversial focus block: captures consecutive dialogue paragraphs containing opposing logical connectives and with high sentiment polarity differences;
[0034] Task distribution block: Identify compound sentences that include resource allocation statements and task acceptance criteria.
[0035] Optionally, the S4 specifically includes:
[0036] S41, parsing the tree-like hierarchical structure of the preset agenda template and extracting the semantic slot constraints under each node;
[0037] S42, matching the summary core block with the semantic slot, including:
[0038] Task distribution block matching: If the summary block type is "task distribution" and there is a responsible subject field, the task distribution block is matched to the execution node based on the consistency of the responsible subject;
[0039] Decision statement block mapping: If the summary block is "decision statement" and has a timestamp, the decision statement block is mapped to the corresponding agenda stage node according to the decision timestamp range;
[0040] Inserting a derived node into the disputed block: If the summary block type is "disputed", a derived node is created based on its original node, a conflict marker is added to the disputed block, and the derived node is inserted;
[0041] S43, generate summary documents according to the template hierarchy: reorganize the successfully mapped summary core blocks into a three-level structure of “timeline-issue-decision”.
[0042] Beneficial effects of the present invention:
[0043] The present invention uses array microphone beamforming and K-means++ voiceprint clustering algorithm to achieve real-time separation and identity binding of multi-speaker voice streams, and automatically inserts high-precision timestamps based on speaker switching events to achieve accurate synchronization between original speech and transcribed text. Compared with traditional linear transcription methods based on speech recognition, it significantly improves the mapping accuracy between speech content and participant identity, providing a solid foundation for subsequent semantic segmentation and decision tracing.
[0044] This invention proposes an "acoustic + semantic" fusion intention strength calculation model, generates an intention strength curve through multimodal features such as fundamental frequency change rate and emotional polarity, and combines mutation point detection with semantic integrity scoring to realize dynamic boundary judgment of semantic paragraphs; and constructs a conflict value model through keyword overlap rate and emotional vector angle to judge and merge paragraph conflicts, significantly improving the semantic aggregation ability and coherence of automatic summarization, and avoiding the problems of fragmented output or semantic fragmentation.
[0045] Based on the preset agenda template, the present invention combines the multi-head attention mechanism with the semantic density scoring model to extract the three core blocks of the summary: "decision / dispute / task", and dynamically reorganizes them into interactive summary documents according to the three-level structure of "timeline-topic-block type". It supports source tracing, conflict visualization and node insertion, and constructs a structured content presentation form that can be embedded in the meeting management system, enhancing the practical value of meeting summaries in collaborative management, review and information retrieval scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 Schematic diagram of the generation method flow in an embodiment of the present invention;
[0048] Figure 2 Schematic diagram of the intention intensity calculation model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It is also noted that, to provide a more detailed description, the following embodiments are best and preferred embodiments, and those skilled in the art may employ alternative methods for implementing certain known technologies. Furthermore, the accompanying drawings are intended only to provide a more detailed description of the embodiments and are not intended to limit the present invention.
[0050] It should be noted that references in the specification to "one embodiment," "an embodiment," "exemplary embodiments," "some embodiments," etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment will include such specific features, structures, or characteristics. Furthermore, when specific features, structures, or characteristics are described in conjunction with an embodiment, it is within the knowledge of persons skilled in the relevant art to implement such features, structures, or characteristics in conjunction with other embodiments (whether or not explicitly described).
[0051] In general, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending at least in part on the context, allow for the presence of other factors that are not necessarily explicitly described.
[0052] like Figure 1-Figure 2 As shown, the intelligent conference automatic meeting record and summary generation method includes the following steps:
[0053] S1: Capture multi-speaker speech streams in real time through a directional microphone array, generate a raw text stream with time series tags based on real-time voiceprint clustering, and simultaneously extract intent intensity parameters from the speech stream;
[0054] S2: Performs intent-driven dynamic segmentation on the original text stream, including:
[0055] Semantic paragraphs are divided according to the mutation points of the intention intensity parameter;
[0056] Merge paragraphs and calibrate boundaries based on the semantic conflict between adjacent paragraphs;
[0057] S3: Perform agenda-aware summary block generation on the segmented text, including extracting decision statements that meet the semantic density threshold as summary core blocks;
[0058] S4: Dynamically map the summary core blocks to the preset agenda template, and assemble the summary core blocks into a structured summary document according to the hierarchical structure of the agenda template.
[0059] S1 specifically includes:
[0060] S11, through the beamforming algorithm to control the directional parameters of the microphone array, lock the current speaker's voice signal source, suppress the voice interference in the non-target direction, and achieve real-time voice signal separation. The beamforming algorithm uses the target direction angle As input, dynamically adjust the array weight , forming a narrowband directional response, expressed as:
[0061] ;
[0062] in, For the The original input signal of the microphone, is the adaptive weight corresponding to the microphone channel, is the total number of microphone channels, It is the directional output signal after formation.
[0063] S12, performs real-time voiceprint clustering on the continuous speech stream, and uses the K-means++ clustering algorithm to cluster the voiceprint feature vectors. Group and add voiceprint similarity constraints , minimize the objective function, expressed as:
[0064] ;
[0065] in, Represents the voiceprint feature vector, extracted from Mel-frequency cepstral coefficients (MFCC) and transition frame dynamic features, It is Voiceprint clusters, Represents a cluster The center vector of Indicates the speaker similarity distance within the voiceprint cluster, is the similarity penalty coefficient.
[0066] Each speaker leaves a unique "voiceprint" feature when speaking, and the Mel-frequency cepstral coefficient (MFCC) acoustic features are extracted from the speech to form a high-dimensional feature vector After these feature vectors are continuously extracted in time, multiple point sets that may come from different speakers are formed. In order to distinguish who these points belong to, a K-means++ clustering algorithm is used to cluster them. K-means++ is superior to traditional K-means in that the initial center selection is more reasonable, which helps to converge faster. The addition of similarity constraints: The item ensures that the voiceprints in the same cluster are as similar as possible, suppresses misclassification, minimizes the objective function, and for each cluster center , all voiceprint vectors close to it are grouped into the same cluster , until the overall error is minimized. After clustering is completed, several subsets (clusters) can be obtained. Each subset represents a speaker's speech segment during the entire meeting. For each clustered voiceprint subset, a unique identity code IDk is automatically bound. The binding process is as follows:
[0067] Set up a global identity code pool (e.g. ID001, ID002, etc.);
[0068] Each time a cluster Ck is completed, an unused ID is assigned from the pool as an identifier. This ID will be used as the "speaker label" of each speech segment in the original text stream and is also bound to the timestamp to form a complete data structure.
[0069] S13, embeds a time stamping engine in the speech transcription process. When a speaker identity switch is detected, a delimiter with a timestamp is automatically inserted, which is expressed as:
[0070] ;
[0071] in, Indicates the The unique identification code of the speaker, 、 Indicates the start and end timestamps of the speech segment. Represents the complete set of speaker-timing tags.
[0072] S14, synchronously extracting intent-related features of the voice stream, including:
[0073] S141, Acoustic Intent Features:
[0074] Fundamental frequency change rate: ;
[0075] Energy mutation gradient: ;
[0076] in, Indicates time The fundamental frequency, Indicates the fundamental frequency change rate, that is, the fundamental frequency change between adjacent frames, represents the time interval between frames, Indicates time The speech energy is the frame energy or logarithmic energy, Represents the energy mutation gradient, that is, the absolute value of the energy change;
[0077] S142, semantic intent features:
[0078] Logical connective density: ;
[0079] Sentiment polarity strength (using BERT + sentiment classification head):
[0080] ;
[0081] in, Indicates time The number of occurrences of logical conjunctions in the text segment, Indicates time The total number of words in the text segment, It represents the density of logical connectives, that is, the frequency of logical words contained in each unit word number. Indicates that the text is at time Positive emotion confidence (such as positive, supportive, and approving), Indicates that the text is at time Negative emotion confidence (such as denial, opposition, anger), Indicates the intensity of emotional polarity, measuring the magnitude of the difference between positive and negative emotions.
[0082] Input the above multi-dimensional features into the pre-trained intention intensity calculation model , output the intention intensity parameter curve changing with time: ,in, , represents the fused intent feature vector, Represents a time series network based on the attention mechanism, which is used to guide the dynamic segmentation process of subsequent semantic paragraphs.
[0083] Intention intensity calculation model The details are as follows:
[0084] 1. Model structure: Multimodal time series network based on attention mechanism
[0085] Adopting the structure of multimodal fusion + double-layer attention mechanism + temporal modeling, such as Figure 2 As shown:
[0086] 2. Input construction: At each time step , construct a fusion input vector , consists of acoustic intent features and semantic intent features: ;
[0087] Multiple time steps constitute the input sequence: .
[0088] 3. Feature encoding layer (independent encoding): The acoustic features (the first two dimensions) and semantic features (the last two dimensions) are input into two MLP or GRU encoding modules respectively:
[0089] Acoustic Coding Module: ;
[0090] Semantic encoding module: .
[0091] Indicates time The acoustic feature encoding vector of Indicates time Semantic feature encoding vector of ;
[0092] 4. Feature fusion layer (multimodal fusion with attention): Uses a bimodal attention mechanism, that is, calculates attention weights for acoustic features and semantic features separately:
[0093] ;
[0094] in, is the acoustic attention weight matrix, used to map To the attention score space, Semantic attention weight matrix, used to map To the attention score space, Indicates time The acoustic attention coefficient reflects the importance of acoustic features in fusion. For the moment The semantic attention coefficient reflects the importance of semantic features in fusion. represents the time step index used for normalization (i.e., the summation variable of the softmax denominator), For the moment The multimodal fusion vector is composed of the acoustic and semantic feature vectors weighted by attention;
[0095] Weighted fusion of the two attention parts: .
[0096] 5. Time series modeling layer: the fused representation sequence Enter Bi-LSTM (bidirectional LSTM model) to capture temporal context features: .
[0097] 6. Output layer (intent strength prediction): Finally, the context state of each time step is Input a linear layer + Sigmoid activation function and output the current intention strength value: ,in, , indicating the current time The closer to 1, the stronger the speaker's intention. are the parameters of the output layer, is the Sigmoid function.
[0098] 7. Model training and output:
[0099] Training objective: Use the manually labeled "intent intensity" as the label and perform regression training using mean square error.
[0100] The output is an intention intensity parameter curve with a time resolution equal to the number of audio frames. , used to drive the dynamic division of subsequent semantic paragraphs.
[0101] S2 specifically includes:
[0102] S21, semantic paragraph segmentation:
[0103] Mutation boundary point detection: intention intensity parameter curve Perform first-order differences to compute the rate of change of intensity over successive time windows: , when satisfied When, the moment Marked as mutation boundary points, recorded as a set: ,in, is the first-order rate of change of intention intensity, The first threshold is used to initially screen mutation points. Setting it to 0.25-0.35 can effectively capture emotional fluctuations or topic transitions and avoid false triggering of boundaries due to small disturbances. is the sampling time window interval, is the set of mutation boundary points.
[0104] Candidate paragraph boundary screening: for each mutation boundary point , extract the duration before and after Text snippets within , and input it into the bidirectional LSTM model to calculate its semantic completeness score:
[0105] , if: , then it is retained as a candidate paragraph boundary, where is the window length for context extraction (adjustable), Mutation point The corresponding context fragment, is the semantic completeness score, The second threshold is used to filter semantically complete segments. The value range is 0.70-0.80. A value above 70% can ensure that the filtered text segments have a good semantic closure and are suitable as paragraph boundaries.
[0106] The word "_" represents a scored bidirectional LSTM model. For text segments before and after the mutation point, the model is fed into the bidirectional LSTM model for semantic encoding. The model extracts the contextual dependencies of word sequences from both the forward and backward directions to capture whether the segment has a complete logical loop (such as a complete sentence structure, natural semantic convergence, and a valid subject-verb-object relationship). Finally, a fully connected + sigmoid scoring layer is added to the encoding result, outputting a score between 0 and 1 to indicate whether the paragraph is semantically complete. A higher score indicates that the paragraph is more likely to be the boundary of a natural semantic unit.
[0107] S22, paragraph merging and boundary calibration: for all adjacent candidate paragraph pairs Perform conflict degree calculation:
[0108] S221, extract paragraph pairs:
[0109] Keyword collection: ;
[0110] Sentiment polarity vector: ;
[0111] S222, define conflict indicators:
[0112] Indicator 1. Keyword overlap rate : ;
[0113] Indicator 2. Sentiment vector angle (Differences in viewpoints): ;
[0114] Combining the above indicators, the comprehensive conflict value is calculated: , when: , then the paragraphs are considered to have no conflicts and a merge operation is performed.
[0115] in, is the keyword overlap rate, is the emotional vector angle (5 radians), is the topological difference of the logical graph, is the conflict fusion weight coefficient (satisfying ), The third threshold, the comprehensive conflict tolerance, is used to control the conflict tolerance level of paragraph merging. Its value changes with the progress of the meeting:
[0116] ;
[0117] In the initial discussion stage, high differences (such as clashes of opinions) should be tolerated, and in the later stage, consistency and logical closure should be emphasized, and the threshold should be gradually lowered.
[0118] S23, boundary calibration and time stamp update: When the paragraphs are merged, the boundary points are mutated. As a benchmark, expand forward / backward until encountering: logical conjunctions ("but", "therefore", etc.) or detecting a change in the responsible party (such as "someone proposed", "the financial group responded", "the leader pointed out", "the group responded", etc.), and determine the new paragraph boundary point. , insert the enhanced time mark at its location, expressed as:
[0119] ,in, is the start time of the segment after merging and calibration, It is the original mutation boundary point, and MergeID represents the unique number of the current merge operation, which is used to build the decision chain traceability record.
[0120] S3 specifically includes:
[0121] S31, build a decision-making model map and set three types of core extraction rules:
[0122] Decision statement block: Matches the triple structure of the following three types of elements that appear simultaneously in the sentence:
[0123] action verbs (e.g., "approved," "rejected");
[0124] Responsible party (e.g., "person in charge", "meeting group");
[0125] Time node (such as "next Monday", "before a certain date in a certain month");
[0126] Controversial focus block: identifies consecutive sentences containing opposing conjunctions (such as "but" and "on the contrary") and with a sentiment polarity difference greater than 0.6;
[0127] Task distribution block: matches compound sentences containing resource allocation entities (such as "budget" and "manpower") and task acceptance criteria (such as "must be completed within ××").
[0128] S32, comprehensive density score calculation: for each candidate text block , calculate its semantic density score :
[0129] ;
[0130] in, It is the sentence length of the text block that is automatically counted. The shorter the text length, the higher the score. is the corresponding spokesperson decision-making level coefficient, ordinary 0.2, middle 0.5, high-level 0.8, is the length penalty coefficient, the value is 0.2, is the decision-making layer weight coefficient, which takes a value of 0.3. For the Attention heads for text blocks The relevance weight to the current agenda topic comes from the attention weight in the BERT model, ranging from [0, 1], with an average of 0.1 to 0.3. is the number of attention heads, as follows:
[0131] 1. (a sentence or paragraph) is encoded into several tokens and fed into the Transformer encoder to encode the title or brief description text of the current agenda node (such as "task assignment", "decision confirmation") into a query sentence vector. , calculated as follows:
[0132] 2. Extracting the original representation from the Transformer model: getting the text block The Key and Value vectors of each token in the ,get the agenda topic vector as the global Query vector;
[0133] 3. Standard attention mechanism calculates relevance: for each token (in a text block), calculate:
[0134] After normalizing the attention weights of all tokens, take the mean or maximum value as the overall attention score of the text block .
[0135] S33, if the following conditions are met: , then Selected as a summary core block with the following tag information: ,in, is the timestamp chain of the original paragraph, This is the speaker identification number. Type label for the summary block (decision / dispute / task), It is a dynamic semantic density threshold that automatically adjusts with the agenda node. Specifically:
[0136] Opening Notes: Values range from 0.65 to 0.70 (strict screening);
[0137] The value of the previous round summary review node is 0.60;
[0138] The value of the free discussion node is 0.50;
[0139] The value of the special report node is 0.45;
[0140] The decision suggestion formation node takes a value of 0.40;
[0141] The value of the task division and execution instruction node is 0.40 (loose screening, retaining details);
[0142] The value of the voting and result confirmation node is 0.45.
[0143] S4 specifically includes:
[0144] S41, parsing the tree-like hierarchical structure of the preset agenda template and extracting the semantic slot constraints under each node;
[0145] S42, matching the summary core block with the semantic slot, including:
[0146] Task distribution block matching: If the summary block type is "task distribution" and there is a responsible subject field, the task distribution block is matched to the execution node based on the responsible subject consistency;
[0147] Decision statement block mapping: If the summary block is "decision statement" and has a timestamp, the decision statement block is mapped to the corresponding agenda stage node according to the decision timestamp range;
[0148] Inserting a derived node into the disputed block: If the summary block type is "disputed", a derived node is created based on its original node, a conflict marker is added to the disputed block, and the derived node is inserted;
[0149] S43, generate summary documents according to the template hierarchy: reorganize the successfully mapped summary core blocks into a three-level structure of "timeline-topic-decision"; attach a source tag to each summary core block, and jump to the original meeting segment timestamp when clicked.
[0150] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.
[0151] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. Intelligent conference automated meeting record and summary generation method, characterized by: The following steps are involved: S1: Capture multi-speaker speech streams in real time through a directional microphone array, generate a raw text stream with time series tags based on real-time voiceprint clustering, and simultaneously extract intent intensity parameters from the speech stream; S2: performing intent-driven dynamic segmentation processing on the original text stream, including: Semantic paragraphs are divided according to the mutation points of the intention intensity parameter; Merge paragraphs and calibrate boundaries based on the semantic conflict between adjacent paragraphs; S3: Perform agenda-aware summary block generation on the segmented text, including extracting candidate text that meets the semantic density threshold as summary core blocks. This includes constructing a decision pattern map and defining candidate text block extraction rules. The semantic density values of the candidate text blocks are calculated using a multi-head attention mechanism to calculate the relevance weight between the sentence and the agenda topic. The sentence length penalty factor and decision hierarchy coefficient are integrated to generate a semantic density score. Candidate text blocks with semantic density scores exceeding the dynamic semantic density threshold are selected and marked as summary core blocks. The timestamp chain of the source paragraph and the speaker identity label are then injected. The candidate text block extraction rules include: Decision statement block: extract triples of action verbs, responsible parties, and time nodes; Controversial focus block: captures consecutive dialogue paragraphs containing opposing logical connectives and with high sentiment polarity differences; Task distribution block: Identify compound sentences including resource allocation statements and task acceptance criteria; S4: Dynamically map the summary core blocks to the preset agenda template, and assemble the summary core blocks into a structured summary document according to the hierarchical structure of the agenda template.
2. The method for generating automatic meeting minutes and summaries for intelligent conferences according to claim 1, characterized in that: Said S1 specifically includes: S11, controls the directional parameters of the microphone array through the beamforming algorithm to lock the voice signal source of the current speaker; S12, performs real-time voiceprint clustering on the continuous speech stream, uses the K-means++ algorithm to divide the voiceprint feature vector into speaker subsets, and binds each speaker subset to a unique identity code; S13, embedding a time stamp engine in the speech transcription process, inserting a timestamp delimiter when a speaker switch is detected, and generating an original text stream with an identification code and time stamps; S14, synchronously extract the acoustic intent features and semantic intent features of the speech stream, input the acoustic intent features and semantic intent features into the pre-trained intention strength calculation model, and output the intention strength parameter curve that changes with time.
3. The method for generating automatic meeting minutes and summaries for intelligent conferences according to claim 2, characterized in that: The acoustic intent features include fundamental frequency change rate and energy mutation gradient; the semantic intent features include logical connective density and sentiment polarity intensity.
4. The method for generating automatic meeting minutes and summaries for intelligent conferences according to claim 3, characterized in that: The intention intensity calculation model includes a two-dimensional feature coding layer, a feature fusion layer, a time series modeling layer and an output layer. The two-dimensional feature coding layer includes an acoustic coding module and a semantic coding module. The feature fusion layer uses a bimodal attention mechanism to calculate the attention weights for the acoustic intent features and the semantic intent features respectively, and weightedly fuses the two parts of attention to obtain a fused feature representation. The time series modeling layer inputs a bidirectional LSTM model based on the fused feature representation to capture the time context features. The output layer outputs the intention intensity value at each moment, and then draws the intention intensity parameter curve.
5. The method for generating automatic meeting minutes and summaries for intelligent conferences according to claim 4, characterized in that: The S2 specifically includes: S21, semantic paragraph segmentation: perform differential calculation on the intent intensity parameter curve. When the intensity change rate of adjacent time windows exceeds a first threshold, it is marked as a mutation boundary point. Extract text segments within a preset time length before and after the mutation boundary point, calculate the semantic integrity score of the segment through a bidirectional LSTM model, and select segments with scores above the second threshold as candidate paragraph boundaries. S22, calculating the semantic conflict degree of adjacent candidate paragraphs and generating a comprehensive conflict value; S23, paragraph merging and calibration: When the comprehensive conflict value is lower than the third threshold, adjacent paragraphs are merged and the boundaries are recalculated, including expanding to both sides based on the mutation boundary point until a logical connective or responsible subject change event is detected; Insert enhanced time markers at the beginning of the merged segments to record the original segmentation points and the merge decision chain.
6. The method for generating automatic meeting minutes and summaries for intelligent conferences according to claim 5, characterized in that: The generation of the comprehensive conflict value in S22 includes: S221, extracting the key decision word set and sentiment polarity vector of each paragraph; S222, integrating the keyword overlap rate and the sentiment polarity vector angle to generate a comprehensive conflict value.
7. The method for generating automatic meeting minutes and summaries for intelligent conferences according to claim 6, characterized in that: After the segment merging and calibration in S23, the process further includes inserting a reinforced time mark at the beginning of the merged segment to record the original segmentation point and the merge decision chain.
8. The method for generating automatic meeting minutes and summaries for intelligent conferences according to claim 1, characterized in that: The S4 specifically includes: S41, parsing the tree-like hierarchical structure of the preset agenda template and extracting the semantic slot constraints under each node; S42, matching the summary core block with the semantic slot, including: Task distribution block matching: If the summary block type is task distribution and there is a responsible subject field, the task distribution block is matched to the execution node based on the responsible subject consistency; Decision statement block mapping: If the summary block is a decision statement and has a timestamp, the decision statement block is mapped to the corresponding agenda stage node according to the decision timestamp range; Inserting a derived node into the dispute focus block: If the summary block type is dispute focus, a derived node is created based on its original node, a conflict marker is added to the dispute focus block, and the derived node is inserted; S43, generate summary documents according to the template hierarchy: reorganize the successfully mapped summary core blocks according to timeline-topic-decision.
Citation Information
Patent Citations
Intelligent conference summary generation method and system
CN110717031A
Domain speech recognition method and system based on RAG
CN119296516A