Intelligent conference summary generation method and device, equipment and storage medium

By combining a multi-channel microphone array with a large language model, the problems of low efficiency and inconsistent formatting in meeting minutes generation have been solved, achieving efficient, accurate, and natural minutes generation, thus improving the quality of decision-making and service levels in the financial and medical fields.

CN121884818APending Publication Date: 2026-04-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing meeting minutes generation technologies suffer from problems such as low efficiency, easy omission of key information, redundant content, and inconsistent formats, especially in the financial and medical fields, where they pose compliance risks and information distortion.

Method used

Audio preprocessing is performed using a multi-channel microphone array. A speech recognition model is used to generate structured text with timestamps. A large language model is then combined to perform semantic analysis and style transfer, generating structured summary text that conforms to the preset text style, and humanized features are injected.

Benefits of technology

It enables efficient, accurate, natural, and operable generation of meeting minutes, improves the traceability of information and the professionalism and readability of the text, ensures the consistency and standardization of output, and enhances the accuracy of decision-making and action information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884818A_ABST
    Figure CN121884818A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent conference summary generation method and device, equipment and a storage medium, and relates to the technical field of natural language processing. According to the method provided by the invention, clear capture and accurate identification of audios are realized by using the multi-channel microphone array; and inputting the audio with the speaker identifier into the speech recognition model, and converting the audio into a structured initial summary text with a timestamp. Performing semantic analysis on the initial summary text through a large language model, and extracting key text information; a rule engine and a large language model are combined for combined judgment, so that the information accuracy is ensured; the initial summary text is converted into a first summary text in a preset text style, so that the text specialty is improved; humanized feature injection enhances the language naturalness of the second summary text; and the second summary text is converted into the structured target summary text, so that the conference summary better meets the requirements of professional scenes such as financial risk control conferences and medical consultation conferences, and the accuracy and availability of intelligent generation of the conference summary are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular to an intelligent method, apparatus, device and storage medium for generating meeting minutes. Background Technology

[0002] With the deepening of digital transformation, meeting minutes, as a core carrier for corporate knowledge accumulation and decision tracking, directly impact organizational operational efficiency in terms of generation efficiency and quality. Traditional meeting minutes generation heavily relies on manual recording, resulting in low efficiency, easy omission of key information, redundant content, and inconsistent formats.

[0003] In the financial sector, scenarios such as investment decision-making meetings and risk control meetings involve a large number of professional terms, compliance requirements, and regulatory-sensitive information. Manual recording is prone to omitting risk warning clauses or miswriting fund amounts, leading to serious compliance risks. In the medical field, multidisciplinary team (MDT) consultations and discussions of difficult cases involve patient privacy data, ethical evaluation of treatment plans, and determination of medical liability. Manually compiling these is not only time-consuming and labor-intensive, but may also lead to medical disputes due to recording errors.

[0004] Although existing technologies have incorporated techniques such as Automatic Speech Recognition (ASR), keyword extraction, and large model summarization, they generally suffer from the following drawbacks: 1. Fragmented process: Each module operates in isolation, lacking collaborative design, and unable to form a complete closed loop of "audio → text → content extraction → formalized output"; 2. Information distortion: Direct summarization after ASR transcription ignores the informality and ambiguity of spoken expression, leading to the weakening or misinterpretation of key decisions; 3. Lack of "human touch": AI-generated text is too formal and logically rigorous, lacking the rhythm and leaps of thought of natural language, making it difficult to use for high-level reports; 4. Inconsistent format: The output document formats are diverse, making it difficult to integrate with the enterprise OA system or email workflow.

[0005] Therefore, the financial and medical fields urgently need an intelligent system that can integrate multimodal processing, deep semantic understanding, style transfer, and structured generation to solve the technical problems of existing meeting minutes generation technologies, such as low efficiency, easy omission of key information, content redundancy, and inconsistent formats. This system would enable efficient, accurate, natural, and operable generation of meeting minutes, improve the accuracy of intelligent meeting minutes generation, and thus enhance the quality of financial decision-making and the level of medical services. Summary of the Invention

[0006] This application provides an intelligent meeting minutes generation method, apparatus, device, and storage medium, aiming to solve the technical problems of low efficiency, easy omission of key information, redundant content, and inconsistent formats in existing meeting minutes generation technologies, so as to achieve efficient, accurate, natural, and operable generation of meeting minutes and improve the accuracy of intelligent meeting minutes generation.

[0007] In a first aspect, this application provides an intelligent meeting minutes generation method, which includes the following steps: preprocessing multi-channel audio collected by a multi-channel microphone array to generate an audio stream with speaker identifiers; inputting the audio stream with speaker identifiers into a speech recognition model, and converting the audio stream into structured text with timestamps through the speech recognition model to obtain initial minutes text for at least one speaker; performing semantic analysis on the initial minutes text corresponding to each speaker through a large language model to extract key text information; based on the key text information, jointly judging the decision content and task elements in the initial minutes text through a rule engine and a large language model to generate decision and action information; performing style transfer on the initial minutes text based on the key text information and the decision and action information to convert the initial minutes text into a first minutes text with a preset text style; injecting humanized features into the first minutes text to generate a second minutes text; and filling the second minutes text into a target meeting minutes template to generate a structured target minutes text.

[0008] Secondly, this application also provides an intelligent meeting minutes generation device, comprising: an audio preprocessing module for preprocessing multi-channel audio collected by a multi-channel microphone array to generate an audio stream with speaker identifiers; a speech conversion module for inputting the audio stream with speaker identifiers into a speech recognition model, and converting the audio stream into structured text with timestamps through the speech recognition model to obtain initial minutes text for at least one speaker; a text semantic parsing module for performing semantic analysis on the initial minutes text corresponding to each speaker through a large language model to extract key text information; and a joint judgment module. The system comprises the following modules: a segmentation module, which uses a rule engine and a large language model to jointly judge the decision content and task elements in the initial minutes text based on the key text information, and generates decision and action information; a style transfer module, which performs style transfer on the initial minutes text based on the key text information and the decision and action information, converting the initial minutes text into a first minutes text with a preset text style; a personalized feature injection module, which performs humanized feature injection processing on the first minutes text to generate a second minutes text; and a structured text generation module, which fills the second minutes text into the target meeting minutes template to generate a structured target minutes text.

[0009] Thirdly, this application also provides a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the intelligent meeting minutes generation method described above.

[0010] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the intelligent meeting minutes generation method described above.

[0011] This application provides an intelligent meeting minutes generation method, apparatus, computer equipment, and storage medium. The method utilizes a multi-channel microphone array for audio preprocessing, ensuring that the voice of each speaker in the audio stream is clearly captured and accurately identified. The audio with speaker identification is input into a speech recognition model and converted into structured text with timestamps, preserving the original speaking order and time information, thus improving information traceability and processing efficiency. Deep semantic analysis of the initial minutes text is performed using a large language model to extract key textual information, enabling the minutes to accurately reflect the core issues and decision points of the meeting. A rule engine and the large language model are combined to jointly judge decision content and task elements, further ensuring the accuracy and executability of decision and action information. Style transfer transforms the initial minutes text into a first minutes text conforming to a preset text style, making the generated text more professional and readable. Humanized feature injection makes the second minutes text more natural and closer to human expression habits, enhancing the naturalness and credibility of the text. The second minutes text is then filled into a target meeting minutes template to generate a structured target minutes text, ensuring the consistency and standardization of the output. This application achieves natural and accurate generation of meeting minutes text through precise sound processing, accurate speech recognition, in-depth semantic analysis, rigorous decision-making, professional text stylization, and humanized expression, thus comprehensively improving the text quality and usability of meeting minutes. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram of an application environment for an intelligent meeting minutes generation method according to an embodiment of the present invention; Figure 2 A flowchart illustrating an embodiment of an intelligent meeting minutes generation method provided in this application; Figure 3 This is a schematic diagram of the structure of an embodiment of an intelligent meeting minutes generation device provided in this application. Figure 4 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.

[0014] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0016] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the described order. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0017] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] The intelligent meeting minutes generation method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. When the server receives a meeting minutes generation request from the client, it can preprocess the multi-channel audio collected by a multi-channel microphone array to generate an audio stream with speaker identifiers. The audio stream with speaker identifiers is then input into a speech recognition model, which converts it into structured text with timestamps, obtaining initial minutes text for at least one speaker. A large language model is used to perform semantic analysis on the initial minutes text corresponding to each speaker, extracting key text information. Based on this key text information, a rule engine and the large language model jointly determine the decision content and task elements in the initial minutes text, generating decision and action information. Based on the key text information and the decision and action information, style transfer is performed on the initial minutes text, converting it into a first minutes text with a preset text style. Humanized feature injection is performed on the first minutes text to generate a second minutes text. The second minutes text is then filled into a target meeting minutes template to generate a structured target minutes text.

[0020] In this invention, to address the technical problems of low efficiency, easy omission of key information, redundant content, and inconsistent formats in the generation of meeting minutes in the financial and medical fields, a multi-channel microphone array can be used for audio preprocessing to ensure that the voice of each speaker in the audio stream can be clearly captured and accurately identified. The audio with speaker identification is then input into a speech recognition model and converted into structured text with timestamps, preserving the original order and time information of the speeches, thereby improving the traceability and processing efficiency of the information. Deep semantic analysis of the initial minutes text using a large language model extracts key textual information, ensuring the minutes accurately reflect the core issues and decision points of the meeting. A combination of a rule engine and the large language model jointly assesses decision content and task elements, further ensuring the accuracy and feasibility of decision and action information. Style transfer transforms the initial minutes text into a first minutes text conforming to a preset style, making the generated text more professional and readable. Humanized feature injection makes the second minutes text more natural and closer to human expression habits, enhancing the text's naturalness and credibility. The second minutes text is then populated into the target meeting minutes template to generate a structured target minutes text, ensuring consistency and standardization of the output.

[0021] The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0022] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of an intelligent meeting minutes generation method provided in this application.

[0023] like Figure 2 As shown, the intelligent meeting minutes generation method includes steps S101 to S107.

[0024] S101. By performing audio preprocessing on the multi-channel audio collected by the multi-channel microphone array, an audio stream with speaker identifiers is generated.

[0025] In one embodiment, multiple microphone arrays are used to capture audio data during the meeting, ensuring coverage of all speakers. Exemplarily, a circular or linear multi-channel (e.g., four-channel or higher) microphone array can be used, employing beamforming technology for spatial filtering and directional enhancement of speaker voice. The multi-channel audio is captured synchronously via an audio controller or network audio protocol.

[0026] Audio preprocessing operations are performed on multi-channel audio, including noise reduction, echo cancellation, and speaker separation, to generate a high-quality speech stream with speaker identification, providing reliable input for subsequent processing. Specifically, the audio preprocessing of multi-channel audio includes: applying a deep learning model for noise reduction to remove background noise and improve speech clarity; using adaptive filter technology to eliminate echoes in the conference room to ensure audio quality; and using speaker recognition technology to separate the speech of different speakers, providing a foundation for subsequent processing.

[0027] Furthermore, based on a deep learning-based multi-channel speech enhancement model, background noise is eliminated from the multi-channel audio to obtain a noise-reduced audio stream; an adaptive filter algorithm is used to eliminate echo noise in the noise-reduced audio stream to obtain echo-cancelled audio; based on a voiceprint recognition and clustering algorithm, audio separation is performed on the echo-cancelled audio, and the speaker identifier of each separated audio stream is labeled to obtain the audio stream with the speaker identifier.

[0028] In one embodiment, a deep learning-based speech enhancement model, such as a convolutional recurrent network (CRN) or a temporal audio separation network (TasNet), can be used to suppress noise in the frequency or temporal domains. Specifically, the deep learning-based speech enhancement model employs a multi-channel extended variant of the temporal audio separation network (Conv-TasNet). The encoder of the temporal audio separation network uses a one-dimensional convolutional layer to convert the temporal waveform into a 256-dimensional deep feature representation. This process avoids explicit Fourier transform, thus preventing phase information loss and providing better suppression of non-stationary noise (such as keyboard clicks and door / window opening / closing sounds). The separation network introduces spatial covariance matrix features in the channel dimension; for each time step, the cross-correlation matrix between channels (of shape [H×H], where H is the number of channels) is calculated, flattened, and used as a spatial feature vector, which is concatenated with the deep acoustic features and input into an 8-layer temporal convolutional network (TCN). Each TCN layer contains dilated convolutions and residual connections, with a receptive field covering the context, effectively capturing the harmonic structure of long-term periodic noises such as air conditioning noise. The decoder reconstructs the time-domain waveform by element-wise multiplying the mask with the encoder output using transposed convolution. During training, a scale-invariant signal-to-noise ratio loss function is employed, along with a spectral contrast loss to enhance the ability to distinguish typical noises in a meeting setting (such as the sound of turning pages and breathing). During inference, a dynamic noise type detector is used to determine the dominant noise category (keyboard / air conditioning / traffic) in real time, adaptively switching the weights of sub-networks within the model to achieve online adjustment of noise suppression strength.

[0029] In one embodiment, an adaptive filtering algorithm (such as NLMS or Kalman filtering) combined with a two-talk detection (DTD) mechanism is used to eliminate acoustic echo between the microphone and the speaker. Specifically, a statistical model based on the probability of speech presence can be used. Two filters are run in parallel through this statistical model. The main filter converges normally, while the reference filter converges at a slower speed. When the difference between the outputs of the two filters exceeds a statistical threshold, it is determined to be a two-talk state, and the coefficient update of the main filter is immediately frozen to prevent near-end speech from being misjudged as echo and attenuated. The filter output is then cascaded into a residual echo neural network suppressor. The input of this network is the differential feature between the filtered audio and the far-end reference signal, and the output is a time-frequency mask, which specifically eliminates nonlinear residual echoes (such as harmonic echoes) that are not fully modeled by the filter.

[0030] Based on voiceprint recognition and clustering algorithms, speaker embedding vectors are extracted using x-vector or ECAPA-TDNN. For multi-channel audio, time difference of arrival (TDOA) and direction of arrival (DOA) features are utilized. The GCC-PHAT algorithm is used to calculate the time difference between each pair of microphones, converting it into a 256-dimensional spatial feature vector. This vector is then concatenated with the voiceprint embedding to form a 448-dimensional joint feature vector. When voiceprint similarity is high (e.g., family members attending a meeting), spatial location information is used to achieve reliable differentiation. Speaker segments are processed sequentially in a streaming manner. Unsupervised speaker segmentation and clustering are performed using spectral clustering or variational autoencoder (VAE), generating metadata tags with speaker ID, gender, and role (host / attendant). During clustering, a speaker fingerprint database is maintained, and historical voiceprint IDs are automatically matched for repeated speakers (e.g., departmental meetings) to avoid duplicate labeling. When dual-speaking or multiple speakers are detected simultaneously, the speech separation front-end is activated. A pre-trained overlapping speech separation model is used to separate overlapping speech, outputting two independent audio tracks from which voiceprint features are extracted.

[0031] This embodiment significantly improves the signal-to-noise ratio and clarity of conference audio through the synergistic effect of multi-channel deep learning speech enhancement, adaptive echo cancellation, and speaker recognition clustering technology. It achieves accurate separation and identity labeling of multiple speakers, effectively solves the problem of overlapping speech separation, greatly reduces speaker confusion rate, provides high-quality, timestamped, and spatially located structured audio stream input for subsequent ASR transcription, improves speech recognition accuracy, and lays a solid foundation for intelligent meeting minutes generation.

[0032] S102. Input the audio stream with speaker identifiers into the speech recognition model, and convert the audio stream into structured text with timestamps through the speech recognition model to obtain the initial summary text of at least one speaker.

[0033] In one embodiment, an end-to-end multi-task Automatic Speech Recognition (ASR) model, such as Whisper-large-v3, is employed to convert audio streams into timestamped structured text. This model, based on the Transformer architecture, possesses powerful multilingual recognition capabilities and contextual understanding. Pre-trained, the model can simultaneously perform multiple tasks, including speech recognition (ASR), speaker identification (Diarization), and language recognition (LanguageID). For meeting scenarios, the model undergoes domain-adaptive fine-tuning, and its parameters are further optimized using an internal enterprise meeting corpus (containing various meeting types and different speaker styles), enhancing the accuracy in recognizing technical jargon and industry slang. The output is structured text with timestamps, speaker IDs, and confidence scores.

[0034] The ASR model's encoder consists of 12 Conformer modules, each integrating an 8-head self-attention mechanism and depthwise separable convolutions. It also introduces Rotated Position Encoding (RoPE) to replace traditional absolute position encoding, enhancing its ability to model long-distance dependencies in long-duration conference audio. The decoder uses a two-layer unidirectional LSTM prediction network. The joint network fuses the acoustic representation output by the encoder with the decoder's historical prediction vectors, outputting character-level posterior probabilities.

[0035] In one embodiment, the ASR model can support multi-speaker identification and start / end time stamping, and assign a unique ID to each speaker. Specifically, a start / end timestamp (accurate ±0.1 seconds) is appended to each identified word or sentence, and the speaker ID is recorded. The timestamp information can be used for text-to-audio alignment, playback localization, and speech duration statistics.

[0036] Specifically, timestamp generation employs a frame-level forced alignment algorithm. The earliest and latest acoustic frame indices for each output character are maintained within the Transducer decoding path, and these are precisely converted into absolute time values ​​using the audio sampling rate and frame shift. To eliminate decoding jitter, a Kalman filter is applied to the timestamp sequence of consecutive characters for smoothing, and hard boundary correction is performed based on speaker switching points: when a speaker ID change is detected, a sentence-ending punctuation mark (period or question mark) is forcibly inserted before the switch, and the end timestamp of the previous speaker is aligned with the start timestamp of the next speaker to the same moment, avoiding temporal overlap. For long sentences (e.g., >10 seconds), they are automatically segmented into multiple clauses based on prosodic pauses (e.g., pauses >0.5s), with each clause independently timestamped to facilitate accurate location of speech segments during subsequent viewpoint extraction.

[0037] In one specific embodiment, a preprocessed multi-channel audio stream containing speaker IDs and signal-to-noise ratio metadata is received. A sliding window with a frame length of 25ms and a frame shift of 10ms can be used to extract 80-dimensional log-Mel features and 30-dimensional prosodic features (fundamental frequency, jitter). The latter helps distinguish between interrogative and declarative sentences, improving punctuation prediction accuracy. An end-to-end Conformer-Transducer ASR model is used, with a 12-layer Conformer encoder (with 8-head self-attention and rotational position encoding) and a 2-layer LSTM decoder. The joint network outputs the posterior probability of characters, balancing accuracy and streaming inference capabilities. A Speaker-AwareAttention layer can be added to the top layer of the ASR model's encoder to embed the speaker ID as a 128-dimensional vector and dynamically inject it, allowing the model to adapt to specific accents and terminology preferences. Simultaneously, the ASR model supports LoRA incremental learning; adding a new speaker only requires 5 minutes of samples to generate a dedicated adapter. To address the issue of mixed Chinese and English text, sub-word units and Chinese characters can be integrated into a pre-defined vocabulary, and a Code-SwitchingDetection sub-network can be deployed. Language switching points are identified through three frames of acoustic features, and the language model weights are dynamically adjusted. The ASR model employs a streaming inference approach, with the encoder inferring forward in 500ms increments and the decoder maintaining a 2-second dynamic alignment buffer. Character output is triggered when the confidence level exceeds 0.85. For the timestamp sequence of output characters, frame-level forced alignment is used to record the earliest and latest acoustic frame indices of the characters, followed by Kalman filtering to smooth and eliminate jitter. When a change in speaker is detected, sentence-end punctuation is forcibly inserted at the speaker switching point, and the boundaries are corrected. The final output is structured text with timestamps, yielding an initial summary text with speaker identifiers.

[0038] The ASR model takes a preprocessed audio stream as input and outputs a sequence of text tokens with timestamps and confidence scores. The transcription results of the ASR model are output in JSON or Protobuf format, containing the following fields: speaker unique identifier (speaker_id), timestamp (start_time / end_time), transcribed text, confidence score, and language (supporting mixed Chinese and English recognition).

[0039] In one embodiment, during the ASR model transcription process, the identified filler words (such as “um”, “ah”, “that”) can be encoded as semantic placeholders instead of being deleted as noise, and their functions in the sentence can be labeled, such as the type and duration of hesitation, to provide clues for subsequent semantic understanding and sentiment analysis.

[0040] In one embodiment, an acoustic-semantic alignment error propagation mechanism can also be established. When the model's confidence in recognizing a word or phrase falls below a preset threshold (e.g., 0.6), the error vector of that word or phrase and its context is passed to the subsequent semantic understanding module. The error vector contains information about the model's uncertainty in recognizing the word or phrase, which is used to trigger context-compensated inference, i.e., re-evaluating the semantic contribution of the word or phrase in a broader context, thereby improving the overall accuracy of semantic understanding.

[0041] In one embodiment, a speaker's speech heatmap can be generated based on acoustic features such as sound intensity, speech rate, and spectral energy to identify areas of active thought. These areas typically correspond to the speaker's key viewpoints, emotional expressions, or logical reasoning during the discussion. By analyzing these areas, subsequent semantic attention mechanisms can be driven more accurately, enabling the model to focus more on these key aspects when processing semantic information.

[0042] Specifically, features such as sound intensity, speech rate, and spectral energy are extracted from the audio signal. A speech heatmap is generated based on these features, highlighting areas of active thought. This heatmap information is then input as additional attention weights into the semantic understanding model, guiding it to focus more on the semantic content of these active areas. During the semantic understanding stage, the model's attention mechanism is adjusted using information obtained from the speech heatmap mapping. Specifically, when processing each word or phrase, the model dynamically adjusts its level of attention based on its position and weight in the heatmap. This allows the model to focus more on the parts of the speech heatmap marked as areas of active thought, thereby improving its ability to capture key information.

[0043] S103. Using a large language model, perform semantic analysis on the initial summary texts corresponding to each speaker to extract key text information.

[0044] In one embodiment, the initial minutes text can be preprocessed before being input into a large language model for semantic analysis to ensure text quality and analysis accuracy. Specifically, this involves: removing irrelevant characters, redundant spaces, and special symbols from the initial minutes text through text clarity; standardizing punctuation usage; dividing the text into sentences and paragraphs based on punctuation and semantic pauses (such as long pauses and line breaks) to simulate natural block division in human reading; associating speaker IDs with corresponding text paragraphs to provide contextual information for subsequent analysis; and ensuring that each sentence or key phrase in the text has a corresponding timestamp to facilitate tracking the specific time when information appeared during the meeting.

[0045] Choose a large language model suitable for long text processing and deep semantic understanding, such as GPT-4, BERT General Questions, or their variants. Fine-tune the large language model to better handle formal documents, recognize technical terms, and understand complex business contexts, taking into account the characteristics of meeting minutes. For example, further training the model can be done using internal company meeting minutes, industry reports, and related documents to enhance its understanding of domain-specific terminology and context. The model can be configured to perform multiple sub-tasks, including entity recognition, relation extraction, sentiment analysis, and summary generation, to comprehensively extract key information from the text.

[0046] Specifically, the entity recognition subtask identifies entities in the text, such as names of people, places, organizations, projects, and technical terms; the relation extraction subtask analyzes the relationships between entities, such as "who is responsible for what" and "which project is executed by which department," to facilitate understanding of the meeting content; the viewpoint summarization subtask extracts the main viewpoints and suggestions of each speaker, especially those involving decision-making and action items; and the sentiment analysis subtask assesses the speaker's sentiment orientation, such as positive, negative, or neutral, which helps in understanding the speaker's stance and attitude.

[0047] Using a finely tuned large language model, we perform granular analysis on the content of each speaker, identify core viewpoints, extract keywords and topic tags, and form a "viewpoint-tag" mapping.

[0048] Furthermore, based on the speaker identifiers corresponding to each audio stream, the initial summary texts corresponding to the same speaker identifier are input into the large language model; through the large language model, the tasks of summarizing core viewpoints, extracting keywords, and classifying topic tags are executed sequentially to obtain the initial key text; through the dialogue state tracking mechanism, the initial key text is ambiguously resolved to obtain the key text information.

[0049] First, the initial minutes text can be reconstructed based on speaker identifiers, stitching together all speaking segments from the same speaker in the meeting according to timestamps to form a speaker-level continuous discourse. To avoid information overload, a sliding window strategy can be used, with adjacent windows retaining 50% overlap to ensure semantic coherence. Before inputting into the large language model, a structured prompt template can be constructed, including: task instructions ("Please analyze the following speaking content"), meeting agenda context (current discussion topic and summaries of the previous 3 relevant speaking segments), enterprise knowledge base (relevant project background, organizational structure), and speaker identity metadata (name, department, job title, and historical speaking style tags).

[0050] The multi-task execution architecture of a large language model can adopt a single-model multi-task framework based on prompting engineering, which serially executes multiple sub-tasks (such as core idea summarization, keyword extraction, and topic tag classification tasks) on the same input text. Each task is decoupled from the output format constraints through delimiter marking.

[0051] In one embodiment, a large language model (such as GPT-4 or Tongyi 1000 Questions) can employ chain-of-thought reasoning to first generate an internal thought chain and then output structured JSON. Attention weights are shared between tasks, allowing keyword extraction to feed back into summary generation (e.g., key nouns appearing preferentially in the summary), thus improving overall consistency.

[0052] Specifically, summary generation can follow a subject-verb-object structured constraint, forcing the output to include complete semantic units containing the subject, action, and object. To prevent the model from generating non-deterministic expressions such as "I think" or "maybe," a deterministic reinforcement instruction is injected into the prompt: "Only extract clearly stated factual viewpoints and eliminate speculative content." A length penalty mechanism is also introduced, using `logit_bias` to suppress the tendency to generate summaries exceeding 50 characters. For long texts, a recursive summarization strategy can be adopted, i.e., first generating sub-summaries for each window, then performing secondary aggregation on the sub-summaries, and finally outputting a concise overall summary. Summary quality is automatically evaluated by a semantic coverage metric, calculating the cosine similarity between the summary vector and the original text vector; if it is below 0.75, regeneration is triggered.

[0053] Keyword extraction integrates statistical significance and semantic importance. The large language model first calculates the TF-IDF (Term Frequency-Inverse Document Frequency) score for each noun phrase to identify high-frequency and discriminative candidate words. Then, through attention weight backtracking, entities that received high attention weights during the summary generation process (such as names of people, projects, and amounts) are extracted. The final keyword list is sorted by weighted score (e.g., ...). ,in, Indicates the weighted score. This represents the attention score. (This represents the word frequency score). To ensure the operability of keywords, it is mandatory to include at least one executable object (such as "risk control system upgrade") and one quantitative indicator (such as "2 million" or "Q3"). Otherwise, the "keyword completion subtask" is triggered, allowing the model to actively inquire about implicit elements.

[0054] The topic tags are drawn from a predefined tag library of enterprises (typically 200-500 tags), organized hierarchically (e.g., "budget approval / marketing budget / additional request"). The model employs a hierarchical classification strategy: first predicting the primary tag ("budget approval"), and then predicting secondary tags within its child nodes. During classification, the maximum cosine similarity between the candidate tag vector and the text semantic vector is calculated, supplemented by speaker role matching (e.g., a CFO's speech is preferentially matched with "budget approval" rather than "technology R&D").

[0055] In one embodiment, during the initial key text extraction process, sentiment analysis technology can also be introduced to identify the speaker's emotional inclination and provide additional information for opinion extraction.

[0056] Specifically, acoustic features (such as prosodic features, spectral features, and sound quality features) are extracted from each frame of speech signal. These acoustic features can then be input into a lightweight convolutional neural network (CNN) to extract the probability distribution of emotions, such as positive, negative, neutral, anxious, excited, hesitant, and angry.

[0057] After ASR transcription is completed, each sentence-level text fragment can be input into the domain sentiment BERT model. This model adds business sentiment corpora (such as earnings call transcripts and internal review records) to the sentiment analysis task, enhancing its ability to identify complex sentiments such as "cautiously optimistic," "prudently reserved," and "strongly opposed." The output is a sentiment polarity score (-1.0 to +1.0) and sentiment intensity (0.0 to 1.0), along with sentiment keywords (such as "risk," "challenge," and "opportunity").

[0058] The sentiment probability distribution vector output by the audio sentiment CNN is concatenated with the sentence vector output by the text sentiment BERT, and then fed into a two-layer bidirectional LSTM fusion network to learn the dynamic evolution of sentiment across time steps. The output is a frame-level fused sentiment vector, which is then aggregated into a speech segment-level sentiment representation through attention pooling. This architecture can capture inconsistencies between audio and text sentiment, such as a speaker's hesitant tone but positive word choice ("Hmm...this project should succeed, right?"), and automatically calibrates it to "low-confidence positive" after fusion.

[0059] The fused sentiment results are accompanied by a sentiment confidence score. During the opinion extraction stage, sentiment information is injected into the prompt of the large language model as external prior knowledge. When generating summaries, the large language model applies attention enhancement to negative sentiment words ("risk", "uncertainty", "worry"), so that they are preferentially retained in the core opinion summary, and automatically adjusts the summary tendency to risk warning type (such as "warning about Q3 cash flow risk" rather than "suggesting implementation plan").

[0060] If a speaker's emotion is strongly negative (e.g., anger, intensity > 0.8), but their opinion summary is neutral or positive, an emotion-opinion conflict detection is triggered. The system automatically marks it as "potential irony or repressive expression," retains the original emotion label ("[strong dissatisfaction]") in the minutes, and suggests manual review. Conversely, if the emotion is excitement and the opinion contains words such as "suggestion" or "promote," the action item is automatically prioritized one level.

[0061] For each speaker, a sentiment timeline curve is plotted (horizontal axis = meeting time, vertical axis = sentiment polarity), identifying sentiment abrupt changes (such as a sudden drop from positive to negative), and automatically labeling them as "sensitive moments in the agenda," driving downstream modules to perform high-density semantic analysis on the text during that period. For example, when the curve shows a negative peak during the discussion of "budget cuts," all keywords for that period are extracted, and a "risk snapshot" is generated and appended to the minutes.

[0062] In one embodiment, the group emotional resonance can be calculated by aggregating the emotional vectors of all speakers:

[0063] in, Indicates the degree of emotional resonance within a group. This indicates the average emotional level of the speaker. This indicates the speaker's emotional standard deviation. Indicates the first The speaker's sentiment value; Indicates the total number of speakers.

[0064] Group emotional resonance A high value indicates that the group's emotions are consistent (e.g., everyone actively supports the group), and the group's emotional resonance is high. A low value indicates significant disagreement. When the group's emotional resonance is low (e.g., ... Furthermore, when the issue involves decision-making, the minutes will automatically generate an alert for “group disagreements that need attention” and list the speakers with opposing sentiments to help managers identify potential conflicts.

[0065] In this embodiment, sentiment analysis technology, through dual-path sentiment recognition that integrates audio prosody features and text semantics, and establishing a sentiment-opinion joint modeling mechanism, can accurately capture the speaker's true intentions, potential risks, and ironic expressions, significantly improving the accuracy and depth of opinion extraction. At the same time, through sentiment temporal evolution analysis and group sentiment profiling, it can dynamically monitor emotional shifts and group disagreements in meetings, providing managers with decision-making warnings. The final generated sentiment-tagged minutes text is more humanized and cautionary, and the model is continuously optimized through a closed loop of human feedback, significantly enhancing the understanding and robustness of the intelligent meeting minutes system in complex business scenarios.

[0066] In one embodiment, the Dialogue State Tracking (DST) mechanism maintains a session-level memory vector, recording the topics discussed, consensus reached, and actions taken so far in the meeting. For ambiguous references in the initial key text (such as "this plan" or "his suggestion"), DST performs reference resolution. Specifically, it retrieves the most recently appearing candidate entity from the memory vector (such as the previously discussed "Q3 marketing plan"), calculates the semantic distance (such as cosine similarity) between the current word and the candidate entity, and automatically replaces it with the full entity name if the distance is <0.3 and the grammatical roles match (subject / object). For elliptical sentences (such as "agree"), DST automatically completes the omitted components, restoring it to "agree [the aforementioned Q3 budget increase plan]". This dialogue state tracking mechanism ensures that each speaker's key text remains independently readable even outside the session context.

[0067] S104. Based on the key text information, the decision content and task elements in the initial minutes text are jointly judged by the rule engine and the large language model to generate decision and action information.

[0068] In one embodiment, a rule engine is used to remove non-deterministic expressions based on preset rules for judging negative words and deterministic semantics. A large language model is then used to perform in-depth analysis on the non-deterministic expressions output by the rule engine. The joint judgment by the rule engine and the large language model can retain definite decision-making matters and actionable actions.

[0069] Specifically, the rule engine performs an initial filtering of the original text transcribed from ASR based on a pre-defined dictionary of negative expressions and deterministic semantic judgment rules. The dictionary includes vague expressions such as "I think," "maybe," "suggest," and "tend to," as well as their synonym extensions (e.g., "perhaps," "probably," "it seems"). Candidate statements are identified through regular expression matching and dependency parsing. The rule engine also captures strong semantic verbs ("decide," "approve," "require") and their subject-verb-object structures. Statements that satisfy the pattern "subject (decision-maker) + strong verb + object (matter)" are marked as deterministic candidates, while statements containing vague words but without strong verbs are marked as uncertain candidates. For boundary cases that are difficult to determine by the rules (e.g., "agree in principle, but the plan needs to be refined"), pending candidates are generated and submitted to a large language model for in-depth analysis.

[0070] For each candidate statement, a structured input prompt is constructed, which contains four core components: the original text of the candidate statement, the speaker's identity and role (such as "Zhang Wei, the CFO"), the meeting agenda context (the current discussion topic and a summary of previous speeches), and knowledge of corporate decision-making norms (such as "budget approval requires clear definition of amount and implementing department").

[0071] After receiving a prompt, a large language model (such as GPT-4 or Tongyi Qianwen) activates its chain-of-thought capability, performing a three-step logical deduction: First, it deconstructs the speech act of the statement to identify whether it belongs to "commitment," "request," or "comment"; second, it analyzes the sources of semantic ambiguity, identifying uncertain expressions such as "can" and "let's see the results" and their degree of weakening of enforceability; finally, it assesses the binding force of the statement in the meeting context, determining whether there are implicit conditions or reservations. The model generates internal reasoning traces, such as: "In this statement, 'can' indicates permission but not mandatory; 'let's see the results' introduces uncertain conditions, lacks a clear purpose and responsible department, and does not meet the enterprise budget approval norms, therefore it is classified as 'non-deterministic decision'." The model output strictly follows the JSON schema specification. When a decision is determined to be deterministic, the output fields must include the three elements of the decision: subject (who makes the decision), object (the content of the decision), and effect (implementation requirements). When a decision is determined to be non-deterministic, the output must clearly state the reasons for rejection (such as "semantic ambiguity", "lack of implementing subject", "uncertainty of conditions") and list the missing elements (such as ["responsible department", "deadline"]).

[0072] Furthermore, based on preset statement filtering rules, the key text information is filtered by the rule engine to obtain target statements and candidate statements; based on the large language model, semantic reasoning and confidence scoring are performed on the candidate statements filtered by the rule engine, and the binary classification results corresponding to the candidate statements are output; a joint decision-making mechanism is adopted to vote on the target statements, the candidate statements, and the binary classification results corresponding to the candidate statements to obtain decision information; action item information matching the decision information is extracted from the key text information to obtain the decision and action information.

[0073] In one embodiment, the rule engine maintains a hot-loadable rule knowledge base, which includes regular expression rules, dependency syntax rules, and synonym expansions. The rule engine performs progressive filtering on key text information, and each statement can be divided into three levels based on the degree of matching: target statement, candidate statement, and non-decision statement.

[0074] Specifically, the target statement refers to a statement that satisfies a strong decision-making pattern and is directly identified as a valid decision. The rule engine uses dependency parsing to match pre-defined decision templates, such as [Subject: Decision Maker] + [Core Verbs: {"decide", "approve", "veto", "require"}] + [Object: Matter] + [Optional: Decision Basis]. A statement must contain at least one quantifiable element (amount, time, quantity) and an implementing entity (a person's name or department identified through NER); otherwise, it is downgraded to a candidate statement. If the sentiment analysis result is hesitant (anxiety / uncertainty) and the intensity is >0.6, even if the syntax matches, it is forcibly downgraded to a candidate statement, requiring verification of its decision-making effectiveness by a large language model.

[0075] Candidate statements refer to statements that meet the weak decision-making pattern or contain vague expressions, which require semantic reasoning by a large language model. Examples include vague verbs (such as "suggest", "believe", etc.), conditional clauses (such as "if", "premise", etc.), negative expressions (such as "I think", "maybe", etc.), and authority mismatch (such as the speaker's role (such as an intern) not matching the authority of the decision-making matter, requiring confirmation from a superior).

[0076] Non-decision statements refer to pure statements, questions, greetings, etc., and are filtered out directly.

[0077] For each candidate statement selected by the rule engine, a multi-dimensional Prompt is constructed, which includes: task instructions, candidate original text and rule engine judgment, global meeting context, speaker behavior profile, enterprise decision-making norms, output format constraints, etc.

[0078] The large language model uses chain reasoning to first classify candidate statements by speech act, identifying their type, such as promises, requests, suggestions, or comments. Then, it checks the completeness of elements against corporate management principles. For example, if the model finds that "approve 2 million first" lacks "purpose" and "deadline," it classifies the candidate statement as an incomplete decision. Finally, it mines implicit intentions, questioning the true intent behind tentative words like "try it first" and "see the effect." The model infers that "this is a pilot proposal rather than a final decision," generating an internal explanation: "This statement is a low-risk probe; a final decision requires subsequent effect evaluation, therefore it does not currently constitute an executable task."

[0079] Furthermore, the semantic features and contextual features of the candidate statements are analyzed using the large language model to calculate the confidence score corresponding to the candidate statements. When the confidence score is less than or equal to a preset confidence threshold, the binary classification result corresponding to the candidate statement is determined to be an invalid decision. When the confidence score is greater than the confidence threshold, the binary classification result corresponding to the candidate statement is determined to be a valid decision, and the candidate statement is marked as the target statement.

[0080] Before performing semantic reasoning analysis on candidate statements, semantic features of the statements are first extracted using a large language model. Specifically, candidate statements are input into a pre-trained large language model, such as BERT or GPT series, to obtain an initial encoded representation of the statements. Combining the contextual information of the statements in the conference text, attention mechanisms or masked language models are used to enable the model to capture the contextual dependencies of the statements. Based on the encoding, the semantic roles (such as subject and object), syntactic structures (such as dependency relations), and potential semantic patterns (such as conditional sentences and hypothetical sentences) of the statements are further extracted.

[0081] For each candidate statement, the large language model can output a confidence score, which reflects the model's confidence that the statement constitutes a valid decision. The confidence score can be calculated based on the probability distribution of the model's output, or it can be generated directly through the model's internal attention weights or other mechanisms.

[0082] For example, the confidence score can integrate three sub-dimensions: semantic confidence, contextual consistency, and normative fit. Semantic confidence can be based on the probability distribution of the model's softmax output, reflecting the degree to which the statement conforms to the decision-making pattern. Contextual consistency can be obtained by calculating the cosine similarity between the semantic vector of the candidate statement and the agenda topic vector, preventing misjudgment across topics. Normative fit can be obtained by matching and scoring the statement elements against the enterprise's decision-making norm template; missing key elements (such as the responsible person) will result in a deduction of 0.1 points for each missing element.

[0083] Final confidence score The calculation is as follows:

[0084] in, Indicates semantic confidence. Indicates contextual consistency. Indicates the degree of canonical matching; , , These are the weight coefficients corresponding to semantic confidence, contextual consistency, and specification matching, respectively.

[0085] A confidence threshold can be set. When the confidence score is less than or equal to the confidence threshold, the large language model determines that the candidate statement does not constitute a valid decision and outputs "invalid decision". When the confidence score is greater than the confidence threshold, the large language model determines that the candidate statement constitutes a valid decision, outputs "valid decision", and marks the statement as the target statement.

[0086] The confidence threshold needs to be adjusted based on the model's performance on the validation set to balance precision and recall. A confidence threshold that is too low may lead to too many false positives, while a confidence threshold that is too high may miss valid decisions. Typically, the initial setting of the confidence threshold can be based on a certain quantile of the probability distribution output by the model (such as 0.7 or 0.8), and then fine-tuned according to the needs of the actual application scenario.

[0087] After both the rule engine and the large language model have determined the decision statement, a joint decision is made based on their outputs. The joint decision-making mechanism aims to integrate information from different sources, including the target statement, candidate statements, and the binary classification results provided by the large language model, to improve the accuracy and reliability of the decision information.

[0088] Specifically, voting weights can be assigned to the target statement, candidate statements, and the binary classification results provided by the large language model. The target statement, having already been filtered by the rule engine, has high credibility and is assigned a higher initial weight (e.g., 0.6); candidate statements, not yet fully validated, have a lower initial weight (e.g., 0.2); and the binary classification results, derived from the analysis of the large language model, have their weights dynamically adjusted based on the model's confidence score. If the confidence score is higher than a threshold, the weight is increased; if it is lower than the threshold, the weight is decreased.

[0089] Based on the aforementioned weights, a weighted vote is applied to the decision results of each statement. If the total weight exceeds a preset threshold (e.g., 0.7), the decision is accepted as valid. When the decision results of the target statement and the candidate statements conflict, a more conservative strategy is adopted, submitting the decision for manual review or further analysis. For example, for decision results close to the threshold, a manual review process is introduced to ensure the accuracy of the decision; or, the decision results are compared with similar historical cases to verify the consistency and rationality of the decision.

[0090] The action items are extracted from key textual information to match the decision-making information, forming a complete decision and action information framework. Action item elements include task content, responsible person, deadline, and priority. Specifically, task content refers to the concrete action description extracted from the decision statement, such as "complete the risk control system upgrade plan"; the responsible person refers to the executor mentioned in the decision statement, or inferred from organizational structure and historical data; the deadline refers to the time expression in the decision statement, such as converting "before the end of the month" into a specific date; and the priority is assigned based on the importance and urgency of the decision.

[0091] The extracted action item information is structured into an easy-to-understand and execute format. Decision information and action item information are integrated to form a complete decision execution plan. Specifically, each decision can be mapped to one or more action items, clarifying the execution path; dependencies between action items are defined to ensure a reasonable execution order.

[0092] This joint decision-making mechanism and action item information extraction method ensures that the decision-making information extracted from meeting minutes is not only accurate and reliable, but also executable, thereby improving the efficiency and effectiveness of meeting decisions.

[0093] S105. Based on the key text information and the decision and action information, perform style transfer on the initial minutes text to convert the initial minutes text into a first minutes text with a preset text style.

[0094] In one embodiment, a generative large model using instruction tuning is employed to perform style transfer on the initial minutes text, transforming colloquial, impromptu remarks into formal, concise written language suitable for high-level reporting.

[0095] Specifically, the initial minutes text, along with key textual information extracted from upstream (core viewpoints, keywords, and topic tags) and decision-making and action information (clear decision-making matters and task elements), are input into a generative large model that has been fine-tuned by instructions. Through multi-dimensional style control (including formality, simplicity, power distance, and domain adaptability), the model is guided to optimize vocabulary selection, reconstruct sentence structure, and enhance logical coherence while preserving the original meaning. Natural interjections and appropriate logical jumps are also injected to reduce AI traces. The final output is a first minutes text that is structurally consistent, highly professional, and possesses natural human expression characteristics, thereby meeting the reporting needs of different levels of the enterprise and improving the readability and credibility of the minutes.

[0096] Furthermore, based on preset text style features, the preset text style is modeled, and a target style vector is calculated; the target style vector, the key text information, and the decision and action information are used as style constraints, and a multi-level style transfer is performed on the initial summary text using the cross-attention mechanism of text style and text content to generate the first summary text.

[0097] In one embodiment, a preset text style can be quantified into a multi-dimensional continuous feature space, with each dimension taking values ​​in the range [0,1], forming a target style vector, including formality, conciseness, power distance, domain adaptability, emotional tone, logical realism, information hierarchy, political sensitivity, etc. The target style vector can be dynamically generated based on real-time context through agenda-driven, speaker-aware, and decision-level mapping. The calculation process employs a meta-controller: the input is a one-hot encoding of meeting type, reporting recipient, decision importance, and emotional label; the output is a target style vector through a small neural network (2-layer MLP), achieving adaptive configuration of style parameters.

[0098] In addition to the target style vector, two other types of information are encoded as constraint tensors: key text information constraint and decision and action information constraint. Specifically, the core abstract, keywords, and topic tags can be encoded as feature vectors to ensure that these high-frequency semantic elements are preferentially retained during style transfer, obtaining the key text information constraint; the decision content and action items (tasks, responsible persons, deadlines) can be encoded as a structured vector sequence to obtain the action constraint vector.

[0099] The target style vector, key text information constraint, and decision and action information constraint are concatenated into a joint constraint vector and projected into the same feature space as the initial summary text encoding through a linear transformation to ensure interactivity in the attention mechanism.

[0100] The decoupled dual encoder can be used to process content and style separately. Specifically, based on the Transformer encoder, the initial summary text (spoken language) is deeply semantically encoded to output the content hidden vector ; local features of the joint constraint vector are extracted through a lightweight CNN to output the style hidden vector .

[0101] At each time step of the decoder, the content hidden vector and the style hidden vector are input into the multi-head cross-attention layer to calculate the style-content attention matrix:

[0102] where, comes from the content encoder, , come from the style encoder, is the style intensity bias matrix, which is dynamically adjusted according to the values of each dimension of the target style vector: the higher the formality, the greater the penalty for the attention weight of spoken words (-∞ mask).

[0103] During the decoding process, the decoder adopts hierarchical progressive generation, imposing style constraints at three levels: vocabulary, syntax, and discourse.

[0104] Specifically, at the vocabulary level, adversarial word replacement and domain term injection are adopted. Maintain the mapping dictionary of the spoken-written language adversarial word list. Before the word list softmax, add the style adversarial loss. Through the gradient reversal layer (GRL), the model is made to learn to map spoken words ("get it done") to the written language space ("complete"), while retaining the word meaning. For the term library activated by the topic tags, the correction value of the corresponding tag is increased during decoding to force the output of professional terms. For example, the generation probability of the term "risk control" is naturally higher in the financial scenario than in the general scenario.

[0105] At the syntactic layer, a TreeLSTM-based syntactic generator is used to transform the colloquial, flat syntactic tree into a written, hierarchical tree. To avoid generating monotonous subject-verb-object structures, a sentence pattern pool (inversion, passive voice, nominalization) can be introduced, randomly switching sentence patterns during decoding. For example, "We have completed the project" can be transformed into the passive "The project is complete".

[0106] At the discourse level, a logical relation classifier trained on the PDTB corpus identifies implicit causal and adversative relationships between sentences and automatically inserts conjunctions such as "therefore" and "however." The insertion probability is dynamically adjusted by the logical explicit dimension of the target style vector. For executive-level minutes, a pyramid principle compression algorithm is used, recursively deleting secondary supporting sentences using sentence importance scores (integrating TF-IDF, positional weights, and speaker rank), retaining only the conclusion and primary arguments.

[0107] In one embodiment, style transfer technology can be introduced to adjust the language style according to different reporting recipients. A multi-dimensional profile of the reporting recipient (including job level, professional background, and historical style preferences) is obtained in real time through the enterprise organizational structure API. Job level is then mapped to a power distance coefficient. The higher the rank The larger the value, the more prominent the humility marker. Based on professional background, the corresponding industry terminology preference library is activated. For example, when reporting to the CTO, "risk control system upgrade" automatically expands to "risk control platform upgrade based on microservice architecture," while for the CFO, the emphasis is on "ROI and compliance." Based on historical style preferences, the report recipient's past approval minutes are analyzed, and their preferred sentence patterns are extracted through style cosine similarity clustering. For example, if a CEO prefers an "elevator pitch" structure (general-specific-general), the weight of the conciseness dimension is dynamically increased.

[0108] In one embodiment, deep syntactic-semantic analysis can be performed on the initial summary text before style transfer. Dependency parsing is used to identify core predicates and their subjects, objects, adverbs, and other components, constructing a syntax tree. A pre-trained BERT-based SRL model is used to label semantic roles (agent, patient, time, place, reason, etc.) to ensure that the transformed text is logically clear and structurally sound.

[0109] Furthermore, after outputting the first summary text, inverse style prediction can be performed on the first summary text based on a style classifier to obtain the actual style vector of the first summary text; the feature similarity between the actual style vector and the target style vector is calculated; when the feature similarity is less than a preset similarity threshold, the style transfer operation is performed again on the initial summary text until the feature similarity between the actual style vector of the generated first summary text and the target style vector is greater than or equal to the similarity threshold.

[0110] In one embodiment, to verify the style consistency of the generated text, a lightweight style classifier decoupled from the generation model can be deployed. This classifier is trained independently on a parallel corpus containing 100,000 pairs of <colloquial text, target style text>, with each pair receiving an eight-dimensional style score (eight dimensions including formality, conciseness, and power distance) through manual annotation. Training employs a multi-task regression loss, predicting continuous values ​​between 0 and 1 for each dimension individually, rather than simple classification, ensuring fine-grained style vector representation.

[0111] After the first summary text is generated, it is input into the style classifier at the sentence level. The text is segmented into sentences (by period, question mark, and exclamation mark), and a style vector is predicted for each sentence separately. Long sentences (>50 characters) can be split into clauses (by commas and semicolons) to prevent style mixing and dilution. The style vectors of all sentences are then weighted and averaged, with the weights determined by the sentence importance (including the weight of sentences with decision / action items × 1.5), ultimately yielding the document-level actual style vector. .

[0112] Actual style vector With the target style vector The similarity is calculated using weighted cosine similarity:

[0113] in, This represents the feature similarity score, with a value range of [0,1]. The closer the value is to 1, the more consistent the generated text style is with the preset target style; if it is lower than the preset threshold (e.g., 0.85), regeneration is triggered. Indicates the first The dynamic weights of each dimension reflect the importance of that dimension in the current scenario. The actual style vector is represented at the th... The dimension value is obtained by the inverse style classifier predicting the generated first summary. The target style vector is represented at the th... The value of the dimension is calculated based on preset conditions such as the reporting recipient and the type of meeting. This value represents the desired intensity of the ideal style.

[0114] The similarity threshold is not a fixed value; multiple levels of dynamic thresholds can be set according to actual application needs.

[0115] When the feature similarity is less than the similarity threshold, the iterative regeneration loop of the first summary text is triggered, and the style transfer operation is re-executed. Specifically, upon the first failure, the top-2 dimensions with the largest deviations (such as conciseness and formality) can be identified, and the weight of these dimensions is adjusted accordingly. If the initial attempt fails, a temporary upgrade is performed, and the vector is regenerated. If the upgrade fails a second time, a backup generation strategy is switched, such as switching from BeamSearch to Sampling mode, introducing higher randomness to escape local optima. If the upgrade fails again, a style fallback mechanism is triggered, bringing the bias dimension values ​​in the target style vector closer to the actual values ​​(e.g., reducing the simplicity from 0.9 to 0.8) to avoid excessive demands that could lead to a collapse in generation quality.

[0116] Set a maximum number of iterations (e.g., 3) and a similarity tolerance threshold (e.g., 0.70). If the maximum number of iterations is reached but the threshold is still not met, but the feature similarity is >0.70, output is allowed with a "style deviation warning" label attached; if the feature similarity is <0.70, manual review is mandatory to prevent low-quality text from being leaked.

[0117] S106. Perform humanized feature injection processing on the first summary text to generate the second summary text.

[0118] In one embodiment, natural interjections, appropriate logical jumps, and reverse consistency mechanisms are introduced into the first summary text to simulate the non-linear expression of human thought and enhance the "human touch" of the text.

[0119] Specifically, a dynamic lexicon of natural tone words can be constructed, including categories such as contrastive tone (e.g., "however," "but"), emphatic tone (e.g., "especially," "the key point is"), summarizing tone (e.g., "generally speaking," "in summary"), and buffering tone (e.g., "if possible," "if it is acceptable"). Furthermore, this dynamic lexicon can adaptively expand according to application scenarios; for example, "to be prudent" is automatically loaded for the financial industry, and "from an engineering perspective" is loaded for the technology industry. Each tone word is assigned a sentiment polarity and formality score to ensure matching with the target style vector.

[0120] The insertion of modal particles follows syntactic dependencies and prosodic rules, rather than being randomly distributed. Specifically, a positional strategy can be employed, prioritizing insertion at the boundaries of main and subordinate clauses (such as after adverbial clauses of reason), at the first sentence of a paragraph (to enhance coherence), and before negative words (to soften the negative tone). Alternatively, a probabilistic model can be used, employing a Bernoulli distribution to control the insertion density, with a global density target of 0.3 particles per 100 words (the baseline value for natural human writing). The specific probability is dynamically adjusted based on the intensity of emotion; for example, the insertion density increases to 0.4 for anxious text and decreases to 0.2 for positive text. If a similar modal particle has already appeared in the preceding text (such as two "however"), the insertion probability of subsequent similar particles is penalized by 50% to avoid repetition.

[0121] A moderate logical jump mechanism can identify jumpable windows using a topic coherence model (based on Sentence-BERT) to enable jumps in scenarios such as omitting transitional reasoning, hard topic switching, and sudden changes in information density. Logical jumps may lead to information loss; therefore, jumping is prohibited for sentences containing key elements such as numerical values, time, and responsible parties. A maximum jump distance is also set: a maximum of 3 sentences can be skipped when semantic similarity > 0.7, and 5 sentences can be skipped when semantic similarity > 0.85, ensuring information integrity.

[0122] The reverse consistency mechanism simulates the human thought process of "saying something wrong first and then correcting it." It introduces minor contradictions through a controlled random process. For example, it can be triggered before a key decision conclusion, inserting a slightly contradictory statement such as, "Initially, the plan appears to be costly (approximately 3 million)," followed immediately by a correction: "However, upon verification, the actual budget is 2 million," finally concluding, "Therefore, it is recommended to approve the plan." The contradiction must be minor and correctable to avoid major factual errors.

[0123] In one embodiment, affective computing technology can be combined to perform multi-dimensional affective analysis and dynamic adjustment on the generated text. First, an NLP model is used to identify the affective polarity, intensity, and affective keywords in the text, and affective mapping relationships are established by combining upstream speech affective analysis results (such as affective tags like anxiety and excitement of the speaker). Then, through an affective-style joint optimization mechanism, the selection of interjections (such as adding buffering expressions like "it is worth noting" and "but caution is needed"), sentence structure (such as converting rigid statements into compound sentences with affective tendencies), and punctuation distribution (appropriately using ellipses and exclamation marks to simulate human hesitation or emphasis) are dynamically adjusted during the text generation stage to match the affective tone with the speaker's true intention and the context of the meeting. Finally, the naturalness of the adjusted text is evaluated through a perplexity model to ensure that affective injection conforms to human expression habits without compromising the accuracy of information, thereby significantly reducing the mechanical feel of AI-generated text and improving the readability and decision-making reference value of the minutes.

[0124] This embodiment uses AI-de-AI enhancement technology to transform AI-generated regular text into a second summary text with real human thought characteristics, rhythm, and credibility, while maintaining the accuracy of information. The generated text is closer to real human expression, improving the acceptance of high-level reports.

[0125] S107. Fill the second minutes text into the target meeting minutes template to generate a structured target minutes text.

[0126] In one embodiment, a template engine and dynamic content mapping can be used to automatically populate the second minutes text, which has been injected with humanized features, into the enterprise's standardized meeting minutes template, generating a target minutes text with a complete structure, unified format, and that can be directly used for archiving and distribution.

[0127] Specifically, a configurable template metamodel is defined, which supports a multi-level nested structure and includes core modules such as meeting header information (topic, time, location, host, list of attendees / absentees), agenda discussion area (discussion summaries, core viewpoints, and decision items divided by topic), decision list (decision content, decision-maker, decision basis, and voting results), action tracking table (task description, person in charge, cooperating departments, deadline, priority, and status flag), attachment index area (automatically extracted document links and image references), and remarks column (risk warnings, explanations of disagreements, and unresolved issues). Each module supports conditional rendering (such as automatically hiding the section when there are no decision items) and loop rendering (traversing multiple action items), and uses placeholder syntax to dynamically bind with the structured data in the second minutes text.

[0128] The population process is driven by a template engine. Upon receiving the second summary text, the engine first performs context binding, mapping JSON fields in the text (such as speaker_id, core_summary, decision_id, task_content, etc.) to the namespace of template variables. Simultaneously, it calls the enterprise organizational structure API and knowledge graph service to automatically complete speaker information (name, department, job title), project background data (budget code, contract number), and terminology definitions, ensuring the completeness and accuracy of the populated content. For dynamically generated content blocks (such as decision lists), the engine employs a paragraph segmentation strategy, dividing the second summary text into independent paragraphs based on semantic boundaries (such as timestamp intervals > 30 seconds or speaker changes). Each paragraph is mapped to a Discussion Block object, containing speaker identifiers, core viewpoints, sentiment tags, and related decisions. These are then automatically rendered in a loop within the template as multi-segment discussion summaries, avoiding the readability degradation caused by simply piling up text.

[0129] The final generated target minutes text is automatically pushed to the enterprise OA system, email server or collaboration platform, and can be automatically attached to the approval flow and automatically entered into the task according to the task requirements, realizing a closed loop from meeting discussion to task execution and improving the efficiency of meeting results implementation.

[0130] This embodiment provides an intelligent meeting minutes generation method. The method utilizes a multi-channel microphone array for audio preprocessing, ensuring that the voice of each speaker in the audio stream is clearly captured and accurately identified. The audio with speaker identification is input into a speech recognition model and converted into structured text with timestamps, preserving the original speaking order and time information, thus improving information traceability and processing efficiency. Deep semantic analysis of the initial minutes text is performed using a large language model to extract key text information, enabling the minutes to accurately reflect the core issues and decision points of the meeting. A rule engine and the large language model are combined to jointly judge decision content and task elements, further ensuring the accuracy and executability of decision and action information. Style transfer transforms the initial minutes text into a first minutes text conforming to a preset text style, making the generated text more professional and readable. Humanized feature injection makes the second minutes text more natural and closer to human expression habits, enhancing the naturalness and credibility of the text. The second minutes text is then filled into the target meeting minutes template to generate a structured target minutes text, ensuring the consistency and standardization of the output.

[0131] Please see Figure 3 , Figure 3 This is a schematic diagram of the current embodiment of an intelligent meeting minutes generation device provided in this application. The intelligent meeting minutes generation device is used to execute the aforementioned intelligent meeting minutes generation method.

[0132] like Figure 3 As shown, the intelligent meeting minutes generation device 200 includes: an audio preprocessing module 201, a speech conversion module 202, a text semantic parsing module 203, a joint judgment module 204, a style transfer module 205, a personalized feature injection module 206, and a structured text generation module 207.

[0133] The audio preprocessing module 201 is used to perform audio preprocessing on multi-channel audio acquired by a multi-channel microphone array to generate an audio stream with speaker identifiers. The speech conversion module 202 is used to input the audio stream with speaker identifiers into the speech recognition model, and convert the audio stream into structured text with timestamps through the speech recognition model to obtain the initial summary text of at least one speaker; The text semantic parsing module 203 is used to perform semantic analysis on the initial summary text corresponding to each speaker through a large language model and extract key text information. The joint judgment module 204 is used to perform joint judgment on the decision content and task elements in the initial minutes text based on the key text information, through a rule engine and a large language model, to generate decision and action information. Style transfer module 205 is used to perform style transfer on the initial minutes text based on the key text information and the decision and action information, and convert the initial minutes text into a first minutes text with a preset text style; Personalized feature injection module 206 is used to perform humanized feature injection processing on the first summary text to generate the second summary text; The structured text generation module 207 is used to fill the second minutes text into the target meeting minutes template to generate the structured target minutes text.

[0134] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the device and each module described above can be referred to the corresponding processes in the aforementioned embodiments of the intelligent meeting minutes generation method, and will not be repeated here.

[0135] The apparatus provided in the above embodiments can be implemented as a computer program, which can be used in, for example... Figure 4 It runs on the computer device shown.

[0136] Please see Figure 4 , Figure 4 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.

[0137] See Figure 4 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0138] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any intelligent meeting minutes generation method.

[0139] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0140] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When these computer programs are executed by a processor, the processor can execute any intelligent method for generating meeting minutes.

[0141] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 4The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0142] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0143] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: By preprocessing the multi-channel audio collected by a multi-channel microphone array, an audio stream with speaker identifiers is generated. The audio stream with speaker identifiers is input into a speech recognition model, which converts the audio stream into structured text with timestamps to obtain the initial summary text of at least one speaker. Using a large language model, semantic analysis is performed on the initial summary texts corresponding to each speaker to extract key textual information; Based on the key text information, the decision content and task elements in the initial minutes text are jointly judged by the rule engine and the large language model to generate decision and action information. Based on the key text information and the decision and action information, the initial minutes text is style-transferred to convert the initial minutes text into a first minutes text with a preset text style. The first summary text is processed by injecting human-like features to generate the second summary text; Fill the target meeting minutes template with the second minutes text to generate the structured target minutes text.

[0144] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the intelligent meeting minutes generation methods provided in the embodiments of this application.

[0145] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard equipped on the computer device.

[0146] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating intelligent meeting minutes, characterized in that, The method includes: By preprocessing the multi-channel audio collected by a multi-channel microphone array, an audio stream with speaker identifiers is generated. The audio stream with speaker identifiers is input into a speech recognition model, which converts the audio stream into structured text with timestamps to obtain the initial summary text of at least one speaker. Using a large language model, semantic analysis is performed on the initial summary texts corresponding to each speaker to extract key textual information; Based on the key text information, the decision content and task elements in the initial minutes text are jointly judged by the rule engine and the large language model to generate decision and action information. Based on the key text information and the decision and action information, the initial minutes text is style-transferred to convert the initial minutes text into a first minutes text with a preset text style. The first summary text is processed by injecting human-like features to generate the second summary text; Fill the target meeting minutes template with the second minutes text to generate the structured target minutes text.

2. The intelligent meeting minutes generation method according to claim 1, characterized in that, The process involves using a large language model to perform semantic analysis on the initial summary text corresponding to each speaker, extracting key textual information, including: Based on the speaker identifiers corresponding to each audio stream, the initial summary text corresponding to the same speaker identifier is input into the large language model; Using a large language model, the tasks of summarizing core viewpoints, extracting keywords, and classifying topic tags are performed sequentially to obtain the initial key text. The initial key text is ambiguously resolved through a dialogue state tracking mechanism to obtain the key text information.

3. The intelligent meeting minutes generation method according to claim 1, characterized in that, Based on the key text information, the decision-making content and task elements in the initial summary text are jointly judged through a rule engine and a large language model to generate decision and action information, including: Based on preset statement filtering rules, the key text information is filtered by the rule engine to obtain target statements and candidate statements; Based on the large language model, semantic reasoning and confidence scoring are performed on the candidate statements selected by the rule engine, and the binary classification results corresponding to the candidate statements are output. A joint decision-making mechanism is adopted to vote on the target statement, the candidate statements, and the binary classification results corresponding to the candidate statements to obtain decision information. The decision and action information is obtained by extracting action item information that matches the decision information from the key text information.

4. The intelligent meeting minutes generation method according to claim 3, characterized in that, Based on the large language model, the process involves semantic reasoning and confidence scoring of the candidate statements selected by the rule engine, outputting binary classification results corresponding to the candidate statements, including: Using the large language model, the semantic features and contextual features of the candidate statements are analyzed and reasoned to calculate the confidence score corresponding to the candidate statements. When the confidence score is less than or equal to a preset confidence threshold, the binary classification result corresponding to the candidate statement is determined to be an invalid decision. When the confidence score is greater than the confidence threshold, the binary classification result corresponding to the candidate statement is determined to be a valid decision, and the candidate statement is marked as the target statement.

5. The intelligent meeting minutes generation method according to claim 1, characterized in that, The step of performing style transfer on the initial minutes text based on the key text information and the decision and action information, converting the initial minutes text into a first minutes text with a preset text style, includes: Based on preset text style features, the preset text style is modeled, and the target style vector is calculated; Using the target style vector, the key text information, and the decision and action information as style constraints, and leveraging the cross-attention mechanism between text style and text content, multi-level style transfer is performed on the initial summary text to generate the first summary text.

6. The intelligent meeting minutes generation method according to claim 5, characterized in that, After using the target style vector, the key text information, and the decision and action information as style constraints, and employing a cross-attention mechanism between text style and text content to perform multi-level style transfer on the initial summary text to generate the first summary text, the process further includes: Based on the style classifier, reverse style prediction is performed on the first summary text to obtain the actual style vector of the first summary text; Calculate the feature similarity between the actual style vector and the target style vector; When the feature similarity is less than a preset similarity threshold, the initial summary text is subjected to style transfer operation again until the feature similarity between the actual style vector of the generated first summary text and the target style vector is greater than or equal to the similarity threshold.

7. The intelligent meeting minutes generation method according to claim 1, characterized in that, The step of preprocessing multi-channel audio acquired by a multi-channel microphone array to generate an audio stream with speaker identifiers includes: A deep learning-based multi-channel speech enhancement model is used to remove background noise from the multi-channel audio to obtain noise-reduced audio. An adaptive filter algorithm is used to eliminate echo noise in the noise reduction frequency to obtain echo-cancelled audio. Based on voiceprint recognition and clustering algorithms, the echo-cancelled audio is separated into audio streams, and the speaker identifier of each separated audio stream is labeled to obtain the audio stream with the speaker identifier.

8. An intelligent meeting minutes generation device, characterized in that, The intelligent meeting minutes generation device includes: The audio preprocessing module is used to preprocess the multi-channel audio acquired by the multi-channel microphone array to generate an audio stream with speaker identification. The speech conversion module is used to input the audio stream with speaker identifiers into the speech recognition model, and convert the audio stream into structured text with timestamps through the speech recognition model to obtain the initial summary text of at least one speaker; The text semantic parsing module is used to perform semantic analysis on the initial summary text corresponding to each speaker using a large language model, and to extract key text information. The joint judgment module is used to jointly judge the decision content and task elements in the initial minutes text based on the key text information, through a rule engine and a large language model, to generate decision and action information. The style transfer module is used to perform style transfer on the initial summary text based on the key text information and the decision and action information, and to convert the initial summary text into a first summary text with a preset text style; A personalized feature injection module is used to perform humanized feature injection processing on the first summary text to generate the second summary text. The structured text generation module is used to fill the second minutes text into the target meeting minutes template to generate the structured target minutes text.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the intelligent meeting minutes generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the intelligent meeting minutes generation method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Voice processing method and system for multi-site parallel speech

    CN122135727A