Conference summary method based on AI large model
By combining speaker separation and speech transcription algorithms with a large language model, the confusion problem of transcribed text in multi-person conversation scenarios is solved, and efficient and clear meeting minutes summaries and customized content distribution are achieved.
Patent Information
- Application Number
- CN202510832236.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing large models easily confuse speakers in multi-person conversation scenarios, resulting in messy and poorly readable transcribed text information, making it difficult to use directly for analysis and decision-making.
A speaker separation algorithm is used to identify speakers and split audio segments. A speech transcription algorithm and a large language model are combined for text transcription and summarization. A speech detection algorithm is used to remove redundant time periods, dynamically limit the length of audio segments, and use a punctuation completion algorithm and multi-round iterative merging and summarization technology.
It improves the resolution and clarity of audio content, optimizes the clarity and readability of transcribed text, and is suitable for efficient meeting minutes summarization and multi-concurrency scenarios, supporting multi-level customized summary content distribution.
Smart Images

Figure CN120708619A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and in particular to a method for summarizing meeting minutes based on an AI large model. Background Art
[0002] Current research on large-scale models for transcribing meeting minutes focuses primarily on the integration of natural language processing and speech recognition. This involves automatically transcribing meeting audio into text, extracting key information from the transcript, and generating indirect meeting summaries. This reduces the workload of manual recording, improves the efficiency and accuracy of meeting minutes, and helps users quickly access and understand key information.
[0003] However, current large models still face challenges in transcribing meeting minutes. For example, in scenarios where multiple people are talking, it is easy to confuse the speakers, resulting in the output transcript being unable to clearly identify the source of the sound (i.e., the speaker). As a result, the final transcript information is messy, unreadable, and difficult to use directly for analysis and decision-making, so there is room for improvement. Summary of the Invention
[0004] In order to optimize the information clarity and text readability of the transcribed text, this application provides a meeting minutes summary method based on an AI big model.
[0005] First, this application provides a meeting minutes summary method based on an AI big model, which adopts the following technical solutions: Acquire audio information, identify all speakers involved in the audio information using a preset speaker separation algorithm, and split the audio information into a plurality of audio segments according to the speakers; Using the preset speech transcription algorithm, each audio segment is analyzed and processed to generate a transcribed text; Based on a preset large language model, the transcribed text is refined and summarized, and summary content for describing the audio information is generated and output so that the user can know the summary content.
[0006] By adopting the above technical solution, the speaker separation algorithm is used to distinguish the speakers in the audio information, and the audio information is split according to the speakers, thereby establishing a correspondence between the speakers and the audio segments, thereby improving the resolution and clarity of the audio content. Then, the speech transcription algorithm and the large language model are used to perform text transcription and text summary operations on the audio segments respectively, and finally the summary content used to describe the audio information is extracted, thereby achieving efficient summary processing of the meeting minutes and optimizing the information clarity and text readability of the transcribed text.
[0007] Optionally, the step of using a preset speaker separation algorithm to identify all speakers involved in the audio information further includes: A preset voice detection algorithm is used to detect and remove redundant time periods in the audio information; wherein the redundant time periods refer to audio time periods in the audio information that do not contain voice content.
[0008] By adopting the above-mentioned technical solution, before identifying and distinguishing the speaker from the audio information, this application will pre-process the audio information using a speech detection algorithm to eliminate audio periods (i.e., redundant periods) in the audio information that do not contain specific speech content, thereby reducing the subsequent time spent on reasoning about the audio information and facilitating efficient processing of the audio information.
[0009] Optionally, the method of using a preset speech transcription algorithm to analyze and process each audio segment separately to generate a transcribed text includes: Each audio segment is input into the preset SR algorithm and PR algorithm for recognition, and forced alignment technology is used to generate a transcript with accurate timestamps; The method utilizes a preset speech transcription algorithm to analyze and process each audio segment to generate a transcribed text, and also includes: When the audio information is split into a number of audio segments according to the speaker, the length of the audio segment is dynamically limited so that the length of the audio segment matches the maximum input of the SR algorithm and the PR algorithm.
[0010] By adopting the above technical solution, when splitting the audio segments, the length of the audio segments is dynamically limited to not exceed the maximum input length of the SR algorithm and the PR algorithm, and then the SR algorithm and the PR algorithm are used to respectively perform recognition processing on each audio segment, and a transcribed text with an accurate timestamp is generated by the forced alignment algorithm, that is, the joint forced alignment of the SR algorithm and the PR algorithm is used to improve the timestamp progress, so that the summary content finally obtained in this application can be applicable to derivative needs such as speech retrieval and pronunciation evaluation.
[0011] Optionally, the method further includes: After generating the audio segments, an execution model is matched for each audio segment; all the execution models are used to synchronously execute the following operations on the corresponding audio segments: generating transcripts and refining and summarizing the transcripts; and each execution model also includes two parallel sub-models, which are used to perform SR and PR recognition operations in parallel; and the outputs of the two parallel sub-models generate precise timestamps through a forced alignment algorithm.
[0012] By adopting the above technical solution, an execution model is configured for each audio segment, and all execution models execute corresponding operations synchronously to achieve first-level parallelism (i.e., cross-segment parallelism), and each execution model has two embedded parallel sub-models to achieve parallel SR and PR operations, that is, to achieve second-level parallelism (i.e., intra-segment parallelism). This two-level parallel architecture can further reduce the processing time of audio information and significantly improve the audio processing efficiency, so that the meeting minutes summary method of this application can be applied to high-concurrency scenarios.
[0013] Optionally, the refining and summarizing of the transcribed text based on a preset large language model may also include: Whenever a transcribed text is generated, the transcribed text is padded with punctuation marks using a preset punctuation completion algorithm.
[0014] By adopting the above technical solution, the transcribed text often lacks punctuation marks, which affects the readability of the transcribed text and the accuracy of the subsequent large language model's summary of the transcribed text. Therefore, before refining and summarizing the transcribed text, a punctuation completion algorithm is used to complete the transcribed text with punctuation marks (such as periods, question marks, exclamation marks, line breaks, etc.) to optimize and improve the text structure of the transcribed text.
[0015] Optionally, the refining and summarizing of the transcribed text based on a preset large language model may also include: Based on a preset segmentation basis, all transcribed texts are segmented to generate a number of transcribed texts in the form of paragraphs; wherein the preset segmentation basis at least includes paragraph segmentation based on punctuation marks.
[0016] By adopting the above technical solution, the transcribed text is further segmented to obtain several paragraph-shaped transcribed texts. For example, punctuation marks are used as natural boundaries for paragraph segmentation, so that the long transcribed text is segmented into transcribed texts with less text content and logical coherence. This facilitates the subsequent more refined refinement and summary of the transcribed text based on the large language model, avoids missing key summary content, and optimizes the accuracy and comprehensiveness of the final summary content.
[0017] Optionally, refining and summarizing the transcribed text based on a preset large language model to generate and output summary content for describing the audio information includes: Using the preset large language model, all transcribed texts are iteratively merged and refined and summarized for multiple rounds until the number of rounds of iterative merging and re-refining and summarizing reaches the preset hyperparameters, and then a summary content describing the audio information is generated and output.
[0018] By adopting the above technical solution, combined with the foregoing, it can be seen that the present application divides the transcribed text into several paragraph forms, so that it can cooperate with the large language model to perform multiple rounds of summarization on the transcribed texts in all paragraph forms (that is, all transcribed texts are independently input into the large language model to generate the initial summary content, and then the initial summary content is grouped and merged, and then the grouped and merged contents are respectively input into the large language model for secondary refinement and summary to generate secondary summary content, and then the secondary summary content is further grouped and merged, and each group of merged contents is input into the large language model for refinement and summary, and so on, until N summary contents consistent with the hyperparameters (such as N) are generated, that is, multiple rounds of iterative merging and re-refining and summarizing are completed), and the content output from the last round is integrated to form the summary content; in summary, the above solution solves the information overload problem in the summary content of long text through the dynamic compression + dynamic iteration solution, avoids the loss of key content in the audio information, and ensures that the key content is retained in the final summary content, and compared with the single refinement and summary, the iterative method of the present application can also effectively reduce the CPU memory usage.
[0019] Secondly, this application provides a meeting minutes summary system based on an AI big model, which adopts the following technical solutions: An audio preprocessing and segmentation module is used to obtain audio information, identify all speakers involved in the audio information using a preset speaker separation algorithm, and split the audio information into several audio segments based on the speakers; The audio multimodal recognition and transcription module is used to analyze and process each audio segment using a preset speech transcription algorithm to generate a transcribed text; The audio transcription text extraction and summarization module is used to extract and summarize the transcription text based on a preset large language model, generate and output summary content for describing the audio information, so that the user can know the summary content.
[0020] In a third aspect, the present application provides a meeting minutes summarizing device based on an AI big model, comprising a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute any of the methods described in the first aspect.
[0021] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and execute any of the methods described in the first aspect.
[0022] In summary, this application includes at least one of the following beneficial technical effects: 1. In this application, a voice detection algorithm (VAD algorithm) is used to remove redundant segments from audio information, thereby reducing the time required to process the audio information. A speaker separation algorithm (SD algorithm) is used to distinguish different speakers in the audio information and split the audio information according to the speakers, thereby improving the resolution and clarity of the audio content. A voice transcription algorithm (SR algorithm + PR algorithm) and a large language model are then used to transcribe and summarize the audio segments in sequence, ultimately extracting summary content that describes the audio information. This allows for efficient summary processing of meeting minutes and optimizes the clarity and readability of the transcribed text. 2. Furthermore, a punctuation completion algorithm is used to restore the punctuation of the transcribed text and improve the text structure of the transcribed text; the transcribed text is segmented into paragraphs based on punctuation marks, and the long text is refined into several short and logically coherent texts to facilitate subsequent accurate extraction and summary. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 It is a flow chart of the meeting minutes summarizing method based on the AI big model disclosed in the embodiment of this application.
[0025] Figure 2 This is a structural block diagram of the meeting minutes summary system based on the AI big model disclosed in the embodiment of this application.
[0026] Explanation of the accompanying drawings: 201, audio preprocessing and segmentation module; 202, audio multimodal recognition and transcription module; 203, audio transcription text extraction and summary module. DETAILED DESCRIPTION
[0027] The following is combined with Figure 1-2 This application is described in further detail.
[0028] The embodiment of this application discloses a method for summarizing meeting minutes based on an AI big model (hereinafter referred to as the summarizing method), which aims to perform clear speech recognition and text-readable summaries on conference audio information. The execution subject of the summarizing method is a meeting minutes summarizing system based on an AI big model (hereinafter referred to as the summarizing system). Figure 1 The specific execution steps of the summary method by the summary system are elaborated in detail.
[0029] S101, obtaining audio information, using a preset speaker separation algorithm to identify all speakers involved in the audio information, and splitting the audio information into several audio segments according to the speakers; Optionally, before using a preset speaker separation algorithm to identify all speakers involved in the audio information, the following steps are further included: A preset speech detection algorithm is used to detect and remove redundant time periods in the audio information; wherein, redundant time periods refer to audio time periods in the audio information that do not contain speech content.
[0030] In implementation, for example, a user can access the summary system in the form of a URL and upload audio information based on the upload window on the preset access page of the summary system. In the embodiment of the present application, the audio information refers to the audio content recorded during the meeting. The summary system is used to obtain the audio information uploaded by the user, and then use a preset voice detection algorithm (such as a neural network-based VAD model (specifically, it can be a WebRTC or TEN VAD algorithm, Silero VAD algorithm)) to detect the voice period in the audio information. The voice period refers to the period in the audio information when the speaker speaks, and the corresponding redundant period is the period when no speaker speaks. The summary system is used to propose the redundant period in the audio information to achieve the update of the audio information, that is, only retain the voice period, so as to reduce the time spent on subsequent processing of the audio information.
[0031] Next, the summary system uses a speaker separation algorithm (such as the SD algorithm (specifically, DeepClustering or TasNet)) to identify and distinguish the different speakers in the updated audio information. This determines all speakers involved in the audio information and then uses the speaker as the basis for segmentation, dividing the audio information into audio segments corresponding to different speakers. Each audio segment is labeled with a speaker to establish a correspondence between speakers and audio segments. This solution allows for the differentiation of speakers and their speech content (i.e., the corresponding audio segments).
[0032] S102, using a preset speech transcription algorithm, analyzing and processing each audio segment to generate a transcribed text; S102 specifically includes the following sub-steps: Each audio segment is input into the preset SR algorithm and PR algorithm for recognition, and forced alignment technology is used to generate a transcript with accurate timestamps; The following steps are also included before S102: When splitting the audio information into several audio segments according to the speaker, the length of the audio segment is dynamically limited so that the length of the audio segment matches the maximum input of the SR algorithm and the PR algorithm.
[0033] In implementation, the preset speech transcription algorithm includes a preset SR algorithm and a PR algorithm. The SR (Speech Recognition) algorithm is used to convert the audio segment into text content (i.e., the transcribed text mentioned above), while the PR (Phoneme Recognition) algorithm is used to identify the phonemes of the audio segment, analyze the pronunciation details, and output the phoneme sequence. Accordingly, the SR algorithm can adopt an end-to-end model (such as Whisper), and the PR algorithm can adopt a phoneme-level CTC model. The summary system is used to input each audio segment into the SR algorithm and the PR algorithm respectively, so that the SR algorithm and the PR algorithm perform the above-mentioned processing on the same audio segment respectively. The summary system is then used to use forced alignment technology (such as Montreal Forced Aligner) to perform timestamp alignment processing on the content output by the SR algorithm and the PR algorithm, and finally generate a transcribed text with a precise timestamp. It is assumed in this application that the timestamp after the forced alignment processing is the precise timestamp described herein.
[0034] S103: Based on a preset large language model, the transcribed text is refined and summarized, and summary content for describing the audio information is generated and output for the user to obtain the summary content.
[0035] In addition, the following steps are also included before S103: Whenever a transcript is generated, the punctuation marks are filled in using a preset punctuation completion algorithm.
[0036] Based on a preset segmentation basis, all the transcribed texts are segmented to generate a plurality of transcribed texts in the form of paragraphs; wherein the preset segmentation basis at least includes paragraph segmentation based on punctuation marks; And S103 specifically includes the following sub-steps: Using the preset large language model, all transcribed texts are iteratively merged and refined and summarized for multiple rounds until the number of rounds of iterative merging and re-refining and summarizing reaches the preset hyperparameters, and then a summary content describing the audio information is generated and output.
[0037] In implementation, after all audio segments are converted into transcribed texts with precise timestamps using a preset speech transcription algorithm, the summary system is used to fill in the punctuation marks on each transcribed text using a preset punctuation completion algorithm to improve the readability of the transcribed text. Specifically, a punctuation prediction model (such as one based on BERT or bidirectional LSTM) can be used to insert punctuation marks (such as periods, commas, semicolons, line breaks, etc.) using the "SEP" tag. Then, the summary system is used to further segment each transcribed text based on a preset segmentation basis, so that each transcribed text is further divided into several paragraphs. The preset segmentation basis disclosed in the embodiment of the present application is paragraph segmentation based on punctuation marks. If two periods appear consecutively or whenever a line break appears, the text is divided into paragraphs, and each paragraph is then treated as a separate transcribed text to update the transcribed text.
[0038] After completing the above processing, the summary system is used to input all the transcribed texts into a preset large language model (such as an LLM model) for multiple rounds of iterative merging and re-refining summarization. The specific implementation method is: first, all the transcribed texts are input into the LLM model respectively, and the initial summary content is obtained for each transcribed text by the LLM model. Then, based on the preset merging rules, the initial summary content is grouped, and the initial summary content of the same group is merged to obtain new text content. Then, all the new text contents are input into the LLM model respectively for refining and summarizing to obtain secondary summary content. And so on to implement iteration, group the i-th summary content, merge the same group to obtain new text content, and then input the new text content into the LLM model for refining and summarizing to obtain the i+1-th summary content, until the N-th summary content is finally obtained and the iteration is stopped. Here, N can be considered as a preset hyperparameter, 1≤i+1≤N.
[0039] For example, the preset merging rule can be based on text semantic similarity (such as TF-IDF or BERT embedding vector clustering) or topic consistency (such as the LDA topic model), merging similar summaries into a group to form new text content, which is then input into the LLM model. Specifically, the semantic similarity between the summaries can be calculated, and then the cosine distance of the text embeddings can be calculated using a pre-trained language model (such as Sentence-BERT). If the similarity exceeds a preset threshold, the content is considered similar and can be grouped together for merging.
[0040] After completing the iteration and re-summarization, the summary system is used to merge and integrate the summary content finally output by LLM and output it for display so that the user who uploaded the corresponding audio information can know it.
[0041] Optionally, the summarizing method further includes the following steps: After generating the audio segments, an execution model is matched for each audio segment; all the execution models are used to synchronously execute the following operations on the corresponding audio segments: generating transcripts and refining and summarizing the transcripts; and each execution model also includes two parallel sub-models, which are used to perform SR and PR recognition operations in parallel; and the outputs of the two parallel sub-models generate precise timestamps through a forced alignment algorithm.
[0042] In implementation, in order to improve the processing efficiency of all audio segments, the present application proposes to match the execution model for each split audio segment separately, and all execution models will execute synchronously: generate transcripts and perform multiple rounds of iterative summarization and re-refinement of the transcripts. When performing the operation of generating transcripts, the two parallel sub-models contained in the execution model will synchronously execute the SR and PR recognition operations respectively, that is, one of the parallel sub-models is used to input the audio segment into SR for processing, and the other parallel sub-model is used to input the audio segment into PR for processing, and the two parallel sub-models will perform the aforementioned processing operations at the same time. Finally, the execution model will force the output results of the two parallel sub-models to align to obtain an accurate timestamp. In summary, the synchronous parallel processing of all audio segments is achieved through the parallel operation of the execution model, and the synchronous execution operation of the parallel sub-model is used to achieve synchronous recognition of SR and PR for the same audio segment, thereby reducing processing time.
[0043] Optionally, considering that the summary content of traditional conference audio information is static and cannot adapt to the information needs of audiences with different roles (such as executives and ordinary employees), some sensitive information that can only be viewed by audiences with high-level access permissions (such as executives) is leaked to audiences with lower access permissions (such as ordinary employees); and the only way to rely on manual tagging and filtering of sensitive information in the above summary content is inefficient and prone to omissions. Therefore, the present application further proposes the following solution, and the corresponding S103 specifically includes the following steps: S1031, all the transcribed texts are input into the LLM model respectively, and the LLM is used to extract and summarize each transcribed text to obtain the initial summary content.
[0044] S1032. For each initial summary content, mark the minimum permission level of all semantic units (such as sentences) according to a preset permission level reference library; wherein the permission level reference library includes different preset permission levels (such as public, department, and executive), and keyword groups corresponding to each preset permission level; by extracting keywords from each semantic unit contained in each initial summary content, the preset permission level corresponding to the extracted keywords in the permission level reference library is used as the minimum permission level of the corresponding semantic unit.
[0045] S1033. For all minimum permission levels of all initial summary contents and their corresponding marks, generate an audience type set (such as ordinary employees, department managers, and CEOs) according to a preset audience permission library; wherein the audience permission library includes several preset audience types and the permission level corresponding to each audience type; each audience type included in the audience type set satisfies: the minimum permission level of the mark corresponding to the initial summary content is not higher than the permission level corresponding to the audience type in the audience type set.
[0046] S1034: For each audience type, independently perform multiple rounds of iterative merging and re-summarizing until the number of iterations reaches a preset hyperparameter. The specific steps of independently performing multiple rounds of iterative merging and re-summarizing for each audience type are as follows: Extract all semantic units whose minimum permission level is not higher than the permission level corresponding to the target audience type from all initial summary content to generate a content subset; use the LLM model to perform multiple rounds of iterative merging and re-summarizing on the content subset (which can be combined with the multiple rounds of merging and re-summarizing scheme described above) to generate summary content corresponding to the target audience type; where the target audience type is any audience type in the audience type set.
[0047] S1035: Integrate the summary contents generated corresponding to all audience types in the audience type set to obtain summary contents of a multi-level structure, wherein each level corresponds to an audience type in the type set, and establish an association relationship between the level and the audience type; store the corresponding audio information and its corresponding summary contents of the multi-level structure, and generate an ID number for distinguishing different audio information; S1036: Upon receiving a search request from a viewer, the viewer's identity is verified, the viewer's corresponding audience type is determined, and based on the audio information included in the search request, the summary content corresponding to the viewer's corresponding audience type is retrieved and output from the multi-level summary content corresponding to the audio information, so that the viewer can review the output summary content. It should be noted that by default, the summary system stores the identities (e.g., name, position, etc.) of all users and their corresponding audience types.
[0048] The above technical solution provides customized summary content tailored to each individual, with the type of audience used as the basis for customization, thus preventing ordinary employees from reviewing executive strategic discussions. Furthermore, the specific operation of customized summary content is integrated into the multi-round iteration and summarization process based on LLM. There is no need to reorganize the summary content based on permissions based on the summary content output by LLM. Instead, the issue of reader permissions is taken into account during the process of refining the summary content. Before refining the summary content, all audiences of the current audio information are first determined. The reading permissions of all determined audiences are then used as the basis for LLM's multi-round iterative merging and refining of the summary. The result is a summary content with a multi-level structure. Through permission-driven hierarchical refinement and dynamic audience-content association mapping, multi-granular secure distribution of summary content is achieved.
[0049] The present application also discloses a meeting minutes summary system based on an AI big model. Figure 2 ,include: The audio preprocessing and segmentation module 201 is used to obtain audio information, identify all speakers involved in the audio information using a preset speaker separation algorithm, and split the audio information into several audio segments according to the speakers; The audio multimodal recognition and transcription module 202 is used to analyze and process each audio segment using a preset speech transcription algorithm to generate a transcribed text; The audio transcription text extraction and summarization module 203 is used to extract and summarize the transcription text based on a preset large language model, generate and output summary content for describing the audio information, so that the user can know the summary content.
[0050] Optionally, the audio preprocessing and segmentation module 201 is further configured to detect and remove redundant time periods in the audio information using a preset speech detection algorithm; wherein the redundant time periods refer to audio time periods in the audio information that do not contain speech content.
[0051] Optionally, the audio multimodal recognition and transcription module 202 is also used to dynamically limit the length of the audio segment when splitting the audio information into several audio segments according to the speaker, so that the audio segment length can match the maximum input of the SR algorithm and the PR algorithm; it is also used to input each audio segment into the preset SR algorithm and PR algorithm for recognition, and use forced alignment technology to generate a transcription text with precise timestamps.
[0052] Optionally, it also includes an audio parallel processing module, which is used to match the execution model for each audio segment after generating the audio segment; all the execution models are used to synchronously execute the following operations on the corresponding audio segments: generating transcription text and refining and summarizing the transcription text; and each execution model also includes two parallel sub-models, and the two parallel sub-models are used to perform SR and PR recognition operations in parallel; and the outputs of the two parallel sub-models generate precise timestamps through a forced alignment algorithm.
[0053] Optionally, the audio transcription text extraction and summarization module 203 is also used to segment all transcribed texts based on preset segmentation criteria to generate transcribed texts in the form of several paragraphs; wherein the preset segmentation criteria at least include paragraph segmentation based on punctuation marks; and is also used to fill in the transcribed text with punctuation marks using a preset punctuation completion algorithm whenever a transcribed text is generated.
[0054] Optionally, the audio transcription text refinement and summarization module 203 is also used to use a preset large language model to perform multiple rounds of iterative merging and re-refining and summarizing on all transcription texts until the number of rounds of iterative merging and re-refining and summarizing reaches a preset hyperparameter, and generate and output summary content for describing the audio information.
[0055] An embodiment of the present application also discloses a meeting minutes summarizing device based on an AI big model. The meeting minutes summarizing device based on an AI big model includes a memory and a processor. The memory stores a computer program that can be loaded by the processor and executes the above-mentioned meeting minutes summarizing method based on the AI big model.
[0056] An embodiment of the present application also discloses a computer-readable storage medium, which stores a computer program that can be loaded by a processor and execute the above-mentioned meeting minutes summarizing method based on the AI large model. The computer-readable storage medium includes, for example: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0057] It should be noted that, in this document, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0058] The above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of protection of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on these embodiments, all other embodiments obtained by persons of ordinary skill in the art without inventive effort are also within the scope of protection to be protected by this application.
Claims
1. A meeting minutes summary method based on AI big model, characterized by: include: Acquire audio information, identify all speakers involved in the audio information using a preset speaker separation algorithm, and split the audio information into a plurality of audio segments according to the speakers; Using the preset speech transcription algorithm, each audio segment is analyzed and processed to generate a transcribed text; Based on a preset large language model, the transcribed text is refined and summarized, and summary content for describing the audio information is generated and output so that the user can know the summary content.
2. The method for summarizing meeting minutes based on the AI big model according to claim 1 is characterized in that: The method of using a preset speaker separation algorithm to identify all speakers involved in the audio information may also include: A preset voice detection algorithm is used to detect and remove redundant time periods in the audio information; wherein the redundant time periods refer to audio time periods in the audio information that do not contain voice content.
3. The method for summarizing meeting minutes based on AI big model according to claim 1 is characterized in that: The method of using a preset speech transcription algorithm to analyze and process each audio segment to generate a transcribed text includes: Each audio segment is input into the preset SR algorithm and PR algorithm for recognition, and forced alignment technology is used to generate a transcript with accurate timestamps; The method utilizes a preset speech transcription algorithm to analyze and process each audio segment to generate a transcribed text, and also includes: When the audio information is split into a number of audio segments according to the speaker, the length of the audio segment is dynamically limited so that the length of the audio segment matches the maximum input of the SR algorithm and the PR algorithm.
4. The method for summarizing meeting minutes based on the AI big model according to claim 3 is characterized in that: The method further comprises: After generating the audio segments, an execution model is matched for each audio segment; all the execution models are used to synchronously execute the following operations on the corresponding audio segments: generating transcripts and refining and summarizing the transcripts; and each execution model also includes two parallel sub-models, which are used to perform SR and PR recognition operations in parallel; and the outputs of the two parallel sub-models generate precise timestamps through a forced alignment algorithm.
5. The method for summarizing meeting minutes based on AI big model according to claim 1 is characterized in that: The transcribed text is refined and summarized based on the preset large language model, and the following steps are also included before the transcribed text is refined and summarized: Whenever a transcribed text is generated, the transcribed text is padded with punctuation marks using a preset punctuation completion algorithm.
6. The method for summarizing meeting minutes based on AI big model according to claim 5 is characterized in that: The transcribed text is refined and summarized based on the preset large language model, and the following steps are also included before the transcribed text is refined and summarized: Based on a preset segmentation basis, all transcribed texts are segmented to generate a number of transcribed texts in the form of paragraphs; wherein the preset segmentation basis at least includes paragraph segmentation based on punctuation marks.
7. The method for summarizing meeting minutes based on AI big model according to claim 6 is characterized in that: The method of refining and summarizing the transcribed text based on a preset large language model, and generating and outputting summary content for describing the audio information, includes: Using the preset large language model, all transcribed texts are iteratively merged and refined and summarized for multiple rounds until the number of rounds of iterative merging and re-refining and summarizing reaches the preset hyperparameters, and then a summary content describing the audio information is generated and output.
8. A meeting minutes summary system based on AI big model, characterized by: include, An audio preprocessing and segmentation module (201) is used to obtain audio information, identify all speakers involved in the audio information using a preset speaker separation algorithm, and split the audio information into a number of audio segments according to the speakers; An audio multimodal recognition and transcription module (202) is used to analyze and process each audio segment using a preset speech transcription algorithm to generate a transcribed text; The audio transcription text extraction and summarization module (203) is used to extract and summarize the transcription text based on a preset large language model, generate and output summary content for describing the audio information, so that the user can know the summary content.
9. A meeting minutes summary device based on AI big model, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Intelligent summary automatic generation method and system for recruitment communication scene
CN121234881A
Intelligent summary automatic generation method and system for recruitment communication scenarios
CN121234881B
LLM enhancement-based multi-speaker voice recognition and voiceprint matching system
CN121306145A