Model fine tuning and template prompt reasoning method for meeting summary generation
Through multi-dimensional meeting minutes sample data quality evaluation and model fine-tuning, combined with template prompt reasoning methods, the problems of low readability and difficulty in extracting key information in the existing technology are solved, and high-quality meeting minutes are achieved.
Patent Information
- Application Number
- CN202411910593.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-06-06
AI Technical Summary
The existing meeting minutes generation system cannot deeply structure key information, and the generated minutes are low readability, making it difficult to efficiently extract key points of the meeting.
The multi-dimensional meeting minutes sample data quality evaluation method is used to screen high-quality fine-tuning samples, and high-quality meeting minutes are generated through model fine-tuning, and template prompt reasoning methods are designed to guide the model to generate minutes that meet the format and content constraints.
It significantly improves the quality of conference minutes generation, improves the readability and user understanding efficiency of conference minutes.
Smart Images

Figure CN120106022A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of automatic generation of meeting minutes based on artificial intelligence-natural language processing, and in particular to a model fine-tuning and template prompt reasoning method for meeting minutes generation. Background Art
[0002] In recent years, natural language processing (NLP) technology has achieved rapid development, especially in the fields of automated office and information management, and the application of intelligent systems has gradually increased. Among them, the generation of meeting minutes has always been an important part of the office work of enterprises and organizations. Existing meeting minutes generation systems, such as Tencent's meeting minutes generation function, although they have certain capabilities in generating meeting minutes, still have some shortcomings. For example, the generated minutes cannot be deeply structured, there is no unified display of key information, and a large amount of original recorded text is prone to lack of thematic summary of the text. As a result, the readability of the meeting minutes is low, and users still need to spend a lot of time to understand the key points of the meeting.
[0003] With the rapid development of information technology and Internet technology, the frequency of meetings has increased significantly. However, the long duration and complex content of meetings make it difficult to extract key information efficiently. In addition, different meeting minutes have specific output formats. These are still issues that need to be urgently addressed in order to improve the quality of meeting minutes generation. Summary of the invention
[0004] The present application provides a model fine-tuning and template prompt reasoning method for meeting minutes generation, which adopts a multi-dimensional meeting minutes sample data quality assessment method to screen out high-quality fine-tuning samples to improve the effectiveness of the fine-tuning model; based on the prompt template designed for model reasoning, the model is guided to generate minutes that meet the format and content constraints as required, thereby effectively improving the generation quality of meeting minutes.
[0005] The technical solution adopted in this application is:
[0006] A model fine-tuning and template prompting reasoning method for meeting minutes generation, the method comprising the following steps:
[0007] Step 1: Input the meeting minutes dataset to construct a fine-tuning sample dataset;
[0008] Among them, the meeting minutes dataset includes the original meeting record file, meeting minutes file, meeting record time, meeting topic and meeting attribute information (such as meeting number, meeting document number), etc.
[0009] Step 2: Perform a meeting minutes quality assessment on the fine-tuning sample data set to screen out fine-tuning sample data that meets the quality assessment requirements, perform supervised fine-tuning on the large language model used to generate meeting minutes, and generate a fine-tuned large language model;
[0010] The evaluation items for the quality assessment of meeting minutes include:
[0011] Content similarity measures the semantic consistency between the original meeting record file and the meeting minutes file;
[0012] Entity coverage, which measures the proportion of entities in the meeting minutes document to those in the original meeting record document;
[0013] Text fluency, assessing the grammatical correctness and fluency of the meeting minutes document;
[0014] Step 3: Based on the set inference template of the large language model, guide the fine-tuned large language model to generate meeting minutes that meet the format and content constraints as required;
[0015] The inference template of the large language model set includes original text and feature information, output format and output examples.
[0016] Furthermore, in step 2, the content similarity is set to:
[0017] The original meeting record file and the meeting minutes file are processed separately to obtain two sets with sentences as elements;
[0018] Based on the sentence embedding representation acquisition model, the embedding representation of each sentence in each set is extracted to obtain the sentence vector of each sentence. Define c j Represents any sentence vector of the original conference record file, and defines g i An arbitrary sentence vector representing a meeting minutes document;
[0019] For each sentence vector in the meeting minutes file, calculate the sentence similarity between it and each sentence vector in the original meeting record file, and use the maximum value as the content similarity evaluation value of the current sentence vector in the meeting minutes file;
[0020] Then, based on the average of the content similarity evaluation values of all sentence vectors in the meeting minutes file, the content similarity measurement value between the original meeting record file and the meeting minutes file is obtained.
[0021] Furthermore, the calculation formula for sentence similarity is:
[0022] Furthermore, we construct a similarity matrix with the sentence index i of the meeting minutes file as the column and the sentence index j of the original record file as the column. Each element of the matrix is used to represent any sentence vector g i With c j The similarity of sentences between
[0023] The maximum value of each row of the similarity matrix is averaged to obtain the content similarity measurement value between the original meeting record file and the meeting minutes file.
[0024] Furthermore, the entity coverage is set as:
[0025] Extract the entity original_entity of the original meeting record file i and the frequency of the corresponding entity original_frequency i , forming a dictionary {original_entity i :original_frequency i (i∈[1~n])};
[0026] Extract the summary_entity entity of the meeting minutes file j , forming an entity set {summary_entity j (j∈[1~m])};
[0027] The entity coverage between the original meeting record file and the meeting minutes file is calculated according to the formula:
[0028]
[0029] Among them, subscript i is the entity index of the original meeting record file, n is the entity number of the original meeting record file; subscript j is the entity index of the meeting minutes file, and m is the entity number of the meeting minutes file.
[0030] Furthermore, the text fluency is set to:
[0031] The meeting minutes file is processed into sentences to obtain the text sequence T = {t 1 ,t 2 ,…,t n′}, where t 1 ,t 2 ,…,t n1 represents n′ sentences;
[0032] Input the text sequence T into the pre-trained natural speech processing model based on the Transformers framework, use the loss value output by the model as the perplexity of the text, and get the corresponding perplexity sequence F ′ ={f 1 ,f 2 ,…,f n′};
[0033] The confusion sequence F ′ The mean confusion of is taken as the overall confusion of the meeting minutes document;
[0034] The overall confusion is inverted using an inverse proportional function to obtain the text fluency of the meeting minutes document.
[0035] Furthermore, when obtaining text fluency, the natural speech processing model based on the Transformers framework is the Bert-Base-Uncased model.
[0036] Furthermore, the original text and feature information, output format and output example of the inference template of the large language model are as follows:
[0037] The original text and feature information include: filling rules, original text and feature information;
[0038] Among them, the feature information is extracted from the text content of the original meeting record file using several text feature extraction models;
[0039] The filling rule is to fill the original text, feature information and output format content into the inference template by replacing characters;
[0040] Output format, including meeting minutes framework, generation format and task description.
[0041] Furthermore, the output format adopts a three-paragraph structure. The first paragraph contains basic information about the meeting, such as the meeting name, meeting time, meeting location, organizing unit, participants, etc.; the second paragraph is the meeting theme and summary information of the corresponding theme; the third paragraph contains the meeting results, meeting summary and other scattered content of the meeting.
[0042] Furthermore, the meeting minutes framework includes the minutes name, meeting time, meeting organizer information, participant information, meeting content / topic, several meeting results, other content, and meeting minutes signature information.
[0043] Furthermore, feature information is extracted from the text content of the original conference record file using several text feature extraction models, including:
[0044] Topic feature extraction: The meeting record text in the original meeting record file is segmented to obtain the segmentation set W = (w 1 ,w 2 ,…,w n″ ), where n″ is the number of segmentations in the segmentation set W; and a corpus is created through the unique index of the segmentation set W and the frequency of the segmentation; the corpus is input into the topic extraction model to obtain several topics and the probability value corresponding to each topic, and the maximum probability value is selected as the result of topic extraction;
[0045] Keyword feature extraction: Input the meeting record text in the original meeting record file into the keyword recognition model to obtain the keyword recognition results;
[0046] Core sentence feature extraction: The meeting record text in the original meeting record file is divided into sentences, and then each sentence is segmented, and the words in the same sentence are regarded as a group of words; the feature vector (i.e. TF-IDF vector) of each sentence is calculated according to the term frequency-inverse document frequency algorithm; then the k-means algorithm is used for clustering, and according to the number of clusters, the sentences closest to the cluster center are selected as the core sentences (i.e. representative sentences) of each category.
[0047] The technical solution provided by this application brings at least the following beneficial effects:
[0048] This application first screens out high-quality meeting minutes in the data set based on the proposed multi-dimensional meeting minutes fine-tuning sample data quality assessment method to achieve fine-tuning of the large language model; at the same time, the traditional text processing method is used to extract the feature information of the original meeting records, and a three-part inference template is designed. Finally, the original record text and feature information are filled into the inference template, and the fine-tuned large language model is input to stably generate meeting minutes.
[0049] The method proposed in this application solves the existing problems of automatically generating meeting minutes, such as long meeting time, complex content, and difficulty in efficiently extracting key information;
[0050] This application is designed to help users quickly understand and design a special reasoning template, which includes three parts: original text and feature information, and output format. The original text and feature information also include filling rules, original text, and feature information. The extraction of feature information can be achieved through methods such as TF-IDF+Kmeans, Bert, and Lda, and the output format is completed by defining the meeting minutes framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0052] Figure 1 A flowchart of a model fine-tuning and template prompting reasoning method for generating meeting minutes provided in an embodiment of the present application;
[0053] Figure 2 A detailed flow chart of a model fine-tuning and template prompt reasoning method for generating meeting minutes according to an embodiment of the present application;
[0054] Figure 3 A schematic diagram of a meeting minutes quality assessment framework according to an embodiment of the present application;
[0055] Figure 4 This is a partial schematic diagram of the reasoning template - original text and feature information of an embodiment of the present application;
[0056] Figure 5 This is a partial schematic diagram of the inference template-output format of an embodiment of the present application;
[0057] Figure 6 A schematic diagram of the meeting minutes framework of an embodiment of the present application;
[0058] Figure 7 This is a schematic diagram of the reasoning template-generation format of an embodiment of the present application;
[0059] Figure 8 A schematic diagram of an example of an original record file in an embodiment of the present application;
[0060] Fig. 9 A schematic diagram of a content similarity algorithm according to an embodiment of the present application;
[0061] Fig.10 A schematic diagram of an entity coverage algorithm according to an embodiment of the present application;
[0062] Fig.11 A schematic diagram of a text fluency algorithm according to an embodiment of the present application;
[0063] Fig.12 This is a schematic diagram of characteristic information of an embodiment of the present application. DETAILED DESCRIPTION
[0064] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions of the embodiments of the present application will be described in detail and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described with reference to the drawings are exemplary and are intended to be used to explain the present application, and cannot be understood as limiting the present application.
[0065] This application embodiment proposes a model fine-tuning and template prompt reasoning method for generating meeting minutes. Figure 1 and Figure 2 , which specifically includes the following steps:
[0066] Step S1, constructing a data set, that is, constructing a fine-tuning sample data set;
[0067] Step S2, multi-dimensional evaluation of data samples, including recognition accuracy (entity coverage), text coherence (i.e., text fluency), and topic similarity (i.e., content similarity); based on the weighted fusion results of the three evaluation indicators, high-quality meeting minutes samples that meet the established threshold are screened out, and then the large language model is fine-tuned in a supervised manner based on the screened samples to generate a fine-tuned large language model;
[0068] Step S3, based on the set inference template of the large language model, guide the large language model fine-tuned in step S2 to generate meeting minutes that meet the format and content constraints as required. That is, the target text and feature information are input into the large language model fine-tuned in step S2, and the inference template of the set large language model is used as the output format, and the automatically generated meeting minutes are obtained through the fine-tuned large language model.
[0069] That is, this application first selects high-quality meeting minutes in the data set based on a multi-dimensional meeting minutes fine-tuning sample data quality assessment method, and then uses it to fine-tune the large language model. At the same time, traditional text processing methods (such as TF-IDF+Kmeans, Bert, Lda, etc.) are used to extract the feature information of the original meeting records, and a three-part inference template is designed. Finally, the original record text and feature information are filled into the inference template, and the fine-tuned large language model is input to stably generate meeting minutes.
[0070] Among them, the multi-dimensional meeting minutes sample data quality assessment method scores the meeting minutes as a whole according to the text flow degree, content similarity, entity coverage and corresponding weights, judges the quality of the meeting minutes according to the score threshold, and selects high-quality meeting minutes to fine-tune the large language model. The inference template consists of three parts, namely the original text and feature information, and the output format. The output format is a description of the meeting minutes framework and a complete meeting minutes.
[0071] The quality of meeting minutes generated by large language models varies greatly, and how to evaluate the quality of generated meeting minutes is also a blank in current research. For a high-quality meeting minute, there should be corresponding quality requirements in terms of content similarity, entity coverage, and text fluency. In this embodiment, All-MiniLM-L6-V2, Bert-Base-Uncased and other models can be used to score in multiple dimensions, and the quality of the meeting minutes can be judged according to the threshold of the score.
[0072] See also Figure 3 The multi-dimensional meeting minutes fine-tuning sample data quality assessment method of this application aims to solve the problems of uneven quality of meeting minutes, difficult to quantify subjective evaluations, and lack of unified measurement standards. It quantifies the text fluency, entity coverage, and content similarity of the meeting minutes, and then assigns weights to calculate the scores. Finally, the quality of the meeting minutes is assessed based on the score range. Usually, it can be divided into three levels, that is, the score range is divided into three data sub-ranges, and the evaluation results are obtained based on the data sub-ranges corresponding to the calculated scores: excellent, good, or poor.
[0073] The large model reasoning template for meeting minutes generation is as follows Figure 4 and Figure 5, in order to solve the problems of incoherent generated content, unstable generated format, inaccurate information extraction, etc., the quality of meeting minutes generation is improved by defining reasoning templates combined with traditional methods, such as using Bert-Bilstm-Crf, TF-IDF+Kmeans, Lda and other algorithms to complete the feature extraction of the original text, and the reasoning template adopts the original text and feature information, and the output format is designed in two parts, so as to improve the quality of meeting minutes generation. Figure 5 A schematic diagram of the meeting minutes framework of an embodiment of the present application; Figure 7 This is a partial schematic diagram of the inference template-generation format of an embodiment of the present application.
[0074] In one embodiment, the specific implementation steps of a model fine-tuning and template prompt reasoning method for generating meeting minutes provided in the embodiment of the present application include:
[0075] Step S 00 , build a fine-tuning sample dataset:
[0076] This embodiment constructs a meeting minutes dataset, which contains original meeting record files. Figure 8 This is an example of an original record file. Meeting minutes file, meeting record time, meeting theme, meeting number, and meeting document number. All of the above information is crawled from the official website (https: / / www.un.org / zh / ga / sessions / regular.shtml). The link that constitutes the dataset is: https: / / github.com / lusong12 / xiao-lun-wen, which contains the dataset files and website crawler files for the 56th to 78th sessions. Due to the different webpage UIs for the 56th to 61st, 62nd to 67th, and 68th to 78th sessions, the crawled webpage files are slightly different. Among them, the 56th to 61st sessions were too early, and the webpages did not have information such as the conference theme, which was used in the model inference stage. Table 1 gives the statistical information of the meeting minutes dataset.
[0077] Table 1 Meeting minutes dataset and metadata
[0078]
[0079] It should be noted that the statistical indicators shown in Table 1 are only examples, and in specific applications, they can be expanded and set based on actual application scenarios.
[0080] Step S 01 ,Meeting minutes sample data evaluation: This application proposes a method for meeting minutes quality evaluation, and the meeting minutes quality evaluation framework is as follows Figure 3As shown in the figure, based on the multi-dimensional quantitative evaluation method, the meeting minutes sample dataset is comprehensively evaluated in terms of content similarity, entity coverage, and text fluency. High-quality meeting minutes are selected to perform supervised fine-tuning on the large language model to generate a fine-tuned large language model.
[0081] Further, in step S 01 In the paper, the sample data quality assessment method is fine-tuned based on multi-dimensional meeting minutes, including:
[0082] S 01-1 Content similarity: It aims to measure the semantic consistency between the original meeting record and the meeting minutes. Fig. 9 As shown in the figure, the two files are first processed into sentences. Using the period as the separator, two sets of sentences are generated. Then the All-MiniLM-L6-V2 model is used for sentence embedding, and the sentence vectors corresponding to the two articles are output. The sentence vector of the original conference record file is c i , the sentence vector set is C = {c 1 ,c 2 ,…,c m}; The sentence vector of the meeting minutes is g i , the sentence vector set is G = {g 1 ,g 2 ,…,g n}. When using the formula
[0083]
[0084] Calculate the similarity between two sentences, with a value range of 0-1. In the formula, · represents the dot product of the vectors, ‖g i ‖ and ||c j || represents the Euclidean norm of the vector, i takes values from 1 to n, and j takes values from 1 to m. By calculating the cosine similarity between two sentences, we can form the familiarity matrix S (i,j) :
[0085] c1 c2 c3 c4 c5 …… g1 <![CDATA[S (1,1) ]]> <![CDATA[S (1,2) ]]> <![CDATA[S (1,3) ]]> <![CDATA[S (1,4) ]]> <![CDATA[S (1,5) ]]> …… g2 <![CDATA[S (2,1) ]]> <![CDATA[S (2,2) ]]> <![CDATA[S (2,3) ]]> <![CDATA[S (2,4) ]]> <![CDATA[S (2,5) ]]> …… g3 <![CDATA[S (3,1) ]]> <![CDATA[S (3,2) ]]> <![CDATA[S (3,3) ]]> <![CDATA[S (3,4) ]]> <![CDATA[S (3,5) ]]> …… g4 <![CDATA[S (4,1) ]]> <![CDATA[S (4,2) ]]> <![CDATA[S (4,3) ]]> <![CDATA[S (4,4) ]]> <![CDATA[S (4,5) ]]> …… g5 <![CDATA[S (5,1) ]]> <![CDATA[S (5,2) ]]> <![CDATA[S (5,3) ]]> <![CDATA[S (5,4) ]]> <![CDATA[S (5,5) ]]> …… …… …… …… …… …… …… ……
[0086] Finally, the maximum value of each row is averaged. The calculation formula is as follows:
[0087]
[0088] Get the similarity S between the original meeting record file and the meeting minutes file in the content. The value range is between 0-1.
[0089] S 01-2 Entity coverage: This measures the proportion of entities in the minutes to those in the original record. Fig.10First, extract the entity original_entity of the original meeting record file i and the frequency of the corresponding entity original_frequency i , forming a dictionary {original_entity i :original_frequency i (i∈[1~n])}. And the entity summary_entity of the meeting minutes file j . The entity set {summmary_entity j (j∈[1~m])}. According to the following formula:
[0090]
[0091] Get the entity coverage rate N of the meeting minutes file in the original meeting record file. The value range is between 0-1. Represents the sum of the frequencies of meeting minutes entities appearing in the original meeting record entity set. Represents the sum of all frequencies of the original meeting record entity set.
[0092] S 01-3 Text Fluency: Methods such as Fig.11 As shown in Figure 2. This process only processes the meeting minutes file, which is used to evaluate the grammatical correctness and fluency of the generated text. The Bert-Base-Uncased model is used. Since the model input token has an upper limit, the original text is first processed into sentences to obtain the text sequence T = {t 1 ,t 2 ,…,t n Then input the text sequence into the Bert-Base-Uncased model, and use the loss value output by the model as the confusion of the text to obtain the confusion sequence F ′ ={f 1 ,f 2 ,…,f n}, and then calculate the mean as the overall confusion of the meeting minutes. Finally, due to the confusion of the text f i The value may reach several thousand. Before it is brought in, it needs to be reduced to between (0-1). The inverse proportional function is used to invert the confusion F to get the fluency F′, the formula is:
[0093]
[0094] Get the text fluency F, which ranges from 0 to 1. Used to characterize the confusion of meeting minutes, mid F It is the mean value of confusion, which is usually 500.
[0095] S 01-4 Comprehensive evaluation score: After calculating the content familiarity S, entity coverage N, and text fluency F, the score is obtained by weighted summing them according to the weight distribution of α, β, and γ.
[0096] score = α × S + β × N + γ × F
[0097] Among them, α, β, and γ are the weights of content familiarity, entity coverage, and text fluency, respectively. It is required that α+β+γ=1. The value range of content familiarity S, entity coverage N, and text fluency F are all between 0-1. The comprehensive evaluation score is also between 0-1.
[0098] S 01-5 Meeting minutes: 01-4 The score is obtained from the calculation, and the value range of score is between 0 and 1. 1 is the dividing point between poor and good scores, x 2 It is the dividing point between good and excellent scores.
[0099] 1) score∈(0~x 1 ) where 0 <x 1 <1. When the score is between (0-x 1 ), the quality of the meeting minutes was rated as poor, indicating that most of the entity recognition was inaccurate, the text was not smooth, and it was unreadable. The generated content was too semantically different from the original content.
[0100] 2) score∈(x 1 ~x 2 ) where 0 <x 1 <x 2 <1When the score is in (x 1 -x 2 ), the quality of the meeting minutes is rated as good, indicating that the entity recognition accuracy is mostly correct, the text is readable, and can be understood by humans. The generated content is semantically similar to the original content.
[0101] 3) score∈(x 2 ~1) where 0 <x 2 <1When the score is in (x 2 -1), the quality of the meeting minutes is rated as excellent, indicating that the entity recognition accuracy is high, the text is fluent, and it can be understood by humans and language models. The generated content has very little semantic difference from the original content.
[0102] Step S 02 ,Inference template for large language models: The inference template contains the following key parts, such as Figure 4As shown in the figure. First, the original text and feature information, including filling rules, original text and feature information. Then comes the output format part, which includes the generation format and task description. The generation format is customized by the meeting minutes framework to ensure that the content generated by the large model conforms to the specification in structure. Finally, a complete meeting minutes example is given to clarify the output content of the large language model.
[0103] Further, in step S 02 Inference templates for large language models include:
[0104] S 02-0 Original text and feature information: The first part is the filling rules, where __ and ** represent the parts to be filled in, __ is the original meeting record filling part, and ** is the feature information filling part. The corresponding time, place, task and other entities need to be filled in the corresponding ** position.
[0105] Secondly, in the original text, due to the upper limit of the tocken input by the large model, some original meeting records that are too long cannot be input into the inference template at one time. Using the large language model chatglm3-6b, the upper limit of the tocken that the model can process at one time is 8192. Therefore, the long meeting record text needs to be processed in blocks, with 5000 tockens as the critical point. Then fill it into the inference template.
[0106] Finally, feature information is required. Before filling in feature information, the original meeting minutes text needs to be feature extracted. Using traditional text processing methods, we first select the corresponding algorithm or model to extract feature information according to the specific subdivision task. After this process, the model can better generate meeting minutes content that conforms to the context based on the original text and feature information. Fig.12 It is the feature information part in the inference template.
[0107] Table 2 Example of feature information content
[0108]
[0109] Further, in step S 02-0 The feature information part of the input description is for the extraction of feature information, including:
[0110] S 02-0-0 Topic information: It depends on probability models such as Lda (Latent Dirichlet Allocation). First, the original conference record text is segmented to obtain the segmentation set W = (w 1 ,w 2 ,…,w n), create a corpus by using the unique index of W and the frequency of word segmentation. Then input the corpus into the Lda model, and the model will output multiple topics and the probability value corresponding to each topic. The maximum probability value is selected as the result of topic extraction and filled into the inference template as part of the feature information.
[0111] S 02-0-1 Keywords: For the original meeting minutes text, the Bert-Bilstm-CRF model is used to identify key information. The original meeting minutes text is input, and the BERT-Bilstm-CRF model outputs information such as time, location, and people, which is filled into the inference template as part of the feature information.
[0112] S 02-0-2 Core sentences: A method combining TF-IDF (Term Frequency-Inverse Document Frequency) and K-means clustering is used. First, the original text is divided into sentences. Then, each sentence is segmented, and the words in the same sentence are regarded as a group of words. Then, the feature vector of each sentence is calculated according to the TF-IDF algorithm. The TF-IDF vector of each sentence is actually the representation of the sentence in the vocabulary space. Clustering is performed using the k-means algorithm. According to the number of clusters, the sentence closest to the cluster center is selected as the representative sentence of the class. Finally, the representative sentence output is filled into the inference template as part of the feature information.
[0113] S 02-1 Output format: In order to make the format of the content generated by the large language model conform to the meeting minutes framework, how to convert the meeting minutes framework into the generated format is of utmost importance. Figure 7 It is the meeting minutes framework. It mainly contains 3 paragraphs. The first paragraph contains basic information of the meeting, such as meeting name, meeting time, meeting location, organization unit, participants, etc. The second paragraph mainly contains the meeting theme and the summary information of the corresponding theme. The third paragraph contains the meeting results, meeting summary and other scattered contents of the meeting. Then the meeting minutes framework is converted into the generated format part. The content related to the semantic information and entity information of the meeting minutes is replaced by xx.
[0114] Secondly, in the task description, it is clearly stated that the large model uses the character replacement method to replace the identifier "XX" in the output format based on the original text and feature information of the previous part of the reasoning template, combined with the generation format of this part, and referring to the complete meeting minutes of the next part, such as Figure 7 As shown, minutes of the meeting are formed.
[0115] The last part of the inference template provides a minutes example as a reference. The semantics of the meeting minutes are independent of the semantics of the original record. This part provides a clear output target for the model, making the generation process more targeted and controllable.
[0116] S 02-2 Large model generates meeting minutes: Before the inference template is input into the large language model for inference. The inference template content needs to be filled, and the inference template contains a variety of identifiers, such as *, _, XX, etc. The underscore "_" represents the original text identifier. The asterisk "*" represents the feature information identifier. "XX" represents the identifier of the content to be generated. The original meeting record is used to replace the underscore _ in the inference template, and the traditional method is used to extract feature information to replace the asterisk * in the inference template. Finally, the filled in inference template is used as input and combined with the fine-tuned large language model to jointly generate the meeting minutes.
[0117] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0118] Any process or method description described in a flowchart or otherwise in this specification may be understood to represent a module, fragment or portion of code that includes one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0119] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0120] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A model fine-tuning and template prompting reasoning method for meeting minutes generation, characterized in that: The following steps are involved: Step 1: Input the meeting minutes dataset to construct a fine-tuning sample dataset; Among them, the meeting minutes dataset includes the original meeting record file, meeting minutes file, meeting record time, meeting topic and meeting attribute information; Step 2: Perform a meeting minutes quality assessment on the fine-tuning sample data set to screen out fine-tuning sample data that meets the quality assessment requirements, perform supervised fine-tuning on the large language model used to generate meeting minutes, and generate a fine-tuned large language model; The evaluation items for the quality assessment of meeting minutes include: Content similarity measures the semantic consistency between the original meeting record file and the meeting minutes file; Entity coverage, which measures the proportion of entities in the meeting minutes document to those in the original meeting record document; Text fluency, assessing the grammatical correctness and fluency of the meeting minutes document; Step 3: Based on the set inference template of the large language model, guide the fine-tuned large language model to generate meeting minutes that meet the format and content constraints as required; The inference template of the large language model set includes original text and feature information, output format and output examples.
2. The method according to claim 1, characterized in that In step 2, the content similarity is set to: The original meeting record file and the meeting minutes file are processed separately to obtain two sets with sentences as elements; Based on the sentence embedding representation acquisition model, the embedding representation of each sentence in each set is extracted to obtain the sentence vector of each sentence. Define c j Represents any sentence vector of the original conference record file, and defines g i An arbitrary sentence vector representing a meeting minutes document; For each sentence vector in the meeting minutes file, calculate the sentence similarity between it and each sentence vector in the original meeting record file, and use the maximum value as the content similarity evaluation value of the current sentence vector in the meeting minutes file; Then, based on the average of the content similarity evaluation values of all sentence vectors in the meeting minutes file, the content similarity measurement value between the original meeting record file and the meeting minutes file is obtained.
3. The method according to claim 2, characterized in that The formula for calculating sentence similarity is:
4. The method according to claim 2, characterized in that The similarity matrix is constructed with the sentence index i of the meeting minutes file as the column and the sentence index j of the original record file as the column. Each element of the matrix is used to represent any sentence vector g. i With c j The similarity of sentences between The maximum value of each row of the similarity matrix is averaged to obtain the content similarity measurement value between the original meeting record file and the meeting minutes file.
5. The method according to claim 2, characterized in that The entity coverage is set to: Extract the entity original_entity of the original meeting record file i and the frequency of the corresponding entity original_frequency i , forming a dictionary {original_entity i :original_frequency i ,i∈[1~n]}; Extract the summary_entity entity of the meeting minutes file j , forming an entity set {summary_entity j ,j∈[1~m]}; The entity coverage between the original meeting record file and the meeting minutes file is calculated according to the formula: Among them, subscript i is the entity index of the original meeting record file, n is the entity number of the original meeting record file; subscript j is the entity index of the meeting minutes file, and m is the entity number of the meeting minutes file.
6. The method according to claim 2, characterized in that The text fluency settings are: The meeting minutes file is processed into sentences to obtain the text sequence T = {t1, t2, ..., t n′ }, where t1, t2, …, t n′ represents n′ sentences; Input the text sequence T into the pre-trained natural speech processing model based on the Transformers framework, use the loss value output by the model as the perplexity of the text, and get the corresponding perplexity sequence F ′ ={f1,f2,…,f n′ }; The confusion sequence F ′ The mean confusion of is taken as the overall confusion of the meeting minutes document; The overall confusion is inverted using an inverse proportional function to obtain the text fluency of the meeting minutes document.
7. The method according to claim 1, characterized in that The original text and feature information, output format, and output examples of the inference template of the large language model are as follows: The original text and feature information include: filling rules, original text and feature information; Among them, the feature information is extracted from the text content of the original meeting record file using several text feature extraction models; The filling rule fills the original text, feature information and output format content into the inference template by replacing characters; Output format, including meeting minutes framework, generation format and task description.
8. The method according to claim 7, characterized in that The output format adopts a three-paragraph structure. The first paragraph contains basic information about the meeting; the second paragraph is the meeting theme and summary information of the corresponding theme; the third paragraph contains the meeting results, meeting summary and other scattered contents of the meeting.
9. The method according to claim 8, characterized in that The meeting minutes framework includes the minutes name, meeting time, meeting organizer information, participant information, meeting content / topic, several meeting results, other content, and meeting minutes signature information.
10. The method according to claim 7, characterized in that The feature information is extracted from the text content of the original meeting record file using several text feature extraction models, including: Topic feature extraction: The meeting record text in the original meeting record file is segmented to obtain the segmentation set W = (w1, w2, ..., w n″ ), where n″ is the number of segmentations in the segmentation set W; and a corpus is created through the unique index of the segmentation set W and the frequency of the segmentation; the corpus is input into the topic extraction model to obtain several topics and the probability value corresponding to each topic, and the maximum probability value is selected as the result of topic extraction; Keyword feature extraction: Input the meeting record text in the original meeting record file into the keyword recognition model to obtain the keyword recognition results; Core sentence feature extraction: The meeting record text in the original meeting record file is divided into sentences, and then each sentence is segmented. The words in the same sentence are regarded as a group of words; the feature vector of each sentence is calculated according to the word frequency-inverse document frequency algorithm; then the k-means algorithm is used for clustering, and according to the number of clusters, the sentences closest to the cluster center are selected as the core sentences of each category.