Conference data construction method and device, electronic equipment and program product
By identifying and preprocessing the original audio of the conference, combining the determination of the conference type and the extraction and correction of elements, the conference data is constructed, and the problem of low accuracy in the construction of conference data in the existing technology is solved, and a higher quality conference minutes is achieved.
Patent Information
- Application Number
- CN202411983555.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The accuracy of constructing conference data in the prior art is low. Traditional keyword extraction methods cannot ensure that conference minutes cover all key points, and it is easy to miss or ignore low-frequency but important information, affecting the accuracy of conference minutes.
By obtaining the original audio of the meeting for identification, the original text of the meeting is obtained and preprocessed, the meeting type is determined, the text is divided, the text is extracted and corrected, and the text, the target element and the prompt word will be spliced to construct the meeting data.
It improves the accuracy and comprehensiveness of conference data construction, ensures the structural integrity and logical clarity of conference minutes, covers all important points, and enhances the quality of conference minutes.
Smart Images

Figure CN119940536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for constructing conference data, a device thereof, an electronic device, and a program product. Background Art
[0002] Meeting minutes are the summary of the meeting, which is of great significance to the participants and organizers. They can not only record the discussion of the meeting in a timely and comprehensive manner, collect opinions and suggestions from all parties, but also effectively promote the execution and implementation of the meeting results and improve work efficiency. They are also an important basis for enterprises or organizations to make decisions, manage and evaluate. With the continuous expansion of the scale of meetings and the continuous increase in the frequency of meetings, traditional manual writing of meeting minutes can no longer meet the needs of efficiency, accuracy and timeliness.
[0003] With the development of artificial intelligence technology and big language models, the automatic generation of meeting minutes has been realized. The meeting minutes generation technology based on big language models is mainly divided into two types: one is the integration of templates and information points, which generates basic minutes by intelligently identifying professional terms and the emotional color and attitude of the content in the meeting information, and then classifies and integrates them to generate summaries, and finally merges them with customized minutes templates to produce meeting minutes with complete structure and clear logic; the second is the extraction of key information, which uses big language models to identify themes, viewpoints, keywords, words and sentences in the initial meeting text to realize the writing of meeting minutes.
[0004] However, the accuracy of large language models is still limited by the quality and diversity of training data. Therefore, building high-quality training datasets is crucial to improving the generation effect of the model.
[0005] In the related art, the construction scheme of meeting minutes data mainly relies on the method of extracting leading keywords to assist in the subsequent minutes summary. However, this method has the following problems: the initial meeting text input into the large model is the text obtained based on automatic speech recognition, which has the problem of low recognition accuracy, and the output results of the meeting minutes content are often unstable due to insufficient training data or lack of clear template guidance, and the structural integrity and logical clarity are also low. At the same time, the process from keyword extraction to meeting minutes generation, although aimed at improving efficiency, is complicated and difficult due to numerous links and unclear definitions. In addition, the method relying on keywords cannot ensure that the meeting minutes cover all the key points, omitting or ignoring low-frequency but important information, affecting the accuracy of the meeting minutes.
[0006] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0007] The embodiment of the present invention provides a method for constructing conference data and a device thereof, an electronic device and a program product, so as to at least solve the technical problem of low accuracy in constructing conference data in the related art.
[0008] According to one aspect of an embodiment of the present invention, a method for constructing conference data is provided, including: acquiring original conference audio, recognizing the original conference audio to obtain original conference text, and preprocessing the original conference text to obtain conference processing text; determining the conference type based on the conference processing text; segmenting the conference processing text based on the conference type to obtain a preset number of conference processing sub-texts, and extracting elements from the conference processing sub-texts based on a first preset prompt word to obtain an initial element; correcting the initial element based on the conference processing text to obtain a first target element; and splicing the conference processing text, the first target element, and the second preset prompt word to complete the construction of the conference data.
[0009] Furthermore, the step of preprocessing the original conference text to obtain a conference processed text includes: determining the order of speakers in the original conference text; determining the speaker number corresponding to each speaker based on the speaker order; replacing the inventor name of each speaker in the original conference text with the speaker number, and deleting the speech time information corresponding to each speaker number to obtain the conference processed text, wherein the conference processed text includes: multiple speech contents associated with the speaker number.
[0010] Further, the conference types include at least: single-person scenario conference type and multi-person scenario conference type, the single-person scenario conference types include at least: single-person manuscript scenario type and single-person no-script scenario type, the single-person manuscript scenario type is a scenario type in which a single person speaks based on a manuscript, and the single-person no-script scenario type is a scenario type in which a single person speaks without a manuscript. The step of determining the conference type based on the conference processing text includes: determining a preset number based on the conference processing text, wherein the preset number is the number of speaker numbers; when the preset number is less than a preset threshold, determining that the conference type is the single-person scenario conference type; when the conference type is the single-person scenario conference type, determining whether a preset language feature exists in the conference processing text; when the preset language feature exists in the conference processing text, determining that the conference type is the single-person manuscript scenario type; when the preset language feature does not exist in the conference processing text, determining that the conference type is the single-person no-script scenario conference type.
[0011] Furthermore, the initial element at least includes: a conference title and key content, the key content includes: multiple titles and text content corresponding to each title, and the initial element is corrected based on the conference processing text to obtain the first target element, including: determining a content standard strategy for the initial element, and based on the content standard strategy, correcting the key content of the initial element to obtain the target key content, wherein the content standard strategy at least includes: a strategy for determining whether the key content of the initial element is consistent with the content of the conference processing text, a strategy for determining whether the sentence order of the initial element is consistent with the sentence order of the conference processing text, a strategy for determining that there are no repeated sentences in the initial element, and a strategy for determining that there are no pronouns in the initial element; if there are numbers in the conference title, the numbers are deleted to obtain the target conference title; based on a preset format, the format of the initial element is adjusted to obtain the target format; based on the target key content, the target conference title and the target format, the first target element is determined.
[0012] Furthermore, based on the content standard strategy, the key content of the initial element is corrected to obtain the target key content, including: when the key content of the initial element is inconsistent with the content in the conference processing text, the key content inconsistent with the content in the conference processing text is modified to the content in the conference processing text; when the sentence order of the initial element is inconsistent with the sentence order in the conference processing text, the sentences in the initial element are reordered based on the sentence order in the conference processing text; when there are repeated sentences in the initial element, the repeated sentences are merged; when the pronoun exists in the initial element, the pronoun is replaced; the number of words in the conference processing text is compared with the number of words in the key content in the initial element, and when it is detected that the number of words in the key content is less than a preset word count threshold, the number of words in the key content is supplemented to obtain the target key content.
[0013] Furthermore, the method for constructing the conference data also includes: when the conference type is the single-person manuscript scene type, the conference processing text is hierarchically divided to obtain multiple text titles of different levels and conference processing text contents corresponding to different levels; when the number of the text titles at the first level is greater than the preset number of titles, each of the text titles at the first level is characterized as the title under the key content; or, when the number of the text titles at the first level is less than the preset number of titles, each of the text titles at the second level is characterized as the title under the key content, wherein the level of the first level is greater than that of the second level; based on each of the titles, key phrases are extracted from the conference processing text content to determine the text content corresponding to each of the titles, and based on all the titles and all the text contents, the key content is determined; based on the key content, the conference title and the full text summary of the conference processing text are determined; based on the key content, the second target element is determined.
[0014] Furthermore, the initial elements also include: a summary of the full text, a to-do list, and the second preset prompt words include: a single-element prompt word, a double-element prompt word, a three-element prompt word, and a four-element prompt word. The meeting processing text, the first target element, and the second preset prompt word are spliced to complete the step of constructing the meeting data, including: splicing the single-element prompt word, the meeting processing text, and the single element to obtain the training data of the single element, wherein the single element is any one of the first target element and the second target element, and the single-element prompt word is the prompt word of any one of the elements; splicing the double-element prompt word, the meeting processing text, and the double elements to obtain the training data of the double elements, wherein the double elements are any two different elements of the first target element and the second target element, and the double-element prompt word is the prompt word of the any two different elements; splicing the three-element prompt word, the meeting processing text, And three elements are spliced to obtain the training data of the three elements, wherein the three elements are any three different elements of the first target element and the second target element, and the three-element prompt word is the prompt word of the any three different elements; the four-element prompt word, the conference processing text, and the four elements are spliced to obtain the training data of the four elements, wherein the four elements are any four different elements of the first target element and the second target element, and the four-element prompt word is the prompt word of the any four different elements, each of the elements has a sequence level in the prompt word, and the sequence level is that the level of the conference title is greater than the level of the full-text summary, the level of the full-text summary is greater than the level of the key content, and the level of the key content is greater than the level of the to-do list; based on the single-element training data, the double-element training data, the three-element training data and the four-element training data, the conference data is constructed.
[0015] Furthermore, after the meeting processing text, the first target element and the second preset prompt word are spliced to complete the construction of the meeting data, it also includes: constructing a mixed data set based on the training data and the preset data; using the mixed data set to train the preset model, and using text expansion technology to adjust the length of the processed text output by the preset model until the number of training rounds is equal to the preset round threshold, thereby obtaining a preset meeting minutes model, wherein the preset meeting minutes model generates meeting minutes based on the original meeting text.
[0016] According to another aspect of an embodiment of the present invention, a device for constructing conference data is also provided, including: a processing unit, used to obtain original conference audio, identify the original conference audio, obtain original conference text, and pre-process the original conference text to obtain conference processing text; a determination unit, used to determine the type of conference based on the conference processing text; an extraction unit, used to segment the conference processing text based on the conference type to obtain a preset number of conference processing sub-texts, and extract elements from the conference processing sub-texts based on a first preset prompt word to obtain initial elements; a correction unit, used to correct the initial elements based on the conference processing text to obtain a first target element; and a splicing unit, used to splice the conference processing text, the first target element and the second preset prompt word to complete the construction of the conference data.
[0017] Furthermore, the processing unit includes: a first determination module, used to determine the order of speakers in the original text of the meeting; a second determination module, used to determine the speaker number corresponding to each speaker based on the speaker order; a first processing module, used to replace the inventor name of each speaker in the original text of the meeting with the speaker number, and delete the speech time information corresponding to each speaker number to obtain the meeting processing text, wherein the meeting processing text includes: multiple speech contents associated with the speaker number.
[0018] Further, the conference type includes at least: a single-person scenario conference type and a multi-person scenario conference type, the single-person scenario conference type includes at least: a single-person manuscript scenario type and a single-person no-manuscript scenario type, the single-person manuscript scenario type is a scenario type in which a single person speaks based on a manuscript, and the single-person no-manuscript scenario type is a scenario type in which a single person speaks without a manuscript, and the determination unit includes: a third determination module, used to determine the number of preset numbers based on the conference processing text, wherein the number of preset numbers is the number of speaker numbers; a fourth determination module, used to determine that the conference type is the single-person scenario conference type when the number of preset numbers is less than a preset threshold; a fifth determination module, used to determine whether there is a preset language feature in the conference processing text when the conference type is the single-person scenario conference type; a sixth determination module, used to determine that the conference type is the single-person manuscript scenario type when the preset language feature exists in the conference processing text; a seventh determination module, used to determine that the conference type is the single-person no-manuscript scenario conference type when the preset language feature does not exist in the conference processing text.
[0019] Furthermore, the initial element at least includes: a conference title and key content, the key content includes: multiple titles and text content corresponding to each title, and the correction unit includes: a first correction module, used to determine the content standard strategy of the initial element, and based on the content standard strategy, correct the key content of the initial element to obtain the target key content, wherein the content standard strategy at least includes: a strategy for determining whether the key content of the initial element is consistent with the content of the conference processing text, a strategy for determining whether the sentence order of the initial element is consistent with the sentence order of the conference processing text, a strategy for determining that there are no repeated sentences in the initial element, and a strategy for determining that there are no pronouns in the initial element; a first deletion module, used to delete the numbers if there are numbers in the conference title to obtain the target conference title; a first adjustment module, used to adjust the format of the initial element based on a preset format to obtain the target format; an eighth determination module, used to determine the first target element based on the target key content, the target conference title and the target format.
[0020] Furthermore, the first correction module includes: a first modification submodule, which is used to modify the key content inconsistent with the content in the conference processing text to the content in the conference processing text when the key content of the initial element is inconsistent with the content in the conference processing text; a first sorting submodule, which is used to reorder the sentences in the initial element based on the sentence order in the conference processing text when the sentence order of the initial element is inconsistent with the sentence order in the conference processing text; a first merging submodule, which is used to merge the repeated sentences when there are repeated sentences in the initial element; a first replacement submodule, which is used to replace the pronoun when there is the pronoun in the initial element; a first supplementing submodule, which is used to compare the number of words in the conference processing text with the number of words in the key content in the initial element, and when it is detected that the number of words in the key content is less than a preset word count threshold, supplement the number of words of the key content to obtain the target key content.
[0021] Furthermore, the construction device also includes: a first division module, which is used to hierarchically divide the conference processing text when the conference type is the single-person manuscript scene type, and obtain multiple text titles of different levels and conference processing text contents corresponding to different levels; a first characterization module, which is used to characterize each of the text titles of the first level as the title under the key content when the number of the text titles of the first level is greater than the preset number of titles; a second characterization module, which is used to characterize each of the text titles of the second level as the title under the key content when the number of the text titles of the first level is less than the preset number of titles, wherein the level of the first level is greater than that of the second level; a ninth determination module, which is used to extract key phrases from the conference processing text content based on each of the titles, determine the text content corresponding to each of the titles, and determine the key content based on all the titles and all the text contents; a tenth determination module, which is used to determine the conference title and the full text summary of the conference processing text based on the key content; an eleventh determination module, which is used to determine the second target element based on the key content, the conference title and the full text summary.
[0022] Furthermore, the initial elements also include: a summary of the full text, a to-do list, the second preset prompt words include: a single-element prompt word, a double-element prompt word, a triple-element prompt word, and a quadruple-element prompt word, and the splicing unit includes: a first splicing module, used to splice the single-element prompt word, the conference processing text, and a single element to obtain training data of the single element, wherein the single element is any one of the first target element and the second target element, and the single-element prompt word is a prompt word of any one of the elements; a second splicing module, used to splice the double-element prompt word, the conference processing text, and the double elements to obtain training data of the double elements, wherein the double elements are any two different elements of the first target element and the second target element, and the double-element prompt word is a prompt word of any two different elements; a third splicing module, used to splice the triple-element prompt word, the conference processing text, and the three elements to obtain the triple-element prompt word The invention relates to training data of the conference process text, wherein the three elements are any three different elements of the first target element and the second target element, and the three-element prompt word is the prompt word of the any three different elements; a fourth splicing module is used to splice the four-element prompt word, the conference processing text, and the four elements to obtain the training data of the four elements, wherein the four elements are any four different elements of the first target element and the second target element, and the four-element prompt word is the prompt word of the any four different elements, and each of the elements has a sequence level in the prompt word, and the sequence level is that the level of the conference title is greater than the level of the full-text summary, the level of the full-text summary is greater than the level of the key content, and the level of the key content is greater than the level of the to-do list; a first construction module is used to construct the conference data based on the single-element training data, the double-element training data, the three-element training data, and the four-element training data.
[0023] Furthermore, the construction device also includes: a second construction module, which is used to construct a mixed data set based on the training data and the preset data after splicing the meeting processing text, the first target element and the second preset prompt word to complete the construction of the meeting data; a first training module, which is used to use the mixed data set to train the preset model, and use text expansion technology to adjust the length of the processed text output by the preset model until the number of training rounds is equal to the preset round threshold, so as to obtain a preset meeting minutes model, wherein the preset meeting minutes model generates meeting minutes based on the original text of the meeting.
[0024] According to another aspect of an embodiment of the present invention, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for constructing conference data described in any one of the above is implemented.
[0025] According to another aspect of an embodiment of the present invention, there is also provided an electronic device, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above-mentioned conference data construction methods.
[0026] In the present invention, the original audio of the meeting is obtained, the original audio of the meeting is recognized to obtain the original text of the meeting, and the original text of the meeting is preprocessed to obtain the conference processing text. Based on the conference processing text, the type of meeting is determined. Based on the conference type, the conference processing text is segmented to obtain a preset number of conference processing sub-texts, and based on the first preset prompt word, elements of the conference processing sub-text are extracted to obtain initial elements. Based on the conference processing text, the initial elements are corrected to obtain the first target element. The conference processing text, the first target element and the second preset prompt word are spliced to complete the construction of the conference data, thereby solving the technical problem of low accuracy in constructing conference data in the related art.
[0027] In the present invention, the original text of the meeting can be obtained by identifying the acquired original audio of the meeting, and the original text of the meeting is preprocessed to obtain the processing text of the meeting. According to the processing text of the meeting, the type of the meeting can be determined. According to the determined type of meeting, the processing text of the meeting can be segmented to generate a preset number of conference processing sub-texts to ensure that each segment contains an appropriate amount of information. Then, according to the first preset prompt word, the meeting processing sub-text is extracted to include elements such as the meeting title, full text summary, main content and to-do items to obtain the initial element. Thereafter, according to the conference processing text, the initial element is corrected to improve the details of the initial element to obtain the first target element. The conference processing text, the first target element and the second preset prompt word are spliced to complete the construction of the conference data, optimize each link of the construction of the conference data, and thus achieve the technical effect of improving the accuracy of the construction of the meeting minutes data. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0029] Figure 1is a flowchart of an optional method for constructing conference data according to an embodiment of the present invention;
[0030] Figure 2 is an optional flow chart for constructing conference data according to an embodiment of the present invention;
[0031] Figure 3 is a schematic diagram of an optional device for constructing conference data according to an embodiment of the present invention;
[0032] Figure 4 The present invention is a hardware structure block diagram of an electronic device (or mobile device) for a method for constructing conference data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) collected and involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data are in compliance with the relevant laws, regulations and standards of the relevant regions, necessary confidentiality measures are taken, and public order and good customs are not violated, and corresponding operation entrances are provided for users to choose to authorize or refuse. For example, an interface is set between the system and the relevant users or organizations. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or organization.
[0036] In the present invention, comprehensive coverage of the content of meeting minutes is ensured through clear scene classification, construction of precise prompt words, and construction of standard meeting minutes strategies. At the same time, customized output templates and standardized format requirements further improve the accuracy and comprehensiveness of meeting minutes data, providing more accurate and comprehensive training data for the automated writing of meeting minutes.
[0037] The present invention is described in detail below in conjunction with various embodiments.
[0038] Embodiment 1
[0039] According to an embodiment of the present invention, an embodiment of a method for constructing conference data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] Figure 1 is a flowchart of an optional method for constructing conference data according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:
[0041] Step S101, obtaining the original conference audio, recognizing the original conference audio to obtain the original conference text, and preprocessing the original conference text to obtain the conference processed text.
[0042] Optionally, the original audio of the meeting can be obtained through offline meetings and online meetings. For example, for offline meetings, the audio file of each meeting can be exported through the recording equipment in the conference room, and for online meetings, the audio in the video can be extracted by downloading the meeting screen recording file.
[0043] In this embodiment, the audio transcription and recognition technology can be used to obtain the ASR (Automatic Speech Recognition) text of the meeting (i.e., the original text of the meeting), and the ASR text of each meeting is preprocessed (such as removing the speaker's name and timestamp), and the text format is standardized to obtain the processed text of the meeting.
[0044] Step S102: Determine the conference type based on the conference processing text.
[0045] In this embodiment, by analyzing the characteristics of the conference processing text, the conference type (such as multi-person scenario conference type, single-person scenario conference type) can be determined. The single-person scenario conference type focuses on the speaker's coherent expression, while the multi-person scenario conference type emphasizes the interaction and collision of opinions between different speakers. Accurate scene classification helps to subsequently determine the text and prompt words.
[0046] Step S103, based on the conference type, the conference processing text is segmented (eg, each 2,000 words is a part) to obtain a preset number of conference processing sub-texts, and based on the first preset prompt word, the conference processing sub-text is extracted to obtain initial elements.
[0047] Optionally, according to the meeting type (such as multi-person scenario meeting type and single-person no-script scenario type), the meeting processing text can be segmented to obtain a preset number of meeting processing sub-texts, ensuring that each sub-text contains sufficient information for element extraction.
[0048] In this embodiment, based on the first preset prompt word (pre-set prompt word), a large model can be used to extract elements from the conference processing subtext to obtain initial elements (such as conference title, full text summary, key content and to-do items).
[0049] For example, the first preset prompt word may be:
[0050] {ASR text fragment}, now you need to understand the conference conversation content provided above, fully understand the key information in the conference conversation, and generate the main content of the paragraph {round}, requirements:
[0051] (1) The main content generated should include the opinions and issues expressed by the speaker;
[0052] (2) List the viewpoints and issues involved in an ordered list format, as follows: "1. Summary of a certain issue or viewpoint\nThe main content of the issue or viewpoint;" The main content of the issue or viewpoint should be written in one paragraph, not in a list format, and no other symbols should be added before the paragraph;
[0053] (3) The main contents should be organized and avoid repetition;
[0054] (4) Do not output content other than the main content of the meeting;
[0055] (5) Please limit your summary to approximately 500 words.
[0056] Please fully understand the given meeting dialogue, fully understand the generation requirements, and then output the main content of the meeting in paragraph {round}: {main content}.
[0057] Step S104: based on the conference processing text, correct the initial element to obtain a first target element.
[0058] Optionally, the extracted elements can be corrected to avoid erroneous information and missing content, ensuring that the representation of all elements is consistent with the meeting content.
[0059] In this embodiment, the initial element is compared with the conference processing text, and the content and format (such as the number of words) of the initial element are corrected according to the conference processing text, so as to obtain the first target element.
[0060] Step S105, concatenating the conference processing text, the first target element and the second preset prompt word to complete the construction of the conference data.
[0061] In this embodiment, the conference processing text, the first target element and the second preset prompt word (ie, a pre-set word (such as a word linking the conference processing text and the first target element)) can be spliced to complete the construction of the conference data.
[0062] For example, please understand the following meeting dialogue, fully understand the key information in the meeting dialogue, and generate a meeting summary. \nThe content of the meeting is: \n{ASR text}\n\nPlease generate a meeting summary based on the content of the meeting dialogue. The meeting summary contains four elements: "title, full text summary, main content, and to-do items". The output format is as follows: ##Title
[0063] {title}
[0064] ##Full text summary
[0065] {Full text summary}
[0066] Main content
[0067] {Main content}.
[0068] In this embodiment, the titles of the four element types, namely title, full text summary, main content, and to-do list, must be displayed as secondary titles in Markdown (i.e., a lightweight markup language) format, and cannot be followed by line breaks, semicolons, colons, spaces, and other redundant information. Each element must be separated by two line breaks to ensure the uniformity of the output format.
[0069] In summary, by recognizing the acquired original audio of the meeting, the original text of the meeting can be obtained, and by preprocessing the original text of the meeting, the processed text of the meeting can be obtained. According to the processed text of the meeting, the type of meeting can be determined. When the meeting type is a multi-person scene meeting type or a single-person no-script scene type, the conference processing text can be segmented to obtain a preset number of conference processing sub-texts, and elements of the conference processing sub-text can be extracted according to the first preset prompt word to obtain the initial element, and then the initial element can be corrected according to the conference processing text to obtain the first target element. Thereafter, the conference processing text, the first target element and the second preset prompt word are spliced to complete the construction of the conference data, thereby solving the technical problem of low accuracy in constructing conference data in related technologies.
[0070] In order to accurately obtain the conference processing text, in the method for constructing conference data provided in Example 1 of the present application, the order of speakers in the original conference text is determined; based on the order of speakers, the speaker number corresponding to each speaker is determined; the inventor name of each speaker in the original conference text is replaced with the speaker number, and the speech time information corresponding to each speaker number is deleted to obtain the conference processing text, wherein the conference processing text includes: multiple speech contents associated with the speaker number.
[0071] Optionally, irrelevant background noise or interference information may be filtered out, including but not limited to removing the speaker's name and time stamp, etc.
[0072] In this embodiment, the order of speakers can be determined from the original text of the meeting (such as Zhang San in front and Li Si in the back), and the speaker number corresponding to each speaker can be determined according to the order of speakers (such as Zhang San is numbered 1 and Li Si is numbered 2), and the inventor name of each speaker in the original text of the meeting can be replaced with the speaker number (such as Zhang San is replaced with 1 and Li Si is replaced with 2), and the speaking time information corresponding to each speaker number can be deleted to obtain the processed text of the meeting.
[0073] Exemplarily, "Zhang San (00:02:38): Foundry module chip or not? Li Si (00:38:25): It's a bit troublesome." is converted into "Speaker 1: Foundry module chip or not? Speaker 2: It's a bit troublesome." Through preprocessing, the text is made more standardized, which is convenient for subsequent element extraction and content correction.
[0074] In this embodiment, the conference types include at least: single-person scenario conference type and multi-person scenario conference type. The single-person scenario conference types include at least: single-person manuscript scenario type and single-person no-manuscript scenario type. The single-person manuscript scenario type is a scenario type in which a single person speaks based on a manuscript, and the single-person no-manuscript scenario type is a scenario type in which a single person speaks without a manuscript. In order to accurately determine the conference type, in the conference data construction method provided in Example 1 of the present application, based on the conference processing text, a preset number number is determined, wherein the preset number number is the number of speaker numbers; when the preset number number is less than a preset threshold, the conference type is determined to be a single-person scenario conference type; when the conference type is a single-person scenario conference type, it is determined whether there is a preset language feature in the conference processing text; when the preset language feature exists in the conference processing text, the conference type is determined to be a single-person manuscript scenario type; when the preset language feature does not exist in the conference processing text, the conference type is determined to be a single-person no-manuscript scenario conference type.
[0075] Optionally, the meeting types may include: single-person scenario meeting type and multi-person scenario meeting type. The single-person scenario meeting type may include: single-person manuscript scenario type (i.e., the scenario type in which a single person speaks based on a manuscript) and single-person no-manuscript scenario type (i.e., the scenario type in which a single person speaks without a manuscript).
[0076] In this embodiment, a preliminary analysis is first performed on the conference processing text to determine the number of preset numbers (i.e., the number of speaker numbers). Based on the number of preset numbers, the type of conference (such as a single-person scenario conference type and a multi-person scenario conference type) can be determined. When the number of preset numbers is less than a preset threshold (such as 2), the conference type can be determined to be a single-person scenario conference type. When the conference type is a single-person scenario conference type, it is determined whether there are preset language features in the conference processing text (such as standardized grammatical structure, coherent paragraph division, and clear title). When the preset language features exist in the conference processing text, the conference type can be determined to be a single-person manuscript scenario type. When the preset language features do not exist in the conference processing text, the conference type can be determined to be a single-person no-manuscript scenario conference type.
[0077] In this embodiment, the initial element includes at least: a meeting title, key content, and the key content includes: multiple titles and text content corresponding to each title. In order to accurately determine the first target element, in the method for constructing conference data provided in Example 1 of the present application, the content standard strategy of the initial element is determined, and based on the content standard strategy, the key content of the initial element is corrected to obtain the target key content, wherein the content standard strategy at least includes: a strategy for determining whether the key content of the initial element is consistent with the content of the conference processing text, a strategy for determining whether the sentence order of the initial element is consistent with the sentence order of the conference processing text, a strategy for determining that there are no repeated sentences in the initial element, and a strategy for determining that there are no pronouns in the initial element; if there are numbers in the conference title, delete the numbers to obtain the target conference title; based on the preset format, adjust the format of the initial element to obtain the target format; determine the first target element based on the target key content, the target conference title and the target format.
[0078] Optionally, the initial elements may include: meeting title, key content, and the key content may include: multiple titles and text content corresponding to each title. The initial elements and the meeting processing text may be compared to check the consistency of the two contents, and deviations may be discovered and corrected to ensure a true reflection of the content of the meeting minutes. At the same time, calibration of the sentence order ensures the logical coherence of the meeting minutes, follows the actual process of the meeting discussion, enables readers to understand the meeting agenda in chronological order, and improves the readability and comprehensibility of the information.
[0079] In this embodiment, according to the content standard strategy, the key content of the initial element can be corrected to obtain the target key content. The content standard strategy includes but is not limited to: a strategy for determining whether the key content of the initial element is consistent with the content of the conference processing text, a strategy for determining whether the sentence order of the initial element is consistent with the sentence order of the conference processing text, a strategy for determining that there are no repeated sentences in the initial element, and a strategy for determining that there are no pronouns in the initial element.
[0080] In this embodiment, the conference title is a summary of the theme of the entire ASR content, and no number-related expressions (such as specific time words, year, month, day, etc.) can appear. If there are numbers in the conference title, the target conference title can be obtained by deleting the numbers.
[0081] In this embodiment, the format of the initial element is adjusted according to a preset format (such as markdown format) to obtain the target format. For example, each output field is preceded by a double pound sign and left a space (such as ##title, ##full text summary, ##main content, ##to-do list), and a colon is not added after it. The main content (i.e., the target key content) includes multiple level 1 titles represented by Arabic numerals. The numerical serial number of the level 1 title should be followed by a space, and each level 1 title needs to be wrapped with double asterisks (such as "**1. Macro environment, market situation analysis and company summary**"). The to-do list uses Arabic serial numbers to distinguish the key points.
[0082] In this embodiment, the first target element may be determined according to the target key content, the target meeting title, and the target format.
[0083] In this embodiment, the corresponding text can also be found one by one in the conference processing text according to the key points in the main content (i.e., key content) of the initial element, and marked in blue text font in the conference processing text. The key words found can also be bolded. The sentence range where the key words are located can be marked in blue font, indicating that the sentence or paragraph has summarized the relevant key points. If there are errors or irrelevant content in the result generated by the large model (i.e., the initial element), the irrelevant content can be deleted. There are also unmarked color and bold fragments in the conference processing text. At this time, the unmarked black font parts need to be checked in turn to determine whether there is unsummarized key information. If there is unsummarized key information, the key information can be marked in red in the conference processing text, and the key information can be added to the main content or to-do list.
[0084] In order to improve the accuracy of determining the target key content, in the method for constructing conference data provided in Example 1 of the present application, when the key content of the initial element is inconsistent with the content in the conference processing text, the key content inconsistent with the content in the conference processing text is modified to the content in the conference processing text; when the sentence order of the initial element is inconsistent with the sentence order in the conference processing text, the sentences in the initial element are reordered based on the sentence order in the conference processing text; when there are repeated sentences in the initial element, the repeated sentences are merged; when there are pronouns in the initial element, the pronouns are replaced; the number of words in the conference processing text is compared with the number of words in the key content in the initial element, and when it is detected that the number of words in the key content is less than the preset word count threshold, the number of words in the key content is supplemented to obtain the target key content.
[0085] In this embodiment, if the key content of the initial element is inconsistent with the content in the meeting processing text, the key content inconsistent with the content in the meeting processing text will be modified to the content in the meeting processing text (for example, if the original meeting processing text contains vague time expressions such as today, tomorrow, this year, and next year, the key content of the initial element cannot be output as a specific date).
[0086] In this embodiment, when the sentence order of the initial element is inconsistent with the sentence order in the conference processing text, the sentences in the initial element can be reordered according to the sentence order in the conference processing text to ensure that the order in which the sentences in the initial element appear is consistent with that in the conference processing text. When there are repeated sentences in the initial element, the repeated sentences can be merged.
[0087] In this embodiment, when there are pronouns in the initial elements, the pronouns can be replaced. If a specific person or thing needs to be referred to, a specific name or description can be used to replace the pronoun. The speaker number (such as speaker n) in the conference processing text can also be deleted to avoid unclear reference or confusion.
[0088] In this embodiment, the number of words in the meeting processing text can be compared with the number of words in the key content in the initial element. When it is detected that the number of words in the key content is less than a preset word count threshold (such as the number of words in the meeting processing text*0.8*0.3), the number of words in the key content is supplemented to obtain the target key content.
[0089] In this embodiment, the number of words under each level 1 title may be determined according to the total number of words in the conference processing text and the number of level 1 titles of key contents.
[0090] In order to accurately determine the second target element, in the method for constructing conference data provided in Example 1 of the present application, when the conference type is a single-person manuscript scenario type, the conference processing text is hierarchically divided to obtain multiple text titles of different levels and conference processing text contents corresponding to different levels; when the number of text titles of the first level is greater than the preset number of titles, each text title of the first level is characterized as a title under the key content; or, when the number of text titles of the first level is less than the preset number of titles, each text title of the second level is characterized as a title under the key content, wherein the level of the first level is greater than that of the second level; based on each title, key phrases are extracted from the conference processing text content to determine the text content corresponding to each title, and based on all titles and all text contents, the key content is determined; based on the key content, the meeting title and the full text summary of the conference processing text are determined; based on the key content, the second target element is determined.
[0091] In this embodiment, when the meeting type is a single-person manuscript scenario type, there is a clear main sentence in the meeting processing text, and the meeting processing text can be divided into levels to obtain multiple text titles at different levels and meeting processing text contents corresponding to different levels (such as first dividing the first level, determining the text content range included in the first-level title, and then splitting out multiple second-level titles within the text content range included in each first-level title).
[0092] In this embodiment, when the number of text titles at the first level is greater than the preset number of titles (the preset number of titles can be set according to the actual text content), each text title at the first level can be characterized as a title under the key content, and when the number of text titles at the first level is less than the preset number of titles, each text title at the second level can be characterized as a title under the key content.
[0093] In this embodiment, based on each title, key phrase extraction can be performed on the conference processing text content (i.e., keywords and phrases closely related to the title are extracted) to obtain the text content corresponding to each title, and key content can be constructed based on all titles and all text contents (i.e., each title and the text content fragments corresponding to the title are connected into a coherent whole). Based on the key content, the meeting title and full-text summary of the conference processing text can be obtained. In this way, the second target element can be determined based on the key content, the meeting title and the full-text summary.
[0094] The initial elements also include: a summary of the full text, a to-do list, and the second preset prompt words include: a single-element prompt word, a double-element prompt word, a three-element prompt word, and a four-element prompt word. In order to accurately construct the conference data, in the conference data construction method provided in Example 1 of the present application, the single-element prompt word, the conference processing text, and the single element are spliced to obtain single-element training data, wherein the single element is any element of the first target element and the second target element, and the single-element prompt word is the prompt word of any element; the double-element prompt word, the conference processing text, and the double elements are spliced to obtain double-element training data, wherein the double elements are any two different elements of the first target element and the second target element, and the double-element prompt word is the prompt word of any two different elements; the three-element prompt word, the conference processing text, and the three-element prompt word are spliced to obtain double-element training data, wherein the double elements are any two different elements of the first target element and the second target element, and the double-element prompt word is the prompt word of any two different elements. Elements are spliced to obtain three-element training data, wherein the three elements are any three different elements of the first target element and the second target element, and the three-element prompt words are the prompt words of any three different elements; four-element prompt words, conference processing text, and four elements are spliced to obtain four-element training data, wherein the four elements are any four different elements of the first target element and the second target element, and the four-element prompt words are the prompt words of any four different elements, and each element has a sequence level in the prompt words, and the sequence level is that the level of the conference title is greater than the level of the full text summary, the level of the full text summary is greater than the level of the key content, and the level of the key content is greater than the level of the to-do list; based on the single-element training data, the double-element training data, the three-element training data, and the four-element training data, the conference data is constructed.
[0095] Optionally, the initial elements may also include: full text summary, to-do list, and the second preset prompt words may include: single-element prompt words (i.e., prompt words for any element), double-element prompt words (i.e., prompt words for any two different elements), triple-element prompt words (i.e., prompt words for any three different elements), and quadruple-element prompt words (i.e., prompt words for any four different elements).
[0096] In this embodiment, in order to enhance the comprehensiveness of the conference data (i.e., to provide high-quality training data for subsequent models), the single-element prompt words, the conference processing text, and the single element can be spliced together to obtain the training data of the single element (i.e., any one of the first target element and the second target element); the double-element prompt words, the conference processing text, and the double elements can be spliced together to obtain the training data of the double elements (i.e., any two different elements of the first target element and the second target element); the three-element prompt words, the conference processing text, and the three elements can be spliced together to obtain the training data of the three elements (i.e., any three different elements of the first target element and the second target element); the four-element prompt words, the conference processing text, and the four elements can be spliced together to obtain the training data of the four elements (i.e., any four different elements of the first target element and the second target element).
[0097] In this embodiment, each element has a sequential level in the prompt word (such as the level of the meeting title is higher than the level of the full text summary, the level of the full text summary is higher than the level of the key content, and the level of the key content is higher than the level of the to-do list). Meeting data can be constructed based on single-element training data, double-element training data, three-element training data, and four-element training data.
[0098] Exemplarily, different element prompt words under the single-person scenario conference type may be as shown in Table 1.
[0099] Table 1
[0100]
[0101] Exemplarily, different element prompt words under the multi-person scenario conference type may be as shown in Table 2.
[0102] Table 2
[0103]
[0104]
[0105] In order to accurately obtain the preset meeting minutes model, in the method for constructing meeting data provided in Example 1 of the present application, a mixed data set is constructed based on training data and preset data; the preset model is trained using the mixed data set, and the text expansion technology is used to adjust the length of the processed text output by the preset model until the number of training rounds is equal to the preset round threshold, so as to obtain the preset meeting minutes model, wherein the preset meeting minutes model generates meeting minutes based on the original text of the meeting.
[0106] In this embodiment, in order to prevent the model from overfitting the meeting minutes data and losing the general understanding ability of the large language model, a mixed data set (including training data and preset data (i.e., a general data set not related to meeting minutes)) can be constructed, and the mixed data set can be used to train the preset model (such as a pre-trained general large language model (i.e., a model trained with large-scale unlabeled text data)), and text expansion technology (such as NTK (Neural TangentKernel) extrapolation method) is used to adjust the length of the processed text output by the preset model until the number of training rounds is equal to the preset round threshold, so as to obtain a preset meeting minutes model (i.e., a model that can generate meeting minutes based on the original text of the meeting).
[0107] Exemplarily, the preset model can process texts with a length of 8192 tokens (ie, the smallest unit into which the text is decomposed). By adopting the text expansion technology, the text generation capability of the preset meeting minutes model is expanded to 16384 tokens.
[0108] Figure 2 is an optional flow chart for constructing conference data according to an embodiment of the present invention, such as Figure 2 As shown, first, the initial file can be obtained online or offline, and then the obtained initial file is preprocessed to obtain the conference processing text. Then, according to the conference processing text, the scene classification is performed, and when the conference type is a single-person no-script scene type or a multi-person scene meeting type, a prompt (i.e., a prompt word) is constructed to call the large model to generate the initial elements, and then a construction strategy can be formulated, wherein the construction strategy may include: conference minutes output strategy, conference minutes correction strategy, single-person verbatim script scene (i.e., single-person manuscript scene type) construction strategy, and the conference minutes output strategy includes: content requirements (such as general content requirements, general language expression requirements, specific content requirements and expression requirements of each element), word count requirements, and format requirements. According to these construction strategies, the conference minutes can be output, and the training data can be determined according to the different elements in the conference minutes. Then, the preset model is trained using the training data to obtain the preset conference minutes model.
[0109] In an optional embodiment, the following may be used: Figure 2 The process shown executes the method of constructing conference data.
[0110] In an embodiment of the present invention, by distinguishing different types of conference scenarios (such as single-person conference scenario types (including scenarios with manuscripts and scenarios without manuscripts) and multi-person conference scenario types), conference data can be constructed according to different conference scenario types, and the initial elements generated by the large model can be corrected to ensure that the minutes are detailed and the format is consistent, balancing accuracy and efficiency. The standardized content standard strategy clarifies the minutes structure, ensures stable output, comprehensive coverage of key points, clear logic, and achieves accurate conversion from spoken to written language, thereby optimizing the generation of meeting minutes and improving the quality of meeting data.
[0111] The following is a detailed description in conjunction with another embodiment.
[0112] Embodiment 2
[0113] A device for constructing conference data provided in this embodiment includes multiple implementation units, each of which corresponds to each implementation step in the above-mentioned embodiment 1.
[0114] Figure 3 is a schematic diagram of an optional device for constructing conference data according to an embodiment of the present invention, such as Figure 3 As shown, the device for constructing conference data may include: a processing unit 30 , a determining unit 31 , an extracting unit 32 , a correcting unit 33 , and a splicing unit 34 .
[0115] The processing unit 30 is used to obtain the original audio of the conference, recognize the original audio of the conference, obtain the original text of the conference, and pre-process the original text of the conference to obtain the processed text of the conference;
[0116] A determination unit 31, configured to determine a conference type based on the conference processing text;
[0117] The extraction unit 32 is used to segment the conference processing text based on the conference type to obtain a preset number of conference processing sub-texts, and extract elements from the conference processing sub-texts based on the first preset prompt word to obtain initial elements;
[0118] A correction unit 33, configured to correct the initial element based on the conference processing text to obtain a first target element;
[0119] The splicing unit 34 is used to splice the conference processing text, the first target element and the second preset prompt word to complete the construction of the conference data.
[0120] The above-mentioned device for constructing conference data can obtain the original audio of the conference through the processing unit 30, recognize the original audio of the conference to obtain the original text of the conference, and pre-process the original text of the conference to obtain the conference processing text. The determination unit 31 determines the type of conference based on the conference processing text. The extraction unit 32 segments the conference processing text based on the conference type to obtain a preset number of conference processing sub-texts, and extracts elements from the conference processing sub-text based on the first preset prompt word to obtain the initial element. The correction unit 33 corrects the initial element based on the conference processing text to obtain the first target element. The splicing unit 34 splices the conference processing text, the first target element and the second preset prompt word to complete the construction of the conference data.
[0121] Optionally, the processing unit includes: a first determination module, used to determine the order of speakers in the original text of the meeting; a second determination module, used to determine the speaker number corresponding to each speaker based on the speaker order; a first processing module, used to replace the inventor name of each speaker in the original text of the meeting with the speaker number, and delete the speech time information corresponding to each speaker number to obtain a meeting processing text, wherein the meeting processing text includes: multiple speech contents associated with the speaker number.
[0122] Optionally, the conference types include at least: single-person scenario conference type and multi-person scenario conference type, the single-person scenario conference types include at least: single-person manuscript scenario type and single-person no-manuscript scenario type, the single-person manuscript scenario type is a scenario type in which a single person speaks based on a manuscript, and the single-person no-manuscript scenario type is a scenario type in which a single person speaks without a manuscript, and the determination unit includes: a third determination module, used to determine the number of preset numbers based on the conference processing text, wherein the preset number of numbers is the number of speaker numbers; a fourth determination module, used to determine that the conference type is a single-person scenario conference type when the preset number of numbers is less than a preset threshold; a fifth determination module, used to determine whether there is a preset language feature in the conference processing text when the conference type is a single-person scenario conference type; a sixth determination module, used to determine that the conference type is a single-person manuscript scenario type when the preset language feature exists in the conference processing text; a seventh determination module, used to determine that the conference type is a single-person no-manuscript scenario conference type when the preset language feature does not exist in the conference processing text.
[0123] Optionally, the initial element includes at least: a meeting title, key content, the key content includes: multiple titles and text content corresponding to each title, and the correction unit includes: a first correction module, used to determine the content standard strategy of the initial element, and based on the content standard strategy, correct the key content of the initial element to obtain the target key content, wherein the content standard strategy at least includes: a strategy for determining whether the key content of the initial element is consistent with the content of the meeting processing text, a strategy for determining whether the sentence order of the initial element is consistent with the sentence order of the meeting processing text, a strategy for determining that there are no repeated sentences in the initial element, and a strategy for determining that there are no pronouns in the initial element; a first deletion module, used to delete numbers when there are numbers in the meeting title to obtain the target meeting title; a first adjustment module, used to adjust the format of the initial element based on a preset format to obtain the target format; an eighth determination module, used to determine the first target element based on the target key content, the target meeting title and the target format.
[0124] Optionally, the first correction module includes: a first modification submodule, which is used to modify the key content inconsistent with the content in the conference processing text to the content in the conference processing text when the key content of the initial element is inconsistent with the content in the conference processing text; a first sorting submodule, which is used to reorder the sentences in the initial element based on the sentence order in the conference processing text when the sentence order of the initial element is inconsistent with the sentence order in the conference processing text; a first merging submodule, which is used to merge the repeated sentences when there are repeated sentences in the initial element; a first replacement submodule, which is used to replace the pronouns when there are pronouns in the initial element; and a first supplementing submodule, which is used to compare the number of words in the conference processing text with the number of words in the key content in the initial element, and when it is detected that the number of words in the key content is less than a preset word count threshold, supplement the number of words in the key content to obtain the target key content.
[0125] Optionally, the construction device also includes: a first division module, which is used to hierarchically divide the conference processing text when the conference type is a single-person manuscript scene type, and obtain multiple text titles of different levels and conference processing text contents corresponding to different levels; a first characterization module, which is used to characterize each first-level text title as a title under key content when the number of text titles at the first level is greater than the preset number of titles; a second characterization module, which is used to characterize each second-level text title as a title under key content when the number of text titles at the first level is less than the preset number of titles, wherein the level of the first level is greater than that of the second level; a ninth determination module, which is used to extract key phrases from the conference processing text content based on each title, determine the text content corresponding to each title, and determine the key content based on all titles and all text contents; a tenth determination module, which is used to determine the meeting title and full-text summary of the conference processing text based on the key content; and an eleventh determination module, which is used to determine the second target element based on the key content, the meeting title and the full-text summary.
[0126] Optionally, the initial element also includes: a summary of the full text, a to-do list, the second preset prompt words include: a single-element prompt word, a double-element prompt word, a three-element prompt word, and a four-element prompt word, and the splicing unit includes: a first splicing module, used to splice the single-element prompt word, the conference processing text, and the single element to obtain single-element training data, wherein the single element is any element of the first target element and the second target element, and the single-element prompt word is the prompt word of any element; a second splicing module, used to splice the double-element prompt word, the conference processing text, and the double elements to obtain double-element training data, wherein the double elements are any two different elements of the first target element and the second target element, and the double-element prompt word is the prompt word of any two different elements; a third splicing module, used to splice the three-element prompt word, the conference processing text, and the three elements, A three-element training data is obtained, wherein the three elements are any three different elements of the first target element and the second target element, and the three-element prompt words are prompt words of any three different elements; a fourth splicing module is used to splice the four-element prompt words, the conference processing text, and the four elements to obtain the four-element training data, wherein the four elements are any four different elements of the first target element and the second target element, and the four-element prompt words are prompt words of any four different elements, and each element has a sequence level in the prompt words, and the sequence level is that the level of the conference title is greater than the level of the full text summary, the level of the full text summary is greater than the level of the key content, and the level of the key content is greater than the level of the to-do list; a first construction module is used to construct conference data based on the single-element training data, the double-element training data, the three-element training data, and the four-element training data.
[0127] Optionally, the construction device also includes: a second construction module, which is used to construct a mixed data set based on the training data and the preset data after splicing the meeting processing text, the first target element and the second preset prompt word to complete the construction of the meeting data; a first training module, which is used to train the preset model using the mixed data set, and use text expansion technology to adjust the length of the processed text output by the preset model until the number of training rounds is equal to the preset round threshold, so as to obtain a preset meeting minutes model, wherein the preset meeting minutes model generates meeting minutes based on the original text of the meeting.
[0128] The above-mentioned conference data construction device may also include a processor and a memory. The above-mentioned processing unit 30, determination unit 31, extraction unit 32, correction unit 33, splicing unit 34, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.
[0129] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the conference processing text, the first target element and the second preset prompt word are spliced by adjusting the kernel parameters to complete the construction of the conference data.
[0130] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip.
[0131] According to another aspect of an embodiment of the present invention, a computer program product is also provided, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and the computer program, when executed by a processor, implements any of the above-mentioned methods for constructing conference data.
[0132] When the computer program product is executed on a data processing device, it is suitable for executing an initialization program having the following method steps: obtaining the original audio of the meeting, recognizing the original audio of the meeting to obtain the original text of the meeting, and preprocessing the original text of the meeting to obtain the processed text of the meeting, determining the type of meeting based on the meeting processed text, segmenting the meeting processing text based on the meeting type to obtain a preset number of meeting processing sub-texts, and extracting elements from the meeting processing sub-texts based on a first preset prompt word to obtain an initial element, correcting the initial element based on the meeting processing text to obtain a first target element, splicing the meeting processing text, the first target element and the second preset prompt word to complete the construction of the meeting data.
[0133] According to another aspect of an embodiment of the present invention, there is also provided an electronic device, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement the above-mentioned method for constructing conference data.
[0134] Figure 4 1 is a hardware structure block diagram of an electronic device (or mobile device) for a method for constructing conference data according to an embodiment of the present invention. Figure 4 As shown, the electronic device may include one or more processors (e.g., Figure 4 The processor 402a, processor 402b, ..., processor 402n, etc., which may include but are not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 404 for storing data. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a keyboard, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 4 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 4 More or fewer components as shown, or with Figure 4 Different configurations are shown.
[0135] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0136] The embodiments or examples of the present disclosure are not exhaustive, but are only illustrative of some embodiments or examples, and are not intended to be specific limitations on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment or example can be implemented as an independent example, and the steps can be combined arbitrarily. For example, the scheme after removing some steps in a certain embodiment or example can also be implemented as an independent example, and the order of the steps in a certain embodiment or example can be arbitrarily exchanged. In addition, the optional methods or optional examples in a certain embodiment or example can be combined arbitrarily; in addition, the various embodiments or examples can be combined arbitrarily, for example, some or all steps of different embodiments or examples can be combined arbitrarily, and a certain embodiment or example can be combined arbitrarily with the optional methods or optional examples of other embodiments or examples.
[0137] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0138] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0139] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0140] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0142] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for constructing conference data, characterized in that: include: Acquire original conference audio, recognize the original conference audio to obtain original conference text, and pre-process the original conference text to obtain conference processed text; Determining the type of meeting based on the meeting processing text; Based on the conference type, the conference processing text is segmented to obtain a preset number of conference processing sub-texts, and based on the first preset prompt word, elements of the conference processing sub-texts are extracted to obtain initial elements; Based on the conference processing text, correcting the initial element to obtain a first target element; The conference processing text, the first target element and the second preset prompt word are spliced to complete the construction of the conference data.
2. The method for constructing conference data according to claim 1, characterized in that: The step of preprocessing the original conference text to obtain the conference processed text includes: determining the order of speakers in the original transcript of said conference; Based on the speaker order, determining a speaker number corresponding to each speaker; The inventor name of each speaker in the original conference text is replaced with the speaker number, and the speech time information corresponding to each speaker number is deleted to obtain the conference processing text, wherein the conference processing text includes: multiple speech contents associated with the speaker number.
3. The method for constructing conference data according to claim 1, characterized in that: The conference types include at least: a single-person scenario conference type and a multi-person scenario conference type, the single-person scenario conference types include at least: a single-person manuscript scenario type and a single-person no-manuscript scenario type, the single-person manuscript scenario type is a scenario type in which a single person speaks based on a manuscript, and the single-person no-manuscript scenario type is a scenario type in which a single person speaks without a manuscript. Based on the conference processing text, the step of determining the conference type includes: Based on the conference processing text, determining a preset number number, wherein the preset number number is the number of speaker numbers; When the number of the preset numbers is less than the preset threshold, determining the conference type as the single-person scenario conference type; In the case where the conference type is the single-person scenario conference type, determining whether there is a preset language feature in the conference processing text; In the case where the preset language feature exists in the conference processing text, determining the conference type as the single-person manuscript scene type; When the preset language feature does not exist in the conference processing text, the conference type is determined to be the single-person, non-script scenario conference type.
4. The method for constructing conference data according to claim 1, characterized in that: The initial element at least includes: a conference title and key content, wherein the key content includes: a plurality of titles and text content corresponding to each title. Based on the conference processing text, the initial element is corrected to obtain a first target element, including: Determine the content standard strategy of the initial element, and based on the content standard strategy, correct the key content of the initial element to obtain the target key content, wherein the content standard strategy at least includes: a strategy for determining whether the key content of the initial element is consistent with the content of the conference processing text, a strategy for determining whether the sentence sequence of the initial element is consistent with the sentence sequence of the conference processing text, a strategy for determining that there are no repeated sentences in the initial element, and a strategy for determining that there are no pronouns in the initial element; If there is a number in the conference title, delete the number to obtain the target conference title; Based on a preset format, adjusting the format of the initial element to obtain a target format; The first target element is determined based on the target key content, the target meeting title, and the target format.
5. The method for constructing conference data according to claim 4, characterized in that: Based on the content standard strategy, the key content of the initial element is corrected to obtain the target key content, including: In the case where the key content of the initial element is inconsistent with the content in the conference processing text, modifying the key content that is inconsistent with the content in the conference processing text to the content in the conference processing text; In the case where the sentence order of the initial element is inconsistent with the sentence order in the conference processing text, reordering the sentences in the initial element based on the sentence order in the conference processing text; In the case where there are repeated sentences in the initial element, merging the repeated sentences; If the pronoun exists in the initial element, replace the pronoun; The number of words in the conference processing text is compared with the number of words in the key content in the initial element. When it is detected that the number of words in the key content is less than a preset word count threshold, the number of words in the key content is supplemented to obtain the target key content.
6. The method for constructing conference data according to claim 3, characterized in that: The method for constructing conference data also includes: In the case where the conference type is the single-person manuscript scenario type, the conference processing text is hierarchically divided to obtain a plurality of text titles at different levels and conference processing text contents corresponding to the different levels; In the case that the number of the text titles at the first level is greater than the preset number of titles, each of the text titles at the first level is characterized as the title under the key content; or, In the case that the number of the text titles at the first level is less than the preset number of titles, characterizing each of the text titles at the second level as the title under the key content, wherein the level of the first level is greater than that of the second level; Based on each of the titles, extract key phrases from the conference processing text content to determine the text content corresponding to each of the titles, and determine the key content based on all of the titles and all of the text content; Based on the key content, determine the conference title and full text summary of the conference processing text; A second target element is determined based on the key content, the conference title, and the full text summary.
7. The method for constructing conference data according to claim 6, characterized in that: The initial element also includes: a full text summary and a to-do list; the second preset prompt words include: a single-element prompt word, a double-element prompt word, a triple-element prompt word, and a quadruple-element prompt word; the conference processing text, the first target element, and the second preset prompt word are spliced to complete the steps of constructing the conference data, including: splicing the single element prompt word, the conference processing text, and a single element to obtain training data of the single element, wherein the single element is any one of the first target element and the second target element, and the single element prompt word is a prompt word of any one of the elements; splicing the dual-element prompt word, the conference processing text, and the dual elements to obtain the dual-element training data, wherein the dual elements are any two different elements of the first target element and the second target element, and the dual-element prompt word is the prompt word of the any two different elements; splicing the three-element prompt word, the conference processing text, and the three elements to obtain the training data of the three elements, wherein the three elements are any three different elements of the first target element and the second target element, and the three-element prompt word is the prompt word of the any three different elements; The four-element prompt word, the conference processing text, and the four elements are spliced to obtain the training data of the four elements, wherein the four elements are any four different elements of the first target element and the second target element, the four-element prompt word is the prompt word of the any four different elements, and each of the elements has a sequence level in the prompt word, and the sequence level is that the level of the conference title is greater than the level of the full text summary, the level of the full text summary is greater than the level of the key content, and the level of the key content is greater than the level of the to-do list; The conference data is constructed based on the single-element training data, the double-element training data, the triple-element training data, and the quadruple-element training data.
8. The method for constructing conference data according to claim 7, characterized in that: After the conference processing text, the first target element and the second preset prompt word are spliced to complete the construction of the conference data, the method further includes: Constructing a mixed data set based on the training data and the preset data; The preset model is trained using the mixed data set, and the text expansion technology is used to adjust the length of the processed text output by the preset model until the number of training rounds is equal to the preset round threshold, thereby obtaining a preset meeting minutes model, wherein the preset meeting minutes model generates meeting minutes based on the original meeting text.
9. A device for constructing conference data, characterized in that: include: A processing unit, used to obtain original conference audio, recognize the original conference audio to obtain original conference text, and pre-process the original conference text to obtain conference processed text; A determination unit, configured to determine a conference type based on the conference processing text; An extraction unit, configured to segment the conference processing text based on the conference type to obtain a preset number of conference processing sub-texts, and extract elements from the conference processing sub-texts based on a first preset prompt word to obtain initial elements; A correction unit, configured to correct the initial element based on the conference processing text to obtain a first target element; The splicing unit is used to splice the conference processing text, the first target element and the second preset prompt word to complete the construction of the conference data.
10. A computer program product, characterized in that It comprises a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for constructing conference data according to any one of claims 1 to 8 is implemented.
11. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for constructing conference data as described in any one of claims 1 to 8.