Meeting minutes processing method and related device

By identifying and correcting hot words in meeting minutes, high-quality target meeting minutes are generated, solving the problem of erroneous words in existing technologies and improving the accuracy of meeting minutes and user experience.

WO2026065238A1PCT designated stage Publication Date: 2026-04-02BOE TECHNOLOGY GROUP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

The existing meeting minutes contain many erroneous words, affecting their quality.

Method used

By acquiring the target text of the original meeting transcript, identifying and determining hot words, and using these hot words to correct the original meeting transcript, a high-quality target meeting transcript is generated.

Benefits of technology

It improved the accuracy of meeting minutes, reduced word errors, and enhanced the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024122346_02042026_PF_FP_ABST
    Figure CN2024122346_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a meeting minutes processing method and a related device, for use in improving the quality of meeting minutes, so that the meeting minutes have fewer wrong words. The meeting minutes processing method may comprise: acquiring original meeting minutes, wherein the original meeting minutes are obtained by performing speech-to-text conversion on the basis of a target meeting audio; acquiring a target text, wherein the target text is a dialog text for the original meeting minutes; acquiring a first hot word on the basis of the target text; and correcting part or all of the original meeting minutes on the basis of a target hot word to obtain target meeting minutes, wherein the target hot word comprises the first hot word.
Need to check novelty before this filing date? Find Prior Art

Description

Conference record processing method and related device TECHNICAL FIELD

[0001] The present application relates to the technical field of conference, in particular to a conference record processing method and related device. BACKGROUND

[0002] Conference record is a kind of objective documentary material, records the speech of the participants. Usually by the record personnel collation, form conference record. With the development of intelligent speech technology, the conference audio is transcribed, and the conference record can be obtained, but there are many error words in the conference record.

[0003] SUMMARY

[0004] The present application provides a conference record processing method and related device, which can improve the quality of conference record and has fewer error words.

[0005] In a first aspect, the present application provides a conference record processing method, which can include:

[0006] Obtaining an original conference record, the original conference record is transcribed based on target conference audio;

[0007] Obtaining a target text, the target text is a dialogue text of the original conference record;

[0008] Obtaining a first hot word based on the target text;

[0009] Based on the target hot word, the original conference record is partially or wholly corrected to obtain a target conference record, and the target hot word includes the first hot word.

[0010] In the present application, the dialogue text can include a question text or an instruction text. The dialogue text can be provided by a user, which can reflect the user's attention to the part of the original conference record. The hot word is extracted from the content of the user's attention, and the original conference record is corrected to obtain the target conference record with improved quality and fewer error words.

[0011] In a possible implementation, the present application provides a conference record processing method, which further includes:

[0012] Based on the target conference record, the related content of the dialogue text is output.

[0013] In the present application, the related content corresponding to the dialogue text is output by using the corrected target conference record, which can improve the quality of the output related content and improve the user experience.

[0014] In a possible implementation, the conference record processing method provided by the embodiments of the present application comprises the following steps.

[0015] The target text is processed by word segmentation to obtain a plurality of words.

[0016] In the plurality of words, a word of a preset word type is determined as a candidate word.

[0017] According to a selection operation of the first candidate word, the first candidate word is determined as the first hot word, the first candidate word being one or more of the determined candidate words.

[0018] In a possible implementation, the conference record processing method provided by the embodiments of the present application comprises the following steps.

[0019] The target text is processed by word segmentation to obtain a plurality of words.

[0020] In the plurality of words, a word of a preset word type is determined as the first hot word.

[0021] In a possible implementation, the conference record processing method provided by the embodiments of the present application comprises the following steps.

[0022] The dialogue audio of the target conference is obtained, and dialogue text corresponding to the dialogue audio is determined; or,

[0023] The target text is received.

[0024] In a possible implementation, the conference record processing method provided by the embodiments of the present application comprises the following steps.

[0025] A second hot word is obtained based on the original conference record, and the target hot word further comprises the second hot word.

[0026] In a possible implementation, the conference record processing method provided by the embodiments of the present application comprises the following steps.

[0027] At least one group of similar words in the original conference record is determined, the pinyin of each candidate word in a group of similar words is the same except for tone, or the pinyin of each candidate word in a group of similar words meets a preset similar word rule, or the audio of each candidate word in a group of similar words is similar.

[0028] Based on any one group of similar words, a second hot word corresponding to the any one group of similar words is determined.

[0029] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following steps.

[0030] Displaying the candidate words in any group of similar words;

[0031] According to the selection operation on the second candidate word, the second candidate word is determined as the second hot word, the second candidate word being one or more candidate words in the any group of similar words.

[0032] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following steps.

[0033] According to the input word operation, the input word is determined as the second hot word.

[0034] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following steps.

[0035] Indicating the position of each candidate word in the any group of similar words in the original conference record.

[0036] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following steps.

[0037] Performing new word discovery on the original conference record, and taking the new word as the second hot word.

[0038] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following steps.

[0039] Performing hot word identification on the relevant content of the dialogue text in the original conference record to obtain the second hot word.

[0040] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following steps.

[0041] Based on the target hot word, the relevant text of the target text in the original conference record is modified to obtain the target conference record.

[0042] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following steps.

[0043] transcribe the target conference audio based on the target hotword to obtain the target conference record.

[0044] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes:

[0045] generating target audio of the target hotword;

[0046] determining similarity between the target audio and an audio segment corresponding to a target word in the original conference record, wherein the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio;

[0047] if the similarity between the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, changing the target word in the text corresponding to the audio segment to the target hotword in the original conference record to obtain the target conference record.

[0048] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes that the target word has a similar relationship with the target hotword. The target word has a similar relationship with the target hotword can mean that the pinyin of the target word and the target hotword is the same except for tone, or the pinyin of the target word and the target hotword conforms to a preset similar word rule.

[0049] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes:

[0050] based on the target hotword, performing correction processing on the original conference record by a first correction mode to obtain a first conference record;

[0051] based on the target hotword, performing correction processing on the original conference record by a second correction mode to obtain a second conference record;

[0052] comparing whether the sentences at each position of the first conference record and the second conference record are consistent;

[0053] For the inconsistent position of the sentence, the fluency of the sentence of the position in the first conference record and the sentence of the position in the second conference record is respectively evaluated, and the target conference record is obtained based on the most fluent sentence of the sentence of the position in the first conference record and the sentence of the position in the second conference record.

[0054] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following two modification methods:

[0055] Method one, based on the target hot word, the relevant text of the dialogue text in the original conference record is modified to obtain the first conference record or the second conference record.

[0056] Method two, based on the target hot word, the target conference audio is transcribed to obtain the first conference record or the second conference record.

[0057] Method three, a target audio of the target hot word is generated, and a similarity of an audio segment corresponding to a target word in the target audio is determined, wherein the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio, the target word and the target hot word have a close relationship; and if the similarity of the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, the target word in the text corresponding to the audio segment in the original conference record is changed to the target hot word to obtain the first conference record or the second conference record.

[0058] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following two modification methods:

[0059] The method further includes:

[0060] Receiving a target file;

[0061] Performing hot word recognition on the target file to obtain a plurality of candidate words;

[0062] According to a selection operation of one or more word terms in the plurality of candidate words, the one or more word terms are determined as the preconfigured hot words.

[0063] In a possible implementation, the conference record processing method provided by the embodiment of the present application includes the following two modification methods:

[0064] display a first interaction interface, the first interaction interface including a first display area and a second display area, the first display area being configured to display a conference record, and the second display area being configured to display a hot word;

[0065] display a sentence containing any hot word in the original conference record in the first display area, and indicate the any hot word;

[0066] display at least one group of similar words to which the any hot word belongs in the second display area, each candidate word in a group of similar words being a word in the original conference record, the pinyin of each candidate word in a group of similar words being identical except for tones, or the pinyin of each candidate word in a group of similar words meeting a preset similar word rule, or the audio of each candidate word in a group of similar words being similar.

[0067] In a possible implementation, the conference record processing method provided by the embodiment of the present application further includes that the first interaction interface further includes a hot word input box; and the method further includes:

[0068] obtaining content in the hot word input box, the target hot word further including the content in the hot word input box.

[0069] In a second aspect, the embodiment of the present application further provides a conference record processing system, configured to:

[0070] obtain an original conference record, the original conference record being obtained based on target conference audio by speech transcription.

[0071] obtain a target text, the target text being a dialogue text of the original conference record;

[0072] obtain a first hot word based on the target text;

[0073] correct part or all of the original conference record based on the target hot word, and obtain a target conference record, the target hot word including the first hot word.

[0074] In a possible implementation, the conference record processing system is further configured to:

[0075] output related content of the dialogue text based on the target conference record.

[0076] In a possible implementation, in the operation of obtaining the first hot word based on the target text, the conference record processing system is specifically configured to:

[0077] perform word segmentation processing on the target text to obtain a plurality of words;

[0078] determine a word with a preset word type as a candidate word in the plurality of words.

[0079] According to the selection operation on the first candidate word, the first candidate word is determined as the first hot word, and the first candidate word is one or more candidate words in the determined candidate words.

[0080] In a possible implementation, in the operation of obtaining the first hot word based on the target text, the conference record processing system is specifically configured to:

[0081] perform a word segmentation processing on the target text to obtain a plurality of words;

[0082] In the plurality of words, a word of a preset word type is determined as the first hot word.

[0083] In a possible implementation, in the operation of obtaining the target text, the conference record processing system is specifically configured to:

[0084] obtain a dialogue audio of the target conference and determine a dialogue text corresponding to the dialogue audio; or

[0085] receive the target text.

[0086] In a possible implementation, the conference record processing system is further configured to:

[0087] obtain a second hot word based on the original conference record, and the target hot word further includes the second hot word.

[0088] In a possible implementation, in the operation of obtaining the second hot word based on the original conference record, the conference record processing system is specifically configured to:

[0089] determine at least one group of similar words in the original conference record, the pinyin of each candidate word in a group of similar words is the same except for tone, or the pinyin of each candidate word in a group of similar words meets a preset similar word rule, or the audio of each candidate word in a group of similar words is similar;

[0090] determine a second hot word corresponding to any group of similar words based on the any group of similar words.

[0091] In a possible implementation, in the operation of determining the second hot word corresponding to the any group of similar words based on the any group of similar words, the conference record processing system is specifically configured to:

[0092] display a candidate word in any group of similar words;

[0093] According to the selection operation on the second candidate word, the second candidate word is determined as the second hot word, and the second candidate word is one or more candidate words in the any group of similar words.

[0094] In a possible implementation, in the operation of determining the second hotword corresponding to the any set of similar words based on the any set of similar words, the conference record processing system is specifically configured to:

[0095] determine the input word as the second hotword according to the operation of the input word.

[0096] In a possible implementation, the conference record processing system is further configured to:

[0097] indicate the position of each candidate word in the any set of similar words in the original conference record.

[0098] In a possible implementation, in the operation of obtaining the second hotword based on the original conference record, the conference record processing system is specifically configured to:

[0099] perform new word discovery on the original conference record, and take the new word as the second hotword.

[0100] In a possible implementation, in the operation of obtaining the second hotword based on the original conference record, the conference record processing system is specifically configured to:

[0101] perform hotword recognition on the relevant content of the dialogue text in the original conference record to obtain the second hotword.

[0102] In a possible implementation, in the operation of modifying part or all of the original conference record based on the target hotword to obtain the target conference record, the conference record processing system is specifically configured to:

[0103] modify the relevant text of the target text in the original conference record based on the target hotword to obtain the target conference record.

[0104] In a possible implementation, in the operation of modifying the original conference record based on the target hotword to obtain the target conference record, the conference record processing system is specifically configured to:

[0105] perform speech transcription on the target conference audio based on the target hotword to obtain the target conference record.

[0106] In a possible implementation, in the operation of modifying part or all of the original conference record based on the target hotword to obtain the target conference record, the conference record processing system is specifically configured to:

[0107] generate target audio of the target hotword.

[0108] determine a similarity between the target audio and an audio segment corresponding to a target word in the original conference record, wherein the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio, and the target word has a close relationship with the target hot word;

[0109] if the similarity between the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, change the target word in the text corresponding to the audio segment to the target hot word in the original conference record to obtain the target conference record.

[0110] In a possible implementation, in the operation of correcting part or all of the original conference record based on the target hot word to obtain the target conference record, the conference record processing system is specifically configured to:

[0111] correct the original conference record based on the target hot word by using a first correction manner to obtain a first conference record;

[0112] correct the original conference record based on the target hot word by using a second correction manner to obtain a second conference record;

[0113] compare whether the sentences at each position of the first conference record and the second conference record are consistent;

[0114] for the position where the sentences are inconsistent, respectively perform fluency evaluation on the sentences at the position in the first conference record and the sentences at the position in the second conference record, and take the most fluent sentence from the sentences at the position in the first conference record and the sentences at the position in the second conference record as the sentence at the position in the third conference record to obtain the target conference record, wherein the third conference record is the first conference record or the second conference record.

[0115] In a possible implementation, the first correction manner and the second correction manner are any two of the following correction manners:

[0116] manner one, based on the target hot word, correct the relevant text of the dialogue text in the original conference record to obtain the first conference record or the second conference record;

[0117] manner two, based on the target hot word, perform speech transcription on the target conference audio to obtain the first conference record or the second conference record;

[0118] In a third mode, a target audio of the target hotword is generated; similarity of an audio segment corresponding to a target word to the target audio is determined, the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio, the target word has a close relationship with the target hotword; and if the similarity of the audio segment to the target audio satisfies a preset relationship with a preset similarity threshold, the target word in the text corresponding to the audio segment is changed to the target hotword in the original conference record, to obtain the first conference record or the second conference record.

[0119] In a possible implementation, the original conference record is obtained by performing speech transcription on the target conference audio based on a preconfigured hotword, and the target hotword further includes the preconfigured hotword.

[0120] receiving a target file;

[0121] performing hotword recognition on the target file to obtain a plurality of candidate words;

[0122] determining one or more words in the plurality of candidate words as the preconfigured hotword according to a selection operation on the one or more words.

[0123] In a possible implementation, the conference record processing system further includes:

[0124] displaying a first interactive interface, the first interactive interface including a first display area and a second display area, the first display area being configured to display a conference record, and the second display area being configured to display a hotword;

[0125] displaying a sentence containing any hotword in the original conference record in the first display area, and indicating the any hotword;

[0126] displaying at least one group of close words to which the any hotword belongs in the second display area, each candidate word in each group of close words being a word in the original conference record, a pinyin of each candidate word in a group of close words being identical except for a tone, or a pinyin of each candidate word in a group of close words meeting a preset close word rule, or an audio of each candidate word in a group of close words being similar.

[0127] In a possible implementation, the first interactive interface further includes a hotword input box, and the conference record processing system further includes:

[0128] obtaining content in the hotword input box, and the target hotword further includes the content in the hotword input box.

[0129] In a third aspect, an electronic device is provided, including a memory and a processor.

[0130] The memory stores computer program instructions.

[0131] The processor executes the instructions to implement the steps in the conference recording processing method as described in the first aspect and any possible implementation manner thereof.

[0132] In a fourth aspect, a computer storage medium is provided, where the instructions in the computer storage medium, when executed by a processor, cause the processor to perform the steps in the conference recording processing method as described in the first aspect and any possible implementation manner thereof. BRIEF DESCRIPTION OF DRAWINGS

[0133] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0134] FIG. 1 is a flowchart of a conference recording processing method provided by an embodiment of the present application;

[0135] FIG. 2 is a schematic diagram of an interactive interface provided by an embodiment of the present application;

[0136] FIG. 3a is a schematic diagram of an interactive interface provided by an embodiment of the present application;

[0137] FIG. 3b is a schematic diagram of an interactive interface provided by an embodiment of the present application;

[0138] FIG. 4 is a flowchart of a conference recording processing method provided by an embodiment of the present application;

[0139] FIG. 5 is a schematic diagram of a working process of a conference recording system provided by an embodiment of the present application;

[0140] FIG. 6 is a schematic diagram of a structure of a conference recording system provided by an embodiment of the present application;

[0141] FIG. 7 is a schematic diagram of a structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0142] At least one of the embodiments of the present application includes one or more; wherein the plurality means greater than or equal to two. In addition, it needs to be understood that in the description of the present application, the terms "first", "second", etc. are used only for the purpose of distinguishing the described from each other, and cannot be understood as indicating or implying relative importance. Also, it cannot be understood as indicating or implying an order. In the description of the present application, "A and / or B" includes three schemes, respectively, including A; including B; and including A and B.

[0143] In the above embodiments, according to the context, the term "when" or "after" can be interpreted as meaning "if" or "after" or "in response to determining" or "in response to detecting". Similarly, according to the context, the phrase "upon determining" or "if detecting (the stated condition or event)" can be interpreted as meaning "if determining" or "in response to determining" or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)". In addition, in the above embodiments, relational terms such as first, second, etc. are used to distinguish one entity from another, without limiting any actual relationship and order between the entities.

[0144] In the description of the present application, the reference to "one embodiment" or "some embodiments" means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in additional some embodiments" and the like appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0145] The "hot word" in the conference record processing method provided by the present application can refer to a keyword or phrase that needs to be accurately identified in a specific field, business, or profession. Generally, hot words have low usage frequency in scenarios other than specific fields, businesses, or professions, and may be rare words. However, in the context of specific fields, businesses, or professions, the usage frequency of hot words is higher than that in other contexts.

[0146] The "hot word" in the conference record processing method provided by the present application can be understood as a popular word. As a kind of word phenomenon, it reflects the problems and things that people in a country or a region generally pay attention to at a certain period of time. Hot words can have era characteristics and can reflect hot topics and livelihood issues at a certain period of time.

[0147] In a scenario of generating a conference record by a voice recognition manner, generally, a trained recognition model is used for voice recognition. The recognition effect often depends on the quantity of training data of the trained recognition model and the words in the training data. There can be words (or phrases) in the to-be-recognized voice that appear less frequently or not at all in the training data, such as some domain-specific terms, names in a user address book, and the like.

[0148] For conference audio, the generated conference record depends on the accuracy of the recognition model. The accuracy of the recognition model has a great influence on user experience or downstream tasks.

[0149] In view of this, the present application provides a conference record processing method and related equipment, which can improve the quality of the conference record.

[0150] FIG. 1 exemplarily shows a conference record processing method, which can be executed by a conference record processing system, and can include the following steps:

[0151] S101, obtaining an original conference record, which is obtained by voice transcription based on target conference audio.

[0152] S102, obtaining target text, which is a dialogue text of the original conference record.

[0153] S103, obtaining a first hotword based on the target text.

[0154] S104, correcting part or all of the original conference record based on the target hotword to obtain a target conference record, wherein the target hotword includes the first hotword.

[0155] In the embodiments of the present application, the original conference record can refer to a text form obtained by voice transcription based on the target conference audio. In some application scenarios, the target conference audio can be all audio of the target conference. It can be seen that the target conference can be a conference that has ended. In other application scenarios, the target conference audio can be audio collected in the target conference. The target conference can be an ongoing conference or an unfinished conference. The original conference record can include voice transcription content of all collected audio in the target conference.

[0156] In the embodiments of the present application, the target text is a dialogue text of the original conference record. The dialogue text can reflect the user's demand for the original conference record. For example, the demand for obtaining data from the conference record, the demand for generating a chart, the demand for finding the answer to a question, and the demand for finding a specified content.

[0157] The dialogue text can include question text and instruction text. The question text can be generally understood as text with a question nature, i.e., a question or a query text. For example, "What is the next generation communication technology?", "What is the total output value last year?". The instruction text can be understood as text with an instruction nature, i.e., an instruction or a command text. For example, "Generate a chart according to the historical total output value data in the meeting", "Find the content about the latest communication technology", etc.

[0158] In the embodiments of the present application, the related text of the target text can be the text in the meeting record. The related content of the target text can include the related text of the target text. The related content of the target text can include content generated based on the related text of the target text, such as generated text, chart, video, picture, audio, etc.

[0159] In some examples, the dialogue text can be implemented as prompt information input to a large language model. The large language model can output the related content of the target text. For example, the target text is a question text, and the related content of the target text can include the related text corresponding to the question text in the meeting record, or the related content of the target text can include content generated based on the answer text of the question text in the meeting record. For another example, the target text is an instruction text, and the related content of the target text can include the related text corresponding to the instruction text in the meeting record, or the related content of the target text can include a chart, a picture, an audio, a video, etc. generated based on the text related to the instruction text.

[0160] In other examples, the meeting record system can match the dialogue text with the meeting record according to a preset algorithm to obtain the related text of the dialogue text in the meeting record. Optionally, the meeting record system can generate the related content of the dialogue text based on the related text of the dialogue text in the meeting record.

[0161] In some possible designs, in the meeting record processing method in the above embodiments, the target text in step S102 is the dialogue text of the original meeting record. In some application scenarios, the target text can not be the dialogue text of the original meeting record, and the target text can be obtained from other acquisition manners. For example, the target text can be a text input by a user. For another example, the target text can be part or all of the content in the meeting record of a target meeting related conference. For another example, the target text can be the dialogue text of the related conference.

[0162] In a possible implementation, based on the above meeting record processing method, a meeting record processing system (hereinafter referred to as system) can use, but is not limited to, any one of the following ways to acquire the target text:

[0163] A1, obtaining dialogue audio of the target meeting, and determining dialogue text corresponding to the dialogue audio.

[0164] The system can have the function of collecting or receiving audio. The system can obtain dialogue audio of the target meeting. The system can perform speech transcription on the dialogue audio to obtain dialogue text.

[0165] A2, receiving the target text.

[0166] The system can have the function of receiving text. The system can provide a display interface. The user can input text through the dialogue input box of the display interface. The system can take the text input by the user in the dialogue input box as the target text.

[0167] In the above conference recording processing method, the system can obtain a first hot word based on the target text. The first hot word can represent a hot word obtained based on the target text. The number of first hot words can be one or more. Extracting hot words from the target text can automatically obtain hot words, and the target conference record obtained by correcting the original conference record using the hot words has high quality.

[0168] In one possible implementation, based on any of the above conference recording processing methods, in the operation of obtaining a first hot word based on the target text by the system, the following steps can be performed:

[0169] performing word segmentation processing on the target text to obtain a plurality of words;

[0170] determining a word with a preset part of speech as a candidate word in the plurality of words;

[0171] determining the first candidate word as the first hot word according to the selection operation of the first candidate word, the first candidate word being one or more of the determined candidate words.

[0172] The system can perform word segmentation processing on the target text to obtain a plurality of words. The system can perform part-of-speech tagging processing on the words. The plurality of words obtained after word segmentation processing can include words with a preset part of speech as candidate words. Optionally, the preset part of speech can include nouns and verbs.

[0173] For example, the target text can be implemented as "Please summarize the patent-related innovation points discussed today." The candidate words can include "today", "discussion", "patent", "innovation point", and "summary".

[0174] The user can select words from the candidate words, and the selected words are taken as the first hot word. The system can display all candidate words, and the user can select one or more candidate words therefrom. The selected candidate words are denoted as first candidate words. The system can determine the first candidate words as the first hot word.

[0175] In a possible implementation, based on any of the conference recording processing methods described above, in the operation of obtaining the first hot word based on the target text, the system can perform the following steps:

[0176] performing word segmentation on the target text to obtain a plurality of words;

[0177] determining, among the plurality of words, a word of a preset word type as the first hot word.

[0178] In a possible implementation, the target hot word in step S104 can also include a second hot word. The conference recording processing method can also include obtaining the second hot word based on the original conference recording. The system can obtain the second hot word based on the original conference recording after performing step S101 and before performing step S104.

[0179] The second hot word can represent a hot word obtained based on the original conference recording. The number of second hot words can be one or more. By performing hot word identification on the original conference recording and hot word identification on the target text, the target conference recording obtained by correcting the original conference recording with the identified hot words can reduce the occurrence of word errors.

[0180] In a possible design, in the operation of obtaining the second hot word based on the original conference recording, the system can perform the following steps:

[0181] determining at least one group of similar words in the original conference recording, the pinyin of each candidate word in a group of similar words excluding the tone part is the same, or the pinyin of each candidate word in a group of similar words meets a preset similar word rule, or the audio of each candidate word in a group of similar words is similar;

[0182] determining a second hot word corresponding to any of the groups of similar words based on the any of the groups of similar words.

[0183] In the at least one group of similar words, each candidate word is a word in the original conference recording.

[0184] In some examples, in the at least one group of similar words, the pinyin of each candidate word in a group of similar words excluding the tone part can be the same. The pinyin is composed of an initial, a medial, a final, and a tone.

[0185] In a possible case, the pinyin of some or all candidate words in a group of similar words excluding the tone part is the same, and the tone is also the same. The some or all candidate words are homophones. For example, the pinyin of word A and word B excluding the tone part is the same, and the tone is also the same. Word A and word B are homophones.

[0186] In another possible scenario, the part of the pinyin of some or all of the candidate words in a group of similar words, excluding the tone, is the same, and the tones are different. The some or all of the candidate words are homophonic words. For example, the part of the pinyin of word A and word B, excluding the tone, is the same, and the tones are also different. Word A and word B are homophonic words.

[0187] It can be seen that in a group of similar words, homophonic words and / or homophonic words of any candidate word can be included.

[0188] The system can perform word segmentation on the original meeting record, and form a character word list based on the word segmentation result. Optionally, the system can record the position of each word in the original meeting record. For example, the position of a word can be represented by the number of the word in the original meeting record.

[0189] The system can perform pinyin transcription on the words in the character word list to obtain a pinyin word list. The system can determine, according to the pinyin word list and the character word list, multiple words that are homophonic to each other and words that are homophonic to each other.

[0190] In some examples, in the at least one group of similar words, the pinyin of each candidate word in a group of similar words can be in accordance with a preset similar word rule. The pinyin is composed of an initial, a medial, a final, and a tone. The preset similar word rule can include an initial similar rule, and / or a final similar rule.

[0191] In the initial similar rule, the pronunciation of two initials having a similar relationship is similar. For example, in the initial similar rule, “r” and “l” have a similar relationship, and “zh” and “z” have a similar relationship. In the final similar rule, the pronunciation of two finals having a similar relationship is similar. For example, in the final similar rule, “en” and “eng” have a similar relationship, and “uang” and “ang” have a similar relationship. For another example, the pinyin of the word “innovation” is “chuang xin”, and the pinyin of the word “constant new” is “chang xin”. The two words are in accordance with the preset similar word rule. Optionally, the final similar rule can also be configured in combination with the tone, which is not limited in the embodiments of the present application.

[0192] The system can determine, according to the preset similar word rule, the pinyin word list, and the character word list, multiple words in accordance with the preset similar word rule.

[0193] In some examples, in the at least one group of similar words, the audio of each candidate word in a group of similar words can be similar.

[0194] Optionally, the system can synthesize the audio of the words in the original meeting record, or the system can extract the audio of the words from the target meeting audio. The audio of the words is recorded as audio segments for easy identification. The audio of the words in the original meeting record is compared for similarity. The acoustic features of the audio of the two words are calculated for similarity. The acoustic features can include, but are not limited to, any one of the following: fundamental frequency (F0), formants, short-time energy, mel-frequency cepstral coefficients (MFCC), etc. Optionally, the similarity of the acoustic features of the audio of the two words can be characterized by, but is not limited to, any one of the following parameters: Euclidean distance, cosine similarity, Manhattan distance, etc.

[0195] If the similarity of the acoustic features of the audio of the two words is greater than or equal to a preset similarity threshold, the audio of the two words has a similar relationship, i.e., the audio of the two words is similar. If the similarity of the acoustic features of the audio of the two words is less than the preset similarity threshold, the audio of the two words does not have a similar relationship, i.e., the audio of the two words is not similar.

[0196] In one possible design, when the system synthesizes the audio of the words in the original meeting record, the system can synthesize the audio of the words according to a preset timbre, or the system can synthesize the audio of the words according to the timbre of the speaker in the target meeting audio. The system can use voiceprint recognition technology to extract the timbre of the speaker from the target meeting audio.

[0197] In some examples, in the operation of determining, by the system, a second hotword corresponding to any one of the sets of similar words based on the any one of the sets of similar words, the following steps can be performed:

[0198] Displaying a candidate word in any one of the sets of similar words;

[0199] Determining, as the second hotword, a second candidate word in the any one of the sets of similar words according to a selection operation on the second candidate word, the second candidate word being one or more candidate words in the any one of the sets of similar words.

[0200] The system can display a candidate word in any one of the sets of similar words in a display interface. The user can perform a word selection operation through the display interface. For example, in a currently displayed set of similar words in the display interface, the selected word is a second candidate word. The user can select one or more words in the currently displayed set of similar words. The system can determine the second candidate word as the second hotword.

[0201] Optionally, the system can also indicate the position of each candidate word in the original meeting record in the any group of similar words. The system can display the sentence of each candidate word in the any group of similar words, and can indicate the any candidate word by one or more of the following indication manners: bold, enlarged font size, changed font, changed word color, changed word background color (highlight), etc. so that the user can know the word being currently corrected.

[0202] In a possible case, none of the candidate words in the group of similar words can be the word intended to be expressed by the user. The user can input the word intended to be expressed. The system can determine the input word as the second hot word according to the operation of inputting the word.

[0203] In a possible design, the system can perform the following operation in the operation of obtaining the second hot word based on the original meeting record: performing new word discovery on the original meeting record, and taking the new word as the second hot word.

[0204] In an actual application scenario, a new word can refer to a word created with the development of time, and can describe a newly emerged thing, phenomenon or concept. The word "green" and the word "economy" are two independent words, and each word can describe an implication. With the emergence of a new concept, a new word "green economy" is generated. The new word discovery can be understood as a processing manner of discovering a new word from a text. The process of performing the new word discovery in the embodiments of the present application can adopt any existing new word discovery algorithm, and the embodiments of the present application are not limited by multiple times.

[0205] The system can perform a word segmentation processing on the original meeting record to obtain a word list. Two words adjacent in position in a sentence are combined as a candidate word, or three words adjacent in position are combined as a candidate word. Based on mutual information and left-right entropy algorithm, an evaluation score of each candidate word is determined.

[0206] For a candidate word, the probability p(x, y) that two words (for example, word x and word y) in the candidate word appear together, and the probability of appearing alone, are respectively denoted as p(x) and p(y), and the mutual information of the candidate word is

[0207] The left entropy and the right entropy of the candidate word are calculated, and the minimum value of the left entropy and the right entropy is taken as the left-right entropy result. The left-right entropy of a word can be implemented by using an existing algorithm. For example, it can be assumed that the candidate word includes "word Pre word W". The left entropy of the two words is wherein p(PreW) represents the probability that the word Pre is the prefix of the word W. The greater the left-right entropy value of a word is, the greater the freedom degree of the word is, and the more likely the word is an independent word.

[0208] The mutual information of the candidate word is added with the sum of the left and right entropy results as an evaluation score of the candidate word. In some examples, the candidate word with an evaluation score greater than or equal to a preset score threshold is determined as a new word. The candidate word with an evaluation score less than the preset score threshold is not a new word. In other examples, the evaluation scores of all candidate words are sorted from large to small, and the candidate words in the top N orders are determined as new words.

[0209] In a possible design, in the operation of obtaining the second hot word based on the original meeting record, the system can perform: performing hot word identification on the relevant content of the dialogue text in the original meeting record to obtain the second hot word.

[0210] Based on the foregoing embodiments, the target text is the dialogue text about the original meeting record. The dialogue text can include the question text and the instruction text. The system can find the relevant text corresponding to the dialogue text in the original meeting record, for example, the answer text corresponding to the question text, or the relevant text corresponding to the instruction text.

[0211] In some examples, the system can input the dialogue text into the large language model, and the large language model can output the relevant text corresponding to the dialogue text.

[0212] In other examples, the system can match the dialogue text with the original meeting record. For example, the system can perform parsing processing on the dialogue text. The parsing processing can include, but is not limited to, word segmentation processing, stop word removal processing, sentence vector determination, word vector determination, and the like. Similarly, the system can perform parsing processing on the original meeting record. The system can match the parsing processing result of the dialogue text with the parsing processing result of the original meeting record, and take the sentence where the matched word is located or the matched sentence as the relevant text of the dialogue text in the original meeting record.

[0213] In other examples, the system can perform text slicing on the original meeting record. The similarity between the dialogue text and each sliced text is calculated respectively. If the similarity between a sliced text and the dialogue text is greater than or equal to a preset similarity threshold, the sliced text is the relevant text of the dialogue text. If the similarity between a sliced text and the dialogue text is less than the preset similarity threshold, the sliced text is not the relevant text of the dialogue text.

[0214] In other examples, the system can include a pre-trained text model. The training data used when training the text model includes the dialogue text and the position of the relevant text of the dialogue text in the preset text. The text model is trained using such training data. The system can input any dialogue text and any meeting record into the text model, and the text model can output the position of the relevant text of the any dialogue text in the any meeting record.

[0215] It can be seen that the system can input the original meeting record and the dialogue text of the original meeting record into a pre-trained text model, and the text model can output the position of the relevant text of the dialogue text in the original meeting record. The system can extract the relevant text corresponding to the dialogue text according to the position.

[0216] The system can perform hotword recognition on the relevant text of the dialogue text in the original meeting record, and the recognized hotword is the second hotword. Optionally, the system can use one or more of the following methods:

[0217] Method B1, performing word segmentation processing on the relevant text of the dialogue text to obtain a plurality of words. In the obtained plurality of words, the words of a preset part of speech are determined as the second hotword. Optionally, the preset part of speech can include nouns and / or verbs.

[0218] Method B2, performing word segmentation processing on the relevant text of the dialogue text to obtain a plurality of words. In the obtained plurality of words, the words of a preset part of speech are determined as candidate words. According to the selection operation of at least one candidate word, the selected candidate word is determined as the second hotword.

[0219] Method B3, determining at least one group of similar words in the relevant text of the dialogue text. The pinyin of each candidate word in a group of similar words, except for the tone, is the same, or the pinyin of each candidate word in a group of similar words meets a preset similar word rule, or the audio of each candidate word in a group of similar words is similar; based on any one group of similar words, the second hotword corresponding to the any one group of similar words is determined.

[0220] Method B4, determining at least one group of similar words in the relevant text of the dialogue text; displaying the candidate words in any one group of similar words; according to the selection operation of at least one candidate word, the selected candidate word is determined as the second hotword.

[0221] Method B5, performing new word discovery on the relevant text of the dialogue text, and taking the new word as the second hotword.

[0222] In some examples, in the operation of modifying the original meeting record based on the target hotword to obtain the target meeting record, the relevant text of the dialogue text in the original meeting record can be modified based on the target hotword to obtain the target meeting record.

[0223] In one possible implementation, the system can further perform the following steps:

[0224] Displaying a sentence containing any hotword in the original meeting record and indicating the any hotword.

[0225] The system displays at least one group of similar words to which the any hot word belongs, each candidate word in the group of similar words being a word in the original conference record, the pinyin of each candidate word in a group of similar words being identical except for tones, or the pinyin of each candidate word in a group of similar words meeting a preset similar word rule, or the audio of each candidate word in a group of similar words being similar.

[0226] According to the selection operation on the third candidate word, the third candidate word is determined as the target hot word, the third candidate word being the selected one.

[0227] The system can display the sentence of any hot word, and can indicate the any hot word by one or more indication manners such as bold, large font, font change, word color change, word background color change (highlight), etc. so that the user can know the word currently being corrected.

[0228] The system can display at least one group of similar words to which the any hot word belongs, that is, can display other candidate words having any of the following relationships with the any hot word: the pinyin being identical except for tones, the pinyin meeting a preset similar word rule, or the audio being similar. This facilitates the user to select a word to replace the any hot word from the words. The system can determine the selected word as the target hot word, and correct the sentence of the any hot word in the original conference record to the third candidate word.

[0229] In a possible implementation, as shown in FIG. 2, the system can display a first interaction interface including a first display area and a second display area, the first display area being used to display a conference record, and the second display area being used to display a hot word.

[0230] The system can display the sentence containing any hot word in the original conference record in the first display area, and indicate the any hot word. Optionally, the any hot word can be the first hot word or the second hot word, or a candidate word in the groups of similar words.

[0231] The system can display at least one group of similar words to which the any hot word belongs in the second display area, each candidate word in the group of similar words being a word in the original conference record, the pinyin of each candidate word in a group of similar words being identical except for tones, or the pinyin of each candidate word in a group of similar words meeting a preset similar word rule, or the audio of each candidate word in a group of similar words being similar. In other words, the system can display other candidate words having any of the following relationships with the any hot word in the second display area: the pinyin being identical except for tones, the pinyin meeting a preset similar word rule, or the audio being similar. This facilitates the user to select a word to replace the any hot word from the words.

[0232] Optionally, the first interaction interface further includes a third display area, the third display area is configured to display the conversation text and / or related content of the conversation text in the target meeting record.

[0233] Optionally, the system further displays a hotword input box in the first interaction interface. The user can input a word in the hotword input box, and the input word can be used to replace any hotword. The system can obtain the content in the hotword input box and determine the content as a target hotword.

[0234] According to the meeting record processing method provided in any of the above embodiments, the original meeting record can be obtained by performing speech transcription on the target meeting audio based on the preconfigured hotword. The target hotword can further include the preconfigured hotword. The system can further receive a target file. The target file is subjected to hotword recognition to obtain a plurality of candidate words. According to the selection operation of one or more words in the plurality of candidate words, the one or more words are determined as the preconfigured hotword. The user can select one or more words from the plurality of candidate words identified based on the target file, and the selected one or more words are used as the preconfigured hotword.

[0235] Please refer to FIG. 3a, the system can display a second interaction interface. The system can display a file upload window in the second interaction interface. The file upload window is configured to receive file information of a target file uploaded by the user. The system can obtain the target file and perform hotword recognition on the target file to obtain candidate words. Please refer to FIG. 3b, the system can display a third interaction interface. The third interaction interface includes a candidate word display area and a configured hotword display area. The system can display the candidate words identified based on the target file in the candidate word display area. The user can select one or more candidate words in the candidate word display area. The system can display the selected candidate words in the configured hotword display area. Optionally, the third interaction interface further includes a configured hotword input box. The user can input a hotword in the configured hotword input box. The system can determine the text in the configured hotword input box as a configured hotword. The system can display the input hotword in the configured hotword display area. Optionally, after the user inputs a word in the configured hotword input box, the user can click an “input” button to trigger the text in the configured hotword input box as the configured hotword. Then the user can input new text in the configured hotword input box.

[0236] According to the meeting record processing method provided in any of the above embodiments, the system can modify part or all of the original meeting record based on the target hotword. In the operation of obtaining the target meeting record, part of the original meeting record can be modified based on the target hotword, or all of the original meeting record can be modified based on the target hotword.

[0237] In a possible application scenario, the system corrects part or all of the original conference record based on the target hot word, and when obtaining the target conference record, the system can correct the relevant text of the target text in the original conference record based on the target hot word to obtain the target conference record.

[0238] In the operation of correcting the original conference record by the system, only the relevant text of the dialogue text can be corrected, and the target conference record can be obtained after correction. In this design, the part concerned by the user can be corrected, and the entire text of the original conference record is not traversed for correction, which can meet the use needs of the user and reduce the amount of processed data.

[0239] In another possible application scenario, the system corrects all of the original conference record based on the target hot word. The system can re-voice transcribe the target conference audio based on the target hot word, and the transcribed result is used as the target conference record.

[0240] In another possible application scenario, the system corrects part or all of the original conference record based on the target hot word, and the following operations can be performed when obtaining the target conference record:

[0241] The target audio of the target hot word is generated. The similarity between the target audio and an audio segment corresponding to a target word in the original conference record is determined, wherein the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio. If the similarity between the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, the target word in the text corresponding to the audio segment in the original conference record is changed to the target hot word to obtain the target conference record.

[0242] The system can generate the audio of the target hot word by using a speech synthesis technology or a text-to-speech manner. The audio of the target hot word is recorded as the target audio for convenience.

[0243] In a possible scenario, the target word is selected from the original conference record according to a preset selection manner. In another possible scenario, the target word can be any word in the original conference record. In yet another possible scenario, the system can record a word in the original conference record that has a close relationship with the target hot word as the target word. The target word and the target hot word have the same part of the pinyin without tone, which can reflect that the target word has a close relationship with the target hot word. Alternatively, the pinyin of the target word and the target hot word conforms to a preset close word rule, which can reflect that the target word has a close relationship with the target hot word. Alternatively, the audio of the target hot word and the target word is similar, which can reflect that the target word has a close relationship with the target hot word.

[0244] In some examples, the system can generate the audio of the target word based on the target word, for the convenience of distinguishing, and record the audio segment corresponding to the target word. In some examples, the system can obtain the target conference audio and the original conference record, determine the time stamp of the target word in the target audio according to the position of the target word in the original conference record, and extract the audio segment of the target word according to the time stamp of the target word.

[0245] The similarity between the audio segment corresponding to the target word and the target audio and the preset similarity threshold satisfies a preset relationship, which can reflect that the audio of the target word is similar to the target hot word, and the target hot word is easily misrecognized as the target word in the speech transcription process. The system can change the target word in the text corresponding to the audio segment of the target word in the original conference record to the target hot word, to obtain the target conference record. The system can correct the target word with speech recognition error to the target hot word.

[0246] Optionally, the similarity between the audio segment corresponding to the target word and the target audio is greater than or equal to the preset similarity threshold, which can reflect that the two satisfy the preset relationship. The similarity between the audio segment corresponding to the target word and the target audio is less than the preset similarity threshold, which can reflect that the two do not satisfy the preset relationship.

[0247] In a possible implementation, the system corrects part or all of the original conference record based on the target hot word to obtain the target conference record, and can perform the following steps in the operation:

[0248] The system corrects the original conference record based on the target hot word by using a first correction manner to obtain a first conference record.

[0249] The system corrects the original conference record based on the target hot word by using a second correction manner to obtain a second conference record.

[0250] compare whether the sentences at the positions of the first meeting record and the second meeting record are consistent;

[0251] For the positions where the sentences are inconsistent, the system performs fluency evaluation on the sentences at the positions in the first meeting record and the second meeting record respectively, and obtains the target meeting record based on the sentence with the best fluency among the sentences at the positions in the first meeting record and the second meeting record.

[0252] The system can use multiple correction methods to correct the original meeting record respectively, and different correction methods can obtain different meeting records. The first meeting record can be obtained by correcting the original meeting record by using a first correction method. The second meeting record can be obtained by correcting the original meeting record by using a second correction method. It should be understood that the same target part in the original meeting record is corrected by the two correction methods. Optionally, the target part can be part of the text or the entire text of the original meeting record.

[0253] The system can compare whether the sentences at the target positions in the first meeting record and the second meeting record are consistent (or the same). The target positions can be any positions. Optionally, the target positions can represent the order of the sentences in the meeting record.

[0254] If the sentences are inconsistent, the system can perform fluency evaluation on the sentences at the target positions in the first meeting record to obtain a first evaluation parameter. The system can perform fluency evaluation on the sentences at the target positions in the second meeting record to obtain a second evaluation parameter. In some scenarios, the first evaluation parameter and the second evaluation parameter are numerical values. The system can take the sentence corresponding to the maximum value between the first evaluation parameter and the second evaluation parameter as the most fluent sentence. Optionally, if the first evaluation parameter and the second evaluation parameter are the same, the sentence corresponding to any evaluation parameter is taken as the most fluent sentence.

[0255] In other scenarios, the first evaluation parameter and the second evaluation parameter are fluency categories. The fluency categories can include, but are not limited to, not fluent, poor, general, good, and fluent. The order of the fluency levels from low to high is not fluent, poor, general, good, and fluent. The system can take the sentence corresponding to the highest fluency level between the first evaluation parameter and the second evaluation parameter as the most fluent sentence. Optionally, if the first evaluation parameter and the second evaluation parameter are the same, the sentence corresponding to any evaluation parameter is taken as the most fluent sentence.

[0256] Optionally, the system can determine the fluency of any sentence by using an existing text fluency evaluation algorithm. Alternatively, the system can determine the fluency of any sentence by using a pre-trained evaluation model. The evaluation model can be a large model or a pre-trained model such as a BERT model. The evaluation model can be trained to have the ability to evaluate the fluency of text. Embodiments of the present application do not make further introductions in this regard.

[0257] In one possible design, the system can modify the first conference record, and the system can take the most fluent sentence from the sentence at the target position in the first conference record and the sentence at the target position in the second conference record as the sentence at the target position in the first conference record. That is, the system can modify the first conference record by taking the most fluent sentence from the sentence at the target position in the first conference record and the sentence at the target position in the second conference record as the sentence at the target position in the first conference record.

[0258] Based on the foregoing process, after the system traverses the sentences at each position in the first conference record and the second conference record, the system can take the modified first conference record as the target conference record.

[0259] In another possible design, the system can modify the second conference record, and the system can take the most fluent sentence from the sentence at the target position in the first conference record and the sentence at the target position in the second conference record as the sentence at the target position in the second conference record. That is, the system can modify the second conference record by taking the most fluent sentence from the sentence at the target position in the first conference record and the sentence at the target position in the second conference record as the sentence at the target position in the second conference record.

[0260] Based on the foregoing process, after the system traverses the sentences at each position in the first conference record and the second conference record, the system can take the modified second conference record as the target conference record.

[0261] In some examples, the first modification manner and the second modification manner are any two of the following modification manners:

[0262] C1, based on the target hotword, modifying the relevant text of the dialogue text in the original conference record to obtain the first conference record or the second conference record;

[0263] C2, based on the target hotword, performing speech transcription on the target conference audio to obtain the first conference record or the second conference record;

[0264] C3, generating a target audio of the target hotword; determining a similarity between the target audio and an audio segment corresponding to a target word, wherein the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio; and if the similarity between the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, changing the target word in a text corresponding to the audio segment to the target hotword in the original conference record to obtain the first conference record or the second conference record.

[0265] For example, the first correction manner can be manner C1 to obtain the first conference record, and the second correction manner can be manner C2 to obtain the second conference record. The first correction manner can be manner C1 to obtain the first conference record, and the second correction manner can be manner C3 to obtain the second conference record. The first correction manner can be manner C2 to obtain the first conference record, and the second correction manner can be manner C3 to obtain the second conference record.

[0266] In a possible implementation, based on the conference record processing method provided in any one of the above embodiments, the system can output related content corresponding to the dialogue text based on the target conference record after obtaining the target conference record. For example, when the dialogue text is a question text, the system can output an answer text corresponding to the question text based on the target conference record. For another example, when the dialogue text is an instruction text, the system can generate corresponding content such as a chart, a picture, an audio, a video, a text, etc. based on related text corresponding to the instruction text in the target conference record.

[0267] FIG. 4 exemplarily shows a conference record processing method, which can be executed by a conference record processing system. The method can include the following steps:

[0268] S201, obtaining an original conference record, which is obtained by speech transcription based on target conference audio.

[0269] S202, obtaining a target text, which is dialogue text of the original conference record.

[0270] S203, obtaining a first hotword based on the target text.

[0271] S204, correcting part or all of the original conference record based on the target hotword to obtain a target conference record, wherein the target hotword at least includes the first hotword.

[0272] S205, generating a conference minutes based on the target conference record.

[0273] In this embodiment, the operations in steps S201 to S204 can refer to the related descriptions in the foregoing embodiments, which will not be repeated here.

[0274] Compared with the meeting minutes generated based on the original meeting records, the system generates meeting minutes based on the corrected meeting records, which are more accurate and of higher quality.

[0275] In some examples, the system can generate meeting minutes directly based on the target meeting records. The system can extract part of the content from the target meeting records in an extraction manner to generate the meeting minutes. Alternatively, the system can summarize the content in the target meeting records in a summarization manner to generate the meeting minutes.

[0276] The system can take the relevant text corresponding to the dialogue text as a focus point. The focus point is added to the meeting minutes, or the meeting minutes can reflect the focus point.

[0277] In some examples, the system can determine the relevant text corresponding to the dialogue text based on the target meeting records. The system can extract part of the content from the target meeting records in an extraction manner. The system can generate meeting minutes based on the extracted content and the relevant text corresponding to the dialogue text. For example, the system can input the extracted content and the relevant text corresponding to the dialogue text to a large model after deduplication processing to generate the meeting minutes.

[0278] In other examples, the system can determine the relevant text corresponding to the dialogue text based on the target meeting records. The system can summarize the content in the target meeting records in a summarization manner to generate a first meeting minutes. The system can add the relevant text corresponding to the dialogue text to the first meeting minutes to obtain the corrected meeting minutes.

[0279] In other examples, the system can determine the relevant text corresponding to the dialogue text based on the target meeting records. The system can utilize a large model to summarize the content in the target meeting records to generate the meeting minutes. The prompt information input into the large model can indicate that the large model focuses on the relevant text corresponding to the dialogue text.

[0280] FIG. 5 exemplarily shows a structural schematic diagram of a meeting record processing system. The meeting processing system can include a server (or cloud) and one or more terminals. The terminals can include, but are not limited to, all-in-one machines, mobile phones, computers, tablet computers, and other electronic devices.

[0281] In some application scenarios, the server in the meeting record processing system executes the meeting record processing method provided in any one of the foregoing embodiments.

[0282] In some application scenarios, the terminal in the meeting record processing system executes the meeting record processing method provided in any one of the foregoing embodiments.

[0283] In some application scenarios, the conference recording processing system, the server and the terminal cooperatively implement the conference recording processing method provided in any one of the foregoing embodiments. In one possible implementation, please refer to FIG. 6. The terminal and the server can respectively obtain the target conference audio. The terminal can perform speech transcription on the target conference audio to obtain a first original conference record. The server can perform speech transcription on the target conference audio to obtain a second original conference record.

[0284] Due to the performance difference between the terminal and the server, the first original conference record and the second original conference record are different. In some scenarios, the performance difference between the terminal and the server is small, and the difference between the first original conference record and the second original conference record is small.

[0285] The terminal can perform hotword recognition based on the first original conference record to obtain a fourth hotword. For example, the terminal can determine at least one group of similar words based on the first original conference record. The pinyin of each candidate word in a group of similar words is the same except the tone, or the pinyin of each candidate word in a group of similar words meets a preset similar word rule, or the audio of each candidate word in a group of similar words is similar.

[0286] Optionally, the terminal can perform new word discovery on the first original conference record, and take the new word as the fourth hotword. The terminal can perform deduplication processing on the obtained fourth hotword.

[0287] The server can perform hotword recognition based on the second original conference record to obtain a fifth hotword. For example, the server can determine at least one group of similar words based on the second original conference record. The pinyin of each candidate word in a group of similar words is the same except the tone, or the pinyin of each candidate word in a group of similar words meets a preset similar word rule, or the audio of each candidate word in a group of similar words is similar.

[0288] Optionally, the server can perform new word discovery on the second original conference record, and take the new word as the fifth hotword. The server can perform deduplication processing on the obtained fifth hotword.

[0289] One of the server or the terminal can integrate, for example, perform deduplication processing on the fourth hotword after deduplication processing of the terminal and the fifth hotword after deduplication processing of the server, to implement the function of determining the second hotword based on the original conference record in the conference processing system.

[0290] In a possible scenario, the server can collect the audio of the target conference through an audio collection device to obtain the audio of the target conference. One terminal corresponds to one audio collection device. Optionally, the terminal and the corresponding audio collection device are independent. Alternatively, the terminal and the corresponding audio collection device are integrated.

[0291] In another possible implementation, the server can obtain the audio of the target conference, the terminal can obtain the audio of the target conference, and the terminal cannot obtain the audio of the target conference. In a possible scenario, the server can collect audio through multiple audio collection devices, and integrate the collected audio to obtain the audio of the target conference. The terminal can not have a corresponding audio collection device, or the terminal can have a corresponding audio collection device. Optionally, the terminal and the corresponding audio collection device are independent, or the terminal and the corresponding audio collection device are integrated. The terminal can obtain the audio collected by the corresponding audio collection device. The terminal can perform speech transcription on the audio to obtain the first original conference record.

[0292] Based on the same idea, the embodiments of the present application further provide a conference record processing system, which is used for:

[0293] obtaining an original conference record, the original conference record being obtained by performing speech transcription on target conference audio.

[0294] obtaining target text, the target text being dialogue text of the original conference record;

[0295] obtaining a first hot word based on the target text;

[0296] correcting part or all of the original conference record based on the target hot word to obtain a target conference record, the target hot word including the first hot word.

[0297] In a possible implementation, the conference record processing system is further used for:

[0298] outputting related content of the dialogue text based on the target conference record.

[0299] In a possible implementation, in the operation of obtaining the first hot word based on the target text, the conference record processing system is specifically used for:

[0300] performing word segmentation processing on the target text to obtain a plurality of words;

[0301] determining a word with a preset word type as a candidate word in the plurality of words;

[0302] According to the selection operation on the first candidate word, the first candidate word is determined as the first hot word, and the first candidate word is one or more candidate words in the determined candidate words.

[0303] In a possible implementation, in the operation of obtaining the first hot word based on the target text, the conference record processing system is specifically configured to:

[0304] perform a word segmentation processing on the target text to obtain a plurality of words;

[0305] In the plurality of words, a word of a preset word type is determined as the first hot word.

[0306] In a possible implementation, in the operation of obtaining the target text, the conference record processing system is specifically configured to:

[0307] obtain a dialogue audio of the target conference and determine a dialogue text corresponding to the dialogue audio; or

[0308] receive the target text.

[0309] In a possible implementation, the conference record processing system is further configured to:

[0310] obtain a second hot word based on the original conference record, and the target hot word further includes the second hot word.

[0311] In a possible implementation, in the operation of obtaining the second hot word based on the original conference record, the conference record processing system is specifically configured to:

[0312] determine at least one group of similar words in the original conference record, the pinyin of each candidate word in a group of similar words is the same except for tone, or the pinyin of each candidate word in a group of similar words meets a preset similar word rule, or the audio of each candidate word in a group of similar words is similar;

[0313] determine a second hot word corresponding to any group of similar words based on the any group of similar words.

[0314] In a possible implementation, in the operation of determining the second hot word corresponding to the any group of similar words based on the any group of similar words, the conference record processing system is specifically configured to:

[0315] display a candidate word in any group of similar words;

[0316] According to the selection operation on the second candidate word, the second candidate word is determined as the second hot word, and the second candidate word is one or more candidate words in the any group of similar words.

[0317] In a possible implementation, in the operation of determining the second hotword corresponding to the any set of similar words based on the any set of similar words, the conference record processing system is specifically configured to:

[0318] determine the input word as the second hotword according to the operation of the input word.

[0319] In a possible implementation, the conference record processing system is further configured to:

[0320] indicate the position of each candidate word in the any set of similar words in the original conference record.

[0321] In a possible implementation, in the operation of obtaining the second hotword based on the original conference record, the conference record processing system is specifically configured to:

[0322] perform new word discovery on the original conference record, and take the new word as the second hotword.

[0323] In a possible implementation, in the operation of obtaining the second hotword based on the original conference record, the conference record processing system is specifically configured to:

[0324] perform hotword recognition on the relevant content of the dialogue text in the original conference record to obtain the second hotword.

[0325] In a possible implementation, in the operation of modifying the original conference record based on the target hotword to obtain the target conference record, the conference record processing system is specifically configured to:

[0326] modify the relevant text of the target text in the original conference record based on the target hotword to obtain the target conference record.

[0327] In a possible implementation, in the operation of modifying the original conference record based on the target hotword to obtain the target conference record, the conference record processing system is specifically configured to:

[0328] perform speech transcription on the target conference audio based on the target hotword to obtain the target conference record.

[0329] In a possible implementation, in the operation of modifying the original conference record based on the target hotword to obtain the target conference record, the conference record processing system is specifically configured to:

[0330] generate target audio of the target hotword;

[0331] determine a similarity between the target audio and an audio segment corresponding to a target word in the original conference record, wherein the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio, and the target word has a close relationship with the target hot word;

[0332] if the similarity between the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, change the target word in the text corresponding to the audio segment to the target hot word in the original conference record to obtain the target conference record.

[0333] In a possible implementation, in the operation of correcting part or all of the original conference record based on the target hot word to obtain the target conference record, the conference record processing system is specifically configured to:

[0334] correct the original conference record based on the target hot word by using a first correction manner to obtain a first conference record;

[0335] correct the original conference record based on the target hot word by using a second correction manner to obtain a second conference record;

[0336] compare whether the sentences at each position of the first conference record and the second conference record are consistent;

[0337] for the position where the sentences are inconsistent, respectively perform fluency evaluation on the sentences at the position in the first conference record and the sentences at the position in the second conference record, and take the most fluent sentence from the sentences at the position in the first conference record and the sentences at the position in the second conference record as the sentence at the position in the third conference record to obtain the target conference record, wherein the third conference record is the first conference record or the second conference record.

[0338] In a possible implementation, the first correction manner and the second correction manner are any two of the following correction manners:

[0339] Manner one: based on the target hot word, correct the relevant text of the dialogue text in the original conference record to obtain the first conference record or the second conference record;

[0340] Manner two: based on the target hot word, perform speech transcription on the target conference audio to obtain the first conference record or the second conference record;

[0341] In a third mode, a target audio of the target hotword is generated; similarity of an audio segment corresponding to a target word to the target audio is determined, the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio, the target word has a close relationship with the target hotword; and if the similarity of the audio segment to the target audio satisfies a preset relationship with a preset similarity threshold, the target word in the text corresponding to the audio segment is changed to the target hotword in the original conference record, to obtain the first conference record or the second conference record.

[0342] In a possible implementation, the original conference record is obtained by performing speech transcription on the target conference audio based on a preconfigured hotword, and the target hotword further includes the preconfigured hotword.

[0343] receiving a target file;

[0344] performing hotword recognition on the target file to obtain a plurality of candidate words;

[0345] determining one or more words in the plurality of candidate words as the preconfigured hotword according to a selection operation on the one or more words.

[0346] In a possible implementation, the conference record processing system further includes:

[0347] displaying a first interactive interface, the first interactive interface including a first display area and a second display area, the first display area being configured to display a conference record, and the second display area being configured to display a hotword;

[0348] displaying a sentence containing any hotword in the original conference record in the first display area, and indicating the any hotword;

[0349] displaying at least one group of close words to which the any hotword belongs in the second display area, each candidate word in each group of close words being a word in the original conference record, a pinyin of each candidate word in a group of close words being identical except for a tone, or a pinyin of each candidate word in a group of close words meeting a preset close word rule, or an audio of each candidate word in a group of close words being similar.

[0350] In a possible implementation, the first interactive interface further includes a hotword input box, and the conference record processing system further includes:

[0351] obtaining content in the hotword input box, and the target hotword further includes the content in the hotword input box.

[0352] Based on the same idea, FIG. 7 provides a structural schematic diagram of an electronic device 700 according to an embodiment of the present application. The electronic device 700 can execute the conference recording processing method provided in any one of the preceding embodiments, or the electronic device 700 can implement part or all of the functions of the conference recording processing system in any one of the preceding embodiments. As shown in FIG. 7, the electronic device 700 can include one or more processors 701, one or more memories 702, a communication interface 703, and one or more computer programs 704, which can be connected through one or more communication buses 705. The one or more computer programs 704 are stored in the memory 702 and configured to be executed by the one or more processors 701, and the one or more computer programs 704 include instructions. For example, the instructions can be used to execute the steps in the conference recording processing method in the corresponding embodiments as described above. The communication interface 703 is configured to implement communication with other devices (such as terminals), for example, the communication interface can be a transceiver.

[0353] In addition, the embodiments of the present application provide a computer-readable storage medium for storing a computer program, which, when executed on a computer, causes the computer to perform the steps in any one of the conference recording processing methods described above.

[0354] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product.

[0355] The embodiments of the present application further provide a computer program product. The computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the computer program instructions produce, wholly or partially, the flow or function described in the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by the computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, Solid State Disk (SSD)), etc. The solutions of the above embodiments can be combined for use without conflict.

Claims

1. A method of processing a conference record, wherein, The method comprises: obtaining an original conference record, the original conference record being obtained by voice transcription based on target conference audio; obtaining target text, the target text being dialogue text of the original conference record; obtaining a first hot word based on the target text; correcting part or all of the original conference record based on the target hot word to obtain target conference record, the target hot word comprising the first hot word.

2. The method of claim 1, wherein, The method further comprises: outputting related content of the dialogue text based on the target conference record.

3. The method of claim 1, wherein, The method further comprises: performing word segmentation on the target text to obtain a plurality of words; determining words of a preset word type as candidate words in the plurality of words; determining the first candidate word as the first hot word according to a selection operation of the first candidate word, the first candidate word being one or more of the determined candidate words.

4. The method of claim 1, wherein, The method further comprises: performing word segmentation on the target text to obtain a plurality of words; determining words of a preset word type as the first hot word in the plurality of words.

5. The method of claim 1, wherein, The method further comprises: obtaining dialogue audio of the target conference and determining dialogue text corresponding to the dialogue audio; or receiving the target text.

6. The method of claim 1, wherein, The method further comprises: obtaining a second hot word based on the original conference record, the target hot word further comprising the second hot word. The method further comprises:

7. The method of claim 6, wherein, determining at least one group of similar words in the original conference record, the pinyin of each candidate word in a group of similar words being the same except for tone, or the pinyin of each candidate word in a group of similar words meeting a preset similar word rule, or the audio of each candidate word in a group of similar words being similar; determining a second hot word corresponding to any group of similar words based on the any group of similar words. The method further comprises:

8. The method of claim 7, wherein, displaying candidate words in any group of similar words; determining the second candidate word as the second hot word according to a selection operation of the second candidate word, the second candidate word being one or more of the candidate words in the any group of similar words. The method further comprises:

9. The method of claim 7, wherein, determining an input word as the second hot word according to an input word operation. The method further comprises:

10. The method of claim 6, wherein, indicating the position of each candidate word in the original conference record. The method further comprises:

11. The method of claim 6, wherein, performing new word discovery on the original conference record, and taking a new word as the second hot word. The method further comprises:

12. The method of claim 6, wherein, performing hot word identification on related content of the dialogue text in the original conference record to obtain the second hot word. The method further comprises:

13. The method of claim 6, wherein, ​ The target text in the original conference record is corrected based on the target hot word, and the target conference record is obtained.

14. The method of claim 1, wherein, The original conference record is corrected based on the target hot word to obtain the target conference record, and the method comprises the following steps: The target conference record is obtained by performing speech transcription on the target conference audio based on the target hot word.

15. The method of claim 1, wherein, The original conference record is partially or wholly corrected based on the target hot word to obtain the target conference record, and the method comprises the following steps: A target audio of the target hot word is generated. A similarity between the target audio and an audio segment corresponding to a target word in the original conference record is determined, wherein the audio segment corresponding to the target word is generated based on the target word, or the audio segment corresponding to the target word is extracted from the target conference audio. If the similarity between the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, the target word in the text corresponding to the audio segment in the original conference record is changed to the target hot word to obtain the target conference record.

16. The method of claim 15, wherein, The target word and the target hot word have a similar relationship, wherein the pinyin of the target word and the pinyin of the target hot word are the same except for tone, or the pinyin of the target word and the pinyin of the target hot word comply with a preset similar word rule.

17. The method of claim 1, wherein, The original conference record is partially or wholly corrected based on the target hot word to obtain the target conference record, and the method comprises the following steps: The original conference record is corrected by a first correction method based on the target hot word to obtain a first conference record. The original conference record is corrected by a second correction method based on the target hot word to obtain a second conference record. Whether the sentences at each position of the first conference record and the second conference record are consistent is compared. For the position where the sentences are inconsistent, the sentences at the position in the first conference record and the sentences at the position in the second conference record are respectively evaluated for fluency, and the target conference record is obtained based on the most fluent sentences among the sentences at the position in the first conference record and the sentences at the position in the second conference record.

18. The method of claim 17, wherein, The first correction method and the second correction method are any two of the following correction methods: Method one: the target text in the original conference record is corrected based on the target hot word to obtain the first conference record or the second conference record. Method two: the target conference audio is transcribed based on the target hot word to obtain the first conference record or the second conference record. Method three: a target audio of the target hot word is generated. determine a similarity between the target audio and an audio segment corresponding to a target keyword, wherein the audio segment corresponding to the target keyword is generated based on the target keyword, or the audio segment corresponding to the target keyword is extracted from the target conference audio; and if the similarity between the audio segment and the target audio satisfies a preset relationship with a preset similarity threshold, change the target keyword to the target hotword in text corresponding to the audio segment in the original conference record, to obtain the first conference record or the second conference record.

19. The method of claim 1, wherein, The original conference record is obtained by performing speech transcription on the target conference audio based on a preconfigured hotword, and the target hotword further includes the preconfigured hotword. The method further includes: receiving a target file; performing hotword recognition on the target file to obtain a plurality of candidate keywords; determining one or more keywords in the plurality of candidate keywords as the preconfigured hotword based on a selection operation on the one or more keywords.

20. The method of claim 1 or 6, wherein, The method further includes: displaying a first interaction interface, the first interaction interface including a first display area and a second display area, the first display area being configured to display a conference record, and the second display area being configured to display a hotword; displaying a sentence containing any hotword in the original conference record in the first display area, and indicating the any hotword; displaying at least one group of similar keywords to which the any hotword belongs in the second display area, each group of similar keywords including candidate keywords in the original conference record, the pinyin of each candidate keyword in a group of similar keywords being identical except for tone, or the pinyin of each candidate keyword in a group of similar keywords conforming to a preset similar keyword rule, or the audio of each candidate keyword in a group of similar keywords being similar.

21. The method of claim 20, wherein, The first interaction interface further includes a hotword input box, and the method further includes: obtaining content in the hotword input box, and the target hotword further includes the content in the hotword input box.

22. A meeting record processing system, wherein, The conference record processing system is configured to: obtain an original conference record, the original conference record being obtained by performing speech transcription on target conference audio; obtain a target text, the target text being a dialogue text of the original conference record; obtain a first hotword based on the target text; correct part or all of the original conference record based on a target hotword to obtain a target conference record, the target hotword including the first hotword.

23. An electronic device, comprising: The electronic device includes a memory and a processor; The memory stores computer program instructions; The processor executes the instructions to implement steps in the conference record processing method according to any one of claims 1-21.

24. A computer storage medium, wherein, When the instructions in the computer storage medium are executed by the processor, the processor performs steps in the conference record processing method according to any one of claims 1-21.

Citation Information

Patent Citations

  • Text correction method and device, intelligent equipment and readable memory medium

    CN108804414A

  • Conference record generation method, electronic device and storage medium

    CN110134756A

  • Method and device for generating conference summary and electronic equipment

    CN117316161A

  • Automatically generating conference minutes

    US20210407499A1