Conference summary generation method and device, computer readable medium and electronic equipment

By converting conference audio into candidate text sequences and using trained language models to generate conference minutes, the problem of low speech recognition accuracy in the prior art is solved, and the accuracy and comprehensiveness of conference minutes are improved.

CN120340475APending Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410061663.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The accuracy of existing speech recognition technology is not high, resulting in low accuracy of generated meeting minutes and failure to effectively utilize relevant content in the presentation.

Method used

By converting the conference audio into a candidate text sequence and evaluating the probability of the candidate text sequence using a language model trained based on the conference record text and presentation text, a text sequence matching the target meeting is generated as the first text sequence, and the meeting minutes are finally generated.

Benefits of technology

It improves the accuracy of speech recognition, enhances the accuracy and comprehensiveness of generating meeting minutes, and reduces the workload of users to record meeting minutes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340475A_ABST
    Figure CN120340475A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a conference summary generation method and device, a computer readable medium and electronic equipment, and the method comprises the steps: converting the conference audio of a target conference into a plurality of candidate text sequences according to a voice recognition model, and inputting each candidate text sequence into a language model trained according to conference related texts, the probability of each candidate text sequence is obtained through evaluation of the language model, so that the voice recognition model outputs the candidate text sequence matched with the target conference as a first text sequence according to the probability; the conference related text comprises at least one of the following items: a conference record text of the conference and a text extracted from a presentation file of the conference; and generating a conference summary of the target conference according to the first text sequence. According to the embodiment of the invention, the context understanding and vocabulary selection of the language model can more closely surround the current meeting needing to generate the summary, the accuracy of voice recognition is improved, and the accuracy of generating the meeting summary is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology. Specifically, it relates to a method, apparatus, computer-readable medium, and electronic device for generating meeting minutes. Background Art

[0002] With the development of artificial intelligence technology, some Internet companies have begun to use speech recognition technology to generate meeting minutes of multimedia meetings.

[0003] However, the accuracy of most current speech recognition technologies is not high, which results in a low accuracy of the generated meeting minutes. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, computer-readable medium, and electronic device for generating meeting minutes, which can at least to some extent improve the accuracy of speech recognition, and thus improve the accuracy of the generated meeting minutes.

[0005] Other features and advantages of the present application will become apparent through the following detailed description, or will be partially learned through the practice of the present application.

[0006] According to one aspect of the embodiments of the present application, a method for generating meeting minutes is provided. The method includes: converting the meeting audio of a target meeting into a plurality of candidate text sequences according to a speech recognition model, and inputting each of the candidate text sequences into a language model trained according to meeting-related texts. The language model evaluates each of the candidate text sequences to obtain the probability of each of the candidate text sequences, so that the speech recognition model outputs the candidate text sequence that matches the target meeting as the first text sequence according to the probabilities of the candidate text sequences; wherein, the meeting-related texts include at least one of the following: the meeting record text of the meeting, the text extracted from the presentation of the meeting; generating the meeting minutes of the target meeting according to the first text sequence, and the meeting minutes are an overview of the meeting content of the target meeting.

[0007] According to one aspect of the embodiments of the present application, a device for generating meeting minutes is provided. The device includes: a conversion unit configured to convert the meeting audio of a target meeting into a plurality of candidate text sequences according to an automatic speech recognition model, and input each of the candidate text sequences into a language model trained according to meeting-related texts, so that the language model evaluates each of the candidate text sequences to obtain the probabilities of each of the candidate text sequences, so that the automatic speech recognition model outputs the candidate text sequence that matches the target meeting as the first text sequence according to the probabilities of each of the candidate text sequences; wherein, the meeting-related texts include at least one of the following: the meeting record text of the meeting, the text extracted from the presentation of the meeting; a generation unit configured to generate the meeting minutes of the target meeting according to the first text sequence, and the meeting minutes are an overview of the meeting content of the target meeting.

[0008] In some embodiments of the present application, based on the foregoing solution, the device further includes a meeting record text acquisition unit; before generating the meeting minutes of the target meeting according to the first text sequence, the meeting record text acquisition unit is configured to: acquire the meeting record text of the target meeting; the generation unit is configured to: input the first text sequence and the meeting record text into a pre-trained meeting minutes generation model to obtain the meeting minutes of the target meeting generated by the meeting minutes generation model.

[0009] In some embodiments of the present application, based on the foregoing solution, the device further includes an extraction unit; before generating the meeting minutes of the target meeting according to the first text sequence, the extraction unit is configured to: extract a second text sequence from the presentation of the target meeting; the generation unit is configured to: input the first text sequence, the second text sequence, and the meeting record text into a pre-trained meeting minutes generation model to obtain the meeting minutes of the target meeting generated by the meeting minutes generation model.

[0010] In some embodiments of the present application, based on the foregoing solution, the device further includes a model training unit; before inputting the first text sequence and the meeting record text into a pre-trained meeting minutes generation model, the model training unit is configured to: convert the meeting audio of a plurality of meetings into third text sequences respectively based on an automatic speech recognition model; extract fourth text sequences from the presentations of the plurality of meetings respectively; acquire the meeting minutes and meeting record texts of the plurality of meetings; and use the third text sequence, the fourth text sequence, the meeting minutes, and the meeting record text of each meeting as a training sample, and train a meeting minutes generation model based on the training samples corresponding to the plurality of meetings.

[0011] In some embodiments of the present application, based on the foregoing solution, the model training unit is configured to: construct input prompt information according to the third text sequence, the fourth text sequence, and the meeting record text in the training sample; use the meeting summary in the training sample as the input and label of the training sample; and train a meeting summary generation model based on the input prompt information, input, and label of each training sample.

[0012] In some embodiments of the present application, based on the foregoing solution, the model training unit is configured to: for each training sample, randomly construct the input prompt information corresponding to the training sample in one of the following ways: use the third text sequence and the fourth text sequence in the training sample as the input prompt information; or, use the third text sequence, the fourth text sequence, and the meeting record text in the training sample as the input prompt information.

[0013] In some embodiments of the present application, based on the foregoing solution, the meeting summary generation model includes multiple Transformer layers, and each Transformer layer includes a masked multi-head attention layer. The vector corresponding to the token of each word in the input of the training sample can perform attention calculation with the vector corresponding to the token of each word in the input prompt information corresponding to the training sample; the vector corresponding to the token of the current word in the input of the training sample can only perform attention calculation with the vector corresponding to the token of the current word and the vectors corresponding to the tokens of the words before the current word.

[0014] In some embodiments of the present application, based on the foregoing solution, the extraction unit is configured to: identify a second text sequence from the presentation image of the target meeting based on an optical character recognition algorithm.

[0015] In some embodiments of the present application, based on the foregoing solution, the language model is trained based on the meeting record text of the target meeting and the text extracted from the presentation of the target meeting.

[0016] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the method for generating a meeting summary as described in the above embodiments.

[0017] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for generating a meeting summary as described in the above embodiments.

[0018] According to one aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for generating meeting minutes as described in the above embodiments.

[0019] In the technical solutions provided by some embodiments of the present application, several candidate text sequences are first obtained according to a speech recognition model, and each candidate text sequence is input into a language model. The language model evaluates each candidate text sequence to obtain the probability of each candidate text sequence. The speech recognition model will then output the candidate text sequence that matches the target meeting as the first text sequence according to the probabilities of each candidate text sequence, and then generate the meeting minutes of the target meeting according to the first text sequence. Since the language model for speech recognition is trained based on at least one of the meeting record text of the meeting and the text extracted from the presentation of the meeting, this will make the context understanding and vocabulary selection of the language model more closely centered around the meeting for which the minutes are to be generated currently, improving the accuracy of speech recognition and thus the accuracy of generating meeting minutes.

[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts. In the drawings:

[0022] Figure 1 Shows a flowchart of a method for generating meeting minutes in the related art;

[0023] Figure 2 Shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied;

[0024] Figure 3 Shows a flowchart of a method for generating meeting minutes according to an embodiment of the present application;

[0025] Figure 4 Shows a training architecture diagram of a meeting minutes generation model according to an embodiment of the present application;

[0026] Figure 5Shows the steps before step 380 and the details of step 380 in an embodiment according to the present application; Figure 3 Flowchart of the steps before step 380 and the details of step 380 in the embodiment;

[0027] Figure 6 Shows the steps before step 380 and the details of step 380' in an embodiment according to the present application; Figure 5 Flowchart of the steps before step 380 and the details of step 380' in the embodiment;

[0028] Figure 7 Shows the steps before step 380' in an embodiment according to the present application; Figure 5 Flowchart of the steps before step 380';

[0029] Figure 8 Shows the flowchart of training a meeting minutes generation model based on training samples corresponding to multiple meetings according to an embodiment of the present application;

[0030] Figure 9 Shows a schematic diagram of a mask matrix according to an embodiment of the present application;

[0031] Figure 10 Shows a block diagram of a device for generating meeting minutes according to an embodiment of the present application;

[0032] Figure 11 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0033] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0034] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0035] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0036] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0037] The flowcharts shown in the drawings are only exemplary descriptions, and do not necessarily include all the content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0038] With the development of network technology, more and more people participate in meetings through multimedia conferences.

[0039] Figure 1 The flowchart of the method for generating meeting minutes in the related art is shown. Please refer to Figure 1 As shown, in the related art, the method for generating meeting minutes mainly includes the following processes: Step 110, during the multimedia conference through the multimedia conference terminal, receive the instruction input by the user using the above multimedia conference terminal; Step 120, record the voice input by the above user after inputting the above instruction; Step 130, recognize the above voice as text information, and generate the meeting minutes of the multimedia conference according to the above text information.

[0040] It can be seen that the solution of the related art mainly uses traditional speech recognition technology to recognize speech as text information and generate the meeting minutes of the multimedia conference according to the text information. Although this method can obtain the meeting minutes of the multimedia conference, due to the low performance of the traditional speech recognition technology, that is, the low accuracy of speech recognition, the accuracy of generating the meeting minutes is also very low; in addition, the solution of the related art only uses the audio modality and does not consider the relevant content in the presentation image and the text information recorded by the meeting recorder, which will lead to information loss when generating the meeting minutes.

[0041] To this end, the present application first provides a method for generating meeting minutes. The method for generating meeting minutes provided by the embodiments of the present application can overcome the above-mentioned defects, improve the accuracy of speech recognition, and further improve the accuracy of generating meeting minutes.

[0042] Figure 2 FIG. shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied. As Figure 2 shown, the system architecture 200 may include a first user terminal 210, a second user terminal 220, a third user terminal 230, and a cloud 240. Each user terminal is communicatively connected to the cloud 240. An online meeting system is deployed on the cloud 240, and a client of the online meeting system is installed on each user terminal. The online meeting system includes a language model, a speech recognition model, and a pre-trained meeting minutes generation model. The cloud 240 may be the execution subject of the solution of the embodiments of the present application. When a method for generating meeting minutes provided by the embodiments of the present application is applied to Figure 2 the shown system architecture, a process may be as follows: First, users of each user terminal access the online meeting system deployed on the cloud 240 through the client of the online meeting system on the user terminal and hold an online meeting. When the meeting ends, at least one user of the user terminal will send the meeting record text recorded by them to the online meeting system, and the online meeting system will also save the audio and video of the online meeting. Then, when a user of a certain user terminal (such as the user of the first user terminal 210) sends a request to the online meeting system to request the generation of meeting minutes, the online meeting system will extract the text from the presentation played in the video of the online meeting and train the language model based on at least one of the text extracted from the presentation and the meeting record text to obtain a trained language model. Finally, the online meeting system uses the speech recognition model and the trained language model to convert the audio of the online meeting into a first text sequence and inputs the first text sequence into the pre-trained meeting minutes generation model, and the meeting minutes generation model generates the meeting minutes of the online meeting.

[0043] In some embodiments of the present application, the meeting minutes generation model is pre-trained based on the text sequences corresponding to the audio of multiple other meetings and the corresponding meeting minutes.

[0044] In some embodiments of the present application, the text sequences corresponding to the audio of multiple other meetings are extracted from the audio of the corresponding meetings using the speech recognition model, and the language model is trained using the meeting record text of the corresponding meetings.

[0045] In some embodiments of the present application, the speech recognition model utilized for the text sequences extracted from the audio of multiple other meetings is different from the speech recognition model in the online meeting system.

[0046] In some embodiments of the present application, the online meeting system further includes a multi-modal model. The multi-modal model generates text based on the video frames of the online meeting, and the online meeting system also trains the language model based on the text generated from the video frames.

[0047] It should be understood that Figure 2 the number of each user terminal and the number of clouds in are merely illustrative. According to actual needs, there can be any number of user terminals and clouds, that is, the number of user terminals can be less than 3 or more than 3, and there can be multiple clouds.

[0048] It should be noted that Figure 2 only one embodiment of the present application is shown. Although in the Figure 2 solution of the embodiment, each user terminal is a laptop computer and the execution entity is the cloud, in other embodiments of the present application, the user terminal can also be various types of terminal devices such as desktop computers, smart phones, tablet computers, vehicle-mounted terminals, portable wearable devices, workstations, etc., and the execution entity can also be an ordinary server, and the types of different user terminals can be different; although in the Figure 2 solution of the embodiment, the meeting minutes are generated after the meeting ends, but in other embodiments of the present application, the meeting minutes can also be generated in real time during the meeting; although in the Figure 2 solution of the embodiment, the meeting minutes are generated according to the user's request, but in other embodiments of the present application, the meeting minutes can also be generated automatically; although in the Figure 2 solution of the embodiment, the language model, speech recognition model, and meeting minutes generation model are all part of the online meeting system, but in other embodiments of the present application, the language model, speech recognition model, and meeting minutes generation model can also be located outside the online meeting system. The embodiments of the present application do not make any limitations in this regard, and the protection scope of the present application should not be restricted thereby.

[0049] It is easy to understand that the method for generating meeting minutes provided by the embodiments of the present application is generally executed by a server. Correspondingly, the device for generating meeting minutes is generally set in the server. However, in other embodiments of the present application, the terminal device can also have a similar function to the server, so as to execute the solution for generating meeting minutes provided by the embodiments of the present application.

[0050] Therefore, the embodiments of the present application can be applied to a terminal or a server. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0051] The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below:

[0052] Figure 3 The flowchart of the method for generating a meeting minutes according to an embodiment of the present application is shown. The method for generating the meeting minutes can be executed by various devices capable of computing and processing, such as a user terminal or a cloud server. The user terminal includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, a smart watch, etc. Please refer to Figure 3 As shown, the method for generating the meeting minutes at least includes the following steps:

[0053] In step 350, according to the speech recognition model, the meeting audio of the target meeting is converted into a plurality of candidate text sequences, and each candidate text sequence is input into the language model trained according to the meeting-related text. The language model evaluates the probabilities of each candidate text sequence to obtain the probabilities of each candidate text sequence, so that the speech recognition model outputs the candidate text sequence that matches the target meeting as the first text sequence according to the probabilities of each candidate text sequence; wherein, the meeting-related text includes at least one of the following: the meeting record text of the meeting, the text extracted from the presentation of the meeting.

[0054] The speech recognition model is the Automatic Speech Recognition (ASR) model, which is a technology that converts human speech into computer-editable text. It analyzes the frequency, time domain and other characteristics of the speech signal and converts it into digital codes so that the computer can process and edit it. The speech recognition model can include an acoustic model and a decoder. The acoustic model is used to convert the speech signal into a candidate text sequence, while the language model is used to evaluate the probability of the candidate text sequence, and the decoder is used to convert the text sequence into the final recognition result according to the probability of the candidate text sequence. The language model is the domain language model.

[0055] In speech recognition, the role of a language model is to provide an understanding of the language context and grammar of the speech input to assist in more accurately transcribing the speech into text. Specifically, the role of the language model in speech recognition includes: 1) Context understanding: The language model helps the recognition system better understand the speech input by considering the language context before and after. It can predict the next possible word based on the previous words and sentence structure, thus providing a more accurate transcription result. 2) Word selection: The speech recognition system may face problems such as polysemy of words or recognition of rare words. The language model, by learning a large amount of corpus data and understanding the probability distribution of words, helps the system select the most likely word in case of uncertainty. 3) Error correction: The speech recognition system may make errors during the transcription process, such as word order errors or substitution errors. The language model can help the system correct these errors by modeling language rules and common phrases, making the transcription result more accurate and fluent. Generally speaking, the role of the language model in speech recognition is to provide support for language context, word selection, and error correction to improve the accuracy and comprehensibility of speech transcription. It is an essential part of the speech recognition system and can provide a more natural and accurate text output. The language model is usually implemented using an n-gram model or a neural network model. The n-gram model is a statistics-based model that predicts the probability of the next word or character by calculating the occurrence probability of n consecutive words or characters in a text sequence. The neural network model is a deep learning-based model that predicts the probability of the next word or character by learning the relationships between words or characters in a text sequence.

[0056] The meeting record text of a meeting is the text obtained by the user's recording of the meeting content, which can include, for example, the key information of the meeting; text can be extracted from the meeting presentation through Optical Character Recognition (OCR) technology. The meeting presentation, that is, the slides, usually contains text; the meeting presentation usually exists in the form of video frames of the meeting video. Therefore, OCR technology can be used to recognize it. Of course, if the system can directly obtain the meeting presentation, such as obtaining a presentation file in the format of.ppt or.pptx, the text can be directly extracted from it. In addition, even if the meeting presentation can be directly obtained, a video file of the presentation file can be obtained by opening the presentation and recording, and then recognized. Optical Character Recognition (OCR) refers to the process of analyzing and recognizing an image file of text materials to obtain the text and layout information.

[0057] In one embodiment of the present application, the language model is trained based on the meeting record text of the target meeting and the text extracted from the presentation of the target meeting.

[0058] That is to say, in the embodiment of the present application, the language model in the speech recognition model for text conversion of the meeting audio of a certain meeting is trained based on the meeting record text and the presentation text of the same meeting. In this way, the training text of the language model is directly related to the current meeting, enabling the speech recognition model to focus on the content of this meeting during speech recognition, thereby greatly improving the speech recognition effect of the meeting.

[0059] Certainly, the language model can also be further trained based on the meeting record text and presentation text of other meetings. For example, the meeting record text and presentation text of other meetings related to or similar to the theme of the target meeting can be used for training the language model. The relevance or similarity of other meetings to the target meeting can be judged by the overlap degree of the keywords in the meeting record text and presentation text of other meetings with those in the meeting record text and presentation text of the target meeting. That is, by training the language model, the context understanding and vocabulary selection of the language model are more closely centered around the meeting for which the minutes are to be generated currently, thus ensuring the speech recognition effect of the presentation audio, especially for some specific words mentioned in the meeting, making the context of the speech recognition text more fitting.

[0060] Since the training duration of the language model is very short, even if the language model is trained after the meeting ends, it will not significantly increase the time-consuming for generating the meeting minutes.

[0061] In addition, the video frames of the target meeting may contain video images, such as the video pictures played during the meeting. Therefore, the video images can also be input into the multi-modal large model, and the multi-modal large model can generate corresponding explanatory text based on the video images. The language model can also be further trained based on the explanatory text of the target meeting or other meetings. In this way, the text used for training the language model can be further enriched, and the relevance between the language model and the meeting can be improved.

[0062] Figure 4 Shows the training architecture diagram of the meeting minutes generation model according to an embodiment of the present application. Please refer to Figure 4As shown, although it is an architecture used in training a meeting minutes generation model, some of its content is consistent with the usage principle of the meeting minutes generation model. For example, first, use an OCR model to identify text from the presentation of the meeting, and then use the meeting notes and the OCR-identified files to train a domain language model, and embed the domain language model into the ASR model. The ASR model recognizes the speech audio of the meeting to obtain the ASR recognition text. Of course, it is easy to understand that the meeting here is not the target meeting, but the meeting corresponding to the training data in the training stage of the meeting minutes generation model.

[0063] Please continue to refer to Figure 3 , in step 380, generate the meeting minutes of the target meeting according to the first text sequence, and the meeting minutes are an overview of the meeting content of the target meeting.

[0064] The meeting minutes are a summary of the overall meeting content of the target meeting.

[0065] Please continue to refer to Figure 4 As shown, whether in the training stage or the usage stage of the meeting minutes generation model, it is necessary to input the text recognized from the audio of the meeting into the meeting minutes generation model.

[0066] Figure 5 shows a Figure 3 flowchart of the steps before step 380 and the details of step 380 in an embodiment according to the present application. Please refer to Figure 5 As shown, before generating the meeting minutes of the target meeting according to the first text sequence, the method for generating the meeting minutes may further include the following steps:

[0067] In step 360, obtain the meeting record text of the target meeting.

[0068] The meeting record text of the target meeting may be text manually recorded by one or more participants of the target meeting, such as recording the key information of the meeting.

[0069] Please continue to refer to Figure 5 As shown, generating the meeting minutes of the target meeting according to the first text sequence may specifically include the following steps:

[0070] In step 380', input the first text sequence and the meeting record text into a pre-trained meeting minutes generation model to obtain the meeting minutes of the target meeting generated by the meeting minutes generation model.

[0071] The meeting minutes generation model can be used to fuse the first text sequence and the meeting record text. The meeting minutes generation model will summarize and analyze the first text sequence and the meeting record text to obtain the meeting minutes of the target meeting. By fusing the first text sequence and the meeting record text to obtain the meeting minutes, the generated meeting minutes can be made more accurate and complete.

[0072] Figure 6 shows a flowchart of the details of the steps before step 380 and step 380' in an embodiment according to the present application. Please refer to Figure 5 As shown, before generating the meeting minutes of the target meeting according to the first text sequence, the method for generating the meeting minutes may further include the following steps: Figure 6 As shown, before generating the meeting minutes of the target meeting according to the first text sequence, the method for generating the meeting minutes may further include the following steps:

[0073] In step 370, a second text sequence is extracted from the presentation of the target meeting.

[0074] As described above, the OCR model can be used to extract the second text sequence from the video frames corresponding to the presentation of the target meeting. The OCR model can recognize the character information from the video frames and splice the recognized character information in the order of coordinates from top to bottom and from left to right to form the second text sequence corresponding to the presentation. In addition, for the case where the presentation file can be directly obtained, the second text sequence can be directly extracted from the presentation. No restrictions are imposed on the OCR model here, and any available model can be used. The OCR model may include a detection module or an identification module, and the OCR model may also be an end-to-end model. This model structure enables the two subtasks of detection and identification to share the features learned by the convolutional layer, thereby saving computing power and time.

[0075] Please continue to refer to Figure 6 As shown, the first text sequence and the meeting record text are input into a pre-trained meeting minutes generation model to obtain the meeting minutes of the target meeting generated by the meeting minutes generation model, which may specifically include the following steps:

[0076] In step 380", the first text sequence, the second text sequence, and the meeting record text are input into a pre-trained meeting minutes generation model to obtain the meeting minutes of the target meeting generated by the meeting minutes generation model.

[0077] The meeting minutes generation model can be used to fuse the first text sequence, the second text sequence, and the meeting record text.

[0078] In the embodiment of the present application, by using the first text sequence, the second text sequence, and the meeting record text, these three types of texts related to the target meeting to generate the meeting minutes, the accuracy and comprehensiveness of generating the meeting minutes are improved.

[0079] Of course, in other embodiments of the present application, the meeting minutes of the target meeting can also be obtained by fusing the first text sequence and the second text sequence.

[0080] In addition, as mentioned above, the explanatory text of the target meeting can be generated based on the video image of the target meeting. Therefore, the explanatory text of the target meeting can be further fused with the first text sequence, the second text sequence, the meeting record text, etc. to generate the meeting minutes of the target meeting.

[0081] Figure 7 Shows a Figure 5 flowchart of the steps before step 380' in an embodiment of the present application. Please refer to Figure 7 As shown, before inputting the first text sequence and the meeting record text into the pre-trained meeting minutes generation model, the method for generating the meeting minutes may further include the following steps:

[0082] In step 310, the meeting audio of multiple meetings is respectively converted into a third text sequence based on a speech recognition model.

[0083] The multiple meetings are meetings used to provide training samples for the training of the meeting minutes generation model. The multiple meetings are usually meetings that have been completed before the target meeting. The topics of the multiple meetings can be the same as or different from the topic of the target meeting.

[0084] Please continue to refer to Figure 4 As shown, in the training stage of the meeting minutes generation model, the domain language model in the speech recognition model also needs to be trained based on the notes and the OCR recognition text of the relevant meetings. Specifically, first, the text is recognized from the presentation of the meeting by using the OCR model, and then the domain language model is trained by using the notes of the meeting and the OCR recognition file, and the domain language model is embedded into the ASR model. The ASR model recognizes the speech audio of the meeting to obtain the ASR recognition text, that is, the third text sequence in this step.

[0085] In step 320, a fourth text sequence is respectively extracted from the presentations of multiple meetings.

[0086] Although Figure 4 not shown, in practical applications, the meeting minutes generation model can be trained in combination with the text in the presentation of the meeting.

[0087] In step 330, the meeting minutes and the meeting record text of multiple meetings are obtained.

[0088] Please continue to refer to Figure 4As shown, the training of the meeting minutes generation model also requires the use of meeting minutes and meeting record texts (i.e., the notes of the meeting). Here, the meeting minutes are manually summarized and recorded by experts or specialized meeting moderators based on experience. The meeting record texts can also be recorded by the participants.

[0089] In step 340, using the third text sequence, the fourth text sequence, the meeting minutes, and the meeting record text of each meeting as a training sample, a meeting minutes generation model is trained based on the training samples corresponding to multiple meetings.

[0090] The meeting minutes generation model can adopt various models that can generate corresponding summary information according to the text. For example, a large language model can be adopted.

[0091] Figure 8 The flowchart of training a meeting minutes generation model based on the training samples corresponding to multiple meetings according to an embodiment of the present application is shown. Please refer to Figure 8 As shown, training a meeting minutes generation model based on the training samples corresponding to multiple meetings may specifically include the following steps:

[0092] In step 810, input prompt information is constructed according to the third text sequence, the fourth text sequence, and the meeting record text in the training sample.

[0093] The input prompt information is the information that needs to be input into the model and encoded by the model.

[0094] In an embodiment of the present application, constructing input prompt information according to the third text sequence, the fourth text sequence, and the meeting record text in the training sample includes:

[0095] For each training sample, randomly adopt one of the following methods to construct the input prompt information corresponding to the training sample:

[0096] Use the third text sequence and the fourth text sequence in the training sample as the input prompt information; or

[0097] Use the third text sequence, the fourth text sequence, and the meeting record text in the training sample as the input prompt information.

[0098] Specifically, if the combination of the third text sequence and the fourth text sequence is M1, and the combination of the third text sequence, the fourth text sequence, and the meeting record text is M2, that is, M2 = {M1, T}, where T is the meeting record text. Therefore, when constructing the corresponding input prompt information for each training sample, randomly select M1 or M2 corresponding to the training sample as the input prompt information.

[0099] During the actual meeting process, the meeting record text may not exist. In the embodiments of the present application, by randomly selecting M1 or M2 as the input prompt information corresponding to the training sample during training, the situation of missing meeting record text can be effectively simulated, and the performance of model training can be improved.

[0100] In step 820, the meeting minutes in the training sample are used as the input and label of the training sample.

[0101] The label is the training target of the model, that is, the correct output sequence of the model. During the training process, the input of the model is the real sequence, while the output is the sequence generated by the model itself. By using the correct output sequence as the training target of the model, the model can learn the correct output sequence faster.

[0102] In step 830, a meeting minutes generation model is trained based on the input prompt information, input, and label of each training sample.

[0103] As Figure 4 shown, during the training process of the model, the meeting minutes are also input into the meeting minutes generation model.

[0104] Specifically, the input prompt information is input into the encoder of the model, encoded into a hidden state by the encoder. The decoder uses the hidden state from the encoder as context information and gradually predicts and generates the target output sequence based on the context information and the input. The target output sequence is compared with the label and the loss is calculated, and the parameters of the model are adjusted according to the loss, thereby training the model.

[0105] In an embodiment of the present application, the meeting minutes generation model includes multiple Transformer layers, and each Transformer layer contains a masked multi-head attention layer. The vector corresponding to the token of each word in the input of the training sample can perform attention calculation with the vector corresponding to the token of each word in the input prompt information corresponding to the training sample; the vector corresponding to the token of the current word in the input of the training sample can only perform attention calculation with the vector corresponding to the token of the current word and the vectors corresponding to the tokens of the words before the current word.

[0106] The meeting minutes generation model can be stacked by L layers of Transformer layers to improve the accommodation and learning ability of the model. A single Transformer layer contains masked multi-head attention, feed forward, and layer normalization.

[0107] Figure 9 shows a schematic diagram of a mask matrix according to an embodiment of the present application. Please refer to Figure 9As shown, assume that the sequence of tokens of each word in the input of the training sample is M, and the sequence of tokens of each word in the corresponding input prompt information of the training sample is S. The filled squares in the mask matrix represent masked, and the blank squares in the mask matrix represent unmasked. The upper left corner of the mask matrix is a blank square, which means that the tokens in sequence M are visible to each other, that is, the current token can see all the tokens in M. Visible means that attention calculation can be performed; the upper right corner of the mask matrix is filled with shading, which means that all tokens in sequence M cannot see any tokens in sequence S; the lower left corner of the mask matrix is a blank square, which means that all tokens in sequence S can see all tokens in sequence M; only the squares below the diagonal in the lower left part of the mask matrix are blank squares, which means that the current token can only see historical tokens, that is, the current token can only perform attention calculation with the vectors corresponding to the tokens that the decoder has already output, and cannot see future tokens.

[0108] Of course, in other embodiments of the present application, the first text sequence, the second text sequence, and the meeting record text can also be fused based on the meeting minutes generation model to obtain the meeting minutes of the target meeting. In addition, the interpretation information of the meeting video generated by using the multimodal large model can be further input into the meeting minutes generation model, so as to further generate the meeting minutes comprehensively and accurately.

[0109] In summary, according to the method for generating meeting minutes provided by the embodiments of the present application, multi-modal information such as images, audio, and text is used and fused and summarized, so as to automatically generate meeting minutes. Users no longer need to manually record meeting minutes, which not only greatly reduces the workload of users, but also improves the comprehensiveness and accuracy of meeting minutes generation, and can be effectively applied to the automatic generation task of meeting minutes for online meetings; since the language model in the speech recognition model is trained using other texts of relevant meetings, the context of the speech recognition text can better fit the content of the current meeting, thereby further improving the accuracy of generating meeting minutes.

[0110] The following introduces the device embodiments of the present application, which can be used to execute the method for generating meeting minutes in the above embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the embodiments of the method for generating meeting minutes in the above of the present application.

[0111] Figure 10 The block diagram of the device for generating meeting minutes according to an embodiment of the present application is shown.

[0112] Refer to Figure 10As shown, a device 1000 for generating a meeting minutes according to an embodiment of the present application includes: a conversion unit 1010 and a generation unit 1020. Among them, the conversion unit 1010 is configured to convert the meeting audio of a target meeting into a plurality of candidate text sequences according to a speech recognition model, and input each of the candidate text sequences into a language model trained according to meeting-related texts, and the language model evaluates each of the candidate text sequences to obtain the probability of each of the candidate text sequences, so that the speech recognition model outputs the candidate text sequence that matches the target meeting as the first text sequence according to the probability of each of the candidate text sequences; wherein, the meeting-related texts include at least one of the following: the meeting record text of the meeting, the text extracted from the presentation of the meeting; the generation unit 1020 is configured to generate the meeting minutes of the target meeting according to the first text sequence, and the meeting minutes is an overview of the meeting content of the target meeting.

[0113] In some embodiments of the present application, based on the foregoing solution, the device further includes a meeting record text acquisition unit; before generating the meeting minutes of the target meeting according to the first text sequence, the meeting record text acquisition unit is configured to: acquire the meeting record text of the target meeting; the generation unit 1020 is configured to: input the first text sequence and the meeting record text into a pre-trained meeting minutes generation model to obtain the meeting minutes of the target meeting generated by the meeting minutes generation model.

[0114] In some embodiments of the present application, based on the foregoing solution, the device further includes an extraction unit; before generating the meeting minutes of the target meeting according to the first text sequence, the extraction unit is configured to: extract a second text sequence from the presentation of the target meeting; the generation unit 1020 is configured to: input the first text sequence, the second text sequence and the meeting record text into a pre-trained meeting minutes generation model to obtain the meeting minutes of the target meeting generated by the meeting minutes generation model.

[0115] In some embodiments of the present application, based on the foregoing solution, the device further includes a model training unit; before inputting the first text sequence and the meeting record text into a pre-trained meeting minutes generation model, the model training unit is configured to: convert the meeting audio of a plurality of meetings into third text sequences respectively based on a speech recognition model; extract fourth text sequences from the presentations of the plurality of meetings respectively; acquire the meeting minutes and meeting record texts of the plurality of meetings; use the third text sequence, the fourth text sequence, the meeting minutes and the meeting record text of each meeting as a training sample, and train a meeting minutes generation model based on the training samples corresponding to the plurality of meetings.

[0116] In some embodiments of the present application, based on the foregoing solution, the model training unit is configured to: construct input prompt information according to the third text sequence, the fourth text sequence, and the meeting record text in the training sample; use the meeting minutes in the training sample as the input and label of the training sample; and train a meeting minutes generation model based on the input prompt information, input, and label of each training sample.

[0117] In some embodiments of the present application, based on the foregoing solution, the model training unit is configured to: for each training sample, randomly construct the input prompt information corresponding to the training sample in one of the following ways: use the third text sequence and the fourth text sequence in the training sample as the input prompt information; or, use the third text sequence, the fourth text sequence, and the meeting record text in the training sample as the input prompt information.

[0118] In some embodiments of the present application, based on the foregoing solution, the meeting minutes generation model includes multiple Transformer layers, and each Transformer layer includes a masked multi-head attention layer. The vector corresponding to the token of each word in the input of the training sample can perform attention calculation with the vector corresponding to the token of each word in the input prompt information corresponding to the training sample; the vector corresponding to the token of the current word in the input of the training sample can only perform attention calculation with the vector corresponding to the token of the current word and the vectors corresponding to the tokens of the words before the current word.

[0119] In some embodiments of the present application, based on the foregoing solution, the extraction unit is configured to: recognize the second text sequence from the presentation image of the target meeting based on an optical character recognition algorithm.

[0120] In some embodiments of the present application, based on the foregoing solution, the language model is trained based on the meeting record text of the target meeting and the text extracted from the presentation of the target meeting.

[0121] Figure 11 The structure diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0122] It should be noted that Figure 11 The computer system 1100 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0123] As Figure 11As shown, the computer system 1100 includes a Central Processing Unit (CPU) 1101, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 1102 or a program loaded from a storage section 1108 into a Random Access Memory (RAM) 1103, such as executing the methods described in the above embodiments. In the RAM 1103, various programs and data required for system operation are also stored. The CPU 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An Input / Output (I / O) interface 1105 is also connected to the bus 1104.

[0124] The following components are connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from it can be installed into the storage section 1108 as needed.

[0125] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1109, and / or installed from the removable medium 1111. When the computer program is executed by a Central Processing Unit (CPU) 1101, various functions defined in the system of the present application are executed.

[0126] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0128] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the unit itself.

[0129] As one aspect, the present application also provides a computer-readable medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.

[0130] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0131] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0132] It can be understood that in the specific embodiments of the present application, data related to the meeting content is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0133] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0134] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for generating meeting minutes, characterized in that, The method includes: Converting the conference audio of a target conference into a number of candidate text sequences according to a speech recognition model, and inputting each of the candidate text sequences into a language model trained according to conference-related texts. The language model evaluates each of the candidate text sequences to obtain the probabilities of each of the candidate text sequences, so that the speech recognition model outputs the candidate text sequence that matches the target conference as the first text sequence according to the probabilities of each of the candidate text sequences; wherein, the conference-related texts include at least one of the following: the meeting record text of the conference, the text extracted from the presentation of the conference. Generating a meeting summary of the target conference according to the first text sequence, where the meeting summary is an overview of the meeting content of the target conference.

2. The method for generating the meeting minutes according to claim 1, wherein Before generating the meeting summary of the target conference according to the first text sequence, the method further includes: Obtaining the meeting record text of the target conference. The generating the meeting summary of the target conference according to the first text sequence includes: Inputting the first text sequence and the meeting record text into a pre-trained meeting summary generation model to obtain the meeting summary of the target conference generated by the meeting summary generation model.

3. The method for generating a meeting minutes according to claim 2, wherein, Before generating the meeting summary of the target conference according to the first text sequence, the method further includes: Extracting a second text sequence from the presentation of the target conference. The inputting the first text sequence and the meeting record text into a pre-trained meeting summary generation model to obtain the meeting summary of the target conference generated by the meeting summary generation model includes: Inputting the first text sequence, the second text sequence, and the meeting record text into a pre-trained meeting summary generation model to obtain the meeting summary of the target conference generated by the meeting summary generation model.

4. The method for generating the meeting minutes according to claim 2, wherein Before inputting the first text sequence and the meeting record text into a pre-trained meeting summary generation model, the method further includes: Converting the conference audio of multiple conferences into third text sequences respectively based on a speech recognition model. Extracting fourth text sequences from the presentations of the multiple conferences respectively. Obtaining the meeting summaries and meeting record texts of the multiple conferences. Using the third text sequence, the fourth text sequence, the meeting minutes, and the meeting record text of each meeting as a training sample, a meeting minutes generation model is trained based on the training samples corresponding to the multiple meetings. 。 5. The method for generating a meeting summary according to claim 4, wherein The training the meeting summary generation model based on the training samples corresponding to the multiple conferences includes: Constructing input prompt information according to the third text sequence, the fourth text sequence, and the meeting record text in the training samples. Using the meeting summary in the training samples as the input and label of the training samples. Training a meeting summary generation model based on the input prompt information, input, and label of each training sample.

6. The method for generating the meeting minutes according to claim 5, wherein, The constructing the input prompt information according to the third text sequence, the fourth text sequence, and the meeting record text in the training samples includes: For each training sample, randomly constructing the input prompt information corresponding to the training sample in one of the following ways: Using the third text sequence and the fourth text sequence in the training sample as the input prompt information; or Use the third text sequence, the fourth text sequence, and the meeting record text in the training samples as input prompt information.

7. The method for generating the meeting minutes according to claim 4, characterized in that The meeting minutes generation model includes multiple Transformer layers, and each Transformer layer contains a masked multi-head attention layer. The vector corresponding to the token of each word in the input of the training sample can perform an attention calculation with the vector corresponding to the token of each word in the input prompt information corresponding to the training sample; the vector corresponding to the token of the current word in the input of the training sample can only perform an attention calculation with the vector corresponding to the token of the current word and the vectors corresponding to the tokens of the words before the current word.

8. The method for generating a meeting minutes according to claim 3, wherein, The extracting the second text sequence from the presentation of the target meeting includes: Identifying the second text sequence from the presentation image of the target meeting based on an optical character recognition algorithm.

9. The method for generating a meeting minutes according to any one of claims 1-8, characterized in that The language model is trained based on the meeting record text of the target meeting and the text extracted from the presentation of the target meeting.

10. A device for generating meeting minutes, characterized in that, The device includes: A conversion unit configured to convert the meeting audio of a target meeting into a plurality of candidate text sequences according to a speech recognition model, and input each of the candidate text sequences into a language model trained according to meeting-related text. The language model evaluates each of the candidate text sequences to obtain the probabilities of each of the candidate text sequences, so that the speech recognition model outputs, according to the probabilities of each of the candidate text sequences, the candidate text sequence that matches the target meeting as the first text sequence; wherein, the meeting-related text includes at least one of the following: the meeting record text of the meeting, the text extracted from the presentation of the meeting. A generating unit configured to generate meeting minutes of the target meeting according to the first text sequence, where the meeting minutes are an overview of the meeting content of the target meeting.

11. A computer-readable medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method for generating meeting minutes according to any one of claims 1 to 9.

12. An electronic device, characterized in that, including: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method for generating meeting minutes according to any one of claims 1 to 9.

13. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for generating meeting minutes according to any one of claims 1 to 9.

Citation Information

Cited By

  • Audio transfer and summary generation method and device for conference

    CN120636411A