Method and apparatus for generating meeting minutes, electronic device and medium

By performing speech recognition and semantic understanding on meeting audio, meeting minutes are generated, solving the problem of low efficiency in existing technologies and improving the accuracy of meeting minutes and user experience.

CN116756367BActive Publication Date: 2026-05-22APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
Filing Date
2023-06-09
Publication Date
2026-05-22

Smart Images

  • Figure CN116756367B_ABST
    Figure CN116756367B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device, electronic device, computer readable storage medium and computer program product for generating meeting minutes, relates to the field of natural language processing, and particularly relates to the field of speech recognition and keyword extraction. The implementation scheme is: obtaining first conference audio, wherein the first conference audio includes information indicating that a request execution person is requested to execute a to-do item; obtaining second conference audio according to the first conference audio, wherein the second conference audio is associated with the execution person, and the second conference audio includes information indicating whether the execution person commits to execute the to-do item; and in response to determining that the execution person commits to execute the to-do item according to the second conference audio, generating a target meeting minutes based on the first conference audio, wherein the target meeting minutes indicate that the to-do item is executed by the execution person.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of natural language processing, and more particularly to the fields of speech recognition and keyword extraction, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating meeting minutes. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] In related technologies, meeting minutes are obtained by manually recording the tasks to be done in a meeting and the person responsible for executing those tasks. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating meeting minutes.

[0005] According to one aspect of this disclosure, a method for generating meeting minutes is provided, comprising: acquiring a first meeting audio, wherein the first meeting audio includes information instructing a requester to perform a task; acquiring a second meeting audio based on the first meeting audio, wherein the second meeting audio is associated with the requester and includes information instructing the requester whether they have committed to performing the task; and generating a target meeting minute based on the first meeting audio in response to determining, based on the second meeting audio, that the target meeting minute instructs the requester to perform the task.

[0006] According to another aspect of this disclosure, a method for training a keyword extraction model is provided, comprising: acquiring multiple historical meeting texts, wherein each historical meeting text indicates a request for a historical executor to perform a historical task, each historical meeting text including a seventh keyword and an eighth keyword, wherein the seventh keyword indicates the historical executor and the eighth keyword indicates the historical task; determining the semantic structure of each historical text data; performing semantic annotation on the seventh keyword and the eighth keyword in the historical text data; and training an initial model based on the annotated multiple historical text data and the corresponding semantic structure of each historical text data to obtain the keyword extraction model.

[0007] According to another aspect of this disclosure, an apparatus for generating meeting minutes is provided, comprising: a first acquisition unit configured to acquire a first meeting audio, wherein the first meeting audio includes information instructing a requester to perform a task; a second acquisition unit configured to acquire a second meeting audio based on the first meeting audio, wherein the second meeting audio is associated with the requester and includes information instructing the requester whether they have committed to performing the task; and a generation unit configured to generate a target meeting minute based on the first meeting audio in response to determining, based on the second meeting audio, that the requester has committed to performing the task, wherein the target meeting minute instructs the requester to perform the task.

[0008] According to another method of this disclosure, a training apparatus for a keyword extraction model is provided, comprising: a third acquisition unit configured to acquire a plurality of historical meeting texts, wherein each historical meeting text indicates a request for a historical executor to perform a historical task, each historical meeting text including a seventh keyword and an eighth keyword, wherein the seventh keyword indicates the historical executor and the eighth keyword indicates the historical task; a processing unit configured to determine the semantic structure of each historical text data; and to perform semantic annotation on the seventh keyword and the eighth keyword in the historical text data; and a training unit configured to train an initial model based on the annotated plurality of historical text data and the corresponding semantic structure of each historical text data to obtain the keyword extraction model.

[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method for generating meeting minutes and the method for training a keyword extraction model as described above.

[0010] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods for generating meeting minutes and training a keyword extraction model as described above.

[0011] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method for generating meeting minutes and the method for training a keyword extraction model as described above.

[0012] According to one or more embodiments of this disclosure, by performing speech recognition and semantic understanding on the first meeting audio of the speaker who raises the to-do item and the second meeting audio of the person designated to perform the to-do item, it is possible to determine in a timely and accurate manner, based on the question-and-answer content of the two, whether meeting minutes of the corresponding to-do item need to be generated and the content of the generated meeting minutes, thereby improving the user experience of the participants.

[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0014] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0015] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;

[0016] Figure 2 An exemplary flowchart of a method for generating meeting minutes according to embodiments of the present disclosure is shown;

[0017] Figure 3 A partial exemplary flowchart of a method for generating meeting minutes according to embodiments of the present disclosure is shown;

[0018] Figure 4 Another exemplary flowchart of a method for generating meeting minutes according to embodiments of the present disclosure is shown;

[0019] Figure 5 Further exemplary flowcharts of a method for generating meeting minutes according to embodiments of the present disclosure are shown;

[0020] Figure 6 An exemplary flowchart of a training method for a keyword extraction model according to an embodiment of the present disclosure is shown;

[0021] Figure 7 A structural block diagram of an apparatus for generating meeting minutes according to an embodiment of the present disclosure is shown.

[0022] Figure 8 A structural block diagram of a training apparatus for a keyword extraction model according to embodiments of the present disclosure is shown; and

[0023] Figure 9 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0025] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0026] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0027] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0028] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0029] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of methods for generating meeting minutes.

[0030] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) network.

[0031] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0032] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to generate meeting minutes for to-do lists. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to users through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0033] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0034] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0035] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0036] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0037] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0038] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0039] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and text files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0040] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0041] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0042] In related technologies, the process of manually recording the tasks proposed in a meeting and the person responsible for executing those tasks to obtain the meeting minutes is inefficient and inaccurate.

[0043] to this end, Figure 2 An exemplary flowchart of a method for generating meeting minutes according to embodiments of the present disclosure is shown.

[0044] like Figure 2 As shown, an embodiment of this disclosure provides a method 200 for generating meeting minutes, comprising: acquiring a first meeting audio, wherein the first meeting audio includes information instructing a requester to perform a pending task (step 210); acquiring a second meeting audio based on the first meeting audio, wherein the second meeting audio is associated with the executor and includes information instructing the executor whether they have committed to performing the pending task (step 220); and generating a target meeting minutes based on the first meeting audio in response to determining, based on the second meeting audio, that the executor has committed to performing the pending task (step 230).

[0045] By performing speech recognition and semantic understanding on the first meeting audio of the speaker who raised the task and the second meeting audio of the person designated to carry out the task, the system can determine in a timely and accurate manner whether to generate meeting minutes for the corresponding task and the content of the generated meeting minutes based on the question-and-answer content of the two meetings, thereby improving the user experience for participants.

[0046] In step 210, a first conference audio is obtained, wherein the first conference audio includes information instructing the requester to perform the pending task.

[0047] In some embodiments, audio data of participants' speeches can be collected through a microphone device, and the collected audio data can be converted into text data after speech recognition. Then, semantic understanding can be performed on the text data to determine that the collected audio data is the aforementioned first meeting audio when the current speech content is identified as a request for an executor to perform a to-do item.

[0048] Figure 3 A partial example flowchart of a method for generating meeting minutes according to embodiments of the present disclosure is shown.

[0049] According to some embodiments, such as Figure 3As shown, step 210 includes: acquiring a third meeting audio (step 310); performing speech recognition on the third meeting audio to obtain a first meeting text (step 320); extracting keywords from the first meeting text using a target keyword extraction model, wherein the keyword extraction model is configured to identify at least a first keyword and a second keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, and wherein the first keyword indicates an executor and the second keyword indicates a to-do item (step 330); and in response to identifying the first keyword and the second keyword from the meeting text, using the third meeting audio as the first meeting audio (step 340).

[0050] In step 310, the audio of the third meeting is acquired.

[0051] In some embodiments, the aforementioned third conference audio may be all or part of the audio data acquired in real time in a streaming manner. In other embodiments, the aforementioned third conference audio may also be all or part of the recorded audio data of the conference acquired after the conference has ended.

[0052] In step 320, speech recognition is performed on the third meeting audio to obtain the first meeting text.

[0053] In some embodiments, a speech recognition model can be used to perform speech recognition for a third meeting. The aforementioned speech recognition model can be built based on machine learning models such as UBM-GMM and SVM, or on neural networks such as DNN, CNN, LSTM, Conformer, and TDNN. Understandably, other types of network structures can be used for the speech recognition model, and no limitations are imposed here.

[0054] In step 330, the target keyword extraction model is used to extract keywords from the first meeting text, wherein the keyword extraction model is configured to identify at least a first keyword and a second keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, and wherein the first keyword indicates the executor and the second keyword indicates the to-do item.

[0055] In some embodiments, the relationship between to-do items and executors can be one-to-one, one-to-many, or many-to-many. The target keyword extraction model described above can determine the keywords for to-do items (first keywords), the keywords for executors (second keywords), and whether there is an execution-being-executed relationship between the two based on the semantic structure of the meeting text and the meaning of each word, thus accurately generating meeting minutes for to-do items.

[0056] In one example, if the meeting text is "Zhang San, you book the flight", the target keyword extraction model can extract "Zhang San" as the executor and "book the flight" as the to-do item, and then associate them.

[0057] In another example, if the meeting text is “Zhang San and Li Si, you need to book flights and hotels,” the target keyword extraction model can extract “Zhang San” and “Li Si” as executors, and “book flights” and “book hotels” as to-do items.

[0058] In some embodiments, to-do items in the meeting text may need to be performed immediately. For such to-do items, the target keyword extraction model can discard keywords in real time based on semantic understanding to further improve the accuracy of the generated target meeting minutes.

[0059] In one example, the meeting text is "Wang Wu, can you share your screen?" In this case, "share your screen" can be regarded as a to-do item. However, such items need to be completed immediately, so there is no need to generate meeting minutes, and they can be discarded.

[0060] In some embodiments, the target keyword extraction model can also achieve event retrieval. Event retrieval refers to a situation where the semantic structure of the meeting text is not the conventional "executor + task" format, but rather a more specific "task + executor". In this case, the target keyword extraction model can extract the subsequent executor keywords and associate them with the preceding task keywords based on the determined semantic structure, so as to make the identified keywords and the associations between keywords more accurate.

[0061] In one example, the meeting text is "We need to book a meeting room for a meeting. Li Si, you can book it." In this case, according to the text order, the target keyword extraction model first extracts the task keyword "book a meeting room" and then extracts the executor keyword "Li Si". Based on the semantic structure of the event retrieval, it can be determined that there is an execution and being executed relationship between the two. Therefore, they can be used as two keywords associated with the meeting minutes of the same task.

[0062] Step 340: In response to the identification of the first keyword and the second keyword from the meeting text, the third meeting audio is used as the first meeting audio.

[0063] In some embodiments, a meeting often needs to discuss multiple topics. By using speech recognition and keyword extraction, it is possible to accurately and effectively identify whether there is content in the third meeting audio that instructs the requester to perform a task, that is, whether it can be used as the first meeting audio of the requesting speaker, so as to determine whether it is necessary to obtain the corresponding second meeting audio of the responding speaker based on the content of the request.

[0064] In step 220, a second meeting audio is obtained based on the first meeting audio, wherein the second meeting audio is associated with the executor and includes information indicating whether the executor has committed to performing the pending tasks.

[0065] In some embodiments, when a speaker raises a task and designates an executor, the executor typically responds immediately. In one example, the conference audio can be segmented based on the speaker; after identifying the first audio segment, the audio segment of the next speaker following the first segment is used as the second audio segment. In another example, for streaming conference audio data captured in a real-time conference, if it is determined that the current speaker's audio data contains information requesting an executor to perform a task, the audio data of the next speaker can be acquired and used as the second audio segment.

[0066] In some embodiments, before the executor responds to whether they agree to perform the task, other speakers may make other statements. In this case, a second meeting audio can be obtained based on the executor's meeting account and / or voiceprint information.

[0067] Figure 4 Another exemplary flowchart of a method for generating meeting minutes according to embodiments of the present disclosure is shown.

[0068] According to some embodiments, such as Figure 4 As shown, the above meeting is an online meeting. Step 220 includes: obtaining the meeting account used by the executor to log in to the online meeting based on the first keyword (step 410); and obtaining the fourth meeting audio associated with the meeting account within the first time period after the speaking time of the first meeting audio as the second meeting audio (step 420).

[0069] In some embodiments, each participant has and uses their own unique meeting account to attend and speak at the meeting. Therefore, after identifying the executor's keywords, their meeting account can be determined based on their identity information. This allows for the retrieval of the second meeting audio recording, where the executor responds to the to-do items in the first meeting audio recording. Based on this, the second meeting audio recording can be obtained efficiently and accurately, improving the efficiency of generating meeting minutes for to-do items.

[0070] According to some embodiments, the keyword extraction model described above is further configured to extract a third keyword and a fourth keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text. The third keyword indicates the execution time of the to-do item, and the fourth keyword indicates the execution location of the to-do item. Step 230 includes generating a target meeting minutes based on the first keyword, the second keyword, the third keyword, and the fourth keyword.

[0071] In some embodiments, to-do items are typically associated with time and location. The keyword extraction model described above can also extract time keywords (third keywords) and location keywords (fourth keywords) from the meeting text, thereby enabling the generated meeting minutes to include richer information.

[0072] In some embodiments, the executor's account nickname can be further added to the meeting minutes, and online reminders can be sent to the executor to improve the user experience.

[0073] In one example, Zhang San's meeting account nickname is "Xiao Zhang", and the meeting text is "Zhang San, go to the first meeting room at 3 pm to set up the meeting". Based on this, the generated to-do meeting minutes can be "Zhang San@Xiao Zhang, 3 pm, first meeting room, set up the meeting".

[0074] In some embodiments, participants speak in offline meetings, or participate in and speak in online meetings using other people's meeting accounts. Based on this, it is not possible to determine the second meeting audio through the meeting account.

[0075] In response, Figure 5 An exemplary flowchart of another part of a method for generating meeting minutes according to embodiments of the present disclosure is shown.

[0076] According to some embodiments, such as Figure 5 As shown, the above meeting can be an online meeting or an offline meeting. Step 220 includes: obtaining the voiceprint information of the executor based on the first keyword (step 510); and obtaining the fifth meeting audio associated with the voiceprint information in the second time period after the speaking time of the first meeting audio as the second meeting audio (step 520).

[0077] In some embodiments, the voiceprint information of the executor can be collected in advance and a voiceprint information database can be established, so that the second meeting audio can be accurately determined based on the individual's unique voiceprint information, thereby improving the effectiveness of generating meeting minutes for to-do items.

[0078] According to some embodiments, the keyword extraction model described above is further configured to extract a fifth keyword and a sixth keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text. The fifth keyword indicates the execution time of the to-do item, and the sixth keyword indicates the execution location of the to-do item. Step 230 includes generating a target meeting minutes based on the first keyword, the second keyword, the fifth keyword, and the sixth keyword.

[0079] The relevant content of the above embodiments can be referred to here, and will not be repeated here.

[0080] Figure 6An exemplary flowchart of a training method for a keyword extraction model according to an embodiment of the present disclosure is shown.

[0081] like Figure 6 As shown, according to an embodiment of this disclosure, a training method 600 for a keyword extraction model is provided, comprising: acquiring multiple historical meeting texts, wherein each historical meeting text indicates a request for a historical executor to perform a historical task, each historical meeting text including a seventh keyword and an eighth keyword, wherein the seventh keyword indicates the historical executor and the eighth keyword indicates the historical task (step 610); for each historical text data, determining the semantic structure of the historical text data (step 621); and performing semantic annotation on the seventh keyword and the eighth keyword in the historical text data (step 622); and training an initial model based on the annotated multiple historical text data and the corresponding semantic structure of each historical text data to obtain a keyword extraction model (step 630).

[0082] In some embodiments, based on the keyword extraction model trained as described above, keywords for tasks and keywords for the executor of the task can be identified from the meeting text, thereby enabling the accurate generation of meeting minutes for the task.

[0083] Figure 7 A structural block diagram of an apparatus for generating meeting minutes according to an embodiment of the present disclosure is shown.

[0084] like Figure 7 As shown, an apparatus for generating meeting minutes is provided according to an embodiment of the present disclosure, comprising: a first acquisition unit 710 configured to acquire a first meeting audio, wherein the first meeting audio includes information instructing a requester to perform a pending task; a second acquisition unit 720 configured to acquire a second meeting audio based on the first meeting audio, wherein the second meeting audio is associated with a requester and includes information instructing the requester whether they have committed to performing the pending task; and a generation unit 730 configured to generate a target meeting minutes based on the first meeting audio in response to determining, based on the second meeting audio, that the requester has committed to performing the pending task, wherein the target meeting minutes instruct the requester to perform the pending task.

[0085] Here, the operation of each of the above-mentioned units 710 to 630 of the device 700 for generating meeting minutes is similar to the operation of steps 210 to 230 described above, and will not be repeated here.

[0086] Figure 8 A structural block diagram of a training method for a keyword extraction model according to an embodiment of the present disclosure is shown.

[0087] like Figure 8As shown, an embodiment of this disclosure provides a training apparatus for a keyword extraction model, comprising: a third acquisition unit 810 configured to acquire multiple historical meeting texts, wherein each historical meeting text indicates a request for a historical executor to perform a historical pending task, each historical meeting text including a seventh keyword and an eighth keyword, wherein the seventh keyword indicates the historical executor and the eighth keyword indicates the historical pending task; a processing unit 820 configured to determine the semantic structure of each historical text data; and to perform semantic annotation on the seventh keyword and the eighth keyword in the historical text data; and a training unit 830 configured to train an initial model based on the annotated multiple historical text data and the corresponding semantic structure of each historical text data to obtain a keyword extraction model.

[0088] Here, the operation of each of the above units 810 to 830 of the keyword extraction model training device 800 is similar to the operation of steps 610 to 630 described above, and will not be repeated here.

[0089] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0090] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0091] refer to Figure 9 The present invention describes a structural block diagram of an electronic device 900 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0092] like Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded into a random access memory (RAM) 903 from a storage unit 908. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0093] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, output unit 907, storage unit 908, and communication unit 909. Input unit 906 can be any type of device capable of inputting information to electronic device 900. Input unit 906 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 907 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 908 can include, but is not limited to, hard disk and optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0094] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning network algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as method 999. For example, in some embodiments, methods 200 and 600 can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of method 200 or method 600 described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform method 200 or method 600 by any other suitable means (e.g., by means of firmware).

[0095] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0096] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0097] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0099] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0100] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0101] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0102] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A method for generating meeting minutes, comprising: Acquiring first conference audio, wherein the first conference audio includes information instructing the requester to perform the pending tasks, wherein acquiring the first conference audio includes: Obtain the audio from the third meeting; Speech recognition is performed on the third conference audio to obtain the first conference text; The first meeting text is used to extract keywords using a target keyword extraction model, wherein the keyword extraction model is configured to identify at least a first keyword and a second keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, and wherein the first keyword indicates the executor and the second keyword indicates the to-do item; and In response to the identification of the first keyword and the second keyword from the meeting text, the third meeting audio is used as the first meeting audio; Based on the executor's meeting account and / or voiceprint information, obtain a second meeting audio within a preset time period following the speaking time of the first meeting audio, wherein the second meeting audio is associated with the executor and includes information indicating whether the executor has committed to performing the to-do item; and In response to determining, based on the second meeting audio, that the executor has committed to performing the to-do item, a target meeting minutes are generated based on the first meeting audio, wherein the target meeting minutes instruct the executor to perform the to-do item.

2. The method according to claim 1, wherein, The meeting is an online meeting, and obtaining the second meeting audio based on the first meeting audio includes: Obtain the executor's meeting account used to log in to the online meeting based on the first keyword; and The fourth meeting audio associated with the meeting account is obtained within a first time period after the speaking time of the first meeting audio and used as the second meeting audio.

3. The method according to claim 2, wherein, The keyword extraction model is further configured to extract a third keyword and a fourth keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, wherein the third keyword indicates the execution time of the to-do item and the fourth keyword indicates the execution location of the to-do item; Furthermore, the step of generating the target meeting minutes based on the first meeting audio includes: The target meeting minutes are generated based on at least the first keyword, the second keyword, the third keyword, and the fourth keyword.

4. The method according to claim 1, wherein, The meeting can be online or offline. Obtaining the second meeting audio based on the first meeting audio includes: Obtain the voiceprint information of the executor based on the first keyword; and The fifth conference audio, associated with the voiceprint information, is obtained during a second time period following the speaking time of the first conference audio and used as the second conference audio.

5. The method according to claim 4, wherein, The keyword extraction model is further configured to extract a fifth keyword and a sixth keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, wherein the fifth keyword indicates the execution time of the to-do item and the sixth keyword indicates the execution location of the to-do item; Furthermore, the step of generating the target meeting minutes based on the first meeting audio includes: The target meeting minutes are generated based on at least the first keyword, the second keyword, the fifth keyword, and the sixth keyword.

6. A training method for a keyword extraction model, comprising: Obtain multiple historical meeting texts, wherein each historical meeting text indicates a request for a historical executor to perform a historical to-do item, each historical meeting text includes a seventh keyword and an eighth keyword, wherein the seventh keyword indicates the historical executor and the eighth keyword indicates the historical to-do item; For each historical text data, Determine the semantic structure of the historical text data, wherein the semantic structure indicates the execution association between the seventh keyword as the historical executor and the eighth keyword as the historical pending item; and Perform semantic annotation on the seventh and eighth keywords in the historical text data; and An initial model is trained based on the labeled historical text data and the corresponding semantic structure of each historical text data to obtain the keyword extraction model. The trained keyword extraction model is used to implement the method described in any one of claims 1-5.

7. An apparatus for generating meeting minutes, comprising: A first acquisition unit is configured to acquire first conference audio, wherein the first conference audio includes information instructing a requester to perform a pending task, wherein the first acquisition unit includes: The first acquisition subunit is configured to acquire the audio from the third conference. The recognition subunit is configured to perform speech recognition on the third conference audio to obtain the first conference text; An extraction subunit is configured to extract keywords from the first meeting text using a target keyword extraction model, wherein the keyword extraction model is configured to identify at least a first keyword and a second keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, and wherein the first keyword indicates the executor and the second keyword indicates the to-do item; and The processing subunit is configured to, in response to recognizing the first keyword and the second keyword from the conference text, use the third conference audio as the first conference audio; The second acquisition unit is configured to acquire, based on the executor's meeting account and / or voiceprint information, a second meeting audio within a preset time period following the speaking time of the first meeting audio, wherein the second meeting audio is associated with the executor and includes information indicating whether the executor has committed to performing the to-do item; and The generation unit is configured to generate a target meeting minutes based on the first meeting audio in response to determining, based on the second meeting audio, that the executor has committed to performing the to-do item, wherein the target meeting minutes indicate that the executor shall perform the to-do item.

8. The apparatus according to claim 7, wherein, The meeting is an online meeting, and the second acquisition unit includes: The second acquisition subunit is configured to acquire the executor's meeting account used to log in to the online meeting based on the first keyword; and The third acquisition subunit is configured to acquire a fourth conference audio associated with the conference account within a first time period after the speaking time of the first conference audio as the second conference audio.

9. The apparatus according to claim 8, wherein, The keyword extraction model is further configured to extract a third keyword and a fourth keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, wherein the third keyword indicates the execution time of the to-do item and the fourth keyword indicates the execution location of the to-do item; Furthermore, the generation unit includes: The first generation subunit is configured to generate the target meeting minutes based at least on the first keyword, the second keyword, the third keyword, and the fourth keyword.

10. The apparatus according to claim 7, wherein, The meeting can be an online meeting or an offline meeting, and the second acquisition unit includes: The fourth acquisition subunit is configured to acquire the voiceprint information of the executor based on the first keyword; and The fifth acquisition subunit is configured to acquire, within a second time period following the speaking time of the first conference audio, a fifth conference audio associated with the voiceprint information as the second conference audio.

11. The apparatus according to claim 10, wherein, The keyword extraction model is further configured to extract a fifth keyword and a sixth keyword based on the semantic structure of the meeting text and the meaning of each word in the meeting text, wherein the fifth keyword indicates the execution time of the to-do item and the sixth keyword indicates the execution location of the to-do item; Furthermore, the step of generating the target meeting minutes based on the first meeting audio includes: The target meeting minutes are generated based on at least the first keyword, the second keyword, the fifth keyword, and the sixth keyword.

12. A training device for a keyword extraction model, comprising: The third acquisition unit is configured to acquire multiple historical meeting texts, wherein each historical meeting text indicates a request for a historical executor to perform a historical to-do item, each historical meeting text includes a seventh keyword and an eighth keyword, and wherein the seventh keyword indicates the historical executor and the eighth keyword indicates the historical to-do item; The processing unit is configured to process each piece of historical text data. Determine the semantic structure of the historical text data, wherein the semantic structure indicates the execution association between the seventh keyword as the historical executor and the eighth keyword as the historical pending item; and Perform semantic annotation on the seventh and eighth keywords in the historical text data; and The training unit is configured to train an initial model based on the labeled historical text data and the corresponding semantic structure of each historical text data to obtain the keyword extraction model. The trained keyword extraction model is used to implement the method described in any one of claims 1-5.

13. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.