Conference record generation method and device, server, and storage medium
By updating meeting records in real time within the video conferencing system and utilizing a combination of audio captions and text cache data, the problem of incomplete semantics in video conferencing records is solved, enabling the real-time generation of high-quality meeting records and meeting users' needs for real-time performance and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MIGU CO LTD
- Filing Date
- 2022-12-29
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, the segmentation processing of video conference recordings cannot accurately reflect the speaker's content, resulting in incomplete semantics and making it impossible to generate high-quality meeting records in real time.
By obtaining the current audio captions from the client speaking at the meeting, and combining them with the target text cache data to update the local meeting data on the server, and saving the updated record data as the final record data when a sentence end marker is detected, the semantic integrity of the meeting record is ensured.
It enables real-time generation of semantically complete meeting minutes, supports mid-session download and viewing of historical records, meets users' needs in latency-sensitive scenarios, and improves the accuracy and real-time performance of meeting minutes.
Smart Images

Figure CN116013306B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, server, and storage medium for generating meeting minutes. Background Technology
[0002] In related technologies, real-time speech recognition technology is commonly used in video conferencing to convert the audio streams of participants into text in real time, thereby generating subtitles and meeting minutes. The subtitles are then transmitted to the client for display. However, meeting minutes, as a relatively formal document, cannot be formatted as loosely as subtitles; they need to be segmented according to the speaker's context.
[0003] However, in the existing scheme, the processing of interrupted sentences in meeting minutes cannot accurately reflect the content of the speaker's speech, thus failing to guarantee semantic integrity and preventing the generation of meeting minutes in real time.
[0004] Application content
[0005] The main purpose of this application is to provide a method, apparatus, server, and storage medium for generating meeting minutes, aiming to solve the technical problem that meeting minutes cannot be generated in real time.
[0006] To achieve the above objectives, this application provides a method for generating meeting minutes, used on a server in a meeting system. The method includes:
[0007] Obtain the current audio captions from the currently speaking client in the meeting; where the current audio captions are obtained by converting the short audio data of the currently speaking client into text;
[0008] Determine the target text cache data corresponding to the currently speaking client;
[0009] Based on the current audio captions and target text cache data, update the first current meeting data on the local server to obtain the updated first current meeting data;
[0010] The updated first current meeting data is sent to the client so that the client can update the current record data in the meeting minutes document based on the first current meeting data;
[0011] If the current audio subtitle has a sentence-end marker, a preset marker is sent to all clients so that the clients save the updated current recording data as the final recording data in the meeting minutes document and generate new current recording data according to preset rules.
[0012] In one possible embodiment of this application, before updating the first current meeting data locally on the server based on the current audio captions and target text cache data, the method further includes:
[0013] Extract the result type identifier of the current speech caption;
[0014] If the result type is identified as a short sentence end marker and does not have a full sentence end marker, then the target text cache data is updated based on the current audio subtitles.
[0015] In one possible embodiment of this application, after extracting the result type identifier of the current speech caption, the method further includes:
[0016] If the result type is identified by a short sentence end marker and also has a full sentence end marker, then clear the target text cache data.
[0017] In one possible embodiment of this application, after extracting the result type identifier of the current speech caption, the method further includes:
[0018] Determine whether the target text cache data is blank;
[0019] If the data is blank, then determine whether the current audio subtitle has the sentence end marker.
[0020] In one possible embodiment of this application, updating the first current meeting data locally on the server based on the current audio captions and target text cache data to obtain the updated first current meeting data includes:
[0021] If the current audio caption does not have a sentence-end marker, then based on the current audio caption and the target text cache data, update the first current meeting data on the server to obtain the current meeting intermediate data;
[0022] If the result type is identified as having a sentence end marker, then based on the current audio captions and target text cache data, the first current meeting data on the local server is updated to obtain the finalized meeting data, which carries a preset marker.
[0023] In one possible embodiment of this application, after determining the target text cache data corresponding to the currently speaking client, the method further includes:
[0024] Determine whether all text cache data except the target text cache data are blank.
[0025] If not all data are blank, then the client corresponding to at least one other text cache data that is not blank will be the interrupted client.
[0026] Add a short sentence end marker and a preset interruption marker to the most recent voice subtitles of the interrupted client, update other text cache data that are not blank, and obtain the interruption text cache data;
[0027] Extract at least one timestamp of interrupted text cache data;
[0028] Based on the timestamp, obtain the sorting result of at least one interrupted text cache data;
[0029] Based on the sorting results, at least one interrupted text cache data is sent to the client in sequence, so that the client updates the current record data in the meeting minutes document based on the interrupted text cache data, saves the updated current record data as the final record data in the meeting minutes document, and generates new current record data according to preset rules.
[0030] In one possible embodiment of this application, after determining whether all text cache data other than the target text cache data are blank data, the method further includes:
[0031] If all data is blank, then based on the current audio captions and target text cache data, update the first current meeting data on the server to obtain the updated first current meeting data.
[0032] Secondly, this application also provides a meeting record generation device, configured on a server of a meeting system, the device comprising:
[0033] The text acquisition module is used to acquire the current audio captions of the currently speaking client in the conference; the current text information is obtained by converting the audio data of the currently speaking client into text.
[0034] The cache determination module is used to determine the target text cache data corresponding to the currently speaking client;
[0035] The meeting data update module is used to update the first current meeting data on the local server based on the current audio captions and target text cache data, and obtain the updated first current meeting data.
[0036] The record sending module is used to send the updated first current meeting data to all clients, so that all clients can update their local meeting record documents based on the first current meeting data;
[0037] The recording and saving module is used to send a preset identifier to all clients if the current audio subtitle has a sentence end identifier, so that the clients save the updated current recording data as the final recording data in the meeting minutes document, and generate new current recording data according to preset rules.
[0038] Thirdly, this application also provides a server, including: a processor, a memory, and a conference record generation program stored in the memory, wherein the conference record generation program is executed by the processor to implement the steps of the above-described conference record generation method.
[0039] Fourthly, this application also provides a computer-readable storage medium storing a meeting record generation program, which, when executed by a processor, implements the meeting record generation method described above.
[0040] This application proposes a meeting record generation method for use in a server within a meeting system. The method includes: obtaining the current audio captions of the currently speaking client; determining the target text cache data corresponding to the currently speaking client; updating the first current meeting data locally on the server based on the current audio captions and the target text cache data to obtain the updated first current meeting data; sending the updated first current meeting data to the client so that the client updates the current record data in the meeting record document based on the first current meeting data; if the current audio captions have a sentence end marker, sending a preset marker to all clients so that the clients save the updated current record data as the final record data in the meeting record document and generate new current record data according to preset rules.
[0041] Therefore, in this embodiment, the server is configured with a cache area for each client. The cache area stores text cache data, including historical audio short phrase subtitles up to the current moment, i.e., short phrases generated by the user during continuous speaking, i.e., incomplete sentences. When a new current audio subtitle is received, the text corresponding to the subtitle and the incomplete sentence are sent to the client together to update the meeting record. Thus, the formation of subtitles is real-time during the meeting, and this application records based on subtitles, so the meeting record is also updated in real-time. If the newly received text has a sentence-end marker, the client is prompted to save the updated current record data as the final record data in the meeting record document, and a new current record data is generated according to preset rules, i.e., a new sentence is started, thereby ensuring the semantic integrity of the meeting record generated by this application. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the architecture of an embodiment of the conference system of this application;
[0043] Figure 2 This is a schematic diagram of the server structure of the hardware operating environment involved in the embodiments of this application.
[0044] Figure 3 This is a flowchart illustrating the first embodiment of the meeting minutes generation method of this application;
[0045] Figure 4 This is a flowchart illustrating the second embodiment of the meeting minutes generation method of this application;
[0046] Figure 5This is a flowchart illustrating the third embodiment of the meeting minutes generation method of this application.
[0047] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0048] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0049] In related technologies, with the development of artificial intelligence, especially speech recognition, not only has the accuracy of speech-to-text conversion made great progress, but it can also process the generated text by segmenting sentences and adding punctuation based on the speaker's semantics and context. However, segmenting sentences based on complete semantics is difficult to apply in real-time scenarios, especially in video conferencing. Users in video conferencing or live streaming scenarios are very sensitive to latency, so subtitles cannot have significant delays. Existing solutions generally generate meeting minutes in real time based on received subtitles or process and generate meeting minutes after the meeting ends.
[0050] However, meeting minutes, as a relatively formal form of meeting record, require neat formatting, accurate recording of participants' speeches, and the ability to reflect the speakers' meanings. Existing solutions, while meeting the requirement of real-time processing, fail to accurately reflect the content of the speakers' speech due to the handling of sentence breaks. It takes time for a person to finish speaking; if the entire sentence is spoken before being transcribed, real-time processing is compromised. Forcibly splitting the speech into two or more sentences in the middle results in meeting minutes that do not fully reflect the speakers' meanings.
[0051] Furthermore, for scenarios involving post-meeting processing of meeting minutes, meeting minutes can only be generated after the meeting ends. Current solutions cannot meet the needs of participants who want to download meeting minutes after leaving mid-meeting, those who want to view previous meeting minutes, or those who want to review past meeting minutes if they didn't understand a speaker's statement.
[0052] To address this, this application provides a solution that records text based on real-time updated subtitles. The meeting minutes are also updated in real time. If newly received text contains a sentence-end marker, the subtitle cache is cleared, and the client is prompted to save the updated current record data as the final record data in the meeting minutes document. New current record data is generated according to preset rules, thereby ensuring the semantic integrity of the meeting minutes generated by this application.
[0053] The inventive concept of this application is further illustrated below with reference to some specific embodiments.
[0054] The following embodiments of this application will describe the conference system used in the technical implementation of this application:
[0055] Reference Figure 1 , Figure 1 This is a schematic diagram of the architecture of a conference system provided in an exemplary embodiment. For example... Figure 1 As shown, the conference system may include a server 11, a network 12, a client 13, and a speech recognition server 15.
[0056] Server 11 can be a physical server containing a single host, or it can be a virtual server hosted in a host cluster. During operation, server 11 can run server-side programs for online conferencing applications to implement the relevant business functions of the application, such as creating new video conferences and sending the speaking data of the currently speaking client to other clients.
[0057] Network 12 may include various types of wired or wireless networks. In one embodiment, network 12 may include the Public Switched Telephone Network (PSTN) and the Internet. Client 13 can interact with server 11 through network 12, and voice recognition server 15 can interact with server 11 through network 12.
[0058] Client 13 may include electronic devices such as smartphones, tablets, laptops, and PDAs (Personal Digital Assumptions), etc., and one or more embodiments in this specification are not intended to limit this. During operation, client 13 may run video conferencing client-side programs to implement the relevant business functions of the application.
[0059] The speech recognition server 15 can acquire the speech collected by the client and convert it into text using speech recognition technology. It is worth mentioning that with the development of speech recognition technology, not only has the accuracy of speech-to-text conversion made great progress, but it can also process the generated text by segmenting sentences and adding punctuation based on the speaker's semantics and contextual information.
[0060] It is understandable that, in some embodiments, the speech recognition server may also be configured as a module in the server or in the client.
[0061] Reference Figure 2 , Figure 2 This is a schematic diagram of the server structure of the hardware operating environment involved in the embodiments of this application.
[0062] like Figure 2As shown, the server may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless Fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0063] Those skilled in the art will understand that Figure 2 The structure shown does not constitute a limitation on the server and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0064] like Figure 2 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a key function configuration program.
[0065] exist Figure 2 In the electromechanical equipment shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the server of this application can be set in the server. The server calls the meeting record generation program stored in the memory 1005 through the processor 1001 and executes the meeting record generation method provided in the embodiment of this application.
[0066] Based on, but not limited to, the above hardware structure, this application provides a first embodiment of a meeting record generation method. (Refer to...) Figure 3 , Figure 3 A flowchart illustrating a first embodiment of the method for generating application meeting minutes is shown.
[0067] It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0068] In this embodiment, a method for generating meeting minutes includes:
[0069] Step S101: Obtain the current audio captions of the currently speaking client in the conference; wherein, the current audio captions are obtained by converting the short audio data of the currently speaking client into text.
[0070] Specifically, in this embodiment, each speaker acquires audio through a corresponding client, transmitting it as an audio stream in real time to the speech recognition server for transcription. The speech recognition server then transcribes the received multiple audio streams into corresponding subtitles in real time and transmits them to the conference system's server. Thus, the server obtains the current audio subtitles from the currently speaking client in real time.
[0071] It's worth noting that the speech recognition server determines in real-time whether the current subtitle is an intermediate result (a subtitle before a sentence is finished) or a final result (a short sentence). The final result, or the corrected subtitle, is based on the audio content and semantics. Simultaneously, punctuation is added to the transcribed text according to the semantics. Therefore, as the speaker continues to utter a complete sentence, many intermediate results are transmitted in real-time to ensure real-time performance. These intermediate results are then fully displayed in the subtitles on various clients, thus meeting the real-time requirements for subtitle processing.
[0072] In one example, the current audio captions are sent to the server by the speech recognition server in the form of caption messages. In this case, the content of the caption message includes, but is not limited to: (1) the speaker's captions; (2) the caption timestamp; and (3) whether it is an intermediate result.
[0073] In some embodiments, whether a result is an intermediate result can be indicated by the result type identifier resultType: resultType = 0 indicates that the preceding audio subtitle is an intermediate result, and resultType = 1 indicates that it is the final result. That is, resultType = 1 can be considered as a sentence end identifier.
[0074] In one example, if the current audio caption is "We go first", then resultType = 0; if the current audio caption is "The weather is nice today", then resultType = 1.
[0075] Understandably, a speaker's sentence may be quite long, with occasional pauses. When translating this into subtitles, commas and other symbols are added. Ultimately, a complete sentence contains one or more final results, meaning it includes one or more short sentences. Therefore, the presentation of meeting minutes cannot simply mirror the subtitle presentation. Thus, in this embodiment, subtitles and meeting minutes are recorded separately. It's worth noting that, to meet the real-time requirements of subtitles, the current audio subtitles can be used in the subtitle field displayed on the client side. The client subtitles simply display the original subtitle content transcribed by the speech recognition component in real time.
[0076] Step S102: Determine the target text cache data corresponding to the currently speaking client.
[0077] The server has multiple cache areas, each corresponding to a client. The cache areas store text cache data, which includes historical audio clips and subtitles up to the current moment.
[0078] Understandably, in the conference system provided in this embodiment, each participant has a dedicated subtitle cache, which is also known as text cache data. This text cache data can be used to record previous historical audio phrases as subtitles. That is, when a speaker's sentence is relatively long, for example, if the speaker pauses slightly while speaking, commas or other symbols will be added when converting it to subtitles. Ultimately, a complete sentence will contain one or more final results, that is, it will include one or more short phrases. At this time, these short phrases can be stored in the cache area to obtain the text cache data.
[0079] Step S103: Based on the current audio captions and target text cache data, update the first current meeting data on the local server to obtain the updated first current meeting data.
[0080] Specifically, the server updates the first current meeting data locally based on the latest received current audio captions and target text cache data.
[0081] Understandably, when the target text cache data is blank, the current audio captions will be used directly as the text in the current meeting data.
[0082] Alternatively, if the target text cache data contains multiple short sentences, the current audio subtitles are appended to the last short sentence of the target text cache data to form the first current meeting data.
[0083] In one example, the current audio caption is "Buy snacks.", and the target text cache data stores multiple short phrases such as "The weather is so nice today," and "Let's go to the supermarket first.", and the timestamp of "The weather is so nice today," is earlier than that of "Let's go to the supermarket first.", then "Buy snacks." is concatenated after "Let's go to the supermarket first," resulting in the first current meeting data being "The weather is so nice today, let's go to the supermarket first, buy snacks.".
[0084] Step S104: Send the updated first current meeting data to the client so that the client can update the current record data in the meeting record document based on the first current meeting data.
[0085] After obtaining the first current meeting data, the server sends the first current meeting data to all clients, so that all clients update the current record data in the meeting record document based on the first current meeting data.
[0086] The client-side meeting minutes can be pre-set with templates, which include, but are not limited to, finalized minutes and current minutes. Finalized minutes consist of completed and semantically intact statements, or statements that have been interrupted. Current minutes are updated in real time and change as the subtitles change.
[0087] In one example, the speaker says, "The weather is so nice today, let's go to the supermarket to buy snacks." During this process, the current recorded data is continuously updated from the initial "today" until it becomes: "The weather is so nice today, let's go to the supermarket to buy snacks." Of course, it's understandable that the subtitle at this point is "buy snacks." It's also understandable that this sentence contains a period, i.e., a sentence-ending marker; at this point, the current recorded data will be converted to final recorded data and saved.
[0088] Step S105: If the current audio subtitle has a sentence-end marker, send a preset marker to all clients so that the clients save the updated current recording data as the final recording data in the meeting minutes document and generate new current recording data according to preset rules.
[0089] Specifically, sentence-ending markers include, but are not limited to, punctuation marks such as periods, question marks, or exclamation marks added to the subtitles by the speech recognition server based on speech recognition technology. It is worth noting that sentence-ending markers can be in Chinese or English; there are no restrictions here.
[0090] In other words, in this embodiment, if the newly received text contains a sentence-ending marker, the client is prompted to save the updated current record data as the final record data in the meeting minutes document, thereby ensuring the semantic integrity of the meeting minutes generated by this application.
[0091] Furthermore, to facilitate subsequent recording, this embodiment will prompt the client to generate new current record data according to preset rules. These preset rules can be configured in advance by the server or the client.
[0092] In one example, if the default rule is that adjacent semantically complete sentences need to be represented in segments, the client will save the current recorded data as the final recorded data in the meeting minutes document, which needs to be broken into lines so that the new updated current recorded data can be displayed on a new line.
[0093] For example, if a speaker's continuous sentence is: "The weather is so nice today, let's go to the supermarket first and buy some snacks. After we've bought everything, let's take a car together to go camping in the countryside!", when the current audio subtitle is "buy snacks.", the client will save "The weather is so nice today, let's go to the supermarket first and buy some snacks." as the final recorded data in the meeting minutes document. Then, it will start a new line and display "snacks" - "after we've bought everything," - "after we've bought everything, let's take a car together" - "after we've bought everything, let's take a car together to go camping in the countryside!" before starting another new line.
[0094] Understandably, in some examples, the default rule could be that the final draft record data is not highlighted, while the current record data is highlighted. Alternatively, the default rule could be that the final draft record data is in black font, while the current record data is in green font, and so on.
[0095] It is worth mentioning that the current recorded data in this embodiment can be blank data.
[0096] Alternatively, as an alternative to this embodiment, the steps of updating the first current meeting data locally on the server based on the current audio captions and target text cache data to obtain the updated first current meeting data include:
[0097] (1) If the current audio subtitle does not have a sentence end marker, then based on the current audio subtitle and the target text cache data, update the first current meeting data on the local server to obtain the current meeting intermediate data;
[0098] (2) If the result type is identified as having a sentence end identifier, then based on the current audio subtitles and target text cache data, update the first current meeting data on the local server to obtain the finalized meeting data, which carries a preset identifier.
[0099] The preset marker can be a special marker carried in the message sent by the server to the client. That is, in this embodiment, the current intermediate meeting data and the finalized meeting data can be distinguished by different marker information. When the client receives the current intermediate meeting data, it continues to update the current record data in real time. When the finalized meeting data is obtained, the client immediately saves the updated current record data as the finalized record data in the meeting record document after updating the current record data, and generates new current record data according to preset rules, thereby ensuring the semantic integrity of the meeting records generated in this application.
[0100] Therefore, in this embodiment, the server is configured with a cache area for each client. The cache area stores text cache data, including historical audio short phrase subtitles up to the current moment, i.e., short phrases generated by the user during continuous speaking, which are incomplete sentences. When a new current audio subtitle is received, the text corresponding to the subtitle and the incomplete sentence are sent to the client to update the meeting record. Thus, the formation of subtitles is real-time during the meeting, and this embodiment records the meeting record based on the subtitles, so the meeting record is also updated in real time. If the newly received text has a sentence-ending marker, the client is prompted to save the updated current record data as the final record data in the meeting record document, and a new current record data is generated according to preset rules, i.e., a new sentence is started, thereby ensuring the semantic integrity of the meeting record generated in this embodiment.
[0101] In addition, in this embodiment, the meeting minutes are stored on the client side, and the client side can also query historical meeting minutes.
[0102] Based on the above embodiments, a second embodiment of the meeting record generation method of this application is provided. (Refer to...) Figure 4 , Figure 4 A flowchart illustrating a second embodiment of the method for generating application meeting minutes is shown.
[0103] It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0104] In this embodiment, a method for generating meeting minutes includes:
[0105] Step S201: Obtain the current audio captions of the currently speaking client in the conference; wherein, the current audio captions are obtained by converting the short audio data of the currently speaking client into text.
[0106] Step S202: Determine the target text cache data corresponding to the currently speaking client.
[0107] Step S203: Extract the result type identifier of the current audio subtitle;
[0108] Step S204: If the result type identifier is a short sentence end identifier and does not have a full sentence end identifier, then update the target text cache data based on the current audio subtitles.
[0109] Step S205: If the result type identifier is a short sentence end identifier and has a whole sentence end identifier, then clear the target text cache data.
[0110] Step S206: Based on the current audio captions and target text cache data, update the first current meeting data on the local server to obtain the updated first current meeting data.
[0111] Step S207: Send the updated first current meeting data to the client so that the client can update the current record data in the meeting record document based on the first current meeting data.
[0112] Step S208: If the current audio subtitle has a sentence-end marker, send a preset marker to all clients so that the clients save the updated current recording data as the final recording data in the meeting minutes document and generate new current recording data according to preset rules.
[0113] Specifically, in this embodiment, the current audio subtitles are sent to the server by the speech recognition server in the form of a subtitle message. The content of the subtitle message includes, but is not limited to: (1) the speaker's subtitles; (2) the subtitle timestamp; and (3) whether it is an intermediate result.
[0114] Whether a result is intermediate can be indicated by the result type identifier `resultType`: `resultType = 0` indicates that the current audio subtitle is an intermediate result, and `resultType = 1` indicates the final result. In other words, `resultType = 1` can be considered a sentence end identifier. For example, in one example, if the current audio subtitle is "We go first," then `resultType = 0`; if the current audio subtitle is "The weather is really nice today," then `resultType = 1`.
[0115] At this point, if the current audio subtitle is an intermediate result, it will not be cached.
[0116] When the current audio subtitle is the final result, i.e. a short sentence, and there is no punctuation, i.e. it is not the end of a semantically complete sentence, the previously cached final result for the user, i.e. the target text cache data, and the current final result, i.e. the current audio subtitle, are superimposed and updated in the corresponding cache area of the client to obtain the updated target text cache data.
[0117] In this embodiment, the target text cache data stores only short sentences to avoid generating a large amount of invalid data during real-time updates to the cache based on the subtitles, especially intermediate results. Furthermore, it is understandable that storing only short sentences in the target text cache data also improves the accuracy of meeting record generation.
[0118] When the original subtitle is the final result and has been segmented into sentences (i.e., has a sentence-ending marker), the target text cache data is cleared. For example, in one instance, if the speaker's continuous sentence is: "The weather is so nice today, let's go to the supermarket to buy snacks.", when the current audio subtitle is "Today," then `resultType` = 0. It is not cached until the current audio subtitle becomes "The weather is so nice today,". At this point, `resultType` = 1, but there is no sentence-ending marker like a period, so the target text cache data is updated to "The weather is so nice today,". Of course, when the current audio subtitle is "Buy snacks.", although `resultType` = 1, it has a sentence-ending marker like a period, so the cache area is cleared. The target text cache data becomes a blank document.
[0119] Therefore, in this embodiment, the cache area is cleared after each semantically complete sentence is generated. The target text cache data is a blank document. This ensures that the amount of data in the cache area is small, avoiding the impact of a large amount of historical data on the generation of subsequent meeting minutes, thereby improving the accuracy of meeting minute generation.
[0120] That is, in this embodiment, the historical audio short phrase subtitles are short phrases that do not include those distributed between two semantically complete sentences.
[0121] As an example, before step S208, the method further includes:
[0122] (1) Determine whether the target text cache data is blank;
[0123] (2) If the data is blank, determine whether the current audio subtitle has a sentence end marker.
[0124] Specifically, if the current audio subtitle is the final result and the target text cache data is not blank, then step S204 is executed normally. If the current audio subtitle is the final result and the target text cache data is determined to be blank, then it is directly determined whether the current audio subtitle has a sentence-end marker.
[0125] Based on the above embodiments, a third embodiment of the meeting record generation method of this application is provided. (Refer to...) Figure 5 , Figure 5 A flowchart illustrating a third embodiment of the method for generating application meeting minutes is shown.
[0126] It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0127] In this embodiment, the method includes:
[0128] Step S301: Obtain the current audio captions of the currently speaking client in the conference; wherein, the current audio captions are obtained by converting the short audio data of the currently speaking client into text.
[0129] Step S302: Determine the target text cache data corresponding to the currently speaking client.
[0130] Step S303: Determine whether all text cache data other than the target text cache data are blank data.
[0131] Step S304: If not all data are blank, then the client corresponding to at least one other text cache data that is not blank is selected as the interrupted client.
[0132] Step S305: Add a short sentence end marker and a preset interruption marker to the most recent voice subtitle of the interrupted client, update other text cache data that are not blank, and obtain the interruption text cache data.
[0133] Step S306: Extract timestamps from other text cache data that are not blank.
[0134] Step S307: Based on the timestamp, obtain the sorting results of other text cache data that are not blank;
[0135] Step S308: According to the sorting result, send the other text cache data that is not blank to the client in sequence, so that the client updates the current record data in the meeting minutes document based on the other text cache data that is not blank, saves the updated current record data as the final record data in the meeting minutes document, and generates new current record data according to preset rules.
[0136] Step S309: If all data is blank, then update the first current meeting data on the server based on the current audio captions and target text cache data, and obtain the updated first current meeting data.
[0137] Step S310: Send the updated first current meeting data to the client so that the client can update the current record data in the meeting record document based on the first current meeting data.
[0138] Step S311: If the current audio subtitle has a sentence end marker, send a preset marker to all clients so that the clients save the updated current recording data as the final recording data in the meeting minutes document and generate new current recording data according to preset rules.
[0139] Specifically, in this embodiment, the current audio captions are sent to the server by the speech recognition server in the form of caption messages. At this time, the content of the caption message includes, but is not limited to: (1) the speaker's captions; (2) the caption timestamp; (3) whether it is an intermediate result; (4) the speaker's ID; (5) the speaker's name; and (6) the meeting ID.
[0140] At this point, the server retrieves the IDs of all participants in the current meeting, excluding the speaking client, based on the speaker's ID or name in the current audio captions. Then, it determines the other text cache data besides the target text cache data. Since the corresponding cache area is cleared each time a semantically complete sentence is generated, if all other text cache data besides the target text cache data are blank, it indicates that no one else has spoken recently, meaning the current speaker has not interrupted anyone, and the process continues to steps 309 to 311. For the specific implementation of steps 309 to 311, please refer to steps S103 to S105 or steps S206 to S207 in the above embodiment; they will not be repeated here.
[0141] If the current speaker interrupts the previous speaker, the speech recognition server does not add a sentence-end marker to the previous speaker's audio subtitles, and consequently, the corresponding cache area for that speaker is not cleared. Therefore, at this point, all other cached text data besides the target text is not empty, confirming that a speech interruption has occurred.
[0142] Since the previous speaker's speech was incomplete—meaning the most recent audio caption closest in time to the previous speaker's speech is an intermediate result, not the final result—the server's cache area corresponding to the previous speaker does not yet contain this most recent audio caption. Therefore, in this embodiment, after determining that a speech interruption has occurred, a short sentence end marker and a preset interruption marker are added to the previous speaker's most recent audio caption. This caches the most recent audio caption in the cache area, updating other non-blank text cache data to obtain the interrupted text cache data. It is understandable that the preset interruption marker is not only used to reflect the actual situation of a speech interruption in the meeting minutes, but also, during the meeting minutes generation process, can be considered, or equivalent to, a sentence end marker, indicating the end of the previous speaker's speech. Therefore, the server also needs to clear the interrupted text cache data after the update.
[0143] After extracting the interrupted text cache data, since the current audio subtitles belong to the current speaker, not the previous speaker, the meeting minutes can be generated solely based on the extracted interrupted text cache data of the previous speaker, without needing to combine it with the current audio subtitles. The specific process for generating the meeting minutes involves sending the interrupted text cache data to all clients. All clients update the current record data in their meeting minutes documents and save it as the final record data. At this point, the current record data may contain text corresponding to the previous speaker; this text is then completed and converted into the final record data.
[0144] It is worth mentioning that during this process, the current speaker continues to speak. Therefore, while the server is generating the meeting minutes of the previous speaker, it is also performing operations such as caching the text of the current speaker's client. Since the text caching of the current speaker's client depends on the current speaker's speaking speed, the generation of the meeting minutes of the previous speaker and the text caching of the current speaker's client do not affect each other.
[0145] In one example, when the previous speaker says "The weather is so nice today, let's go first," they are interrupted by the current speaker's "Let's go play badminton!" At this point, the server has cached the text data "The weather is so nice today," in the cache area corresponding to the previous speaker, but "Let's go first," as an intermediate result, is not cached. The server then retrieves the current audio caption "Let's go play badminton!" and determines that the previous speaker's cached text data is not blank, thus confirming an interruption. The server then retrieves the most recent audio caption "Let's go first" from the caption server, message cache, or the client's caption cache, adds a short sentence end marker and a preset interruption marker to it, updating the previous speaker's cached text data to "The weather is so nice today, let's go first (interrupted)." The server then generates the finalized meeting data based on "The weather is so nice today, let's go first (interrupted)" and sends it to the client. The client updates "The weather is so nice today," in the current recorded data to "The weather is so nice today, let's go first (interrupted)" and saves it as the finalized record data. Then a newline is added.
[0146] Of course, in actual meetings, there may be consecutive interruptions, meaning there might be two or more clients interrupted. These interrupted clients might be speaking sequentially. In this case, the text cache data of all interrupted clients is updated to obtain all the interrupted text cache data. Then, the interrupted text cache data is sorted according to its timestamp. It's worth noting that the timestamp of the interrupted text cache data could be the generation timestamp of the subtitle corresponding to the first short sentence of each interrupted text cache data, which more accurately reflects the speaking order. After sorting, multiple interrupted text cache data are sent to the clients sequentially to generate the final draft record data.
[0147] In this embodiment, the meeting minutes can also perfectly reflect complex contexts such as one person speaking and one or more people interrupting, ensuring the accuracy of the meeting minutes.
[0148] Based on the same inventive concept, in a second aspect, this application also provides a meeting record generation device, configured on a server of a meeting system, the device comprising:
[0149] The text acquisition module is used to acquire the current audio captions of the currently speaking client in the conference; the current text information is obtained by converting the audio data of the currently speaking client into text.
[0150] The cache determination module is used to determine the target text cache data corresponding to the currently speaking client;
[0151] The meeting data update module is used to update the first current meeting data on the local server based on the current audio captions and target text cache data, and obtain the updated first current meeting data.
[0152] The record sending module is used to send the updated first current meeting data to all clients, so that all clients can update their local meeting record documents based on the first current meeting data;
[0153] The recording and saving module is used to send a preset identifier to all clients if the current audio subtitle has a sentence end identifier, so that the clients save the updated current recording data as the final recording data in the meeting minutes document, and generate new current recording data according to preset rules.
[0154] It should be noted that the various implementations of the meeting record generation device in this embodiment and the technical effects they achieve can be referred to the various implementations of the meeting record generation method in the foregoing embodiments, and will not be repeated here.
[0155] Furthermore, embodiments of this application also propose a computer storage medium storing a meeting record generation program. When executed by a processor, the meeting record generation program implements the steps of the meeting record generation method described above. Therefore, it will not be repeated here. Additionally, the beneficial effects of using the same method will not be repeated here either. For technical details not disclosed in the computer-readable storage medium embodiments of this application, please refer to the description of the method embodiments of this application. As an example, program instructions can be deployed to execute on a single computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0156] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0157] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, and of course, it can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memory, special components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0159] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for generating meeting minutes, characterized in that, The method, used in a conference system, comprises: Obtain the current audio captions of the currently speaking client in the meeting; wherein the current audio captions are obtained by converting the audio data of the currently speaking client into text; Determine the target text cache data corresponding to the currently speaking client; Based on the current audio captions and the target text cache data, update the first current meeting data locally on the server to obtain the updated first current meeting data; The updated first current meeting data is sent to the client, so that the client updates the current record data in the meeting record document based on the first current meeting data; If the current audio subtitle has a sentence-end marker, a preset marker is sent to the client so that the client saves the updated current recording data as the final recording data in the meeting record document and generates new current recording data according to preset rules; After determining the target text cache data corresponding to the currently speaking client, the method further includes: Determine whether all text cache data other than the target text cache data are blank. If not all of them are blank data, then the client corresponding to at least one of the other text cache data that is not blank data will be the interrupted client. Add a short sentence end marker and a preset interruption marker to the most recent voice subtitles of the interrupted client, update the other text cache data that is not blank, and obtain the interrupted text cache data; The interrupted text cache data is sent to the client so that the client updates the current record data in the meeting minutes document based on the interrupted text cache data, saves the updated current record data as the final record data in the meeting minutes document, and generates new current record data according to preset rules.
2. The meeting record generation method according to claim 1, characterized in that, Before updating the first current meeting data locally on the server based on the current audio captions and the target text cache data, and obtaining the updated first current meeting data, the method further includes: Extract the result type identifier of the current voice caption; If the result type identifier is a short sentence end identifier and does not have a full sentence end identifier, then the target text cache data is updated based on the current audio subtitle.
3. The meeting record generation method according to claim 2, characterized in that, After extracting the result type identifier of the current speech subtitle, the method further includes: If the result type identifier is the short sentence end identifier and has the whole sentence end identifier, then the target text cache data is cleared.
4. The meeting record generation method according to claim 3, characterized in that, After extracting the result type identifier of the current speech subtitle, the method further includes: Determine whether the target text cache data is blank; If the data is blank, then determine whether the current audio subtitle has the sentence end marker.
5. The meeting record generation method according to claim 1, characterized in that, The step of updating the first current meeting data locally on the server based on the current audio captions and the target text cache data to obtain the updated first current meeting data includes: If the current audio subtitle does not have a sentence-end identifier, then based on the current audio subtitle and the target text cache data, update the first current meeting data on the server to obtain the current meeting intermediate data; If the current audio subtitle has a sentence-end identifier, then based on the current audio subtitle and the target text cache data, the first current meeting data on the local server is updated to obtain the finalized meeting data, which carries the preset identifier.
6. The meeting record generation method according to claim 1, characterized in that, After adding a short sentence end marker and a preset interruption marker to the most recent audio subtitles of the interrupted client, updating the other text cache data that is not blank, and obtaining the interrupted text cache data, the method further includes: Extract at least one timestamp of the interrupted text cache data; Based on the timestamp, obtain at least one sorting result of the interrupted text cache data; The step of sending the interrupted text cache data to the client, so that the client updates the current record data in the meeting minutes document based on the interrupted text cache data, saves the updated current record data as the final record data in the meeting minutes document, and generates new current record data according to preset rules, includes: Based on the sorting results, at least one of the interrupted text cache data is sent to the client in sequence, so that the client updates the current record data in the meeting minutes document based on the interrupted text cache data, saves the updated current record data as the final record data in the meeting minutes document, and generates new current record data according to preset rules.
7. The meeting record generation method according to claim 6, characterized in that, After determining whether all text cache data other than the target text cache data are blank, the method further includes: If all data is blank, then the process of updating the first current meeting data on the server based on the current audio captions and the target text cache data is performed to obtain the updated first current meeting data.
8. A meeting minutes generation device, characterized in that, The server configured in the conference system includes: The text acquisition module is used to acquire the current audio captions of the currently speaking client in the conference; wherein the current audio captions are obtained by converting the audio data of the currently speaking client into text; The cache determination module is used to determine the target text cache data corresponding to the currently speaking client; The meeting data update module is used to update the first current meeting data on the server based on the current audio captions and the target text cache data, so as to obtain the updated first current meeting data; The record sending module is used to send the updated first current meeting data to all the clients, so that all the clients update the current record data in the meeting record document based on the first current meeting data; The recording and saving module is used to send a preset identifier to all the clients if the current audio subtitle has a sentence end identifier, so that the clients save the updated current recording data as the final recording data in the meeting record document, and generate new current recording data according to preset rules; The meeting record generation device is further configured to, after determining the target text cache data corresponding to the currently speaking client, determine whether all other text cache data besides the target text cache data are blank data; if not all are blank data, then at least one client corresponding to the other text cache data that is not blank data is designated as the interrupted speaking client, a short sentence end marker and a preset interruption marker are added to the most recent audio subtitle of the interrupted speaking client, the other text cache data that is not blank data is updated, and interrupted text cache data is obtained; the interrupted text cache data is sent to the client, so that the client updates the current record data in the meeting record document based on the interrupted text cache data, saves the updated current record data as the final record data in the meeting record document, and generates new current record data according to preset rules.
9. A server, characterized in that, include: A processor, a memory, and a meeting record generation program stored in the memory, the meeting record generation program being executed by the processor to implement the steps of the meeting record generation method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a meeting record generation program, which, when executed by a processor, implements the meeting record generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Subtitle display method and device, terminal and machine readable storage medium
CN114143591A
Apparatus and Method for Generating Subtitles and Meeting Minutes based on Voice Recognition
KR1020220130490A