Conference processing method and related apparatus

By acquiring meeting audio and automatically generating meeting minutes using a large language model, the problem of low efficiency in manual recording is solved, and efficient and accurate meeting minutes generation is achieved.

WO2026091860A1PCT designated stage Publication Date: 2026-05-07BOE TECHNOLOGY GROUP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BOE TECHNOLOGY GROUP CO LTD
Filing Date
2025-09-03
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The generation of meeting minutes relies on manual recording, which is inefficient and difficult to guarantee in terms of quality, and is prone to information omissions.

Method used

By acquiring meeting audio, a large language model is used to generate meeting minutes. Combined with topic switching and voice-to-text length judgment conditions, meeting minutes are automatically generated and can be viewed and modified in real time.

Benefits of technology

It enables timely generation and high-quality recording of meeting minutes, improving meeting efficiency and the accuracy of the minutes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025118842_07052026_PF_FP_ABST
    Figure CN2025118842_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a conference processing method and a related apparatus, which enable automatic generation of conference minutes, and support generation of conference minutes during conferences, thus facilitating viewing and modification. The conference processing method comprises: acquiring audio of a target conference; on the basis of speech text corresponding to the audio, determining whether a conference minutes generation condition is satisfied; and, if the conference minutes generation condition is satisfied, generating conference minutes of the target conference.
Need to check novelty before this filing date? Find Prior Art

Description

Meeting processing methods and related devices

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411545846.3, filed on October 31, 2024, entitled "Meeting Processing Method and Related Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of electronic technology, and in particular to conference processing methods and related devices. Background Technology

[0004] In meeting settings, meeting minutes are an important record of the meeting. They typically rely on meeting recorders or relevant personnel to take notes during the meeting, which is manpower-dependent and suffers from low efficiency, difficulty in ensuring record quality, and the potential for information omissions. Summary of the Invention

[0005] This application provides a meeting processing method and related apparatus that can automatically generate meeting minutes, support the generation of meeting minutes during the meeting, and facilitate viewing and modification.

[0006] In a first aspect, embodiments of this application provide a meeting processing method, including:

[0007] Obtain the audio of the target meeting;

[0008] Based on the audio text corresponding to the audio, determine whether the conditions for generating meeting minutes are met;

[0009] If the conditions for generating meeting minutes are met, the meeting minutes of the target meeting are generated.

[0010] In practice, during the meeting, the electronic device can determine whether the conditions for generating meeting minutes are met. If the conditions are met, the meeting minutes of the target meeting are generated, enabling timely generation of meeting minutes for users to view and modify, thereby improving meeting efficiency and enhancing the quality of meeting minutes.

[0011] In one possible design, this application embodiment provides a meeting processing method where the meeting minutes generation conditions are one or more of the following:

[0012] A topic shift occurred;

[0013] The number of characters in the audio-text correspondence is greater than or equal to a preset value.

[0014] In practice, the electronic device can determine whether a topic switch has occurred and promptly generate meeting minutes when such a switch happens, allowing participants who have just joined the target meeting to quickly understand the content discussed in the previous meeting. Alternatively, the electronic device can determine whether the number of characters in the audio-text correspondence is greater than or equal to a preset value. The electronic device can utilize a large language model to generate meeting minutes. The preset value can be determined based on the maximum number of characters that the large language model can process in a single session. Generally, if the number of characters input into the large language model exceeds the preset value, the quality of the output result is difficult to guarantee.

[0015] In one possible design, this application embodiment provides a meeting processing method in which the meeting minutes generation conditions include a topic switch;

[0016] The generation of the meeting minutes for the target meeting includes:

[0017] Generate meeting minutes for the target meeting before the topic switch.

[0018] In practice, the electronic device can acquire the audio of the target meeting. Based on the corresponding audio transcript, it determines whether a topic change has occurred. If a topic change is detected, it generates meeting minutes of the target meeting prior to the topic change. The electronic device can determine whether a topic change has occurred based on the audio transcript. During the meeting, the electronic device can promptly generate meeting minutes upon receiving a topic change notification, allowing users to view and modify them, thus improving meeting efficiency and the quality of the meeting minutes.

[0019] In one possible design, the meeting processing method provided in this application embodiment further includes:

[0020] Display the generated meeting minutes.

[0021] In one possible design, embodiments of this application provide a meeting processing method, further comprising:

[0022] Collect images of each speaker;

[0023] Based on the images of each speaker, generate virtual avatars for each speaker.

[0024] In practice, electronic devices can generate virtual images for participants in the target meeting, thereby enhancing the meeting's effectiveness.

[0025] In one possible design, embodiments of this application provide a meeting processing method, further comprising:

[0026] Based on the audio text corresponding to any speaker's audio in the target meeting, generate the meeting highlights for that speaker;

[0027] Based on the key points of the meeting from any of the speakers, generate the audio to be synthesized;

[0028] Based on the virtual image of any speaker and the audio to be synthesized, generate a meeting video corresponding to any speaker.

[0029] In practice, electronic devices can generate virtual images of each speaker and broadcast videos of the key points of their speeches at the meeting, facilitating review of the meeting content and providing a good display effect.

[0030] In one possible design, this application embodiment provides a conference processing method in which the audio-corresponding speech text includes the speech content of each speaker in order of speaking sequence.

[0031] The step of determining whether a topic switch has occurred based on the voice text corresponding to the audio includes:

[0032] According to the sorting, calculate the first similarity between the vector representation of the first speech content and the vector representation of the second speech content, wherein the second speech content includes the first n speech contents of the first speech content, and n is an integer greater than or equal to 1;

[0033] Based on the first similarity and the similarity threshold, it is determined whether a topic switch has occurred.

[0034] In one possible design, this application embodiment provides a meeting processing method in which a topic switch occurs if the first similarity is less than the similarity threshold.

[0035] If the first similarity is greater than or equal to the similarity threshold, then no topic switching has occurred.

[0036] In one possible design, embodiments of this application provide a meeting processing method in which generating meeting minutes of the target meeting before topic switching includes:

[0037] The target speech text is segmented, and the target speech text includes the text before the topic change in the speech text corresponding to the audio, or the target speech text includes the text before the topic change in the speech text corresponding to the audio and the text in the speech text of the previously acquired audio that was not used to generate meeting minutes;

[0038] The segmented paragraphs and prompt word information are input into the large language model, which then generates a summary of each paragraph.

[0039] Based on the summaries of each paragraph, the meeting minutes corresponding to the target topic are generated.

[0040] In practice, the speech text is segmented and processed, and each segment is input into the large language model. This can improve the quality and efficiency of the minutes output by the large language model, which is beneficial to improving the quality of meeting minutes.

[0041] In one possible design, embodiments of this application provide a conference processing method in which the segmentation of the target speech text includes:

[0042] The target speech text is segmented into topics to obtain text on one or more topics;

[0043] If the number of words in the text of any topic is less than or equal to the maximum word count threshold for a paragraph, then the text of that topic is treated as a single paragraph.

[0044] If the number of words in the text of any topic exceeds the maximum word count threshold of the paragraph, then the text of any topic is divided into paragraphs.

[0045] In practice, the audio text is first divided into paragraphs according to the topic, ensuring that content on the same topic is contained within a single paragraph, so that the generated paragraph summaries better reflect that paragraph. If the word count of a topic exceeds the maximum word count threshold for a paragraph, the text is divided into paragraphs again to control the word count of the paragraphs input into the large language model. This ensures the efficiency and quality of the large language model, resulting in high-quality paragraph summaries output by the model, which is beneficial for improving the quality of meeting minutes.

[0046] In one possible design, embodiments of this application provide a meeting processing method, wherein the text of any topic includes the speaking content of each speaker in order of speaking sequence;

[0047] The paragraph division of text on any of the topics includes:

[0048] Create a new first paragraph, and use the speech of the speaker corresponding to the first order in the sorting as the content of the first paragraph;

[0049] For the third statement in any order other than the first order in the sorting:

[0050] Determine the second similarity between the vector representation of the third speech content and the vector representation of the fourth speech content, wherein the fourth speech content is the speech content preceding the third speech content;

[0051] If the segmentation conditions are met, a new second paragraph will be created, and the third speech content will be used as the content of the second paragraph. The segmentation conditions are that the second similarity is less than the similarity threshold, or the number of words in the paragraph containing the fourth speech content is greater than or equal to the maximum number of words in the paragraph.

[0052] If the segmentation conditions are not met, the third speech content will be included in the segment containing the fourth speech content.

[0053] Secondly, embodiments of this application also provide a conference system, including:

[0054] Audio acquisition device, used to acquire audio from the target meeting;

[0055] The processing device is used to determine whether a topic switch has occurred based on the audio text corresponding to the audio, and when a topic switch is determined to have occurred, to generate meeting minutes of the target meeting before the topic switch.

[0056] Display device for displaying the generated meeting minutes.

[0057] Thirdly, embodiments of this application also provide an electronic device, including a memory and a processor, wherein:

[0058] The memory stores computer program instructions;

[0059] The processor executes the computer program instructions to perform the steps of the conference processing method as described in the first aspect and any possible design thereof.

[0060] Fourthly, embodiments of this application also provide a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program that, when the computer program is run on a computer, causes the computer to perform the steps of the conference processing method as described in the first aspect and any possible design thereof.

[0061] Furthermore, the technical effects brought about by the second to fourth aspects can be found in the technical effects brought about by the different designs in the first aspect, and will not be repeated here. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1A is a flowchart illustrating a meeting processing method provided in an embodiment of this application;

[0064] Figure 1B is a schematic flowchart of a meeting processing method provided in an embodiment of this application;

[0065] Figure 2 is a schematic diagram of the process of generating meeting minutes provided in an embodiment of this application;

[0066] Figure 3 is a schematic diagram illustrating the effect of discourse normalization processing according to an exemplary embodiment;

[0067] Figure 4A is a schematic diagram of a paragraph segmentation process provided in an embodiment of this application;

[0068] Figure 4B is a schematic diagram of a paragraph segmentation process provided in an embodiment of this application;

[0069] Figure 4C is a schematic diagram of a topic segmentation process provided in an embodiment of this application;

[0070] Figure 5 is a schematic diagram of a paragraph segmentation process provided in an embodiment of this application;

[0071] Figure 6 is a schematic diagram of a process for generating a conference video corresponding to a speaker, provided in an embodiment of this application.

[0072] Figure 7 is a schematic diagram of a conference system provided in an embodiment of this application;

[0073] Figure 8 is a schematic diagram of the structure of the electronic device provided in the embodiment of this application. Detailed Implementation

[0074] The embodiments of this application involve at least one, including one or more; wherein, multiple means two or more. Furthermore, it should be understood that in the description of this application, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance, nor as indicating or implying order. In the description of this application, "A and / or B" includes three options: including A; including B; and including both A and B.

[0075] As used in the embodiments of this application, the terms "when..." or "after..." can be interpreted, depending on the context, as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, the phrases "when..." or "if (the stated condition or event) is detected" can be interpreted, depending on the context, as meaning "if...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)". Furthermore, in the above embodiments, relational terms such as "first" and "second" are used to distinguish one entity from another, without limiting any actual relationship or order between these entities.

[0076] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0077] Figure 1A is a schematic flowchart of a meeting processing method provided in an embodiment of this application. This method can be executed by an electronic device and may include the following steps:

[0078] S10, acquire the audio of the target meeting.

[0079] S11, Based on the voice text corresponding to the audio, determine whether the conditions for generating meeting minutes are met.

[0080] S12, if the meeting minutes generation conditions are met, generate the meeting minutes of the target meeting.

[0081] In practice, during the meeting, the electronic device can determine whether the conditions for generating meeting minutes are met. If the conditions are met, the meeting minutes of the target meeting are generated, enabling timely generation of meeting minutes for users to view and modify, thereby improving meeting efficiency and enhancing the quality of meeting minutes.

[0082] In some application scenarios, the meeting minutes generation conditions are one or more of the following:

[0083] Condition 1: A topic change occurs;

[0084] Clause 2: The number of characters in the audio-text corresponding to the audio is greater than or equal to a preset value.

[0085] Regarding condition 1, the electronic device can determine whether a topic change has occurred. If a topic change has occurred, condition 1 is satisfied; otherwise, condition 1 is not satisfied. When condition 1 is included in the meeting minutes generation conditions, if a topic change is detected, it indicates that the meeting minutes generation condition is met. The electronic device can determine whether a topic change has occurred and promptly generate meeting minutes when such a change occurs, allowing participants who have just attended the target meeting to quickly understand the content discussed in the previous meeting.

[0086] Regarding condition 2, the electronic device can determine whether the number of characters in the audio-text correspondence is greater than or equal to a preset value. If the number of characters is greater than or equal to the preset value, condition 2 is satisfied; if the number of characters is less than the preset value, condition 2 is not satisfied. When condition 2 is included in the meeting minutes generation conditions, if the electronic device determines that the audio-text correspondence is greater than or equal to the preset value, it indicates that the meeting minutes generation conditions are met. In some possible implementations, the electronic device can use a large language model to generate meeting minutes. The preset value can be determined based on the maximum number of characters that the large language model can process in a single operation. Generally, if the number of characters input into the large language model exceeds the preset value, the quality of the output result is difficult to guarantee.

[0087] In some examples, the conditions for generating meeting minutes may include condition 1. In some examples, the conditions for generating meeting minutes may include condition 2. In other examples, the conditions for generating meeting minutes may include both condition 1 and condition 2. Meeting minutes generation is considered met if either condition 1 or condition 2 is satisfied. Meeting minutes generation is not considered met if neither condition 1 nor condition 2 is satisfied.

[0088] In practical applications, the conditions for generating meeting minutes can also include other conditions depending on the application requirements.

[0089] In one possible implementation, the meeting minutes generation conditions include a topic switch; during the operation of the electronic device generating the meeting minutes of the target meeting, the meeting minutes of the target meeting before the topic switch can be generated.

[0090] In one possible scenario, the electronic device determines whether the conditions for generating meeting minutes are met based on the audio-text correspondence. If the conditions are not met, the electronic device can repeatedly execute steps S10-S12, enabling real-time detection of whether the target meeting meets the conditions for generating meeting minutes. When the conditions are met, the electronic device generates or updates the meeting minutes.

[0091] Figure 1B is a schematic flowchart of a meeting processing method provided in an embodiment of this application. This method can be executed by an electronic device and may include the following steps:

[0092] S101, acquire the audio of the target meeting.

[0093] Electronic devices can acquire audio from a target meeting. For example, an electronic device can have audio capture capabilities to capture the audio of the target meeting. Alternatively, an electronic device can receive audio from a target meeting provided by an audio capture device. The audio from the target meeting can be real-time audio from the target meeting.

[0094] S102, based on the voice text corresponding to the audio, determine whether a topic switch has occurred.

[0095] Electronic devices can perform speech recognition on audio to identify the corresponding spoken text. Based on the spoken text, the electronic device determines whether a topic change has occurred. A topic change can include a shift from one topic to another, or a situation where any spoken text in the audio does not belong to the same topic as the spoken text preceding it.

[0096] S103, when a topic switch is detected, generate the meeting minutes of the target meeting before the topic switch.

[0097] Electronic devices can generate meeting minutes for the target meeting prior to the topic change after detecting or confirming a topic switch, such as generating meeting minutes for one or more previous topics. In some examples, when an electronic device detects that the target meeting has changed from topic one to topic two, it can generate meeting minutes for topic one, or generate all meeting minutes prior to topic two.

[0098] In one possible scenario, the electronic device determines whether a topic switch has occurred based on the audio-text correspondence. If no topic switch is detected, the electronic device can repeatedly execute steps S101-S103, enabling real-time detection of topic switches in the target meeting. Upon detecting a topic switch, it generates meeting minutes for the topics before the switch, or meeting minutes for all topics before the switch.

[0099] In one possible design, after generating meeting minutes, the electronic device can display the generated meeting minutes. The electronic device can display the minutes of the most recent topic before the topic switch. Alternatively, the electronic device can display the minutes of several recent topics before the topic switch.

[0100] Optionally, the electronic device can support the modification or editing of meeting minutes. For example, the electronic device can display the meeting minutes and receive user modifications, editing, or adjustments. This design can improve the quality of the meeting minutes. The timely display of meeting minutes by the electronic device facilitates timely review of the content by participants or recorders, improving review efficiency and enhancing the quality of the meeting minutes.

[0101] Based on any of the above embodiments, the electronic device can have audio processing capabilities. In some examples, the electronic device can preprocess the audio data of the target conference to extract human voices or speech portions from the audio of the target conference. For example, the electronic device can obtain noise-reduced speech data through automatic gain control, microphone array beamforming, and adaptive speech noise reduction, and then extract human voices through speech endpoint detection.

[0102] Electronic devices can perform speech recognition. They can convert human voices or speech portions in audio into text information that computers can recognize. This text information includes the content of the speech and a timestamp. The speech content can also be understood as the spoken words or language content, representing what the speaker expresses through language.

[0103] Optionally, the electronic device can pre-store the voiceprint feature information of one or more users, as well as the identity identifier of each user. The electronic device can extract the voiceprint features of different speakers from the human voice or speech portion and match (or compare) them with the pre-stored voiceprint feature information. The identity identifier corresponding to the matched voiceprint feature information can be used as the speaker's identity identifier. When the electronic device matches voiceprint feature information according to a preset matching algorithm, if the matching probability of voiceprint feature A and voiceprint feature B is greater than a preset matching probability threshold, it can be determined that voiceprint feature A and voiceprint feature B are a match.

[0104] Based on any of the above embodiments, the speech text corresponding to the audio may include the speech content (turns) of speakers in each order of speaking. In some possible scenarios, the same speaker may speak multiple times, for example, in the order of user A, user B, user C, user A. The speech text corresponding to the audio may include the speech content of user A in the first order (which is also the speech content of user A's first speech in the current audio), the speech content of user B in the second order, the speech content of user C in the third order, and the speech content of user A in the fourth order (which is also the speech content of user A's second speech in the current audio).

[0105] When an electronic device determines whether a topic switch has occurred based on the audio-text correspondence, it can calculate a first similarity between the vector representations of the first and second speech contents according to the aforementioned order. The second speech content may include the first n speech contents of the first speech content, where n is an integer greater than or equal to 1. The electronic device can determine whether a topic switch has occurred based on the first similarity and a similarity threshold. If the first similarity is less than the similarity threshold, a topic switch has occurred; if the first similarity is greater than or equal to the similarity threshold, no topic switch has occurred.

[0106] In some examples, electronic devices can input the speech content (turn) of each speaker in the audio-text speech to a pre-trained language model (Masked and Permutation Language Model, MPNet) and obtain the embedding vector representation of each turn as the vector representation of each turn.

[0107] For any given turn, the electronic device can calculate the cosine similarity between that turn and its predecessor. If the cosine similarity is less than a similarity threshold, a topic switch occurs. This turn can be considered the first statement in a new topic, and the previous turn can be considered the last statement in a topic preceding the new one. If the cosine similarity is greater than or equal to the similarity threshold, it can be determined that no topic switch has occurred, indicating that this turn and its predecessor belong to the same discussion topic.

[0108] Based on any of the above embodiments, the electronic device can store the acquired audio of the target meeting, perform speech recognition on the audio, and obtain the corresponding speech text. The electronic device can then generate meeting minutes based on the speech text corresponding to the audio.

[0109] In some examples, given the audio transcripts of a target meeting already acquired by the electronic device, the device can use the transcripts of the preceding topic to generate meeting minutes for that topic. In other examples, given the audio transcripts of a target meeting already acquired by the electronic device, the device can use the transcripts of part or all of the meeting's audio transcripts to generate meeting minutes. Specifically, the process of generating meeting minutes using transcripts is the same or similar. The following example illustrates how an electronic device can generate meeting minutes for a preceding topic (referred to as the target topic) using transcripts of the preceding topic.

[0110] Figure 2 illustrates a flowchart for generating meeting minutes. Generating meeting minutes using electronic devices may include some or all of the following operations:

[0111] S201, segment the target speech text.

[0112] The target speech text includes the text in the speech text corresponding to the audio before the topic switch, or the target speech text includes the text in the speech text corresponding to the audio before the topic switch and the text in the speech text of the previously acquired audio that was not used to generate meeting minutes.

[0113] In some applications, the target speech text can be the speech text corresponding to all audio from any meeting. How can we generate the complete meeting minutes of that meeting in a single operation?

[0114] S202, input the segmented paragraphs and prompt word information into the large language model, so that the large language model generates a summary of each paragraph.

[0115] S203, Based on the summaries of each paragraph, generate the meeting minutes corresponding to the target topic.

[0116] Optionally, before step S201, or after the electronic device determines the speech text of the audio, it performs text preprocessing on the speech text, such as discourse normalization, to reduce the number of modal particles in the speech text. The electronic device can use existing text processing methods to reduce the number of modal particles in the speech text. Alternatively, the electronic device can utilize a large language model to reduce the number of modal particles in the speech text. In some examples, for part or all of the speech text of any speaker, and with pre-configured text preprocessing prompt information, the electronic device can invoke a large language model to perform discourse normalization on part or all of the speaker's speech.

[0117] For example, a text preprocessing prompt could be: "You are an editor focused on improving the clarity and professionalism of an article. Your main task is to remove unnecessary interjections from a conference paper, without adding any content based on subjective speculation, maintaining the objectivity, factuality, and completeness of the article, and ensuring that the article remains coherent and fluent with a clear logical structure after removing interjections. The conference content is as follows: \n' + original conference text." Here, the "original conference text" in the preprocessing prompt specifically refers to part or all of a speaker's audio text. As shown in Figure 3, the text without boxes represents part or all of a speaker's audio text, the text (or punctuation) with horizontal lines represents the removed interjections and punctuation after standardization, and the text (or punctuation) with boxes represents newly added text (or punctuation) in the audio text.

[0118] In some applications, the target speech text can be speech text that has undergone discourse normalization. In such a design, electronic devices can generate meeting minutes based on the speech text from the audio, which can improve the quality and efficiency of the meeting minutes. In other applications, the target speech text can be speech text that has not undergone discourse normalization.

[0119] In step S201, the electronic device segments the voice text of the target topic, which can be done in one or more ways.

[0120] In some implementations, the electronic device can segment the speech text of the target topic according to a maximum word count threshold, such that the number of words in each segment is less than or equal to the maximum word count threshold.

[0121] In other embodiments, the electronic device may be segmented in a manner shown in FIG4A, which may include the following steps:

[0122] S401, perform topic segmentation on the target speech text to obtain text on one or more topics.

[0123] S402, determine whether the number of words in the text of any topic is less than the maximum word count threshold of the paragraph. If yes, proceed to step S403; otherwise, proceed to step S404.

[0124] The initial segmentation process, performed by topic, can be understood as the first segmentation step. The electronic device then determines whether secondary segmentation is needed based on the word count of the topic's text. Specifically, the electronic device compares the word count of the topic's text with a maximum paragraph word count threshold. If the threshold is exceeded, secondary segmentation is performed on the topic's text; otherwise, it is not. In some applications, the maximum paragraph word count threshold is configurable. Optionally, the maximum paragraph word count threshold can be set based on the computing resources of the electronic device.

[0125] S403, treat the text of any of the topics as a paragraph.

[0126] If the number of words in the text of any topic is less than or equal to the maximum word count threshold for a paragraph, then the text of that topic is treated as a single paragraph.

[0127] S404, divide the text of any of the topics into paragraphs.

[0128] If the number of words in any topic text is greater than the maximum word count threshold of the paragraph, then the topic text is divided into paragraphs.

[0129] The text for any given topic may include the speech content (turns) of speakers in each order of speaking. The electronic device can divide the text for any given topic into paragraphs, creating a new first paragraph and using the speech content of the first speaker in the order as the content of the first paragraph.

[0130] For the third statement in any order other than the first order in the sorting: perform the following operations: determine the second similarity between the vector representation of the third statement and the vector representation of the fourth statement, where the fourth statement is the statement preceding the third statement. Then determine if a segmentation condition is met; if the segmentation condition is met, a new second paragraph is created, and the third statement is included in the second paragraph. The segmentation condition can be that the second similarity is less than a similarity threshold, or that the number of characters in the paragraph containing the fourth statement is greater than or equal to the maximum character count threshold of the paragraph. If the segmentation condition is not met, the third statement is included in the paragraph containing the fourth statement.

[0131] Regarding step S401, the electronic device performs topic segmentation on the target speech text to obtain text on one or more topics.

[0132] In one possible design, the electronic device can utilize a pre-trained network model to perform topic segmentation on spoken text. For example, the electronic device can use a pooling network (PoNET) model to model a text sequence, identify, and segment parts with consistent topics.

[0133] The PoNet network uses a pooling mechanism to replace the self-attention mechanism in the traditional Transformer model, making it more efficient when processing long sequences. The PoNet network mainly includes three pooling modules with different granularities: a global pooling module (generally called the GA module), a segmented max-pooling module (generally called the SMP module), and a local max-pooling module (generally called the LMP module). These modules can capture sequence information at different granularities, thus more accurately understanding the topic structure of the text.

[0134] The following are the specific roles and implementation methods of each module in text topic segmentation:

[0135] The Global Pooling module captures global information, that is, the overall features of the entire text sequence. By performing a pooling operation on the entire text sequence, the Global Pooling module extracts global feature representations. These features can reflect the overall theme or sentiment of the text.

[0136] The Segment Max-Pooling (SMP) module is designed to capture local information from different segments of a text. It segments the speech text according to certain rules (such as sentences or paragraphs) and performs max pooling on each segment. In this way, the Segment Max-Pooling module can extract the main features of each segment, which helps the model understand thematic changes in different parts of the text.

[0137] The Local Max-Pooling (LMP) module focuses on local details within the text, such as keywords or phrases. LMP performs max pooling operations on a local scale within a sequence (e.g., a sentence or a few words). This helps the model capture key information points in the text, further enhancing its understanding of the text's topic.

[0138] The global pooling module, SMP module, and LMP module each output feature maps. The feature maps from these three modules are summed to obtain the first data. This first data is then normalized to obtain the second data. The second data is then input into a feedforward neural network for processing to obtain the third data. The third and second data are then normalized together to obtain the fourth data. The fourth data undergoes further linear transformation through a Dense layer (linear activation function + tanh activation function) to obtain the fifth data. The fifth data is then input into an Output Projection layer (dropout + linear), which outputs the category ID corresponding to the text sequence. A text sequence with a category ID of 1 indicates that the text sequence is the ending sentence of the topic. A text sequence with a category ID of 0 indicates that the text sequence is not the ending sentence of the topic.

[0139] Electronic devices can perform topic segmentation based on the category IDs corresponding to each text sequence in the target speech text.

[0140] In another possible design, since large language models have strong semantic understanding capabilities, as shown in Figure 4B, electronic devices can use large language models for topic segmentation.

[0141] The electronic device can perform initial segmentation of the target speech text, with the total number of characters in each segment exceeding a coarse segmentation character threshold. Optionally, the electronic device can divide the target speech text into n equal segments, with the total number of characters in each segment exceeding a coarse segmentation character threshold.

[0142] After the electronic device performs initial segmentation on the target speech text, it obtains multiple segments. The electronic device can then traverse each segment according to their order in the target text, from the first segment to the last. For any segment i, segment i is input into a large language model, which performs topic segmentation on segment i.

[0143] In some cases, the initial segmentation process may result in content on the same topic being scattered across multiple segments. After the large language model performs topic segmentation on segment i, the last small portion of text in segment i may share a single topic. The electronic device can merge the text of the last topic in segment i with the text of segment i+1. The electronic device then uses the large language model to perform subject segmentation on the merged segment.

[0144] For example, Figure 4C shows the sentences and their numbers in a segment. The large language model performs topic segmentation on the segment and can output the sentence numbers at the end of each topic.

[0145] For each topic's text, the electronic device can determine whether the number of words in the text exceeds the maximum paragraph word count threshold. If the number of words in the text exceeds the maximum paragraph word count threshold, the text is divided into multiple paragraphs within the topic, with each paragraph containing fewer than or equal to the maximum paragraph word count threshold. If the number of words in the text is less than the maximum paragraph word count threshold, the text is treated as a single paragraph.

[0146] In one possible implementation, Figure 5 illustrates a process for segmenting text on a given topic into paragraphs, which may include the following steps:

[0147] S501, create a new first paragraph, and use the speech content of the first order in the sorting as the content of the first paragraph.

[0148] The speech content i represents the speech content in the i-th order of the sort, and t represents the total number of all speech contents in the sort. Traverse from 2 to t and execute the operations in steps S502 to S508.

[0149] S502, calculate the similarity s between the vector representation of speech content i and the vector representation of speech content i-1.

[0150] Electronic devices can predetermine the embedding of each speech as the vector representation of each speech content.

[0151] S503, determine whether the segmentation condition is met. If yes, proceed to step S504; otherwise, proceed to step S505.

[0152] The segmentation condition can be that the similarity s is less than the similarity threshold, or that the number of characters in the paragraph containing the speech content i-1 is greater than or equal to the maximum number of characters in the paragraph threshold.

[0153] S504 will then create a new second paragraph, and use the speech content i as the content of the second paragraph.

[0154] If the segmentation condition is met, then proceed to step S504.

[0155] S505, merge the speech content i into the paragraph containing the speech content i-1.

[0156] S506, Is i equal to t? If yes, end; otherwise, proceed to step S507.

[0157] S507, i = i + 1.

[0158] After step S507, the electronic device may then perform the operation in step S502.

[0159] After the electronic device segments the speech text of the target topic, it executes the operations in steps S202 and S203, that is, it inputs the segmented paragraphs and prompt word information into the large language model, so that the large language model generates a summary of each paragraph. Based on the summary of each paragraph, it generates the meeting minutes corresponding to the target topic.

[0160] In one possible implementation, after segmenting the audio text of the target topic, each segment and the prompt words for generating a segment summary can be input into a large language model to generate a segment summary.

[0161] For example, the prompt message for generating paragraph minutes could be: "You are a meeting secretary, focused on organizing the meeting proceedings. Your main task is to extract the meeting discussion process from a meeting document, without adding any subjective speculation, maintaining the objectivity, factuality, and completeness of the proceedings, ensuring coherence and fluency, and a clear logical structure. The meeting content is as follows: \n' + meeting transcript." The "meeting transcript" in the prompt message could be implemented as a paragraph of text.

[0162] For example, the prompt message for generating paragraph summaries could be: "You are a meeting secretary, focused on summarizing meeting tasks. Your main task is to identify segments containing task-related semantics from a meeting document and summarize the tasks without adding any subjective assumptions, maintaining the objectivity, factuality, and completeness of the tasks. The meeting content is as follows: \n' + meeting transcript." In this prompt message, "meeting transcript" can be implemented as a paragraph of text.

[0163] In this embodiment, the prompt words for generating paragraph summaries are used for illustrative purposes and can be adjusted according to actual needs.

[0164] In the process of generating meeting minutes corresponding to the target topic based on the summaries of each paragraph, the electronic device can generate meeting minutes according to a preset template and the summaries of each paragraph. For example, the preset template can be "AA:xxx", where "AA" represents the title and "xxx" represents the specific content.

[0165] Alternatively, electronic devices can use preset prompts for generating meeting minutes and paragraph summaries for each section, inputting them into a large language model, which then outputs the meeting minutes. For example, the prompts for generating meeting minutes could be: "This is a meeting summary extracted from a meeting record. Please organize and merge it and output it in the format 'AA:xxx'."

[0166] Based on the meeting processing method provided in any of the above embodiments, the electronic device can visualize the target meeting.

[0167] In some applications, electronic devices can capture images of each speaker. Based on these images, virtual avatars of each speaker are generated. The identities of each speaker in the audio of the target meeting can be pre-identified, and the images captured by the electronic devices are also the images corresponding to those speaker identities. The virtual avatars of the speakers are also the virtual avatars corresponding to those speaker identities.

[0168] In one possible design, the electronic device can generate the key points of the meeting for any speaker based on the voice text corresponding to the audio of any speaker in the target meeting; generate audio to be synthesized based on the key points of the meeting for any speaker; and generate a meeting video corresponding to any speaker based on the virtual image of the speaker and the audio to be synthesized.

[0169] In one possible implementation, during the process of generating a virtual avatar of a speaker, the electronic device can use PaddleHub to create the virtual avatar (also known as a virtual digital human). PaddleHub technology includes three models: First Order Motion, Text to Speech, and Wav2Lip.

[0170] Referring to Figure 6, the electronic device can input a virtual avatar image of the speaker and a video recording of the speaker's face into the First Order Motion model for facial expression transfer, outputting a video of a virtual avatar with expressions more closely resembling or even approaching those of a real person. Optionally, the virtual avatar image of the speaker can be generated based on an image of the speaker, for example, using any existing method for generating virtual avatar images. For instance, an AI drawing tool can output a virtual avatar image by simultaneously inputting a real-life image and an image of the desired style.

[0171] The electronic device can input the speaker's key points from the target meeting into a Text-to-Speech model, converting the text into audio to obtain the speaker's audio to be synthesized. The electronic device then uses a Wav2Lip model to merge the aforementioned virtual avatar video and the speaker's audio to obtain the corresponding meeting video. Optionally, the electronic device can also adjust the virtual avatar's lip movements in the video based on the audio content, making the virtual avatar more lifelike.

[0172] The electronic device can utilize a pre-trained large language model to perform deep understanding of the input data, extracting key information, topics, and decisions from the meeting. In generating key points for any speaker based on their audio transcript within the target meeting, after each speaker finishes speaking, the device calculates the embedding similarity between their current speech and the speech of the previous k speakers. Based on this embedding similarity, it determines whether a topic change has occurred. If a topic change is discussed, the content of the previous topic is fed into the large language model for key point extraction (the key points include relevant speaker viewpoints).

[0173] In some application scenarios, electronic devices can use a pre-generated virtual avatar of a meeting assistant and a recorded content broadcast with the anchor's face as input into the First Order Motion model to perform facial expression transfer and output a video of the virtual avatar of the meeting assistant.

[0174] The electronic device can input the text of the generated meeting minutes into a Text to Speech model to convert the text into audio, resulting in the audio to be synthesized for the meeting assistant. The electronic device then uses a Wav2Lip model to merge the video of the virtual avatar of the meeting assistant with the audio to be synthesized, resulting in a video of the virtual avatar reading the meeting minutes. Optionally, the electronic device can also adjust the lip movements of the virtual avatar in the video based on the audio content, making the virtual avatar more lifelike.

[0175] Electronic devices can respond to the action of broadcasting meeting minutes by displaying a video of the meeting minutes broadcast by a virtual avatar of a meeting assistant.

[0176] In some application scenarios, electronic devices can have query functions. In some examples, an electronic device can respond to a query interaction and output the query results corresponding to the query content. In some examples, an electronic device can respond to a first query operation in which the first part of a meeting minutes video is played, and find the audio of the target meeting corresponding to that first part. In some examples, an electronic device can respond to a second query operation in which the first part of a meeting minutes video is played, and find the speech text of the target meeting audio corresponding to that first part. For example, an electronic device can use a large language model to find the speech text of the target meeting audio corresponding to the first part of the meeting minutes, and find the audio segment of the target meeting corresponding to the first part based on the timestamp of the speech text of the target meeting audio.

[0177] Based on the meeting processing method provided in any of the above embodiments, this application provides a meeting system that may include the following modules, as shown in Figure 7.

[0178] The audio preprocessing module can preprocess the raw audio data. It obtains the noise-reduced speech data through automatic gain control, microphone array beamforming, and adaptive speech noise reduction. Then, it extracts human voice through speech endpoint detection and transmits it to the speech recognition module and the voiceprint recognition module.

[0179] The recognition module converts the voice information extracted by the audio preprocessing module into text information that a computer can recognize. This text information may include timestamps for each sentence. It also extracts voiceprint features from the voice output by the audio preprocessing module for different users and compares these features with those stored in the user's voice feature database. This comparison determines the probability of a match between the current voiceprint and the voiceprints in the user's voice feature database (which stores user identification information and voiceprint features), thus identifying the speaker.

[0180] The detection module can be used to determine whether a topic change has occurred, and if a topic change has occurred, it instructs the key point extraction module to input data.

[0181] The key point extraction module can use a trained large language model to deeply understand the input data and extract key information, topics, decisions, etc. from the meeting.

[0182] The meeting minutes generation module can automatically generate structured meeting minutes based on information extracted from a large language model. These minutes can include one or more of the following: meeting topic, time, location, participants, discussion content, decision results, meeting progress, and to-do items. It also supports user-defined minutes format templates.

[0183] Optionally, the conference system may also include an image acquisition module, a virtual human module, and a conference content interaction module.

[0184] The image acquisition module can capture images of the attendees.

[0185] The virtual avatar module allows each participant and meeting host to create a corresponding virtual avatar and generate meeting videos featuring the virtual avatars of each speaker.

[0186] The interactive module can generate dynamic notes. It can transmit the speech-to-text results from the target meeting to a large language model in real time, extracting key points and generating dynamic notes. The prompts used by the large language model for key point extraction can be phrased as "Please analyze the meeting text in real time and extract key information." The interactive module allows participants to edit the dynamic meeting notes in real time. During the meeting, all participants can edit and add notes in real time via mobile devices or computers, ensuring the completeness and timeliness of the meeting content.

[0187] The interactive module can have content synchronization and display functions, which can synchronize the dynamic notes edited by users in real time and the meeting minutes automatically generated by the system to the virtual avatar display interface, ensuring that all participants, including latecomers, can quickly understand the meeting content and progress.

[0188] The interactive module can include a knowledge-based question-and-answer function. During the target meeting, the large language model can intelligently recommend relevant materials, documents, or minutes from previous meetings based on the meeting content, helping participants better understand the current discussion topics. Optionally, the large language model can associate the current meeting content with a pre-set knowledge base to extract relevant technical terms and background information. Users can ask questions to the large language model through the interactive module and receive corresponding answers, improving meeting efficiency.

[0189] The interaction module allows users to interact with the virtual avatar of the meeting assistant, as well as engage in Q&A functionality. It supports gesture, voice, and text interaction. After the meeting concludes, the virtual avatar of the meeting assistant can read aloud the meeting minutes. The reading includes key points extracted from the meeting and the corresponding speakers.

[0190] Through the interactive module, users can ask questions about the content in the meeting minutes and view the original audio or the transcribed text of the original audio. For example, a large language model can be used to find the original text for each point in the minutes, and the corresponding audio segment can be found based on the timestamp of the original text.

[0191] Based on the same concept, Figure 8 is a schematic diagram of the structure of an electronic device 800 provided in an embodiment of this application. The electronic device 800 can be the electronic device executing the conference processing method described above. As shown in Figure 8, the electronic device 800 may include: one or more processors 801; one or more memories 802; a communication interface 803; and one or more computer programs 804. These devices can be connected via one or more communication buses 805. The one or more computer programs 804 are stored in the memory 802 and configured to be executed by the one or more processors 801. The one or more computer programs 804 include instructions. For example, the instructions can be used to execute relevant steps of the conference processing method in the corresponding embodiments above, or the functions of one or more modules in the conference system. The communication interface 803 is used to realize communication between the master device and other devices (such as control devices); for example, the communication interface can be a transceiver.

[0192] The methods provided in the embodiments of this application above are described from the perspective of an electronic device as the executing entity. To implement the functions of the methods provided in the embodiments of this application above, the electronic device may include hardware structures and / or software modules, implementing the above functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is executed in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.

[0193] In addition, embodiments of this application provide a computer-readable storage medium for storing a computer program that, when run on a computer, causes the computer to perform steps in any of the above-described conference processing methods, or the functions of one or more modules in a conference system.

[0194] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0195] This application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this invention are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)). Where there is no conflict, the solutions of the above embodiments can be combined.

Claims

1. A meeting processing method, wherein, include: Obtain the audio of the target meeting; Based on the audio text corresponding to the audio, determine whether the conditions for generating meeting minutes are met; If the conditions for generating meeting minutes are met, the meeting minutes of the target meeting are generated.

2. The method as described in claim 1, wherein, The meeting minutes are generated under one or more of the following conditions: A topic shift occurred; The number of characters in the audio-text correspondence is greater than or equal to a preset value.

3. The method as described in claim 1, wherein, The conditions for generating meeting minutes include a topic switch; The generation of the meeting minutes for the target meeting includes: Generate meeting minutes for the target meeting before the topic switch.

4. The method as described in any one of claims 1-3, wherein, The method further includes: Display the generated meeting minutes.

5. The method of claim 1, wherein, The method further includes: Collect images of each speaker; Based on the images of each speaker, generate virtual avatars for each speaker.

6. The method of claim 5, wherein, The method further includes: Based on the audio text corresponding to any speaker's audio in the target meeting, generate the meeting highlights for that speaker; Based on the key points of the meeting from any of the speakers, generate the audio to be synthesized; Based on the virtual image of any speaker and the audio to be synthesized, generate a meeting video corresponding to any speaker.

7. The method of claim 3, wherein, The audio-text packets correspond to the speaker's speech content in the order of their speech sequence. The step of determining whether a topic switch has occurred based on the voice text corresponding to the audio includes: According to the sorting, calculate the first similarity between the vector representation of the first speech content and the vector representation of the second speech content, wherein the second speech content includes the first n speech contents of the first speech content, and n is an integer greater than or equal to 1; Based on the first similarity and the similarity threshold, it is determined whether a topic switch has occurred.

8. The method of claim 7, wherein, If the first similarity is less than the similarity threshold, a topic switch occurs; If the first similarity is greater than or equal to the similarity threshold, then no topic switching has occurred.

9. The method of claim 3, wherein, The generated meeting minutes of the target meeting before the topic switch include: The target speech text is segmented, and the target speech text includes the text before the topic change in the speech text corresponding to the audio, or the target speech text includes the text before the topic change in the speech text corresponding to the audio and the text in the speech text of the previously acquired audio that was not used to generate meeting minutes; The segmented paragraphs and prompt word information are input into the large language model, which then generates a summary of each paragraph. Based on the summaries of each paragraph, a meeting summary corresponding to the target speech text is generated.

10. The method of claim 9, wherein, The segmentation process for the target speech text includes: The target speech text is segmented into topics to obtain text on one or more topics; If the number of words in the text of any topic is less than or equal to the maximum word count threshold for a paragraph, then the text of that topic is treated as a single paragraph. If the number of words in the text of any topic exceeds the maximum word count threshold of the paragraph, then the text of any topic is divided into paragraphs.

11. The method of claim 10, wherein, The text for any topic includes the speaker's speech content in the order of speaking. The paragraph division of text on any of the topics includes: Create a new first paragraph, and use the speech of the speaker corresponding to the first order in the sorting as the content of the first paragraph; For the third statement in any order other than the first order in the sorting: Determine the second similarity between the vector representation of the third speech content and the vector representation of the fourth speech content, wherein the fourth speech content is the speech content preceding the third speech content; If the segmentation conditions are met, a new second paragraph will be created, and the third speech content will be used as the content of the second paragraph. The segmentation conditions are that the second similarity is less than the similarity threshold, or the number of words in the paragraph containing the fourth speech content is greater than or equal to the maximum number of words in the paragraph. If the segmentation conditions are not met, the third speech content will be included in the segment containing the fourth speech content.

12. A conference system, wherein, include: Audio acquisition device, used to acquire audio from the target meeting; The processing device is used to determine whether a topic switch has occurred based on the audio text corresponding to the audio, and when a topic switch is determined to have occurred, to generate meeting minutes of the target meeting before the topic switch. Display device for displaying the generated meeting minutes.

13. An electronic device, wherein, Includes memory and processor, wherein: The memory stores computer program instructions; The processor executes the computer program instructions to perform the steps in the conference processing method as described in any one of claims 1-11.

14. A computer-readable storage medium, wherein, The computer-readable storage medium is used to store a computer program that, when run on a computer, causes the computer to perform the steps of the conference processing method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Automatic recording method, device and electronic equipment for conference minutes

    CN106802885A

  • Automated assistants with conference capabilities

    CN110741601A

  • Conference summary generation method and device, equipment and storage medium

    CN112836016A

  • Text processing method and device, electronic equipment and computer readable storage medium

    CN114492375A

  • Conference proceed apparatus and method for advancing conference

    US20160086605A1