Conference summary generation method and device, equipment, storage medium and program product

By obtaining conference audio data and combining the affiliation between keywords and conference types, selecting a suitable text recognition model to generate conference minutes, solving the problem of low accuracy in conference minutes generation in traditional technology, and achieving more accurate conference minutes generation.

CN120128674APending Publication Date: 2025-06-10CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510295153.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In traditional technology, the accuracy of generating meeting minutes is low, making it difficult to effectively record the meeting content.

Method used

By obtaining audio data in the target meeting, and determining the meeting type using the affiliation between candidate keywords and meeting types, selecting the corresponding text recognition model for audio data conversion, and generating meeting minutes text.

Benefits of technology

Improve the accuracy of meeting minutes generation, making the generated meeting minutes text more accurate and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128674A_ABST
    Figure CN120128674A_ABST
Patent Text Reader

Abstract

The invention relates to a conference summary generation method and device, equipment, a storage medium and a program product. The method comprises the steps of obtaining target audio data of a target conference in response to a conference recording request for the target conference in the proceeding process of the target conference; determining a target conference type of the target conference according to the acquired target audio data, the candidate keyword and a membership relationship between the candidate keyword and the candidate conference type; determining a target text recognition model according to a corresponding relationship between the candidate conference type and a candidate text recognition model and the target conference type; and inputting the target audio data into the target text recognition model to obtain a target conference summary text corresponding to the target audio data. According to the scheme, the corresponding text recognition model is set for each conference type, so that the text recognition model is more targeted during text recognition, and the target conference summary text corresponding to the obtained target audio data is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, storage medium and program product for generating meeting minutes. Background Art

[0002] With the development of Internet technology and the change of the social work mode, remote video conferencing has become an important tool for communication and collaboration, which can provide users with real-time communication and collaboration support. In order to record the meeting content to promote communication and collaboration, it is necessary to generate meeting minutes.

[0003] In the traditional technology, the way to generate meeting minutes is usually to convert the meeting audio and video data into text data, and then generate meeting minutes according to the text data. There is a problem of low accuracy in generating meeting minutes. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, device, equipment, storage medium and program product for generating meeting minutes, which can improve the accuracy of generating meeting minutes.

[0005] In a first aspect, the present application provides a method for generating meeting minutes, including:

[0006] During the progress of a target meeting, in response to a meeting recording request for the target meeting, obtain target audio data of the target meeting;

[0007] According to the obtained target audio data, candidate keywords, and the membership relationship between the candidate keywords and candidate meeting types, determine the target meeting type of the target meeting;

[0008] According to the correspondence between the candidate meeting types and candidate text recognition models, and the target meeting type, determine a target text recognition model;

[0009] Input the target audio data into the target text recognition model to obtain a target meeting minutes text corresponding to the target audio data.

[0010] In one embodiment, the step of determining the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the membership relationship between the candidate keywords and candidate meeting types includes:

[0011] Determine target keywords in the candidate keywords hit by the target audio data and the number of keywords of the target keywords;

[0012] Determine the number of keywords and the number of keyword categories of the keywords that match each candidate meeting type for the target audio data according to the subordination relationship between the candidate keywords and the candidate meeting types, the target keywords, and the number of keywords of the target keywords;

[0013] Determine the target meeting type of the target meeting according to the number of keywords and the number of keyword categories of the keywords that match each candidate meeting type for the target audio data.

[0014] In one embodiment, the determining the target meeting type of the target meeting according to the number of keywords and the number of keyword categories of the keywords that match each candidate meeting type for the target audio data includes:

[0015] For each candidate meeting type, perform weighted processing on the number of keywords and the number of keyword categories of the keywords that match the candidate meeting type for the target audio data to obtain the matching degree between the target audio data and the candidate meeting type;

[0016] Use the candidate meeting type corresponding to the maximum matching degree among the matching degrees as the target meeting type of the target meeting.

[0017] In one embodiment, the inputting the target audio data into a target text recognition model to obtain the target meeting minutes text corresponding to the target audio data includes:

[0018] Convert the target audio data according to a preset format template to obtain first audio data in a preset format; wherein, the preset format template at least includes a start recording time, a speaker identifier, a speaker speaking time, and a speaking data length;

[0019] Input the first audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

[0020] In one embodiment, each candidate text recognition model is established by the following method:

[0021] For each candidate meeting type, obtain the sample data corresponding to the candidate meeting type; wherein, the sample data includes the sample audio data and the sample meeting minutes text corresponding to the candidate meeting type;

[0022] Convert the sample audio data according to a preset format template to obtain second audio data in a preset format;

[0023] Using the second audio data as training data and the sample meeting minutes text as label data, train the initial text recognition model to obtain the candidate text recognition model corresponding to the candidate meeting type.

[0024] In one embodiment, the target text recognition model includes a text recognition unit and a meeting minutes generation unit; the step of inputting the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data includes:

[0025] Input the target audio data into the text recognition unit for text recognition to obtain the initial text data corresponding to the target audio data;

[0026] Input the initial text data into the meeting minutes generation unit for text parsing to obtain the target meeting minutes text in template format; wherein, the template format includes at least one of meeting date, participants, meeting keywords, chapter summary, to-do items, and meeting Q&A.

[0027] In a second aspect, the present application also provides a meeting minutes generation device, including:

[0028] An acquisition module, configured to acquire the target audio data of the target meeting in response to a meeting recording request for the target meeting during the progress of the target meeting;

[0029] A first determination module, configured to determine the target meeting type of the target meeting according to the acquired target audio data, candidate keywords, and the membership relationship between the candidate keywords and the candidate meeting types;

[0030] A second determination module, configured to determine the target text recognition model according to the correspondence between the candidate meeting types and the candidate text recognition models, and the target meeting type;

[0031] A recognition module, configured to input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

[0032] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0033] During the progress of the target meeting, in response to a meeting recording request for the target meeting, acquire the target audio data of the target meeting;

[0034] Determine the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the subordination relationship between the candidate keywords and the candidate meeting types;

[0035] Determine the target text recognition model according to the correspondence between the candidate meeting types and the candidate text recognition models, and the target meeting type;

[0036] Input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

[0037] Fourthly, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0038] During the progress of the target meeting, in response to a meeting recording request for the target meeting, obtain the target audio data of the target meeting;

[0039] Determine the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the subordination relationship between the candidate keywords and the candidate meeting types;

[0040] Determine the target text recognition model according to the correspondence between the candidate meeting types and the candidate text recognition models, and the target meeting type;

[0041] Input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

[0042] Fifthly, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0043] During the progress of the target meeting, in response to a meeting recording request for the target meeting, obtain the target audio data of the target meeting;

[0044] Determine the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the subordination relationship between the candidate keywords and the candidate meeting types;

[0045] Determine the target text recognition model according to the correspondence between the candidate meeting types and the candidate text recognition models, and the target meeting type;

[0046] Input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

[0047] The above-mentioned meeting minutes generation method, device, equipment, storage medium and program product, during the process of a target meeting, in response to a meeting recording request for the target meeting, obtain the target audio data of the target meeting; determine the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the membership relationship between the candidate keywords and the candidate meeting types; determine the target text recognition model according to the corresponding relationship between the candidate meeting types and the candidate text recognition models, and the target meeting type; input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data. In the above solution, a corresponding text recognition model is set for each meeting type, making the text recognition by the text recognition model more targeted, so that the target meeting minutes text corresponding to the obtained target audio data is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0049] Figure 1 It is a schematic flowchart of the meeting minutes generation method in one embodiment;

[0050] Figure 2 It is a schematic diagram of the distributed rooms of a cloud meeting in one embodiment;

[0051] Figure 3 It is a schematic flowchart of determining the target meeting type of the target meeting in one embodiment;

[0052] Figure 4 It is a schematic flowchart of determining the target meeting type of the target meeting in another embodiment;

[0053] Figure 5 It is a schematic flowchart of generating the target meeting minutes text corresponding to the target audio data in one embodiment;

[0054] Figure 6 It is a schematic flowchart of training the candidate text recognition model in one embodiment;

[0055] Figure 7 It is a schematic flowchart of generating the target meeting minutes text corresponding to the target audio data in another embodiment;

[0056] Figure 8 It is a schematic diagram of the structure of the meeting minutes generation system in one embodiment;

[0057] Figure 9 It is a structural block diagram of a meeting minutes generation device in an embodiment;

[0058] Figure 10 It is an internal structure diagram of a computer device in an embodiment. Specific implementation manners

[0059] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0060] The meeting minutes generation method provided by the embodiments of the present application can be applied to the application scenario of converting the audio data corresponding to a meeting into a meeting minutes text in response to a meeting recording request during a remote meeting. This method can be executed by a server or by a terminal with a certain computing power.

[0061] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, a smart glasses, etc.

[0062] Before introducing the solution, relevant terms will be introduced first:

[0063] Web Real-Time Communication (WebRTC) is a technology that supports real-time media communication (audio and video) and data transmission between web browsers and mobile applications. It allows peer-to-peer audio and video communication and data sharing without relying on an intermediary server.

[0064] The MultiPoint Control Unit (MCU) is specifically used to process the data stream convergence in a multi-party video conference. It can centrally manage multiple video and audio streams, merge them into a single stream, and then transmit it to other viewers.

[0065] Automatic Speech Recognition (ASR) technology converts speech into binary input that can be read by a computer.

[0066] Natural Language Processing (NLP) enables computers and digital devices to recognize, understand, and generate text and speech through computer modeling and deep learning.

[0067] nginx-rtc is a transmission system based on the nginx architecture and adopting the WebRTC protocol.

[0068] In an exemplary embodiment, as Figure 1 shown, a meeting minutes generation method is provided. Taking the application of this method to a terminal as an example, it includes the following steps:

[0069] S101, during the target meeting, in response to a meeting recording request for the target meeting, obtain the target audio data of the target meeting.

[0070] Exemplarily, the target meeting can be a cloud meeting, and the audio and video services of the cloud meeting can be implemented based on the nginx system architecture. For example, in terms of the system architecture, nginx-rtc can be used as the server of the cloud meeting. The server of the cloud meeting is transmitted through a reliable transmission protocol, and the average delay can be reduced by 30%-40%; in terms of the network topology, the server of the cloud meeting makes full use of the content delivery network (CDN) nodes covering the whole country for caching and the highly elastic network topology, and performs static cascade transmission and dynamic cascade transmission respectively for data with different real-time requirements to ensure meeting different data transmission needs; in terms of mechanism construction, the server of the cloud meeting applies the distributed room mode, which enables users to access nearby and avoids congestion of the central server.

[0071] Exemplarily, refer to Figure 2 , Figure 2A schematic diagram of a distributed room for cloud meetings is provided. Participants in each cloud meeting can access the in-group aggregation server nearby according to the distance between the participant and the in-group aggregation server. For example, if Participant 1, Participant 2, Participant 3, and Participant 4 are closer to the server in In-Group Aggregation Machine 1, they can access the room with the server in In-Group Aggregation Machine 1. If Participant 5 is closer to the server in In-Group Aggregation Machine 2, it can access the room with the server in In-Group Aggregation Machine 2. Furthermore, both In-Group Aggregation Machine 1 and In-Group Aggregation Machine 2 are connected to the room in the parent layer machine, and then the room in the parent layer machine accesses the central server of the cloud meeting. Compared with connecting all participants to the central server of the cloud meeting, the pressure on the central server is dispersed, thus avoiding congestion of the central server.

[0072] In addition, nginx-rtc also has extremely strong weak network resistance capabilities. In the case of network fluctuations, it can quickly restore audio and video data transmission through a precise bandwidth detection algorithm, and the weak network recovery time is shortened by 80%. Secondly, by further adopting a scalable video coding and decoding algorithm, the data is split into a multi-layer video stream consisting of a base layer and multiple other optional layers that can improve the resolution and frame rate, reducing the bandwidth requirement and greatly enhancing the error resilience and video quality when devices communicate with each other.

[0073] Moreover, nginx-rtc embeds a retransmission algorithm to deeply ensure the reliable transmission of services in extreme environments, with a video and audio packet loss resistance of up to 50%. In addition, for the retransmission strategy, different from traditional packet loss retransmission algorithms, this module retransmits according to the frame type, preferentially retransmitting the I-frame + audio frame of the video to ensure that key information is not lost and maximizing the stability of the service.

[0074] Exemplarily, when a participant has a need to record a meeting, the participant can trigger a meeting recording instruction to send a meeting recording request for the target meeting to the terminal. This meeting recording request can carry information such as the meeting identifier, recording time range, participant identifier, etc. The terminal can obtain the target audio data of the target meeting through the MCU according to information such as the meeting identifier and participant identifier. Among them, the MCU mainly performs processing such as encoding, decoding, transcoding, mixing, and splicing on the audio and video streams of different participants.

[0075] In addition, the MCU also adopts an intelligent layout mode. Through an intelligent data stream merging algorithm, it can dynamically identify the speaking situation of each participant in the meeting, preferentially maximize the display of the sharer's screen, and automatically adjust the display priority according to the importance and participation degree of the speaker. This dynamic layout based on real-time analysis enables the participants in the meeting to focus more on the core content of the discussion and improves the meeting efficiency.

[0076] The MCU also supports multiple video layout options, allowing users to freely choose according to their needs, further enhancing the flexibility of the meeting. Moreover, to better control the load balancing, the MCU uses a cluster mode and introduces etcd for control and scheduling, aiming to achieve unified scheduling of the MCU. etcd is an open-source, distributed, consistent key-value store mainly used for shared configuration, service discovery, and scheduling coordination in distributed systems or computer clusters.

[0077] S102. Determine the target meeting type of the target meeting based on the obtained target audio data, candidate keywords, and the membership relationship between the candidate keywords and the candidate meeting types.

[0078] Exemplarily, the candidate keywords can be a set of keywords under each candidate meeting type. Each meeting type has corresponding keywords. For example, the meeting type of product R & D has keywords such as "algorithm complexity", "algorithm efficiency improvement", "system architecture", "software architecture design", "code", "performance testing", etc.; the meeting type of product marketing has keywords such as "market size", "market share", "consumer pain points", "target audience", "industry trends", "market forecast", "competitive product analysis", "cost performance", and "channel expansion".

[0079] Based on this, text recognition and keyword extraction can be performed on the obtained target audio data. Furthermore, based on the matching situation between the extracted keywords and the candidate keywords, and the membership relationship between the matched candidate keywords and the candidate meeting types, the target meeting type of the target meeting can be determined. For example, the candidate meeting type with the largest number of successfully matched keywords can be used as the target meeting type of the target meeting.

[0080] S103. Determine the target text recognition model based on the corresponding relationship between the candidate meeting types and the candidate text recognition models, and the target meeting type.

[0081] Exemplarily, the sample data of each candidate meeting type can be obtained in advance. The sample data can include the audio data of the meeting and the corresponding meeting minutes text, and the initial text recognition model can be trained using the sample data to obtain the candidate text recognition model corresponding to each candidate meeting type, and establish the corresponding relationship between each candidate meeting type and the candidate text recognition model.

[0082] Based on this, according to the corresponding relationship between the candidate meeting types and the candidate text recognition models, the candidate meeting types that are the same as the target meeting type can be filtered out, and the candidate text recognition model corresponding to the candidate meeting type can be used as the target text recognition model.

[0083] S104. Input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

[0084] Furthermore, the target audio data can be input into the text recognition model for text recognition and text analysis to obtain the target meeting minutes text corresponding to the target audio data.

[0085] In the above method for generating meeting minutes, during the target meeting, in response to a meeting recording request for the target meeting, obtain the target audio data of the target meeting; determine the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the membership relationship between the candidate keywords and the candidate meeting types; determine the target text recognition model according to the corresponding relationship between the candidate meeting types and the candidate text recognition models, and the target meeting type; input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data. In the above solution, a corresponding text recognition model is set for each meeting type, making the text recognition by the text recognition model more targeted, so that the target meeting minutes text corresponding to the target audio data obtained is more accurate.

[0086] In some optional implementation manners, the target meeting type of the target meeting can be determined according to the number of keywords and the number of categories of the keywords that the target audio data corresponding to the target meeting matches with the candidate keywords.

[0087] Exemplarily, refer to Figure 3 , Figure 3 A flowchart for determining the target meeting type of the target meeting is provided, which specifically includes the following steps:

[0088] S301. Determine the target keywords in the candidate keywords hit by the target audio data and the number of keywords of the target keywords.

[0089] Exemplarily, taking the candidate keywords as keyword A, keyword B, keyword C, keyword D, keyword E, keyword F, keyword G, and keyword H as an example for explanation. Assume that the target keywords in the candidate keywords hit by the target audio data are keyword A, keyword B, keyword C, and keyword D, and the number of keywords of keyword A, keyword B, keyword C, and keyword D hit are 1, 2, 5, and 8 respectively.

[0090] S302. Determine the number of keywords and the number of keyword categories of the keywords that the target audio data matches with each candidate meeting type according to the membership relationship between the candidate keywords and the candidate meeting types, the target keywords, and the number of keywords of the target keywords.

[0091] Exemplarily, taking candidate meeting types 1, 2, and 3 as examples for illustration. Suppose the candidate keywords affiliated with candidate meeting type 1 are keyword A, keyword B, keyword F, and keyword G; the candidate keywords affiliated with candidate meeting type 2 are keyword C, keyword E, and keyword F; and the candidate keywords affiliated with candidate meeting type 3 are keyword D, keyword E, and keyword H.

[0092] Exemplarily, based on the affiliation relationship between the candidate keywords and the candidate meeting types, the target keyword, and the number of keywords of the target keyword, the number of keywords and the number of keyword categories of the keywords that the target audio data matches with each candidate meeting type can be determined. For example, the number of keywords of the keywords that the target audio data matches with candidate meeting type 1 is 3, and the number of keyword categories is 2; the number of keywords of the keywords that the target audio data matches with candidate meeting type 2 is 5, and the number of keyword categories is 1; the number of keywords of the keywords that the target audio data matches with candidate meeting type 3 is 8, and the number of keyword categories is 1.

[0093] S303. Determine the target meeting type of the target meeting according to the number of keywords and the number of keyword categories of the keywords that the target audio data matches with each candidate meeting type.

[0094] Furthermore, the target meeting type of the target meeting can be determined according to the number of keywords and the number of keyword categories of the keywords that the target audio data matches with each candidate meeting type. For example, the candidate meeting type corresponding to the largest number of keywords among the number of keywords of the keywords that the target audio data matches with each candidate meeting type can be used as the target meeting type; or the candidate meeting type corresponding to the largest number of keyword categories among the number of keyword categories of the keywords that the target audio data matches with each candidate meeting type can be used as the target meeting type. It is also possible to use the candidate meeting type corresponding to the largest number of keyword categories among the number of keyword categories of the keywords that the target audio data matches with each candidate meeting type as the alternative meeting type, and then use the candidate meeting type corresponding to the largest number of keywords among the number of keywords of the keywords that the target audio data matches with the alternative meeting type as the target meeting type.

[0095] In the embodiments of the present application, the target meeting type is determined according to the number of keywords and the number of categories of the keywords that the target audio data matches with the candidate keywords, so that the determined target meeting type is more accurate.

[0096] In some alternative implementation manners, refer to Figure 4 , Figure 4 Another schematic flowchart for determining the target meeting type of the target meeting is provided, which specifically includes the following steps:

[0097] S401. For each candidate meeting type, perform weighted processing on the number of keywords and the number of keyword categories of the keywords that match the candidate meeting type in the target audio data to obtain the matching degree between the target audio data and the candidate meeting type.

[0098] Exemplarily, correlation analysis can be performed on historical test data to determine the number of matching keywords and the relationship between the number of matching keyword categories and the matching accuracy of the meeting type, and then determine the weights corresponding to the number of keywords and the number of keyword categories respectively.

[0099] Furthermore, for each candidate meeting type, weighted processing can be performed on the number of keywords and the number of keyword categories of the keywords that match the candidate meeting type in the target audio data to obtain the matching degree between the target audio data and the candidate meeting type.

[0100] Exemplarily, taking the weights corresponding to the number of keywords and the number of keyword categories as 0.4 and 0.6 respectively, the example in the previous embodiment is used for illustration. Suppose the number of keywords of the keywords that match candidate meeting type 1 in the target audio data is 3, and the number of keyword categories is 2, and the matching degree between the target audio data and candidate meeting type 1 is 2.4 (3×0.4 + 2×0.6); the number of keywords of the keywords that match candidate meeting type 2 in the target audio data is 5, and the number of keyword categories is 1, and the matching degree between the target audio data and candidate meeting type 2 is 2.6 (5×0.4 + 1×0.6); the number of keywords of the keywords that match candidate meeting type 3 in the target audio data is 8, and the number of keyword categories is 1, and the matching degree between the target audio data and candidate meeting type 3 is 3.8 (8×0.4 + 1×0.6).

[0101] S402. Use the candidate meeting type corresponding to the maximum matching degree among the matching degrees as the target meeting type of the target meeting.

[0102] Furthermore, the candidate meeting type corresponding to the maximum matching degree among the matching degrees can be used as the target meeting type of the target meeting. For example, candidate meeting type 3 can be used as the target meeting type.

[0103] In the embodiments of the present application, weighted processing is performed on the number of keywords and the number of keyword categories of the keywords that match the candidate meeting type in the target audio data, and the target meeting type is screened according to the obtained matching degree, which can fully consider the influence of the number of keywords and the number of keyword categories of the matching keywords on screening the target meeting type, and thus make the determined target meeting type more reliable.

[0104] In some alternative implementation manners, in order to improve the accuracy of generating the target meeting minutes text, the target audio data can be preprocessed according to a preset format template.

[0105] Exemplarily, refer to Figure 5 , Figure 5 a flowchart showing a process for generating a target meeting minutes text corresponding to target audio data, which specifically includes the following steps:

[0106] S501. Convert the target audio data according to a preset format template to obtain first audio data in the preset format.

[0107] Exemplarily, the target audio data can be converted according to a preset format template. For example, the target audio data can be segmented and labeled according to the start recording time, speaker identifier, speaker speaking time, and speech data length to obtain first audio data in the preset format. Among them, the preset format template at least includes the start recording time, speaker identifier, speaker speaking time, and speech data length. For example, the target audio data can be converted in a segmented manner, and each segment is converted according to the preset format template. For example, the duration corresponding to each segment of audio data can be set according to the actual situation, for example, it can be set to 30 minutes.

[0108] Exemplarily, refer to Table 1. Table 1 provides a data format of first audio data in a preset format, where the first audio data can be audio data labeled according to the preset format, so as to facilitate subsequent text conversion of the audio data according to the annotation information.

[0109] Table 1:

[0110]

[0111] S502. Input the first audio data into a target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

[0112] Furthermore, input the first audio data into a target text recognition model to obtain the target meeting minutes text corresponding to the target audio data. Since information such as the start recording time, speaker identifier, speaker speaking time, and speech data length is labeled in the first audio data, it is convenient for the target text recognition model to accurately parse the first audio data according to the annotation information, thereby improving the accuracy of target meeting minutes text recognition.

[0113] In some alternative implementation manners, each candidate text recognition model can be pre-trained based on sample data corresponding to each candidate meeting type.

[0114] Exemplarily, refer to Figure 6 , Figure 6 a flowchart showing a process for training a candidate text recognition model, which specifically includes the following steps:

[0115] S601. For each candidate meeting type, obtain the sample data corresponding to the candidate meeting type.

[0116] Exemplarily, for each candidate meeting type, an MCU can be used to obtain the sample data corresponding to the candidate meeting type, where the sample data includes the sample audio data and the sample meeting minutes text corresponding to the candidate meeting type.

[0117] S602. Convert the sample audio data according to a preset format template to obtain second audio data in the preset format.

[0118] Furthermore, the sample data can also be labeled to improve the model training efficiency and accuracy. For example, the sample audio data can be converted according to a preset format template to obtain second audio data in the preset format. The preset format template includes at least the start recording time, the speaker identifier, the speaker's speaking time, and the speaking data length. In this way, the second audio data is also labeled according to the preset format template, so as to facilitate the initial text recognition model to extract features and recognize text from the audio data.

[0119] S603. Use the second audio data as the training data and the sample meeting minutes text as the label data to train the initial text recognition model to obtain a candidate text recognition model corresponding to the candidate meeting type.

[0120] Furthermore, the second audio data can be divided into a training set and a test set. Using the sample meeting minutes text as the label data, train the initial text model. When the preset number of training times is reached, or the text recognition accuracy reaches the set threshold, stop training the initial text model to obtain a candidate text recognition model corresponding to the candidate meeting type.

[0121] In the embodiments of the present application, on the one hand, training a corresponding candidate text recognition model for each candidate meeting type can improve the accuracy of the candidate text recognition model in generating the target meeting minutes text corresponding to the candidate meeting type; on the other hand, labeling the sample data according to the preset format template facilitates the initial text recognition model to extract features and recognize text from the sample audio data, improving the model training speed and accuracy.

[0122] In some optional implementation manners, the target text recognition model may include a text recognition unit and a meeting minutes generation unit. Among them, the text recognition unit can be implemented by ASR, and the meeting minutes generation unit can be implemented by NLP.

[0123] Based on this, see Figure 7 , Figure 7 A flowchart of generating the target meeting minutes text corresponding to the target audio data is provided, which specifically includes the following steps:

[0124] S701: Input the target audio data into a text recognition unit for text recognition to obtain initial text data corresponding to the target audio data.

[0125] Exemplarily, the target audio data may be input into a text recognition unit for text recognition, and the text recognition unit converts the audio data into text data to obtain initial text data corresponding to the target audio data.

[0126] S702, inputting the initial text data into the meeting minutes generating unit for text parsing to obtain the target meeting minutes text in a template format.

[0127] Furthermore, the initial text data can be input into the meeting minutes generation unit for text parsing. The meeting minutes generation unit can identify entities, keywords and syntactic structures in the text, analyze and process the identified text, extract key information, such as extracting information such as main topics and decision-making content, and can integrate the extracted information to obtain a meeting minutes text in a template format. The template format includes at least one of the meeting date, participants, meeting keywords, chapter summaries, to-do items and meeting questions and answers. In this way, the target meeting minutes text is generated according to the template format, which improves the readability of the target meeting minutes text.

[0128] In an embodiment of the present application, by setting a text recognition unit and a meeting minutes generation unit to generate target meeting minutes data based on target audio data, the accuracy and efficiency of generating meeting minutes can be improved, and the readability of the target meeting minutes can be improved.

[0129] For example, see Figure 8 , Figure 8 A structural diagram of a meeting minutes generation system is provided, which may include an nginx-rtc module, an MCU confluence module, and a text recognition module. During the target meeting, the client initiates a meeting recording request and transmits the request to the nginx-rtc module. The nginx-rtc module responds to the meeting recording request and, after scheduling, forwards the request to the MCU confluence module; the MCU confluence module joins the corresponding room according to the meeting ID and participant ID in the meeting recording request. After the MCU confluence module joins the room, it performs a stream pulling operation on other participants in the room and saves the audio data locally. After the meeting ends, it uploads the audio data to the content server and notifies the text recognition module to perform text recognition. Finally, the text recognition module automatically generates the target meeting minutes text.

[0130] In the embodiments of the present application, on the one hand, through an intelligent data stream merging algorithm, efficient communication is promoted; and intelligent layout optimization is adopted to enhance the visual experience, accurately identify the speaking state, and provide a basis for subsequent confluence. On the other hand, through end-to-end integration, the problems of data landing and annotation are solved. After the MCU centrally processes multi-party audio and video streams, the meeting data is recorded in real time and transmitted to the text recognition module, enabling rapid integration and analysis of information, completing in-depth matching of the meeting scenario, and thus improving the accuracy of generating the target meeting minutes text.

[0131] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown in the direction of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0132] Based on the same inventive concept, the embodiments of the present application also provide a meeting minutes generation device for implementing the above-mentioned meeting minutes generation method. The implementation solutions for solving problems provided by this device are similar to those recorded in the above method. Therefore, the specific limitations in one or more embodiments of the meeting minutes generation device provided below can refer to the limitations on the meeting minutes generation method in the above text, and will not be repeated here.

[0133] In an exemplary embodiment, as Figure 9 shown, a meeting minutes generation device is provided, including:

[0134] An acquisition module 10, configured to obtain target audio data of a target meeting in response to a meeting recording request for the target meeting during the progress of the target meeting;

[0135] A first determination module 20, configured to determine the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the membership relationship between the candidate keywords and the candidate meeting types;

[0136] A second determination module 30, configured to determine the target text recognition model according to the correspondence between the candidate meeting types and the candidate text recognition models, and the target meeting type;

[0137] An identification module 40, configured to input target audio data into a target text recognition model to obtain a target meeting minutes text corresponding to the target audio data.

[0138] During the progress of the target meeting, the above-mentioned meeting minutes generation device responds to a meeting recording request for the target meeting, obtains the target audio data of the target meeting; determines the target meeting type of the target meeting according to the obtained target audio data, candidate keywords, and the membership relationship between the candidate keywords and the candidate meeting types; determines the target text recognition model according to the corresponding relationship between the candidate meeting types and the candidate text recognition models, and the target meeting type; inputs the target audio data into the target text recognition model to obtain a target meeting minutes text corresponding to the target audio data. In the above solution, a corresponding text recognition model is set for each meeting type, making the text recognition by the text recognition model more targeted, so that the target meeting minutes text corresponding to the target audio data obtained is more accurate.

[0139] In one embodiment, the first determination module 20 specifically includes:

[0140] A first determination unit, configured to determine the target keyword in the candidate keywords hit by the target audio data and the number of keywords of the target keyword;

[0141] A second determination unit, configured to determine the number of keywords and the number of keyword categories of the keywords that match each candidate meeting type for the target audio data according to the membership relationship between the candidate keywords and the candidate meeting types, the target keyword, and the number of keywords of the target keyword;

[0142] A third determination unit, configured to determine the target meeting type of the target meeting according to the number of keywords and the number of keyword categories of the keywords that match each candidate meeting type for the target audio data.

[0143] In one embodiment, the third determination unit is specifically configured to:

[0144] For each candidate meeting type, perform a weighted process on the number of keywords and the number of keyword categories of the keywords that match the candidate meeting type for the target audio data to obtain the matching degree between the target audio data and the candidate meeting type; use the candidate meeting type corresponding to the largest matching degree among the matching degrees as the target meeting type of the target meeting.

[0145] In one embodiment, the identification module 40 is specifically configured to:

[0146] Convert the target audio data according to a preset format template to obtain first audio data in the preset format; wherein, the preset format template at least includes a start recording time, a speaker identifier, a speaker's speaking time, and a speaking data length; input the first audio data into a target text recognition model to obtain a target meeting minutes text corresponding to the target audio data.

[0147] In one embodiment, the apparatus further includes a training module for:

[0148] For each candidate meeting type, obtain sample data corresponding to the candidate meeting type; wherein, the sample data includes sample audio data corresponding to the candidate meeting type and a sample meeting minutes text; convert the sample audio data according to the preset format template to obtain second audio data in the preset format; use the second audio data as training data and the sample meeting minutes text as label data to train an initial text recognition model to obtain a candidate text recognition model corresponding to the candidate meeting type.

[0149] In one embodiment, the target text recognition model includes a text recognition unit and a meeting minutes generation unit; the recognition module 40 is specifically configured to:

[0150] Input the target audio data into the text recognition unit for text recognition to obtain initial text data corresponding to the target audio data; input the initial text data into the meeting minutes generation unit for text parsing to obtain a target meeting minutes text in a template format; wherein, the template format includes at least one of a meeting date, participants, meeting keywords, chapter summary, to-do items, and meeting Q&A.

[0151] Each module in the above meeting minutes generation apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in a processor in a computer device in hardware form or be independent of it, or can be stored in a memory in a computer device in software form so that the processor can call and execute operations corresponding to the above respective modules.

[0152] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a method for generating meeting minutes. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0153] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0154] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the method for generating meeting minutes described in any of the above embodiments are implemented.

[0155] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for generating meeting minutes described in any of the above embodiments are implemented.

[0156] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the method for generating meeting minutes described in any of the above embodiments are implemented.

[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0158] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0159] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0160] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. A method for generating meeting minutes, characterized in that: The method comprises: During the target conference, in response to a conference recording request for the target conference, acquiring target audio data of the target conference; Determining a target conference type of the target conference according to the acquired target audio data, candidate keywords, and affiliation between the candidate keywords and the candidate conference types; Determining a target text recognition model according to the correspondence between the candidate conference types and the candidate text recognition models, and the target conference type; The target audio data is input into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

2. The method according to claim 1, characterized in that The determining the target conference type of the target conference according to the acquired target audio data, the candidate keywords, and the affiliation between the candidate keywords and the candidate conference types includes: Determine a target keyword and a keyword quantity of the target keyword among candidate keywords hit by the target audio data; Determine the number of keywords and the number of keyword categories of the keywords that match the target audio data with each candidate conference type according to the affiliation between the candidate keywords and the candidate conference types, the target keywords, and the keyword quantity of the target keywords; The target conference type of the target conference is determined according to the keyword quantity and keyword category quantity of the keywords that match the target audio data with each candidate conference type.

3. The method according to claim 2, characterized in that The determining the target conference type of the target conference according to the keyword quantity and keyword category quantity of the keywords matching the target audio data with each candidate conference type includes: For each candidate conference type, weighting the number of keywords and the number of keyword categories of the keywords that match the target audio data with the candidate conference type is performed to obtain a matching degree between the target audio data and the candidate conference type; The candidate conference type corresponding to the maximum matching degree among the matching degrees is used as the target conference type of the target conference.

4. The method according to claim 1, characterized in that: The step of inputting the target audio data into the target text recognition model to obtain a target meeting minutes text corresponding to the target audio data includes: According to a preset format template, the target audio data is converted to obtain first audio data in a preset format; wherein the preset format template at least includes a start recording time, a speaker identifier, a speaker's speaking time, and a length of speech data; The first audio data is input into a target text recognition model to obtain a target meeting minutes text corresponding to the target audio data.

5. The method according to claim 1, characterized in that Each candidate text recognition model is established in the following way: For each candidate meeting type, obtain sample data corresponding to the candidate meeting type; wherein the sample data includes sample audio data and sample meeting minutes text corresponding to the candidate meeting type; Convert the sample audio data according to a preset format template to obtain second audio data in a preset format; The initial text recognition model is trained using the second audio data as training data and the sample meeting minutes text as label data to obtain a candidate text recognition model corresponding to the candidate meeting type.

6. The method according to claim 1, characterized in that The target text recognition model includes a text recognition unit and a meeting minutes generation unit; the target audio data is input into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data, including: Inputting the target audio data into the text recognition unit for text recognition to obtain initial text data corresponding to the target audio data; The initial text data is input into the meeting minutes generation unit for text parsing to obtain a target meeting minutes text in a template format; wherein the template format includes at least one of the meeting date, participants, meeting keywords, chapter summaries, to-do items, and meeting questions and answers.

7. A device for generating meeting minutes, characterized in that: The device comprises: An acquisition module, configured to acquire target audio data of a target conference in response to a conference recording request for the target conference during the target conference; A first determination module, configured to determine a target conference type of the target conference according to the acquired target audio data, candidate keywords, and affiliation between the candidate keywords and the candidate conference types; A second determination module, configured to determine a target text recognition model according to a correspondence between candidate conference types and candidate text recognition models, and the target conference type; The recognition module is used to input the target audio data into the target text recognition model to obtain the target meeting minutes text corresponding to the target audio data.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.