Method, apparatus, device, and readable storage medium for generating meeting minutes

By receiving and analyzing the sound signals collected by multiple mobile terminals, generating user opinion information and conference topic information, the problem of lack of intelligence and high cost when generating conference records in the existing conference system is solved, and low-cost and high-intelligence conference records are achieved.

CN111626061BActive Publication Date: 2025-06-27WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010464020.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-27
Publication Date
2025-06-27
Estimated Expiration
2040-05-27

AI Technical Summary

Technical Problem

When generating meeting records, existing conference systems lack intelligence, rely on human search and addition, and have high costs or poor radio effects.

Method used

By receiving the sound signals collected by multiple mobile terminals, identifying and semantic analysis, user opinion information and conference topic information are generated, and intelligent conference records are formed.

Benefits of technology

A low-cost and highly intelligent conference system is realized, which avoids the problems of poor audio effects of single audio equipment and high conversion costs of multiple audio equipment, and improves the efficiency and intelligence of conference record generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111626061B_ABST
    Figure CN111626061B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and readable storage medium for generating meeting records. The method for generating meeting records is applied to a meeting system, and the meeting system is formed based on a plurality of mobile terminals added to the meeting. The method includes: receiving voice signals collected by each mobile terminal in the meeting, and respectively recognizing each voice signal to obtain multiple pieces of text information; performing semantic recognition on the multiple pieces of text information to generate user opinion information corresponding to each mobile terminal and meeting theme information; generating the multiple pieces of text information, each piece of user opinion information and the meeting theme information into a meeting record. By using a plurality of mobile terminals as sound collection devices to form a meeting system, the present invention ensures the sound collection effect while avoiding the transformation of the meeting system; and performs semantic recognition on multiple pieces of text information to automatically generate user opinion information and meeting theme information, improving the intelligence of meeting record generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial technology (Fintech), and particularly to a method, apparatus, device, and readable storage medium for generating meeting records. Background Art

[0002] With the continuous development of financial technology (Fintech), especially Internet technology finance, more and more technologies (such as artificial intelligence, big data, cloud storage, etc.) are applied in the financial field. However, the financial field also poses higher requirements for various technologies, such as requiring the improvement of the intelligence of the meeting system.

[0003] Currently, meeting systems mainly exist in two modes. One is to use a single sound collection device for sound collection, and the sound collection device is moved for sound collection. The sound collection effect is poor and it is not convenient to use. The other is to use multiple sound collection devices for sound collection, but multiple sound collection devices require a large amount of additional hardware to transform the meeting system, and the transformation cost is high. And whether it is a meeting system with a single sound collection device or multiple sound collection devices, the meeting records are formed by converting the audio data recorded by the meeting system into text information for subsequent viewing. However, this meeting record is a simple conversion of the audio data, and the content related to the meeting depends on manual search and addition, and the process of generating the meeting record is not intelligent enough.

[0004] Therefore, how to form a highly intelligent meeting system at low cost is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method, apparatus, device, and readable storage medium for generating meeting records, aiming to solve the technical problem of how to form a highly intelligent meeting system at low cost in the prior art.

[0006] To achieve the above object, the present invention provides a method for generating meeting records. The method for generating meeting records is applied to a meeting system, and the meeting system is formed based on multiple mobile terminals added to the meeting. The method for generating meeting records includes the following steps:

[0007] Receiving the sound signals collected by each mobile terminal in the meeting, and respectively identifying each of the sound signals to obtain multiple pieces of text information;

[0008] Performing semantic recognition on the multiple pieces of text information to generate user opinion information corresponding to each mobile terminal and meeting theme information;

[0009] Generating the multiple pieces of text information, each piece of user opinion information, and the meeting theme information into a meeting record.

[0010] Optionally, the step of performing semantic recognition on the multiple pieces of text information to generate user opinion information corresponding to each mobile terminal and conference theme information includes:

[0011] Based on a preset theme model, perform semantic recognition on the multiple pieces of text information to generate the conference theme information;

[0012] Classify the multiple pieces of text information according to the user identifiers corresponding to the respective voice signals to generate classified text information corresponding to each mobile terminal;

[0013] Based on a preset theme model, perform semantic recognition on each piece of classified text information to generate user opinion information corresponding to each mobile terminal.

[0014] Optionally, after the step of receiving voice signals collected by each mobile terminal in the conference, the method further includes:

[0015] Extract the voiceprint information to be recognized from the voice signals collected by each mobile terminal, and determine whether each piece of voiceprint information to be recognized is valid according to the voiceprint information associated with each mobile terminal;

[0016] If each piece of voiceprint information to be recognized is valid, filter each voice signal according to the voiceprint information associated with each mobile terminal to update each voice signal.

[0017] Optionally, after the step of determining whether each piece of voiceprint information to be recognized is valid, the method further includes:

[0018] If there is invalid voiceprint information to be recognized among each piece of voiceprint information to be recognized, find the target mobile terminal corresponding to the invalid voiceprint information to be recognized;

[0019] Determine whether the target mobile terminal carries an authorization identifier. If it carries an authorization identifier, perform the step of filtering each voice signal according to the voiceprint information associated with each mobile terminal;

[0020] If it does not carry an authorization identifier, remove the target mobile terminal from the conference.

[0021] Optionally, the step of generating a conference record from the multiple pieces of text information, each piece of user opinion information, and the conference theme information includes:

[0022] Arrange the multiple pieces of text information according to the time information corresponding to the multiple pieces of text information;

[0023] Add each piece of user opinion information and the conference theme information to the arranged multiple pieces of text information to generate a conference record.

[0024] Optionally, before the step of receiving the voice signals collected by each mobile terminal in the meeting, the method further includes:

[0025] Collecting voiceprint information of the holders of each mobile terminal based on each mobile terminal;

[0026] Forming an association relationship between each mobile terminal and the voiceprint information of the holder of each mobile terminal, and adding each association relationship to the voiceprint database for storage.

[0027] Optionally, after the step of generating the meeting record from the multiple text information, each user opinion information, and the meeting theme information, the method further includes:

[0028] Sending the meeting record to each mobile terminal for display.

[0029] Furthermore, to achieve the above object, the present invention further provides a meeting record generation device, and the meeting record generation device includes:

[0030] A receiving module, configured to receive the voice signals collected by each mobile terminal in the meeting, and respectively identify each voice signal to obtain multiple text information;

[0031] An identification module, configured to perform semantic identification on the multiple text information to generate user opinion information corresponding to each mobile terminal and meeting theme information;

[0032] A generation module, configured to generate a meeting record from the multiple text information, each user opinion information, and the meeting theme information.

[0033] Furthermore, to achieve the above object, the present invention further provides a meeting system, and the meeting system includes a memory, a processor, and a meeting record generation program stored on the memory and executable on the processor. When the meeting record generation program is executed by the processor, the steps of the meeting record generation method as described above are implemented.

[0034] Furthermore, to achieve the above object, the present invention further provides a readable storage medium, and a meeting record generation program is stored on the readable storage medium. When the meeting record generation program is executed by a processor, the steps of the meeting record generation method as described above are implemented.

[0035] The method, device, equipment and computer-readable storage medium for generating meeting records according to the present invention. The method for generating meeting records is applied to a meeting system formed by adding multiple mobile terminals to a meeting. Each mobile terminal is used for sound collection to collect the voice signals of the holders of each mobile terminal. After the meeting system receives the voice signals collected by each mobile terminal in the meeting, it performs recognition and conversion on each voice signal to obtain multiple text information; and performs semantic recognition on the multiple text information to obtain user opinion information representing the opinions of the holders of each mobile terminal and the meeting theme information of this meeting; and then generates the multiple text information, each user opinion information and the meeting theme information into a meeting record. By using multiple mobile terminals as sound collection devices to form a meeting system, on the one hand, it avoids the problems of poor sound collection effect and inconvenient use of the meeting system due to a single sound collection device; on the other hand, it avoids the problem of high cost for the transformation of the meeting system by multiple sound collection devices; and performs semantic recognition on the multiple text information to automatically generate user opinion information and meeting theme information, avoiding manual search and addition, and improving the efficiency and intelligence of generating meeting records. Therefore, a meeting system with high intelligence is formed at low cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic structural diagram of the hardware operating environment of the equipment related to the embodiment of the meeting system of the present invention;

[0037] Figure 2 It is a schematic flowchart of the first embodiment of the method for generating meeting records according to the present invention;

[0038] Figure 3 It is a schematic diagram of the functional modules of the preferred embodiment of the device for generating meeting records according to the present invention;

[0039] Figure 4 It is an architecture diagram of the meeting system to which the method for generating meeting records according to the present invention is applied.

[0040] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] The present invention provides a meeting system. Refer to Figure 1 , Figure 1 It is a schematic structural diagram of the hardware operating environment of the equipment related to the embodiment of the meeting system of the present invention.

[0043] As Figure 1As shown in the figure, the conference system may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to implement the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0044] Those skilled in the art can understand that Figure 1 the hardware structure of the conference system shown in the figure does not limit the conference system, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0045] As Figure 1 shown, the memory 1005, as a readable storage medium, may include an operating system, a network communication module, a user interface module, and a conference record generation program. Among them, the operating system is a program for managing and controlling the conference system and software resources, and supports the operation of the network communication module, the user interface module, the conference record generation program, and other programs or software; the network communication module is used to manage and control the network interface 1004; the user interface module is used to manage and control the user interface 1003.

[0046] In Figure 1 the hardware structure of the conference system shown in the figure, the network interface 1004 is mainly used to connect to the background server and communicate with the background server for data; the user interface 1003 is mainly used to connect to the client (user side) and communicate with the client for data; the processor 1001 can call the conference record generation program stored in the memory 1005 and perform the following operations:

[0047] Receive the voice signals collected by each mobile terminal in the conference, and respectively identify each of the voice signals to obtain multiple pieces of text information;

[0048] Perform semantic recognition on the multiple pieces of text information to generate user opinion information corresponding to each mobile terminal and conference theme information;

[0049] Generate the multiple pieces of text information, each of the user opinion information, and the conference theme information into a conference record.

[0050] Further, the step of performing semantic recognition on the multiple pieces of text information to generate user opinion information corresponding to each mobile terminal and conference theme information includes:

[0051] Based on a preset theme model, perform semantic recognition on the multiple pieces of text information to generate the conference theme information;

[0052] Classify the multiple pieces of text information according to the user identifiers corresponding to the respective voice signals to generate classified text information corresponding to each mobile terminal;

[0053] Based on a preset theme model, perform semantic recognition on each piece of classified text information to generate user opinion information corresponding to each mobile terminal.

[0054] Further, after the step of receiving the voice signals collected by each mobile terminal in the conference, the processor 1001 may call the conference record generation program stored in the memory 1005 and perform the following operations:

[0055] Extract the voiceprint information to be recognized from the voice signals collected by each mobile terminal, and determine whether each piece of voiceprint information to be recognized is valid according to the voiceprint information associated with each mobile terminal;

[0056] If each piece of voiceprint information to be recognized is valid, filter each voice signal according to the voiceprint information associated with each mobile terminal to update each voice signal.

[0057] Further, after the step of determining whether each piece of voiceprint information to be recognized is valid, the processor 1001 may call the conference record generation program stored in the memory 1005 and perform the following operations:

[0058] If there is invalid voiceprint information to be recognized among each piece of voiceprint information to be recognized, find the target mobile terminal corresponding to the invalid voiceprint information to be recognized;

[0059] Determine whether the target mobile terminal carries an authorization identifier. If it carries an authorization identifier, perform the step of filtering each voice signal according to the voiceprint information associated with each mobile terminal;

[0060] If it does not carry an authorization identifier, remove the target mobile terminal from the conference.

[0061] Further, the step of generating a conference record from the multiple pieces of text information, each piece of user opinion information, and the conference theme information includes:

[0062] Arrange the multiple pieces of text information according to the time information corresponding to the multiple pieces of text information;

[0063] Add each of the user view information and the meeting theme information to the arranged multiple text information to generate a meeting record.

[0064] Further, before the step of receiving the voice signals collected by each mobile terminal in the meeting, the processor 1001 may call the meeting record generation program stored in the memory 1005 and perform the following operations:

[0065] Collect voiceprint information of the holders of each mobile terminal based on each mobile terminal;

[0066] Form an association relationship between each mobile terminal and the voiceprint information of the holder of each mobile terminal, and add each association relationship to the voiceprint database for storage.

[0067] Further, after the step of generating a meeting record from the multiple text information, each user view information, and the meeting theme information, the processor 1001 may call the meeting record generation program stored in the memory 1005 and perform the following operations:

[0068] Send the meeting record to each mobile terminal for display.

[0069] The specific implementation manner of the meeting system of the present invention is basically the same as that of each embodiment of the following meeting record generation method, and will not be described in detail here.

[0070] The present invention also provides a meeting record generation method.

[0071] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the meeting record generation method of the present invention.

[0072] The embodiments of the present invention provide embodiments of a meeting record generation method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here. Specifically, the meeting record generation method in this embodiment is applied to a meeting system formed based on multiple mobile terminals added to the meeting. The meeting record generation method includes:

[0073] Step S10, receive the voice signals collected by each mobile terminal in the meeting, and respectively identify each voice signal to obtain multiple text information;

[0074] The meeting record generation method in this embodiment is applied to a meeting system, and the meeting system is formed by multiple mobile terminals added to the meeting. The mobile terminals are intelligent terminals used by meeting participants such as mobile phones and tablets. First, install client software that can access the meeting in the mobile terminal, and connect to the server through the software to form a meeting system. Please refer to Figure 4 , Figure 4 which shows the architecture of the meeting system. Figure 4 The server of the meeting system in [the figure] forms the service layer, and each mobile terminal accesses the service layer through the network layer. The network layer can be formed by calling the external network cloud or by private deployment of the internal network. The service layer provides algorithm services and management services of the background system to the accessed mobile terminals. The provided algorithm services include but are not limited to speech recognition, voiceprint recognition, etc., and the associated services of the background system at least include meeting processing, meeting management, and user management, etc. Among them, meeting processing is mainly for generating meeting records, meeting management is mainly for scheduling and convening meetings, and user management is mainly for connecting users to the meeting system.

[0075] When a meeting needs to be held through the meeting system, the meeting organizer or the meeting system administrator applies to start the meeting system, and the meeting number is fed back to each meeting participant. Each meeting participant applies to join the meeting system through the meeting number. The mobile terminal held by the meeting participant serves as the sound collection device in the meeting, collects the sound signal and transmits it to the server of the meeting system for processing, so as to avoid adding additional sound collection devices, and each meeting participant uses their own mobile terminal for sound collection, ensuring the sound collection effect of each meeting participant.

[0076] It can be understood that there are many participants in the meeting, and the meeting participants may speak one by one or speak simultaneously in a group discussion during the meeting. For a meeting where participants speak one by one, there may be environmental noise, resulting in noise in the sound signal collected by the mobile terminal of the speaker; for simultaneous speaking, the mobile terminals of multiple speakers will inevitably collect the sound signals of other speakers, forming noise in the sound signals collected by each mobile terminal. Therefore, after the meeting system receives the sound signals collected by each mobile terminal through the sound collection of each mobile terminal, it distinguishes each sound signal to distinguish the speech content of each meeting participant in the meeting, and converts it into real-time text information for display on each mobile terminal for each meeting participant to view. And in this embodiment, each sound signal is distinguished through the pre-set voiceprint information bound to the mobile terminal. First, search for the voiceprint information bound to the mobile terminal, then analyze the sound signal collected by the mobile terminal, determine the sound signal that matches the voiceprint information, and extract this sound signal as the speech content of the mobile terminal holder in the meeting.

[0077] Further, after distinguishing the voice signals collected by each mobile terminal, the voice signals are converted and recognized, and the converted text information is displayed on each mobile terminal one by one according to the collection time of each voice signal and the user name of the meeting participant, so as to realize conversion while collecting for real-time viewing by the meeting participants. Specifically, a speech recognition algorithm for recognizing and converting voice signals is pre-set in the server of the meeting system. After the meeting system receives the voice signals collected by each mobile terminal, it calls this recognition algorithm to recognize each voice signal, converts each voice signal into its respective text information, and obtains multiple text information. Then, according to the collection time of each voice signal, the multiple text information and their respective voice signals are displayed item by item in correspondence, which is convenient for each meeting participant to view.

[0078] Step S20: Perform semantic recognition on the multiple text information to generate user opinion information corresponding to each mobile terminal and meeting theme information.

[0079] Furthermore, a preset theme model for analyzing the theme, such as a probabilistic theme model, is pre-set in the server of the meeting system. While recognizing each voice signal and obtaining the display of multiple text information, semantic recognition is also performed on the multiple text information through the preset theme model to obtain the themes reflected by each of the multiple text information and the theme reflected by the multiple text information as a whole. Among them, the themes reflected by each of the multiple text information reflect the opinions of each meeting participant and are the user opinion information corresponding to each mobile terminal; the theme reflected as a whole reflects the theme issues of the meeting as a whole and is the meeting theme information.

[0080] Step S30: Generate a meeting record from the multiple text information, each user opinion information, and the meeting theme information.

[0081] Further, after converting the voice signals collected by each mobile terminal in the meeting into multiple text information and extracting the user opinion information of each meeting participant and the meeting theme information from the multiple text information, the multiple text information, each user opinion information, and the meeting theme information are formed into a meeting record, which is convenient for subsequent viewing of the meeting content. Specifically, the steps of generating a meeting record from the multiple text information, each user opinion information, and the meeting theme information include:

[0082] Step S31: Arrange the multiple text information according to the time information corresponding to the multiple text information.

[0083] Step S32: Add each user opinion information and the meeting theme information to the arranged multiple text information to generate a meeting record.

[0084] Furthermore, different meeting record templates are preset for different meeting topics. During the process of generating meeting records, the meeting record template that matches the meeting topic information is called first. Moreover, the acquisition time of each voice signal that generates multiple text messages is found, and this acquisition time is used as the time information corresponding to the multiple text messages. Then, according to this time information, the multiple text messages are added to the meeting record template for arrangement; the text messages with earlier time information are arranged in the front, and the text messages with later time information are arranged in the back, forming multiple text messages arranged in chronological order, which reflects the speech content of each meeting participant in chronological order. After that, each user's opinion information and the meeting topic information are added to the multiple text messages after arrangement. The added position can be a preset position in the meeting record template or a custom position. In this way, the final meeting record is formed and sent to each mobile terminal for visual display, which not only reflects the specific speech content of the meeting participants in the meeting but also facilitates quickly viewing the opinions of each meeting participant and the meeting topic; it avoids manually adding the meeting topic, and the intelligence of the meeting system is higher.

[0085] The meeting record generation method of the present invention is applied to a meeting system formed by adding multiple mobile terminals to a meeting. Each mobile terminal is used for sound collection to collect the voice signals of the holders of each mobile terminal. After the meeting system receives the voice signals collected by each mobile terminal in the meeting, it performs recognition and conversion on each voice signal to obtain multiple text messages; and performs semantic recognition on the multiple text messages to obtain user opinion information representing the opinions of the holders of each mobile terminal and the meeting topic information of this meeting; then generates a meeting record from the multiple text messages, each user opinion information, and the meeting topic information. By forming a meeting system with multiple mobile terminals as sound collection devices, on the one hand, it avoids the problems of poor sound collection effect and inconvenient use of the meeting system due to a single sound collection device; on the other hand, it avoids the high cost problem of transforming the meeting system with multiple sound collection devices; and performs semantic recognition on multiple text messages to automatically generate user opinion information and meeting topic information, avoiding manual search and addition, and improving the efficiency and intelligence of meeting record generation. Therefore, a meeting system with high intelligence is formed at low cost.

[0086] Further, based on the first embodiment of the meeting record generation method of the present invention, a second embodiment of the meeting record generation method of the present invention is proposed.

[0087] The difference between the second embodiment of the meeting record generation method and the first embodiment of the meeting record generation method is that the step of performing semantic recognition on the multiple text messages to generate user opinion information corresponding to each mobile terminal and meeting topic information includes:

[0088] Step S21: Based on a preset topic model, perform semantic recognition on multiple pieces of the text information to generate the meeting topic information;

[0089] Step S22: According to the user identifiers corresponding to the respective voice signals, classify the multiple pieces of the text information to generate classified text information corresponding to each of the mobile terminals;

[0090] Step S23: Based on the preset topic model, perform semantic recognition on each of the classified text information to generate user opinion information corresponding to each of the mobile terminals.

[0091] In this embodiment, a preset topic model is used to generate user opinion information and meeting topic information. Specifically, an initial model is pre-trained with a large number of training texts until the initial model can accurately recognize the semantics of the large number of training texts, extract the topic information therein, and then the initial model is generated into a preset topic model. For the multiple pieces of text information obtained through conversion, the preset topic model is called to perform semantic recognition on it, extract the topic words therein, and calculate the scores of each topic word. The topic word with the largest score is determined as the meeting topic information to represent the meeting topic. It should be noted that for the case where multiple topic words all have relatively high scores and the scores are not much different from each other, it indicates that the meeting may have multiple topics. At this time, these multiple topic words can be set as the topic information together to reflect the topics of the meeting in multiple aspects.

[0092] It can be understood that the multiple pieces of text information are generated according to the speaking time sequence of the meeting participants. The same meeting participant has different speaking contents at different times, resulting in more than one piece of text information for the same meeting participant in the multiple pieces of text information. Therefore, in order to accurately generate user opinion information, this embodiment classifies the speaking contents of the same meeting participant. Specifically, each meeting participant collects voice signals through their respective mobile terminals. The collected voice signals carry the identifier representing the source mobile terminal, and the text information obtained through voice signal conversion also carries the identifier of the mobile terminal. Thus, classification can be performed according to the identifiers carried by each piece of text information. The identifier carried by each voice signal is used as the user identifier corresponding to each voice signal, and according to each user identifier, the multiple pieces of text information are classified. The text information with the same user identifier is grouped into the same category, while the text information with different user identifiers is grouped into different categories. After the classification is completed, the number of classified categories is the same as the number of mobile terminals, and each category of text information forms the classified text information corresponding to each mobile terminal.

[0093] Further, semantic recognition is performed on each classified text information through a preset topic model. For each classified text information, its respective topic words are extracted, and the scores of its respective topic words are calculated. The topic word with the largest score is determined as the target topic word for each type of text information. These target topic words reflect the core viewpoints of the speeches of each meeting participant, so they are generated into user viewpoint information corresponding to each mobile terminal. Similarly, for a certain classified text information, if multiple generated topic words all have relatively high scores and the scores are not much different from each other, it indicates that the corresponding meeting participant has multiple core viewpoints. Therefore, these multiple topic words are jointly generated into user viewpoint information corresponding to its mobile terminal to reflect viewpoints in multiple aspects.

[0094] In this embodiment, through a preset topic model, the meeting theme information is extracted from multiple text information, reflecting the meeting theme, avoiding relying on the meeting organizer to determine the meeting theme, and the intelligence of the meeting system is higher. Moreover, by classifying multiple text information to generate user viewpoint information corresponding to each mobile terminal, it is beneficial to accurately reflect the viewpoints of each meeting participant.

[0095] Further, based on the first or second embodiment of the meeting record generation method of the present invention, a third embodiment of the meeting record generation method of the present invention is proposed.

[0096] The difference between the third embodiment of the meeting record generation method and the first or second embodiment of the meeting record generation method is that before the step of receiving the voice signals collected by each mobile terminal in the meeting, the following steps are further included:

[0097] Step a1, based on each of the mobile terminals, collect the voiceprint information of the holders of each of the mobile terminals;

[0098] Step a2, form an association relationship between each of the mobile terminals and the voiceprint information of the holders of each of the mobile terminals, and add each of the association relationships to the voiceprint database for storage.

[0099] In this embodiment, the mobile terminal and the voiceprint information of the mobile terminal holder are pre-bound to extract each voice signal through the bound voiceprint information. Specifically, the audio signals of the holders of each mobile terminal are collected through each mobile terminal, and feature extraction is performed on the audio signals to form the voiceprint information of the holders of each mobile terminal; among them, the extracted features include but are not limited to the number, trend, and frequency of formants, etc. Furthermore, an association relationship is formed between each mobile terminal and the voiceprint information of the holder of each mobile terminal, and each association relationship is added to the voiceprint database docked with the meeting system for storage. And the association relationship can exist in the form of a key_value key-value pair, where the key is the mobile terminal and the value is the voiceprint information of the mobile terminal holder.

[0100] Further, after receiving the voice signals collected by each mobile terminal, the noise in each voice signal can be removed through the associated relationship bound in the voiceprint database, and the effective voice signals can be extracted therefrom, so as to update the voice signals collected by each mobile terminal for recognition and convert them into text information.

[0101] Furthermore, considering that there may be cases of impostors among some participants in a meeting. For example, some qualified participants do not want to or do not have time to participate in the meeting, and let other unqualified participants use their mobile terminals to participate. Therefore, in order to avoid such situations, this embodiment is provided with a verification mechanism, which is started immediately after receiving the voice signals collected by each mobile terminal; specifically, after the step of receiving the voice signals collected by each mobile terminal in the meeting, the following steps are further included:

[0102] Step b1, respectively extract the voiceprint information to be recognized from the voice signals collected by each of the mobile terminals, and determine whether each of the voiceprint information to be recognized is valid according to the voiceprint information associated with each of the mobile terminals;

[0103] Step b2, if each of the voiceprint information to be recognized is valid, filter each of the voice signals according to the voiceprint information associated with each of the mobile terminals, so as to update each of the voice signals.

[0104] Further, respectively extract the voiceprint information to be recognized from the voice signals collected by each mobile terminal, search for the associated relationship of each mobile terminal in the voiceprint database, and use the voiceprint information bound in each associated relationship as the voiceprint information associated with each mobile terminal. Then, for each mobile terminal, compare the voiceprint information associated with it with the voiceprint information to be recognized extracted from the voice signals collected by this mobile terminal. Determine whether the voiceprint information and the voiceprint information to be recognized are consistent, or the similarity is greater than a preset threshold. If they are consistent or the similarity is greater than the preset threshold, it is determined that the voiceprint information to be recognized is valid, indicating that the meeting participant and the mobile terminal match. Then, filter the voice signals according to the voiceprint information associated with the mobile terminal, remove the noise formed by other meeting participants in the voice signals, retain the voice signals of the meeting participants matching the mobile terminal for extraction, and obtain effective updated voice signals for recognition and conversion.

[0105] It can be understood that in the process of judging the validity of the voiceprint information to be recognized, there may be a situation where it is judged to be invalid. At this time, it means that there is an impostor among the meeting participants, and further authorization determination is required. Specifically, after the step of determining whether each voiceprint information to be recognized is valid, the following steps are further included:

[0106] Step b3, if there is invalid voiceprint information to be recognized among each of the voiceprint information to be recognized, search for the target mobile terminal corresponding to the invalid voiceprint information to be recognized;

[0107] Step b4, determining whether the target mobile terminal carries an authorization identifier, and if so, performing a step of filtering each of the sound signals according to the voiceprint information associated with each of the mobile terminals;

[0108] Step b5: If the target mobile terminal does not carry the authorization identifier, the target mobile terminal is removed from the conference.

[0109] Furthermore, if it is determined through comparison that there is any invalid voiceprint information to be identified in each voiceprint information to be identified, that is, the similarity between a certain voiceprint information to be identified and its corresponding voiceprint information is less than a preset threshold, indicating that the two have a large difference, then the voiceprint information to be identified is determined to be invalid; the voiceprint information to be identified collected by the mobile terminal is not the voiceprint information bound to the mobile terminal, and the participant currently using the mobile terminal to participate in the meeting is not the holder of the mobile terminal. At this time, first determine the target mobile terminal corresponding to the voiceprint information to be identified, that is, the mobile terminal used by the fraudulent user, based on the identifier of the mobile terminal carried by the sound signal from which the voiceprint information to be identified comes. Then detect whether the target mobile terminal carries an authorized identifier. Among them, the authorized identifier is an identifier formed by the holder of the mobile terminal applying to the conference system when the participant who uses the mobile terminal to participate in the meeting is inconsistent with the holder of the mobile terminal. Before the meeting starts, the holder of the mobile terminal initiates a conference replacement application to the conference system through its mobile terminal, and the conference system sends the replacement application to the manager or conference organizer; the conference manager or conference organizer returns an instruction to agree or disagree to the application to the conference system based on the substitute person information in the replacement application. After receiving the application approval instruction, the conference system allocates an authorization identifier to the mobile terminal that initiated the conference replacement application, so that the mobile terminal carries the authorization identifier.

[0110] Furthermore, if it is determined that the target terminal carries an authorization mark, it means that although the participant using the mobile terminal to participate in the meeting is not the same as the holder of the mobile terminal, the participant is effectively authorized and qualified to participate in the meeting, so the sound signal is filtered according to the voiceprint information associated with the mobile terminal. On the contrary, if it is determined that the target mobile terminal does not carry an authorization mark, it indicates that the meeting cannot be participated in by other users using the mobile terminal instead of the mobile terminal holder, or the mobile terminal holder has not initiated a meeting replacement application. At this time, the target mobile terminal is removed from the meeting to prohibit non-mobile terminal holders from participating in the meeting and avoid leakage of meeting content.

[0111] It should be noted that since the conference participants are not the same as the mobile terminal holders, the associated voiceprint information is inconsistent with the voiceprint information in the sound signals collected by the mobile terminal. At this time, if the original associated voiceprint information of the mobile terminal is still used to filter the sound signals, all the sound signals collected by the mobile terminal will be filtered out. Therefore, when the mobile terminal holder applies for a meeting replacement, the audio signals of the participants who use the mobile terminal to participate in the meeting collected by the mobile terminal are uploaded to the conference system. After receiving the approval application instruction, the conference system extracts the voiceprint information from the audio signals and forms a new association relationship with the mobile terminal and transmits it to the voiceprint database for storage, and adds a temporary identifier for the new association relationship, so as to delete the association relationship representing the voiceprint information of non-mobile terminal holders with temporary identifiers after the end of this meeting. In this way, for the situation where the conference participants are not the same as the mobile terminal holders, in the process of filtering the sound signals according to the voiceprint information associated with the mobile terminal, the new voiceprint information associated with the mobile terminal is determined through the new association relationship, so as to filter the sound signals according to the new voiceprint information, avoid filtering all the sound signals collected by the mobile terminal, and ensure that accurate sound signals are extracted for recognition and conversion.

[0112] In this embodiment, the voiceprint information of the mobile terminal and the mobile terminal holder is bound to form an association relationship. By comparing the consistency between the voiceprint information in the sound signals collected by the mobile terminal and the bound voiceprint information, it is determined whether there is an impersonation situation for the participants who use the mobile terminal to participate in the meeting, ensuring that the conference participants have the qualification to participate and avoiding the leakage of conference content. At the same time, for the cleaning of the need to replace the participation in the meeting, an authorization mechanism is set up, which is flexible while ensuring the security of the meeting.

[0113] The present invention also provides a conference record generating device.

[0114] Refer to Figure 3 , Figure 3 which is a schematic diagram of the functional modules of the first embodiment of the conference record generating device of the present invention.

[0115] The conference record generating device includes:

[0116] A receiving module 10, configured to receive the sound signals collected by each mobile terminal in the conference, and respectively identify each of the sound signals to obtain multiple pieces of text information;

[0117] An identification module 20, configured to perform semantic identification on the multiple pieces of text information to generate user opinion information corresponding to each mobile terminal and conference theme information;

[0118] A generating module 30, configured to generate a conference record from the multiple pieces of text information, each piece of user opinion information, and the conference theme information.

[0119] Further, the recognition module 20 further includes:

[0120] A recognition unit, configured to perform semantic recognition on multiple copies of the text information based on a preset topic model to generate the conference topic information;

[0121] A classification unit, configured to classify multiple copies of the text information according to the user identifiers corresponding to the respective voice signals to generate classified text information corresponding to each of the mobile terminals;

[0122] A generation unit, configured to perform semantic recognition on each of the classified text information based on a preset topic model to generate user opinion information corresponding to each of the mobile terminals.

[0123] Further, the conference record generation device further includes:

[0124] An extraction module, configured to respectively extract to-be-recognized voiceprint information from the voice signals collected by each of the mobile terminals, and determine whether each of the to-be-recognized voiceprint information is valid according to the voiceprint information associated with each of the mobile terminals;

[0125] A filtering module, configured to, if each of the to-be-recognized voiceprint information is valid, filter each of the voice signals according to the voiceprint information associated with each of the mobile terminals to update each of the voice signals.

[0126] Further, the conference record generation device further includes:

[0127] A search module, configured to, if there is invalid to-be-recognized voiceprint information among each of the to-be-recognized voiceprint information, search for a target mobile terminal corresponding to the invalid to-be-recognized voiceprint information;

[0128] A judgment module, configured to judge whether the target mobile terminal carries an authorization identifier, and if it carries an authorization identifier, execute the step of filtering each of the voice signals according to the voiceprint information associated with each of the mobile terminals;

[0129] A removal module, configured to, if it does not carry an authorization identifier, remove the target mobile terminal from the conference.

[0130] Further, the generation module 30 further includes:

[0131] An arrangement unit, configured to arrange multiple copies of the text information according to the time information corresponding to the multiple copies of the text information;

[0132] An addition unit, configured to add each of the user opinion information and the conference topic information to the arranged multiple copies of the text information to generate a conference record.

[0133] Further, the conference record generation device further includes:

[0134] A collection module, configured to collect voiceprint information of the holders of the mobile terminals based on each of the mobile terminals;

[0135] An adding module, configured to form an association relationship between each of the mobile terminals and the voiceprint information of the holders of the mobile terminals, and add each of the association relationships to a voiceprint database for storage.

[0136] Further, the meeting record generating device further includes:

[0137] A display module, configured to send the meeting record to each of the mobile terminals for display.

[0138] The specific implementation manner of the meeting record generating device of the present invention is basically the same as that of the embodiments of the above-mentioned meeting record generating method, and will not be described in detail herein.

[0139] In addition, an embodiment of the present invention further provides a readable storage medium.

[0140] A meeting record generation program is stored on the readable storage medium. When the meeting record generation program is executed by a processor, the steps of the above-mentioned meeting record generating method are implemented.

[0141] The readable storage medium of the present invention may be a computer-readable storage medium. The specific implementation manner thereof is basically the same as that of the embodiments of the above-mentioned meeting record generating method, and will not be described in detail herein.

[0142] The embodiments of the present invention are described above with reference to the drawings. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, all fall within the protection scope of the present invention.

Claims

1. A method for generating meeting minutes, characterized in that, The method for generating meeting minutes is applied to a meeting system, which is formed based on a plurality of mobile terminals added to the meeting. The method for generating meeting minutes includes the following steps: Receiving voice signals collected by each mobile terminal in the meeting; Respectively extracting the voiceprint information to be recognized from the voice signals collected by each mobile terminal, and judging whether each voiceprint information to be recognized is valid according to the voiceprint information associated with each mobile terminal; If each voiceprint information to be recognized is valid, determining that the corresponding mobile terminal matches the meeting participant, removing the noise formed by the meeting participants who do not match the mobile terminal from the voice signals according to the voiceprint information associated with each mobile terminal, and retaining the voice signals of the meeting participants who match the mobile terminal to update each voice signal; If there is invalid voiceprint information to be recognized among each voiceprint information to be recognized, searching for the target mobile terminal corresponding to the invalid voiceprint information to be recognized; Judging whether the target mobile terminal carries an authorization identifier. If it carries an authorization identifier, performing the step of removing the noise formed by the meeting participants who do not match the mobile terminal from the voice signals according to the voiceprint information associated with each mobile terminal and retaining the voice signals of the meeting participants who match the mobile terminal; If it does not carry an authorization identifier, removing the target mobile terminal from the meeting; Respectively recognizing each voice signal to obtain multiple pieces of text information; Based on a preset topic model, performing semantic recognition on the multiple pieces of text information, extracting the topic words of the text information, calculating the scores of each topic word, and taking the topic word with the largest score as the meeting topic information; Classifying the multiple pieces of text information according to the user identifiers corresponding to each voice signal to generate classified text information corresponding to each mobile terminal respectively; Based on a preset topic model, performing semantic recognition on each classified text information to generate user opinion information corresponding to each mobile terminal respectively; Generating meeting minutes from the multiple pieces of text information, each user opinion information, and the meeting topic information; 2. The method for generating a meeting record according to claim 1, wherein The step of generating meeting minutes from the multiple pieces of text information, each user opinion information, and the meeting topic information includes: Arranging the multiple pieces of text information according to the time information corresponding to the multiple pieces of text information; Adding each user opinion information and the meeting topic information to the arranged multiple pieces of text information to generate meeting minutes; 3. The meeting record generation method according to claim 1, characterized in that, Before the step of receiving voice signals collected by each mobile terminal in the meeting, the method further includes: Collecting the voiceprint information of the holders of each mobile terminal based on each mobile terminal; Forming an association relationship between each mobile terminal and the voiceprint information of the holder of each mobile terminal, and adding each association relationship to the voiceprint database for storage; 4. The method for generating meeting minutes according to claim 1, characterized in that After the step of generating meeting minutes from the multiple pieces of text information, each user opinion information, and the meeting topic information, the method further includes: Sending the meeting minutes to each mobile terminal for display; 5. A meeting record generating device, characterized in that, The meeting minute generation device includes: A receiving module, configured to receive voice signals collected by each mobile terminal in a meeting; An extraction module, configured to respectively extract to-be-recognized voiceprint information from the voice signals collected by each of the mobile terminals, and determine whether each of the to-be-recognized voiceprint information is valid according to the voiceprint information associated with each of the mobile terminals; A filtering module, configured to, if each of the to-be-recognized voiceprint information is valid, determine that the corresponding mobile terminal matches the meeting participant, and remove the noise formed by the meeting participants who do not match the mobile terminal from the voice signals according to the voiceprint information associated with each of the mobile terminals, and retain the voice signals of the meeting participants who match the mobile terminal, so as to update each of the voice signals; A searching module, configured to, if there is invalid to-be-recognized voiceprint information among each of the to-be-recognized voiceprint information, search for a target mobile terminal corresponding to the invalid to-be-recognized voiceprint information; A judging module, configured to judge whether the target mobile terminal carries an authorization identifier, and if it carries the authorization identifier, execute the step of removing the noise formed by the meeting participants who do not match the mobile terminal from the voice signals according to the voiceprint information associated with each of the mobile terminals, and retaining the voice signals of the meeting participants who match the mobile terminal; A removing module, configured to, if it does not carry the authorization identifier, remove the target mobile terminal from the meeting; Respectively recognize each of the voice signals to obtain multiple pieces of text information; An identification module, the identification module includes: An identification unit, configured to perform semantic identification on the multiple pieces of text information based on a preset topic model, extract the topic words of the text information, calculate the scores of each topic word, and use the topic word with the largest score as the meeting topic information; A classification unit, configured to classify the multiple pieces of text information according to the user identifiers corresponding to each of the voice signals, and generate classified text information corresponding to each of the mobile terminals respectively; A generating unit, configured to perform semantic identification on each of the classified text information based on a preset topic model, and generate user opinion information corresponding to each of the mobile terminals respectively; A generating module, configured to generate the multiple pieces of text information, each of the user opinion information, and the meeting topic information into a meeting record.

6. A conference system, characterized in that, The meeting system is formed based on multiple mobile terminals added to the meeting. The meeting system includes a memory, a processor, and a meeting record generation program stored on the memory and executable on the processor. When the meeting record generation program is executed by the processor, the steps of the meeting record generation method according to any one of claims 1-4 are implemented.

7. A readable storage medium, characterized in that, The readable storage medium stores a meeting record generation program. When the meeting record generation program is executed by a processor, the steps of the meeting record generation method according to any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Conference voice data processing method and device, computer equipment and storage medium

    CN110322872A

  • Method for generating conference record automatically and apparatus thereof

    KR1020170126667A