Minute creation system
The system automates meeting minutes creation by eliminating the need for advance voiceprint registration, enabling accurate and efficient speaker identification and minutes production.
Patent Information
- Application Number
- JP2024047816
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-10-07
AI Technical Summary
Existing minutes-taking devices require advance registration of attendee voice features for speaker identification, which is cumbersome and limits flexibility.
A system that includes a speaker identification memory unit, record memory unit, transcription unit, speaker identification unit, and update unit, allowing for real-time speaker identification and elimination of the need for pre-registration of voiceprints.
Enables accurate and efficient creation of meeting minutes without prior voiceprint registration, reducing effort and ensuring consistent, high-quality minutes distribution.
Smart Images

Figure 2025147527000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system for creating minutes of a meeting. [Background technology]
[0002] Patent Document 1 discloses a minutes-taking device that includes a speech recognition device that extracts words from a speech signal based on its feature parameters and identifies the speaker by comparing the words with the speech feature parameters. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-open No. 2-206825 Summary of the Invention [Problem to be solved by the invention]
[0004] The minutes creating device of Patent Document 1 needs to register attendee voice feature parameters in advance in an attendee voice registration device in order to identify the speaker of the voice signal. The present invention aims to solve such problems, for example. [Means for solving the problem]
[0005] The minutes-taking system includes a speaker identification memory unit that stores information for identifying the speaker of a statement, a record memory unit that stores information representing audio recorded at a conference, a transcription unit that transcribes the statements made at the conference based on the information stored in the record memory unit, a speaker identification unit that identifies the speaker of the audio represented by the information stored in the record memory unit based on the information stored in the speaker identification memory unit, a minutes-taking unit that creates minutes in which the speaker of a statement transcribed by the transcription unit is linked to the speaker of the statement based on the speaker identified by the speaker identification unit, a record playback unit that plays back the audio based on the information stored in the record memory unit, a speaker acquisition unit that acquires information about the speaker of the audio from the audio played back by the record playback unit, and a speaker identification update unit that updates the information stored in the speaker identification memory unit based on the information acquired by the speaker acquisition unit. [Effects of the Invention]
[0006] Based on the information stored in the recording storage unit, the recording / playback unit plays back the audio, the speaker acquisition unit acquires information about the speaker of the audio, and the speaker identification update unit updates the information stored in the speaker identification storage unit, so the speaker identification unit can identify the speech even if the speaker's information is not stored in the speaker identification storage unit before the conference begins. This reduces the effort required for advance voiceprint registration, etc. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a block diagram showing an example of a minutes-taking system. DETAILED DESCRIPTION OF THE INVENTION
[0008] A minutes-taking system 10 will be described with reference to FIG. The minutes creation system 10 is a system that automatically creates minutes of a meeting based on audio and video recordings of the meeting. The minutes taking system 10 is, for example, a single computer or multiple computers. The multiple computers may be connected via a network such as a LAN (narrow area network) or the Internet. The computers may also be virtually configured on the cloud. The minutes taking system 10 is configured, for example, by a computer executing a computer program to realize the functional blocks described below.
[0009] The minutes creation system 10 includes, for example, a record acquisition unit 12, a record storage unit 14, a transcription unit 15, a speaker identification storage unit 24, a speaker identification unit 25, an agenda acquisition unit 32, an agenda storage unit 34, a remark classification unit 35, a minutes creation unit 45, a minutes storage unit 46, a minutes output unit 47, a recording / playback unit 51, a speaker acquisition unit 52, and a speaker identification update unit 55.
[0010] The record acquisition unit 12 acquires a record of a meeting for which minutes are to be created. The record acquired by the record acquisition unit 12 includes at least information representing the audio recorded from the meeting, and may also include information representing video footage of the meeting. The record acquisition unit 12 may acquire audio and video in real time from a microphone or video camera during the meeting, or may acquire audio and video recorded by a microphone or video camera after the meeting has ended.
[0011] The record storage unit 14 stores the records acquired by the record acquisition unit 12.
[0012] The transcription unit 15 transcribes statements made in a meeting based on the records stored in the record storage unit 14. The transcription unit 15 transcribes statements made in a meeting, for example, by performing speech recognition on the audio recorded during the meeting. The transcription unit 15 generates information that links text data representing the content of the transcribed statements with data representing the start and end times of the statements, for example. The transcription unit 15 may perform transcription using a generation AI (artificial intelligence) such as ChatGPT (registered trademark). The transcription unit 15 performs transcription by, for example, generating a prompt that instructs the generation AI to transcribe and instructing the generation AI using the generated prompt. The transcription unit 15 generates a prompt that instructs the generation AI to transcribe by selecting a prompt that matches the content of the meeting from a plurality of prompts that are pre-stored according to the purpose or situation, such as technical, sales, or manufacturing.
[0013] The speaker identification storage unit 24 stores information for identifying the speaker of a utterance. The information stored in the speaker identification storage unit 24 is used, for example, to determine, for each of one or more speakers, whether a certain voice is a utterance from that speaker. The information stored in the speaker identification storage unit 24 includes, for example, voiceprint information of the speaker, a learning model trained using speech of a known speaker as supervised data, and the like.
[0014] The speaker identification unit 25 identifies speakers in the records stored in the record storage unit 14 based on the information stored in the speaker identification storage unit 24. For example, the recorded conference is divided into predetermined periods (for example, 1 second), and for each divided period, the conference participants who spoke within that period are identified. If the speaker cannot be identified with 100% accuracy, the speaker may be identified with a probability, for example, a 60% probability for participant A and a 40% probability for participant B. Since multiple participants may be speaking simultaneously, the total probability may exceed 100%. The speaker identification unit 25 usually identifies speakers based on their voices. However, in addition to voices, a video of the meeting may also be used to supplement speaker identification. For example, in voice-based identification, if there is both a possibility that a speaker is participant A and a possibility that a speaker is participant B, and if the video shows that participant A's mouth is moving but participant B's mouth is not moving, then it can be assumed that the utterance was made by participant A. When making such a determination, the speaker identification storage unit 24 may store information about the speaker's appearance so that it can identify who is in the video. The speaker identification unit 25 may identify the speaker by using, for example, a generation AI, etc. The speaker identification unit 25, for example, generates a prompt that instructs the generation AI to identify the speaker, and identifies the speaker by instructing the generation AI using the generated prompt.
[0015] The agenda acquisition unit 32 acquires information representing the agenda of a meeting. An agenda is the content scheduled for a meeting. For example, the meeting organizer inputs the agenda, and the agenda acquisition unit 32 acquires the agenda. Alternatively, the agenda acquisition unit 32 may acquire a meeting notification and extract the agenda from the acquired meeting notification using, for example, a generation AI. For example, the agenda acquisition unit 32 generates a prompt that instructs the generation AI to extract the agenda from the meeting notification, and extracts the agenda by instructing the generation AI using the generated prompt. The agenda acquisition unit 32 usually acquires the agenda before the meeting starts, but it is not limited to this and may acquire the agenda after the meeting starts.
[0016] The agenda storage unit 34 stores the information acquired by the agenda acquisition unit 32.
[0017] The comment classification unit 35 classifies the comments transcribed by the transcription unit 15 based on the information stored in the agenda storage unit 34. For example, the comment classification unit 35 classifies the comments based on whether the comment is related to an agenda, or if there are multiple agendas, which agenda the comment is related to. The utterance classification unit 35 may classify utterances using, for example, a generation AI. The utterance classification unit 35, for example, generates a prompt that instructs the generation AI to classify utterances based on the agenda, and instructs the generation AI using the generated prompt, thereby classifying the utterances. The utterance classification unit 35 generates the prompt, for example, by applying the agenda stored in the agenda storage unit 34 to a template stored in advance.
[0018] The minutes creation unit 45 creates minutes based on the utterances transcribed by the transcription unit 15. The minutes creation unit 45 links speakers to utterances based on the speakers identified by the speaker identification unit 25. For example, for each utterance transcribed by the transcription unit 15, the unit 45 determines the speaker who is most likely to have spoken during the period from the start time to the end time of the utterance, and links the determined speaker as the speaker of that utterance. Furthermore, the minutes creation unit 45 separates the comments by agenda based on the classification by the comment classification unit 35. For example, the background color of the comments displayed in the minutes may be different depending on the agenda. Alternatively, instead of arranging all the comments in chronological order, the comments may be arranged by agenda. This allows you to create minutes that show at a glance which agenda each statement relates to and who made the statement. The minutes creation unit 45 may create the minutes, for example, by using a generation AI etc. The minutes creation unit 45, for example, creates a prompt that instructs the generation AI to create the minutes, and creates the minutes by instructing the generation AI using the generated prompt.
[0019] The minutes storage unit 46 stores information representing the minutes created by the minutes creating unit 45.
[0020] The minutes output unit 47 outputs the minutes created by the minutes creation unit 45 based on the information stored in the minutes storage unit 46. For example, the minutes are sent by email to the participants and related parties of the meeting.
[0021] The recording and playback unit 51 plays back the audio and video of the conference based on the records stored in the record storage unit 14. Note that the recording and playback unit 51 may also display the remarks transcribed by the transcription unit 15 and information about the speakers identified by the speaker identification unit 25, for example, as subtitles superimposed on the video of the conference.
[0022] The speaker acquisition unit 52 acquires information about the speaker of the audio being played back by the recording and playback unit 51. For example, the organizer or participants of the conference listen to the audio or video played back by the recording and playback unit 51, identify the speaker who is speaking at that time, and input this information to the speaker acquisition unit 52, which then acquires the speaker of the audio based on the input information.
[0023] The speaker identification update unit 55 updates the information stored in the speaker identification storage unit 24 based on the information acquired by the speaker acquisition unit 52. For example, the update unit 55 cuts out a portion of speech by a speaker acquired by the speaker acquisition unit 52 from the speech stored in the record storage unit 14, extracts voiceprint information from the cut-out speech, and registers it in the speaker identification storage unit 24. Alternatively, the cut-out speech is used as supervised data and trained in a learning model stored in the speaker identification storage unit 24.
[0024] The accuracy of identification by the speaker identification unit 25 depends on the information stored in the speaker identification storage unit 24. In particular, speakers whose information is not stored in the speaker identification storage unit 24 cannot be identified. The speaker identification update unit 55 updates the information stored in the speaker identification storage unit 24 based on the conference records stored in the record storage unit 14, so that even for speakers such as new participants whose information is not stored in the speaker identification storage unit 24 before the conference begins, the speaker identification update unit 55 updates the information, allowing the speaker identification unit 25 to identify their comments. This eliminates the need to register voiceprints in advance.
[0025] The above-described embodiment is an example for facilitating understanding of the present invention. The present invention is not limited thereto, and includes various modifications, changes, additions, or omissions without departing from the scope defined by the appended claims. This can be easily understood by those skilled in the art from the above description.
[0026] Creating minutes is an essential task for disseminating information about meetings, recording the history of decisions, and sharing information with those absent from the meeting. The minutes-taking system reduces the effort required to take minutes, eliminates variations in the quality and accuracy of minutes due to lack of advance preparation or bias on the part of the creator, and enables minutes to be distributed quickly. By utilizing generation AI, the creation of meeting minutes can be automated. Based on standardized instructions, generation AI can quickly create minutes. As a result, it is possible to reduce the workload of minute writers while ensuring accurate minutes can be distributed quickly to relevant parties. It allows for efficient and accurate speaker identification and recording of speeches according to appropriate agenda items, which allows for accurate reminders to stakeholders and consistent minutes. It also allows for accurate speaker identification of transcriptions, definition of appropriate agenda items, and grouping according to each agenda item. It is possible to identify speakers in the same room. There is no need to have each invitee register their voiceprint in advance, and it is possible to know what each person said. This allows for accurate speaker-identified transcription, definition of appropriate topics, and grouping according to each topic, resulting in quick and accurate reminders to stakeholders and consistent minutes. By using generation AI that guarantees transparency and reliability, such as ChatGPT (registered trademark), it is possible to ensure the accuracy of the minutes generation AI and flexibly customize it to suit society. Speaker identification can be performed by the conference organizer when transcribing. This makes it possible to identify speakers in the same room while eliminating the need for each conference participant to register their voiceprint in advance. Specifically, when transcribing from real-time recorded video, the conference organizer can specify the name of the transcribed speaker in a timely manner. The associated speaker name and transcribed audio are stored in a database (speaker database), enabling automatic speaker identification by AI in subsequent conferences. This reduces the organizer's effort to explain the system to each new invitee and the invitee's effort to register their voiceprint. It also enables accurate speaker database identification as intended by the conference organizer. The meeting organizer can specify the meeting agenda in advance and create minutes that follow that agenda. Instructions for the generation AI can be generated based on a predetermined format, and the transcriptions created by the generation AI can be grouped in accordance with the topics specified in advance by the meeting organizer. As a result, it is possible to provide minutes that are consistent with the organizer's desired topics and the progress of the meeting. Based on the transcript data with speaker identification, the meeting is grouped according to the agenda. This allows the generation AI to quickly create minutes that follow the agenda while retaining speaker information. As a result, it becomes possible to accurately record speech history and resolutions, which is the main purpose of meeting minutes. By applying highly transparent AI that complies with international standards, highly reliable system operation will be possible. Accurate speaker identification allows for reliable storage of speech history. It is possible to create minutes of meetings according to a predefined agenda. It reduces the effort required to prepare minutes. AI makes it possible to quickly create minutes that accurately record what has been said and share them with relevant parties. This reduces variations in the quality and accuracy of minutes due to lack of advance preparation or bias on the part of the creator. [Explanation of symbols]
[0027] 10 Minutes creation system, 12 Record acquisition unit, 14 Record storage unit, 15 Transcription unit, 24 Speaker identification storage unit, 25 Speaker identification unit, 32 Agenda acquisition unit, 34 Agenda storage unit, 35 Speech classification unit, 45 Minutes creation unit, 46 Minutes storage unit, 47 Minutes output unit, 51 Recording and playback unit, 52 Speaker acquisition unit, 55 Speaker identification update unit.
Claims
1. a speaker identification storage unit that stores information for identifying a speaker of a utterance; a recording storage unit for storing information representing recorded audio of a conference; a transcription unit that transcribes statements made in the conference based on the information stored in the record storage unit; a speaker identification unit that identifies a speaker of a voice represented by the information stored in the record storage unit based on the information stored in the speaker identification storage unit; a minutes creation unit that creates minutes in which the utterances transcribed by the transcription unit are linked to the utterances, based on the utterances identified by the utterance identification unit; a recording and reproducing unit that reproduces the audio based on the information stored in the recording and reproducing unit; a speaker acquisition unit that acquires information about a speaker of the voice reproduced by the recording and reproduction unit; a speaker identification update unit that updates the information stored in the speaker identification storage unit based on the information acquired by the speaker acquisition unit; A minutes creation system that includes:
2. an agenda storage unit that stores information representing the agenda of the meeting; a comment classification unit that classifies the comments transcribed by the transcription unit based on the information stored in the agenda storage unit; Further provided with The minutes preparation department creating minutes of the meeting in accordance with the agenda based on the remarks classified by the remark classification unit; The minutes creation system according to claim 1.
3. The message classification unit generating a prompt that instructs the artificial intelligence generator to classify the utterances transcribed by the transcription unit based on the information stored in the agenda storage unit; classifying the utterance by instructing the generating artificial intelligence using the generated prompt; The minutes creation system according to claim 2.
Citation Information
Patent Citations
Device for preparing minutes
JP1990206825A