system

The system addresses inefficiencies in meeting management by converting audio to text, monitoring speaking time and topic consistency, and recording decisions, enhancing facilitation and decision-making.

JP2026063868APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Current meeting management systems, particularly for young facilitators, struggle to effectively manage time, prevent discussions from deviating from the topic, and record final decisions, leading to inefficiencies and potential delays in decision-making.

Method used

A system that acquires audio data, converts it to text, analyzes the text to monitor meeting progress, issues warnings for exceeding speaking time, detects inconsistent topics, and records final decisions, using a server and smart devices to facilitate efficient meeting management.

Benefits of technology

Enables efficient time management, prevents discussion overruns, and ensures clear recording of final decisions, improving meeting efficiency and effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063868000001_ABST
    Figure 2026063868000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for acquiring audio data, A means of converting acquired audio data into text data, A means of analyzing the converted text data to monitor the progress of the meeting, A means to monitor the presenter's speaking time and issue a warning if the set time limit is exceeded, A means of comparing the content of the discussion with the agenda and detecting topics that do not match, A means of issuing warnings based on detected mismatched topics, A means of recording the final decision and the decision-maker, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Currently, in many meetings, especially for young facilitators, it is difficult to effectively manage time and conduct discussions. Especially when there are many senior participants, it is difficult to appropriately pay attention and give guidance to overtime and deviation from the topic. As a result, the efficiency of the meeting decreases, and important decisions may be delayed.

Means for Solving the Problems

[0005] The present invention includes means for acquiring audio data, means for converting the acquired audio data into text data, and means for analyzing the converted text data to monitor the progress of the meeting. It also includes means for monitoring the speaking time of presenters and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting inconsistent topics, means for issuing a warning based on the detected inconsistent topics, and means for recording final decisions and the decision-makers. This enables efficient management of the meeting, prevents time overruns and digressions, and clearly records final decisions.

[0006] "Audio data" refers to data that represents audio signals, including speeches and conversations from meetings, in digital or analog format.

[0007] "Text data" refers to digital information that is obtained by converting audio data into written information, and is expressed in a human-readable form.

[0008] "Monitoring" refers to the act of observing and analyzing the progress of a meeting and the content of the discussions in real time.

[0009] "Monitoring presentation time" means measuring the elapsed time from the moment the presenter starts speaking and continuously measuring the elapsed time to ensure it stays within the set time limit.

[0010] "Issuing a warning" is the act of drawing attention through visual and auditory signals when a set condition (for example, exceeding the speaking time limit) is reached.

[0011] An "agenda" is a pre-set plan that includes the order of proceedings and topics to be discussed at a meeting.

[0012] "Discrepancies" refer to topics or subjects not listed on the agenda that are unexpected in relation to the progress of the meeting.

[0013] "Decisions made" are matters that were agreed upon as a result of discussions and consultations at a meeting and that need to be officially recorded.

[0014] The term "decision-maker" refers to the individual or position that has the final decision-making authority in a meeting. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the language used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a system for a meeting management system that analyzes audio data in real time to appropriately manage the speaking time of presenters and the progress of discussions.

[0037] Program Overview

[0038] The meeting management system will be implemented with the following program structure.

[0039] 1. Acquisition of audio data

[0040] The device records the meeting audio in real time and streams it to the server.

[0041] 2. Speech-to-text conversion

[0042] The server converts the received audio data into text data using a speech recognition engine.

[0043] 3. Monitoring the progress of the meeting

[0044] The server analyzes the converted text data, automatically recognizes the speaker, and monitors the progress of the meeting.

[0045] 4. Timekeeping

[0046] The server monitors the presenter's speaking time and generates a warning if the set time limit is exceeded.

[0047] The device displays an alert to the presenter and provides an audio notification.

[0048] 5. Detecting Derailments in the Discussion

[0049] The server compares the discussion content with the agenda and generates a warning if a topic that does not match is mentioned.

[0050] The device displays a warning to the facilitator and provides an audio notification.

[0051] 6. Records of final decisions and decision-makers

[0052] The server monitors and detects statements containing keywords such as "conclusion" and "decision" in real time, and records their content.

[0053] The device displays an alert prompting the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[0054] Specific example

[0055] Case Study 1: Timekeeping Alerts

[0056] 1. User A begins a 15-minute presentation.

[0057] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[0058] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[0059] 4. The device displays and announces to user A that "3 minutes remaining."

[0060] 5. User A completes the presentation within the allotted time.

[0061] Case Study 2: Detecting Derailments in Discussions

[0062] 1. The meeting agenda is set to "Progress report for Project A".

[0063] 2. User B begins talking at length about "Project B".

[0064] 3. The server analyzes the discussion content as text and detects any discrepancies with the agenda.

[0065] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[0066] 5. The device displays and provides an audio notification to the facilitator stating, "The topic has gone off track. Let's return to the agenda."

[0067] 6. The facilitator encourages the discussion to return to the agenda.

[0068] Case Study 3: Records of Final Decisions and Decision-makers

[0069] 1. During the meeting, comments containing keywords such as "conclusion" are made.

[0070] 2. The server detects the statement and records its content.

[0071] 3. The server sends an alert to the terminal prompting the final decision-maker to confirm.

[0072] 4. The terminal displays the message "Please enter the final decision-maker," and User C enters the decision-maker.

[0073] 5. The entered decision-maker information is sent to the server and recorded along with the decision.

[0074] As described above, this system offers the advantage of enabling junior facilitators to efficiently manage meetings and facilitate time management and agenda achievement.

[0075] The following describes the processing flow.

[0076] Step 1:

[0077] The user launches the meeting management application and clicks the "Start Meeting" button.

[0078] Step 2:

[0079] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[0080] Step 3:

[0081] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[0082] Step 4:

[0083] The device records the meeting audio in real time and streams the audio data to the server.

[0084] Step 5:

[0085] The server uses a speech recognition engine to convert the received audio data into text data.

[0086] Step 6:

[0087] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[0088] Step 7:

[0089] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[0090] Step 8:

[0091] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[0092] Step 9:

[0093] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[0094] Step 10:

[0095] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[0096] Step 11:

[0097] If a derailment is detected, the server generates a warning and sends it to the terminal.

[0098] Step 12:

[0099] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[0100] Step 13:

[0101] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[0102] Step 14:

[0103] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[0104] Step 15:

[0105] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[0106] Step 16:

[0107] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[0108] The above outlines the specific processing steps and details of the actions performed at each step. This system will enable junior facilitators to manage meetings efficiently and effectively.

[0109] (Example 1)

[0110] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0111] Traditional meeting management systems struggle to efficiently manage speaking time and facilitate discussions, leaving younger facilitators, in particular, lacking the tools to ensure smooth meetings. Specifically, they are unable to effectively perform real-time analysis of meeting audio data, monitor presenter speaking time, detect tangents, and record final decisions and decision-makers, resulting in a decline in meeting quality.

[0112] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0113] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected mismatched topics, means for detecting statements containing keywords such as "conclusion" or "decision" and recording their content, and means for prompting confirmation of decision-maker information. This makes it possible to appropriately manage the presenter's speaking time, detect digressions in the discussion, and ensure the smooth progress of the meeting.

[0114] "Means for acquiring audio data" refers to a device or function for recording audio during a meeting in real time and streaming it to a server.

[0115] "Means for converting audio data to text data" refers to a device or function that converts received audio data into corresponding text data using speech recognition technology.

[0116] "Means of analyzing converted text data to monitor the progress of a meeting" refers to a device or function that analyzes text data converted by speech recognition to understand the content of statements and speakers, and to monitor the progress of a meeting.

[0117] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a device or function that measures the elapsed time from the moment the presenter starts speaking and generates a warning if the set time limit is exceeded.

[0118] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a device or function that compares a pre-set agenda with the actual content of a discussion and detects when a topic not included in the agenda is mentioned.

[0119] "Means for issuing warnings based on detected mismatched topics" refers to a device or function that warns that the discussion is going off track when a topic that does not match the agenda is detected.

[0120] "Means for detecting and recording statements containing keywords such as 'conclusion' or 'decision'" refers to a device or function that detects statements containing specific keywords in real time from text data during a meeting and records the content of those statements in a database or similar.

[0121] "Means for prompting confirmation of decision-maker information" refers to a device or function that generates an alert prompting the user to confirm the decision-maker's information regarding the detected final decision, and prompts the user to input that information.

[0122] This invention relates to a system for a meeting management system that analyzes audio data in real time to appropriately manage presenter speaking time and facilitate discussion. This system is primarily implemented through the interaction of a server, terminals, and users.

[0123] First, the device records the meeting audio in real time using its built-in or external microphone as soon as the meeting starts. Standard audio formats (e.g., WAV, MP3) are used for recording. This recorded audio data is transmitted to the server via streaming at regular intervals (e.g., every second).

[0124] The server uses a storage service (e.g., Amazon S3, Google Cloud Storage) to temporarily store the received audio data. Next, the server uses a speech recognition engine, such as the Google Cloud Speech-to-Text API, to convert the audio data into corresponding text data. This converted text data is then stored in a database system (e.g., MySQL).

[0125] Furthermore, the server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the converted text data. This analysis monitors the progress of the meeting and identifies speakers. Identification is based on speaker-specific language patterns and vocal characteristics, and the results are stored in a database.

[0126] Regarding the management of presenter speaking time, the server records the meeting start time and presentation start time and sets a timer. It monitors speaking time, and if the set time is exceeded, the server automatically generates a warning. The server then sends the corresponding warning to the terminal, which alerts the user with a display and audio notification such as "3 minutes remaining."

[0127] For managing the progress of the discussion, the server compares the pre-configured agenda with the actual discussion content. If the server detects a discussion that does not match the agenda, it automatically generates a "discussion derailment" warning and sends it to the terminal. The terminal displays and provides an audio notification to the facilitator with a warning such as, "The topic has gone off track. Let's get back to the agenda."

[0128] To record final decisions, the server detects in real time any statements containing keywords such as "conclusion" or "decision" during the meeting. The detected statements are recorded in the database, and the server then sends an alert to the terminal prompting confirmation of the decision-maker's information. The terminal displays a message such as "Please enter the final decision-maker," and the user enters the decision-maker's information, which is then sent to the server and stored in the database.

[0129] This system enables efficient management of meeting time and agenda completion, and particularly allows junior facilitators to conduct meetings smoothly.

[0130] The following scenarios can be considered as concrete examples:

[0131] Example prompt: "Please describe the processing flow of a system that converts audio data into text data in real time and manages speaking time."

[0132] The above describes specific embodiments for carrying out the present invention.

[0133] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0134] Step 1: Acquire audio data

[0135] The device uses its microphone to record the meeting audio and starts recording in real time as soon as the meeting begins.

[0136] The device streams the recorded audio data to the server at regular time intervals (e.g., every second).

[0137] Input: Real-time audio data from the meeting

[0138] Output: Audio data to be streamed.

[0139] Step 2: Save the audio file.

[0140] The server temporarily stores the audio data received from the terminal in a storage service (e.g., Amazon S3, Google Cloud Storage).

[0141] Input: Streaming transmitted audio data

[0142] Output: Temporarily saved audio data

[0143] Step 3: Convert speech to text

[0144] The server uses the Google Cloud Speech-to-Text API to convert the temporarily stored audio data into text data.

[0145] Input: Temporarily stored audio data

[0146] Output: Converted text data

[0147] Step 4: Save text data

[0148] The server saves the converted text data to a database system (e.g., MySQL).

[0149] Input: Converted text data

[0150] Output: Text data stored in the database

[0151] Step 5: Analyze the progress of the meeting

[0152] The server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the stored text data.

[0153] Identify the progress of the meeting and the speakers.

[0154] Input: Text data stored in the database

[0155] Output: Analyzed meeting progress and speaker information

[0156] Step 6: Monitor speaking time

[0157] The server records the start time of the presentation and sets a timer.

[0158] The system monitors the presenter's speaking time in real time and generates a warning if the set time limit is exceeded.

[0159] Input: Analyzed meeting progress

[0160] Output: Warning based on excess time

[0161] Step 7: Warning Notification

[0162] The server sends the generated warning to the terminal.

[0163] The device will display a warning message as a pop-up and provide an audio notification.

[0164] Input: Generated warning

[0165] Output: Warning messages and audio notifications displayed on the device

[0166] Step 8: Detecting Derailments in the Discussion

[0167] The server compares the meeting discussion content with a pre-configured agenda.

[0168] A warning is generated if a topic unrelated to the agenda is mentioned.

[0169] Input: Analyzed meeting progress, agenda

[0170] Output: Warning based on mismatched topics

[0171] Step 9: Derailment warning notification

[0172] The server sends a derailment warning to the terminal.

[0173] The device displays a derailment warning to the facilitator and provides an audio notification.

[0174] Input: Generated derailment warning

[0175] Output: Derailment warning message and audio notification displayed on the device.

[0176] Step 10: Record and confirm decisions

[0177] The server detects in real time any statements containing keywords such as "conclusion" or "decision," and records their content.

[0178] The server sends an alert to the terminal prompting the decision-maker to confirm.

[0179] The user enters decision-maker information, and the terminal sends the entered information to the server.

[0180] Input: Analyzed text data, decision-maker information

[0181] Output: Recorded decisions and decision-maker information

[0182] The above outlines the specific processing steps of this system. This enables efficient real-time management and progress of meetings.

[0183] (Application Example 1)

[0184] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0185] Efficient progress management is required in discussions and meetings within factories, but preventing deviations from the discussion and managing the progress is difficult. Furthermore, managing presenters' time and meticulously recording important decisions are also critical challenges. Current meeting management systems cannot comprehensively address these issues, thus limiting their potential for improving factory productivity.

[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0187] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected topics that do not match, means for recording the final decisions and decision-makers, means for monitoring the progress of the meeting in real time and managing the progress based on a set process plan, means for monitoring the presenter's speaking time and issuing a warning a certain time before a set time, means for detecting deviations and issuing a warning if the detected content does not match the process plan, and means for installing the progress management system on a smartphone. This enables efficient progress management, time management, and complete recording of important decisions for discussions and meetings within the factory.

[0188] "Means for acquiring audio data" refers to equipment and systems for acquiring audio from meetings and discussions in real time and transmitting it to a server.

[0189] "Means for converting acquired audio data into text data" refers to software or hardware that uses speech recognition technology to convert audio data into text information.

[0190] "Means of analyzing converted text data to monitor the progress of a meeting" refers to algorithms and programs that analyze text data to identify speakers and monitor the progress of a meeting in real time.

[0191] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system or device that measures the elapsed time from the start of the presenter's speech and generates a warning and notification when the set time is exceeded.

[0192] "Means for comparing the content of discussions with the agenda and detecting inconsistent topics" refers to a text analysis system that analyzes the content of discussions based on the purpose and agenda of a meeting and automatically detects topics that do not match the agenda.

[0193] "Means of issuing warnings based on detected mismatched topics" refers to a warning system that notifies users when a discussion deviates from the agenda.

[0194] "Means for recording final decisions and decision-makers" refers to a data recording system for organizing and preserving information about the decisions made during a meeting and the individuals who made those decisions.

[0195] "A means of monitoring the progress of a meeting in real time and managing its progress based on a set schedule" refers to a system for tracking the progress of meetings and discussions in real time and managing their progress according to a pre-set schedule.

[0196] "A means of monitoring the presenter's speaking time and issuing a warning a certain amount of time before the set time" refers to a system that issues a warning in advance when the scheduled speaking time approaches.

[0197] "Means for detecting deviations and issuing warnings when detected content does not match the process plan" refers to a system that detects when meeting content deviates from a pre-set process plan and issues a warning.

[0198] "Methods for installing a progress management system on a smartphone" refers to providing application software for progress management in a form that runs on a smartphone, and the methods and procedures for installing it.

[0199] This invention provides a system for efficiently managing the progress of discussions and meetings held within a factory, managing time, detecting deviations from the discussion, and recording final decisions. To achieve this, it employs a configuration that combines a server, a speech recognition engine, a text analysis system, a timekeeping function, and a warning notification system.

[0200] Hardware and software configuration

[0201] Hardware:

[0202] Smartphone: Equipped with a microphone for acquiring voice and a display for displaying warnings.

[0203] Server: A computer system used for storing and processing data.

[0204] software:

[0205] Python: A programming language for building the system of the present invention.

[0206] SpeechRecognition: A library for converting speech data into text.

[0207] threading: A multithreaded library for timekeeping and warning generation.

[0208] Process Overview

[0209] The server includes the following measures:

[0210] 1. Means of acquiring audio data:

[0211] The smartphone's microphone captures the meeting's audio data in real time and sends it to the server. The server continuously receives and stores this audio data.

[0212] 2. Means for converting acquired audio data into text data:

[0213] The server converts the received audio data into text data using the SpeechRecognition library. This conversion process is performed automatically, and the generated text data is then analyzed.

[0214] 3. Means of analyzing converted text data to monitor the progress of the meeting:

[0215] The server analyzes text data, identifies speakers, and monitors the progress of the meeting in real time. This allows for verification that the discussion is proceeding according to plan.

[0216] 4. A means of monitoring the presenter's speaking time and issuing a warning if the set time limit is exceeded:

[0217] The system monitors the elapsed time from the speaker's start and uses a timer function to issue a warning if the set time is exceeded. A warning message is displayed on the smartphone's screen, and audio notifications are provided as needed.

[0218] 5. Means for comparing the content of the discussion with the agenda and detecting inconsistencies:

[0219] The server compares the text data with the agenda and automatically detects any mismatched topics. This prevents the discussion from going in an unintended direction.

[0220] 6. Means of issuing warnings based on detected mismatched topics:

[0221] If mismatched topics are detected, the server will alert the facilitator and prompt them to quickly return the discussion to the agenda.

[0222] 7. Means for recording final decisions and decision-makers:

[0223] The server monitors statements containing keywords such as "conclusion" or "decision" in real time, and records their content and the information of the decision-maker. This information is stored in a format that can be referenced later.

[0224] 8. Means for monitoring the progress of meetings in real time and managing their progress based on the set schedule:

[0225] Real-time progress management is performed based on the established project plan, and the progress of discussions is monitored to ensure that they proceed as planned.

[0226] 9. A means of monitoring the presenter's speaking time and issuing a warning a certain period of time before the set end time:

[0227] By issuing a warning a certain amount of time in advance as the scheduled speaking time approaches, this system helps speakers organize their information within the allotted time.

[0228] 10. Means for detecting derailments and issuing warnings when the detected content does not match the process plan:

[0229] If a statement is made that does not conform to the project plan, the server will automatically detect the derailment and issue a warning.

[0230] 11. How to install the progress management system on a smartphone:

[0231] We will install the progress management system on smartphones and provide it to users so that they can easily use it.

[0232] Specific example

[0233] For example, when discussions are held within a factory about introducing a new process, a progress management assistant app can be used to manage time, prevent discussions from going off track, and record important decisions in real time. This app is installed on a smartphone and can be used from anywhere within the factory.

[0234] Example of a prompt

[0235] Requirements for a real-time meeting management system:

[0236] A feature that acquires audio data in real time and converts it to text using a speech recognition engine.

[0237] A function to identify the speaker and monitor the progress.

[0238] A feature that monitors speaking time and generates a warning if the set time limit is exceeded.

[0239] A feature that compares the discussion content with the agenda and detects deviations from the topic.

[0240] A feature that records final decisions and the decision-makers in real time.

[0241] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0242] Step 1:

[0243] Acquisition of audio data

[0244] Input: Audio from meetings or discussions

[0245] Output: Audio data acquired in real time

[0246] Specific operation: The device (smartphone) acquires audio from meetings and discussions in real time via its microphone. The acquired audio data is sent directly to the server. During this process, the device maintains a continuous connection with the server to ensure a stable communication environment.

[0247] Step 2:

[0248] Text conversion of acquired audio data

[0249] Input: Acquired audio data

[0250] Output: Text data converted by speech recognition

[0251] Specific operation: The server uses the SpeechRecognition library to convert the acquired audio data into text data. In this process, the speech recognition engine analyzes the audio signal and generates the corresponding string. The converted text data is temporarily stored in the server's memory or database.

[0252] Step 3:

[0253] Monitoring the progress of the meeting

[0254] Input: Converted text data

[0255] Output: Record of progress, identification of speakers

[0256] Specific operation: The server analyzes the converted text data and monitors the progress of the meeting. Specifically, it automatically recognizes speakers from the content of their statements within the text data and records the presentation time and content of each speaker. This allows for real-time tracking of who said what and when.

[0257] Step 4:

[0258] Monitoring and warnings regarding the presenter's speaking time.

[0259] Input: Announcement start time and current time, set time limit

[0260] Output: Time-based warnings as needed

[0261] Specific operation: The server records the presenter's speaking start time and monitors the speaking time based on the set time limit. Using a timer function, it sends a warning to the terminal when the speaking time is a certain amount of time away from the set limit. The warning message is displayed on the terminal's screen and an audio notification is also provided.

[0262] Step 5:

[0263] Comparison of discussion content and agenda

[0264] Input: Parsed text data, pre-configured agenda

[0265] Output: Detection and recording of mismatched topics

[0266] Specific operation: The server compares the converted text data with a pre-configured agenda and automatically detects mismatched topics. A text analysis algorithm is used to determine whether specific keywords or phrases match the agenda. If mismatched topics are detected, their content is recorded, and an alert is issued to the facilitator.

[0267] Step 6:

[0268] Issuing warnings based on unrelated topics

[0269] Input: Detection results for mismatched topics

[0270] Output: Warning to the facilitator

[0271] Specific operation: When the server detects mismatched topics, it issues a warning to the facilitator. The warning is displayed as a text message on the terminal's screen, and an audio notification is also provided. This allows the facilitator to quickly understand that the discussion is deviating and take appropriate action.

[0272] Step 7:

[0273] Record of Final Decisions and Decision Makers

[0274] Input: Converted text data, specific keywords (such as "conclusion", "decision", etc.)

[0275] Output: Record of decisions and decision makers

[0276] Specific operation: The server monitors the converted text data in real time and detects statements containing keywords such as "conclusion" and "decision". The detected content is regarded as a decision item and the information is recorded. In addition, a display prompting the user to confirm the decision maker is shown on the terminal, and the input decision maker information is also sent to the server and recorded.

[0277] Step 8:

[0278] Real-time monitoring of meeting progress and progress management of project plan

[0279] Input: Converted text data, set project plan

[0280] Output: Progress management information

[0281] Specific operation: The server monitors the meeting progress in real time and manages the progress based on the set project plan. It checks whether the progress is proceeding as planned and issues warnings or alerts if necessary. This function is extremely important to ensure that the meeting proceeds as planned.

[0282] Step 9:

[0283] Pre-warning for the speaker's speaking time

[0284] Input: A certain time before the start time of the speech and the set time limit

[0285] Output: Pre-warning

[0286] Specific operation: When a certain period of time has elapsed since the start time of the speech, a pre-warning of the speech time is sent to the terminal. This allows the presenter to summarize the speech while being aware of the remaining time.

[0287] Step 10:

[0288] Detection of deviation from the process plan and sending of warnings

[0289] Input: Converted text data, process plan

[0290] Output: Deviation detection result and warning

[0291] Specific operation: Compare the converted text data with the process plan. If content that does not match is detected, the server detects the deviation and issues a warning. The warning is given through the display and voice notification of the terminal to enable the facilitator to respond promptly

[0292] Step 11:

[0293] Installation of the progress management system on the smartphone

[0294] Input: None (assuming pre-settings)

[0295] Output: Installed progress management system

[0296] Specific operation: Install the progress management system on the smartphone. The installation provides dedicated application software that allows the user to easily download and set up. With this function, the user can use the progress management system from anywhere within the factory.

[0297] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.

[0298] The present invention is a system that acquires voice data during the progress of a meeting and analyzes it in real time. By further combining an emotion engine, it recognizes the emotions of participants and is intended to conduct the meeting operation more smoothly.

[0299] Outline of the program

[0300] This meeting management system can be implemented with the following program configuration.

[0301] 1. Acquisition of voice data

[0302] The terminal records the voice of the meeting in real time and streams and transmits the voice data to the server.

[0303] 2. Text conversion of voice

[0304] The server converts the received voice data into text data using a voice recognition engine.

[0305] 3. Monitoring of the progress of the meeting

[0306] The server analyzes the converted text data, identifies the speaker, and annotates the data to monitor the progress of the meeting. <00009​​​​​​​​​​​​​​​​​​​​ The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[0313] The server generates a derailment detection warning and sends it to the terminal.

[0314] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[0315] 6. Emotion recognition by an emotion engine

[0316] The terminals and servers are equipped with emotion engines that recognize emotions in real time from the voice tone and content of what meeting participants say.

[0317] 7. Adjusting Emotion-Based Warnings

[0318] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[0319] 8. Emotion-based stress management

[0320] The emotion engine monitors the meeting facilitator's stress level in real time and provides alerts to reduce stress as needed.

[0321] 9. Records of final decisions and decision-makers

[0322] The server monitors keywords such as "conclusion" and "decision" in real time and records their content when detected.

[0323] The device displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's name.

[0324] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[0325] Specific example

[0326] Case Study 1: Timekeeping Alerts

[0327] 1. The user begins a 15-minute presentation.

[0328] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[0329] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[0330] 4. The device notifies the user that "3 minutes remaining."

[0331] Case Study 2: Detecting Derailments in Discussions

[0332] 1. The agenda is set to "Project A progress report".

[0333] 2. The user begins talking at length about "Project B".

[0334] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[0335] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[0336] 5. The device notifies the facilitator that "the topic has gone off track."

[0337] Case Study 3: Stress Management Using an Emotional Engine

[0338] 1. The facilitator begins to feel stressed while conducting the meeting.

[0339] 2. The emotion engine monitors the facilitator's stress level in real time.

[0340] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[0341] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[0342] The following describes the processing flow.

[0343] Step 1:

[0344] The user launches the meeting management application and clicks the "Start Meeting" button.

[0345] Step 2:

[0346] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[0347] Step 3:

[0348] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[0349] Step 4:

[0350] The device records the meeting audio in real time and streams the audio data to the server.

[0351] Step 5:

[0352] The server uses a speech recognition engine to convert the received audio data into text data.

[0353] Step 6:

[0354] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[0355] Step 7:

[0356] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[0357] Step 8:

[0358] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[0359] Step 9:

[0360] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[0361] Step 10:

[0362] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[0363] Step 11:

[0364] If a derailment is detected, the server generates a warning and sends it to the terminal.

[0365] Step 12:

[0366] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[0367] Step 13:

[0368] An emotion engine installed on the server analyzes the voice tone and content of meeting participants in real time to recognize their emotions.

[0369] Step 14:

[0370] The server adjusts the warning content and notification tone for presenters and participants based on the emotional state obtained from the emotion engine.

[0371] Step 15:

[0372] If the server, based on the results of its emotion engine's emotional state monitoring, determines that the facilitator's stress level is high, it generates an alert saying, "Let's take a short break," and sends it to the terminal.

[0373] Step 16:

[0374] The device provides the facilitator with an audio notification and a display alert saying, "We recommend you take a short break."

[0375] Step 17:

[0376] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[0377] Step 18:

[0378] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[0379] Step 19:

[0380] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[0381] Step 20:

[0382] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[0383] (Example 2)

[0384] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0385] Many challenges exist in managing the progress of meetings. Existing systems simply converted audio to text and monitored the agenda, resulting in insufficient management of speaking time, prevention of discussion digressions, and recognition of participants' emotions. Furthermore, stress management for meeting facilitators was not considered. A system that comprehensively addresses these challenges is needed.

[0386] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0387] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting mismatched topics, means for issuing a warning based on the detected mismatched topics, means for recognizing the emotions of participants and adjusting the content of warnings and notification tones based on their emotional state, means for monitoring the stress level of the meeting facilitator in real time and providing alerts for stress reduction as needed, and means for recording final decisions and the decision-makers. This makes it possible to comprehensively and in real time perform everything from acquiring audio data to text conversion, monitoring progress, timekeeping, detecting deviations from the discussion, emotion recognition, stress management, and recording final decisions.

[0388] "Audio data" refers to digital data in which sound is recorded electronically.

[0389] "Text data" refers to digital data obtained by converting audio data into written text.

[0390] "Progress status" refers to information indicating the current state of the meeting's progress.

[0391] "Speaking time" refers to the time elapsed since the presenter began speaking.

[0392] An "agenda" is a list of topics or subjects that should be discussed at a meeting.

[0393] A "topic" refers to a specific subject or theme that is discussed in a meeting.

[0394] A "warning" is a notice issued when a set condition is not met.

[0395] "Emotions" refers to the psychological and emotional state of the meeting participants.

[0396] "Emotional state" refers to the specific expression of a meeting participant's emotions at a particular moment.

[0397] "Notification tone" refers to the sound or tone used when issuing warnings or notifications.

[0398] "Stress level" indicates the degree of mental and physical burden on the meeting organizer.

[0399] An "alert" is a warning or notice that appears when certain conditions occur.

[0400] "Decisions made" refer to the specific details and action plans that were ultimately decided upon at the meeting.

[0401] A "decision-maker" is the person or position that ultimately makes the final decision on a particular matter in a meeting.

[0402] Modes for carrying out the invention

[0403] This invention is a system that acquires audio data during a meeting and analyzes it in real time. Furthermore, by combining it with an emotion engine, it recognizes the emotions of participants and facilitates smoother meeting management. This system includes the following components:

[0404] Acquisition of audio data

[0405] When a user starts a meeting, the device uses its microphone to record the meeting's audio data in real time and streams that audio data to the server. The audio data is processed in digital format and transferred to the server via the communication network.

[0406] Speech-to-text conversion

[0407] The server uses a speech recognition engine such as Google Cloud Speech-to-Text to convert the received audio data into text data. The converted text data is temporarily stored on the server and used for subsequent data analysis.

[0408] Monitoring the progress of the meeting

[0409] The server uses a natural language processing engine, such as the Google Cloud Natural Language API, to analyze the text data converted from speech and identify the speaker. It adds annotations to the text data, including the speaker's ID and the content of their speech, and monitors the progress of the meeting.

[0410] Timekeeping

[0411] The server records the start time of each user's speech and monitors the speech duration in real time. Once the set time has elapsed, the server generates an alert and sends it to the terminal. The terminal then notifies the presenter of the remaining time via audio notification and a display alert.

[0412] Derailment detection in discussions

[0413] The server compares the converted text data with pre-configured agenda items and recognizes any discussion of a mismatched topic as a digression. When a digression is detected, the server generates a warning and sends it to the terminal. The terminal then notifies the facilitator with an audio and visual alert stating, "The topic has gone off-topic."

[0414] Emotion recognition by an emotion engine

[0415] The terminals and servers are equipped with emotion engines such as Azure Cognitive Services Emotion API to analyze the voice tone and content of meeting participants. This makes it possible to recognize participants' emotions (e.g., joy, anger, surprise) in real time.

[0416] Adjusting emotion-based warnings

[0417] The server receives the emotional state recognized by the emotion engine and adjusts the warning content and notification tone accordingly. For example, if a participant is feeling stressed, the alert tone will be softened and the notification content will be made gentler.

[0418] Emotion-based stress management

[0419] The emotion engine monitors the meeting facilitator's stress level in real time and provides this information to the server. If the server determines that the facilitator's stress level is high, it generates alerts for breaks or urgent discussions. For example, it might send a message to the facilitator such as, "We recommend you take a short break."

[0420] Records of final decisions and decision-makers

[0421] The server monitors keywords such as "conclusion" and "decision" in the conversation in real time. When a keyword is detected, its content is immediately recorded. The terminal displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's information. The entered decision-maker information is sent from the terminal to the server and recorded along with the final decision.

[0422] Specific example

[0423] Case Study 1: Timekeeping Alerts

[0424] 1. The user begins a 15-minute presentation.

[0425] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[0426] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[0427] 4. The device notifies the user that "3 minutes remaining."

[0428] Prompt example

[0429] "Design a system that records the start time of a 15-minute presentation and generates a '3 minutes remaining' alert 12 minutes later."

[0430] Case Study 2: Detecting Derailments in Discussions

[0431] 1. The agenda is set to "Project progress report".

[0432] 2. The user starts talking at length about "another project."

[0433] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[0434] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[0435] 5. The device notifies the facilitator that "the topic has gone off track."

[0436] Prompt example

[0437] "Design a system that detects off-topic comments during meetings and alerts the facilitator to the derailment."

[0438] Case Study 3: Stress Management Using an Emotional Engine

[0439] 1. The facilitator begins to feel stressed while conducting the meeting.

[0440] 2. The emotion engine monitors the facilitator's stress level in real time.

[0441] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[0442] Prompt example

[0443] "Design a system that monitors the facilitator's stress level during meetings and sends notifications prompting them to take breaks when necessary."

[0444] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[0445] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0446] Step 1:

[0447] The user starts the meeting.

[0448] Specifically, the user launches the meeting application and starts the meeting by clicking the "Start Meeting" button.

[0449] Input: User's meeting start operation

[0450] Output: Meeting start signal

[0451] Step 2:

[0452] The device prepares to record audio.

[0453] Specifically, the device checks the readiness status of the built-in or externally connected microphone and enables the recording function.

[0454] Input: Signal to start meeting

[0455] Output: Signal indicating ready to record

[0456] Step 3:

[0457] The device records audio data in real time and streams it to the server.

[0458] Specifically, the device records the audio during the meeting in real time and streams it to a server via the internet as digital audio data.

[0459] Input: Signal indicating recording is ready.

[0460] Output: Streaming of real-time audio data

[0461] Step 4:

[0462] The server receives the audio data and uses a speech recognition engine to convert the audio data into text data.

[0463] In terms of specific operations, the server uses a speech recognition engine to analyze the audio data and convert it into text data written as characters. A typical example of the software used is a general-purpose speech recognition engine.

[0464] Input: Real-time audio data

[0465] Output: Text data

[0466] Step 5:

[0467] The server temporarily stores the converted text data.

[0468] Specifically, the server temporarily stores the converted text data in storage such as a database.

[0469] Input: Text data

[0470] Output: Temporarily saved text data

[0471] Step 6:

[0472] The server analyzes the text data converted from the speech and monitors the progress using a natural language processing engine.

[0473] In terms of specific operations, the server uses a natural language processing engine to analyze text data, identify the speaker, and annotate the content of the speech. The software used includes a natural language processing engine, among other things.

[0474] Input: Temporarily saved text data

[0475] Output: Annotated text data

[0476] Step 7:

[0477] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[0478] Specifically, the server records the start time of a message and then continuously measures the duration of subsequent messages.

[0479] Input: Annotated text data

[0480] Output: Speech time elapsed data

[0481] Step 8:

[0482] After a set amount of time has elapsed since the start of the announcement, the server generates an alert and sends it to the terminal.

[0483] Specifically, when the server determines that a set time has elapsed, it generates an alert message and sends it to the terminal.

[0484] Input: Elapsed time of speech

[0485] Output: Alert

[0486] Step 9:

[0487] The device will notify the presenter of the remaining time via audio notification and a display alert.

[0488] Specifically, the device will use voice notifications and screen displays based on the received alerts to inform the presenter of the remaining time.

[0489] Input: Alert

[0490] Output: Voice notifications and display alerts

[0491] Step 10:

[0492] The server compares the converted text data with the agenda and detects any mismatched topics.

[0493] Specifically, the server compares the text data with a pre-configured agenda list and automatically detects any mismatches.

[0494] Input: Annotated text data

[0495] Output: Detection results for mismatched topics

[0496] Step 11:

[0497] The server issues a warning based on mismatched topics and sends it to the terminal.

[0498] Specifically, the server generates a warning message based on the detected mismatched topics and sends it to the terminal.

[0499] Input: Detection results for mismatched topics

[0500] Output: Warning message

[0501] Step 12:

[0502] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[0503] Specifically, the device uses the received warning message to provide audio notifications and screen displays, informing the facilitator that the topic has gone off track.

[0504] Input: Warning message

[0505] Output: Voice notifications and display alerts

[0506] Step 13:

[0507] The terminals and servers are equipped with an emotion engine to recognize the emotions of the participants.

[0508] Specifically, the terminal and server use an emotion engine to analyze the participant's emotions in real time based on their voice tone and what they say.

[0509] Input: Real-time audio data

[0510] Output: Emotion recognition result

[0511] Step 14:

[0512] The server adjusts warning content and notification tone based on the emotional state recognized by the emotion engine.

[0513] Specifically, the server adjusts the content and tone of the warning message based on the emotion recognition results.

[0514] Input: Sentiment recognition result

[0515] Output: Adjusted warning message

[0516] Step 15:

[0517] The emotion engine monitors the stress levels of meeting organizers, and the server provides alerts to reduce stress as needed.

[0518] Specifically, the emotion engine constantly analyzes the operator's stress level and generates alerts for breaks and relaxation at appropriate times.

[0519] Input: Real-time audio data and operator sentiment data

[0520] Output: Stress reduction alert

[0521] Step 16:

[0522] The server monitors for "conclusion" and "decision" keywords and records their content.

[0523] Specifically, the server monitors for specific keywords within the converted text data in real time and records their content in the database if detected.

[0524] Input: Annotated text data

[0525] Output: Recorded decisions

[0526] Step 17:

[0527] The terminal prompts the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[0528] Specifically, after a decision is finalized, the terminal displays a pop-up notification prompting the user to enter the decision-maker's information. The entered information is sent to the server and recorded.

[0529] Input: Decision made and user-inputted decision-maker information

[0530] Output: Recorded decision-maker information

[0531] (Application Example 2)

[0532] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0533] Meetings during shift changes in factories require smooth information transfer and efficient discussion. However, insufficient communication on the factory floor and frequent digressions in discussions often lead to problems with efficient shift changes and work handovers. Furthermore, the emotional state of participants significantly impacts the progress of the meeting, making emotional management a crucial issue. A system is needed to appropriately address these problems.

[0534] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting inconsistent topics, means for issuing a warning based on the detected inconsistent topics, and means for recognizing the emotions of meeting participants and adjusting the content of the warning and the tone of the notification based on their emotional state. This enables appropriate information transmission, efficient progress of discussions, and emotional management of participants in factory shift change meetings.

[0535] "Means for acquiring audio data" refers to a device or software for recording speech spoken in real time during meetings or shift changes and saving it as digital data.

[0536] "Means for converting acquired audio data into text data" refers to a speech recognition engine or software that analyzes recorded audio data and automatically converts it into text data.

[0537] "Means of analyzing converted text data to monitor the progress of a meeting" refers to a program or system that analyzes text data to understand who is speaking and the progress of the agenda in real time.

[0538] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system with a timekeeping function that records the start time of the presentation and generates an alert when the set time limit is exceeded.

[0539] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a program or algorithm that compares the content of a statement with a set agenda and detects when a topic that does not match the agenda is mentioned.

[0540] "Means for issuing warnings based on detected mismatched topics" refers to a system that automatically generates warnings related to detected mismatched topics and notifies speakers and moderators.

[0541] "Means for recording final decisions and decision-makers" refers to a system that has the function of recording decisions made during meetings or shift changes, as well as the decision-makers involved, and saving them in a database for later reference.

[0542] "Means for recognizing the emotions of meeting participants and adjusting the tone of warnings and notifications based on their emotional state" refers to an algorithm or engine that analyzes emotions in real time from the tone of voice and content of statements of participants, and changes the content and expression of warnings and notifications according to their emotional state.

[0543] The system that implements this application supports meetings during shift changes in a factory. This system acquires audio data in real time, converts it into text data, and performs meeting progress monitoring, timekeeping, discussion derailment detection, and sentiment recognition.

[0544] Specifically, the program will be implemented with the following structure:

[0545] Hardware and software

[0546] Hardware to use:

[0547] Microphone: For capturing the voices of meeting participants.

[0548] Robot or smartphone: A platform for processing voice data and sending notifications.

[0549] Software to use:

[0550] speech_recognition: A library for performing speech recognition.

[0551] emotion_recognition: A library for performing emotion recognition (uses a pre-trained model).

[0552] datetime: A standard library for time management.

[0553] Process Overview

[0554] 1. Acquisition of audio data:

[0555] The audio of meeting participants can be captured in real time via microphones.

[0556] The voice data is transferred to a robot or smartphone platform.

[0557] 2. Speech-to-text conversion:

[0558] The speech_recognition library installed on the robot or smartphone is used to convert speech data into text data.

[0559] The converted text is sent to the server.

[0560] 3. Monitoring the progress of the meeting:

[0561] The server analyzes the converted text data to identify the speaker and monitor the progress.

[0562] Verify that the meeting is proceeding according to the set agenda.

[0563] 4. Timekeeping:

[0564] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[0565] If the presentation time exceeds the set time, a warning is sent from the server to the robot or smartphone, and the presenter is notified via audio notification and a display alert.

[0566] 5. Detecting deviations from the discussion:

[0567] The server compares the discussion to the agenda and detects deviations if a topic that does not match is mentioned.

[0568] If a derailment is detected, a warning will be displayed to the facilitator.

[0569] 6. Emotion recognition by the emotion engine:

[0570] The robot or smartphone is equipped with an emotion engine that recognizes the participant's emotions in real time based on their voice tone and what they say.

[0571] This allows for monitoring the emotional state of participants and taking appropriate action.

[0572] 7. Adjusting emotion-based warnings:

[0573] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[0574] For example, if emotions are heightened, the notification will be delivered in a calm tone.

[0575] Specific example

[0576] Examples of support for shift change meetings

[0577] 1. Start of the meeting:

[0578] When the shift changes, the leader begins explaining the situation to the next team.

[0579] The microphone acquires audio data in real time and transmits it to the robot or smartphone.

[0580] 2. Speech-to-text conversion:

[0581] The acquired audio data is converted to text using the speech_recognition library and sent to the server.

[0582] 3. Monitoring the progress of the meeting and keeping time:

[0583] The server analyzes the converted text data to identify the speaker, monitor the progress, and keep time.

[0584] After a set amount of time has elapsed since the announcement began, a warning will be sent to the robot or smartphone, triggering an audio notification and a visual alert.

[0585] 4. Detecting deviations from the discussion:

[0586] The server compares the discussion content with the agenda and warns the facilitator if it detects a deviation from the topic.

[0587] 5. Recognition by the emotion engine:

[0588] It monitors future emotional states in real time and adjusts warning content and notification tones as needed.

[0589] Examples of prompt statements to use

[0590] "Record the 5-minute shift change meeting and convert it to text. Perform sentiment analysis using an emotion recognition model, and implement tactical detection and timekeeping."

[0591] This system allows factory shift change meetings to proceed smoothly and efficiently, and because it takes into account the emotional state of the participants, it enables more effective information sharing and discussion.

[0592] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0593] Step 1:

[0594] Acquire audio data.

[0595] A microphone connected to a terminal (robot or smartphone) records the speech of meeting participants in real time. The input is the raw voice of the meeting participants, and the output is digitized audio data. This audio data is streamed to a server.

[0596] Step 2:

[0597] Convert audio data to text data.

[0598] The audio data sent to the server is automatically converted into text data using the speech_recognition library. The input is audio data, and the output is text data in string format. This text data is then passed on to the next parsing step.

[0599] Step 3:

[0600] Monitor the progress of the meeting.

[0601] The server analyzes text data to identify the content and speaker of a statement. The input is text data, and the output is a pair of speaker and statement information. Based on this, it identifies speakers and generates progress graphs.

[0602] Step 4:

[0603] Timekeeping will be performed.

[0604] The server records the presenter's start time and measures the elapsed time in real time. When the set presentation time has elapsed, it generates a warning and sends it to the terminal. The input is the start time, and the output is the elapsed time and the warning message. The terminal notifies the presenter with an audio notification and a display alert.

[0605] Step 5:

[0606] Detects deviations from the discussion.

[0607] The server compares the text data with the previously set agenda and detects any mismatched topics. The input is the text data and agenda information, and the output is a list of mismatched topics. If a digression is detected, the server sends a warning to the facilitator.

[0608] Step 6:

[0609] Recognize your emotions.

[0610] The emotion engine built into the device recognizes the participant's emotions in real time based on their voice tone and speech content. Input consists of voice and text data, and output is the identified emotional state. This information is sent to the server.

[0611] Step 7:

[0612] The tone of warnings and notifications is adjusted based on emotions.

[0613] The server dynamically adjusts the content of warnings and notification tones based on the emotional state obtained from the emotion engine. The input is emotional state data, and the output is the adjusted warning message. For example, if a participant is stressed, the notification will be delivered in a calm tone.

[0614] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0615] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0616] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0617] [Second Embodiment]

[0618] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0619] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0620] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0621] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0622] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0623] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0624] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0625] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0626] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0627] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0628] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0629] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0630] This invention is a system for a meeting management system that analyzes audio data in real time to appropriately manage the speaking time of presenters and the progress of discussions.

[0631] Program Overview

[0632] The meeting management system will be implemented with the following program structure.

[0633] 1. Acquisition of audio data

[0634] The device records the meeting audio in real time and streams it to the server.

[0635] 2. Speech-to-text conversion

[0636] The server converts the received audio data into text data using a speech recognition engine.

[0637] 3. Monitoring the progress of the meeting

[0638] The server analyzes the converted text data, automatically recognizes the speaker, and monitors the progress of the meeting.

[0639] 4. Timekeeping

[0640] The server monitors the presenter's speaking time and generates a warning if the set time limit is exceeded.

[0641] The device displays an alert to the presenter and provides an audio notification.

[0642] 5. Detecting Derailments in the Discussion

[0643] The server compares the discussion content with the agenda and generates a warning if a topic that does not match is mentioned.

[0644] The device displays a warning to the facilitator and provides an audio notification.

[0645] 6. Records of final decisions and decision-makers

[0646] The server monitors and detects statements containing keywords such as "conclusion" and "decision" in real time, and records their content.

[0647] The device displays an alert prompting the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[0648] Specific example

[0649] Case Study 1: Timekeeping Alerts

[0650] 1. User A begins a 15-minute presentation.

[0651] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[0652] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[0653] 4. The device displays and announces to user A that "3 minutes remaining."

[0654] 5. User A completes the presentation within the allotted time.

[0655] Case Study 2: Detecting Derailments in Discussions

[0656] 1. The meeting agenda is set to "Progress report for Project A".

[0657] 2. User B begins talking at length about "Project B".

[0658] 3. The server analyzes the discussion content as text and detects any discrepancies with the agenda.

[0659] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[0660] 5. The device displays and provides an audio notification to the facilitator stating, "The topic has gone off track. Let's return to the agenda."

[0661] 6. The facilitator encourages the discussion to return to the agenda.

[0662] Case Study 3: Records of Final Decisions and Decision-makers

[0663] 1. During the meeting, comments containing keywords such as "conclusion" are made.

[0664] 2. The server detects the statement and records its content.

[0665] 3. The server sends an alert to the terminal prompting the final decision-maker to confirm.

[0666] 4. The terminal displays the message "Please enter the final decision-maker," and User C enters the decision-maker.

[0667] 5. The entered decision-maker information is sent to the server and recorded along with the decision.

[0668] As described above, this system offers the advantage of enabling junior facilitators to efficiently manage meetings and facilitate time management and agenda achievement.

[0669] The following describes the processing flow.

[0670] Step 1:

[0671] The user launches the meeting management application and clicks the "Start Meeting" button.

[0672] Step 2:

[0673] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[0674] Step 3:

[0675] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[0676] Step 4:

[0677] The device records the meeting audio in real time and streams the audio data to the server.

[0678] Step 5:

[0679] The server uses a speech recognition engine to convert the received audio data into text data.

[0680] Step 6:

[0681] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[0682] Step 7:

[0683] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[0684] Step 8:

[0685] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[0686] Step 9:

[0687] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[0688] Step 10:

[0689] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[0690] Step 11:

[0691] If a derailment is detected, the server generates a warning and sends it to the terminal.

[0692] Step 12:

[0693] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[0694] Step 13:

[0695] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[0696] Step 14:

[0697] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[0698] Step 15:

[0699] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[0700] Step 16:

[0701] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[0702] The above outlines the specific processing steps and details of the actions performed at each step. This system will enable junior facilitators to manage meetings efficiently and effectively.

[0703] (Example 1)

[0704] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0705] Traditional meeting management systems struggle to efficiently manage speaking time and facilitate discussions, leaving younger facilitators, in particular, lacking the tools to ensure smooth meetings. Specifically, they are unable to effectively perform real-time analysis of meeting audio data, monitor presenter speaking time, detect tangents, and record final decisions and decision-makers, resulting in a decline in meeting quality.

[0706] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0707] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected mismatched topics, means for detecting statements containing keywords such as "conclusion" or "decision" and recording their content, and means for prompting confirmation of decision-maker information. This makes it possible to appropriately manage the presenter's speaking time, detect digressions in the discussion, and ensure the smooth progress of the meeting.

[0708] "Means for acquiring audio data" refers to a device or function for recording audio during a meeting in real time and streaming it to a server.

[0709] "Means for converting audio data to text data" refers to a device or function that converts received audio data into corresponding text data using speech recognition technology.

[0710] "Means of analyzing converted text data to monitor the progress of a meeting" refers to a device or function that analyzes text data converted by speech recognition to understand the content of statements and speakers, and to monitor the progress of a meeting.

[0711] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a device or function that measures the elapsed time from the moment the presenter starts speaking and generates a warning if the set time limit is exceeded.

[0712] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a device or function that compares a pre-set agenda with the actual content of a discussion and detects when a topic not included in the agenda is mentioned.

[0713] "Means for issuing warnings based on detected mismatched topics" refers to a device or function that warns that the discussion is going off track when a topic that does not match the agenda is detected.

[0714] "Means for detecting and recording statements containing keywords such as 'conclusion' or 'decision'" refers to a device or function that detects statements containing specific keywords in real time from text data during a meeting and records the content of those statements in a database or similar.

[0715] "Means for prompting confirmation of decision-maker information" refers to a device or function that generates an alert prompting the user to confirm the decision-maker's information regarding the detected final decision, and prompts the user to input that information.

[0716] This invention relates to a system for a meeting management system that analyzes audio data in real time to appropriately manage presenter speaking time and facilitate discussion. This system is primarily implemented through the interaction of a server, terminals, and users.

[0717] First, the device records the meeting audio in real time using its built-in or external microphone as soon as the meeting starts. Standard audio formats (e.g., WAV, MP3) are used for recording. This recorded audio data is transmitted to the server via streaming at regular intervals (e.g., every second).

[0718] The server uses a storage service (e.g., Amazon S3, Google Cloud Storage) to temporarily store the received audio data. Next, the server uses a speech recognition engine, such as the Google Cloud Speech-to-Text API, to convert the audio data into corresponding text data. This converted text data is then stored in a database system (e.g., MySQL).

[0719] Furthermore, the server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the converted text data. This analysis monitors the progress of the meeting and identifies speakers. Identification is based on speaker-specific language patterns and vocal characteristics, and the results are stored in a database.

[0720] Regarding the management of presenter speaking time, the server records the meeting start time and presentation start time and sets a timer. It monitors speaking time, and if the set time is exceeded, the server automatically generates a warning. The server then sends the corresponding warning to the terminal, which alerts the user with a display and audio notification such as "3 minutes remaining."

[0721] For managing the progress of the discussion, the server compares the pre-configured agenda with the actual discussion content. If the server detects a discussion that does not match the agenda, it automatically generates a "discussion derailment" warning and sends it to the terminal. The terminal displays and provides an audio notification to the facilitator with a warning such as, "The topic has gone off track. Let's get back to the agenda."

[0722] To record final decisions, the server detects in real time any statements containing keywords such as "conclusion" or "decision" during the meeting. The detected statements are recorded in the database, and the server then sends an alert to the terminal prompting confirmation of the decision-maker's information. The terminal displays a message such as "Please enter the final decision-maker," and the user enters the decision-maker's information, which is then sent to the server and stored in the database.

[0723] This system enables efficient management of meeting time and agenda completion, and particularly allows junior facilitators to conduct meetings smoothly.

[0724] The following scenarios can be considered as concrete examples:

[0725] Example prompt: "Please describe the processing flow of a system that converts audio data into text data in real time and manages speaking time."

[0726] The above describes specific embodiments for carrying out the present invention.

[0727] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0728] Step 1: Acquire audio data

[0729] The device uses its microphone to record the meeting audio and starts recording in real time as soon as the meeting begins.

[0730] The device streams the recorded audio data to the server at regular time intervals (e.g., every second).

[0731] Input: Real-time audio data from the meeting

[0732] Output: Audio data to be streamed.

[0733] Step 2: Save the audio file.

[0734] The server temporarily stores the audio data received from the terminal in a storage service (e.g., Amazon S3, Google Cloud Storage).

[0735] Input: Streaming transmitted audio data

[0736] Output: Temporarily saved audio data

[0737] Step 3: Convert speech to text

[0738] The server uses the Google Cloud Speech-to-Text API to convert the temporarily stored audio data into text data.

[0739] Input: Temporarily stored audio data

[0740] Output: Converted text data

[0741] Step 4: Save text data

[0742] The server saves the converted text data to a database system (e.g., MySQL).

[0743] Input: Converted text data

[0744] Output: Text data stored in the database

[0745] Step 5: Analyze the progress of the meeting

[0746] The server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the stored text data.

[0747] Identify the progress of the meeting and the speakers.

[0748] Input: Text data stored in the database

[0749] Output: Analyzed meeting progress and speaker information

[0750] Step 6: Monitor speaking time

[0751] The server records the start time of the presentation and sets a timer.

[0752] The system monitors the presenter's speaking time in real time and generates a warning if the set time limit is exceeded.

[0753] Input: Analyzed meeting progress

[0754] Output: Warning based on excess time

[0755] Step 7: Warning Notification

[0756] The server sends the generated warning to the terminal.

[0757] The device will display a warning message as a pop-up and provide an audio notification.

[0758] Input: Generated warning

[0759] Output: Warning messages and audio notifications displayed on the device

[0760] Step 8: Detecting Derailments in the Discussion

[0761] The server compares the meeting discussion content with a pre-configured agenda.

[0762] A warning is generated if a topic unrelated to the agenda is mentioned.

[0763] Input: Analyzed meeting progress, agenda

[0764] Output: Warning based on mismatched topics

[0765] Step 9: Derailment warning notification

[0766] The server sends a derailment warning to the terminal.

[0767] The device displays a derailment warning to the facilitator and provides an audio notification.

[0768] Input: Generated derailment warning

[0769] Output: Derailment warning message and audio notification displayed on the device.

[0770] Step 10: Record and confirm decisions

[0771] The server detects in real time any statements containing keywords such as "conclusion" or "decision," and records their content.

[0772] The server sends an alert to the terminal prompting the decision-maker to confirm.

[0773] The user enters decision-maker information, and the terminal sends the entered information to the server.

[0774] Input: Analyzed text data, decision-maker information

[0775] Output: Recorded decisions and decision-maker information

[0776] The above outlines the specific processing steps of this system. This enables efficient real-time management and progress of meetings.

[0777] (Application Example 1)

[0778] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0779] Efficient progress management is required in discussions and meetings within factories, but preventing deviations from the discussion and managing the progress is difficult. Furthermore, managing presenters' time and meticulously recording important decisions are also critical challenges. Current meeting management systems cannot comprehensively address these issues, thus limiting their potential for improving factory productivity.

[0780] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0781] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected topics that do not match, means for recording the final decisions and decision-makers, means for monitoring the progress of the meeting in real time and managing the progress based on a set process plan, means for monitoring the presenter's speaking time and issuing a warning a certain time before a set time, means for detecting deviations and issuing a warning if the detected content does not match the process plan, and means for installing the progress management system on a smartphone. This enables efficient progress management, time management, and complete recording of important decisions for discussions and meetings within the factory.

[0782] "Means for acquiring audio data" refers to equipment and systems for acquiring audio from meetings and discussions in real time and transmitting it to a server.

[0783] "Means for converting acquired audio data into text data" refers to software or hardware that uses speech recognition technology to convert audio data into text information.

[0784] "Means of analyzing converted text data to monitor the progress of a meeting" refers to algorithms and programs that analyze text data to identify speakers and monitor the progress of a meeting in real time.

[0785] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system or device that measures the elapsed time from the start of the presenter's speech and generates a warning and notification when the set time is exceeded.

[0786] "Means for comparing the content of discussions with the agenda and detecting inconsistent topics" refers to a text analysis system that analyzes the content of discussions based on the purpose and agenda of a meeting and automatically detects topics that do not match the agenda.

[0787] "Means of issuing warnings based on detected mismatched topics" refers to a warning system that notifies users when a discussion deviates from the agenda.

[0788] "Means for recording final decisions and decision-makers" refers to a data recording system for organizing and preserving information about the decisions made during a meeting and the individuals who made those decisions.

[0789] "A means of monitoring the progress of a meeting in real time and managing its progress based on a set schedule" refers to a system for tracking the progress of meetings and discussions in real time and managing their progress according to a pre-set schedule.

[0790] "A means of monitoring the presenter's speaking time and issuing a warning a certain amount of time before the set time" refers to a system that issues a warning in advance when the scheduled speaking time approaches.

[0791] "Means for detecting deviations and issuing warnings when detected content does not match the process plan" refers to a system that detects when meeting content deviates from a pre-set process plan and issues a warning.

[0792] "Methods for installing a progress management system on a smartphone" refers to providing application software for progress management in a form that runs on a smartphone, and the methods and procedures for installing it.

[0793] This invention provides a system for efficiently managing the progress of discussions and meetings held within a factory, managing time, detecting deviations from the discussion, and recording final decisions. To achieve this, it employs a configuration that combines a server, a speech recognition engine, a text analysis system, a timekeeping function, and a warning notification system.

[0794] Hardware and software configuration

[0795] Hardware:

[0796] Smartphone: Equipped with a microphone for acquiring voice and a display for displaying warnings.

[0797] Server: A computer system used for storing and processing data.

[0798] software:

[0799] Python: A programming language for building the system of the present invention.

[0800] SpeechRecognition: A library for converting speech data into text.

[0801] threading: A multithreaded library for timekeeping and warning generation.

[0802] Process Overview

[0803] The server includes the following measures:

[0804] 1. Means of acquiring audio data:

[0805] The smartphone's microphone captures the meeting's audio data in real time and sends it to the server. The server continuously receives and stores this audio data.

[0806] 2. Means for converting acquired audio data into text data:

[0807] The server converts the received audio data into text data using the SpeechRecognition library. This conversion process is performed automatically, and the generated text data is then analyzed.

[0808] 3. Means of analyzing converted text data to monitor the progress of the meeting:

[0809] The server analyzes text data, identifies speakers, and monitors the progress of the meeting in real time. This allows for verification that the discussion is proceeding according to plan.

[0810] 4. A means of monitoring the presenter's speaking time and issuing a warning if the set time limit is exceeded:

[0811] The system monitors the elapsed time from the speaker's start and uses a timer function to issue a warning if the set time is exceeded. A warning message is displayed on the smartphone's screen, and audio notifications are provided as needed.

[0812] 5. Means for comparing the content of the discussion with the agenda and detecting inconsistencies:

[0813] The server compares the text data with the agenda and automatically detects any mismatched topics. This prevents the discussion from going in an unintended direction.

[0814] 6. Means of issuing warnings based on detected mismatched topics:

[0815] If mismatched topics are detected, the server will alert the facilitator and prompt them to quickly return the discussion to the agenda.

[0816] 7. Means for recording final decisions and decision-makers:

[0817] The server monitors statements containing keywords such as "conclusion" or "decision" in real time, and records their content and the information of the decision-maker. This information is stored in a format that can be referenced later.

[0818] 8. Means for monitoring the progress of meetings in real time and managing their progress based on the set schedule:

[0819] Real-time progress management is performed based on the established project plan, and the progress of discussions is monitored to ensure that they proceed as planned.

[0820] 9. A means of monitoring the presenter's speaking time and issuing a warning a certain period of time before the set end time:

[0821] By issuing a warning a certain amount of time in advance as the scheduled speaking time approaches, this system helps speakers organize their information within the allotted time.

[0822] 10. Means for detecting derailments and issuing warnings when the detected content does not match the process plan:

[0823] If a statement is made that does not conform to the project plan, the server will automatically detect the derailment and issue a warning.

[0824] 11. How to install the progress management system on a smartphone:

[0825] We will install the progress management system on smartphones and provide it to users so that they can easily use it.

[0826] Specific example

[0827] For example, when discussions are held within a factory about introducing a new process, a progress management assistant app can be used to manage time, prevent discussions from going off track, and record important decisions in real time. This app is installed on a smartphone and can be used from anywhere within the factory.

[0828] Example of a prompt

[0829] Requirements for a real-time meeting management system:

[0830] A feature that acquires audio data in real time and converts it to text using a speech recognition engine.

[0831] A function to identify the speaker and monitor the progress.

[0832] A feature that monitors speaking time and generates a warning if the set time limit is exceeded.

[0833] A feature that compares the discussion content with the agenda and detects deviations from the topic.

[0834] A feature that records final decisions and the decision-makers in real time.

[0835] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0836] Step 1:

[0837] Acquisition of audio data

[0838] Input: Audio from meetings or discussions

[0839] Output: Audio data acquired in real time

[0840] Specific operation: The device (smartphone) acquires audio from meetings and discussions in real time via its microphone. The acquired audio data is sent directly to the server. During this process, the device maintains a continuous connection with the server to ensure a stable communication environment.

[0841] Step 2:

[0842] Text conversion of acquired audio data

[0843] Input: Acquired audio data

[0844] Output: Text data converted by speech recognition

[0845] Specific operation: The server uses the SpeechRecognition library to convert the acquired audio data into text data. In this process, the speech recognition engine analyzes the audio signal and generates the corresponding string. The converted text data is temporarily stored in the server's memory or database.

[0846] Step 3:

[0847] Monitoring the progress of the meeting

[0848] Input: Converted text data

[0849] Output: Record of progress, identification of speakers

[0850] Specific operation: The server analyzes the converted text data and monitors the progress of the meeting. Specifically, it automatically recognizes speakers from the content of their statements within the text data and records the presentation time and content of each speaker. This allows for real-time tracking of who said what and when.

[0851] Step 4:

[0852] Monitoring and warnings regarding the presenter's speaking time.

[0853] Input: Announcement start time and current time, set time limit

[0854] Output: Time-based warnings as needed

[0855] Specific operation: The server records the presenter's speaking start time and monitors the speaking time based on the set time limit. Using a timer function, it sends a warning to the terminal when the speaking time is a certain amount of time away from the set limit. The warning message is displayed on the terminal's screen and an audio notification is also provided.

[0856] Step 5:

[0857] Comparison of discussion content and agenda

[0858] Input: Parsed text data, pre-configured agenda

[0859] Output: Detection and recording of mismatched topics

[0860] Specific operation: The server compares the converted text data with a pre-configured agenda and automatically detects mismatched topics. A text analysis algorithm is used to determine whether specific keywords or phrases match the agenda. If mismatched topics are detected, their content is recorded, and an alert is issued to the facilitator.

[0861] Step 6:

[0862] Issuing warnings based on unrelated topics

[0863] Input: Detection results for mismatched topics

[0864] Output: Warning to the facilitator

[0865] Specific operation: When the server detects mismatched topics, it issues a warning to the facilitator. The warning is displayed as a text message on the terminal's screen, and an audio notification is also provided. This allows the facilitator to quickly understand that the discussion is deviating and take appropriate action.

[0866] Step 7:

[0867] Records of final decisions and decision-makers

[0868] Input: Converted text data, specific keywords (e.g., "conclusion," "decision")

[0869] Output: Records of decisions and decision-makers

[0870] Specific operation: The server monitors the converted text data in real time and detects statements containing keywords such as "conclusion" or "decision." Detected statements are considered decisions, and the information is recorded. Additionally, a prompt for the user to confirm the decision is displayed on the terminal, and the entered decision-maker information is also sent to the server and recorded.

[0871] Step 8:

[0872] Real-time monitoring of meeting progress and project plan progress management.

[0873] Input: Converted text data, configured process plan

[0874] Output: Progress management information

[0875] Specific operation: The server monitors the meeting progress in real time and manages its progress based on the configured schedule. It verifies that the progress is proceeding according to plan and issues warnings and alerts as needed. This function is crucial to ensuring that the meeting proceeds as planned.

[0876] Step 9:

[0877] Pre-scheduled warning of presenter's speaking time

[0878] Input: Start time of speaking and a certain period of time before the set time limit.

[0879] Output: Pre-warning

[0880] Specific operation: A warning about the remaining speaking time is sent to the device a certain amount of time after the scheduled start time. This allows the presenter to organize their remarks while being mindful of the remaining time.

[0881] Step 10:

[0882] Derailment detection and warning issuance for content that does not match the process plan.

[0883] Input: Converted text data, process plan

[0884] Output: Derailment detection results and warnings

[0885] Specific operation: The converted text data is compared with the process plan. If any discrepancies are detected, the server detects a deviation and issues a warning. The warning is displayed on the terminal's screen and an audio notification is provided to allow the facilitator to respond quickly.

[0886] Step 11:

[0887] Installing the progress management system on smartphones

[0888] Input: None (pre-configuration is assumed)

[0889] Output: Installed progress management systems

[0890] Specific actions: Install the progress management system on smartphones. Installation involves providing dedicated application software, making it easy for users to download and configure. This feature allows users to access the progress management system from anywhere within the factory.

[0891] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0892] This invention is a system that acquires audio data during a meeting and analyzes it in real time. Furthermore, by combining it with an emotion engine, it recognizes the emotions of participants and facilitates smoother meeting management.

[0893] Program Overview

[0894] This meeting management system can be implemented with the following program configuration.

[0895] 1. Acquisition of audio data

[0896] The device records the meeting audio in real time and streams the audio data to the server.

[0897] 2. Speech-to-text conversion

[0898] The server converts the received audio data into text data using a speech recognition engine.

[0899] 3. Monitoring the progress of the meeting

[0900] The server analyzes the converted text data, identifies speakers, and annotates the data to monitor the progress of the meeting.

[0901] 4. Timekeeping

[0902] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[0903] After a set amount of time has elapsed since the start of the announcement, the server generates an alert and sends it to the device.

[0904] The device will notify the presenter of the remaining time via audio notification and a display alert.

[0905] 5. Detecting Derailments in the Discussion

[0906] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[0907] The server generates a derailment detection warning and sends it to the terminal.

[0908] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[0909] 6. Emotion recognition by an emotion engine

[0910] The terminals and servers are equipped with emotion engines that recognize emotions in real time from the voice tone and content of what meeting participants say.

[0911] 7. Adjusting Emotion-Based Warnings

[0912] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[0913] 8. Emotion-based stress management

[0914] The emotion engine monitors the meeting facilitator's stress level in real time and provides alerts to reduce stress as needed.

[0915] 9. Records of final decisions and decision-makers

[0916] The server monitors keywords such as "conclusion" and "decision" in real time and records their content when detected.

[0917] The device displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's name.

[0918] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[0919] Specific example

[0920] Case Study 1: Timekeeping Alerts

[0921] 1. The user begins a 15-minute presentation.

[0922] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[0923] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[0924] 4. The device notifies the user that "3 minutes remaining."

[0925] Case Study 2: Detecting Derailments in Discussions

[0926] 1. The agenda is set to "Project A progress report".

[0927] 2. The user begins talking at length about "Project B".

[0928] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[0929] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[0930] 5. The device notifies the facilitator that "the topic has gone off track."

[0931] Case Study 3: Stress Management Using an Emotional Engine

[0932] 1. The facilitator begins to feel stressed while conducting the meeting.

[0933] 2. The emotion engine monitors the facilitator's stress level in real time.

[0934] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[0935] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[0936] The following describes the processing flow.

[0937] Step 1:

[0938] The user launches the meeting management application and clicks the "Start Meeting" button.

[0939] Step 2:

[0940] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[0941] Step 3:

[0942] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[0943] Step 4:

[0944] The device records the meeting audio in real time and streams the audio data to the server.

[0945] Step 5:

[0946] The server uses a speech recognition engine to convert the received audio data into text data.

[0947] Step 6:

[0948] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[0949] Step 7:

[0950] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[0951] Step 8:

[0952] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[0953] Step 9:

[0954] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[0955] Step 10:

[0956] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[0957] Step 11:

[0958] If a derailment is detected, the server generates a warning and sends it to the terminal.

[0959] Step 12:

[0960] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[0961] Step 13:

[0962] An emotion engine installed on the server analyzes the voice tone and content of meeting participants in real time to recognize their emotions.

[0963] Step 14:

[0964] The server adjusts the warning content and notification tone for presenters and participants based on the emotional state obtained from the emotion engine.

[0965] Step 15:

[0966] If the server, based on the results of its emotion engine's emotional state monitoring, determines that the facilitator's stress level is high, it generates an alert saying, "Let's take a short break," and sends it to the terminal.

[0967] Step 16:

[0968] The device provides the facilitator with an audio notification and a display alert saying, "We recommend you take a short break."

[0969] Step 17:

[0970] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[0971] Step 18:

[0972] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[0973] Step 19:

[0974] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[0975] Step 20:

[0976] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[0977] (Example 2)

[0978] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0979] Many challenges exist in managing the progress of meetings. Existing systems simply converted audio to text and monitored the agenda, resulting in insufficient management of speaking time, prevention of discussion digressions, and recognition of participants' emotions. Furthermore, stress management for meeting facilitators was not considered. A system that comprehensively addresses these challenges is needed.

[0980] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0981] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting mismatched topics, means for issuing a warning based on the detected mismatched topics, means for recognizing the emotions of participants and adjusting the content of warnings and notification tones based on their emotional state, means for monitoring the stress level of the meeting facilitator in real time and providing alerts for stress reduction as needed, and means for recording final decisions and the decision-makers. This makes it possible to comprehensively and in real time perform everything from acquiring audio data to text conversion, monitoring progress, timekeeping, detecting deviations from the discussion, emotion recognition, stress management, and recording final decisions.

[0982] "Audio data" refers to digital data in which sound is recorded electronically.

[0983] "Text data" refers to digital data obtained by converting audio data into written text.

[0984] "Progress status" refers to information indicating the current state of the meeting's progress.

[0985] "Speaking time" refers to the time elapsed since the presenter began speaking.

[0986] An "agenda" is a list of topics or subjects that should be discussed at a meeting.

[0987] A "topic" refers to a specific subject or theme that is discussed in a meeting.

[0988] A "warning" is a notice issued when a set condition is not met.

[0989] "Emotions" refers to the psychological and emotional state of the meeting participants.

[0990] "Emotional state" refers to the specific expression of a meeting participant's emotions at a particular moment.

[0991] "Notification tone" refers to the sound or tone used when issuing warnings or notifications.

[0992] "Stress level" indicates the degree of mental and physical burden on the meeting organizer.

[0993] An "alert" is a warning or notice that appears when certain conditions occur.

[0994] "Decisions made" refer to the specific details and action plans that were ultimately decided upon at the meeting.

[0995] A "decision-maker" is the person or position that ultimately makes the final decision on a particular matter in a meeting.

[0996] Modes for carrying out the invention

[0997] This invention is a system that acquires audio data during a meeting and analyzes it in real time. Furthermore, by combining it with an emotion engine, it recognizes the emotions of participants and facilitates smoother meeting management. This system includes the following components:

[0998] Acquisition of audio data

[0999] When a user starts a meeting, the device uses its microphone to record the meeting's audio data in real time and streams that audio data to the server. The audio data is processed in digital format and transferred to the server via the communication network.

[1000] Speech-to-text conversion

[1001] The server uses a speech recognition engine such as Google Cloud Speech-to-Text to convert the received audio data into text data. The converted text data is temporarily stored on the server and used for subsequent data analysis.

[1002] Monitoring the progress of the meeting

[1003] The server uses a natural language processing engine, such as the Google Cloud Natural Language API, to analyze the text data converted from speech and identify the speaker. It adds annotations to the text data, including the speaker's ID and the content of their speech, and monitors the progress of the meeting.

[1004] Timekeeping

[1005] The server records the start time of each user's speech and monitors the speech duration in real time. Once the set time has elapsed, the server generates an alert and sends it to the terminal. The terminal then notifies the presenter of the remaining time via audio notification and a display alert.

[1006] Derailment detection in discussions

[1007] The server compares the converted text data with pre-configured agenda items and recognizes any discussion of a mismatched topic as a digression. When a digression is detected, the server generates a warning and sends it to the terminal. The terminal then notifies the facilitator with an audio and visual alert stating, "The topic has gone off-topic."

[1008] Emotion recognition by an emotion engine

[1009] The terminals and servers are equipped with emotion engines such as the Azure Cognitive Services Emotion API to analyze the voice tone and content of meeting participants. This makes it possible to recognize participants' emotions (e.g., joy, anger, surprise) in real time.

[1010] Adjusting emotion-based warnings

[1011] The server receives the emotional state recognized by the emotion engine and adjusts the warning content and notification tone accordingly. For example, if a participant is feeling stressed, the alert tone will be softened and the notification content will be made gentler.

[1012] Emotion-based stress management

[1013] The emotion engine monitors the meeting facilitator's stress level in real time and provides this information to the server. If the server determines that the facilitator's stress level is high, it generates alerts for breaks or urgent discussions. For example, it might send a message to the facilitator such as, "We recommend you take a short break."

[1014] Records of final decisions and decision-makers

[1015] The server monitors keywords such as "conclusion" and "decision" in the conversation in real time. When a keyword is detected, its content is immediately recorded. The terminal displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's information. The entered decision-maker information is sent from the terminal to the server and recorded along with the final decision.

[1016] Specific example

[1017] Case Study 1: Timekeeping Alerts

[1018] 1. The user begins a 15-minute presentation.

[1019] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[1020] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[1021] 4. The device notifies the user that "3 minutes remaining."

[1022] Prompt example

[1023] "Design a system that records the start time of a 15-minute presentation and generates a '3 minutes remaining' alert 12 minutes later."

[1024] Case Study 2: Detecting Derailments in Discussions

[1025] 1. The agenda is set to "Project progress report".

[1026] 2. The user starts talking at length about "another project."

[1027] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[1028] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[1029] 5. The device notifies the facilitator that "the topic has gone off track."

[1030] Prompt example

[1031] "Design a system that detects off-topic comments during meetings and alerts the facilitator to the derailment."

[1032] Case Study 3: Stress Management Using an Emotional Engine

[1033] 1. The facilitator begins to feel stressed while conducting the meeting.

[1034] 2. The emotion engine monitors the facilitator's stress level in real time.

[1035] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[1036] Prompt example

[1037] "Design a system that monitors the facilitator's stress level during meetings and sends notifications prompting them to take breaks when necessary."

[1038] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[1039] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1040] Step 1:

[1041] The user starts the meeting.

[1042] Specifically, the user launches the meeting application and starts the meeting by clicking the "Start Meeting" button.

[1043] Input: User's meeting start operation

[1044] Output: Meeting start signal

[1045] Step 2:

[1046] The device prepares to record audio.

[1047] Specifically, the device checks the readiness status of the built-in or externally connected microphone and enables the recording function.

[1048] Input: Signal to start meeting

[1049] Output: Recording ready signal

[1050] Step 3:

[1051] The device records audio data in real time and streams it to the server.

[1052] Specifically, the device records the audio during the meeting in real time and streams it to a server via the internet as digital audio data.

[1053] Input: Signal indicating recording is ready.

[1054] Output: Streaming of real-time audio data

[1055] Step 4:

[1056] The server receives the audio data and uses a speech recognition engine to convert the audio data into text data.

[1057] In terms of specific operations, the server uses a speech recognition engine to analyze the audio data and convert it into text data written as characters. A typical example of the software used is a general-purpose speech recognition engine.

[1058] Input: Real-time audio data

[1059] Output: Text data

[1060] Step 5:

[1061] The server temporarily stores the converted text data.

[1062] Specifically, the server temporarily stores the converted text data in storage such as a database.

[1063] Input: Text data

[1064] Output: Temporarily saved text data

[1065] Step 6:

[1066] The server analyzes the text data converted from the speech and monitors the progress using a natural language processing engine.

[1067] In terms of specific operations, the server uses a natural language processing engine to analyze text data, identify the speaker, and annotate the content of the speech. The software used includes a natural language processing engine, among other things.

[1068] Input: Temporarily saved text data

[1069] Output: Annotated text data

[1070] Step 7:

[1071] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[1072] Specifically, the server records the start time of a message and then continuously measures the duration of subsequent messages.

[1073] Input: Annotated text data

[1074] Output: Speech time elapsed data

[1075] Step 8:

[1076] After a set amount of time has elapsed since the start of the announcement, the server generates an alert and sends it to the device.

[1077] Specifically, when the server determines that a set time has elapsed, it generates an alert message and sends it to the terminal.

[1078] Input: Elapsed time of speech

[1079] Output: Alert

[1080] Step 9:

[1081] The device will notify the presenter of the remaining time via audio notification and a display alert.

[1082] Specifically, the device will use voice notifications and screen displays based on the received alerts to inform the presenter of the remaining time.

[1083] Input: Alert

[1084] Output: Voice notifications and display alerts

[1085] Step 10:

[1086] The server compares the converted text data with the agenda and detects any mismatched topics.

[1087] Specifically, the server compares the text data with a pre-configured agenda list and automatically detects any mismatches.

[1088] Input: Annotated text data

[1089] Output: Detection results for mismatched topics

[1090] Step 11:

[1091] The server issues a warning based on mismatched topics and sends it to the terminal.

[1092] Specifically, the server generates a warning message based on the detected mismatched topics and sends it to the terminal.

[1093] Input: Detection results for mismatched topics

[1094] Output: Warning message

[1095] Step 12:

[1096] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[1097] Specifically, the device uses the received warning message to provide audio notifications and screen displays, informing the facilitator that the topic has gone off track.

[1098] Input: Warning message

[1099] Output: Voice notifications and display alerts

[1100] Step 13:

[1101] The terminals and servers are equipped with emotion engines to recognize the emotions of the participants.

[1102] Specifically, the terminal and server use an emotion engine to analyze the participant's emotions in real time based on their voice tone and what they say.

[1103] Input: Real-time audio data

[1104] Output: Emotion recognition result

[1105] Step 14:

[1106] The server adjusts warning content and notification tone based on the emotional state recognized by the emotion engine.

[1107] Specifically, the server adjusts the content and tone of the warning message based on the emotion recognition results.

[1108] Input: Sentiment recognition result

[1109] Output: Adjusted warning message

[1110] Step 15:

[1111] The emotion engine monitors the stress levels of meeting organizers, and the server provides alerts to reduce stress as needed.

[1112] Specifically, the emotion engine constantly analyzes the operator's stress level and generates alerts for breaks and relaxation at appropriate times.

[1113] Input: Real-time audio data and operator sentiment data

[1114] Output: Stress reduction alert

[1115] Step 16:

[1116] The server monitors for "conclusion" and "decision" keywords and records their content.

[1117] Specifically, the server monitors for specific keywords within the converted text data in real time and records their content in the database if detected.

[1118] Input: Annotated text data

[1119] Output: Recorded decisions

[1120] Step 17:

[1121] The terminal prompts the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[1122] Specifically, after a decision is finalized, the terminal displays a pop-up notification prompting the user to enter the decision-maker's information. The entered information is sent to the server and recorded.

[1123] Input: Decision made and user-inputted decision-maker information

[1124] Output: Recorded decision-maker information

[1125] (Application Example 2)

[1126] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1127] Meetings during shift changes in factories require smooth information transfer and efficient discussion. However, insufficient communication on the factory floor and frequent digressions in discussions often lead to problems with efficient shift changes and work handovers. Furthermore, the emotional state of participants significantly impacts the progress of the meeting, making emotional management a crucial issue. A system is needed to appropriately address these problems.

[1128] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting inconsistent topics, means for issuing a warning based on the detected inconsistent topics, and means for recognizing the emotions of meeting participants and adjusting the content of the warning and the tone of the notification based on their emotional state. This enables appropriate information transmission, efficient progress of discussions, and emotional management of participants in factory shift change meetings.

[1129] "Means for acquiring audio data" refers to a device or software for recording speech spoken in real time during meetings or shift changes and saving it as digital data.

[1130] "Means for converting acquired audio data into text data" refers to a speech recognition engine or software that analyzes recorded audio data and automatically converts it into text data.

[1131] "Means of analyzing converted text data to monitor the progress of a meeting" refers to a program or system that analyzes text data to understand who is speaking and the progress of the agenda in real time.

[1132] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system with a timekeeping function that records the start time of the presentation and generates an alert when the set time limit is exceeded.

[1133] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a program or algorithm that compares the content of a statement with a set agenda and detects when a topic that does not match the agenda is mentioned.

[1134] "Means for issuing warnings based on detected mismatched topics" refers to a system that automatically generates warnings related to detected mismatched topics and notifies speakers and moderators.

[1135] "Means for recording final decisions and decision-makers" refers to a system that has the function of recording decisions made during meetings or shift changes, as well as the decision-makers involved, and saving them in a database for later reference.

[1136] "Means for recognizing the emotions of meeting participants and adjusting the tone of warnings and notifications based on their emotional state" refers to an algorithm or engine that analyzes emotions in real time from the tone of voice and content of statements of participants, and changes the content and expression of warnings and notifications according to their emotional state.

[1137] The system that implements this application supports meetings during shift changes in a factory. This system acquires audio data in real time, converts it into text data, and performs meeting progress monitoring, timekeeping, discussion derailment detection, and sentiment recognition.

[1138] Specifically, the program will be implemented with the following structure:

[1139] Hardware and software

[1140] Hardware to use:

[1141] Microphone: For capturing the voices of meeting participants.

[1142] Robot or smartphone: A platform for processing and notifying voice data.

[1143] Software to use:

[1144] speech_recognition: A library for performing speech recognition.

[1145] emotion_recognition: A library for performing emotion recognition (uses a pre-trained model).

[1146] datetime: A standard library for time management.

[1147] Process Overview

[1148] 1. Acquisition of audio data:

[1149] The audio of meeting participants can be captured in real time via microphones.

[1150] The voice data is transferred to a robot or smartphone platform.

[1151] 2. Speech-to-text conversion:

[1152] The speech_recognition library installed on the robot or smartphone is used to convert speech data into text data.

[1153] The converted text is sent to the server.

[1154] 3. Monitoring the progress of the meeting:

[1155] The server analyzes the converted text data to identify the speaker and monitor the progress.

[1156] Verify that the meeting is proceeding according to the set agenda.

[1157] 4. Timekeeping:

[1158] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[1159] If the presentation time exceeds the set time, a warning is sent from the server to the robot or smartphone, and the presenter is notified via audio notification and a display alert.

[1160] 5. Detecting deviations from the discussion:

[1161] The server compares the discussion to the agenda and detects deviations if a topic that does not match is mentioned.

[1162] If a derailment is detected, a warning will be displayed to the facilitator.

[1163] 6. Emotion recognition by the emotion engine:

[1164] The robot or smartphone is equipped with an emotion engine that recognizes the participant's emotions in real time based on their voice tone and what they say.

[1165] This allows for monitoring the emotional state of participants and taking appropriate action.

[1166] 7. Adjusting emotion-based warnings:

[1167] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[1168] For example, if emotions are heightened, the notification will be delivered in a calm tone.

[1169] Specific example

[1170] Examples of support for shift change meetings

[1171] 1. Start of the meeting:

[1172] When the shift changes, the leader begins explaining the situation to the next team.

[1173] The microphone acquires audio data in real time and transmits it to the robot or smartphone.

[1174] 2. Speech-to-text conversion:

[1175] The acquired audio data is converted to text using the speech_recognition library and sent to the server.

[1176] 3. Monitoring the progress of the meeting and keeping time:

[1177] The server analyzes the converted text data to identify the speaker, monitor the progress, and keep time.

[1178] After a set amount of time has elapsed since the announcement began, a warning will be sent to the robot or smartphone, triggering an audio notification and a visual alert.

[1179] 4. Detecting deviations from the discussion:

[1180] The server compares the discussion content with the agenda and warns the facilitator if it detects a deviation from the topic.

[1181] 5. Recognition by the emotion engine:

[1182] It monitors future emotional states in real time and adjusts warning content and notification tones as needed.

[1183] Examples of prompt statements to use

[1184] "Record the 5-minute shift change meeting and convert it to text. Perform sentiment analysis using an emotion recognition model, and implement tactical detection and timekeeping."

[1185] This system allows factory shift change meetings to proceed smoothly and efficiently, and because it takes into account the emotional state of the participants, it enables more effective information sharing and discussion.

[1186] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1187] Step 1:

[1188] Acquire audio data.

[1189] A microphone connected to a terminal (robot or smartphone) records the speech of meeting participants in real time. The input is the raw voice of the meeting participants, and the output is digitized audio data. This audio data is streamed to a server.

[1190] Step 2:

[1191] Convert audio data to text data.

[1192] The audio data sent to the server is automatically converted into text data using the speech_recognition library. The input is audio data, and the output is text data in string format. This text data is then passed on to the next parsing step.

[1193] Step 3:

[1194] Monitor the progress of the meeting.

[1195] The server analyzes text data to identify the content and speaker of a statement. The input is text data, and the output is a pair of speaker and statement information. Based on this, it identifies speakers and generates progress graphs.

[1196] Step 4:

[1197] Timekeeping will be performed.

[1198] The server records the presenter's start time and measures the elapsed time in real time. When the set presentation time has elapsed, it generates a warning and sends it to the terminal. The input is the start time, and the output is the elapsed time and the warning message. The terminal notifies the presenter with an audio notification and a display alert.

[1199] Step 5:

[1200] Detects deviations from the discussion.

[1201] The server compares the text data with the previously set agenda and detects any mismatched topics. The input is the text data and agenda information, and the output is a list of mismatched topics. If a digression is detected, the server sends a warning to the facilitator.

[1202] Step 6:

[1203] Recognize your emotions.

[1204] The emotion engine built into the device recognizes the participant's emotions in real time based on their voice tone and speech content. Input consists of voice and text data, and output is the identified emotional state. This information is sent to the server.

[1205] Step 7:

[1206] The tone of warnings and notifications is adjusted based on emotions.

[1207] The server dynamically adjusts the content of warnings and notification tones based on the emotional state obtained from the emotion engine. The input is emotional state data, and the output is the adjusted warning message. For example, if a participant is stressed, the notification will be delivered in a calm tone.

[1208] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1209] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1210] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1211] [Third Embodiment]

[1212] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1213] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1214] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1215] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1216] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1217] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1218] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1219] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1220] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1221] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1222] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1223] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1224] This invention is a system for a meeting management system that analyzes audio data in real time to appropriately manage the speaking time of presenters and the progress of discussions.

[1225] Program Overview

[1226] The meeting management system will be implemented with the following program structure.

[1227] 1. Acquisition of audio data

[1228] The device records the meeting audio in real time and streams it to the server.

[1229] 2. Speech-to-text conversion

[1230] The server converts the received audio data into text data using a speech recognition engine.

[1231] 3. Monitoring the progress of the meeting

[1232] The server analyzes the converted text data, automatically recognizes the speaker, and monitors the progress of the meeting.

[1233] 4. Timekeeping

[1234] The server monitors the presenter's speaking time and generates a warning if the set time limit is exceeded.

[1235] The device displays an alert to the presenter and provides an audio notification.

[1236] 5. Detecting Derailments in the Discussion

[1237] The server compares the discussion content with the agenda and generates a warning if a topic that does not match is mentioned.

[1238] The device displays a warning to the facilitator and provides an audio notification.

[1239] 6. Records of final decisions and decision-makers

[1240] The server monitors and detects statements containing keywords such as "conclusion" and "decision" in real time, and records their content.

[1241] The device displays an alert prompting the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[1242] Specific example

[1243] Case Study 1: Timekeeping Alerts

[1244] 1. User A begins a 15-minute presentation.

[1245] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[1246] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[1247] 4. The device displays and announces to user A that "3 minutes remaining."

[1248] 5. User A completes the presentation within the allotted time.

[1249] Case Study 2: Detecting Derailments in Discussions

[1250] 1. The meeting agenda is set to "Progress report for Project A".

[1251] 2. User B begins talking at length about "Project B".

[1252] 3. The server analyzes the discussion content as text and detects any discrepancies with the agenda.

[1253] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[1254] 5. The device displays and provides an audio notification to the facilitator stating, "The topic has gone off track. Let's return to the agenda."

[1255] 6. The facilitator encourages the discussion to return to the agenda.

[1256] Case Study 3: Records of Final Decisions and Decision-makers

[1257] 1. During the meeting, comments containing keywords such as "conclusion" are made.

[1258] 2. The server detects the statement and records its content.

[1259] 3. The server sends an alert to the terminal prompting the final decision-maker to confirm.

[1260] 4. The terminal displays the message "Please enter the final decision-maker," and User C enters the decision-maker.

[1261] 5. The entered decision-maker information is sent to the server and recorded along with the decision.

[1262] As described above, this system offers the advantage of enabling junior facilitators to efficiently manage meetings and facilitate time management and agenda achievement.

[1263] The following describes the processing flow.

[1264] Step 1:

[1265] The user launches the meeting management application and clicks the "Start Meeting" button.

[1266] Step 2:

[1267] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[1268] Step 3:

[1269] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[1270] Step 4:

[1271] The device records the meeting audio in real time and streams the audio data to the server.

[1272] Step 5:

[1273] The server uses a speech recognition engine to convert the received audio data into text data.

[1274] Step 6:

[1275] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[1276] Step 7:

[1277] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[1278] Step 8:

[1279] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[1280] Step 9:

[1281] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[1282] Step 10:

[1283] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[1284] Step 11:

[1285] If a derailment is detected, the server generates a warning and sends it to the terminal.

[1286] Step 12:

[1287] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[1288] Step 13:

[1289] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[1290] Step 14:

[1291] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[1292] Step 15:

[1293] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[1294] Step 16:

[1295] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[1296] The above outlines the specific processing steps and details of the actions performed at each step. This system will enable junior facilitators to manage meetings efficiently and effectively.

[1297] (Example 1)

[1298] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1299] Traditional meeting management systems struggle to efficiently manage speaking time and facilitate discussions, leaving younger facilitators, in particular, lacking the tools to ensure smooth meetings. Specifically, they are unable to effectively perform real-time analysis of meeting audio data, monitor presenter speaking time, detect tangents, and record final decisions and decision-makers, resulting in a decline in meeting quality.

[1300] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1301] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected mismatched topics, means for detecting statements containing keywords such as "conclusion" or "decision" and recording their content, and means for prompting confirmation of decision-maker information. This makes it possible to appropriately manage the presenter's speaking time, detect digressions in the discussion, and ensure the smooth progress of the meeting.

[1302] "Means for acquiring audio data" refers to a device or function for recording audio during a meeting in real time and streaming it to a server.

[1303] "Means for converting audio data to text data" refers to a device or function that converts received audio data into corresponding text data using speech recognition technology.

[1304] "Means for analyzing converted text data to monitor the progress of a meeting" refers to a device or function that analyzes text data converted by speech recognition to understand the content of statements and speakers, and to monitor the progress of a meeting.

[1305] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a device or function that measures the elapsed time from the moment the presenter starts speaking and generates a warning if the set time limit is exceeded.

[1306] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a device or function that compares a pre-set agenda with the actual content of a discussion and detects when a topic not included in the agenda is mentioned.

[1307] "Means for issuing warnings based on detected mismatched topics" refers to a device or function that warns that the discussion is going off track when a topic that does not match the agenda is detected.

[1308] "Means for detecting and recording statements containing keywords such as 'conclusion' or 'decision'" refers to a device or function that detects statements containing specific keywords in real time from text data during a meeting and records the content of those statements in a database or similar.

[1309] "Means for prompting confirmation of decision-maker information" refers to a device or function that generates an alert prompting the user to confirm the decision-maker's information regarding the detected final decision, and prompts the user to input that information.

[1310] This invention relates to a system for a meeting management system that analyzes audio data in real time to appropriately manage presenter speaking time and facilitate discussion. This system is primarily implemented through the interaction of a server, terminals, and users.

[1311] First, the device records the meeting audio in real time using its built-in or external microphone as soon as the meeting starts. Standard audio formats (e.g., WAV, MP3) are used for recording. This recorded audio data is transmitted to the server via streaming at regular intervals (e.g., every second).

[1312] The server uses a storage service (e.g., Amazon S3, Google Cloud Storage) to temporarily store the received audio data. Next, the server uses a speech recognition engine, such as the Google Cloud Speech-to-Text API, to convert the audio data into corresponding text data. This converted text data is then stored in a database system (e.g., MySQL).

[1313] Furthermore, the server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the converted text data. This analysis monitors the progress of the meeting and identifies speakers. Identification is based on speaker-specific language patterns and vocal characteristics, and the results are stored in a database.

[1314] Regarding the management of presenter speaking time, the server records the meeting start time and presentation start time and sets a timer. It monitors speaking time, and if the set time is exceeded, the server automatically generates a warning. The server then sends the corresponding warning to the terminal, which alerts the user with a display and audio notification such as "3 minutes remaining."

[1315] For managing the progress of the discussion, the server compares the pre-configured agenda with the actual discussion content. If the server detects a discussion that does not match the agenda, it automatically generates a "discussion derailment" warning and sends it to the terminal. The terminal displays and provides an audio notification to the facilitator with a warning such as, "The topic has gone off track. Let's get back to the agenda."

[1316] To record final decisions, the server detects in real time any statements containing keywords such as "conclusion" or "decision" during the meeting. The detected statements are recorded in the database, and the server then sends an alert to the terminal prompting confirmation of the decision-maker's information. The terminal displays a message such as "Please enter the final decision-maker," and the user enters the decision-maker's information, which is then sent to the server and stored in the database.

[1317] This system enables efficient management of meeting time and agenda completion, and particularly allows junior facilitators to conduct meetings smoothly.

[1318] The following scenarios can be considered as concrete examples:

[1319] Example prompt: "Please describe the processing flow of a system that converts audio data into text data in real time and manages speaking time."

[1320] The above describes specific embodiments for carrying out the present invention.

[1321] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1322] Step 1: Acquire audio data

[1323] The device uses its microphone to record the meeting audio and starts recording in real time as soon as the meeting begins.

[1324] The device streams the recorded audio data to the server at regular time intervals (e.g., every second).

[1325] Input: Real-time audio data from the meeting

[1326] Output: Audio data to be streamed.

[1327] Step 2: Save the audio file.

[1328] The server temporarily stores the audio data received from the terminal in a storage service (e.g., Amazon S3, Google Cloud Storage).

[1329] Input: Streaming transmitted audio data

[1330] Output: Temporarily saved audio data

[1331] Step 3: Convert speech to text

[1332] The server uses the Google Cloud Speech-to-Text API to convert the temporarily stored audio data into text data.

[1333] Input: Temporarily stored audio data

[1334] Output: Converted text data

[1335] Step 4: Save text data

[1336] The server saves the converted text data to a database system (e.g., MySQL).

[1337] Input: Converted text data

[1338] Output: Text data stored in the database

[1339] Step 5: Analyze the progress of the meeting

[1340] The server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the stored text data.

[1341] Identify the progress of the meeting and the speakers.

[1342] Input: Text data stored in the database

[1343] Output: Analyzed meeting progress and speaker information

[1344] Step 6: Monitor speaking time

[1345] The server records the start time of the presentation and sets a timer.

[1346] The system monitors the presenter's speaking time in real time and generates a warning if the set time limit is exceeded.

[1347] Input: Analyzed meeting progress

[1348] Output: Warning based on excess time

[1349] Step 7: Warning Notification

[1350] The server sends the generated warning to the terminal.

[1351] The device will display a warning message as a pop-up and provide an audio notification.

[1352] Input: Generated warning

[1353] Output: Warning messages and audio notifications displayed on the device

[1354] Step 8: Detecting Derailments in the Discussion

[1355] The server compares the meeting discussion content with a pre-configured agenda.

[1356] A warning is generated if a topic unrelated to the agenda is mentioned.

[1357] Input: Analyzed meeting progress, agenda

[1358] Output: Warning based on mismatched topics

[1359] Step 9: Derailment warning notification

[1360] The server sends a derailment warning to the terminal.

[1361] The device displays a derailment warning to the facilitator and provides an audio notification.

[1362] Input: Generated derailment warning

[1363] Output: Derailment warning message and audio notification displayed on the device.

[1364] Step 10: Record and confirm decisions

[1365] The server detects in real time any statements containing keywords such as "conclusion" or "decision," and records their content.

[1366] The server sends an alert to the terminal prompting the decision-maker to confirm.

[1367] The user enters decision-maker information, and the terminal sends the entered information to the server.

[1368] Input: Analyzed text data, decision-maker information

[1369] Output: Recorded decisions and decision-maker information

[1370] The above outlines the specific processing steps of this system. This enables efficient real-time management and progress of meetings.

[1371] (Application Example 1)

[1372] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1373] Efficient progress management is required in discussions and meetings within factories, but preventing deviations from the discussion and managing the progress is difficult. Furthermore, managing presenters' time and meticulously recording important decisions are also critical challenges. Current meeting management systems cannot comprehensively address these issues, thus limiting their potential for improving factory productivity.

[1374] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1375] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected topics that do not match, means for recording the final decisions and decision-makers, means for monitoring the progress of the meeting in real time and managing the progress based on a set process plan, means for monitoring the presenter's speaking time and issuing a warning a certain time before a set time, means for detecting deviations and issuing a warning if the detected content does not match the process plan, and means for installing the progress management system on a smartphone. This enables efficient progress management, time management, and complete recording of important decisions for discussions and meetings within the factory.

[1376] "Means for acquiring audio data" refers to equipment and systems for acquiring audio from meetings and discussions in real time and transmitting it to a server.

[1377] "Means for converting acquired audio data into text data" refers to software or hardware that uses speech recognition technology to convert audio data into text information.

[1378] "Means of analyzing converted text data to monitor the progress of a meeting" refers to algorithms and programs that analyze text data to identify speakers and monitor the progress of a meeting in real time.

[1379] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system or device that measures the elapsed time from the start of the presenter's speech and generates a warning and notification when the set time is exceeded.

[1380] "Means for comparing the content of discussions with the agenda and detecting inconsistent topics" refers to a text analysis system that analyzes the content of discussions based on the purpose and agenda of a meeting and automatically detects topics that do not match the agenda.

[1381] "Means of issuing warnings based on detected mismatched topics" refers to a warning system that notifies users when a discussion deviates from the agenda.

[1382] "Means for recording final decisions and decision-makers" refers to a data recording system for organizing and preserving information about the decisions made during a meeting and the individuals who made those decisions.

[1383] "A means of monitoring the progress of a meeting in real time and managing its progress based on a set schedule" refers to a system for tracking the progress of meetings and discussions in real time and managing their progress according to a pre-set schedule.

[1384] "A means of monitoring the presenter's speaking time and issuing a warning a certain amount of time before the set time" refers to a system that issues a warning in advance when the scheduled speaking time approaches.

[1385] "Means for detecting deviations and issuing warnings when detected content does not match the process plan" refers to a system that detects when meeting content deviates from a pre-set process plan and issues a warning.

[1386] "Methods for installing a progress management system on a smartphone" refers to providing application software for progress management in a form that runs on a smartphone, and the methods and procedures for installing it.

[1387] This invention provides a system for efficiently managing the progress of discussions and meetings held within a factory, managing time, detecting deviations from the discussion, and recording final decisions. To achieve this, it employs a configuration that combines a server, a speech recognition engine, a text analysis system, a timekeeping function, and a warning notification system.

[1388] Hardware and software configuration

[1389] Hardware:

[1390] Smartphone: Equipped with a microphone for acquiring voice and a display for displaying warnings.

[1391] Server: A computer system used for storing and processing data.

[1392] software:

[1393] Python: A programming language for building the system of the present invention.

[1394] SpeechRecognition: A library for converting speech data into text.

[1395] threading: A multithreaded library for timekeeping and warning generation.

[1396] Process Overview

[1397] The server includes the following measures:

[1398] 1. Means of acquiring audio data:

[1399] The smartphone's microphone captures the meeting's audio data in real time and sends it to the server. The server continuously receives and stores this audio data.

[1400] 2. Means for converting acquired audio data into text data:

[1401] The server converts the received audio data into text data using the SpeechRecognition library. This conversion process is performed automatically, and the generated text data is then analyzed.

[1402] 3. Means of analyzing converted text data to monitor the progress of the meeting:

[1403] The server analyzes text data, identifies speakers, and monitors the progress of the meeting in real time. This allows for verification that the discussion is proceeding according to plan.

[1404] 4. A means of monitoring the presenter's speaking time and issuing a warning if the set time limit is exceeded:

[1405] The system monitors the elapsed time from the speaker's start and uses a timer function to issue a warning if the set time is exceeded. A warning message is displayed on the smartphone's screen, and audio notifications are provided as needed.

[1406] 5. Means for comparing the content of the discussion with the agenda and detecting inconsistencies:

[1407] The server compares the text data with the agenda and automatically detects any mismatched topics. This prevents the discussion from going in an unintended direction.

[1408] 6. Means of issuing warnings based on detected mismatched topics:

[1409] If mismatched topics are detected, the server will alert the facilitator and prompt them to quickly return the discussion to the agenda.

[1410] 7. Means for recording final decisions and decision-makers:

[1411] The server monitors statements containing keywords such as "conclusion" or "decision" in real time, and records their content and the information of the decision-maker. This information is stored in a format that can be referenced later.

[1412] 8. Means for monitoring the progress of meetings in real time and managing their progress based on the set schedule:

[1413] Real-time progress management is performed based on the established project plan, and the progress of discussions is monitored to ensure that they proceed as planned.

[1414] 9. A means of monitoring the presenter's speaking time and issuing a warning a certain period of time before the set end time:

[1415] By issuing a warning a certain amount of time in advance as the scheduled speaking time approaches, this system helps speakers organize their information within the allotted time.

[1416] 10. Means for detecting derailments and issuing warnings when the detected content does not match the process plan:

[1417] If a statement is made that does not conform to the project plan, the server will automatically detect the derailment and issue a warning.

[1418] 11. How to install the progress management system on a smartphone:

[1419] We will install the progress management system on smartphones and provide it to users so that they can easily use it.

[1420] Specific example

[1421] For example, when discussions are held within a factory about introducing a new process, a progress management assistant app can be used to manage time, prevent discussions from going off track, and record important decisions in real time. This app is installed on a smartphone and can be used from anywhere within the factory.

[1422] Example of a prompt

[1423] Requirements for a real-time meeting management system:

[1424] A feature that acquires audio data in real time and converts it to text using a speech recognition engine.

[1425] A function to identify the speaker and monitor the progress.

[1426] A feature that monitors speaking time and generates a warning if the set time limit is exceeded.

[1427] A feature that compares the discussion content with the agenda and detects deviations from the topic.

[1428] A feature that records final decisions and the decision-makers in real time.

[1429] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1430] Step 1:

[1431] Acquisition of audio data

[1432] Input: Audio from meetings or discussions

[1433] Output: Audio data acquired in real time

[1434] Specific operation: The device (smartphone) acquires audio from meetings and discussions in real time via its microphone. The acquired audio data is sent directly to the server. During this process, the device maintains a continuous connection with the server to ensure a stable communication environment.

[1435] Step 2:

[1436] Text conversion of acquired audio data

[1437] Input: Acquired audio data

[1438] Output: Text data converted by speech recognition

[1439] Specific operation: The server uses the SpeechRecognition library to convert the acquired audio data into text data. In this process, the speech recognition engine analyzes the audio signal and generates the corresponding string. The converted text data is temporarily stored in the server's memory or database.

[1440] Step 3:

[1441] Monitoring the progress of the meeting

[1442] Input: Converted text data

[1443] Output: Record of progress, identification of speakers

[1444] Specific operation: The server analyzes the converted text data and monitors the progress of the meeting. Specifically, it automatically recognizes speakers from the content of their statements within the text data and records the presentation time and content of each speaker. This allows for real-time tracking of who said what and when.

[1445] Step 4:

[1446] Monitoring and warnings regarding the presenter's speaking time.

[1447] Input: Announcement start time and current time, set time limit

[1448] Output: Time-based warnings as needed

[1449] Specific operation: The server records the presenter's speaking start time and monitors the speaking time based on the set time limit. Using a timer function, it sends a warning to the terminal when the speaking time is a certain amount of time away from the set limit. The warning message is displayed on the terminal's screen and an audio notification is also provided.

[1450] Step 5:

[1451] Comparison of discussion content and agenda

[1452] Input: Parsed text data, pre-configured agenda

[1453] Output: Detection and recording of mismatched topics

[1454] Specific operation: The server compares the converted text data with a pre-configured agenda and automatically detects mismatched topics. A text analysis algorithm is used to determine whether specific keywords or phrases match the agenda. If mismatched topics are detected, their content is recorded, and an alert is issued to the facilitator.

[1455] Step 6:

[1456] Issuing warnings based on unrelated topics

[1457] Input: Detection results for mismatched topics

[1458] Output: Warning to the facilitator

[1459] Specific operation: When the server detects mismatched topics, it issues a warning to the facilitator. The warning is displayed as a text message on the terminal's screen, and an audio notification is also provided. This allows the facilitator to quickly understand that the discussion is deviating and take appropriate action.

[1460] Step 7:

[1461] Records of final decisions and decision-makers

[1462] Input: Converted text data, specific keywords (e.g., "conclusion," "decision")

[1463] Output: Records of decisions and decision-makers

[1464] Specific operation: The server monitors the converted text data in real time and detects statements containing keywords such as "conclusion" or "decision." Detected statements are considered decisions, and the information is recorded. Additionally, a prompt for the user to confirm the decision is displayed on the terminal, and the entered decision-maker information is also sent to the server and recorded.

[1465] Step 8:

[1466] Real-time monitoring of meeting progress and project plan progress management.

[1467] Input: Converted text data, configured process plan

[1468] Output: Progress management information

[1469] Specific operation: The server monitors the meeting progress in real time and manages its progress based on the configured schedule. It verifies that the progress is proceeding according to plan and issues warnings and alerts as needed. This function is crucial to ensuring that the meeting proceeds as planned.

[1470] Step 9:

[1471] Pre-scheduled warning of presenter's speaking time

[1472] Input: Start time of speaking and a certain period of time before the set time limit.

[1473] Output: Pre-warning

[1474] Specific operation: A warning about the remaining speaking time is sent to the device a certain amount of time after the scheduled start time. This allows the presenter to organize their remarks while being mindful of the remaining time.

[1475] Step 10:

[1476] Derailment detection and warning issuance for content that does not match the process plan.

[1477] Input: Converted text data, process plan

[1478] Output: Derailment detection results and warnings

[1479] Specific operation: The converted text data is compared with the process plan. If any discrepancies are detected, the server detects a deviation and issues a warning. The warning is displayed on the terminal's screen and an audio notification is provided to allow the facilitator to respond quickly.

[1480] Step 11:

[1481] Installing the progress management system on smartphones

[1482] Input: None (pre-configuration is assumed)

[1483] Output: Installed progress management systems

[1484] Specific actions: Install the progress management system on smartphones. Installation involves providing dedicated application software, making it easy for users to download and configure. This feature allows users to access the progress management system from anywhere within the factory.

[1485] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1486] This invention is a system that acquires audio data during a meeting and analyzes it in real time. Furthermore, by combining it with an emotion engine, it recognizes the emotions of participants and facilitates smoother meeting management.

[1487] Program Overview

[1488] This meeting management system can be implemented with the following program configuration.

[1489] 1. Acquisition of audio data

[1490] The device records the meeting audio in real time and streams the audio data to the server.

[1491] 2. Speech-to-text conversion

[1492] The server converts the received audio data into text data using a speech recognition engine.

[1493] 3. Monitoring the progress of the meeting

[1494] The server analyzes the converted text data, identifies speakers, and annotates the data to monitor the progress of the meeting.

[1495] 4. Timekeeping

[1496] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[1497] After a set amount of time has elapsed since the start of the announcement, the server generates an alert and sends it to the device.

[1498] The device will notify the presenter of the remaining time via audio notification and a display alert.

[1499] 5. Detecting Derailments in the Discussion

[1500] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[1501] The server generates a derailment detection warning and sends it to the terminal.

[1502] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[1503] 6. Emotion recognition by an emotion engine

[1504] The terminals and servers are equipped with emotion engines that recognize emotions in real time from the voice tone and content of what meeting participants say.

[1505] 7. Adjusting Emotion-Based Warnings

[1506] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[1507] 8. Emotion-based stress management

[1508] The emotion engine monitors the meeting facilitator's stress level in real time and provides alerts to reduce stress as needed.

[1509] 9. Records of final decisions and decision-makers

[1510] The server monitors keywords such as "conclusion" and "decision" in real time and records their content when detected.

[1511] The device displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's name.

[1512] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[1513] Specific example

[1514] Case Study 1: Timekeeping Alerts

[1515] 1. The user begins a 15-minute presentation.

[1516] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[1517] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[1518] 4. The device notifies the user that "3 minutes remaining."

[1519] Case Study 2: Detecting Derailments in Discussions

[1520] 1. The agenda is set to "Project A progress report".

[1521] 2. The user begins talking at length about "Project B".

[1522] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[1523] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[1524] 5. The device notifies the facilitator that "the topic has gone off track."

[1525] Case Study 3: Stress Management Using an Emotional Engine

[1526] 1. The facilitator begins to feel stressed while conducting the meeting.

[1527] 2. The emotion engine monitors the facilitator's stress level in real time.

[1528] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[1529] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[1530] The following describes the processing flow.

[1531] Step 1:

[1532] The user launches the meeting management application and clicks the "Start Meeting" button.

[1533] Step 2:

[1534] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[1535] Step 3:

[1536] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[1537] Step 4:

[1538] The device records the meeting audio in real time and streams the audio data to the server.

[1539] Step 5:

[1540] The server uses a speech recognition engine to convert the received audio data into text data.

[1541] Step 6:

[1542] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[1543] Step 7:

[1544] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[1545] Step 8:

[1546] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[1547] Step 9:

[1548] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[1549] Step 10:

[1550] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[1551] Step 11:

[1552] If a derailment is detected, the server generates a warning and sends it to the terminal.

[1553] Step 12:

[1554] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[1555] Step 13:

[1556] An emotion engine installed on the server analyzes the voice tone and content of meeting participants in real time to recognize their emotions.

[1557] Step 14:

[1558] The server adjusts the warning content and notification tone for presenters and participants based on the emotional state obtained from the emotion engine.

[1559] Step 15:

[1560] If the server, based on the results of its emotion engine's emotional state monitoring, determines that the facilitator's stress level is high, it generates an alert saying, "Let's take a short break," and sends it to the terminal.

[1561] Step 16:

[1562] The device provides the facilitator with an audio notification and a display alert saying, "We recommend you take a short break."

[1563] Step 17:

[1564] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[1565] Step 18:

[1566] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[1567] Step 19:

[1568] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[1569] Step 20:

[1570] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[1571] (Example 2)

[1572] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1573] Many challenges exist in managing the progress of meetings. Existing systems simply converted audio to text and monitored the agenda, resulting in insufficient management of speaking time, prevention of discussion digressions, and recognition of participants' emotions. Furthermore, stress management for meeting facilitators was not considered. A system that comprehensively addresses these challenges is needed.

[1574] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1575] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting mismatched topics, means for issuing a warning based on the detected mismatched topics, means for recognizing the emotions of participants and adjusting the content of warnings and notification tones based on their emotional state, means for monitoring the stress level of the meeting facilitator in real time and providing alerts for stress reduction as needed, and means for recording final decisions and the decision-makers. This makes it possible to comprehensively and in real time perform everything from acquiring audio data to text conversion, monitoring progress, timekeeping, detecting deviations from the discussion, emotion recognition, stress management, and recording final decisions.

[1576] "Audio data" refers to digital data in which sound is recorded electronically.

[1577] "Text data" refers to digital data obtained by converting audio data into written text.

[1578] "Progress status" refers to information indicating the current state of the meeting's progress.

[1579] "Speaking time" refers to the time elapsed since the presenter began speaking.

[1580] An "agenda" is a list of topics or subjects that should be discussed at a meeting.

[1581] A "topic" refers to a specific subject or theme that is discussed in a meeting.

[1582] A "warning" is a notice issued when a set condition is not met.

[1583] "Emotions" refers to the psychological and emotional state of the meeting participants.

[1584] "Emotional state" refers to the specific expression of a meeting participant's emotions at a particular moment.

[1585] "Notification tone" refers to the sound or tone used when issuing warnings or notifications.

[1586] "Stress level" indicates the degree of mental and physical burden on the meeting organizer.

[1587] An "alert" is a warning or notice that appears when certain conditions occur.

[1588] "Decisions made" refer to the specific details and action plans that were ultimately decided upon at the meeting.

[1589] A "decision-maker" is the person or position that ultimately makes the final decision on a particular matter in a meeting.

[1590] Modes for carrying out the invention

[1591] This invention is a system that acquires audio data during a meeting and analyzes it in real time. Furthermore, by combining it with an emotion engine, it recognizes the emotions of participants and facilitates smoother meeting management. This system includes the following components:

[1592] Acquisition of audio data

[1593] When a user starts a meeting, the device uses its microphone to record the meeting's audio data in real time and streams that audio data to the server. The audio data is processed in digital format and transferred to the server via the communication network.

[1594] Speech-to-text conversion

[1595] The server uses a speech recognition engine such as Google Cloud Speech-to-Text to convert the received audio data into text data. The converted text data is temporarily stored on the server and used for subsequent data analysis.

[1596] Monitoring the progress of the meeting

[1597] The server uses a natural language processing engine, such as the Google Cloud Natural Language API, to analyze the text data converted from speech and identify the speaker. It adds annotations to the text data, including the speaker's ID and the content of their speech, and monitors the progress of the meeting.

[1598] Timekeeping

[1599] The server records the start time of each user's speech and monitors the speech duration in real time. Once the set time has elapsed, the server generates an alert and sends it to the terminal. The terminal then notifies the presenter of the remaining time via audio notification and a display alert.

[1600] Derailment detection in discussions

[1601] The server compares the converted text data with pre-configured agenda items and recognizes any discussion of a mismatched topic as a digression. When a digression is detected, the server generates a warning and sends it to the terminal. The terminal then notifies the facilitator with an audio and visual alert stating, "The topic has gone off-topic."

[1602] Emotion recognition by an emotion engine

[1603] The terminals and servers are equipped with emotion engines such as the Azure Cognitive Services Emotion API to analyze the voice tone and content of meeting participants. This makes it possible to recognize participants' emotions (e.g., joy, anger, surprise) in real time.

[1604] Adjusting emotion-based warnings

[1605] The server receives the emotional state recognized by the emotion engine and adjusts the warning content and notification tone accordingly. For example, if a participant is feeling stressed, the alert tone will be softened and the notification content will be made gentler.

[1606] Emotion-based stress management

[1607] The emotion engine monitors the meeting facilitator's stress level in real time and provides this information to the server. If the server determines that the facilitator's stress level is high, it generates alerts for breaks or urgent discussions. For example, it might send a message to the facilitator such as, "We recommend you take a short break."

[1608] Records of final decisions and decision-makers

[1609] The server monitors keywords such as "conclusion" and "decision" in the conversation in real time. When a keyword is detected, its content is immediately recorded. The terminal displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's information. The entered decision-maker information is sent from the terminal to the server and recorded along with the final decision.

[1610] Specific example

[1611] Case Study 1: Timekeeping Alerts

[1612] 1. The user begins a 15-minute presentation.

[1613] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[1614] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[1615] 4. The device notifies the user that "3 minutes remaining."

[1616] Prompt example

[1617] "Design a system that records the start time of a 15-minute presentation and generates a '3 minutes remaining' alert 12 minutes later."

[1618] Case Study 2: Detecting Derailments in Discussions

[1619] 1. The agenda is set to "Project progress report".

[1620] 2. The user starts talking at length about "another project."

[1621] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[1622] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[1623] 5. The device notifies the facilitator that "the topic has gone off track."

[1624] Prompt example

[1625] "Design a system that detects off-topic comments during meetings and alerts the facilitator to the derailment."

[1626] Case Study 3: Stress Management Using an Emotional Engine

[1627] 1. The facilitator begins to feel stressed while conducting the meeting.

[1628] 2. The emotion engine monitors the facilitator's stress level in real time.

[1629] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[1630] Prompt example

[1631] "Design a system that monitors the facilitator's stress level during meetings and sends notifications prompting them to take breaks when necessary."

[1632] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[1633] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1634] Step 1:

[1635] The user starts the meeting.

[1636] Specifically, the user launches the meeting application and starts the meeting by clicking the "Start Meeting" button.

[1637] Input: User's meeting start operation

[1638] Output: Meeting start signal

[1639] Step 2:

[1640] The device prepares to record audio.

[1641] Specifically, the device checks the readiness status of the built-in or externally connected microphone and enables the recording function.

[1642] Input: Signal to start meeting

[1643] Output: Recording ready signal

[1644] Step 3:

[1645] The device records audio data in real time and streams it to the server.

[1646] Specifically, the device records the audio during the meeting in real time and streams it to a server via the internet as digital audio data.

[1647] Input: Signal indicating recording is ready.

[1648] Output: Streaming of real-time audio data

[1649] Step 4:

[1650] The server receives the audio data and uses a speech recognition engine to convert the audio data into text data.

[1651] In terms of specific operations, the server uses a speech recognition engine to analyze the audio data and convert it into text data written as characters. A typical example of the software used is a general-purpose speech recognition engine.

[1652] Input: Real-time audio data

[1653] Output: Text data

[1654] Step 5:

[1655] The server temporarily stores the converted text data.

[1656] Specifically, the server temporarily stores the converted text data in storage such as a database.

[1657] Input: Text data

[1658] Output: Temporarily saved text data

[1659] Step 6:

[1660] The server analyzes the text data converted from the speech and monitors the progress using a natural language processing engine.

[1661] In terms of specific operations, the server uses a natural language processing engine to analyze text data, identify the speaker, and annotate the content of the speech. The software used includes a natural language processing engine, among other things.

[1662] Input: Temporarily saved text data

[1663] Output: Annotated text data

[1664] Step 7:

[1665] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[1666] Specifically, the server records the start time of a message and then continuously measures the duration of subsequent messages.

[1667] Input: Annotated text data

[1668] Output: Speech time elapsed data

[1669] Step 8:

[1670] After a set amount of time has elapsed since the start of the announcement, the server generates an alert and sends it to the device.

[1671] Specifically, when the server determines that a set time has elapsed, it generates an alert message and sends it to the terminal.

[1672] Input: Elapsed time of speech

[1673] Output: Alert

[1674] Step 9:

[1675] The device will notify the presenter of the remaining time via audio notification and a display alert.

[1676] Specifically, the device will use voice notifications and screen displays based on the received alerts to inform the presenter of the remaining time.

[1677] Input: Alert

[1678] Output: Voice notifications and display alerts

[1679] Step 10:

[1680] The server compares the converted text data with the agenda and detects any mismatched topics.

[1681] Specifically, the server compares the text data with a pre-configured agenda list and automatically detects any mismatches.

[1682] Input: Annotated text data

[1683] Output: Detection results for mismatched topics

[1684] Step 11:

[1685] The server issues a warning based on mismatched topics and sends it to the terminal.

[1686] Specifically, the server generates a warning message based on the detected mismatched topics and sends it to the terminal.

[1687] Input: Detection results for mismatched topics

[1688] Output: Warning message

[1689] Step 12:

[1690] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[1691] Specifically, the device uses the received warning message to provide audio notifications and screen displays, informing the facilitator that the topic has gone off track.

[1692] Input: Warning message

[1693] Output: Voice notifications and display alerts

[1694] Step 13:

[1695] The terminals and servers are equipped with emotion engines to recognize the emotions of the participants.

[1696] Specifically, the terminal and server use an emotion engine to analyze the participant's emotions in real time based on their voice tone and what they say.

[1697] Input: Real-time audio data

[1698] Output: Emotion recognition result

[1699] Step 14:

[1700] The server adjusts warning content and notification tone based on the emotional state recognized by the emotion engine.

[1701] Specifically, the server adjusts the content and tone of the warning message based on the emotion recognition results.

[1702] Input: Sentiment recognition result

[1703] Output: Adjusted warning message

[1704] Step 15:

[1705] The emotion engine monitors the stress levels of meeting organizers, and the server provides alerts to reduce stress as needed.

[1706] Specifically, the emotion engine constantly analyzes the operator's stress level and generates alerts for breaks and relaxation at appropriate times.

[1707] Input: Real-time audio data and operator sentiment data

[1708] Output: Stress reduction alert

[1709] Step 16:

[1710] The server monitors for "conclusion" and "decision" keywords and records their content.

[1711] Specifically, the server monitors for specific keywords within the converted text data in real time and records their content in the database if detected.

[1712] Input: Annotated text data

[1713] Output: Recorded decisions

[1714] Step 17:

[1715] The terminal prompts the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[1716] Specifically, after a decision is finalized, the terminal displays a pop-up notification prompting the user to enter the decision-maker's information. The entered information is sent to the server and recorded.

[1717] Input: Decision made and user-inputted decision-maker information

[1718] Output: Recorded decision-maker information

[1719] (Application Example 2)

[1720] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1721] Meetings during shift changes in factories require smooth information transfer and efficient discussion. However, insufficient communication on the factory floor and frequent digressions in discussions often lead to problems with efficient shift changes and work handovers. Furthermore, the emotional state of participants significantly impacts the progress of the meeting, making emotional management a crucial issue. A system is needed to appropriately address these problems.

[1722] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting inconsistent topics, means for issuing a warning based on the detected inconsistent topics, and means for recognizing the emotions of meeting participants and adjusting the content of the warning and the tone of the notification based on their emotional state. This enables appropriate information transmission, efficient progress of discussions, and emotional management of participants in factory shift change meetings.

[1723] "Means for acquiring audio data" refers to a device or software for recording speech spoken in real time during meetings or shift changes and saving it as digital data.

[1724] "Means for converting acquired audio data into text data" refers to a speech recognition engine or software that analyzes recorded audio data and automatically converts it into text data.

[1725] "Means of analyzing converted text data to monitor the progress of a meeting" refers to a program or system that analyzes text data to understand who is speaking and the progress of the agenda in real time.

[1726] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system with a timekeeping function that records the start time of the presentation and generates an alert when the set time limit is exceeded.

[1727] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a program or algorithm that compares the content of a statement with a set agenda and detects when a topic that does not match the agenda is mentioned.

[1728] "Means for issuing warnings based on detected mismatched topics" refers to a system that automatically generates warnings related to detected mismatched topics and notifies speakers and moderators.

[1729] "Means for recording final decisions and decision-makers" refers to a system that has the function of recording decisions made during meetings or shift changes, as well as the decision-makers involved, and saving them in a database for later reference.

[1730] "Means for recognizing the emotions of meeting participants and adjusting the tone of warnings and notifications based on their emotional state" refers to an algorithm or engine that analyzes emotions in real time from the tone of voice and content of statements of participants, and changes the content and expression of warnings and notifications according to their emotional state.

[1731] The system that implements this application supports meetings during shift changes in a factory. This system acquires audio data in real time, converts it into text data, and performs meeting progress monitoring, timekeeping, discussion derailment detection, and sentiment recognition.

[1732] Specifically, the program will be implemented with the following structure:

[1733] Hardware and software

[1734] Hardware to use:

[1735] Microphone: For capturing the voices of meeting participants.

[1736] Robot or smartphone: A platform for processing and notifying voice data.

[1737] Software to use:

[1738] speech_recognition: A library for performing speech recognition.

[1739] emotion_recognition: A library for performing emotion recognition (uses a pre-trained model).

[1740] datetime: A standard library for time management.

[1741] Process Overview

[1742] 1. Acquisition of audio data:

[1743] The audio of meeting participants can be captured in real time via microphones.

[1744] The voice data is transferred to a robot or smartphone platform.

[1745] 2. Speech-to-text conversion:

[1746] The speech_recognition library installed on the robot or smartphone is used to convert speech data into text data.

[1747] The converted text is sent to the server.

[1748] 3. Monitoring the progress of the meeting:

[1749] The server analyzes the converted text data to identify the speaker and monitor the progress.

[1750] Verify that the meeting is proceeding according to the set agenda.

[1751] 4. Timekeeping:

[1752] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[1753] If the presentation time exceeds the set time, a warning is sent from the server to the robot or smartphone, and the presenter is notified via audio notification and a display alert.

[1754] 5. Detecting deviations from the discussion:

[1755] The server compares the discussion to the agenda and detects deviations if a topic that does not match is mentioned.

[1756] If a derailment is detected, a warning will be displayed to the facilitator.

[1757] 6. Emotion recognition by the emotion engine:

[1758] The robot or smartphone is equipped with an emotion engine that recognizes the participant's emotions in real time based on their voice tone and what they say.

[1759] This allows for monitoring the emotional state of participants and taking appropriate action.

[1760] 7. Adjusting emotion-based warnings:

[1761] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[1762] For example, if emotions are heightened, the notification will be delivered in a calm tone.

[1763] Specific example

[1764] Examples of support for shift change meetings

[1765] 1. Start of the meeting:

[1766] When the shift changes, the leader begins explaining the situation to the next team.

[1767] The microphone acquires audio data in real time and transmits it to the robot or smartphone.

[1768] 2. Speech-to-text conversion:

[1769] The acquired audio data is converted to text using the speech_recognition library and sent to the server.

[1770] 3. Monitoring the progress of the meeting and keeping time:

[1771] The server analyzes the converted text data to identify the speaker, monitor the progress, and keep time.

[1772] After a set amount of time has elapsed since the announcement began, a warning will be sent to the robot or smartphone, triggering an audio notification and a visual alert.

[1773] 4. Detecting deviations from the discussion:

[1774] The server compares the discussion content with the agenda and warns the facilitator if it detects a deviation from the topic.

[1775] 5. Recognition by the emotion engine:

[1776] It monitors future emotional states in real time and adjusts warning content and notification tones as needed.

[1777] Examples of prompt statements to use

[1778] "Record the 5-minute shift change meeting and convert it to text. Perform sentiment analysis using an emotion recognition model, and implement tactical detection and timekeeping."

[1779] This system allows factory shift change meetings to proceed smoothly and efficiently, and because it takes into account the emotional state of the participants, it enables more effective information sharing and discussion.

[1780] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1781] Step 1:

[1782] Acquire audio data.

[1783] A microphone connected to a terminal (robot or smartphone) records the speech of meeting participants in real time. The input is the raw voice of the meeting participants, and the output is digitized audio data. This audio data is streamed to a server.

[1784] Step 2:

[1785] Convert audio data to text data.

[1786] The audio data sent to the server is automatically converted into text data using the speech_recognition library. The input is audio data, and the output is text data in string format. This text data is then passed on to the next parsing step.

[1787] Step 3:

[1788] Monitor the progress of the meeting.

[1789] The server analyzes text data to identify the content and speaker of a statement. The input is text data, and the output is a pair of speaker and statement information. Based on this, it identifies speakers and generates progress graphs.

[1790] Step 4:

[1791] Timekeeping will be performed.

[1792] The server records the presenter's start time and measures the elapsed time in real time. When the set presentation time has elapsed, it generates a warning and sends it to the terminal. The input is the start time, and the output is the elapsed time and the warning message. The terminal notifies the presenter with an audio notification and a display alert.

[1793] Step 5:

[1794] Detects deviations from the discussion.

[1795] The server compares the text data with the previously set agenda and detects any mismatched topics. The input is the text data and agenda information, and the output is a list of mismatched topics. If a digression is detected, the server sends a warning to the facilitator.

[1796] Step 6:

[1797] Recognize your emotions.

[1798] The emotion engine built into the device recognizes the participant's emotions in real time based on their voice tone and speech content. Input consists of voice and text data, and output is the identified emotional state. This information is sent to the server.

[1799] Step 7:

[1800] The tone of warnings and notifications is adjusted based on emotions.

[1801] The server dynamically adjusts the content of warnings and notification tones based on the emotional state obtained from the emotion engine. The input is emotional state data, and the output is the adjusted warning message. For example, if a participant is stressed, the notification will be delivered in a calm tone.

[1802] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1803] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1804] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1805] [Fourth Embodiment]

[1806] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1807] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1808] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1809] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1810] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1811] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1812] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1813] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1814] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1815] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1816] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1817] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1818] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1819] This invention is a system for a meeting management system that analyzes audio data in real time to appropriately manage the speaking time of presenters and the progress of discussions.

[1820] Program Overview

[1821] The meeting management system will be implemented with the following program structure.

[1822] 1. Acquisition of audio data

[1823] The device records the meeting audio in real time and streams it to the server.

[1824] 2. Speech-to-text conversion

[1825] The server converts the received audio data into text data using a speech recognition engine.

[1826] 3. Monitoring the progress of the meeting

[1827] The server analyzes the converted text data, automatically recognizes the speaker, and monitors the progress of the meeting.

[1828] 4. Timekeeping

[1829] The server monitors the presenter's speaking time and generates a warning if the set time limit is exceeded.

[1830] The device displays an alert to the presenter and provides an audio notification.

[1831] 5. Detecting Derailments in the Discussion

[1832] The server compares the discussion content with the agenda and generates a warning if a topic that does not match is mentioned.

[1833] The device displays a warning to the facilitator and provides an audio notification.

[1834] 6. Records of final decisions and decision-makers

[1835] The server monitors and detects statements containing keywords such as "conclusion" and "decision" in real time, and records their content.

[1836] The device displays an alert prompting the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[1837] Specific example

[1838] Case Study 1: Timekeeping Alerts

[1839] 1. User A begins a 15-minute presentation.

[1840] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[1841] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[1842] 4. The device displays and announces to user A that "3 minutes remaining."

[1843] 5. User A completes the presentation within the allotted time.

[1844] Case Study 2: Detecting Derailments in Discussions

[1845] 1. The meeting agenda is set to "Progress report for Project A".

[1846] 2. User B begins talking at length about "Project B".

[1847] 3. The server analyzes the discussion content as text and detects any discrepancies with the agenda.

[1848] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[1849] 5. The device displays and provides an audio notification to the facilitator stating, "The topic has gone off track. Let's return to the agenda."

[1850] 6. The facilitator encourages the discussion to return to the agenda.

[1851] Case Study 3: Records of Final Decisions and Decision-makers

[1852] 1. During the meeting, comments containing keywords such as "conclusion" are made.

[1853] 2. The server detects the statement and records its content.

[1854] 3. The server sends an alert to the terminal prompting the final decision-maker to confirm.

[1855] 4. The terminal displays the message "Please enter the final decision-maker," and User C enters the decision-maker.

[1856] 5. The entered decision-maker information is sent to the server and recorded along with the decision.

[1857] As described above, this system offers the advantage of enabling junior facilitators to efficiently manage meetings and facilitate time management and agenda achievement.

[1858] The following describes the processing flow.

[1859] Step 1:

[1860] The user launches the meeting management application and clicks the "Start Meeting" button.

[1861] Step 2:

[1862] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[1863] Step 3:

[1864] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[1865] Step 4:

[1866] The device records the meeting audio in real time and streams the audio data to the server.

[1867] Step 5:

[1868] The server uses a speech recognition engine to convert the received audio data into text data.

[1869] Step 6:

[1870] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[1871] Step 7:

[1872] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[1873] Step 8:

[1874] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[1875] Step 9:

[1876] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[1877] Step 10:

[1878] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[1879] Step 11:

[1880] If a derailment is detected, the server generates a warning and sends it to the terminal.

[1881] Step 12:

[1882] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[1883] Step 13:

[1884] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[1885] Step 14:

[1886] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[1887] Step 15:

[1888] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[1889] Step 16:

[1890] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[1891] The above outlines the specific processing steps and details of the actions performed at each step. This system will enable junior facilitators to manage meetings efficiently and effectively.

[1892] (Example 1)

[1893] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1894] Traditional meeting management systems struggle to efficiently manage speaking time and facilitate discussions, leaving younger facilitators, in particular, lacking the tools to ensure smooth meetings. Specifically, they are unable to effectively perform real-time analysis of meeting audio data, monitor presenter speaking time, detect tangents, and record final decisions and decision-makers, resulting in a decline in meeting quality.

[1895] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1896] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected mismatched topics, means for detecting statements containing keywords such as "conclusion" or "decision" and recording their content, and means for prompting confirmation of decision-maker information. This makes it possible to appropriately manage the presenter's speaking time, detect digressions in the discussion, and ensure the smooth progress of the meeting.

[1897] "Means for acquiring audio data" refers to a device or function for recording audio during a meeting in real time and streaming it to a server.

[1898] "Means for converting audio data to text data" refers to a device or function that converts received audio data into corresponding text data using speech recognition technology.

[1899] "Means for analyzing converted text data to monitor the progress of a meeting" refers to a device or function that analyzes text data converted by speech recognition to understand the content of statements and speakers, and to monitor the progress of a meeting.

[1900] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a device or function that measures the elapsed time from the moment the presenter starts speaking and generates a warning if the set time limit is exceeded.

[1901] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a device or function that compares a pre-set agenda with the actual content of a discussion and detects when a topic not included in the agenda is mentioned.

[1902] "Means for issuing warnings based on detected mismatched topics" refers to a device or function that warns that the discussion is going off track when a topic that does not match the agenda is detected.

[1903] "Means for detecting and recording statements containing keywords such as 'conclusion' or 'decision'" refers to a device or function that detects statements containing specific keywords in real time from text data during a meeting and records the content of those statements in a database or similar.

[1904] "Means for prompting confirmation of decision-maker information" refers to a device or function that generates an alert prompting the user to confirm the decision-maker's information regarding the detected final decision, and prompts the user to input that information.

[1905] This invention relates to a system for a meeting management system that analyzes audio data in real time to appropriately manage presenter speaking time and facilitate discussion. This system is primarily implemented through the interaction of a server, terminals, and users.

[1906] First, the device records the meeting audio in real time using its built-in or external microphone as soon as the meeting starts. Standard audio formats (e.g., WAV, MP3) are used for recording. This recorded audio data is transmitted to the server via streaming at regular intervals (e.g., every second).

[1907] The server uses a storage service (e.g., Amazon S3, Google Cloud Storage) to temporarily store the received audio data. Next, the server uses a speech recognition engine, such as the Google Cloud Speech-to-Text API, to convert the audio data into corresponding text data. This converted text data is then stored in a database system (e.g., MySQL).

[1908] Furthermore, the server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the converted text data. This analysis monitors the progress of the meeting and identifies speakers. Identification is based on speaker-specific language patterns and vocal characteristics, and the results are stored in a database.

[1909] Regarding the management of presenter speaking time, the server records the meeting start time and presentation start time and sets a timer. It monitors speaking time, and if the set time is exceeded, the server automatically generates a warning. The server then sends the corresponding warning to the terminal, which alerts the user with a display and audio notification such as "3 minutes remaining."

[1910] For managing the progress of the discussion, the server compares the pre-configured agenda with the actual discussion content. If the server detects a discussion that does not match the agenda, it automatically generates a "discussion derailment" warning and sends it to the terminal. The terminal displays and provides an audio notification to the facilitator with a warning such as, "The topic has gone off track. Let's get back to the agenda."

[1911] To record final decisions, the server detects in real time any statements containing keywords such as "conclusion" or "decision" during the meeting. The detected statements are recorded in the database, and the server then sends an alert to the terminal prompting confirmation of the decision-maker's information. The terminal displays a message such as "Please enter the final decision-maker," and the user enters the decision-maker's information, which is then sent to the server and stored in the database.

[1912] This system enables efficient management of meeting time and agenda completion, and particularly allows junior facilitators to conduct meetings smoothly.

[1913] The following scenarios can be considered as concrete examples:

[1914] Example prompt: "Please describe the processing flow of a system that converts audio data into text data in real time and manages speaking time."

[1915] The above describes specific embodiments for carrying out the present invention.

[1916] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1917] Step 1: Acquire audio data

[1918] The device uses its microphone to record the meeting audio and starts recording in real time as soon as the meeting begins.

[1919] The device streams the recorded audio data to the server at regular time intervals (e.g., every second).

[1920] Input: Real-time audio data from the meeting

[1921] Output: Audio data to be streamed.

[1922] Step 2: Save the audio file.

[1923] The server temporarily stores the audio data received from the terminal in a storage service (e.g., Amazon S3, Google Cloud Storage).

[1924] Input: Streaming transmitted audio data

[1925] Output: Temporarily saved audio data

[1926] Step 3: Convert speech to text

[1927] The server uses the Google Cloud Speech-to-Text API to convert the temporarily stored audio data into text data.

[1928] Input: Temporarily stored audio data

[1929] Output: Converted text data

[1930] Step 4: Save text data

[1931] The server saves the converted text data to a database system (e.g., MySQL).

[1932] Input: Converted text data

[1933] Output: Text data stored in the database

[1934] Step 5: Analyze the progress of the meeting

[1935] The server uses natural language processing tools (e.g., Spacy, NLTK) to analyze the stored text data.

[1936] Identify the progress of the meeting and the speakers.

[1937] Input: Text data stored in the database

[1938] Output: Analyzed meeting progress and speaker information

[1939] Step 6: Monitor speaking time

[1940] The server records the start time of the presentation and sets a timer.

[1941] The system monitors the presenter's speaking time in real time and generates a warning if the set time limit is exceeded.

[1942] Input: Analyzed meeting progress

[1943] Output: Warning based on excess time

[1944] Step 7: Warning Notification

[1945] The server sends the generated warning to the terminal.

[1946] The device will display a warning message as a pop-up and provide an audio notification.

[1947] Input: Generated warning

[1948] Output: Warning messages and audio notifications displayed on the device

[1949] Step 8: Detecting Derailments in the Discussion

[1950] The server compares the meeting discussion content with a pre-configured agenda.

[1951] A warning is generated if a topic unrelated to the agenda is mentioned.

[1952] Input: Analyzed meeting progress, agenda

[1953] Output: Warning based on mismatched topics

[1954] Step 9: Derailment warning notification

[1955] The server sends a derailment warning to the terminal.

[1956] The device displays a derailment warning to the facilitator and provides an audio notification.

[1957] Input: Generated derailment warning

[1958] Output: Derailment warning message and audio notification displayed on the device.

[1959] Step 10: Record and confirm decisions

[1960] The server detects in real time any statements containing keywords such as "conclusion" or "decision," and records their content.

[1961] The server sends an alert to the terminal prompting the decision-maker to confirm.

[1962] The user enters decision-maker information, and the terminal sends the entered information to the server.

[1963] Input: Analyzed text data, decision-maker information

[1964] Output: Recorded decisions and decision-maker information

[1965] The above outlines the specific processing steps of this system. This enables efficient real-time management and progress of meetings.

[1966] (Application Example 1)

[1967] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1968] Efficient progress management is required in discussions and meetings within factories, but preventing deviations from the discussion and managing the progress is difficult. Furthermore, managing presenters' time and meticulously recording important decisions are also critical challenges. Current meeting management systems cannot comprehensively address these issues, thus limiting their potential for improving factory productivity.

[1969] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1970] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the presenter's speaking time and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting topics that do not match, means for issuing a warning based on the detected topics that do not match, means for recording the final decisions and decision-makers, means for monitoring the progress of the meeting in real time and managing the progress based on a set process plan, means for monitoring the presenter's speaking time and issuing a warning a certain time before a set time, means for detecting deviations and issuing a warning if the detected content does not match the process plan, and means for installing the progress management system on a smartphone. This enables efficient progress management, time management, and complete recording of important decisions for discussions and meetings within the factory.

[1971] "Means for acquiring audio data" refers to equipment and systems for acquiring audio from meetings and discussions in real time and transmitting it to a server.

[1972] "Means for converting acquired audio data into text data" refers to software or hardware that uses speech recognition technology to convert audio data into text information.

[1973] "Means of analyzing converted text data to monitor the progress of a meeting" refers to algorithms and programs that analyze text data to identify speakers and monitor the progress of a meeting in real time.

[1974] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system or device that measures the elapsed time from the start of the presenter's speech and generates a warning and notification when the set time is exceeded.

[1975] "Means for comparing the content of discussions with the agenda and detecting inconsistent topics" refers to a text analysis system that analyzes the content of discussions based on the purpose and agenda of a meeting and automatically detects topics that do not match the agenda.

[1976] "Means of issuing warnings based on detected mismatched topics" refers to a warning system that notifies users when a discussion deviates from the agenda.

[1977] "Means for recording final decisions and decision-makers" refers to a data recording system for organizing and preserving information about the decisions made during a meeting and the individuals who made those decisions.

[1978] "A means of monitoring the progress of a meeting in real time and managing its progress based on a set schedule" refers to a system for tracking the progress of meetings and discussions in real time and managing their progress according to a pre-set schedule.

[1979] "A means of monitoring the presenter's speaking time and issuing a warning a certain amount of time before the set time" refers to a system that issues a warning in advance when the scheduled speaking time approaches.

[1980] "Means for detecting deviations and issuing warnings when detected content does not match the process plan" refers to a system that detects when meeting content deviates from a pre-set process plan and issues a warning.

[1981] "Methods for installing a progress management system on a smartphone" refers to providing application software for progress management in a form that runs on a smartphone, and the methods and procedures for installing it.

[1982] This invention provides a system for efficiently managing the progress of discussions and meetings held within a factory, managing time, detecting deviations from the discussion, and recording final decisions. To achieve this, it employs a configuration that combines a server, a speech recognition engine, a text analysis system, a timekeeping function, and a warning notification system.

[1983] Hardware and software configuration

[1984] Hardware:

[1985] Smartphone: Equipped with a microphone for acquiring voice and a display for displaying warnings.

[1986] Server: A computer system used for storing and processing data.

[1987] software:

[1988] Python: A programming language for building the system of the present invention.

[1989] SpeechRecognition: A library for converting speech data into text.

[1990] threading: A multithreaded library for timekeeping and warning generation.

[1991] Process Overview

[1992] The server includes the following measures:

[1993] 1. Means of acquiring audio data:

[1994] The smartphone's microphone captures the meeting's audio data in real time and sends it to the server. The server continuously receives and stores this audio data.

[1995] 2. Means for converting acquired audio data into text data:

[1996] The server converts the received audio data into text data using the SpeechRecognition library. This conversion process is performed automatically, and the generated text data is then analyzed.

[1997] 3. Means of analyzing converted text data to monitor the progress of the meeting:

[1998] The server analyzes text data, identifies speakers, and monitors the progress of the meeting in real time. This allows for verification that the discussion is proceeding according to plan.

[1999] 4. A means of monitoring the presenter's speaking time and issuing a warning if the set time limit is exceeded:

[2000] The system monitors the elapsed time from the speaker's start and uses a timer function to issue a warning if the set time is exceeded. A warning message is displayed on the smartphone's screen, and audio notifications are provided as needed.

[2001] 5. Means for comparing the content of the discussion with the agenda and detecting inconsistencies:

[2002] The server compares the text data with the agenda and automatically detects any mismatched topics. This prevents the discussion from going in an unintended direction.

[2003] 6. Means of issuing warnings based on detected mismatched topics:

[2004] If mismatched topics are detected, the server will alert the facilitator and prompt them to quickly return the discussion to the agenda.

[2005] 7. Means for recording final decisions and decision-makers:

[2006] The server monitors statements containing keywords such as "conclusion" or "decision" in real time, and records their content and the information of the decision-maker. This information is stored in a format that can be referenced later.

[2007] 8. Means for monitoring the progress of meetings in real time and managing their progress based on the set schedule:

[2008] Real-time progress management is performed based on the established project plan, and the progress of discussions is monitored to ensure that they proceed as planned.

[2009] 9. A means of monitoring the presenter's speaking time and issuing a warning a certain period of time before the set end time:

[2010] By issuing a warning a certain amount of time in advance as the scheduled speaking time approaches, this system helps speakers organize their information within the allotted time.

[2011] 10. Means for detecting derailments and issuing warnings when the detected content does not match the process plan:

[2012] If a statement is made that does not conform to the project plan, the server will automatically detect the derailment and issue a warning.

[2013] 11. How to install the progress management system on a smartphone:

[2014] We will install the progress management system on smartphones and provide it to users so that they can easily use it.

[2015] Specific example

[2016] For example, when discussions are held within a factory about introducing a new process, a progress management assistant app can be used to manage time, prevent discussions from going off track, and record important decisions in real time. This app is installed on a smartphone and can be used from anywhere within the factory.

[2017] Example of a prompt

[2018] Requirements for a real-time meeting management system:

[2019] A feature that acquires audio data in real time and converts it to text using a speech recognition engine.

[2020] A function to identify the speaker and monitor the progress.

[2021] A feature that monitors speaking time and generates a warning if the set time limit is exceeded.

[2022] A feature that compares the discussion content with the agenda and detects deviations from the topic.

[2023] A feature that records final decisions and the decision-makers in real time.

[2024] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[2025] Step 1:

[2026] Acquisition of audio data

[2027] Input: Audio from meetings or discussions

[2028] Output: Audio data acquired in real time

[2029] Specific operation: The device (smartphone) acquires audio from meetings and discussions in real time via its microphone. The acquired audio data is sent directly to the server. During this process, the device maintains a continuous connection with the server to ensure a stable communication environment.

[2030] Step 2:

[2031] Text conversion of acquired audio data

[2032] Input: Acquired audio data

[2033] Output: Text data converted by speech recognition

[2034] Specific operation: The server uses the SpeechRecognition library to convert the acquired audio data into text data. In this process, the speech recognition engine analyzes the audio signal and generates the corresponding string. The converted text data is temporarily stored in the server's memory or database.

[2035] Step 3:

[2036] Monitoring the progress of the meeting

[2037] Input: Converted text data

[2038] Output: Record of progress, identification of speakers

[2039] Specific operation: The server analyzes the converted text data and monitors the progress of the meeting. Specifically, it automatically recognizes speakers from the content of their statements within the text data and records the presentation time and content of each speaker. This allows for real-time tracking of who said what and when.

[2040] Step 4:

[2041] Monitoring and warnings regarding the presenter's speaking time.

[2042] Input: Announcement start time and current time, set time limit

[2043] Output: Time-based warnings as needed

[2044] Specific operation: The server records the presenter's speaking start time and monitors the speaking time based on the set time limit. Using a timer function, it sends a warning to the terminal when the speaking time is a certain amount of time away from the set limit. The warning message is displayed on the terminal's screen and an audio notification is also provided.

[2045] Step 5:

[2046] Comparison of discussion content and agenda

[2047] Input: Parsed text data, pre-configured agenda

[2048] Output: Detection and recording of mismatched topics

[2049] Specific operation: The server compares the converted text data with a pre-configured agenda and automatically detects mismatched topics. A text analysis algorithm is used to determine whether specific keywords or phrases match the agenda. If mismatched topics are detected, their content is recorded, and an alert is issued to the facilitator.

[2050] Step 6:

[2051] Issuing warnings based on unrelated topics

[2052] Input: Detection results for mismatched topics

[2053] Output: Warning to the facilitator

[2054] Specific operation: When the server detects mismatched topics, it issues a warning to the facilitator. The warning is displayed as a text message on the terminal's screen, and an audio notification is also provided. This allows the facilitator to quickly understand that the discussion is deviating and take appropriate action.

[2055] Step 7:

[2056] Records of final decisions and decision-makers

[2057] Input: Converted text data, specific keywords (e.g., "conclusion," "decision")

[2058] Output: Records of decisions and decision-makers

[2059] Specific operation: The server monitors the converted text data in real time and detects statements containing keywords such as "conclusion" or "decision." Detected statements are considered decisions, and the information is recorded. Additionally, a prompt for the user to confirm the decision is displayed on the terminal, and the entered decision-maker information is also sent to the server and recorded.

[2060] Step 8:

[2061] Real-time monitoring of meeting progress and project plan progress management.

[2062] Input: Converted text data, configured process plan

[2063] Output: Progress management information

[2064] Specific operation: The server monitors the meeting progress in real time and manages its progress based on the configured schedule. It verifies that the progress is proceeding according to plan and issues warnings and alerts as needed. This function is crucial to ensuring that the meeting proceeds as planned.

[2065] Step 9:

[2066] Pre-scheduled warning of presenter's speaking time

[2067] Input: Start time of speaking and a certain period of time before the set time limit.

[2068] Output: Pre-warning

[2069] Specific operation: A warning about the remaining speaking time is sent to the device a certain amount of time after the scheduled start time. This allows the presenter to organize their remarks while being mindful of the remaining time.

[2070] Step 10:

[2071] Derailment detection and warning issuance for content that does not match the process plan.

[2072] Input: Converted text data, process plan

[2073] Output: Derailment detection results and warnings

[2074] Specific operation: The converted text data is compared with the process plan. If any discrepancies are detected, the server detects a deviation and issues a warning. The warning is displayed on the terminal's screen and an audio notification is provided to allow the facilitator to respond quickly.

[2075] Step 11:

[2076] Installing the progress management system on smartphones

[2077] Input: None (pre-configuration is assumed)

[2078] Output: Installed progress management systems

[2079] Specific actions: Install the progress management system on smartphones. Installation involves providing dedicated application software, making it easy for users to download and configure. This feature allows users to access the progress management system from anywhere within the factory.

[2080] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[2081] This invention is a system that acquires audio data during a meeting and analyzes it in real time. Furthermore, by combining it with an emotion engine, it recognizes the emotions of participants and facilitates smoother meeting management.

[2082] Program Overview

[2083] This meeting management system can be implemented with the following program configuration.

[2084] 1. Acquisition of audio data

[2085] The device records the meeting audio in real time and streams the audio data to the server.

[2086] 2. Speech-to-text conversion

[2087] The server uses a speech recognition engine to convert the received audio data into text data.

[2088] 3. Monitoring the progress of the meeting

[2089] The server analyzes the converted text data, identifies speakers, and annotates the data to monitor the progress of the meeting.

[2090] 4. Timekeeping

[2091] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[2092] After a set amount of time has elapsed since the start of the announcement, the server generates an alert and sends it to the terminal.

[2093] The device will notify the presenter of the remaining time via audio notification and a display alert.

[2094] 5. Detecting Derailments in the Discussion

[2095] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[2096] The server generates a derailment detection warning and sends it to the terminal.

[2097] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[2098] 6. Emotion recognition by an emotion engine

[2099] The terminals and servers are equipped with emotion engines that recognize emotions in real time from the voice tone and content of what meeting participants say.

[2100] 7. Adjusting Emotion-Based Warnings

[2101] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[2102] 8. Emotion-based stress management

[2103] The emotion engine monitors the meeting facilitator's stress level in real time and provides alerts to reduce stress as needed.

[2104] 9. Records of final decisions and decision-makers

[2105] The server monitors keywords such as "conclusion" and "decision" in real time and records their content when detected.

[2106] The device displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's name.

[2107] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[2108] Specific example

[2109] Case Study 1: Timekeeping Alerts

[2110] 1. The user begins a 15-minute presentation.

[2111] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[2112] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[2113] 4. The device notifies the user that "3 minutes remaining."

[2114] Case Study 2: Detecting Derailments in Discussions

[2115] 1. The agenda is set to "Project A progress report".

[2116] 2. The user begins talking at length about "Project B".

[2117] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[2118] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[2119] 5. The device notifies the facilitator that "the topic has gone off track."

[2120] Case Study 3: Stress Management Using an Emotional Engine

[2121] 1. The facilitator begins to feel stressed while conducting the meeting.

[2122] 2. The emotion engine monitors the facilitator's stress level in real time.

[2123] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[2124] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[2125] The following describes the processing flow.

[2126] Step 1:

[2127] The user launches the meeting management application and clicks the "Start Meeting" button.

[2128] Step 2:

[2129] The server retrieves the meeting schedule, agenda, and attendee list from the database.

[2130] Step 3:

[2131] The server sends the attendee list and meeting information it has acquired to the terminal, and the terminal notifies the attendees.

[2132] Step 4:

[2133] The device records the meeting audio in real time and streams the audio data to the server.

[2134] Step 5:

[2135] The server uses a speech recognition engine to convert the received audio data into text data.

[2136] Step 6:

[2137] The server analyzes the converted text data, identifies the speaker, and annotates the text data.

[2138] Step 7:

[2139] The server records the start time of the presenter's speech and measures the elapsed time of the speech in real time.

[2140] Step 8:

[2141] After a set time has elapsed since the start of the announcement (for example, 12 minutes), the server generates an alert and sends it to the terminal.

[2142] Step 9:

[2143] The device will notify the presenter with an audio message and a visual alert saying, "3 minutes remaining."

[2144] Step 10:

[2145] The server compares the converted text data with the agenda and detects a digression if a topic that does not match is mentioned.

[2146] Step 11:

[2147] If a derailment is detected, the server generates a warning and sends it to the terminal.

[2148] Step 12:

[2149] The device provides the facilitator with an audio notification and a display alert saying, "The topic is going off track. Let's get back to the agenda."

[2150] Step 13:

[2151] An emotion engine installed on the server analyzes the voice tone and content of meeting participants in real time to recognize their emotions.

[2152] Step 14:

[2153] The server adjusts the warning content and notification tone for presenters and participants based on the emotional state obtained from the emotion engine.

[2154] Step 15:

[2155] If the server, based on the results of its emotion engine's emotional state monitoring, determines that the facilitator's stress level is high, it generates an alert saying, "Let's take a short break," and sends it to the terminal.

[2156] Step 16:

[2157] The device provides the facilitator with an audio notification and a display alert saying, "We recommend you take a short break."

[2158] Step 17:

[2159] The server monitors and detects keywords such as "conclusion" and "decision" in real time.

[2160] Step 18:

[2161] If a keyword is detected, the server records its contents and sends an alert to the terminal.

[2162] Step 19:

[2163] The device displays the message "Please enter the final decision-maker," and the user enters the decision-maker's name.

[2164] Step 20:

[2165] The terminal sends the decision-maker information entered by the user to the server, which records it along with the decision made.

[2166] (Example 2)

[2167] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2168] Many challenges exist in managing the progress of meetings. Existing systems simply converted audio to text and monitored the agenda, resulting in insufficient management of speaking time, prevention of discussion digressions, and recognition of participants' emotions. Furthermore, stress management for meeting facilitators was not considered. A system that comprehensively addresses these challenges is needed.

[2169] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[2170] In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if a set time is exceeded, means for comparing the content of the discussion with the agenda and detecting mismatched topics, means for issuing a warning based on the detected mismatched topics, means for recognizing the emotions of participants and adjusting the content of warnings and notification tones based on their emotional state, means for monitoring the stress level of the meeting facilitator in real time and providing alerts for stress reduction as needed, and means for recording final decisions and the decision-makers. This makes it possible to comprehensively and in real time perform everything from acquiring audio data to text conversion, monitoring progress, timekeeping, detecting deviations from the discussion, emotion recognition, stress management, and recording final decisions.

[2171] "Audio data" refers to digital data in which sound is recorded electronically.

[2172] "Text data" refers to digital data obtained by converting audio data into written text.

[2173] "Progress status" refers to information indicating the current state of the meeting's progress.

[2174] "Speaking time" refers to the time elapsed since the presenter began speaking.

[2175] An "agenda" is a list of topics or subjects that should be discussed at a meeting.

[2176] A "topic" refers to a specific subject or theme that is discussed in a meeting.

[2177] A "warning" is a notice issued when a set condition is not met.

[2178] "Emotions" refers to the psychological and emotional state of the meeting participants.

[2179] "Emotional state" refers to the specific expression of a meeting participant's emotions at a particular moment.

[2180] "Notification tone" refers to the sound or tone used when issuing warnings or notifications.

[2181] "Stress level" indicates the degree of mental and physical burden on the meeting organizer.

[2182] An "alert" is a warning or notice that appears when certain conditions occur.

[2183] "Decisions made" refer to the specific details and action plans that were ultimately decided upon at the meeting.

[2184] A "decision-maker" is the person or position that ultimately makes the final decision on a particular matter in a meeting.

[2185] Modes for carrying out the invention

[2186] This invention is a system that acquires audio data during a meeting and analyzes it in real time. Furthermore, by combining it with an emotion engine, it recognizes the emotions of participants and facilitates smoother meeting management. This system includes the following components:

[2187] Acquisition of audio data

[2188] When a user starts a meeting, the device uses its microphone to record the meeting's audio data in real time and streams that audio data to the server. The audio data is processed in digital format and transferred to the server via the communication network.

[2189] Speech-to-text conversion

[2190] The server uses a speech recognition engine such as Google Cloud Speech-to-Text to convert the received audio data into text data. The converted text data is temporarily stored on the server and used for subsequent data analysis.

[2191] Monitoring the progress of the meeting

[2192] The server uses a natural language processing engine, such as the Google Cloud Natural Language API, to analyze the text data converted from speech and identify the speaker. It adds annotations to the text data, including the speaker's ID and the content of their speech, and monitors the progress of the meeting.

[2193] Timekeeping

[2194] The server records the start time of each user's speech and monitors the speech duration in real time. Once the set time has elapsed, the server generates an alert and sends it to the terminal. The terminal then notifies the presenter of the remaining time via audio notification and a display alert.

[2195] Derailment detection in discussions

[2196] The server compares the converted text data with pre-configured agenda items and recognizes any discussion of a mismatched topic as a digression. When a digression is detected, the server generates a warning and sends it to the terminal. The terminal then notifies the facilitator with an audio and visual alert stating, "The topic has gone off-topic."

[2197] Emotion recognition by an emotion engine

[2198] The terminals and servers are equipped with emotion engines such as the Azure Cognitive Services Emotion API to analyze the voice tone and content of meeting participants. This makes it possible to recognize participants' emotions (e.g., joy, anger, surprise) in real time.

[2199] Adjusting emotion-based warnings

[2200] The server receives the emotional state recognized by the emotion engine and adjusts the warning content and notification tone accordingly. For example, if a participant is feeling stressed, the alert tone will be softened and the notification content will be made gentler.

[2201] Emotion-based stress management

[2202] The emotion engine monitors the meeting facilitator's stress level in real time and provides this information to the server. If the server determines that the facilitator's stress level is high, it generates alerts for breaks or urgent discussions. For example, it might send a message to the facilitator such as, "We recommend you take a short break."

[2203] Records of final decisions and decision-makers

[2204] The server monitors keywords such as "conclusion" and "decision" in the conversation in real time. When a keyword is detected, its content is immediately recorded. The terminal displays an alert prompting the user to confirm the decision-maker, and the user enters the decision-maker's information. The entered decision-maker information is sent from the terminal to the server and recorded along with the final decision.

[2205] Specific example

[2206] Case Study 1: Timekeeping Alerts

[2207] 1. The user begins a 15-minute presentation.

[2208] 2. The server records the start time of the presentation and measures the elapsed time in real time.

[2209] 3. After 12 minutes, the server generates a "3 minutes remaining" warning and sends it to the terminal.

[2210] 4. The device notifies the user that "3 minutes remaining."

[2211] Prompt example

[2212] "Design a system that records the start time of a 15-minute presentation and generates a '3 minutes remaining' alert 12 minutes later."

[2213] Case Study 2: Detecting Derailments in Discussions

[2214] 1. The agenda is set to "Project progress report".

[2215] 2. The user starts talking at length about "another project."

[2216] 3. The server analyzes the discussion content as text and detects if it does not match the agenda.

[2217] 4. The server generates a "discussion off-topic" warning and sends it to the terminal.

[2218] 5. The device notifies the facilitator that "the topic has gone off track."

[2219] Prompt example

[2220] "Design a system that detects off-topic comments during meetings and alerts the facilitator to the derailment."

[2221] Case Study 3: Stress Management Using an Emotional Engine

[2222] 1. The facilitator begins to feel stressed while conducting the meeting.

[2223] 2. The emotion engine monitors the facilitator's stress level in real time.

[2224] 3. The server generates a stress management alert and notifies you, "We recommend you take a short break."

[2225] Prompt example

[2226] "Design a system that monitors the facilitator's stress level during meetings and sends notifications prompting them to take breaks when necessary."

[2227] This system enables sophisticated meeting management that considers not only time management and agenda progression, but also the emotional state of participants, allowing younger facilitators in particular to manage meetings efficiently.

[2228] The flow of the specific processing in Example 2 will be explained using Figure 13.

[2229] Step 1:

[2230] The user starts the meeting.

[2231] Specifically, the user launches the meeting application and starts the meeting by clicking the "Start Meeting" button.

[2232] Input: User's meeting start operation

[2233] Output: Meeting start signal

[2234] Step 2:

[2235] The device prepares to record audio.

[2236] Specifically, the device checks the readiness status of the built-in or externally connected microphone and enables the recording function.

[2237] Input: Signal to start meeting

[2238] Output: Signal indicating ready to record

[2239] Step 3:

[2240] The device records audio data in real time and streams it to the server.

[2241] Specifically, the device records the audio during the meeting in real time and streams it to a server via the internet as digital audio data.

[2242] Input: Signal indicating recording is ready.

[2243] Output: Streaming of real-time audio data

[2244] Step 4:

[2245] The server receives the audio data and uses a speech recognition engine to convert the audio data into text data.

[2246] In terms of specific operations, the server uses a speech recognition engine to analyze the audio data and convert it into text data written as characters. A typical example of the software used is a general-purpose speech recognition engine.

[2247] Input: Real-time audio data

[2248] Output: Text data

[2249] Step 5:

[2250] The server temporarily stores the converted text data.

[2251] Specifically, the server temporarily stores the converted text data in storage such as a database.

[2252] Input: Text data

[2253] Output: Temporarily saved text data

[2254] Step 6:

[2255] The server analyzes the text data converted from the speech and monitors the progress using a natural language processing engine.

[2256] In terms of specific operations, the server uses a natural language processing engine to analyze text data, identify the speaker, and annotate the content of the speech. The software used includes a natural language processing engine, among other things.

[2257] Input: Temporarily saved text data

[2258] Output: Annotated text data

[2259] Step 7:

[2260] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[2261] Specifically, the server records the start time of a message and then continuously measures the duration of subsequent messages.

[2262] Input: Annotated text data

[2263] Output: Speech time elapsed data

[2264] Step 8:

[2265] After a set amount of time has elapsed since the start of the announcement, the server generates an alert and sends it to the terminal.

[2266] Specifically, when the server determines that a set time has elapsed, it generates an alert message and sends it to the terminal.

[2267] Input: Elapsed time of speech

[2268] Output: Alert

[2269] Step 9:

[2270] The device will notify the presenter of the remaining time via audio notification and a display alert.

[2271] Specifically, the device will use voice notifications and screen displays based on the received alerts to inform the presenter of the remaining time.

[2272] Input: Alert

[2273] Output: Voice notifications and display alerts

[2274] Step 10:

[2275] The server compares the converted text data with the agenda and detects any mismatched topics.

[2276] Specifically, the server compares the text data with a pre-configured agenda list and automatically detects any mismatches.

[2277] Input: Annotated text data

[2278] Output: Detection results for mismatched topics

[2279] Step 11:

[2280] The server issues a warning based on mismatched topics and sends it to the terminal.

[2281] Specifically, the server generates a warning message based on the detected mismatched topics and sends it to the terminal.

[2282] Input: Detection results for mismatched topics

[2283] Output: Warning message

[2284] Step 12:

[2285] The device provides the facilitator with an audio notification and a visual alert when the conversation goes off-topic.

[2286] Specifically, the device uses the received warning message to provide audio notifications and screen displays, informing the facilitator that the topic has gone off track.

[2287] Input: Warning message

[2288] Output: Voice notifications and display alerts

[2289] Step 13:

[2290] The terminals and servers are equipped with an emotion engine to recognize the emotions of the participants.

[2291] Specifically, the terminal and server use an emotion engine to analyze the participant's emotions in real time based on their voice tone and what they say.

[2292] Input: Real-time audio data

[2293] Output: Emotion recognition result

[2294] Step 14:

[2295] The server adjusts warning content and notification tone based on the emotional state recognized by the emotion engine.

[2296] Specifically, the server adjusts the content and tone of the warning message based on the emotion recognition results.

[2297] Input: Sentiment recognition result

[2298] Output: Adjusted warning message

[2299] Step 15:

[2300] The emotion engine monitors the stress levels of meeting organizers, and the server provides alerts to reduce stress as needed.

[2301] Specifically, the emotion engine constantly analyzes the operator's stress level and generates alerts for breaks and relaxation at appropriate times.

[2302] Input: Real-time audio data and operator sentiment data

[2303] Output: Stress reduction alert

[2304] Step 16:

[2305] The server monitors for "conclusion" and "decision" keywords and records their content.

[2306] Specifically, the server monitors for specific keywords within the converted text data in real time and records their content in the database if detected.

[2307] Input: Annotated text data

[2308] Output: Recorded decisions

[2309] Step 17:

[2310] The terminal prompts the user to confirm the decision-maker and sends the decision-maker information entered by the user to the server.

[2311] Specifically, after a decision is finalized, the terminal displays a pop-up notification prompting the user to enter the decision-maker's information. The entered information is sent to the server and recorded.

[2312] Input: Decision made and user-inputted decision-maker information

[2313] Output: Recorded decision-maker information

[2314] (Application Example 2)

[2315] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2316] Meetings during shift changes in factories require smooth information transfer and efficient discussion. However, insufficient communication on the factory floor and frequent digressions in discussions often lead to problems with efficient shift changes and work handovers. Furthermore, the emotional state of participants significantly impacts the progress of the meeting, making emotional management a crucial issue. A system is needed to appropriately address these problems.

[2317] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for converting the acquired audio data into text data, means for analyzing the converted text data to monitor the progress of the meeting, means for monitoring the speaking time of presenters and issuing a warning if the set time is exceeded, means for comparing the content of the discussion with the agenda and detecting inconsistent topics, means for issuing a warning based on the detected inconsistent topics, and means for recognizing the emotions of meeting participants and adjusting the content of the warning and the tone of the notification based on their emotional state. This enables appropriate information transmission, efficient progress of discussions, and emotional management of participants in factory shift change meetings.

[2318] "Means for acquiring audio data" refers to a device or software for recording speech spoken in real time during meetings or shift changes and saving it as digital data.

[2319] "Means for converting acquired audio data into text data" refers to a speech recognition engine or software that analyzes recorded audio data and automatically converts it into text data.

[2320] "Means of analyzing converted text data to monitor the progress of a meeting" refers to a program or system that analyzes text data to understand who is speaking and the progress of the agenda in real time.

[2321] "A means of monitoring the presenter's speaking time and issuing a warning if the set time is exceeded" refers to a system with a timekeeping function that records the start time of the presentation and generates an alert when the set time limit is exceeded.

[2322] "Means for comparing the content of a discussion with the agenda and detecting inconsistent topics" refers to a program or algorithm that compares the content of a statement with a set agenda and detects when a topic that does not match the agenda is mentioned.

[2323] "Means for issuing warnings based on detected mismatched topics" refers to a system that automatically generates warnings related to detected mismatched topics and notifies speakers and moderators.

[2324] "Means for recording final decisions and decision-makers" refers to a system that has the function of recording decisions made during meetings or shift changes, as well as the decision-makers involved, and saving them in a database for later reference.

[2325] "Means for recognizing the emotions of meeting participants and adjusting the tone of warnings and notifications based on their emotional state" refers to an algorithm or engine that analyzes emotions in real time from the tone of voice and content of statements of participants, and changes the content and expression of warnings and notifications according to their emotional state.

[2326] The system that implements this application supports meetings during shift changes in a factory. This system acquires audio data in real time, converts it into text data, and performs meeting progress monitoring, timekeeping, discussion derailment detection, and sentiment recognition.

[2327] Specifically, the program will be implemented with the following structure:

[2328] Hardware and software

[2329] Hardware to use:

[2330] Microphone: For capturing the voices of meeting participants.

[2331] Robot or smartphone: A platform for processing voice data and sending notifications.

[2332] Software to use:

[2333] speech_recognition: A library for performing speech recognition.

[2334] emotion_recognition: A library for performing emotion recognition (uses a pre-trained model).

[2335] datetime: A standard library for time management.

[2336] Process Overview

[2337] 1. Acquisition of audio data:

[2338] The audio of meeting participants can be captured in real time via microphones.

[2339] The voice data is transferred to a robot or smartphone platform.

[2340] 2. Speech-to-text conversion:

[2341] The speech_recognition library installed on the robot or smartphone is used to convert speech data into text data.

[2342] The converted text is sent to the server.

[2343] 3. Monitoring the progress of the meeting:

[2344] The server analyzes the converted text data to identify the speaker and monitor the progress.

[2345] Verify that the meeting is proceeding according to the set agenda.

[2346] 4. Timekeeping:

[2347] The server records the start time of the presenter's speech and measures the elapsed time in real time.

[2348] If the presentation time exceeds the set time, a warning is sent from the server to the robot or smartphone, and the presenter is notified via audio notification and a display alert.

[2349] 5. Detecting deviations from the discussion:

[2350] The server compares the discussion to the agenda and detects deviations if a topic that does not match is mentioned.

[2351] If a derailment is detected, a warning will be displayed to the facilitator.

[2352] 6. Emotion recognition by the emotion engine:

[2353] The robot or smartphone is equipped with an emotion engine that recognizes the participant's emotions in real time based on their voice tone and what they say.

[2354] This allows for monitoring the emotional state of participants and taking appropriate action.

[2355] 7. Adjusting emotion-based warnings:

[2356] The server adjusts the tone of warnings and notifications based on the emotional state recognized by the emotion engine.

[2357] For example, if emotions are heightened, the notification will be delivered in a calm tone.

[2358] Specific example

[2359] Examples of support for shift change meetings

[2360] 1. Start of the meeting:

[2361] When the shift changes, the leader begins explaining the situation to the next team.

[2362] The microphone acquires audio data in real time and transmits it to the robot or smartphone.

[2363] 2. Speech-to-text conversion:

[2364] The acquired audio data is converted to text using the speech_recognition library and sent to the server.

[2365] 3. Monitoring the progress of the meeting and keeping time:

[2366] The server analyzes the converted text data to identify the speaker, monitor the progress, and keep time.

[2367] After a set amount of time has elapsed since the announcement began, a warning will be sent to the robot or smartphone, triggering an audio notification and a visual alert.

[2368] 4. Detecting deviations from the discussion:

[2369] The server compares the discussion content with the agenda and warns the facilitator if it detects a deviation from the topic.

[2370] 5. Recognition by the emotion engine:

[2371] It monitors future emotional states in real time and adjusts warning content and notification tones as needed.

[2372] Examples of prompt statements to use

[2373] "Record the 5-minute shift change meeting and convert it to text. Perform sentiment analysis using an emotion recognition model, and implement tactical detection and timekeeping."

[2374] This system allows factory shift change meetings to proceed smoothly and efficiently, and because it takes into account the emotional state of the participants, it enables more effective information sharing and discussion.

[2375] The flow of a specific process in Application Example 2 will be...

Claims

1. Means for acquiring audio data, A means of converting acquired audio data into text data, A means of analyzing the converted text data to monitor the progress of the meeting, A means to monitor the presenter's speaking time and issue a warning if the set time limit is exceeded, A means of comparing the content of the discussion with the agenda and detecting topics that do not match, A means of issuing warnings based on detected mismatched topics, A means of recording the final decision and the decision-maker, A system that includes this.

2. The system according to claim 1, wherein the means for issuing a warning is means for providing an audio notification and a display alert.

3. The system according to claim 1, comprising means for monitoring the progress of a meeting in real time and managing its progress based on a set agenda.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A