system

A system that uses natural language input to generate meeting plans, monitor speech, and provide real-time feedback and follow-up to enhance meeting efficiency and productivity by keeping discussions on track and documenting outcomes.

JP2026069144APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Meetings often proceed inefficiently due to topic derailment and insufficient follow-up, leading to wasted time and reduced productivity, as participants struggle to grasp decisions and next actions.

Method used

A system that accepts meeting purpose and time in natural language, generates a meeting plan, monitors speech in real-time, and automatically generates minutes and next actions, ensuring participants stay on topic and follow up effectively.

Benefits of technology

Enhances meeting efficiency by maintaining focus, generating clear minutes, and ensuring timely follow-up on actions, thereby improving productivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069144000001_ABST
    Figure 2026069144000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of accepting input of the purpose and scheduled time of a meeting in natural language, A means of automatically generating a meeting schedule, A method for monitoring speech in real time during a meeting using speech recognition technology, A means to detect when a topic deviates from the subject of the meeting and prompt the participant to return to the subject at an appropriate time, A means to automatically generate meeting minutes and identify the next action for each participant, A system that includes a means of communication to notify each participant of their next action after a meeting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In many modern business environments, meetings proceed inefficiently, resulting in problems such as wasted time and reduced productivity. This can be attributed to factors such as the inability of meetings to focus on the main issues due to topic derailment and insufficient follow-up after meetings. As a result, there is a problem that participants cannot grasp the matters decided in the meetings and the next actions, hindering the achievement of results.

Means for Solving the Problems

[0005] This invention provides a system that accepts the purpose and scheduled time of a meeting as input in natural language, automatically generates a meeting plan based on that information, and improves the efficiency of meetings. It also utilizes speech recognition technology to monitor speech during meetings in real time and has a function to appropriately guide participants back to the main topic when they stray from it. Furthermore, it automatically generates meeting minutes and next actions and notifies participants, thereby strengthening post-meeting follow-up and ensuring that meeting outcomes are reliably implemented.

[0006] "Natural language" refers to the language that humans use on a daily basis, and it allows for intuitive information exchange in interfaces with machines.

[0007] The "purpose of the meeting" refers to the specific goals or themes for conducting the meeting, and serves as a guideline for participants to work towards achieving them.

[0008] "Scheduled time" refers to the time range in which a meeting will take place, and the meeting is planned and conducted based on that time.

[0009] A "meeting plan" is a plan that organizes the agenda and procedures of a meeting in advance, and is designed to ensure that the meeting proceeds efficiently according to the allotted time.

[0010] "Speech recognition technology" is a technology that collects spoken words as audio data and converts it into text or specific commands.

[0011] "Real-time monitoring" refers to the immediate and uninterrupted monitoring and tracking of ongoing processes, enabling immediate responses.

[0012] "Deviating from the topic" refers to a situation where the discussion strays from the original purpose or agenda of the meeting, causing the focus of the meeting to become blurred.

[0013] "Meeting minutes" are documents that record what was said, what was decided, and an overview of the discussions during a meeting, allowing the contents of that meeting to be reviewed later.

[0014] "Next action" refers to the specific actions or tasks that each participant should take next based on the decisions made at the meeting.

[0015] "Communication methods" refer to the technologies and methods used to send and receive information, and include forms such as email. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention is a system for supporting the efficient progress of meetings, and is configured as follows: This system consists of a user, a terminal, and a server, which work together to assist in the progress of the meeting.

[0038] When a user starts a meeting, they input the meeting's purpose and scheduled time into their device using natural language. This collects the basic information necessary for the meeting. The device then converts the audio data into text and sends it to the server.

[0039] The server analyzes the received text data and automatically generates a meeting schedule based on the meeting's purpose and scheduled time. The schedule includes the main agenda items and timeline. This schedule is then used to prepare for and support the meeting's progress.

[0040] During the meeting, the terminal and server utilize speech recognition technology to monitor the meeting content in real time. As the meeting progresses, the content is recorded by transcribing speech into text, and if the topic deviates, the server sends a notification at an appropriate time to prompt the user to return to the meeting's subject. This prompt is displayed on the terminal and, if necessary, communicated to the user verbally.

[0041] Once the meeting concludes, the server automatically generates meeting minutes based on the recorded statements. These minutes include key points, conclusions, and decisions made during the meeting. They also identify and summarize the next steps for each participant. The server organizes this information and sends notifications regarding these next steps to each participant via their terminal. These notifications arrive as emails, ensuring continued follow-up even after the meeting ends.

[0042] For example, a user might define the purpose as "meeting about the annual budget" and set the scheduled time to "1 hour." The server would then set three agenda items: revenue and expenditure forecast, cost reduction proposals, and revenue improvement measures, allocating 20 minutes to each. If participants stray from the agenda during the meeting, the server would send a reminder to bring them back to the main topic, ensuring the meeting runs smoothly.

[0043] In this way, this system improves meeting productivity and helps participants efficiently share information and move on to the next step.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device then converts this audio data into text data.

[0047] Step 2:

[0048] The server receives text data sent from the terminal and uses natural language processing to analyze the purpose and scheduled time of the meeting. It then generates a meeting schedule.

[0049] Step 3:

[0050] The server creates a progress plan, sets the meeting schedule and agenda based on it, and sends it to the terminal.

[0051] Step 4:

[0052] During the meeting, the terminal utilizes speech recognition technology to monitor the user's speech in real time, convert the audio data into text data, and send it to the server.

[0053] Step 5:

[0054] The server analyzes the meeting discussions in real time, and if it detects that the topic has deviated from the set agenda, it sends a notification to the device prompting the participant to return to the topic at an appropriate time.

[0055] Step 6:

[0056] Once the meeting ends, the server aggregates the recorded text data of the speeches and automatically generates meeting minutes. Furthermore, it identifies and organizes the next actions for each participant.

[0057] Step 7:

[0058] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, thereby facilitating follow-up after the meeting.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] In meetings, participants often deviate from the agenda, and meeting minutes are frequently left ambiguous, leading to decreased meeting efficiency and delays in important decisions. Furthermore, there is a problem with insufficient follow-up, as specific next steps are not promptly communicated to participants after the meeting.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes a device that accepts input of agenda items and scheduled times in natural language, a device that converts voice data into text, and a device that analyzes the converted data to automatically generate a meeting plan based on the agenda. This prevents deviations from the agenda during the meeting, enables efficient progress and the generation of clear meeting minutes, and also strengthens post-meeting follow-up by quickly notifying participants of their next steps.

[0064] A "device that accepts input of agenda items and scheduled times in natural language" is a device that allows users to input meeting topics and estimated times verbally or in text, and then converts and processes that information into digital signals.

[0065] A "device that converts audio data to text" is a device that analyzes voice input from a user and accurately transcribes its content into text. This device processes audio signals and converts them into text data.

[0066] A "device that analyzes converted data and automatically generates a progress plan based on the agenda" is a device that analyzes the content of a transcribed meeting and automatically creates the meeting procedure and time allocation based on pre-set objectives.

[0067] A "device that uses speech recognition technology to monitor conversations in real time" is a device that instantly transcribes conversations during a meeting into text, analyzes the content, and monitors deviations from the planned agenda.

[0068] A "device that detects deviations from the agenda and prompts participants to return to the agenda at the appropriate time" is a device that detects statements that deviate from the original agenda from conversations monitored in real time and instructs participants to return to the original agenda.

[0069] A "device that automatically generates meeting minutes and identifies each participant's next action" is a device that records meeting content, organizes important decisions and next steps, and automatically documents them.

[0070] A "communication device for notifying each participant of the minutes and subsequent actions" is a device for transmitting the generated minutes and future instructions to each meeting participant using communication methods such as email.

[0071] This invention is a system that efficiently supports the progress of meetings, in which users, terminals, and servers work together. Specific embodiments are described below.

[0072] When a user starts a meeting, they input the meeting's purpose and scheduled time into the device using natural language. The device receives this as voice input and converts the voice data into text using the Google® Speech-to-Text API. This process ensures that the voice signal is accurately converted into text data.

[0073] Next, the terminal sends the converted text data to the server. The server receives this data and performs analysis using IBM Watson® Natural Language Understanding. This analysis identifies the meeting's objectives and key topics, and the server automatically generates a meeting plan based on this information. The plan includes the agenda and necessary time schedules.

[0074] As the meeting progresses, the terminal and server monitor the conversation in real time using speech recognition technology. If a participant deviates from the agenda during the meeting, the server detects this and sends a reminder to the user via the terminal at an appropriate time to return to the topic. Specifically, this may involve displaying a message on the terminal such as "The current topic has deviated from the scheduled agenda," or an audio notification may also be considered.

[0075] At the end of the meeting, the server automatically generates meeting minutes based on the recorded speech data. These minutes include the meeting's objectives, key decisions, and next steps. The server uses the Microsoft® Word API to format the minutes as a document and provides it to each participant.

[0076] The server then identifies each participant's next action and notifies them via email through their device. This notification, delivered using the Gmail API, includes specific instructions and follow-ups. An example of a prompt sent to a participant might be, "Please consider the next steps regarding the annual budget and prepare your proposal for the next meeting."

[0077] This system provides support to help users efficiently conduct meetings and share information. By optimizing prompt sentences using a generative AI model, even faster and more effective communication becomes possible.

[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0079] Step 1:

[0080] User input of meeting information

[0081] To start a meeting, the user enters the purpose and scheduled time of the meeting into their device using natural language. The entered information is received as voice or text data. This input information becomes the basic data for the system. The user provides data such as "About next year's budget proposal, approximately one hour" using the voice input function of their smartphone.

[0082] Step 2:

[0083] Converting audio data to text

[0084] The device converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to analyze the audio signal and generate data as a string. This step converts the audio information into readable text. The output is the text "Regarding next year's budget proposal, about 1 hour."

[0085] Step 3:

[0086] Text data transmission and analysis

[0087] The terminal sends the converted text data to the server. The server analyzes the received data using IBM Watson Natural Language Understanding. This analysis extracts the meeting's objectives and important keywords, and organizes the information necessary for the meeting's progress. Text data is used as input, and the output is the analyzed information structure.

[0088] Step 4:

[0089] Automatic generation of meeting schedules

[0090] The server automatically generates a meeting schedule based on the analysis results. This takes into account the meeting's topic and time allocation, and the schedule includes information such as each agenda item and the time allocated to each item. The output includes specific procedures such as "budget review," "resource allocation," and "cost reduction measures."

[0091] Step 5:

[0092] Real-time monitoring and reminder sending

[0093] The terminal and server monitor what is said during the meeting in real time. Using speech recognition, the conversation is transcribed into text, and a reminder is sent if it deviates from the agenda. The server compares it to the scheduled items and detects when the topic has gone off track. At the appropriate time, the terminal displays a message saying, "The topic has gone off-topic. Let's return to the main subject."

[0094] Step 6:

[0095] Meeting minutes generation after the meeting

[0096] The server automatically generates meeting minutes based on the recorded conversations during the meeting. This process uses the Microsoft Word API to document the text data and organize important information for participants. The output is a meeting minutes document containing information such as "budget approval," "next meeting date and time," and "action items."

[0097] Step 7:

[0098] Notification of the next action

[0099] The server notifies each participant of the generated meeting minutes and the identified next action. The notification is sent via email through the terminal using the Gmail API. This allows participants to understand the key points of the meeting and receive specific instructions for the next steps. As output, a follow-up email is sent to each participant.

[0100] (Application Example 1)

[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0102] In the face of the need for effective meeting management and operational efficiency in physical stores, traditional methods often fail to keep participants focused on the agenda, resulting in insufficient achievement of meeting objectives. Furthermore, a lack of follow-up and action checks after meetings leads to decreased productivity. To overcome these challenges, there is a need to develop a system that manages meetings efficiently and systematically, while providing appropriate support to participants.

[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0104] In this invention, the server includes an information receiving means for inputting data on the purpose and scheduled time of the meeting in natural language, a planning means for automatically generating a meeting schedule, and an information providing means for providing hints and suggestions based on data in real time during the meeting. This enables participants to concentrate on the meeting's topic, effectively advance discussions, and ensure follow-up after the meeting.

[0105] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format that computer systems can understand and process.

[0106] "Information receiving means" refers to a device or method for receiving the purpose and scheduled time of a meeting, entered by the user, in digital format.

[0107] "Planning means" refers to algorithms or devices for automatically generating a meeting schedule.

[0108] "Speech recognition technology" is a technology that converts speech data into text data in real time and understands the content of what is being said.

[0109] "Monitoring methods" refer to technologies that monitor speeches during a meeting in real time and record and analyze relevant data.

[0110] "Correction measures" are technologies or devices that detect when a meeting deviates from its topic and prompt participants to return to the topic at an appropriate time.

[0111] "Action identification means" refers to a method or device for automatically identifying the next steps for each individual who participated in the meeting.

[0112] "Communication means" refers to electronic methods or devices used to notify participants of an assembly of information.

[0113] "Information provision means" refers to technologies and devices that provide hints and suggestions based on real-time data during a meeting.

[0114] The system for realizing this invention is designed to enable employees to conduct meetings more efficiently. In the in-store meeting support system, the main hardware used is a smartphone or tablet to collect audio data. The information receiving means installed in these devices can receive natural language data entered by the user. Specifically, the software components include speech recognition technology that converts audio data into text data in real time, using Python and its library, SpeechRecognition.

[0115] After receiving text data, the server performs text analysis using the natural language processing library spaCy. Next, a meeting agenda is automatically generated. During the meeting, the content of the discussion is monitored using Google Cloud's Natural Language API. In addition, hints and suggestions are provided in real time through various information channels to help participants focus on the meeting's main topic.

[0116] As an example, consider the use of this system by a store operations team when holding a new product promotion meeting. In this meeting, participants choose "New Product Promotion Plan" as the topic and set a one-hour time limit. Based on this data, the server automatically sets the agenda and time allocation, and notifies participants through corrective means if they deviate from the discussion. Furthermore, after the meeting, each participant receives a notification regarding the next steps.

[0117] Examples of prompt statements to input into a generative AI model are as follows:

[0118] "The purpose of today's meeting is to discuss the 'marketing plan for the new product,' and it will last one hour. Please provide a meeting plan based on this. Also, monitor the progress of each agenda item and send a notification if the discussion goes off track."

[0119] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0120] Step 1:

[0121] The user enters the purpose and scheduled time of the meeting into the device using natural language. The entered information is collected by the device's information receiving mechanism. The input data is converted into text using speech recognition technology and sent to the next step.

[0122] Step 2:

[0123] The terminal converts audio data into text data in real time using the SpeechRecognition library. The converted data is sent to the server via the internet. The output here is text data that includes the purpose and scheduled time of the meeting.

[0124] Step 3:

[0125] The server receives text data and performs natural language processing using the spaCy library. This analyzes the purpose of the meeting and the priority agenda items, and automatically generates a meeting plan. The generated plan includes the time allocation for each agenda item. This plan is used in the next step.

[0126] Step 4:

[0127] During the meeting, the device collects audio data again, transcribes it into text in real time, and sends it to the server. The server uses Google Cloud's Natural Language API to monitor the content of the discussion as the meeting progresses. If the topic deviates, the server automatically detects the deviation and generates a notification prompting the user to correct it.

[0128] Step 5:

[0129] At the end of the meeting, the server automatically generates meeting minutes using the collected data. These minutes will include key points from the meeting and the next steps for participants. The generated minutes will be sent to each participant via their terminal.

[0130] Step 6:

[0131] After the meeting, the server sends participants notifications about the next steps via email or other means of communication. This ensures that each participant is clearly aware of their next actions and can follow up effectively.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention provides a system that enables efficient and emotionally insightful meeting management, and has the following configuration: This system includes a user, a terminal, a server, and an emotion engine, which work together to support the progress of the meeting.

[0134] First, at the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device converts the entered information from voice data to text data and sends it to the server.

[0135] Next, the server analyzes the text data and automatically generates a meeting plan based on the meeting's purpose and duration. This plan includes a time schedule and agenda, enabling efficient meeting management. The meeting plan is then fed back to the user via their terminal.

[0136] During the meeting, the terminal and server use speech recognition technology to monitor the content of the discussion in real time. If the discussion deviates from the set agenda, the server sends a notification to the terminal prompting the user to return to the topic at an appropriate time. Furthermore, once the meeting ends, the server automatically generates meeting minutes based on the content of the discussion, and identifies and organizes the next actions for each participant. This ensures that follow-up is provided even after the meeting has ended.

[0137] The emotion engine, a key feature of this invention, analyzes users' emotions in real time during a meeting. Based on the emotion data obtained from the emotion engine, the server flexibly adjusts the progress of the meeting. For example, if a participant is feeling dissatisfied, the server can adjust the timing of changing topics. Furthermore, the meeting minutes and reports generated after the meeting also include the emotion data, providing emotional evaluations and insights into how the meeting proceeded.

[0138] As a concrete example, a user sets the objective as "Meeting about the launch of a new product" and enters a scheduled time of "2 hours." During the meeting, if the emotion engine detects signs of agitation in some participants, the server takes this into consideration and adjusts the flow of the meeting to help ensure a constructive discussion. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0139] The following describes the processing flow.

[0140] Step 1:

[0141] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device receives the input as voice data and converts it into text data.

[0142] Step 2:

[0143] The terminal sends the converted text data to the server. The server analyzes the purpose and duration of the meeting and automatically generates a meeting schedule.

[0144] Step 3:

[0145] The server generates a progress plan and sends it to the terminal, which then displays it to the user. The progress plan includes the agenda and time schedule.

[0146] Step 4:

[0147] During the meeting, the terminal uses speech recognition technology to transcribe the user's speech into text in real time and send it to the server.

[0148] Step 5:

[0149] The server analyzes the received message data and detects when the topic deviates from the agenda. It then sends a notification to the device prompting the user to return to the topic at an appropriate time.

[0150] Step 6:

[0151] The server uses an emotion engine to analyze the emotions of participants during a meeting in real time and detects emotional data.

[0152] Step 7:

[0153] The server flexibly adjusts the flow of the meeting based on sentiment data and suggests new discussion transitions as needed.

[0154] Step 8:

[0155] Once the meeting ends, the server automatically generates meeting minutes based on the collected speech and sentiment data, and also identifies the next actions for each participant.

[0156] Step 9:

[0157] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, and then follows up after the meeting.

[0158] Step 10:

[0159] The server helps participants gain emotional insights by including emotional data in the meeting report and providing it to them.

[0160] (Example 2)

[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0162] Traditional meeting systems have had problems such as inefficient progress management and meetings sometimes going in an inappropriate direction because participants' feelings are not taken into consideration. Specifically, problems include wasted time due to off-topic discussions, difficulty in tracking meeting content, and the accumulation of participant dissatisfaction.

[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0164] In this invention, the server includes information processing means for analyzing natural language input data received via voice and determining the purpose and scheduled time of the meeting; schedule generation means for automatically generating a meeting plan based on the input data and notifying participants; and speech recognition technology means for monitoring speech in real time during the meeting and supporting the progress of the meeting in accordance with the agenda. This enables efficient management of the meeting and flexible meeting adjustments based on the emotions of the participants.

[0165] "Information processing means" refers to a device or program that has the function of analyzing natural language input data received via voice and determining the purpose and scheduled time of a meeting.

[0166] A "schedule generation means" is a device or program that has the function of automatically generating a meeting schedule based on input data and notifying participants of that schedule.

[0167] "Speech recognition technology means" refers to a device or program that uses technology to monitor speech during a meeting in real time and support the progress of the meeting in accordance with the agenda.

[0168] "Detection and notification means" refers to a device or program that has the function of detecting when a statement deviates from the agenda and prompting a return to the topic at an appropriate time.

[0169] "Document generation means" refers to a device or program that has the function of automatically generating meeting minutes based on audio data and clearly indicating the next action plan for each participant.

[0170] A "sentiment analysis module" is a device or program that analyzes the emotional data of meeting participants and has the function of adjusting the progress of the meeting.

[0171] "Communication technology means" refers to a device or program that has a communication function for informing each participant of the next action to take after a meeting.

[0172] "Data conversion means" refers to a device or program that has the function of converting audio data into text data in real time.

[0173] Modes for carrying out the invention

[0174] This invention is a system for facilitating smooth meeting progress and adjusting based on participants' emotions. The system is implemented through the collaboration of a server, terminals, and users, and utilizes the following hardware and software.

[0175] The device is equipped with speech recognition software that accepts user input using natural language. Specific examples of such software include Google Cloud Speech-to-Text and IBM Watson Speech to Text. This converts speech data into text data.

[0176] The server processes the input text data and automatically generates a meeting agenda using a generative AI model. This generated agenda includes the meeting's time schedule and specific agenda items.

[0177] Furthermore, the server utilizes an emotion engine to analyze participants' emotions in real time based on data from devices such as NeuroSky and Emotiv, thereby adjusting the meeting's progress. If participants are dissatisfied, the server can take action, such as changing the agenda.

[0178] During the meeting, the terminal and server work together to provide a system that monitors speech in real time. If a participant's remarks stray from the agenda, the server will notify them to return to the topic within the allotted time, thus appropriately supporting the progress of the meeting.

[0179] Finally, the server automatically generates meeting minutes based on the audio data of the meeting, including content and sentiment data, and creates a report that clearly communicates the next steps for each participant.

[0180] As a concrete example, a user sets the purpose as "Meeting about the launch of a new product" and enters the scheduled time as "2 hours." During the meeting, if the emotion engine detects that a participant is agitated, the server constructively adjusts the flow of the meeting. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0181] Example of a prompt

[0182] "Please use a generative AI model to create meeting minutes."

[0183] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0184] Step 1:

[0185] At the start of the meeting, the user enters the purpose and scheduled time into the terminal. Specifically, they provide information to the terminal by voice, such as "Meeting regarding the launch of a new product, 2 hours." The entered voice data is converted into text data by the terminal's voice recognition software. This step involves data processing, where voice data is converted into text.

[0186] Step 2:

[0187] The terminal sends the meeting's purpose and scheduled time as data to the server. Secure protocols such as HTTPS are used for transmission. This ensures that user input is safely transmitted to the server. The output data is sent to the server as a text file within the software.

[0188] Step 3:

[0189] The server processes the received text data. Using a generative AI model, it automatically generates a meeting schedule. This involves analyzing the text data and proposing appropriate time schedules and agenda items. This plan is then ready for feedback.

[0190] Step 4:

[0191] The server notifies the user of the generated schedule via the terminal. The terminal presents this information to the user through push notifications or display. In this step, the data output is provided visually through the user interface, preparing the meeting.

[0192] Step 5:

[0193] During the meeting, the terminal and server work together to use speech recognition technology to monitor speech in real time. They check whether the speech is relevant to the agenda, and the server sends notifications to the terminal as needed. The input is the audio data from the meeting, and the output is a notification resulting from the analysis of that data.

[0194] Step 6:

[0195] If a discussion deviates from the topic, the server sends a notification to the user's device prompting them to return to the subject. The device displays the message, "The current topic has deviated from the agenda." This allows users to appropriately steer the meeting back on track.

[0196] Step 7:

[0197] To perform sentiment analysis during meetings, the server uses an emotion engine. It analyzes the emotional data of meeting participants in real time and adjusts the meeting's topic and timing as needed. Input is data from emotion devices, and output is meeting adjustments based on that analysis.

[0198] Step 8:

[0199] After the meeting ends, the server automatically generates meeting minutes based on audio and sentiment data. This process involves analyzing what was said and identifying the next actions for each participant. The output is formalized as a report and provided to all participants.

[0200] Step 9:

[0201] Finally, the server sends the generated meeting minutes and reports to the participants. The data is output via email or cloud storage service, allowing participants to see what their next steps are.

[0202] (Application Example 2)

[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0204] Modern meetings and conferences require not only efficient progress but also an understanding of participants' feelings and appropriate responses. However, many meetings deviate from the agenda and cause dissatisfaction among participants, resulting in inefficient waste of time. Furthermore, effectively capturing and reflecting participants' comments and feelings during a meeting is difficult, hindering the improvement of meeting effectiveness. Similar challenges exist in team meetings on factory floors, highlighting the increasing need for progress management based on real-time sentiment analysis.

[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0206] In this invention, the server includes means having a module for receiving data on the purpose and scheduled time of a meeting in natural language, means having a module for automatically generating a meeting schedule, and means having a mechanism for monitoring speech content in real time using speech recognition technology. This makes it possible to centrally grasp the speech and feelings of attendees and improve the efficiency and quality of the meeting.

[0207] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format suitable for processing by machines.

[0208] A "progress plan" is a schedule or agenda that is automatically generated based on the purpose and duration of a meeting or conference.

[0209] "Speech recognition technology" is a technology that converts speech into text data in real time and is used to monitor speeches in a venue.

[0210] "Topic deviation detection" is a function that detects when a meeting's topic deviates from the designated agenda and prompts the user to return to the original topic at an appropriate time.

[0211] "Automatic meeting minutes generation" is a process that automatically creates a record of a meeting based on the content of the discussions during the meeting.

[0212] The "emotion analysis module" is a system that analyzes the emotions of attendees during a meeting in real time, and incorporates the results into the progress management.

[0213] "Adjusting the meeting's progress" is the process of flexibly changing the meeting's flow, taking into account the participants' emotional state and comments, as the meeting progresses.

[0214] A "communication mechanism" is a system that provides a means of notifying participants of their next course of action after a meeting.

[0215] This invention provides a system that efficiently and emotionally supports meetings and team meetings within factories. This system converts user input and speech from audio data into text data, analyzes its content and emotional state, and dynamically adjusts the meeting plan. The server uses speech recognition technology to capture speech in real time. It receives input of the meeting's purpose and scheduled time in natural language, generates an appropriate meeting plan, and notifies the user via their terminal. If the topic deviates from the meeting's subject, it can notify the user to return to the subject at an appropriate time.

[0216] The server includes an emotion analysis module that analyzes attendees' emotions in real time. This data is used to adjust the meeting plan. For example, if attendees show signs of agitation or dissatisfaction, the server will flexibly revise the meeting to encourage a more constructive discussion. This makes meetings more efficient while also considering the emotions of the participants.

[0217] The device has the ability to transcribe audio data into text in real time and receives the generated progress plan and sentiment analysis results through communication with the server. Because users can monitor the progress through the device, meetings can proceed smoothly.

[0218] As a concrete example, in a factory production team meeting, the team leader sets the objective as "an idea-generating meeting to improve production efficiency" and enters a one-hour time slot. If dissatisfaction among participants is detected during the meeting, the server adjusts the meeting plan based on this information to improve the quality of the exchange of ideas among participants. This system also provides emotional support to participants, making the meeting more productive.

[0219] An example of a prompt to the generative AI model is, "Analyze the participants' emotions from this conversation and suggest the optimal course of action for the meeting." In this way, the present invention aims to improve the efficiency of meeting management and the emotional satisfaction of participants.

[0220] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0221] Step 1:

[0222] The user inputs the purpose and scheduled time of the meeting into the terminal using natural language. The terminal converts the input voice data into text data using speech recognition technology. This text data is then sent to the server.

[0223] Step 2:

[0224] The server generates a meeting schedule based on the received text data. This schedule includes an agenda and timeline to facilitate efficient meeting management. The generated schedule is then fed back to the terminal.

[0225] Step 3:

[0226] As the meeting progresses, the server uses speech recognition technology to monitor the content of the discussion in real time. If a participant's comments deviate from the set agenda, the server sends a notification to the terminal, prompting the user to return to the topic.

[0227] Step 4:

[0228] The server uses an emotion analysis module to analyze the emotional state of meeting attendees in real time. This generates emotional data for the attendees, and the meeting plan is flexibly adjusted as needed. During this process, the emotional data is also analyzed using a generative AI model.

[0229] Step 5:

[0230] At the end of the meeting, the server automatically generates meeting minutes based on the participants' comments. It also identifies and lists each participant's next action. This data is then communicated to participants via their devices.

[0231] Step 6:

[0232] The generated sentiment data and progress plan are integrated to provide emotional evaluations and insights into the meeting's progress. Based on these results, users can clearly identify areas for improvement and plan future meetings.

[0233] Step 7:

[0234] For example, if a meeting is scheduled regarding the launch of a new product, the server will adjust the schedule and conduct the meeting at the optimal time when participants show excitement or interest. An example of a prompt that balances efficiency and emotional satisfaction in meeting management is, "Analyze the participants' emotions from this conversation and propose the optimal way to conduct the meeting."

[0235] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0236] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0237] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0238] [Second Embodiment]

[0239] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0240] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0241] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0242] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0243] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0244] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0245] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0246] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0247] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0248] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0249] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0250] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0251] This invention is a system for supporting the efficient progress of meetings, and is configured as follows: This system consists of a user, a terminal, and a server, which work together to assist in the progress of the meeting.

[0252] When a user starts a meeting, they input the meeting's purpose and scheduled time into their device using natural language. This collects the basic information necessary for the meeting. The device then converts the audio data into text and sends it to the server.

[0253] The server analyzes the received text data and automatically generates a meeting schedule based on the meeting's purpose and scheduled time. The schedule includes the main agenda items and timeline. This schedule is then used to prepare for and support the meeting's progress.

[0254] During the meeting, the terminal and server utilize speech recognition technology to monitor the meeting content in real time. As the meeting progresses, the content is recorded by transcribing speech into text, and if the topic deviates, the server sends a notification at an appropriate time to prompt the user to return to the meeting's subject. This prompt is displayed on the terminal and, if necessary, communicated to the user verbally.

[0255] Once the meeting concludes, the server automatically generates meeting minutes based on the recorded statements. These minutes include key points, conclusions, and decisions made during the meeting. They also identify and summarize the next steps for each participant. The server organizes this information and sends notifications regarding these next steps to each participant via their terminal. These notifications arrive as emails, ensuring continued follow-up even after the meeting ends.

[0256] For example, a user might define the purpose as "meeting about the annual budget" and set the scheduled time to "1 hour." The server would then set three agenda items: revenue and expenditure forecast, cost reduction proposals, and revenue improvement measures, allocating 20 minutes to each. If participants stray from the agenda during the meeting, the server would send a reminder to bring them back to the main topic, ensuring the meeting runs smoothly.

[0257] In this way, this system improves meeting productivity and helps participants efficiently share information and move on to the next step.

[0258] The following describes the processing flow.

[0259] Step 1:

[0260] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device then converts this audio data into text data.

[0261] Step 2:

[0262] The server receives text data sent from the terminal and uses natural language processing to analyze the purpose and scheduled time of the meeting. It then generates a meeting schedule.

[0263] Step 3:

[0264] The server creates a progress plan, sets the meeting schedule and agenda based on it, and sends it to the terminal.

[0265] Step 4:

[0266] During the meeting, the terminal utilizes speech recognition technology to monitor the user's speech in real time, convert the audio data into text data, and send it to the server.

[0267] Step 5:

[0268] The server analyzes the meeting discussions in real time, and if it detects that the topic has deviated from the set agenda, it sends a notification to the device prompting the participant to return to the topic at an appropriate time.

[0269] Step 6:

[0270] Once the meeting ends, the server aggregates the recorded text data of the speeches and automatically generates meeting minutes. Furthermore, it identifies and organizes the next actions for each participant.

[0271] Step 7:

[0272] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, thereby facilitating follow-up after the meeting.

[0273] (Example 1)

[0274] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0275] In meetings, participants often deviate from the agenda, and meeting minutes are frequently left ambiguous, leading to decreased meeting efficiency and delays in important decisions. Furthermore, there is a problem with insufficient follow-up, as specific next steps are not promptly communicated to participants after the meeting.

[0276] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0277] In this invention, the server includes a device that accepts input of agenda items and scheduled times in natural language, a device that converts voice data into text, and a device that analyzes the converted data to automatically generate a meeting plan based on the agenda. This prevents deviations from the agenda during the meeting, enables efficient progress and the generation of clear meeting minutes, and also strengthens post-meeting follow-up by quickly notifying participants of their next steps.

[0278] A "device that accepts input of agenda items and scheduled times in natural language" is a device that allows users to input meeting topics and estimated times verbally or in text, and then converts and processes that information into digital signals.

[0279] A "device that converts audio data to text" is a device that analyzes voice input from a user and accurately transcribes its content into text. This device processes audio signals and converts them into text data.

[0280] A "device that analyzes converted data and automatically generates a progress plan based on the agenda" is a device that analyzes the content of a transcribed meeting and automatically creates the meeting procedure and time allocation based on pre-set objectives.

[0281] The "device for real-time monitoring of conversations by utilizing voice recognition technology" is a device that instantaneously converts conversations during a meeting into text, analyzes the content, and monitors deviations from the scheduled topics.

[0282] The "device for detecting deviations from topics and prompting to return to the topic at an appropriate timing" is a device that detects statements deviating from the original topic from the conversations monitored in real time and instructs the participants to return to the original topic.

[0283] The "device for automatically generating minutes of a meeting and identifying the next actions of each participant" is a device that records the meeting content, organizes important decisions and next steps, and automatically documents them.

[0284] The "communication device for notifying each participant of the minutes of a meeting and the next actions" is a device that transmits the generated minutes of a meeting and future instructions to each meeting participant using communication means such as e-mail.

[0285] The present invention is a system that efficiently supports the progress of a meeting, in which a user, a terminal, and a server cooperate to function. Specific embodiments thereof will be described below.

[0286] When a user starts a meeting, the user inputs the purpose of the meeting and the scheduled time into the terminal in natural language. The terminal receives this as voice input and converts the voice data into text using the Google Speech-to-Text API. Through this process, the voice signal is converted into accurate text data.

[0287] Next, the terminal transmits the converted text data to the server. The server receives this data and performs analysis using IBM Watson Natural Language Understanding. Through this analysis, the purpose of the meeting and important topics are identified, and the server automatically generates a meeting progress plan based on this. The progress plan includes topics and necessary time schedules.

[0288] As the meeting progresses, the terminal and server monitor the conversation in real time using speech recognition technology. If a participant deviates from the agenda during the meeting, the server detects this and sends a reminder to the user via the terminal at an appropriate time to return to the topic. Specifically, this may involve displaying a message on the terminal such as "The current topic has deviated from the scheduled agenda," or an audio notification may also be considered.

[0289] At the end of the meeting, the server automatically generates meeting minutes based on the recorded speech data. These minutes include the meeting's objectives, key decisions, and next steps. The server uses the Microsoft Word API to format the minutes as a document and provides it to each participant.

[0290] The server then identifies each participant's next action and notifies them via email through their device. This notification, delivered using the Gmail API, includes specific instructions and follow-ups. An example of a prompt sent to a participant might be, "Please consider the next steps regarding the annual budget and prepare your proposal for the next meeting."

[0291] This system provides support to help users efficiently conduct meetings and share information. By optimizing prompt sentences using a generative AI model, even faster and more effective communication becomes possible.

[0292] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0293] Step 1:

[0294] User input of meeting information

[0295] To start a meeting, the user enters the purpose and scheduled time of the meeting into their device using natural language. The entered information is received as voice or text data. This input information becomes the basic data for the system. The user provides data such as "About next year's budget proposal, approximately one hour" using the voice input function of their smartphone.

[0296] Step 2:

[0297] Converting audio data to text

[0298] The device converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to analyze the audio signal and generate data as a string. This step converts the audio information into readable text. The output is the text "Regarding next year's budget proposal, about 1 hour."

[0299] Step 3:

[0300] Text data transmission and analysis

[0301] The terminal sends the converted text data to the server. The server analyzes the received data using IBM Watson Natural Language Understanding. This analysis extracts the meeting's objectives and important keywords, and organizes the information necessary for the meeting's progress. Text data is used as input, and the output is the analyzed information structure.

[0302] Step 4:

[0303] Automatic generation of meeting schedules

[0304] The server automatically generates a meeting progress plan based on the analysis results. This is done while considering the meeting topic and time allocation, and the progress plan includes information such as each topic and the time allocated to the topic. As output, specific progress procedures such as "Review of the budget plan", "Allocation of resources", and "Cost reduction measures" are set.

[0305] Step 5:

[0306] Real-time monitoring and reminder sending

[0307] The terminal and the server monitor the speech during the meeting in real time. By using speech recognition, the conversation during the meeting is texturized, and a reminder is sent when it deviates from the progress plan. The server compares with the scheduled items and detects that the topic has deviated. At an appropriate timing, the terminal is displayed with "The topic has deviated from the agenda. Let's return to the main topic."

[0308] Step 6:

[0309] Generation of meeting minutes after the meeting

[0310] The server automatically generates meeting minutes based on the speech content recorded during the meeting. In this process, text data is documented using the Microsoft Word API, and important information for the participants is organized. As output, a meeting minutes document is generated, which includes information such as "Approval of the budget plan", "Next meeting date and time", and "Action items".

[0311] Step 7:

[0312] Notification of the next action

[0313] The server notifies each participant of the generated meeting minutes and the identified next action. The notification is sent via email through the terminal using the Gmail API. This allows participants to understand the key points of the meeting and receive specific instructions for the next steps. As output, a follow-up email is sent to each participant.

[0314] (Application Example 1)

[0315] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0316] In the face of the need for effective meeting management and operational efficiency in physical stores, traditional methods often fail to keep participants focused on the agenda, resulting in insufficient achievement of meeting objectives. Furthermore, a lack of follow-up and action checks after meetings leads to decreased productivity. To overcome these challenges, there is a need to develop a system that manages meetings efficiently and systematically, while providing appropriate support to participants.

[0317] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0318] In this invention, the server includes an information receiving means for inputting data on the purpose and scheduled time of the meeting in natural language, a planning means for automatically generating a meeting schedule, and an information providing means for providing hints and suggestions based on data in real time during the meeting. This enables participants to concentrate on the meeting's topic, effectively advance discussions, and ensure follow-up after the meeting.

[0319] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format that computer systems can understand and process.

[0320] "Information receiving means" refers to a device or method for receiving the purpose and scheduled time of a meeting, entered by the user, in digital format.

[0321] "Planning means" refers to algorithms or devices for automatically generating a meeting schedule.

[0322] "Speech recognition technology" is a technology that converts speech data into text data in real time and understands the content of what is being said.

[0323] "Monitoring methods" refer to technologies that monitor speeches during a meeting in real time and record and analyze relevant data.

[0324] "Correction measures" are technologies or devices that detect when a meeting deviates from its topic and prompt participants to return to the topic at an appropriate time.

[0325] "Action identification means" refers to a method or device for automatically identifying the next steps for each individual who participated in the meeting.

[0326] "Communication means" refers to electronic methods or devices used to notify participants of an assembly of information.

[0327] "Information provision means" refers to technologies and devices that provide hints and suggestions based on real-time data during a meeting.

[0328] The system for realizing this invention is designed to enable employees to conduct meetings more efficiently. In the in-store meeting support system, the main hardware used is a smartphone or tablet to collect audio data. The information receiving means installed in these devices can receive natural language data entered by the user. Specifically, the software components include speech recognition technology that converts audio data into text data in real time, using Python and its library, SpeechRecognition.

[0329] After receiving text data, the server performs text analysis using the natural language processing library spaCy. Next, a meeting agenda is automatically generated. During the meeting, the content of the discussion is monitored using Google Cloud's Natural Language API. In addition, hints and suggestions are provided in real time through various information channels to help participants focus on the meeting's main topic.

[0330] As an example, consider the use of this system by a store operations team when holding a new product promotion meeting. In this meeting, participants choose "New Product Promotion Plan" as the topic and set a one-hour time limit. Based on this data, the server automatically sets the agenda and time allocation, and notifies participants through corrective means if they deviate from the discussion. Furthermore, after the meeting, each participant receives a notification regarding the next steps.

[0331] Examples of prompt statements to input into a generative AI model are as follows:

[0332] "The purpose of today's meeting is to discuss the 'marketing plan for the new product,' and it will last one hour. Please provide a meeting plan based on this. Also, monitor the progress of each agenda item and send a notification if the discussion goes off track."

[0333] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0334] Step 1:

[0335] The user enters the purpose and scheduled time of the meeting into the device using natural language. The entered information is collected by the device's information receiving mechanism. The input data is converted into text using speech recognition technology and sent to the next step.

[0336] Step 2:

[0337] The terminal converts audio data into text data in real time using the SpeechRecognition library. The converted data is sent to the server via the internet. The output here is text data that includes the purpose and scheduled time of the meeting.

[0338] Step 3:

[0339] The server receives text data and performs natural language processing using the spaCy library. This analyzes the purpose of the meeting and the priority agenda items, and automatically generates a meeting plan. The generated plan includes the time allocation for each agenda item. This plan is used in the next step.

[0340] Step 4:

[0341] During the meeting, the device collects audio data again, transcribes it into text in real time, and sends it to the server. The server uses Google Cloud's Natural Language API to monitor the content of the discussion as the meeting progresses. If the topic deviates, the server automatically detects the deviation and generates a notification prompting the user to correct it.

[0342] Step 5:

[0343] At the end of the meeting, the server automatically generates meeting minutes using the collected data. These minutes will include key points from the meeting and the next steps for participants. The generated minutes will be sent to each participant via their terminal.

[0344] Step 6:

[0345] After the meeting, the server sends participants notifications about the next steps via email or other means of communication. This ensures that each participant is clearly aware of their next actions and can follow up effectively.

[0346] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0347] This invention provides a system that enables efficient and emotionally insightful meeting management, and has the following configuration: This system includes a user, a terminal, a server, and an emotion engine, which work together to support the progress of the meeting.

[0348] First, at the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device converts the entered information from voice data to text data and sends it to the server.

[0349] Next, the server analyzes the text data and automatically generates a meeting plan based on the meeting's purpose and duration. This plan includes a time schedule and agenda, enabling efficient meeting management. The meeting plan is then fed back to the user via their terminal.

[0350] During the meeting, the terminal and server use speech recognition technology to monitor the content of the discussion in real time. If the discussion deviates from the set agenda, the server sends a notification to the terminal prompting the user to return to the topic at an appropriate time. Furthermore, once the meeting ends, the server automatically generates meeting minutes based on the content of the discussion, and identifies and organizes the next actions for each participant. This ensures that follow-up is provided even after the meeting has ended.

[0351] The emotion engine, a key feature of this invention, analyzes users' emotions in real time during a meeting. Based on the emotion data obtained from the emotion engine, the server flexibly adjusts the progress of the meeting. For example, if a participant is feeling dissatisfied, the server can adjust the timing of changing topics. Furthermore, the meeting minutes and reports generated after the meeting also include the emotion data, providing emotional evaluations and insights into how the meeting proceeded.

[0352] As a concrete example, a user sets the objective as "Meeting about the launch of a new product" and enters a scheduled time of "2 hours." During the meeting, if the emotion engine detects signs of agitation in some participants, the server takes this into consideration and adjusts the flow of the meeting to help ensure a constructive discussion. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0353] The following describes the processing flow.

[0354] Step 1:

[0355] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device receives the input as voice data and converts it into text data.

[0356] Step 2:

[0357] The terminal sends the converted text data to the server. The server analyzes the purpose and duration of the meeting and automatically generates a meeting schedule.

[0358] Step 3:

[0359] The server generates a progress plan and sends it to the terminal, which then displays it to the user. The progress plan includes the agenda and time schedule.

[0360] Step 4:

[0361] During the meeting, the terminal uses speech recognition technology to transcribe the user's speech into text in real time and send it to the server.

[0362] Step 5:

[0363] The server analyzes the received message data and detects when the topic deviates from the agenda. It then sends a notification to the device prompting the user to return to the topic at an appropriate time.

[0364] Step 6:

[0365] The server uses an emotion engine to analyze the emotions of participants during a meeting in real time and detects emotional data.

[0366] Step 7:

[0367] The server flexibly adjusts the flow of the meeting based on sentiment data and suggests new discussion transitions as needed.

[0368] Step 8:

[0369] Once the meeting ends, the server automatically generates meeting minutes based on the collected speech and sentiment data, and also identifies the next actions for each participant.

[0370] Step 9:

[0371] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, and then follows up after the meeting.

[0372] Step 10:

[0373] The server helps participants gain emotional insights by including emotional data in the meeting report and providing it to them.

[0374] (Example 2)

[0375] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0376] Traditional meeting systems have had problems such as inefficient progress management and meetings sometimes going in an inappropriate direction because participants' feelings are not taken into consideration. Specifically, problems include wasted time due to off-topic discussions, difficulty in tracking meeting content, and the accumulation of participant dissatisfaction.

[0377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0378] In this invention, the server includes information processing means for analyzing natural language input data received via voice and determining the purpose and scheduled time of the meeting; schedule generation means for automatically generating a meeting plan based on the input data and notifying participants; and speech recognition technology means for monitoring speech in real time during the meeting and supporting the progress of the meeting in accordance with the agenda. This enables efficient management of the meeting and flexible meeting adjustments based on the emotions of the participants.

[0379] "Information processing means" refers to a device or program that has the function of analyzing natural language input data received via voice and determining the purpose and scheduled time of a meeting.

[0380] A "schedule generation means" is a device or program that has the function of automatically generating a meeting schedule based on input data and notifying participants of that schedule.

[0381] "Speech recognition technology means" refers to a device or program that uses technology to monitor speech during a meeting in real time and support the progress of the meeting in accordance with the agenda.

[0382] "Detection and notification means" refers to a device or program that has the function of detecting when a statement deviates from the agenda and prompting a return to the topic at an appropriate time.

[0383] "Document generation means" refers to a device or program that has the function of automatically generating meeting minutes based on audio data and clearly indicating the next action plan for each participant.

[0384] A "sentiment analysis module" is a device or program that analyzes the emotional data of meeting participants and has the function of adjusting the progress of the meeting.

[0385] "Communication technology means" refers to a device or program that has a communication function for informing each participant of the next action to take after a meeting.

[0386] "Data conversion means" refers to a device or program that has the function of converting audio data into text data in real time.

[0387] Modes for carrying out the invention

[0388] This invention is a system for facilitating smooth meeting progress and adjusting based on participants' emotions. The system is implemented through the collaboration of a server, terminals, and users, and utilizes the following hardware and software.

[0389] The device is equipped with speech recognition software that accepts user input using natural language. Specific examples of such software include Google Cloud Speech-to-Text and IBM Watson Speech to Text. This converts speech data into text data.

[0390] The server processes the input text data and automatically generates a meeting agenda using a generative AI model. This generated agenda includes the meeting's time schedule and specific agenda items.

[0391] Furthermore, the server utilizes an emotion engine to analyze participants' emotions in real time based on data from devices such as NeuroSky and Emotiv, thereby adjusting the meeting's progress. If participants are dissatisfied, the server can take action, such as changing the agenda.

[0392] During the meeting, the terminal and server work together to provide a system that monitors speech in real time. If a participant's remarks stray from the agenda, the server will notify them to return to the topic within the allotted time, thus appropriately supporting the progress of the meeting.

[0393] Finally, the server automatically generates meeting minutes based on the audio data of the meeting, including content and sentiment data, and creates a report that clearly communicates the next steps for each participant.

[0394] As a concrete example, a user sets the purpose as "Meeting about the launch of a new product" and enters the scheduled time as "2 hours." During the meeting, if the emotion engine detects that a participant is agitated, the server constructively adjusts the flow of the meeting. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0395] Example of a prompt

[0396] "Please use a generative AI model to create meeting minutes."

[0397] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0398] Step 1:

[0399] At the start of the meeting, the user enters the purpose and scheduled time into the terminal. Specifically, they provide information to the terminal by voice, such as "Meeting regarding the launch of a new product, 2 hours." The entered voice data is converted into text data by the terminal's voice recognition software. This step involves data processing, where voice data is converted into text.

[0400] Step 2:

[0401] The terminal sends the meeting's purpose and scheduled time as data to the server. Secure protocols such as HTTPS are used for transmission. This ensures that user input is safely transmitted to the server. The output data is sent to the server as a text file within the software.

[0402] Step 3:

[0403] The server processes the received text data. Using a generative AI model, it automatically generates a meeting schedule. This involves analyzing the text data and proposing appropriate time schedules and agenda items. This plan is then ready for feedback.

[0404] Step 4:

[0405] The server notifies the user of the generated schedule via the terminal. The terminal presents this information to the user through push notifications or display. In this step, the data output is provided visually through the user interface, preparing the meeting.

[0406] Step 5:

[0407] During the meeting, the terminal and server work together to use speech recognition technology to monitor speech in real time. They check whether the speech is relevant to the agenda, and the server sends notifications to the terminal as needed. The input is the audio data from the meeting, and the output is a notification resulting from the analysis of that data.

[0408] Step 6:

[0409] If a discussion deviates from the topic, the server sends a notification to the user's device prompting them to return to the subject. The device displays the message, "The current topic has deviated from the agenda." This allows users to appropriately steer the meeting back on track.

[0410] Step 7:

[0411] To perform sentiment analysis during meetings, the server uses an emotion engine. It analyzes the emotional data of meeting participants in real time and adjusts the meeting's topic and timing as needed. Input is data from emotion devices, and output is meeting adjustments based on that analysis.

[0412] Step 8:

[0413] After the meeting ends, the server automatically generates meeting minutes based on audio and sentiment data. This process involves analyzing what was said and identifying the next actions for each participant. The output is formalized as a report and provided to all participants.

[0414] Step 9:

[0415] Finally, the server sends the generated meeting minutes and reports to the participants. The data is output via email or cloud storage service, allowing participants to see what their next steps are.

[0416] (Application Example 2)

[0417] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0418] Modern meetings and conferences require not only efficient progress but also an understanding of participants' feelings and appropriate responses. However, many meetings deviate from the agenda and cause dissatisfaction among participants, resulting in inefficient waste of time. Furthermore, effectively capturing and reflecting participants' comments and feelings during a meeting is difficult, hindering the improvement of meeting effectiveness. Similar challenges exist in team meetings on factory floors, highlighting the increasing need for progress management based on real-time sentiment analysis.

[0419] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0420] In this invention, the server includes means having a module for receiving data on the purpose and scheduled time of a meeting in natural language, means having a module for automatically generating a meeting schedule, and means having a mechanism for monitoring speech content in real time using speech recognition technology. This makes it possible to centrally grasp the speech and feelings of attendees and improve the efficiency and quality of the meeting.

[0421] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format suitable for processing by machines.

[0422] A "progress plan" is a schedule or agenda that is automatically generated based on the purpose and duration of a meeting or conference.

[0423] "Speech recognition technology" is a technology that converts speech into text data in real time and is used to monitor speeches in a venue.

[0424] "Topic deviation detection" is a function that detects when a meeting's topic deviates from the designated agenda and prompts the user to return to the original topic at an appropriate time.

[0425] "Automatic meeting minutes generation" is a process that automatically creates a record of a meeting based on the content of the discussions during the meeting.

[0426] The "emotion analysis module" is a system that analyzes the emotions of attendees during a meeting in real time, and incorporates the results into the progress management.

[0427] "Adjusting the meeting's progress" is the process of flexibly changing the meeting's flow, taking into account the participants' emotional state and comments, as the meeting progresses.

[0428] A "communication mechanism" is a system that provides a means of notifying participants of their next course of action after a meeting.

[0429] This invention provides a system that efficiently and emotionally supports meetings and team meetings within factories. This system converts user input and speech from audio data into text data, analyzes its content and emotional state, and dynamically adjusts the meeting plan. The server uses speech recognition technology to capture speech in real time. It receives input of the meeting's purpose and scheduled time in natural language, generates an appropriate meeting plan, and notifies the user via their terminal. If the topic deviates from the meeting's subject, it can notify the user to return to the subject at an appropriate time.

[0430] The server includes an emotion analysis module that analyzes attendees' emotions in real time. This data is used to adjust the meeting plan. For example, if attendees show signs of agitation or dissatisfaction, the server will flexibly revise the meeting to encourage a more constructive discussion. This makes meetings more efficient while also considering the emotions of the participants.

[0431] The device has the ability to transcribe audio data into text in real time and receives the generated progress plan and sentiment analysis results through communication with the server. Because users can monitor the progress through the device, meetings can proceed smoothly.

[0432] As a concrete example, in a factory production team meeting, the team leader sets the objective as "an idea-generating meeting to improve production efficiency" and enters a one-hour time slot. If dissatisfaction among participants is detected during the meeting, the server adjusts the meeting plan based on this information to improve the quality of the exchange of ideas among participants. This system also provides emotional support to participants, making the meeting more productive.

[0433] An example of a prompt to the generative AI model is, "Analyze the participants' emotions from this conversation and suggest the optimal course of action for the meeting." In this way, the present invention aims to improve the efficiency of meeting management and the emotional satisfaction of participants.

[0434] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0435] Step 1:

[0436] The user inputs the purpose and scheduled time of the meeting into the terminal using natural language. The terminal converts the input voice data into text data using speech recognition technology. This text data is then sent to the server.

[0437] Step 2:

[0438] The server generates a meeting schedule based on the received text data. This schedule includes an agenda and timeline to facilitate efficient meeting management. The generated schedule is then fed back to the terminal.

[0439] Step 3:

[0440] As the meeting progresses, the server uses speech recognition technology to monitor the content of the discussion in real time. If a participant's comments deviate from the set agenda, the server sends a notification to the terminal, prompting the user to return to the topic.

[0441] Step 4:

[0442] The server uses an emotion analysis module to analyze the emotional state of meeting attendees in real time. This generates emotional data for the attendees, and the meeting plan is flexibly adjusted as needed. During this process, the emotional data is also analyzed using a generative AI model.

[0443] Step 5:

[0444] At the end of the meeting, the server automatically generates meeting minutes based on the participants' comments. It also identifies and lists each participant's next action. This data is then communicated to participants via their devices.

[0445] Step 6:

[0446] The generated sentiment data and progress plan are integrated to provide emotional evaluations and insights into the meeting's progress. Based on these results, users can clearly identify areas for improvement and plan future meetings.

[0447] Step 7:

[0448] For example, if a meeting is scheduled regarding the launch of a new product, the server will adjust the schedule and conduct the meeting at the optimal time when participants show excitement or interest. An example of a prompt that balances efficiency and emotional satisfaction in meeting management is, "Analyze the participants' emotions from this conversation and propose the optimal way to conduct the meeting."

[0449] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0450] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0451] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0452] [Third Embodiment]

[0453] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0454] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0455] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0456] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0457] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0458] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0459] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0460] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0461] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0462] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0463] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0464] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0465] This invention is a system for supporting the efficient progress of meetings, and is configured as follows: This system consists of a user, a terminal, and a server, which work together to assist in the progress of the meeting.

[0466] When a user starts a meeting, they input the meeting's purpose and scheduled time into their device using natural language. This collects the basic information necessary for the meeting. The device then converts the audio data into text and sends it to the server.

[0467] The server analyzes the received text data and automatically generates a meeting schedule based on the meeting's purpose and scheduled time. The schedule includes the main agenda items and timeline. This schedule is then used to prepare for and support the meeting's progress.

[0468] During the meeting, the terminal and server utilize speech recognition technology to monitor the meeting content in real time. As the meeting progresses, the content is recorded by transcribing speech into text, and if the topic deviates, the server sends a notification at an appropriate time to prompt the user to return to the meeting's subject. This prompt is displayed on the terminal and, if necessary, communicated to the user verbally.

[0469] Once the meeting concludes, the server automatically generates meeting minutes based on the recorded statements. These minutes include key points, conclusions, and decisions made during the meeting. They also identify and summarize the next steps for each participant. The server organizes this information and sends notifications regarding these next steps to each participant via their terminal. These notifications arrive as emails, ensuring continued follow-up even after the meeting ends.

[0470] For example, a user might define the purpose as "meeting about the annual budget" and set the scheduled time to "1 hour." The server would then set three agenda items: revenue and expenditure forecast, cost reduction proposals, and revenue improvement measures, allocating 20 minutes to each. If participants stray from the agenda during the meeting, the server would send a reminder to bring them back to the main topic, ensuring the meeting runs smoothly.

[0471] In this way, this system improves meeting productivity and helps participants efficiently share information and move on to the next step.

[0472] The following describes the processing flow.

[0473] Step 1:

[0474] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device then converts this audio data into text data.

[0475] Step 2:

[0476] The server receives text data sent from the terminal and uses natural language processing to analyze the purpose and scheduled time of the meeting. It then generates a meeting schedule.

[0477] Step 3:

[0478] The server creates a progress plan, sets the meeting schedule and agenda based on it, and sends it to the terminal.

[0479] Step 4:

[0480] During the meeting, the terminal utilizes speech recognition technology to monitor the user's speech in real time, convert the audio data into text data, and send it to the server.

[0481] Step 5:

[0482] The server analyzes the meeting discussions in real time, and if it detects that the topic has deviated from the set agenda, it sends a notification to the device prompting the participant to return to the topic at an appropriate time.

[0483] Step 6:

[0484] Once the meeting ends, the server aggregates the recorded text data of the speeches and automatically generates meeting minutes. Furthermore, it identifies and organizes the next actions for each participant.

[0485] Step 7:

[0486] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, thereby facilitating follow-up after the meeting.

[0487] (Example 1)

[0488] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0489] In meetings, participants often deviate from the agenda, and meeting minutes are frequently left ambiguous, leading to decreased meeting efficiency and delays in important decisions. Furthermore, there is a problem with insufficient follow-up, as specific next steps are not promptly communicated to participants after the meeting.

[0490] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0491] In this invention, the server includes a device that accepts input of agenda items and scheduled times in natural language, a device that converts voice data into text, and a device that analyzes the converted data to automatically generate a meeting plan based on the agenda. This prevents deviations from the agenda during the meeting, enables efficient progress and the generation of clear meeting minutes, and also strengthens post-meeting follow-up by quickly notifying participants of their next steps.

[0492] A "device that accepts input of agenda items and scheduled times in natural language" is a device that allows users to input meeting topics and estimated times verbally or in text, and then converts and processes that information into digital signals.

[0493] A "device that converts audio data to text" is a device that analyzes voice input from a user and accurately transcribes its content into text. This device processes audio signals and converts them into text data.

[0494] A "device that analyzes converted data and automatically generates a progress plan based on the agenda" is a device that analyzes the content of a transcribed meeting and automatically creates the meeting procedure and time allocation based on pre-set objectives.

[0495] A "device that uses speech recognition technology to monitor conversations in real time" is a device that instantly transcribes conversations during a meeting into text, analyzes the content, and monitors deviations from the planned agenda.

[0496] A "device that detects deviations from the agenda and prompts participants to return to the agenda at the appropriate time" is a device that detects statements that deviate from the original agenda from conversations monitored in real time and instructs participants to return to the original agenda.

[0497] A "device that automatically generates meeting minutes and identifies each participant's next action" is a device that records meeting content, organizes important decisions and next steps, and automatically documents them.

[0498] A "communication device for notifying each participant of the minutes and subsequent actions" is a device for transmitting the generated minutes and future instructions to each meeting participant using communication methods such as email.

[0499] This invention is a system that efficiently supports the progress of meetings, in which users, terminals, and servers work together. Specific embodiments are described below.

[0500] When a user starts a meeting, they input the meeting's purpose and scheduled time into the device using natural language. The device receives this as voice input and converts the voice data into text using the Google Speech-to-Text API. This process ensures that the voice signal is accurately converted into text data.

[0501] Next, the terminal sends the converted text data to the server. The server receives this data and performs analysis using IBM Watson Natural Language Understanding. This analysis identifies the meeting's objectives and key topics, and the server automatically generates a meeting plan based on this information. The plan includes the agenda and necessary time schedules.

[0502] As the meeting progresses, the terminal and server monitor the conversation in real time using speech recognition technology. If a participant deviates from the agenda during the meeting, the server detects this and sends a reminder to the user via the terminal at an appropriate time to return to the topic. Specifically, this may involve displaying a message on the terminal such as "The current topic has deviated from the scheduled agenda," or an audio notification may also be considered.

[0503] At the end of the meeting, the server automatically generates meeting minutes based on the recorded speech data. These minutes include the meeting's objectives, key decisions, and next steps. The server uses the Microsoft Word API to format the minutes as a document and provides it to each participant.

[0504] The server then identifies each participant's next action and notifies them via email through their device. This notification, delivered using the Gmail API, includes specific instructions and follow-ups. An example of a prompt sent to a participant might be, "Please consider the next steps regarding the annual budget and prepare your proposal for the next meeting."

[0505] This system provides support to help users efficiently conduct meetings and share information. By optimizing prompt sentences using a generative AI model, even faster and more effective communication becomes possible.

[0506] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0507] Step 1:

[0508] User input of meeting information

[0509] To start a meeting, the user enters the purpose and scheduled time of the meeting into their device using natural language. The entered information is received as voice or text data. This input information becomes the basic data for the system. The user provides data such as "About next year's budget proposal, approximately one hour" using the voice input function of their smartphone.

[0510] Step 2:

[0511] Converting audio data to text

[0512] The device converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to analyze the audio signal and generate data as a string. This step converts the audio information into readable text. The output is the text "Regarding next year's budget proposal, about 1 hour."

[0513] Step 3:

[0514] Text data transmission and analysis

[0515] The terminal sends the converted text data to the server. The server analyzes the received data using IBM Watson Natural Language Understanding. This analysis extracts the meeting's objectives and important keywords, and organizes the information necessary for the meeting's progress. Text data is used as input, and the output is the analyzed information structure.

[0516] Step 4:

[0517] Automatic generation of meeting schedules

[0518] The server automatically generates a meeting schedule based on the analysis results. This takes into account the meeting's topic and time allocation, and the schedule includes information such as each agenda item and the time allocated to each item. The output includes specific procedures such as "budget review," "resource allocation," and "cost reduction measures."

[0519] Step 5:

[0520] Real-time monitoring and reminder sending

[0521] The terminal and server monitor what is said during the meeting in real time. Using speech recognition, the conversation is transcribed into text, and a reminder is sent if it deviates from the agenda. The server compares it to the scheduled items and detects when the topic has gone off track. At the appropriate time, the terminal displays a message saying, "The topic has gone off-topic. Let's return to the main subject."

[0522] Step 6:

[0523] Meeting minutes generation after the meeting

[0524] The server automatically generates meeting minutes based on the recorded conversations during the meeting. This process uses the Microsoft Word API to document the text data and organize important information for participants. The output is a meeting minutes document containing information such as "budget approval," "next meeting date and time," and "action items."

[0525] Step 7:

[0526] Notification of the next action

[0527] The server notifies each participant of the generated meeting minutes and the identified next action. The notification is sent via email through the terminal using the Gmail API. This allows participants to understand the key points of the meeting and receive specific instructions for the next steps. As output, a follow-up email is sent to each participant.

[0528] (Application Example 1)

[0529] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0530] In the face of the need for effective meeting management and operational efficiency in physical stores, traditional methods often fail to keep participants focused on the agenda, resulting in insufficient achievement of meeting objectives. Furthermore, a lack of follow-up and action checks after meetings leads to decreased productivity. To overcome these challenges, there is a need to develop a system that manages meetings efficiently and systematically, while providing appropriate support to participants.

[0531] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0532] In this invention, the server includes an information receiving means for inputting data on the purpose and scheduled time of the meeting in natural language, a planning means for automatically generating a meeting schedule, and an information providing means for providing hints and suggestions based on data in real time during the meeting. This enables participants to concentrate on the meeting's topic, effectively advance discussions, and ensure follow-up after the meeting.

[0533] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format that computer systems can understand and process.

[0534] "Information receiving means" refers to a device or method for receiving the purpose and scheduled time of a meeting, entered by the user, in digital format.

[0535] "Planning means" refers to algorithms or devices for automatically generating a meeting schedule.

[0536] "Speech recognition technology" is a technology that converts speech data into text data in real time and understands the content of what is being said.

[0537] "Monitoring methods" refer to technologies that monitor speeches during a meeting in real time and record and analyze relevant data.

[0538] "Correction measures" are technologies or devices that detect when a meeting deviates from its topic and prompt participants to return to the topic at an appropriate time.

[0539] "Action identification means" refers to a method or device for automatically identifying the next steps for each individual who participated in the meeting.

[0540] "Communication means" refers to electronic methods or devices used to notify participants of an assembly of information.

[0541] "Information provision means" refers to technologies and devices that provide hints and suggestions based on real-time data during a meeting.

[0542] The system for realizing this invention is designed to enable employees to conduct meetings more efficiently. In the in-store meeting support system, the main hardware used is a smartphone or tablet to collect audio data. The information receiving means installed in these devices can receive natural language data entered by the user. Specifically, the software components include speech recognition technology that converts audio data into text data in real time, using Python and its library, SpeechRecognition.

[0543] After receiving text data, the server performs text analysis using the natural language processing library spaCy. Next, a meeting agenda is automatically generated. During the meeting, the content of the discussion is monitored using Google Cloud's Natural Language API. In addition, hints and suggestions are provided in real time through various information channels to help participants focus on the meeting's main topic.

[0544] As an example, consider the use of this system by a store operations team when holding a new product promotion meeting. In this meeting, participants choose "New Product Promotion Plan" as the topic and set a one-hour time limit. Based on this data, the server automatically sets the agenda and time allocation, and notifies participants through corrective means if they deviate from the discussion. Furthermore, after the meeting, each participant receives a notification regarding the next steps.

[0545] Examples of prompt statements to input into a generative AI model are as follows:

[0546] "The purpose of today's meeting is to discuss the 'marketing plan for the new product,' and it will last one hour. Please provide a meeting plan based on this. Also, monitor the progress of each agenda item and send a notification if the discussion goes off track."

[0547] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0548] Step 1:

[0549] The user enters the purpose and scheduled time of the meeting into the device using natural language. The entered information is collected by the device's information receiving mechanism. The input data is converted into text using speech recognition technology and sent to the next step.

[0550] Step 2:

[0551] The terminal converts audio data into text data in real time using the SpeechRecognition library. The converted data is sent to the server via the internet. The output here is text data that includes the purpose and scheduled time of the meeting.

[0552] Step 3:

[0553] The server receives text data and performs natural language processing using the spaCy library. This analyzes the purpose of the meeting and the priority agenda items, and automatically generates a meeting plan. The generated plan includes the time allocation for each agenda item. This plan is used in the next step.

[0554] Step 4:

[0555] During the meeting, the device collects audio data again, transcribes it into text in real time, and sends it to the server. The server uses Google Cloud's Natural Language API to monitor the content of the discussion as the meeting progresses. If the topic deviates, the server automatically detects the deviation and generates a notification prompting the user to correct it.

[0556] Step 5:

[0557] At the end of the meeting, the server automatically generates meeting minutes using the collected data. These minutes will include key points from the meeting and the next steps for participants. The generated minutes will be sent to each participant via their terminal.

[0558] Step 6:

[0559] After the meeting, the server sends participants notifications about the next steps via email or other means of communication. This ensures that each participant is clearly aware of their next actions and can follow up effectively.

[0560] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0561] This invention provides a system that enables efficient and emotionally insightful meeting management, and has the following configuration: This system includes a user, a terminal, a server, and an emotion engine, which work together to support the progress of the meeting.

[0562] First, at the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device converts the entered information from voice data to text data and sends it to the server.

[0563] Next, the server analyzes the text data and automatically generates a meeting plan based on the meeting's purpose and duration. This plan includes a time schedule and agenda, enabling efficient meeting management. The meeting plan is then fed back to the user via their terminal.

[0564] During the meeting, the terminal and server use speech recognition technology to monitor the content of the discussion in real time. If the discussion deviates from the set agenda, the server sends a notification to the terminal prompting the user to return to the topic at an appropriate time. Furthermore, once the meeting ends, the server automatically generates meeting minutes based on the content of the discussion, and identifies and organizes the next actions for each participant. This ensures that follow-up is provided even after the meeting has ended.

[0565] The emotion engine, a key feature of this invention, analyzes users' emotions in real time during a meeting. Based on the emotion data obtained from the emotion engine, the server flexibly adjusts the progress of the meeting. For example, if a participant is feeling dissatisfied, the server can adjust the timing of changing topics. Furthermore, the meeting minutes and reports generated after the meeting also include the emotion data, providing emotional evaluations and insights into how the meeting proceeded.

[0566] As a concrete example, a user sets the objective as "Meeting about the launch of a new product" and enters a scheduled time of "2 hours." During the meeting, if the emotion engine detects signs of agitation in some participants, the server takes this into consideration and adjusts the flow of the meeting to help ensure a constructive discussion. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0567] The following describes the processing flow.

[0568] Step 1:

[0569] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device receives the input as voice data and converts it into text data.

[0570] Step 2:

[0571] The terminal sends the converted text data to the server. The server analyzes the purpose and duration of the meeting and automatically generates a meeting schedule.

[0572] Step 3:

[0573] The server generates a progress plan and sends it to the terminal, which then displays it to the user. The progress plan includes the agenda and time schedule.

[0574] Step 4:

[0575] During the meeting, the terminal uses speech recognition technology to transcribe the user's speech into text in real time and send it to the server.

[0576] Step 5:

[0577] The server analyzes the received message data and detects when the topic deviates from the agenda. It then sends a notification to the device prompting the user to return to the topic at an appropriate time.

[0578] Step 6:

[0579] The server uses an emotion engine to analyze the emotions of participants during a meeting in real time and detects emotional data.

[0580] Step 7:

[0581] The server flexibly adjusts the flow of the meeting based on sentiment data and suggests new discussion transitions as needed.

[0582] Step 8:

[0583] Once the meeting ends, the server automatically generates meeting minutes based on the collected speech and sentiment data, and also identifies the next actions for each participant.

[0584] Step 9:

[0585] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, and then follows up after the meeting.

[0586] Step 10:

[0587] The server helps participants gain emotional insights by including emotional data in the meeting report and providing it to them.

[0588] (Example 2)

[0589] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0590] Traditional meeting systems have had problems such as inefficient progress management and meetings sometimes going in an inappropriate direction because participants' feelings are not taken into consideration. Specifically, problems include wasted time due to off-topic discussions, difficulty in tracking meeting content, and the accumulation of participant dissatisfaction.

[0591] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0592] In this invention, the server includes information processing means for analyzing natural language input data received via voice and determining the purpose and scheduled time of the meeting; schedule generation means for automatically generating a meeting plan based on the input data and notifying participants; and speech recognition technology means for monitoring speech in real time during the meeting and supporting the progress of the meeting in accordance with the agenda. This enables efficient management of the meeting and flexible meeting adjustments based on the emotions of the participants.

[0593] "Information processing means" refers to a device or program that has the function of analyzing natural language input data received via voice and determining the purpose and scheduled time of a meeting.

[0594] A "schedule generation means" is a device or program that has the function of automatically generating a meeting schedule based on input data and notifying participants of that schedule.

[0595] "Speech recognition technology means" refers to a device or program that uses technology to monitor speech during a meeting in real time and support the progress of the meeting in accordance with the agenda.

[0596] "Detection and notification means" refers to a device or program that has the function of detecting when a statement deviates from the agenda and prompting a return to the topic at an appropriate time.

[0597] "Document generation means" refers to a device or program that has the function of automatically generating meeting minutes based on audio data and clearly indicating the next action plan for each participant.

[0598] A "sentiment analysis module" is a device or program that analyzes the emotional data of meeting participants and has the function of adjusting the progress of the meeting.

[0599] "Communication technology means" refers to a device or program that has a communication function for informing each participant of the next action to take after a meeting.

[0600] "Data conversion means" refers to a device or program that has the function of converting audio data into text data in real time.

[0601] Modes for carrying out the invention

[0602] This invention is a system for facilitating smooth meeting progress and adjusting based on participants' emotions. The system is implemented through the collaboration of a server, terminals, and users, and utilizes the following hardware and software.

[0603] The device is equipped with speech recognition software that accepts user input using natural language. Specific examples of such software include Google Cloud Speech-to-Text and IBM Watson Speech to Text. This converts speech data into text data.

[0604] The server processes the input text data and automatically generates a meeting agenda using a generative AI model. This generated agenda includes the meeting's time schedule and specific agenda items.

[0605] Furthermore, the server utilizes an emotion engine to analyze participants' emotions in real time based on data from devices such as NeuroSky and Emotiv, thereby adjusting the meeting's progress. If participants are dissatisfied, the server can take action, such as changing the agenda.

[0606] During the meeting, the terminal and server work together to provide a system that monitors speech in real time. If a participant's remarks stray from the agenda, the server will notify them to return to the topic within the allotted time, thus appropriately supporting the progress of the meeting.

[0607] Finally, the server automatically generates meeting minutes based on the audio data of the meeting, including content and sentiment data, and creates a report that clearly communicates the next steps for each participant.

[0608] As a concrete example, a user sets the purpose as "Meeting about the launch of a new product" and enters the scheduled time as "2 hours." During the meeting, if the emotion engine detects that a participant is agitated, the server constructively adjusts the flow of the meeting. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0609] Example of a prompt

[0610] "Please use a generative AI model to create meeting minutes."

[0611] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0612] Step 1:

[0613] At the start of the meeting, the user enters the purpose and scheduled time into the terminal. Specifically, they provide information to the terminal by voice, such as "Meeting regarding the launch of a new product, 2 hours." The entered voice data is converted into text data by the terminal's voice recognition software. This step involves data processing, where voice data is converted into text.

[0614] Step 2:

[0615] The terminal sends the meeting's purpose and scheduled time as data to the server. Secure protocols such as HTTPS are used for transmission. This ensures that user input is safely transmitted to the server. The output data is sent to the server as a text file within the software.

[0616] Step 3:

[0617] The server processes the received text data. Using a generative AI model, it automatically generates a meeting schedule. This involves analyzing the text data and proposing appropriate time schedules and agenda items. This plan is then ready for feedback.

[0618] Step 4:

[0619] The server notifies the user of the generated schedule via the terminal. The terminal presents this information to the user through push notifications or display. In this step, the data output is provided visually through the user interface, preparing the meeting.

[0620] Step 5:

[0621] During the meeting, the terminal and server work together to use speech recognition technology to monitor speech in real time. They check whether the speech is relevant to the agenda, and the server sends notifications to the terminal as needed. The input is the audio data from the meeting, and the output is a notification resulting from the analysis of that data.

[0622] Step 6:

[0623] If a discussion deviates from the topic, the server sends a notification to the user's device prompting them to return to the subject. The device displays the message, "The current topic has deviated from the agenda." This allows users to appropriately steer the meeting back on track.

[0624] Step 7:

[0625] To perform sentiment analysis during meetings, the server uses an emotion engine. It analyzes the emotional data of meeting participants in real time and adjusts the meeting's topic and timing as needed. Input is data from emotion devices, and output is meeting adjustments based on that analysis.

[0626] Step 8:

[0627] After the meeting ends, the server automatically generates meeting minutes based on audio and sentiment data. This process involves analyzing what was said and identifying the next actions for each participant. The output is formalized as a report and provided to all participants.

[0628] Step 9:

[0629] Finally, the server sends the generated meeting minutes and reports to the participants. The data is output via email or cloud storage service, allowing participants to see what their next steps are.

[0630] (Application Example 2)

[0631] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0632] Modern meetings and conferences require not only efficient progress but also an understanding of participants' feelings and appropriate responses. However, many meetings deviate from the agenda and cause dissatisfaction among participants, resulting in inefficient waste of time. Furthermore, effectively capturing and reflecting participants' comments and feelings during a meeting is difficult, hindering the improvement of meeting effectiveness. Similar challenges exist in team meetings on factory floors, highlighting the increasing need for progress management based on real-time sentiment analysis.

[0633] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0634] In this invention, the server includes means having a module for receiving data on the purpose and scheduled time of a meeting in natural language, means having a module for automatically generating a meeting schedule, and means having a mechanism for monitoring speech content in real time using speech recognition technology. This makes it possible to centrally grasp the speech and feelings of attendees and improve the efficiency and quality of the meeting.

[0635] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format suitable for processing by machines.

[0636] A "progress plan" is a schedule or agenda that is automatically generated based on the purpose and duration of a meeting or conference.

[0637] "Speech recognition technology" is a technology that converts speech into text data in real time and is used to monitor speeches in a venue.

[0638] "Topic deviation detection" is a function that detects when a meeting's topic deviates from the designated agenda and prompts the user to return to the original topic at an appropriate time.

[0639] "Automatic meeting minutes generation" is a process that automatically creates a record of a meeting based on the content of the discussions during the meeting.

[0640] The "emotion analysis module" is a system that analyzes the emotions of attendees during a meeting in real time, and incorporates the results into the progress management.

[0641] "Adjusting the meeting's progress" is the process of flexibly changing the meeting's flow, taking into account the participants' emotional state and comments, as the meeting progresses.

[0642] A "communication mechanism" is a system that provides a means of notifying participants of their next course of action after a meeting.

[0643] This invention provides a system that efficiently and emotionally supports meetings and team meetings within factories. This system converts user input and speech from audio data into text data, analyzes its content and emotional state, and dynamically adjusts the meeting plan. The server uses speech recognition technology to capture speech in real time. It receives input of the meeting's purpose and scheduled time in natural language, generates an appropriate meeting plan, and notifies the user via their terminal. If the topic deviates from the meeting's subject, it can notify the user to return to the subject at an appropriate time.

[0644] The server includes an emotion analysis module that analyzes attendees' emotions in real time. This data is used to adjust the meeting plan. For example, if attendees show signs of agitation or dissatisfaction, the server will flexibly revise the meeting to encourage a more constructive discussion. This makes meetings more efficient while also considering the emotions of the participants.

[0645] The device has the ability to transcribe audio data into text in real time and receives the generated progress plan and sentiment analysis results through communication with the server. Because users can monitor the progress through the device, meetings can proceed smoothly.

[0646] As a concrete example, in a factory production team meeting, the team leader sets the objective as "an idea-generating meeting to improve production efficiency" and enters a one-hour time slot. If dissatisfaction among participants is detected during the meeting, the server adjusts the meeting plan based on this information to improve the quality of the exchange of ideas among participants. This system also provides emotional support to participants, making the meeting more productive.

[0647] An example of a prompt to the generative AI model is, "Analyze the participants' emotions from this conversation and suggest the optimal course of action for the meeting." In this way, the present invention aims to improve the efficiency of meeting management and the emotional satisfaction of participants.

[0648] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0649] Step 1:

[0650] The user inputs the purpose and scheduled time of the meeting into the terminal using natural language. The terminal converts the input voice data into text data using speech recognition technology. This text data is then sent to the server.

[0651] Step 2:

[0652] The server generates a meeting schedule based on the received text data. This schedule includes an agenda and timeline to facilitate efficient meeting management. The generated schedule is then fed back to the terminal.

[0653] Step 3:

[0654] As the meeting progresses, the server uses speech recognition technology to monitor the content of the discussion in real time. If a participant's comments deviate from the set agenda, the server sends a notification to the terminal, prompting the user to return to the topic.

[0655] Step 4:

[0656] The server uses an emotion analysis module to analyze the emotional state of meeting attendees in real time. This generates emotional data for the attendees, and the meeting plan is flexibly adjusted as needed. During this process, the emotional data is also analyzed using a generative AI model.

[0657] Step 5:

[0658] At the end of the meeting, the server automatically generates meeting minutes based on the participants' comments. It also identifies and lists each participant's next action. This data is then communicated to participants via their devices.

[0659] Step 6:

[0660] The generated sentiment data and progress plan are integrated to provide emotional evaluations and insights into the meeting's progress. Based on these results, users can clearly identify areas for improvement and plan future meetings.

[0661] Step 7:

[0662] For example, if a meeting is scheduled regarding the launch of a new product, the server will adjust the schedule and conduct the meeting at the optimal time when participants show excitement or interest. An example of a prompt that balances efficiency and emotional satisfaction in meeting management is, "Analyze the participants' emotions from this conversation and propose the optimal way to conduct the meeting."

[0663] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0664] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0665] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0666] [Fourth Embodiment]

[0667] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0668] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0669] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0670] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0671] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0672] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0673] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0674] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0675] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0676] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0677] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0678] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0679] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0680] This invention is a system for supporting the efficient progress of meetings, and is configured as follows: This system consists of a user, a terminal, and a server, which work together to assist in the progress of the meeting.

[0681] When a user starts a meeting, they input the meeting's purpose and scheduled time into their device using natural language. This collects the basic information necessary for the meeting. The device then converts the audio data into text and sends it to the server.

[0682] The server analyzes the received text data and automatically generates a meeting schedule based on the meeting's purpose and scheduled time. The schedule includes the main agenda items and timeline. This schedule is then used to prepare for and support the meeting's progress.

[0683] During the meeting, the terminal and server utilize speech recognition technology to monitor the meeting content in real time. As the meeting progresses, the content is recorded by transcribing speech into text, and if the topic deviates, the server sends a notification at an appropriate time to prompt the user to return to the meeting's subject. This prompt is displayed on the terminal and, if necessary, communicated to the user verbally.

[0684] Once the meeting concludes, the server automatically generates meeting minutes based on the recorded statements. These minutes include key points, conclusions, and decisions made during the meeting. They also identify and summarize the next steps for each participant. The server organizes this information and sends notifications regarding these next steps to each participant via their terminal. These notifications arrive as emails, ensuring continued follow-up even after the meeting ends.

[0685] For example, a user might define the purpose as "meeting about the annual budget" and set the scheduled time to "1 hour." The server would then set three agenda items: revenue and expenditure forecast, cost reduction proposals, and revenue improvement measures, allocating 20 minutes to each. If participants stray from the agenda during the meeting, the server would send a reminder to bring them back to the main topic, ensuring the meeting runs smoothly.

[0686] In this way, this system improves meeting productivity and helps participants efficiently share information and move on to the next step.

[0687] The following describes the processing flow.

[0688] Step 1:

[0689] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device then converts this audio data into text data.

[0690] Step 2:

[0691] The server receives text data sent from the terminal and uses natural language processing to analyze the purpose and scheduled time of the meeting. It then generates a meeting schedule.

[0692] Step 3:

[0693] The server creates a progress plan, sets the meeting schedule and agenda based on it, and sends it to the terminal.

[0694] Step 4:

[0695] During the meeting, the terminal utilizes speech recognition technology to monitor the user's speech in real time, convert the audio data into text data, and send it to the server.

[0696] Step 5:

[0697] The server analyzes the meeting discussions in real time, and if it detects that the topic has deviated from the set agenda, it sends a notification to the device prompting the participant to return to the topic at an appropriate time.

[0698] Step 6:

[0699] Once the meeting ends, the server aggregates the recorded text data of the speeches and automatically generates meeting minutes. Furthermore, it identifies and organizes the next actions for each participant.

[0700] Step 7:

[0701] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, thereby facilitating follow-up after the meeting.

[0702] (Example 1)

[0703] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0704] In meetings, participants often deviate from the agenda, and meeting minutes are frequently left ambiguous, leading to decreased meeting efficiency and delays in important decisions. Furthermore, there is a problem with insufficient follow-up, as specific next steps are not promptly communicated to participants after the meeting.

[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0706] In this invention, the server includes a device that accepts input of agenda items and scheduled times in natural language, a device that converts voice data into text, and a device that analyzes the converted data to automatically generate a meeting plan based on the agenda. This prevents deviations from the agenda during the meeting, enables efficient progress and the generation of clear meeting minutes, and also strengthens post-meeting follow-up by quickly notifying participants of their next steps.

[0707] A "device that accepts input of agenda items and scheduled times in natural language" is a device that allows users to input meeting topics and estimated times verbally or in text, and then converts and processes that information into digital signals.

[0708] A "device that converts audio data to text" is a device that analyzes voice input from a user and accurately transcribes its content into text. This device processes audio signals and converts them into text data.

[0709] A "device that analyzes converted data and automatically generates a progress plan based on the agenda" is a device that analyzes the content of a transcribed meeting and automatically creates the meeting procedure and time allocation based on pre-set objectives.

[0710] A "device that uses speech recognition technology to monitor conversations in real time" is a device that instantly transcribes conversations during a meeting into text, analyzes the content, and monitors deviations from the planned agenda.

[0711] A "device that detects deviations from the agenda and prompts participants to return to the agenda at the appropriate time" is a device that detects statements that deviate from the original agenda from conversations monitored in real time and instructs participants to return to the original agenda.

[0712] A "device that automatically generates meeting minutes and identifies each participant's next action" is a device that records meeting content, organizes important decisions and next steps, and automatically documents them.

[0713] A "communication device for notifying each participant of the minutes and subsequent actions" is a device for transmitting the generated minutes and future instructions to each meeting participant using communication methods such as email.

[0714] This invention is a system that efficiently supports the progress of meetings, in which users, terminals, and servers work together. Specific embodiments are described below.

[0715] When a user starts a meeting, they input the meeting's purpose and scheduled time into the device using natural language. The device receives this as voice input and converts the voice data into text using the Google Speech-to-Text API. This process ensures that the voice signal is accurately converted into text data.

[0716] Next, the terminal sends the converted text data to the server. The server receives this data and performs analysis using IBM Watson Natural Language Understanding. This analysis identifies the meeting's objectives and key topics, and the server automatically generates a meeting plan based on this information. The plan includes the agenda and necessary time schedules.

[0717] As the meeting progresses, the terminal and server monitor the conversation in real time using speech recognition technology. If a participant deviates from the agenda during the meeting, the server detects this and sends a reminder to the user via the terminal at an appropriate time to return to the topic. Specifically, this may involve displaying a message on the terminal such as "The current topic has deviated from the scheduled agenda," or an audio notification may also be considered.

[0718] At the end of the meeting, the server automatically generates meeting minutes based on the recorded speech data. These minutes include the meeting's objectives, key decisions, and next steps. The server uses the Microsoft Word API to format the minutes as a document and provides it to each participant.

[0719] The server then identifies each participant's next action and notifies them via email through their device. This notification, delivered using the Gmail API, includes specific instructions and follow-ups. An example of a prompt sent to a participant might be, "Please consider the next steps regarding the annual budget and prepare your proposal for the next meeting."

[0720] This system provides support to help users efficiently conduct meetings and share information. By optimizing prompt sentences using a generative AI model, even faster and more effective communication becomes possible.

[0721] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0722] Step 1:

[0723] User input of meeting information

[0724] To start a meeting, the user enters the purpose and scheduled time of the meeting into their device using natural language. The entered information is received as voice or text data. This input information becomes the basic data for the system. The user provides data such as "About next year's budget proposal, approximately one hour" using the voice input function of their smartphone.

[0725] Step 2:

[0726] Converting audio data to text

[0727] The device converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to analyze the audio signal and generate data as a string. This step converts the audio information into readable text. The output is the text "Regarding next year's budget proposal, about 1 hour."

[0728] Step 3:

[0729] Text data transmission and analysis

[0730] The terminal sends the converted text data to the server. The server analyzes the received data using IBM Watson Natural Language Understanding. This analysis extracts the meeting's objectives and important keywords, and organizes the information necessary for the meeting's progress. Text data is used as input, and the output is the analyzed information structure.

[0731] Step 4:

[0732] Automatic generation of meeting schedules

[0733] The server automatically generates a meeting schedule based on the analysis results. This takes into account the meeting's topic and time allocation, and the schedule includes information such as each agenda item and the time allocated to each item. The output includes specific procedures such as "budget review," "resource allocation," and "cost reduction measures."

[0734] Step 5:

[0735] Real-time monitoring and reminder sending

[0736] The terminal and server monitor what is said during the meeting in real time. Using speech recognition, the conversation is transcribed into text, and a reminder is sent if it deviates from the agenda. The server compares it to the scheduled items and detects when the topic has gone off track. At the appropriate time, the terminal displays a message saying, "The topic has gone off-topic. Let's return to the main subject."

[0737] Step 6:

[0738] Meeting minutes generation after the meeting

[0739] The server automatically generates meeting minutes based on the recorded conversations during the meeting. This process uses the Microsoft Word API to document the text data and organize important information for participants. The output is a meeting minutes document containing information such as "budget approval," "next meeting date and time," and "action items."

[0740] Step 7:

[0741] Notification of the next action

[0742] The server notifies each participant of the generated meeting minutes and the identified next action. The notification is sent via email through the terminal using the Gmail API. This allows participants to understand the key points of the meeting and receive specific instructions for the next steps. As output, a follow-up email is sent to each participant.

[0743] (Application Example 1)

[0744] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0745] In the face of the need for effective meeting management and operational efficiency in physical stores, traditional methods often fail to keep participants focused on the agenda, resulting in insufficient achievement of meeting objectives. Furthermore, a lack of follow-up and action checks after meetings leads to decreased productivity. To overcome these challenges, there is a need to develop a system that manages meetings efficiently and systematically, while providing appropriate support to participants.

[0746] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0747] In this invention, the server includes an information receiving means for inputting data on the purpose and scheduled time of the meeting in natural language, a planning means for automatically generating a meeting schedule, and an information providing means for providing hints and suggestions based on data in real time during the meeting. This enables participants to concentrate on the meeting's topic, effectively advance discussions, and ensure follow-up after the meeting.

[0748] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format that computer systems can understand and process.

[0749] "Information receiving means" refers to a device or method for receiving the purpose and scheduled time of a meeting, entered by the user, in digital format.

[0750] "Planning means" refers to algorithms or devices for automatically generating a meeting schedule.

[0751] "Speech recognition technology" is a technology that converts speech data into text data in real time and understands the content of what is being said.

[0752] "Monitoring methods" refer to technologies that monitor speeches during a meeting in real time and record and analyze relevant data.

[0753] "Correction measures" are technologies or devices that detect when a meeting deviates from its topic and prompt participants to return to the topic at an appropriate time.

[0754] "Action identification means" refers to a method or device for automatically identifying the next steps for each individual who participated in the meeting.

[0755] "Communication means" refers to electronic methods or devices used to notify participants of an assembly of information.

[0756] "Information provision means" refers to technologies and devices that provide hints and suggestions based on real-time data during a meeting.

[0757] The system for realizing this invention is designed to enable employees to conduct meetings more efficiently. In the in-store meeting support system, the main hardware used is a smartphone or tablet to collect audio data. The information receiving means installed in these devices can receive natural language data entered by the user. Specifically, the software components include speech recognition technology that converts audio data into text data in real time, using Python and its library, SpeechRecognition.

[0758] After receiving text data, the server performs text analysis using the natural language processing library spaCy. Next, a meeting agenda is automatically generated. During the meeting, the content of the discussion is monitored using Google Cloud's Natural Language API. In addition, hints and suggestions are provided in real time through various information channels to help participants focus on the meeting's main topic.

[0759] As an example, consider the use of this system by a store operations team when holding a new product promotion meeting. In this meeting, participants choose "New Product Promotion Plan" as the topic and set a one-hour time limit. Based on this data, the server automatically sets the agenda and time allocation, and notifies participants through corrective means if they deviate from the discussion. Furthermore, after the meeting, each participant receives a notification regarding the next steps.

[0760] Examples of prompt statements to input into a generative AI model are as follows:

[0761] "The purpose of today's meeting is to discuss the 'marketing plan for the new product,' and it will last one hour. Please provide a meeting plan based on this. Also, monitor the progress of each agenda item and send a notification if the discussion goes off track."

[0762] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0763] Step 1:

[0764] The user enters the purpose and scheduled time of the meeting into the device using natural language. The entered information is collected by the device's information receiving mechanism. The input data is converted into text using speech recognition technology and sent to the next step.

[0765] Step 2:

[0766] The terminal converts audio data into text data in real time using the SpeechRecognition library. The converted data is sent to the server via the internet. The output here is text data that includes the purpose and scheduled time of the meeting.

[0767] Step 3:

[0768] The server receives text data and performs natural language processing using the spaCy library. This analyzes the purpose of the meeting and the priority agenda items, and automatically generates a meeting plan. The generated plan includes the time allocation for each agenda item. This plan is used in the next step.

[0769] Step 4:

[0770] During the meeting, the device collects audio data again, transcribes it into text in real time, and sends it to the server. The server uses Google Cloud's Natural Language API to monitor the content of the discussion as the meeting progresses. If the topic deviates, the server automatically detects the deviation and generates a notification prompting the user to correct it.

[0771] Step 5:

[0772] At the end of the meeting, the server automatically generates meeting minutes using the collected data. These minutes will include key points from the meeting and the next steps for participants. The generated minutes will be sent to each participant via their terminal.

[0773] Step 6:

[0774] After the meeting, the server sends participants notifications about the next steps via email or other means of communication. This ensures that each participant is clearly aware of their next actions and can follow up effectively.

[0775] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0776] This invention provides a system that enables efficient and emotionally insightful meeting management, and has the following configuration: This system includes a user, a terminal, a server, and an emotion engine, which work together to support the progress of the meeting.

[0777] First, at the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device converts the entered information from voice data to text data and sends it to the server.

[0778] Next, the server analyzes the text data and automatically generates a meeting plan based on the meeting's purpose and duration. This plan includes a time schedule and agenda, enabling efficient meeting management. The meeting plan is then fed back to the user via their terminal.

[0779] During the meeting, the terminal and server use speech recognition technology to monitor the content of the discussion in real time. If the discussion deviates from the set agenda, the server sends a notification to the terminal prompting the user to return to the topic at an appropriate time. Furthermore, once the meeting ends, the server automatically generates meeting minutes based on the content of the discussion, and identifies and organizes the next actions for each participant. This ensures that follow-up is provided even after the meeting has ended.

[0780] The emotion engine, a key feature of this invention, analyzes users' emotions in real time during a meeting. Based on the emotion data obtained from the emotion engine, the server flexibly adjusts the progress of the meeting. For example, if a participant is feeling dissatisfied, the server can adjust the timing of changing topics. Furthermore, the meeting minutes and reports generated after the meeting also include the emotion data, providing emotional evaluations and insights into how the meeting proceeded.

[0781] As a concrete example, a user sets the objective as "Meeting about the launch of a new product" and enters a scheduled time of "2 hours." During the meeting, if the emotion engine detects signs of agitation in some participants, the server takes this into consideration and adjusts the flow of the meeting to help ensure a constructive discussion. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0782] The following describes the processing flow.

[0783] Step 1:

[0784] At the start of the meeting, the user enters the purpose and scheduled time into the device using natural language. The device receives the input as voice data and converts it into text data.

[0785] Step 2:

[0786] The terminal sends the converted text data to the server. The server analyzes the purpose and duration of the meeting and automatically generates a meeting schedule.

[0787] Step 3:

[0788] The server generates a progress plan and sends it to the terminal, which then displays it to the user. The progress plan includes the agenda and time schedule.

[0789] Step 4:

[0790] During the meeting, the terminal uses speech recognition technology to transcribe the user's speech into text in real time and send it to the server.

[0791] Step 5:

[0792] The server analyzes the received message data and detects when the topic deviates from the agenda. It then sends a notification to the device prompting the user to return to the topic at an appropriate time.

[0793] Step 6:

[0794] The server uses an emotion engine to analyze the emotions of participants during a meeting in real time and detects emotional data.

[0795] Step 7:

[0796] The server flexibly adjusts the flow of the meeting based on sentiment data and suggests new discussion transitions as needed.

[0797] Step 8:

[0798] Once the meeting ends, the server automatically generates meeting minutes based on the collected speech and sentiment data, and also identifies the next actions for each participant.

[0799] Step 9:

[0800] The server generates meeting minutes and sends the next steps to each participant via email through their terminal, and then follows up after the meeting.

[0801] Step 10:

[0802] The server helps participants gain emotional insights by including emotional data in the meeting report and providing it to them.

[0803] (Example 2)

[0804] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0805] Traditional meeting systems have had problems such as inefficient progress management and meetings sometimes going in an inappropriate direction because participants' feelings are not taken into consideration. Specifically, problems include wasted time due to off-topic discussions, difficulty in tracking meeting content, and the accumulation of participant dissatisfaction.

[0806] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0807] In this invention, the server includes information processing means for analyzing natural language input data received via voice and determining the purpose and scheduled time of the meeting; schedule generation means for automatically generating a meeting plan based on the input data and notifying participants; and speech recognition technology means for monitoring speech in real time during the meeting and supporting the progress of the meeting in accordance with the agenda. This enables efficient management of the meeting and flexible meeting adjustments based on the emotions of the participants.

[0808] "Information processing means" refers to a device or program that has the function of analyzing natural language input data received via voice and determining the purpose and scheduled time of a meeting.

[0809] A "schedule generation means" is a device or program that has the function of automatically generating a meeting schedule based on input data and notifying participants of that schedule.

[0810] "Speech recognition technology means" refers to a device or program that uses technology to monitor speech during a meeting in real time and support the progress of the meeting in accordance with the agenda.

[0811] "Detection and notification means" refers to a device or program that has the function of detecting when a statement deviates from the agenda and prompting a return to the topic at an appropriate time.

[0812] "Document generation means" refers to a device or program that has the function of automatically generating meeting minutes based on audio data and clearly indicating the next action plan for each participant.

[0813] A "sentiment analysis module" is a device or program that analyzes the emotional data of meeting participants and has the function of adjusting the progress of the meeting.

[0814] "Communication technology means" refers to a device or program that has a communication function for informing each participant of the next action to take after a meeting.

[0815] "Data conversion means" refers to a device or program that has the function of converting audio data into text data in real time.

[0816] Modes for carrying out the invention

[0817] This invention is a system for facilitating smooth meeting progress and adjusting based on participants' emotions. The system is implemented through the collaboration of a server, terminals, and users, and utilizes the following hardware and software.

[0818] The device is equipped with speech recognition software that accepts user input using natural language. Specific examples of such software include Google Cloud Speech-to-Text and IBM Watson Speech to Text. This converts speech data into text data.

[0819] The server processes the input text data and automatically generates a meeting agenda using a generative AI model. This generated agenda includes the meeting's time schedule and specific agenda items.

[0820] Furthermore, the server utilizes an emotion engine to analyze participants' emotions in real time based on data from devices such as NeuroSky and Emotiv, thereby adjusting the meeting's progress. If participants are dissatisfied, the server can take action, such as changing the agenda.

[0821] During the meeting, the terminal and server work together to provide a system that monitors speech in real time. If a participant's remarks stray from the agenda, the server will notify them to return to the topic within the allotted time, thus appropriately supporting the progress of the meeting.

[0822] Finally, the server automatically generates meeting minutes based on the audio data of the meeting, including content and sentiment data, and creates a report that clearly communicates the next steps for each participant.

[0823] As a concrete example, a user sets the purpose as "Meeting about the launch of a new product" and enters the scheduled time as "2 hours." During the meeting, if the emotion engine detects that a participant is agitated, the server constructively adjusts the flow of the meeting. In this way, the system aims to support participants emotionally as well, providing a more effective meeting experience.

[0824] Example of a prompt

[0825] "Please use a generative AI model to create meeting minutes."

[0826] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0827] Step 1:

[0828] At the start of the meeting, the user enters the purpose and scheduled time into the terminal. Specifically, they provide information to the terminal by voice, such as "Meeting regarding the launch of a new product, 2 hours." The entered voice data is converted into text data by the terminal's voice recognition software. This step involves data processing, where voice data is converted into text.

[0829] Step 2:

[0830] The terminal sends the meeting's purpose and scheduled time as data to the server. Secure protocols such as HTTPS are used for transmission. This ensures that user input is safely transmitted to the server. The output data is sent to the server as a text file within the software.

[0831] Step 3:

[0832] The server processes the received text data. Using a generative AI model, it automatically generates a meeting schedule. This involves analyzing the text data and proposing appropriate time schedules and agenda items. This plan is then ready for feedback.

[0833] Step 4:

[0834] The server notifies the user of the generated schedule via the terminal. The terminal presents this information to the user through push notifications or display. In this step, the data output is provided visually through the user interface, preparing the meeting.

[0835] Step 5:

[0836] During the meeting, the terminal and server work together to use speech recognition technology to monitor speech in real time. They check whether the speech is relevant to the agenda, and the server sends notifications to the terminal as needed. The input is the audio data from the meeting, and the output is a notification resulting from the analysis of that data.

[0837] Step 6:

[0838] If a discussion deviates from the topic, the server sends a notification to the user's device prompting them to return to the subject. The device displays the message, "The current topic has deviated from the agenda." This allows users to appropriately steer the meeting back on track.

[0839] Step 7:

[0840] To perform sentiment analysis during meetings, the server uses an emotion engine. It analyzes the emotional data of meeting participants in real time and adjusts the meeting's topic and timing as needed. Input is data from emotion devices, and output is meeting adjustments based on that analysis.

[0841] Step 8:

[0842] After the meeting ends, the server automatically generates meeting minutes based on audio and sentiment data. This process involves analyzing what was said and identifying the next actions for each participant. The output is formalized as a report and provided to all participants.

[0843] Step 9:

[0844] Finally, the server sends the generated meeting minutes and reports to the participants. The data is output via email or cloud storage service, allowing participants to see what their next steps are.

[0845] (Application Example 2)

[0846] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0847] Modern meetings and conferences require not only efficient progress but also an understanding of participants' feelings and appropriate responses. However, many meetings deviate from the agenda and cause dissatisfaction among participants, resulting in inefficient waste of time. Furthermore, effectively capturing and reflecting participants' comments and feelings during a meeting is difficult, hindering the improvement of meeting effectiveness. Similar challenges exist in team meetings on factory floors, highlighting the increasing need for progress management based on real-time sentiment analysis.

[0848] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0849] In this invention, the server includes means having a module for receiving data on the purpose and scheduled time of a meeting in natural language, means having a module for automatically generating a meeting schedule, and means having a mechanism for monitoring speech content in real time using speech recognition technology. This makes it possible to centrally grasp the speech and feelings of attendees and improve the efficiency and quality of the meeting.

[0850] "Natural language" refers to the language that humans use on a daily basis, which can be converted into a format suitable for processing by machines.

[0851] A "progress plan" is a schedule or agenda that is automatically generated based on the purpose and duration of a meeting or conference.

[0852] "Speech recognition technology" is a technology that converts speech into text data in real time and is used to monitor speeches in a venue.

[0853] "Topic deviation detection" is a function that detects when a meeting's topic deviates from the designated agenda and prompts the user to return to the original topic at an appropriate time.

[0854] "Automatic meeting minutes generation" is a process that automatically creates a record of a meeting based on the content of the discussions during the meeting.

[0855] The "emotion analysis module" is a system that analyzes the emotions of attendees during a meeting in real time, and incorporates the results into the progress management.

[0856] "Adjusting the meeting's progress" is the process of flexibly changing the meeting's flow, taking into account the participants' emotional state and comments, as the meeting progresses.

[0857] A "communication mechanism" is a system that provides a means of notifying participants of their next course of action after a meeting.

[0858] This invention provides a system that efficiently and emotionally supports meetings and team meetings within factories. This system converts user input and speech from audio data into text data, analyzes its content and emotional state, and dynamically adjusts the meeting plan. The server uses speech recognition technology to capture speech in real time. It receives input of the meeting's purpose and scheduled time in natural language, generates an appropriate meeting plan, and notifies the user via their terminal. If the topic deviates from the meeting's subject, it can notify the user to return to the subject at an appropriate time.

[0859] The server includes an emotion analysis module that analyzes attendees' emotions in real time. This data is used to adjust the meeting plan. For example, if attendees show signs of agitation or dissatisfaction, the server will flexibly revise the meeting to encourage a more constructive discussion. This makes meetings more efficient while also considering the emotions of the participants.

[0860] The device has the ability to transcribe audio data into text in real time and receives the generated progress plan and sentiment analysis results through communication with the server. Because users can monitor the progress through the device, meetings can proceed smoothly.

[0861] As a concrete example, in a factory production team meeting, the team leader sets the objective as "an idea-generating meeting to improve production efficiency" and enters a one-hour time slot. If dissatisfaction among participants is detected during the meeting, the server adjusts the meeting plan based on this information to improve the quality of the exchange of ideas among participants. This system also provides emotional support to participants, making the meeting more productive.

[0862] An example of a prompt to the generative AI model is, "Analyze the participants' emotions from this conversation and suggest the optimal course of action for the meeting." In this way, the present invention aims to improve the efficiency of meeting management and the emotional satisfaction of participants.

[0863] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0864] Step 1:

[0865] The user inputs the purpose and scheduled time of the meeting into the terminal using natural language. The terminal converts the input voice data into text data using speech recognition technology. This text data is then sent to the server.

[0866] Step 2:

[0867] The server generates a meeting schedule based on the received text data. This schedule includes an agenda and timeline to facilitate efficient meeting management. The generated schedule is then fed back to the terminal.

[0868] Step 3:

[0869] As the meeting progresses, the server uses speech recognition technology to monitor the content of the discussion in real time. If a participant's comments deviate from the set agenda, the server sends a notification to the terminal, prompting the user to return to the topic.

[0870] Step 4:

[0871] The server uses an emotion analysis module to analyze the emotional state of meeting attendees in real time. This generates emotional data for the attendees, and the meeting plan is flexibly adjusted as needed. During this process, the emotional data is also analyzed using a generative AI model.

[0872] Step 5:

[0873] At the end of the meeting, the server automatically generates meeting minutes based on the participants' comments. It also identifies and lists each participant's next action. This data is then communicated to participants via their devices.

[0874] Step 6:

[0875] The generated sentiment data and progress plan are integrated to provide emotional evaluations and insights into the meeting's progress. Based on these results, users can clearly identify areas for improvement and plan future meetings.

[0876] Step 7:

[0877] For example, if a meeting is scheduled regarding the launch of a new product, the server will adjust the schedule and conduct the meeting at the optimal time when participants show excitement or interest. An example of a prompt that balances efficiency and emotional satisfaction in meeting management is, "Analyze the participants' emotions from this conversation and propose the optimal way to conduct the meeting."

[0878] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0879] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0880] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0881] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0882] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0883] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0884] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0885] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0886] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0887] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0888] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0889] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0890] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0891] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0892] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0893] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0894] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0895] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0896] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0897] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0898] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0899] The following is further disclosed regarding the embodiments described above.

[0900] (Claim 1)

[0901] A means of accepting input of the purpose and scheduled time of a meeting in natural language,

[0902] A means of automatically generating a meeting schedule,

[0903] A method for monitoring speech in real time during a meeting using speech recognition technology,

[0904] A means to detect when a topic deviates from the subject of the meeting and prompt the participant to return to the subject at an appropriate time,

[0905] A means to automatically generate meeting minutes and identify the next action for each participant,

[0906] A system that includes a means of communication to notify each participant of their next action after a meeting.

[0907] (Claim 2)

[0908] The system according to claim 1, further comprising means for automatically generating meeting report materials.

[0909] (Claim 3)

[0910] The system according to claim 1, further comprising means for transcribing meeting audio data into text in real time.

[0911] "Example 1"

[0912] (Claim 1)

[0913] A device that accepts input of agenda items and scheduled times in natural language,

[0914] A device that converts audio data into text,

[0915] A device that analyzes converted data and automatically generates a progress plan based on the agenda,

[0916] A device that uses speech recognition technology to monitor conversations in real time,

[0917] A device that detects deviations from the agenda and prompts participants to return to the agenda at the appropriate time,

[0918] A device that automatically generates meeting minutes and identifies the next actions of each participant,

[0919] A system including meeting minutes and a communication device to notify each participant of the next course of action.

[0920] (Claim 2)

[0921] The system according to claim 1, further comprising a function for automatically generating meeting documents.

[0922] (Claim 3)

[0923] The system according to claim 1, further comprising a function to transcribe conversational audio data into text in real time.

[0924] "Application Example 1"

[0925] (Claim 1)

[0926] An information receiving device that inputs data on the purpose and scheduled time of a meeting using natural language,

[0927] A planning tool that automatically generates a meeting schedule,

[0928] A monitoring method that uses speech recognition technology to monitor speech during a meeting in real time,

[0929] A corrective mechanism to detect when the topic deviates from the subject of the meeting and prompts participants to return to the subject at an appropriate time,

[0930] A means of automatically generating meeting records and identifying the next steps for each participant,

[0931] A means of communication to notify each participant of the next steps after the meeting,

[0932] A means of providing information that offers hints and suggestions based on data in real time during the meeting,

[0933] A system that includes this.

[0934] (Claim 2)

[0935] The system according to claim 1, further comprising a means for automatically generating meeting report documents.

[0936] (Claim 3)

[0937] The system according to claim 1, further comprising a data conversion means for converting audio data of a meeting into text data in real time.

[0938] "Example 2 of combining an emotion engine"

[0939] (Claim 1)

[0940] An information processing means for analyzing natural language input data received via voice to determine the purpose and scheduled time of a meeting,

[0941] A schedule generation means that automatically generates a meeting schedule based on the aforementioned input data and notifies the participants,

[0942] A speech recognition technology means for monitoring speech in real time during a meeting and supporting the progress of the meeting in accordance with the agenda,

[0943] A means of detection and notification to detect when a statement deviates from the topic and to prompt a return to the subject,

[0944] A document generation method that automatically generates meeting minutes based on audio data and clearly outlines the next action plan for each participant,

[0945] An emotion analysis module analyzes the emotional data of meeting participants and provides control means to adjust the progress of the meeting.

[0946] A system that includes communication technologies to inform each participant of their next action after a meeting.

[0947] (Claim 2)

[0948] The system according to claim 1, comprising means for generating meeting reports and providing participants with emotional insights by including emotional data.

[0949] (Claim 3)

[0950] The system according to claim 1, comprising data conversion means for converting audio data into text data in real time.

[0951] "Application example 2 when combining with an emotional engine"

[0952] (Claim 1)

[0953] A means comprising a module that accepts data on the purpose and scheduled time of a meeting in natural language,

[0954] A means having a module that automatically generates a meeting schedule,

[0955] A means having a mechanism for monitoring the content of speech in real time using speech recognition technology,

[0956] A means having a module that recognizes when a conversation deviates from a specified topic and sends a message to return to the topic at an appropriate time,

[0957] A means comprising a device that automatically generates meeting minutes and determines the next action for each participant,

[0958] A means of providing a communication mechanism to notify participants of their next steps after the meeting,

[0959] A means comprising a module that uses an emotion analysis module to analyze the feelings of attendees in real time and enables flexible adjustment of the proceedings based on the emotional data,

[0960] A means having a module that provides a dataset for evaluating the progress of a meeting based on the sentiment data of attendees,

[0961] A means having a module that integrates generated sentiment data and a progress plan to optimize the progress of a meeting,

[0962] A system that includes this.

[0963] (Claim 2)

[0964] The system according to claim 1, further comprising a module for automatically generating meeting report materials.

[0965] (Claim 3)

[0966] The system according to claim 1, comprising technology for converting audio data of a meeting into text data in real time. [Explanation of Symbols]

[0967] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of accepting input of the purpose and scheduled time of a meeting in natural language, A means of automatically generating a meeting schedule, A method for monitoring speech in real time during a meeting using speech recognition technology, A means to detect when a topic deviates from the subject of the meeting and prompt the participant to return to the subject at an appropriate time, A means to automatically generate meeting minutes and identify the next action for each participant, A system that includes a means of communication to notify each participant of their next action after a meeting.

2. The system according to claim 1, further comprising means for automatically generating meeting report materials.

3. The system according to claim 1, further comprising means for converting meeting audio data into text in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A