System

A system automates the recording, text conversion, summarization, and analysis of one-on-one meetings to generate actionable items and best practices, addressing the inefficiencies in manual management and enhancing productivity.

JP2026014237APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024115234
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Efficiently recording and analyzing the content of one-on-one meetings in large organizations is difficult, leading to time-consuming manual management and hindering employee growth and work efficiency due to the inability to reliably link issues and tasks to concrete actions.

Method used

A system that records meeting content, converts it into text, summarizes, analyzes, and generates action items using speech recognition and natural language processing, and proposes know-how and best practices.

Benefits of technology

Enables efficient management of one-on-one meeting content, linking it to specific actions, reducing manual effort and enhancing employee productivity by automating the process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014237000001_ABST
    Figure 2026014237000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for initiating a recording of a meeting; means for transmitting recorded data to a server; means for the server to perform speech recognition on the recorded data and convert the recorded data to text; means for the server to summarize the text data; means for the server to analyze the summarized data; and means for the server to generate and suggest an action item.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern companies, one-on-one meetings play an important role in confirming employees' work progress and challenges, and proposing support measures. However, efficiently recording and analyzing the content of these meetings and linking them to concrete action is difficult. Large organizations, in particular, hold a huge number of meetings, so manually managing all of their content is time-consuming and labor-intensive. Furthermore, if the issues and tasks discussed in meetings cannot be reliably linked to action, employee growth and work efficiency are hindered. [Means for solving the problem]

[0005] The present invention relates to a system that records the content of one-on-one meetings, sends the data to a server, and converts it into text, summarizes, and analyzes it. This system includes a means for starting the recording of the meeting, a means for sending the recorded data to the server, a means for the server to recognize the recorded data and convert it into text, a means for the server to summarize the text data, a means for the server to analyze the summarized data, and a means for the server to generate and propose action items. It also includes a means for saving the summarized data in a database and a means for extracting know-how and best practices from a knowledge base and generating a proposal list. This provides a system that can efficiently manage the content of one-on-one meetings and lead to concrete actions.

[0006] "Method for starting a meeting recording" refers to the functionality of the device or software that a user uses to start recording a one-on-one meeting.

[0007] "Means for transmitting recorded data to a server" refers to a mechanism including network communications and protocols for uploading recorded audio data to a server in real time or in batches.

[0008] "Means by which the server recognizes the recorded data and converts it into text" refers to the process by which the server uses voice recognition technology to convert the recorded data into text data.

[0009] "Means for the server to summarize text data" refers to the function by which the server uses natural language processing (NLP) algorithms to extract important information from long text data and summarize it in a concise form.

[0010] The "means by which the server analyzes the summarized data" refers to the process by which the server analyzes the summarized data, extracts issues and themes, and performs statistical analysis.

[0011] "Means for the server to generate and propose action items" refers to a function in which the server generates executable tasks and action plans based on the analysis results and notifies the user of them.

[0012] "Means for storing the summarized data in a database" refers to the functionality of the data storage system or database management system used to manage and store the summarized text data.

[0013] "Means for extracting know-how and best practices from the knowledge base and generating a list of proposals" refers to the process or algorithm by which the server searches and extracts relevant information from the accumulated knowledge database and proposes it to the user as reference material.

[0014] "Database" means a data storage system that organizes and manages collected data and makes it easily accessible.

[0015] "Natural Language Processing (NLP) algorithms" refers to a set of algorithms that include techniques and methods for computers to understand, analyze, and generate human language. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, mainly including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices. The following describes in detail the embodiments of the present invention.

[0038] 1. System Configuration

[0039] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[0040] 2. Description of main functions

[0041] Voice recording and data transmission

[0042] User starts recording:

[0043] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[0044] The device sends the audio data to the server:

[0045] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[0046] Converting audio data to text

[0047] Server receives audio data:

[0048] The server receives the voice data sent from the terminal and stores it securely in data storage.

[0049] Audio to text conversion:

[0050] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[0051] Text summary generation

[0052] The server parses the text data:

[0053] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[0054] Summary generation:

[0055] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[0056] Data analysis and action item generation

[0057] Summary data analysis:

[0058] The server aggregates multiple summaries to identify common issues and trends.

[0059] Generate action items:

[0060] The server uses past meeting data and success stories to generate action items for the issues, such as "reviewing tasks" and "adjusting additional resources."

[0061] Proposal of know-how and best practices

[0062] Search our knowledge base:

[0063] The server searches and extracts relevant know-how and best practices from the knowledge base.

[0064] Generate a list of suggestions:

[0065] The server organizes the search results and creates a specific list of suggestions for the user.

[0066] Notification and confirmation

[0067] Send notifications:

[0068] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[0069] User sees action item:

[0070] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[0071] Specific examples

[0072] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[0073] 1. Start recording:

[0074] Person A presses the recording button on the device to start the meeting.

[0075] 2. Sending audio data to the server:

[0076] During the meeting, the device transmits audio data to the server in real time.

[0077] 3. Text:

[0078] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[0079] 4. Summary generation:

[0080] Using an NLP algorithm, summarize it as "Project Y problem."

[0081] 5. Data analysis and action generation:

[0082] The server analyzes similar past data and generates action items such as "set up a meeting to identify problems."

[0083] 6. Know-how proposal:

[0084] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[0085] 7. User confirms:

[0086] Persons A and B check this information on their devices and create an action plan.

[0087] In this way, the system can efficiently manage the content of one-on-one meetings and link them to concrete actions.

[0088] The processing flow will be explained below.

[0089] Step 1:

[0090] User starts recording:

[0091] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[0092] Step 2:

[0093] The device captures audio data:

[0094] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[0095] Step 3:

[0096] The device sends the audio data to the server:

[0097] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[0098] Step 4:

[0099] Server receives audio data:

[0100] The server receives the voice data sent from the terminal and stores the data in storage.

[0101] Step 5:

[0102] The server converts the audio data to text:

[0103] The server uses a speech recognition API to convert the voice data into text data in real time or in batches, and the converted text data is stored in a database.

[0104] Step 6:

[0105] The server summarizes the text data:

[0106] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[0107] Step 7:

[0108] The server stores the summary data:

[0109] The summarized text data is stored in a database and used for subsequent analysis and statistical processing.

[0110] Step 8:

[0111] The server analyzes the summary data:

[0112] The server aggregates multiple summaries of data and uses an analytics engine to identify issues and trends.

[0113] Step 9:

[0114] Server generates action item:

[0115] Based on the analysis results, the server generates specific action items, referencing past history and best practices.

[0116] Step 10:

[0117] Server searches knowledge base:

[0118] The server searches databases and knowledge bases to extract relevant know-how and best practices.

[0119] Step 11:

[0120] Server generates suggestion list:

[0121] Based on the extracted information, a list of specific action suggestions and best practices is generated to provide to the user.

[0122] Step 12:

[0123] Server sends notification:

[0124] Generated action items and suggestion lists are notified to the user's device, including via pop-up notifications and email notifications.

[0125] Step 13:

[0126] User sees action item:

[0127] The user uses the device to check the notification and view the generated action items and suggested know-how.

[0128] Step 14:

[0129] User performs an action:

[0130] The user then executes specific tasks based on the action items they have confirmed, such as setting up a project meeting or arranging for additional resources.

[0131] Example 1

[0132] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0133] Conventional conference management systems have had difficulty efficiently managing conference content and linking it to specific actions. Furthermore, converting conference recordings into text, summarizing them, and generating action items require a lot of manual work, which takes time and effort. Therefore, there is a demand for a system that can achieve efficient conference management and link it to specific actions.

[0134] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0135] In this invention, the server includes a means for [starting recording of the meeting], a means for [sending the recorded data to the server], a means for [the server converting the recorded data into text data using speech recognition technology], a means for [the server summarizing the text data using a natural language processing algorithm], a means for [the server analyzing the summarized data using a data analysis algorithm], and a means for [the server generating and proposing action items based on past success stories]. This makes it possible to convert the contents of the meeting into text and summarize it in real time, and automatically generate and propose specific action items.

[0136] The "means for starting recording of a conference" refers to the device or software that a user uses to record a conference, specifically a recording button.

[0137] "Means for transmitting recorded data to a server" refers to a function or device that transmits voice data recorded by a user to a server in real time or in batch mode.

[0138] "Means by which the server converts recorded data into text data using speech recognition technology" refers to the functions and processes by which the server converts audio data into text data using a speech recognition API or algorithm.

[0139] "Means by which the server summarizes text data using a natural language processing algorithm" refers to the function or algorithm by which the server uses natural language processing technology to extract important information from text data and summarize it concisely.

[0140] "Means for the server to analyze the summarized data using data analysis algorithms" means the data analysis algorithms used by the server to analyze the summarized text data and identify recurring issues or trends.

[0141] "Means for the server to generate and suggest action items based on past success stories" refers to the process or function in which the server refers to past meeting data and best practices, generates specific action items for specific issues, and suggests them to the user.

[0142] "Means for storing abstract data in a database" refers to the functions and processes for securely storing abstract data generated by the server in a database.

[0143] "Means for the server to extract relevant knowledge and best practices from the knowledge base and generate a list of suggestions" refers to the process or function by which the server searches an existing knowledge base, extracts useful information and best practices that meet the user's needs, and provides them as a list.

[0144] This invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices.

[0145] System Configuration

[0146] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[0147] Hardware and Software

[0148] 1. Device: The smartphone or computer used by the user.

[0149] 2. Server: Cloud server that processes data.

[0150] 3. Database: A database for storing meeting data, analysis results, etc.

[0151] 4. Speech recognition APIs: Google Speech-to-Text, Amazon Transcribe, etc.

[0152] 5. Natural Language Processing (NLP) algorithms: Open source libraries such as SpaCy and NLTK.

[0153] System functions and examples

[0154] Voice recording and data transmission

[0155] User starts recording:

[0156] The user opens the dedicated app on their device and presses the record button. By pressing the record button, the app activates the device's microphone and begins capturing audio. For example, employee A uses the smartphone app to record a one-on-one meeting.

[0157] The device sends the audio data to the server:

[0158] The device uploads the captured audio data to the server in regular segments in real time. For example, the audio data is divided into segments every 10 seconds and sent to the server sequentially.

[0159] Converting audio data to text

[0160] Server receives audio data:

[0161] The server receives the voice data sent from the device and stores it securely in data storage, for example, as an audio file in cloud storage.

[0162] The server converts the audio data to text:

[0163] The server converts the received voice data into text using a speech recognition API (e.g., Google Speech-to-Text). For example, a speech saying "Project X's deadline is behind schedule" is obtained as text data.

[0164] Text summary generation

[0165] The server parses the text data:

[0166] The server applies NLP algorithms (e.g., SpaCy) to extract important keywords and context from the text data.

[0167] Server generates summary:

[0168] The server summarizes the original text based on the extracted information. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[0169] Data analysis

[0170] The server aggregates the summary data:

[0171] The server aggregates multiple summaries generated from past meetings and identifies common issues and trends. For example, "late delivery" emerges as a common issue across multiple projects.

[0172] Generate action items

[0173] Server generates action item:

[0174] The server generates specific action items for issues based on past success stories and knowledge bases. For example, it suggests "setting up an emergency meeting" or "arranging additional resources" as a way to address delivery delays.

[0175] Proposal of know-how and best practices

[0176] Server searches for know-how:

[0177] The server searches the knowledge base for relevant know-how and best practices, for example, "project management best practices."

[0178] Server generates suggestion list:

[0179] The server creates a list of suggestions based on the search results and provides it to the user. Specifically, it displays a list of past success stories and procedures based on those stories.

[0180] Notification and confirmation

[0181] Server sends notification:

[0182] The server notifies the user of the generated action items and suggestion list via a pop-up notification on their smartphone or email.

[0183] User sees action item:

[0184] Users check notifications on their devices and view suggested action items and know-how. They then plan specific actions and put them into action. For example, Person A and Person B discuss specific measures based on the information displayed on their devices and create an action plan.

[0185] Prompt Sentence Examples

[0186] "Please summarize the content of the 1-on-1 meeting and propose specific action items and know-how."

[0187] "Convert the recorded audio data into text and extract the key points."

[0188] This system allows you to efficiently manage the content of one-on-one meetings and link them to concrete actions.

[0189] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0190] Step 1:

[0191] User starts recording

[0192] The user opens the dedicated app on their device and presses the record button. The input is the user clicking the record button, and the device starts capturing audio. Specifically, the smartphone's microphone is activated and audio data is captured in real time. The output is the audio data being recorded.

[0193] Step 2:

[0194] The device sends the voice data to the server

[0195] The device uploads captured audio data to the server in regular segments in real time. The input is the audio data being recorded, divided into segments, such as 10 seconds. Specifically, the device sends the audio data segments to the server using HTTP requests. The output is the audio data stored on the server.

[0196] Step 3:

[0197] The server receives the audio data

[0198] The server receives the voice data sent from the terminal and safely stores it in the data storage. The input is the voice data sent from the terminal, and the output is the voice file stored in the data storage. Specifically, the server executes a process to store the voice file in a database.

[0199] Step 4:

[0200] The server converts the voice data into text

[0201] The server uses a speech recognition API (e.g., Google Speech-to-Text) to convert the voice data into text. The input is a saved audio file, and when the server calls the API, the API analyzes the voice data and returns it as text data. The output is text data. For example, the text generated might say, "The delivery date for Project X is behind schedule."

[0202] Step 5:

[0203] The server analyzes the text data

[0204] The server applies an NLP algorithm (e.g., SpaCy) to extract important keywords and context from the text data. The input is text data, which the server analyzes by running the NLP algorithm. The output is important keywords and context data. For example, the important keyword "delayed delivery" is extracted.

[0205] Step 6:

[0206] Server generates summary

[0207] The server summarizes the original text based on the extracted information. The input is important keywords and contextual data, and the server generates a summary by running a summary generation algorithm. The output is summarized text data. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[0208] Step 7:

[0209] The server aggregates the summary data

[0210] The server aggregates multiple summaries generated from past meetings and identifies frequently occurring issues and trends. The input is multiple summarized text data, which the server analyzes by applying a data aggregation algorithm. The output is the identification of issues and trends. For example, "delayed delivery" is identified as a common issue across multiple projects.

[0211] Step 8:

[0212] Server generates action items

[0213] The server generates specific action items for issues based on past success stories and a knowledge base. The input is the results of identifying issues and trends, and the server generates specific action items by running an action item generation algorithm. The output is a specific action item. For example, it may suggest measures such as "setting up an emergency meeting" or "adjusting additional resources" to address delivery delays.

[0214] Step 9:

[0215] The server searches for know-how

[0216] The server searches the knowledge base to find relevant know-how and best practices. The input is summary data and action items, and the server runs a knowledge base search algorithm to extract relevant information. The output is know-how and best practice information.

[0217] Step 10:

[0218] The server generates a list of suggestions

[0219] The server creates a proposal list based on the search results and provides it to the user. The input is know-how and best practice information, and the server generates the list by running a proposal list generation algorithm. The output is the proposal list provided to the user.

[0220] Step 11:

[0221] The server sends a notification

[0222] The server notifies the user's terminal of the generated action items and suggestion list. The input is the action items and suggestion list, and the server executes the notification process to send the notification to the user. The output is the notification displayed on the user's terminal.

[0223] Step 12:

[0224] User confirms action item

[0225] The user receives a notification on their device and checks the proposed action items and know-how. The input is the notification displayed on the user's device, and the specific action is confirmed by the user checking the operation. The output is the creation of an action plan. For example, Person A and Person B discuss specific measures and create an action plan.

[0226] (Application example 1)

[0227] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0228] There is a need to efficiently manage the content of one-on-one meetings held in the field and link it to specific action items and best practices. However, the current situation is such that the content of meetings is not fully utilized, resulting in a waste of time and resources. In addition, it is a heavy burden for administrators to continue to manage the content of each meeting individually. For this reason, there is a need for a more efficient and automated system.

[0229] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0230] In this invention, the server includes means for starting recording of the meeting, means for transmitting the recorded data to the server, means for the server to recognize the voice of the recorded data and convert it into text, means for the server to summarize the text data, means for the server to analyze the summarized data, means for the server to generate and propose action items, means for automatically generating a proposal list using a generative AI model, and means for notifying the action items and proposal list. This makes it possible to efficiently manage the content of one-on-one meetings and link them to specific actions.

[0231] "Means for starting recording of a meeting" refers to a device or function that allows a user to record the audio of a meeting by pressing a button or other operation.

[0232] "Means for transmitting recorded data to a server" refers to communication functions or programs for uploading recorded audio data to a server in real time or by batch processing.

[0233] "Means for the server to recognize the recorded data and convert it into text" refers to the process of analyzing the voice data on the server and converting it into text data using voice recognition technology.

[0234] "Means for the server to summarize text data" refers to the function of the server using natural language processing (NLP) algorithms, etc. to extract important information from text data and summarize it concisely.

[0235] "Means for the server to analyze the summary data" refers to a data analysis function that analyzes the summary data generated by the server and identifies significant keywords and patterns.

[0236] "Means for the server to generate and suggest action items" refers to the function by which the server automatically generates specific actions and suggestions that the user should take based on the analysis results.

[0237] "Means for automatically generating a suggestion list using a generative AI model" refers to an algorithm or program for automatically creating a suggestion list for a user using a generative AI model.

[0238] "Means for notifying action items and suggestion lists" refers to the system's functionality for notifying the user's terminal of generated action items and suggestion lists.

[0239] The present invention provides a system for efficiently managing the content of one-on-one meetings and linking them to specific actions. The following describes in detail the embodiments of the present invention.

[0240] System Configuration

[0241] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[0242] Voice recording and data transmission

[0243] Yu

[0244] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0245] Step 1:

[0246] When a user starts a one-on-one meeting, they press the recording button on their device, and the device begins capturing audio.

[0247] Input: User presses record button.

[0248] Output: Recorded audio data.

[0249] Specific operation: Captures audio data using the device's microphone function.

[0250] Step 2:

[0251] The device transmits the recorded audio data to the server in real time.

[0252] Input: Pre-recorded audio data.

[0253] Output: The audio data sent to the server.

[0254] Specific operation: Audio data is divided and uploaded to the server via data communication.

[0255] Step 3:

[0256] The server receives the recording and stores it securely.

[0257] Input: Audio data sent from the device.

[0258] Output: Audio data stored on the server.

[0259] What it does: Receives audio data and stores it securely in a database or storage.

[0260] Step 4:

[0261] The server converts the recorded data into text data using voice recognition technology.

[0262] Input: Recorded audio data.

[0263] Output: Text data.

[0264] Specific operation: Uses a speech recognition API to convert voice data into text.

[0265] Step 5:

[0266] The server applies natural language processing (NLP) algorithms to summarize the text data.

[0267] Input: Text data.

[0268] Output: Summarized text data.

[0269] Specific operation: Using NLP algorithms, important keywords and context are extracted from text data and a summary is generated.

[0270] Step 6:

[0271] The server analyzes the summary data and generates action items based thereon.

[0272] Input: Summarized text data.

[0273] Output: Action items.

[0274] Specific Behavior: Analyzes summary data, identifies issues and trends, and generates corresponding action items.

[0275] Step 7:

[0276] The server uses a generative AI model to automatically generate a list of suggestions including relevant know-how and best practices.

[0277] Inputs: Action items and knowledge base data.

[0278] Output: A list of suggestions.

[0279] What it does: It applies generative AI models to extract useful know-how and best practices from a knowledge base and generate a list.

[0280] Step 8:

[0281] The server notifies the user's terminal of the action items and the suggestion list.

[0282] Input: Action items and suggestion lists.

[0283] Output: Notification displayed on the user's device.

[0284] Specific behavior: Using the notification system, the generated action items and suggestion list are sent to the user's device via push notification or email.

[0285] For example, if a factory manager discusses a problem with a new production line during a one-on-one meeting with a field staff member, the server will summarize the content and suggest specific action items such as setting up a trouble-shooting team meeting. In this way, the system can efficiently manage the content of meetings and lead to actual actions.

[0286] For example, an example of a prompt for a generative AI model is: "Please summarize the following text: There is a problem with the new production line. Specifically, there is a problem with the machines frequently stopping. The cause has not yet been identified, but manual intervention may be necessary."

[0287] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0288] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[0289] 1. System Configuration

[0290] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[0291] 2. Description of main functions

[0292] Voice recording and data transmission

[0293] User starts recording:

[0294] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[0295] The device sends the audio data to the server:

[0296] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[0297] Converting audio data to text

[0298] Server receives audio data:

[0299] The server receives the voice data sent from the terminal and stores it securely in data storage.

[0300] Audio to text conversion:

[0301] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[0302] Text summary generation

[0303] The server parses the text data:

[0304] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[0305] Summary generation:

[0306] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[0307] emotion recognition

[0308] The server uses an emotion engine to recognize the user's emotions:

[0309] The server analyzes the user's emotions from the voice and text data, including algorithms that identify emotions from voice tone and text content.

[0310] Data storage and analysis

[0311] Summary and sentiment data storage:

[0312] The server stores the summarized text data and the sentiment analysis results in a database.

[0313] Data Analysis:

[0314] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[0315] Action item generation and suggestions

[0316] Server generates action item:

[0317] The server generates specific action items based on the analysis results and sentiment data, referring to past history and best practices.

[0318] Search our knowledge base:

[0319] The server searches and extracts relevant know-how and best practices from the knowledge base.

[0320] Generate a list of suggestions:

[0321] The server generates specific action suggestions and a list of best practices for the user based on the search results.

[0322] Notification and confirmation

[0323] Send notifications:

[0324] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[0325] User sees action item:

[0326] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[0327] Specific examples

[0328] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[0329] 1. Start recording:

[0330] Person A presses the recording button on the device to start the meeting.

[0331] 2. Sending audio data to the server:

[0332] During the meeting, the device transmits audio data to the server in real time.

[0333] 3. Text:

[0334] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[0335] 4. Summary generation:

[0336] Using an NLP algorithm, summarize it as "Project Y problem."

[0337] 5. Emotion recognition:

[0338] The emotion engine recognizes from the tone of the voice and the content of the text that Person A is feeling stressed about the problem.

[0339] 6. Data Retention:

[0340] The server stores the text data and emotion data in a database.

[0341] 7. Data analysis and action generation:

[0342] The server analyzes similar past data and sentiment data and generates action items such as "set up a meeting to identify problems."

[0343] 8. Know-how proposal:

[0344] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[0345] 9. User confirms:

[0346] Persons A and B check this information on their devices and create an action plan.

[0347] In this way, by combining an emotion recognition engine, it is possible to provide a management system that deepens the content of one-on-one meetings and leads to concrete actions.

[0348] The processing flow will be explained below.

[0349] Step 1:

[0350] User starts recording:

[0351] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[0352] Step 2:

[0353] The device captures audio data:

[0354] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[0355] Step 3:

[0356] The device sends the audio data to the server:

[0357] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[0358] Step 4:

[0359] Server receives audio data:

[0360] The server receives the voice data sent from the device and stores it in a secure data storage.

[0361] Step 5:

[0362] The server converts the audio data to text:

[0363] The server calls the speech recognition API to convert the voice data into text data, which is then stored in a database.

[0364] Step 6:

[0365] The server summarizes the text data:

[0366] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[0367] Step 7:

[0368] The server performs emotion recognition using the emotion engine:

[0369] The server uses an emotion engine to analyze the user's emotions from the voice and text data, using algorithms that identify emotions from voice tone and text content.

[0370] Step 8:

[0371] The server stores the emotion recognition results and summary data:

[0372] The server stores the generated summary data and emotion recognition results in a database for later analysis and statistical processing.

[0373] Step 9:

[0374] The server analyzes the summary data and sentiment data:

[0375] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[0376] Step 10:

[0377] Server generates action item:

[0378] The server generates specific action items based on the analysis results and sentiment data, referencing past history and best practices.

[0379] Step 11:

[0380] Server searches knowledge base:

[0381] The server searches and extracts relevant know-how and best practices from the knowledge base.

[0382] Step 12:

[0383] Server generates suggestion list:

[0384] Based on the extracted information, specific action suggestions and a list of best practices are generated for the user.

[0385] Step 13:

[0386] Server sends notification:

[0387] Generated action items and suggestion lists are notified to the user's device in the form of a popup or email.

[0388] Step 14:

[0389] User sees action item:

[0390] The user uses the device to check the notification and view the generated action items and suggested know-how.

[0391] Step 15:

[0392] User performs an action:

[0393] Based on the action items confirmed by the user, the system performs specific tasks, such as setting up a project meeting or coordinating additional resources.

[0394] Example 2

[0395] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0396] In traditional one-on-one meetings, simply recording the conversation and then manually converting it into text and summarizing it later takes time and effort. Nuances such as the speaker's emotions and tone during the meeting are not recorded, often resulting in insufficient understanding and analysis. Furthermore, opportunities for business improvement tend to be missed because specific action items and effective best practices are not generated.

[0397] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for starting recording of the conference, means for transmitting recorded data to a computer, means for the computer to perform speech recognition on the recorded data and convert it into text data, means for the computer to summarize the text data, means for the computer to analyze the summary data and emotion data, means for the computer to generate and suggest action items, and means for recognizing the user's emotions using an emotion recognition engine. This makes it possible to effectively record and analyze the contents of the conference, obtain a detailed understanding including emotional nuances, and suggest specific improvement actions.

[0398] "Means for starting a meeting recording" refers to a feature that allows users to record one-on-one meetings and other meetings on their devices.

[0399] The "means for transmitting recorded data to a computer" is a function for uploading audio data recorded by a terminal to a server in real time or in batch processing.

[0400] "Means for a computer to recognize the recorded data and convert it into text data" refers to the process by which a server uses a speech recognition API to convert the voice data into text data.

[0401] "Means for computer-generated summarization of text data" refers to the process in which the server uses a natural language processing algorithm to extract important information from the converted text data and summarize it concisely.

[0402] "Means for a computer to analyze summary data and emotion data" refers to a function in which the server uses the stored summary and emotion data to analyze past history and patterns and identify trends and issues.

[0403] "Means for computer-generated and proposed action items" refers to a function in which the server generates specific action items based on the results of data analysis and best practices, and proposes them to the user.

[0404] "Means for recognizing the user's emotions using an emotion recognition engine" refers to the process in which the server analyzes the user's emotions from voice data and text data and records them as emotion data.

[0405] "Storing the summary data and emotion data in an information storage device" refers to the process in which the server safely stores the generated summary text data and emotion data in storage such as a database.

[0406] "Extracting know-how and best practices from the knowledge base and generating a proposal list" refers to the process in which the server searches the knowledge base, extracts relevant information, and creates a specific proposal list for the user.

[0407] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[0408] System Configuration

[0409] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[0410] Main Features

[0411] Voice recording and data transmission

[0412] When a user presses the record button on the device, the device begins capturing audio using the built-in microphone. The device then sends the captured audio data to the server at regular intervals (for example, every 10 seconds). When the device sends the recorded data to the server, it uses Wi-Fi or mobile data communication and sends the audio data via an HTTP POST request.

[0413] Converting audio data to text

[0414] The server receives the voice data sent from the device and stores it in a secure data storage (e.g., Amazon S3).The server then converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).Through this conversion process, all conversation content is obtained as text data.

[0415] Text summary generation

[0416] The server uses natural language processing (NLP) algorithms to extract important keywords and context from the acquired text data and summarize it concisely. For example, the server analyzes the text data using an NLP library such as the Natural Language Toolkit (NLTK) and generates a summary.

[0417] emotion recognition

[0418] The server uses an emotion recognition engine to analyze the user's emotions from the voice data and text data. To determine emotions from the tone of the voice or specific words, the server sends the voice data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to obtain emotion data.

[0419] Data storage and analysis

[0420] The server generates summary text data and stores the sentiment data in a database (e.g., MySQL or PostgreSQL). The server uses the stored data to analyze past history and patterns and identify issues and trends. When analyzing the data, the server uses Structured Query Language (SQL).

[0421] Action item generation and suggestions

[0422] The server generates specific action items based on the analysis results and sentiment data. In addition, the server searches a knowledge base to extract relevant information and generate a list of specific suggestions for the user. The server uses a rule-based engine and machine learning algorithms to generate appropriate action items and match them with the knowledge base.

[0423] Notification and confirmation

[0424] The server notifies the user device of the generated action items and suggestion list. Notification methods include pop-up notifications, emails, and in-app notifications. The server sends real-time notifications to the user's smartphone using services such as Firebase Cloud Messaging (FCM). The user receives the notification on their device, checks the specific action items and suggested know-how, and can then take the next step.

[0425] Specific examples

[0426] For example, consider the case where employee A has a one-on-one meeting with line manager B. A presses the record button on their device to begin the meeting. During the meeting, the device sends audio data to the server in real time. The server uses a speech recognition API to convert the audio data into text: "We're experiencing problems with project Y." The server then uses an NLP algorithm to summarize the text as "Problem with project Y." The server's emotion engine recognizes from the tone of the voice and the content of the text that A is feeling stressed about the problem. The server stores the text data and emotion data in a database, analyzes similar past data and emotion data, and generates action items such as "Arrange a meeting to identify problems." The server extracts and suggests "problem-solving methods that were previously effective on project Y" from its knowledge base. Finally, A and B review this information on their devices and create an action plan.

[0427] Examples of prompt statements

[0428] "Generate a prompt sentence to describe the above 1-on-1 meeting management system."

[0429] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0430] A detailed explanation of the system program processing flow and processing steps

[0431] Step 1:

[0432] User starts recording

[0433] Input: A user starts a 1-on-1 meeting.

[0434] How it works: When a user presses the record button on their device, the device's built-in microphone begins capturing audio.

[0435] Output: Audio data is generated and captured.

[0436] Step 2:

[0437] The device sends the voice data to the server

[0438] Input: Audio data captured by the device.

[0439] How it works: Audio data is sent to the server at regular intervals (e.g., every 10 seconds). The audio data is sent via an HTTP POST request using Wi-Fi or mobile data.

[0440] Output: Audio data uploaded to the server.

[0441] Step 3:

[0442] The server receives and stores the audio data

[0443] Input: The server receives the audio data sent from the device.

[0444] How it works: The server receives the audio data and stores it in secure data storage (e.g. Amazon S3).

[0445] Output: Saved audio data.

[0446] Step 4:

[0447] The server converts the voice data into text

[0448] Input: Stored audio data.

[0449] How it works: The server uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the audio data into text, sends the audio file to the API, and stores the returned text.

[0450] Output: The generated text data.

[0451] Step 5:

[0452] The server summarizes the text data

[0453] Input: Generated text data.

[0454] How it works: The server uses NLP algorithms (e.g., the NLTK library) to analyze the text data, extract important keywords and context, and generate a summary.

[0455] Output: The generated summary.

[0456] Step 6:

[0457] The server recognizes emotions

[0458] Input: Audio and text data.

[0459] How it works: The server uses an emotion recognition API (e.g. IBM Watson Tone Analyzer) to extract emotion data from voice tone and text content.

[0460] Output: Emotion data.

[0461] Step 7:

[0462] The server saves and analyzes the summary and emotion data

[0463] Input: Summary sentences and sentiment data.

[0464] How it works: The server saves the summary sentences and emotion data in a database and analyzes past history and patterns. SQL is used to save the data in the database and perform the analysis.

[0465] Output: Analysis of issues and trends.

[0466] Step 8:

[0467] Server generates and suggests action items

[0468] Input: Analysis results and sentiment data.

[0469] How it works: The server generates specific action items using a rule-based engine and machine learning algorithms, extracts relevant information from a knowledge base, and generates a list of suggestions.

[0470] Output: Generated action items and suggestion list.

[0471] Step 9:

[0472] The server notifies the user

[0473] Input: Generated list of action items and suggestions.

[0474] How it works: The server uses Firebase Cloud Messaging (FCM) or similar to send real-time notifications to the user's smartphone.

[0475] Output: A notification message to the user.

[0476] Step 10:

[0477] User confirms action item

[0478] Input: Notification message.

[0479] Action: The user receives a notification on their device, sees an action item, or suggests a know-how, and takes specific steps based on that.

[0480] Output: User confirmed action items and implementation plans.

[0481] (Application example 2)

[0482] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0483] Current meeting recording and data analysis systems are limited to converting audio data into text, providing a brief summary, and suggesting action items. As a result, they lack insight into the emotional state of participants during meetings and specific, real-time countermeasures based on that information. Furthermore, because they are unable to process data in real time, they lack the ability to respond to emergencies and provide immediate response.

[0484] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0485] In this invention, the server includes means for [analyzing summary data and text data using an emotion recognition engine], means for [acquiring emotion data in real time and proposing specific action items such as setting up an emergency response meeting based on the analysis results], and means for [converting audio data into text in real time while recording a meeting, performing emotion analysis, and immediately reflecting the results]. This makes it possible [to grasp the emotional state of participants in a meeting in real time and immediately propose specific measures].

[0486] The "means for starting recording of a conference" is a function that a user operates to record the audio of a conference or meeting.

[0487] The "means for transmitting recorded data to a server" is a function for uploading recorded voice data to a server in real time or later.

[0488] "Means for the server to recognize the voice of the recorded data and convert it into text" refers to a function that enables the server to use voice recognition technology to convert the recorded voice data into text data.

[0489] The "means for the server to summarize text data" is a function for the server to summarize the text data acquired by the server using a natural language processing algorithm and extract important points.

[0490] The "means for the server to analyze the summarized data and text data using an emotion recognition engine" is a function that enables the server to analyze summarized text data and unsummarized text data using an emotion recognition algorithm to identify the emotional states of conference participants.

[0491] "Means for the server to generate and suggest action items" is a function that allows the server to generate specific action items based on the analysis results and suggest them to the user.

[0492] "A means of acquiring emotional data in real time and proposing specific action items such as setting up emergency response meetings based on the analysis results" is a function in which the server acquires emotional data in real time during a meeting and quickly reflects the analysis results to propose actions such as setting up emergency response meetings.

[0493] "Means for converting audio data into text in real time while recording a meeting, performing sentiment analysis, and immediately displaying the results" refers to a function that enables a robot or system to convert the audio data of a meeting into text in real time while recording it, perform sentiment analysis based on the text data, and immediately present the results to the user.

[0494] The "means for saving summary data and emotion data in a database" is a function that enables the server to safely save summary data and emotion recognition results in a database.

[0495] "Means of inputting a prompt sentence into a generative AI model and proposing optimal best practices from past data" is a function that enables the server to input a specific prompt (input sentence) into the generative AI model, search for optimal best practices from past data, and propose them.

[0496] This invention is a system for making work improvement meetings in factories more efficient, and has the functions of converting voice data into text, recognizing emotions, conducting real-time analysis, and generating and proposing action items.

[0497] System Configuration

[0498] The system includes terminals used in factories, a server for processing recorded data, a database, and an emotion recognition engine. The terminals are primarily used for recording meetings, while the server performs speech recognition and analysis. The database stores text data and emotion data. The emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state.

[0499] Program processing

[0500] The system begins when a user presses a recording button on their device to start a meeting. The recorded audio data is sent in real time to a server, which then converts it into text using speech recognition technology. The server then analyzes the text data using a natural language processing algorithm to generate a summary of the meeting. An emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state and reflects the results in real time.

[0501] Novel Features

[0502] 1. Real-time emotional data acquisition and analysis: The server acquires emotional data in real time during the meeting and, based on the analysis results, proposes specific action items such as setting up an emergency response meeting.

[0503] 2. Instant text conversion and sentiment analysis during meeting recording: Audio data is converted into text in real time during meeting recording, and the text data is instantly analyzed for sentiment.

[0504] 3. Use of generative AI models: By inputting prompt statements into generative AI models, the models can suggest optimal best practices based on past data.

[0505] Specific usage

[0506] The device is equipped with a microphone (e.g., a USB microphone) and captures voice data. The recorded data is sent to a server, where it is converted into text data using the Google Speech Recognition API. The text data is then summarized using a natural language processing algorithm (e.g., BERT), and sentiment analysis is performed using an emotion recognition engine. The analysis results are stored in a database and notified to the user in real time. Examples of prompts include "Please print a summary of this meeting" and "Perform emotion recognition and summarize the issues."

[0507] Specific examples

[0508] For example, when factory worker A holds a one-on-one meeting with manager B, he presses the recording button on his device to start the meeting. During the meeting, audio data is sent to the server in real time and instantly converted into text data. The server analyzes the generated text data, and if worker A says, "There has been an increase in product defects this month," it generates a summary saying, "Investigate the cause of product defects," and recognizes through sentiment analysis that worker A is frustrated. As a result, the server suggests "Set up an emergency meeting to address product defects" and generates specific action items.

[0509] This system makes it possible to identify problems in real time during meetings and respond immediately, which is expected to lead to more efficient work improvements within the factory.

[0510] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0511] Step 1:

[0512] The user presses the record button on the device to start the conference.

[0513] Input: User actions

[0514] How it works: By pressing the record button, the device will begin recording the audio data during the meeting.

[0515] Output: Audio data begins to accumulate on the device.

[0516] Step 2:

[0517] The device sends the recorded data to the server in real time.

[0518] Input: Recorded audio data

[0519] How it works: Recorded data is uploaded to the server in real time, segment by segment.

[0520] Output: Audio data sent to the server

[0521] Step 3:

[0522] The server converts the voice data into text data using a voice recognition API.

[0523] Input: Audio data sent to the server

[0524] How it works: Uses the Google Speech Recognition API to convert audio data into text.

[0525] Output: Text data (e.g., "Product defects are increasing this month.")

[0526] Step 4:

[0527] The server summarizes the text data using natural language processing algorithms.

[0528] Input: Text data converted by speech recognition

[0529] How it works: It uses natural language processing algorithms such as BERT to extract key points from text data and generate summaries.

[0530] Output: Summary data (e.g., "Investigation into the cause of product defects")

[0531] Step 5:

[0532] The server analyzes the summary data and text data using an emotion recognition engine.

[0533] Input: Abstract and text data

[0534] How it works: An emotion recognition engine is used to recognize and analyze user emotions from text data and voice tones.

[0535] Output: Emotion data (e.g., stress and irritation are recognized)

[0536] Step 6:

[0537] The server stores the summary data and the emotion data in a database.

[0538] Input: Summary data and sentiment data

[0539] What it does: Stores data securely in a database.

[0540] Output: Summary data and sentiment data stored in a database

[0541] Step 7:

[0542] The server generates and suggests action items.

[0543] Input: Summary data and sentiment data

[0544] How it works: You input a prompt into the generative AI model, which will then use past data to suggest best practices. For example, "Please provide a summary of this meeting." "Please use emotion recognition to summarize the issues."

[0545] Output: Suggested action items and best practices (e.g., "Schedule an emergency product defect resolution meeting")

[0546] Step 8:

[0547] The user reviews the proposed action items and creates an implementation plan.

[0548] Input: Action items and best practices notified by the server

[0549] How it works: The user sees specific action items on their device and creates an action plan based on them.

[0550] Output: A concrete action plan (e.g., "Schedule an emergency meeting on ____ day to discuss problem-solving methods")

[0551] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0552] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0553] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0554] [Second embodiment]

[0555] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0556] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0557] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0558] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0559] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0560] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0561] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0562] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0563] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0564] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0565] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0566] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0567] The present invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, mainly including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices. The following describes in detail the embodiments of the present invention.

[0568] 1. System Configuration

[0569] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[0570] 2. Description of main functions

[0571] Voice recording and data transmission

[0572] User starts recording:

[0573] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[0574] The device sends the audio data to the server:

[0575] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[0576] Converting audio data to text

[0577] Server receives audio data:

[0578] The server receives the voice data sent from the terminal and stores it securely in data storage.

[0579] Audio to text conversion:

[0580] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[0581] Text summary generation

[0582] The server parses the text data:

[0583] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[0584] Summary generation:

[0585] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[0586] Data analysis and action item generation

[0587] Summary data analysis:

[0588] The server aggregates multiple summaries to identify common issues and trends.

[0589] Generate action items:

[0590] The server uses past meeting data and success stories to generate action items for the issues, such as "reviewing tasks" and "adjusting additional resources."

[0591] Proposal of know-how and best practices

[0592] Search our knowledge base:

[0593] The server searches and extracts relevant know-how and best practices from the knowledge base.

[0594] Generate a list of suggestions:

[0595] The server organizes the search results and creates a specific list of suggestions for the user.

[0596] Notification and confirmation

[0597] Send notifications:

[0598] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[0599] User sees action item:

[0600] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[0601] Specific examples

[0602] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[0603] 1. Start recording:

[0604] Person A presses the recording button on the device to start the meeting.

[0605] 2. Sending audio data to the server:

[0606] During the meeting, the device transmits audio data to the server in real time.

[0607] 3. Text:

[0608] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[0609] 4. Summary generation:

[0610] Using an NLP algorithm, summarize it as "Project Y problem."

[0611] 5. Data analysis and action generation:

[0612] The server analyzes similar past data and generates action items such as "set up a meeting to identify problems."

[0613] 6. Know-how proposal:

[0614] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[0615] 7. User confirms:

[0616] Persons A and B check this information on their devices and create an action plan.

[0617] In this way, the system can efficiently manage the content of one-on-one meetings and link them to concrete actions.

[0618] The processing flow will be explained below.

[0619] Step 1:

[0620] User starts recording:

[0621] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[0622] Step 2:

[0623] The device captures audio data:

[0624] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[0625] Step 3:

[0626] The device sends the audio data to the server:

[0627] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[0628] Step 4:

[0629] Server receives audio data:

[0630] The server receives the voice data sent from the terminal and stores the data in storage.

[0631] Step 5:

[0632] The server converts the audio data to text:

[0633] The server uses a speech recognition API to convert the voice data into text data in real time or in batches, and the converted text data is stored in a database.

[0634] Step 6:

[0635] The server summarizes the text data:

[0636] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[0637] Step 7:

[0638] The server stores the summary data:

[0639] The summarized text data is stored in a database and used for subsequent analysis and statistical processing.

[0640] Step 8:

[0641] The server analyzes the summary data:

[0642] The server aggregates multiple summaries of data and uses an analytics engine to identify issues and trends.

[0643] Step 9:

[0644] Server generates action item:

[0645] Based on the analysis results, the server generates specific action items, referencing past history and best practices.

[0646] Step 10:

[0647] Server searches knowledge base:

[0648] The server searches databases and knowledge bases to extract relevant know-how and best practices.

[0649] Step 11:

[0650] Server generates suggestion list:

[0651] Based on the extracted information, a list of specific action suggestions and best practices is generated to provide to the user.

[0652] Step 12:

[0653] Server sends notification:

[0654] Generated action items and suggestion lists are notified to the user's device, including via pop-up notifications and email notifications.

[0655] Step 13:

[0656] User sees action item:

[0657] The user uses the device to check the notification and view the generated action items and suggested know-how.

[0658] Step 14:

[0659] User performs an action:

[0660] The user then executes specific tasks based on the action items they have confirmed, such as setting up a project meeting or arranging for additional resources.

[0661] Example 1

[0662] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0663] Conventional conference management systems have had difficulty efficiently managing conference content and linking it to specific actions. Furthermore, converting conference recordings into text, summarizing them, and generating action items require a lot of manual work, which takes time and effort. Therefore, there is a demand for a system that can achieve efficient conference management and link it to specific actions.

[0664] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0665] In this invention, the server includes a means for [starting recording of the meeting], a means for [sending the recorded data to the server], a means for [the server converting the recorded data into text data using speech recognition technology], a means for [the server summarizing the text data using a natural language processing algorithm], a means for [the server analyzing the summarized data using a data analysis algorithm], and a means for [the server generating and proposing action items based on past success stories]. This makes it possible to convert the contents of the meeting into text and summarize it in real time, and automatically generate and propose specific action items.

[0666] The "means for starting recording of a conference" refers to the device or software that a user uses to record a conference, specifically a recording button.

[0667] "Means for transmitting recorded data to a server" refers to a function or device that transmits voice data recorded by a user to a server in real time or in batch mode.

[0668] "Means by which the server converts recorded data into text data using speech recognition technology" refers to the functions and processes by which the server converts audio data into text data using a speech recognition API or algorithm.

[0669] "Means by which the server summarizes text data using a natural language processing algorithm" refers to the function or algorithm by which the server uses natural language processing technology to extract important information from text data and summarize it concisely.

[0670] "Means for the server to analyze the summarized data using data analysis algorithms" means the data analysis algorithms used by the server to analyze the summarized text data and identify recurring issues or trends.

[0671] "Means for the server to generate and suggest action items based on past success stories" refers to the process or function in which the server refers to past meeting data and best practices, generates specific action items for specific issues, and suggests them to the user.

[0672] "Means for storing abstract data in a database" refers to the functions and processes for securely storing abstract data generated by the server in a database.

[0673] "Means for the server to extract relevant knowledge and best practices from the knowledge base and generate a list of suggestions" refers to the process or function by which the server searches an existing knowledge base, extracts useful information and best practices that meet the user's needs, and provides them as a list.

[0674] This invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices.

[0675] System Configuration

[0676] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[0677] Hardware and Software

[0678] 1. Device: The smartphone or computer used by the user.

[0679] 2. Server: Cloud server that processes data.

[0680] 3. Database: A database for storing meeting data, analysis results, etc.

[0681] 4. Speech recognition APIs: Google Speech-to-Text, Amazon Transcribe, etc.

[0682] 5. Natural Language Processing (NLP) algorithms: Open source libraries such as SpaCy and NLTK.

[0683] System functions and examples

[0684] Voice recording and data transmission

[0685] User starts recording:

[0686] The user opens the dedicated app on their device and presses the record button. By pressing the record button, the app activates the device's microphone and begins capturing audio. For example, employee A uses the smartphone app to record a one-on-one meeting.

[0687] The device sends the audio data to the server:

[0688] The device uploads the captured audio data to the server in regular segments in real time. For example, the audio data is divided into segments every 10 seconds and sent to the server sequentially.

[0689] Converting audio data to text

[0690] Server receives audio data:

[0691] The server receives the voice data sent from the device and stores it securely in data storage, for example, as an audio file in cloud storage.

[0692] The server converts the audio data to text:

[0693] The server converts the received voice data into text using a speech recognition API (e.g., Google Speech-to-Text). For example, a speech saying "Project X's deadline is behind schedule" is obtained as text data.

[0694] Text summary generation

[0695] The server parses the text data:

[0696] The server applies NLP algorithms (e.g., SpaCy) to extract important keywords and context from the text data.

[0697] Server generates summary:

[0698] The server summarizes the original text based on the extracted information. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[0699] Data analysis

[0700] The server aggregates the summary data:

[0701] The server aggregates multiple summaries generated from past meetings and identifies common issues and trends. For example, "late delivery" emerges as a common issue across multiple projects.

[0702] Generate action items

[0703] Server generates action item:

[0704] The server generates specific action items for issues based on past success stories and knowledge bases. For example, it suggests "setting up an emergency meeting" or "arranging additional resources" as a way to address delivery delays.

[0705] Proposal of know-how and best practices

[0706] Server searches for know-how:

[0707] The server searches the knowledge base for relevant know-how and best practices, for example, "project management best practices."

[0708] Server generates suggestion list:

[0709] The server creates a list of suggestions based on the search results and provides it to the user. Specifically, it displays a list of past success stories and procedures based on those stories.

[0710] Notification and confirmation

[0711] Server sends notification:

[0712] The server notifies the user of the generated action items and suggestion list via a pop-up notification on their smartphone or email.

[0713] User sees action item:

[0714] Users check notifications on their devices and view suggested action items and know-how. They then plan specific actions and put them into action. For example, Person A and Person B discuss specific measures based on the information displayed on their devices and create an action plan.

[0715] Prompt Sentence Examples

[0716] "Please summarize the content of the 1-on-1 meeting and propose specific action items and know-how."

[0717] "Convert the recorded audio data into text and extract the key points."

[0718] This system allows you to efficiently manage the content of one-on-one meetings and link them to concrete actions.

[0719] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0720] Step 1:

[0721] User starts recording

[0722] The user opens the dedicated app on their device and presses the record button. The input is the user clicking the record button, and the device starts capturing audio. Specifically, the smartphone's microphone is activated and audio data is captured in real time. The output is the audio data being recorded.

[0723] Step 2:

[0724] The device sends the voice data to the server

[0725] The device uploads captured audio data to the server in regular segments in real time. The input is the audio data being recorded, divided into segments, such as 10 seconds. Specifically, the device sends the audio data segments to the server using HTTP requests. The output is the audio data stored on the server.

[0726] Step 3:

[0727] The server receives the audio data

[0728] The server receives the voice data sent from the terminal and safely stores it in the data storage. The input is the voice data sent from the terminal, and the output is the voice file stored in the data storage. Specifically, the server executes a process to store the voice file in a database.

[0729] Step 4:

[0730] The server converts the voice data into text

[0731] The server uses a speech recognition API (e.g., Google Speech-to-Text) to convert the voice data into text. The input is a saved audio file, and when the server calls the API, the API analyzes the voice data and returns it as text data. The output is text data. For example, the text generated might say, "The delivery date for Project X is behind schedule."

[0732] Step 5:

[0733] The server analyzes the text data

[0734] The server applies an NLP algorithm (e.g., SpaCy) to extract important keywords and context from the text data. The input is text data, which the server analyzes by running the NLP algorithm. The output is important keywords and context data. For example, the important keyword "delayed delivery" is extracted.

[0735] Step 6:

[0736] Server generates summary

[0737] The server summarizes the original text based on the extracted information. The input is important keywords and contextual data, and the server generates a summary by running a summary generation algorithm. The output is summarized text data. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[0738] Step 7:

[0739] The server aggregates the summary data

[0740] The server aggregates multiple summaries generated from past meetings and identifies frequently occurring issues and trends. The input is multiple summarized text data, which the server analyzes by applying a data aggregation algorithm. The output is the identification of issues and trends. For example, "delayed delivery" is identified as a common issue across multiple projects.

[0741] Step 8:

[0742] Server generates action items

[0743] The server generates specific action items for issues based on past success stories and a knowledge base. The input is the results of identifying issues and trends, and the server generates specific action items by running an action item generation algorithm. The output is a specific action item. For example, it may suggest measures such as "setting up an emergency meeting" or "adjusting additional resources" to address delivery delays.

[0744] Step 9:

[0745] The server searches for know-how

[0746] The server searches the knowledge base to find relevant know-how and best practices. The input is summary data and action items, and the server runs a knowledge base search algorithm to extract relevant information. The output is know-how and best practice information.

[0747] Step 10:

[0748] The server generates a list of suggestions

[0749] The server creates a proposal list based on the search results and provides it to the user. The input is know-how and best practice information, and the server generates the list by running a proposal list generation algorithm. The output is the proposal list provided to the user.

[0750] Step 11:

[0751] The server sends a notification

[0752] The server notifies the user's terminal of the generated action items and suggestion list. The input is the action items and suggestion list, and the server executes the notification process to send the notification to the user. The output is the notification displayed on the user's terminal.

[0753] Step 12:

[0754] User confirms action item

[0755] The user receives a notification on their device and checks the proposed action items and know-how. The input is the notification displayed on the user's device, and the specific action is confirmed by the user checking the operation. The output is the creation of an action plan. For example, Person A and Person B discuss specific measures and create an action plan.

[0756] (Application example 1)

[0757] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0758] There is a need to efficiently manage the content of one-on-one meetings held in the field and link it to specific action items and best practices. However, the current situation is such that the content of meetings is not fully utilized, resulting in a waste of time and resources. In addition, it is a heavy burden for administrators to continue to manage the content of each meeting individually. For this reason, there is a need for a more efficient and automated system.

[0759] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0760] In this invention, the server includes means for starting recording of the meeting, means for transmitting the recorded data to the server, means for the server to recognize the voice of the recorded data and convert it into text, means for the server to summarize the text data, means for the server to analyze the summarized data, means for the server to generate and propose action items, means for automatically generating a proposal list using a generative AI model, and means for notifying the action items and proposal list. This makes it possible to efficiently manage the content of one-on-one meetings and link them to specific actions.

[0761] "Means for starting recording of a meeting" refers to a device or function that allows a user to record the audio of a meeting by pressing a button or other operation.

[0762] "Means for transmitting recorded data to a server" refers to communication functions or programs for uploading recorded audio data to a server in real time or by batch processing.

[0763] "Means for the server to recognize the recorded data and convert it into text" refers to the process of analyzing the voice data on the server and converting it into text data using voice recognition technology.

[0764] "Means for the server to summarize text data" refers to the function of the server using natural language processing (NLP) algorithms, etc. to extract important information from text data and summarize it concisely.

[0765] "Means for the server to analyze the summary data" refers to a data analysis function that analyzes the summary data generated by the server and identifies significant keywords and patterns.

[0766] "Means for the server to generate and suggest action items" refers to the function by which the server automatically generates specific actions and suggestions that the user should take based on the analysis results.

[0767] "Means for automatically generating a suggestion list using a generative AI model" refers to an algorithm or program for automatically creating a suggestion list for a user using a generative AI model.

[0768] "Means for notifying action items and suggestion lists" refers to the system's functionality for notifying the user's terminal of generated action items and suggestion lists.

[0769] The present invention provides a system for efficiently managing the content of one-on-one meetings and linking them to specific actions. The following describes in detail the embodiments of the present invention.

[0770] System Configuration

[0771] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[0772] Voice recording and data transmission

[0773] Yu

[0774] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0775] Step 1:

[0776] When a user starts a one-on-one meeting, they press the recording button on their device, and the device begins capturing audio.

[0777] Input: User presses record button.

[0778] Output: Recorded audio data.

[0779] Specific operation: Captures audio data using the device's microphone function.

[0780] Step 2:

[0781] The device transmits the recorded audio data to the server in real time.

[0782] Input: Pre-recorded audio data.

[0783] Output: The audio data sent to the server.

[0784] Specific operation: Audio data is divided and uploaded to the server via data communication.

[0785] Step 3:

[0786] The server receives the recording and stores it securely.

[0787] Input: Audio data sent from the device.

[0788] Output: Audio data stored on the server.

[0789] What it does: Receives audio data and stores it securely in a database or storage.

[0790] Step 4:

[0791] The server converts the recorded data into text data using voice recognition technology.

[0792] Input: Recorded audio data.

[0793] Output: Text data.

[0794] Specific operation: Uses a speech recognition API to convert voice data into text.

[0795] Step 5:

[0796] The server applies natural language processing (NLP) algorithms to summarize the text data.

[0797] Input: Text data.

[0798] Output: Summarized text data.

[0799] Specific operation: Using NLP algorithms, important keywords and context are extracted from text data and a summary is generated.

[0800] Step 6:

[0801] The server analyzes the summary data and generates action items based thereon.

[0802] Input: Summarized text data.

[0803] Output: Action items.

[0804] Specific Behavior: Analyzes summary data, identifies issues and trends, and generates corresponding action items.

[0805] Step 7:

[0806] The server uses a generative AI model to automatically generate a list of suggestions including relevant know-how and best practices.

[0807] Inputs: Action items and knowledge base data.

[0808] Output: A list of suggestions.

[0809] What it does: It applies generative AI models to extract useful know-how and best practices from a knowledge base and generate a list.

[0810] Step 8:

[0811] The server notifies the user's terminal of the action items and the suggestion list.

[0812] Input: Action items and suggestion lists.

[0813] Output: Notification displayed on the user's device.

[0814] Specific behavior: Using the notification system, the generated action items and suggestion list are sent to the user's device via push notification or email.

[0815] For example, if a factory manager discusses a problem with a new production line during a one-on-one meeting with a field staff member, the server will summarize the content and suggest specific action items such as setting up a trouble-shooting team meeting. In this way, the system can efficiently manage the content of meetings and lead to actual actions.

[0816] For example, an example of a prompt for a generative AI model is: "Please summarize the following text: There is a problem with the new production line. Specifically, there is a problem with the machines frequently stopping. The cause has not yet been identified, but manual intervention may be necessary."

[0817] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0818] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[0819] 1. System Configuration

[0820] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[0821] 2. Description of main functions

[0822] Voice recording and data transmission

[0823] User starts recording:

[0824] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[0825] The device sends the audio data to the server:

[0826] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[0827] Converting audio data to text

[0828] Server receives audio data:

[0829] The server receives the voice data sent from the terminal and stores it securely in data storage.

[0830] Audio to text conversion:

[0831] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[0832] Text summary generation

[0833] The server parses the text data:

[0834] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[0835] Summary generation:

[0836] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[0837] emotion recognition

[0838] The server uses an emotion engine to recognize the user's emotions:

[0839] The server analyzes the user's emotions from the voice and text data, including algorithms that identify emotions from voice tone and text content.

[0840] Data storage and analysis

[0841] Summary and sentiment data storage:

[0842] The server stores the summarized text data and the sentiment analysis results in a database.

[0843] Data Analysis:

[0844] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[0845] Action item generation and suggestions

[0846] Server generates action item:

[0847] The server generates specific action items based on the analysis results and sentiment data, referring to past history and best practices.

[0848] Search our knowledge base:

[0849] The server searches and extracts relevant know-how and best practices from the knowledge base.

[0850] Generate a list of suggestions:

[0851] The server generates specific action suggestions and a list of best practices for the user based on the search results.

[0852] Notification and confirmation

[0853] Send notifications:

[0854] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[0855] User sees action item:

[0856] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[0857] Specific examples

[0858] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[0859] 1. Start recording:

[0860] Person A presses the recording button on the device to start the meeting.

[0861] 2. Sending audio data to the server:

[0862] During the meeting, the device transmits audio data to the server in real time.

[0863] 3. Text:

[0864] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[0865] 4. Summary generation:

[0866] Using an NLP algorithm, summarize it as "Project Y problem."

[0867] 5. Emotion recognition:

[0868] The emotion engine recognizes from the tone of the voice and the content of the text that Person A is feeling stressed about the problem.

[0869] 6. Data Retention:

[0870] The server stores the text data and emotion data in a database.

[0871] 7. Data analysis and action generation:

[0872] The server analyzes similar past data and sentiment data and generates action items such as "set up a meeting to identify problems."

[0873] 8. Know-how proposal:

[0874] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[0875] 9. User confirms:

[0876] Persons A and B check this information on their devices and create an action plan.

[0877] In this way, by combining an emotion recognition engine, it is possible to provide a management system that deepens the content of one-on-one meetings and leads to concrete actions.

[0878] The processing flow will be explained below.

[0879] Step 1:

[0880] User starts recording:

[0881] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[0882] Step 2:

[0883] The device captures audio data:

[0884] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[0885] Step 3:

[0886] The device sends the audio data to the server:

[0887] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[0888] Step 4:

[0889] Server receives audio data:

[0890] The server receives the voice data sent from the device and stores it in a secure data storage.

[0891] Step 5:

[0892] The server converts the audio data to text:

[0893] The server calls the speech recognition API to convert the voice data into text data, which is then stored in a database.

[0894] Step 6:

[0895] The server summarizes the text data:

[0896] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[0897] Step 7:

[0898] The server performs emotion recognition using the emotion engine:

[0899] The server uses an emotion engine to analyze the user's emotions from the voice and text data, using algorithms that identify emotions from voice tone and text content.

[0900] Step 8:

[0901] The server stores the emotion recognition results and summary data:

[0902] The server stores the generated summary data and emotion recognition results in a database for later analysis and statistical processing.

[0903] Step 9:

[0904] The server analyzes the summary data and sentiment data:

[0905] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[0906] Step 10:

[0907] Server generates action item:

[0908] The server generates specific action items based on the analysis results and sentiment data, referencing past history and best practices.

[0909] Step 11:

[0910] Server searches knowledge base:

[0911] The server searches and extracts relevant know-how and best practices from the knowledge base.

[0912] Step 12:

[0913] Server generates suggestion list:

[0914] Based on the extracted information, specific action suggestions and a list of best practices are generated for the user.

[0915] Step 13:

[0916] Server sends notification:

[0917] Generated action items and suggestion lists are notified to the user's device in the form of a popup or email.

[0918] Step 14:

[0919] User sees action item:

[0920] The user uses the device to check the notification and view the generated action items and suggested know-how.

[0921] Step 15:

[0922] User performs an action:

[0923] Based on the action items confirmed by the user, the system performs specific tasks, such as setting up a project meeting or coordinating additional resources.

[0924] Example 2

[0925] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0926] In traditional one-on-one meetings, simply recording the conversation and then manually converting it into text and summarizing it later takes time and effort. Nuances such as the speaker's emotions and tone during the meeting are not recorded, often resulting in insufficient understanding and analysis. Furthermore, opportunities for business improvement tend to be missed because specific action items and effective best practices are not generated.

[0927] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for starting recording of the conference, means for transmitting recorded data to a computer, means for the computer to perform speech recognition on the recorded data and convert it into text data, means for the computer to summarize the text data, means for the computer to analyze the summary data and emotion data, means for the computer to generate and suggest action items, and means for recognizing the user's emotions using an emotion recognition engine. This makes it possible to effectively record and analyze the contents of the conference, obtain a detailed understanding including emotional nuances, and suggest specific improvement actions.

[0928] "Means for starting a meeting recording" refers to a feature that allows users to record one-on-one meetings and other meetings on their devices.

[0929] The "means for transmitting recorded data to a computer" is a function for uploading audio data recorded by a terminal to a server in real time or in batch processing.

[0930] "Means for a computer to recognize the recorded data and convert it into text data" refers to the process by which a server uses a speech recognition API to convert the voice data into text data.

[0931] "Means for computer-generated summarization of text data" refers to the process in which the server uses a natural language processing algorithm to extract important information from the converted text data and summarize it concisely.

[0932] "Means for a computer to analyze summary data and emotion data" refers to a function in which the server uses the stored summary and emotion data to analyze past history and patterns and identify trends and issues.

[0933] "Means for computer-generated and proposed action items" refers to a function in which the server generates specific action items based on the results of data analysis and best practices, and proposes them to the user.

[0934] "Means for recognizing the user's emotions using an emotion recognition engine" refers to the process in which the server analyzes the user's emotions from voice data and text data and records them as emotion data.

[0935] "Storing the summary data and emotion data in an information storage device" refers to the process in which the server safely stores the generated summary text data and emotion data in storage such as a database.

[0936] "Extracting know-how and best practices from the knowledge base and generating a proposal list" refers to the process in which the server searches the knowledge base, extracts relevant information, and creates a specific proposal list for the user.

[0937] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[0938] System Configuration

[0939] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[0940] Main Features

[0941] Voice recording and data transmission

[0942] When a user presses the record button on the device, the device begins capturing audio using the built-in microphone. The device then sends the captured audio data to the server at regular intervals (for example, every 10 seconds). When the device sends the recorded data to the server, it uses Wi-Fi or mobile data communication and sends the audio data via an HTTP POST request.

[0943] Converting audio data to text

[0944] The server receives the voice data sent from the device and stores it in a secure data storage (e.g., Amazon S3).The server then converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).Through this conversion process, all conversation content is obtained as text data.

[0945] Text summary generation

[0946] The server uses natural language processing (NLP) algorithms to extract important keywords and context from the acquired text data and summarize it concisely. For example, the server analyzes the text data using an NLP library such as the Natural Language Toolkit (NLTK) and generates a summary.

[0947] emotion recognition

[0948] The server uses an emotion recognition engine to analyze the user's emotions from the voice data and text data. To determine emotions from the tone of the voice or specific words, the server sends the voice data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to obtain emotion data.

[0949] Data storage and analysis

[0950] The server generates summary text data and stores the sentiment data in a database (e.g., MySQL or PostgreSQL). The server uses the stored data to analyze past history and patterns and identify issues and trends. When analyzing the data, the server uses Structured Query Language (SQL).

[0951] Action item generation and suggestions

[0952] The server generates specific action items based on the analysis results and sentiment data. In addition, the server searches a knowledge base to extract relevant information and generate a list of specific suggestions for the user. The server uses a rule-based engine and machine learning algorithms to generate appropriate action items and match them with the knowledge base.

[0953] Notification and confirmation

[0954] The server notifies the user device of the generated action items and suggestion list. Notification methods include pop-up notifications, emails, and in-app notifications. The server sends real-time notifications to the user's smartphone using services such as Firebase Cloud Messaging (FCM). The user receives the notification on their device, checks the specific action items and suggested know-how, and can then take the next step.

[0955] Specific examples

[0956] For example, consider the case where employee A has a one-on-one meeting with line manager B. A presses the record button on their device to begin the meeting. During the meeting, the device sends audio data to the server in real time. The server uses a speech recognition API to convert the audio data into text: "We're experiencing problems with project Y." The server then uses an NLP algorithm to summarize the text as "Problem with project Y." The server's emotion engine recognizes from the tone of the voice and the content of the text that A is feeling stressed about the problem. The server stores the text data and emotion data in a database, analyzes similar past data and emotion data, and generates action items such as "Arrange a meeting to identify problems." The server extracts and suggests "problem-solving methods that were previously effective on project Y" from its knowledge base. Finally, A and B review this information on their devices and create an action plan.

[0957] Examples of prompt statements

[0958] "Generate a prompt sentence to describe the above 1-on-1 meeting management system."

[0959] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0960] A detailed explanation of the system program processing flow and processing steps

[0961] Step 1:

[0962] User starts recording

[0963] Input: A user starts a 1-on-1 meeting.

[0964] How it works: When a user presses the record button on their device, the device's built-in microphone begins capturing audio.

[0965] Output: Audio data is generated and captured.

[0966] Step 2:

[0967] The device sends the voice data to the server

[0968] Input: Audio data captured by the device.

[0969] How it works: Audio data is sent to the server at regular intervals (e.g., every 10 seconds). The audio data is sent via an HTTP POST request using Wi-Fi or mobile data.

[0970] Output: Audio data uploaded to the server.

[0971] Step 3:

[0972] The server receives and stores the audio data

[0973] Input: The server receives the audio data sent from the device.

[0974] How it works: The server receives the audio data and stores it in secure data storage (e.g. Amazon S3).

[0975] Output: Saved audio data.

[0976] Step 4:

[0977] The server converts the voice data into text

[0978] Input: Stored audio data.

[0979] How it works: The server uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the audio data into text, sends the audio file to the API, and stores the returned text.

[0980] Output: The generated text data.

[0981] Step 5:

[0982] The server summarizes the text data

[0983] Input: Generated text data.

[0984] How it works: The server uses NLP algorithms (e.g., the NLTK library) to analyze the text data, extract important keywords and context, and generate a summary.

[0985] Output: The generated summary.

[0986] Step 6:

[0987] The server recognizes emotions

[0988] Input: Audio and text data.

[0989] How it works: The server uses an emotion recognition API (e.g. IBM Watson Tone Analyzer) to extract emotion data from voice tone and text content.

[0990] Output: Emotion data.

[0991] Step 7:

[0992] The server saves and analyzes the summary and emotion data

[0993] Input: Summary sentences and sentiment data.

[0994] How it works: The server saves the summary sentences and emotion data in a database and analyzes past history and patterns. SQL is used to save the data in the database and perform the analysis.

[0995] Output: Analysis of issues and trends.

[0996] Step 8:

[0997] Server generates and suggests action items

[0998] Input: Analysis results and sentiment data.

[0999] How it works: The server generates specific action items using a rule-based engine and machine learning algorithms, extracts relevant information from a knowledge base, and generates a list of suggestions.

[1000] Output: Generated action items and suggestion list.

[1001] Step 9:

[1002] The server notifies the user

[1003] Input: Generated list of action items and suggestions.

[1004] How it works: The server uses Firebase Cloud Messaging (FCM) or similar to send real-time notifications to the user's smartphone.

[1005] Output: A notification message to the user.

[1006] Step 10:

[1007] User confirms action item

[1008] Input: Notification message.

[1009] Action: The user receives a notification on their device, sees an action item, or suggests a know-how, and takes specific steps based on that.

[1010] Output: User confirmed action items and implementation plans.

[1011] (Application example 2)

[1012] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1013] Current meeting recording and data analysis systems are limited to converting audio data into text, providing a brief summary, and suggesting action items. As a result, they lack insight into the emotional state of participants during meetings and specific, real-time countermeasures based on that information. Furthermore, because they are unable to process data in real time, they lack the ability to respond to emergencies and provide immediate response.

[1014] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1015] In this invention, the server includes means for [analyzing summary data and text data using an emotion recognition engine], means for [acquiring emotion data in real time and proposing specific action items such as setting up an emergency response meeting based on the analysis results], and means for [converting audio data into text in real time while recording a meeting, performing emotion analysis, and immediately reflecting the results]. This makes it possible [to grasp the emotional state of participants in a meeting in real time and immediately propose specific measures].

[1016] The "means for starting recording of a conference" is a function that a user operates to record the audio of a conference or meeting.

[1017] The "means for transmitting recorded data to a server" is a function for uploading recorded voice data to a server in real time or later.

[1018] "Means for the server to recognize the voice of the recorded data and convert it into text" refers to a function that enables the server to use voice recognition technology to convert the recorded voice data into text data.

[1019] The "means for the server to summarize text data" is a function for the server to summarize the text data acquired by the server using a natural language processing algorithm and extract important points.

[1020] The "means for the server to analyze the summarized data and text data using an emotion recognition engine" is a function that enables the server to analyze summarized text data and unsummarized text data using an emotion recognition algorithm to identify the emotional states of conference participants.

[1021] "Means for the server to generate and suggest action items" is a function that allows the server to generate specific action items based on the analysis results and suggest them to the user.

[1022] "A means of acquiring emotional data in real time and proposing specific action items such as setting up emergency response meetings based on the analysis results" is a function in which the server acquires emotional data in real time during a meeting and quickly reflects the analysis results to propose actions such as setting up emergency response meetings.

[1023] "Means for converting audio data into text in real time while recording a meeting, performing sentiment analysis, and immediately displaying the results" refers to a function that enables a robot or system to convert the audio data of a meeting into text in real time while recording it, perform sentiment analysis based on the text data, and immediately present the results to the user.

[1024] The "means for saving summary data and emotion data in a database" is a function that enables the server to safely save summary data and emotion recognition results in a database.

[1025] "Means of inputting a prompt sentence into a generative AI model and proposing optimal best practices from past data" is a function that enables the server to input a specific prompt (input sentence) into the generative AI model, search for optimal best practices from past data, and propose them.

[1026] This invention is a system for making work improvement meetings in factories more efficient, and has the functions of converting voice data into text, recognizing emotions, conducting real-time analysis, and generating and proposing action items.

[1027] System Configuration

[1028] The system includes terminals used in factories, a server for processing recorded data, a database, and an emotion recognition engine. The terminals are primarily used for recording meetings, while the server performs speech recognition and analysis. The database stores text data and emotion data. The emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state.

[1029] Program processing

[1030] The system begins when a user presses a recording button on their device to start a meeting. The recorded audio data is sent in real time to a server, which then converts it into text using speech recognition technology. The server then analyzes the text data using a natural language processing algorithm to generate a summary of the meeting. An emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state and reflects the results in real time.

[1031] Novel Features

[1032] 1. Real-time emotional data acquisition and analysis: The server acquires emotional data in real time during the meeting and, based on the analysis results, proposes specific action items such as setting up an emergency response meeting.

[1033] 2. Instant text conversion and sentiment analysis during meeting recording: Audio data is converted into text in real time during meeting recording, and the text data is instantly analyzed for sentiment.

[1034] 3. Use of generative AI models: By inputting prompt statements into generative AI models, the models can suggest optimal best practices based on past data.

[1035] Specific usage

[1036] The device is equipped with a microphone (e.g., a USB microphone) and captures voice data. The recorded data is sent to a server, where it is converted into text data using the Google Speech Recognition API. The text data is then summarized using a natural language processing algorithm (e.g., BERT), and sentiment analysis is performed using an emotion recognition engine. The analysis results are stored in a database and notified to the user in real time. Examples of prompts include "Please print a summary of this meeting" and "Perform emotion recognition and summarize the issues."

[1037] Specific examples

[1038] For example, when factory worker A holds a one-on-one meeting with manager B, he presses the recording button on his device to start the meeting. During the meeting, audio data is sent to the server in real time and instantly converted into text data. The server analyzes the generated text data, and if worker A says, "There has been an increase in product defects this month," it generates a summary saying, "Investigate the cause of product defects," and recognizes through sentiment analysis that worker A is frustrated. As a result, the server suggests "Set up an emergency meeting to address product defects" and generates specific action items.

[1039] This system makes it possible to identify problems in real time during meetings and respond immediately, which is expected to lead to more efficient work improvements within the factory.

[1040] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1041] Step 1:

[1042] The user presses the record button on the device to start the conference.

[1043] Input: User actions

[1044] How it works: By pressing the record button, the device will begin recording the audio data during the meeting.

[1045] Output: Audio data begins to accumulate on the device.

[1046] Step 2:

[1047] The device sends the recorded data to the server in real time.

[1048] Input: Recorded audio data

[1049] How it works: Recorded data is uploaded to the server in real time, segment by segment.

[1050] Output: Audio data sent to the server

[1051] Step 3:

[1052] The server converts the voice data into text data using a voice recognition API.

[1053] Input: Audio data sent to the server

[1054] How it works: Uses the Google Speech Recognition API to convert audio data into text.

[1055] Output: Text data (e.g., "Product defects are increasing this month.")

[1056] Step 4:

[1057] The server summarizes the text data using natural language processing algorithms.

[1058] Input: Text data converted by speech recognition

[1059] How it works: It uses natural language processing algorithms such as BERT to extract key points from text data and generate summaries.

[1060] Output: Summary data (e.g., "Investigation into the cause of product defects")

[1061] Step 5:

[1062] The server analyzes the summary data and text data using an emotion recognition engine.

[1063] Input: Abstract and text data

[1064] How it works: An emotion recognition engine is used to recognize and analyze user emotions from text data and voice tones.

[1065] Output: Emotion data (e.g., stress and irritation are recognized)

[1066] Step 6:

[1067] The server stores the summary data and the emotion data in a database.

[1068] Input: Summary data and sentiment data

[1069] What it does: Stores data securely in a database.

[1070] Output: Summary data and sentiment data stored in a database

[1071] Step 7:

[1072] The server generates and suggests action items.

[1073] Input: Summary data and sentiment data

[1074] How it works: You input a prompt into the generative AI model, which will then use past data to suggest best practices. For example, "Please provide a summary of this meeting." "Please use emotion recognition to summarize the issues."

[1075] Output: Suggested action items and best practices (e.g., "Schedule an emergency product defect resolution meeting")

[1076] Step 8:

[1077] The user reviews the proposed action items and creates an implementation plan.

[1078] Input: Action items and best practices notified by the server

[1079] How it works: The user sees specific action items on their device and creates an action plan based on them.

[1080] Output: A concrete action plan (e.g., "Schedule an emergency meeting on ____ day to discuss problem-solving methods")

[1081] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1082] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1083] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1084] [Third embodiment]

[1085] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1086] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1087] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1088] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1089] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1090] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1091] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1092] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1093] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1094] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1095] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1096] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1097] The present invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, mainly including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices. The following describes in detail the embodiments of the present invention.

[1098] 1. System Configuration

[1099] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[1100] 2. Description of main functions

[1101] Voice recording and data transmission

[1102] User starts recording:

[1103] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[1104] The device sends the audio data to the server:

[1105] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[1106] Converting audio data to text

[1107] Server receives audio data:

[1108] The server receives the voice data sent from the terminal and stores it securely in data storage.

[1109] Audio to text conversion:

[1110] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[1111] Text summary generation

[1112] The server parses the text data:

[1113] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[1114] Summary generation:

[1115] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[1116] Data analysis and action item generation

[1117] Summary data analysis:

[1118] The server aggregates multiple summaries to identify common issues and trends.

[1119] Generate action items:

[1120] The server uses past meeting data and success stories to generate action items for the issues, such as "reviewing tasks" and "adjusting additional resources."

[1121] Proposal of know-how and best practices

[1122] Search our knowledge base:

[1123] The server searches and extracts relevant know-how and best practices from the knowledge base.

[1124] Generate a list of suggestions:

[1125] The server organizes the search results and creates a specific list of suggestions for the user.

[1126] Notification and confirmation

[1127] Send notifications:

[1128] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[1129] User sees action item:

[1130] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[1131] Specific examples

[1132] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[1133] 1. Start recording:

[1134] Person A presses the recording button on the device to start the meeting.

[1135] 2. Sending audio data to the server:

[1136] During the meeting, the device transmits audio data to the server in real time.

[1137] 3. Text:

[1138] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[1139] 4. Summary generation:

[1140] Using an NLP algorithm, summarize it as "Project Y problem."

[1141] 5. Data analysis and action generation:

[1142] The server analyzes similar past data and generates action items such as "set up a meeting to identify problems."

[1143] 6. Know-how proposal:

[1144] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[1145] 7. User confirms:

[1146] Persons A and B check this information on their devices and create an action plan.

[1147] In this way, the system can efficiently manage the content of one-on-one meetings and link them to concrete actions.

[1148] The processing flow will be explained below.

[1149] Step 1:

[1150] User starts recording:

[1151] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[1152] Step 2:

[1153] The device captures audio data:

[1154] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[1155] Step 3:

[1156] The device sends the audio data to the server:

[1157] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[1158] Step 4:

[1159] Server receives audio data:

[1160] The server receives the voice data sent from the terminal and stores the data in storage.

[1161] Step 5:

[1162] The server converts the audio data to text:

[1163] The server uses a speech recognition API to convert the voice data into text data in real time or in batches, and the converted text data is stored in a database.

[1164] Step 6:

[1165] The server summarizes the text data:

[1166] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[1167] Step 7:

[1168] The server stores the summary data:

[1169] The summarized text data is stored in a database and used for subsequent analysis and statistical processing.

[1170] Step 8:

[1171] The server analyzes the summary data:

[1172] The server aggregates multiple summaries of data and uses an analytics engine to identify issues and trends.

[1173] Step 9:

[1174] Server generates action item:

[1175] Based on the analysis results, the server generates specific action items, referencing past history and best practices.

[1176] Step 10:

[1177] Server searches knowledge base:

[1178] The server searches databases and knowledge bases to extract relevant know-how and best practices.

[1179] Step 11:

[1180] Server generates suggestion list:

[1181] Based on the extracted information, a list of specific action suggestions and best practices is generated to provide to the user.

[1182] Step 12:

[1183] Server sends notification:

[1184] Generated action items and suggestion lists are notified to the user's device, including via pop-up notifications and email notifications.

[1185] Step 13:

[1186] User sees action item:

[1187] The user uses the device to check the notification and view the generated action items and suggested know-how.

[1188] Step 14:

[1189] User performs an action:

[1190] The user then executes specific tasks based on the action items they have confirmed, such as setting up a project meeting or arranging for additional resources.

[1191] Example 1

[1192] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1193] Conventional conference management systems have had difficulty efficiently managing conference content and linking it to specific actions. Furthermore, converting conference recordings into text, summarizing them, and generating action items require a lot of manual work, which takes time and effort. Therefore, there is a demand for a system that can achieve efficient conference management and link it to specific actions.

[1194] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1195] In this invention, the server includes a means for [starting recording of the meeting], a means for [sending the recorded data to the server], a means for [the server converting the recorded data into text data using speech recognition technology], a means for [the server summarizing the text data using a natural language processing algorithm], a means for [the server analyzing the summarized data using a data analysis algorithm], and a means for [the server generating and proposing action items based on past success stories]. This makes it possible to convert the contents of the meeting into text and summarize it in real time, and automatically generate and propose specific action items.

[1196] The "means for starting recording of a conference" refers to the device or software that a user uses to record a conference, specifically a recording button.

[1197] "Means for transmitting recorded data to a server" refers to a function or device that transmits voice data recorded by a user to a server in real time or in batch mode.

[1198] "Means by which the server converts recorded data into text data using speech recognition technology" refers to the functions and processes by which the server converts audio data into text data using a speech recognition API or algorithm.

[1199] "Means by which the server summarizes text data using a natural language processing algorithm" refers to the function or algorithm by which the server uses natural language processing technology to extract important information from text data and summarize it concisely.

[1200] "Means for the server to analyze the summarized data using data analysis algorithms" means the data analysis algorithms used by the server to analyze the summarized text data and identify recurring issues or trends.

[1201] "Means for the server to generate and suggest action items based on past success stories" refers to the process or function in which the server refers to past meeting data and best practices, generates specific action items for specific issues, and suggests them to the user.

[1202] "Means for storing abstract data in a database" refers to the functions and processes for securely storing abstract data generated by the server in a database.

[1203] "Means for the server to extract relevant knowledge and best practices from the knowledge base and generate a list of suggestions" refers to the process or function by which the server searches an existing knowledge base, extracts useful information and best practices that meet the user's needs, and provides them as a list.

[1204] This invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices.

[1205] System Configuration

[1206] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[1207] Hardware and Software

[1208] 1. Device: The smartphone or computer used by the user.

[1209] 2. Server: Cloud server that processes data.

[1210] 3. Database: A database for storing meeting data, analysis results, etc.

[1211] 4. Speech recognition APIs: Google Speech-to-Text, Amazon Transcribe, etc.

[1212] 5. Natural Language Processing (NLP) algorithms: Open source libraries such as SpaCy and NLTK.

[1213] System functions and examples

[1214] Voice recording and data transmission

[1215] User starts recording:

[1216] The user opens the dedicated app on their device and presses the record button. By pressing the record button, the app activates the device's microphone and begins capturing audio. For example, employee A uses the smartphone app to record a one-on-one meeting.

[1217] The device sends the audio data to the server:

[1218] The device uploads the captured audio data to the server in regular segments in real time. For example, the audio data is divided into segments every 10 seconds and sent to the server sequentially.

[1219] Converting audio data to text

[1220] Server receives audio data:

[1221] The server receives the voice data sent from the device and stores it securely in data storage, for example, as an audio file in cloud storage.

[1222] The server converts the audio data to text:

[1223] The server converts the received voice data into text using a speech recognition API (e.g., Google Speech-to-Text). For example, a speech saying "Project X's deadline is behind schedule" is obtained as text data.

[1224] Text summary generation

[1225] The server parses the text data:

[1226] The server applies NLP algorithms (e.g., SpaCy) to extract important keywords and context from the text data.

[1227] Server generates summary:

[1228] The server summarizes the original text based on the extracted information. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[1229] Data analysis

[1230] The server aggregates the summary data:

[1231] The server aggregates multiple summaries generated from past meetings and identifies common issues and trends. For example, "late delivery" emerges as a common issue across multiple projects.

[1232] Generate action items

[1233] Server generates action item:

[1234] The server generates specific action items for issues based on past success stories and knowledge bases. For example, it suggests "setting up an emergency meeting" or "arranging additional resources" as a way to address delivery delays.

[1235] Proposal of know-how and best practices

[1236] Server searches for know-how:

[1237] The server searches the knowledge base for relevant know-how and best practices, for example, "project management best practices."

[1238] Server generates suggestion list:

[1239] The server creates a list of suggestions based on the search results and provides it to the user. Specifically, it displays a list of past success stories and procedures based on those stories.

[1240] Notification and confirmation

[1241] Server sends notification:

[1242] The server notifies the user of the generated action items and suggestion list via a pop-up notification on their smartphone or email.

[1243] User sees action item:

[1244] Users check notifications on their devices and view suggested action items and know-how. They then plan specific actions and put them into action. For example, Person A and Person B discuss specific measures based on the information displayed on their devices and create an action plan.

[1245] Prompt Sentence Examples

[1246] "Please summarize the content of the 1-on-1 meeting and propose specific action items and know-how."

[1247] "Convert the recorded audio data into text and extract the key points."

[1248] This system allows you to efficiently manage the content of one-on-one meetings and link them to concrete actions.

[1249] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1250] Step 1:

[1251] User starts recording

[1252] The user opens the dedicated app on their device and presses the record button. The input is the user clicking the record button, and the device starts capturing audio. Specifically, the smartphone's microphone is activated and audio data is captured in real time. The output is the audio data being recorded.

[1253] Step 2:

[1254] The device sends the voice data to the server

[1255] The device uploads captured audio data to the server in regular segments in real time. The input is the audio data being recorded, divided into segments, such as 10 seconds. Specifically, the device sends the audio data segments to the server using HTTP requests. The output is the audio data stored on the server.

[1256] Step 3:

[1257] The server receives the audio data

[1258] The server receives the voice data sent from the terminal and safely stores it in the data storage. The input is the voice data sent from the terminal, and the output is the voice file stored in the data storage. Specifically, the server executes a process to store the voice file in a database.

[1259] Step 4:

[1260] The server converts the voice data into text

[1261] The server uses a speech recognition API (e.g., Google Speech-to-Text) to convert the voice data into text. The input is a saved audio file, and when the server calls the API, the API analyzes the voice data and returns it as text data. The output is text data. For example, the text generated might say, "The delivery date for Project X is behind schedule."

[1262] Step 5:

[1263] The server analyzes the text data

[1264] The server applies an NLP algorithm (e.g., SpaCy) to extract important keywords and context from the text data. The input is text data, which the server analyzes by running the NLP algorithm. The output is important keywords and context data. For example, the important keyword "delayed delivery" is extracted.

[1265] Step 6:

[1266] Server generates summary

[1267] The server summarizes the original text based on the extracted information. The input is important keywords and contextual data, and the server generates a summary by running a summary generation algorithm. The output is summarized text data. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[1268] Step 7:

[1269] The server aggregates the summary data

[1270] The server aggregates multiple summaries generated from past meetings and identifies frequently occurring issues and trends. The input is multiple summarized text data, which the server analyzes by applying a data aggregation algorithm. The output is the identification of issues and trends. For example, "delayed delivery" is identified as a common issue across multiple projects.

[1271] Step 8:

[1272] Server generates action items

[1273] The server generates specific action items for issues based on past success stories and a knowledge base. The input is the results of identifying issues and trends, and the server generates specific action items by running an action item generation algorithm. The output is a specific action item. For example, it may suggest measures such as "setting up an emergency meeting" or "adjusting additional resources" to address delivery delays.

[1274] Step 9:

[1275] The server searches for know-how

[1276] The server searches the knowledge base to find relevant know-how and best practices. The input is summary data and action items, and the server runs a knowledge base search algorithm to extract relevant information. The output is know-how and best practice information.

[1277] Step 10:

[1278] The server generates a list of suggestions

[1279] The server creates a proposal list based on the search results and provides it to the user. The input is know-how and best practice information, and the server generates the list by running a proposal list generation algorithm. The output is the proposal list provided to the user.

[1280] Step 11:

[1281] The server sends a notification

[1282] The server notifies the user's terminal of the generated action items and suggestion list. The input is the action items and suggestion list, and the server executes the notification process to send the notification to the user. The output is the notification displayed on the user's terminal.

[1283] Step 12:

[1284] User confirms action item

[1285] The user receives a notification on their device and checks the proposed action items and know-how. The input is the notification displayed on the user's device, and the specific action is confirmed by the user checking the operation. The output is the creation of an action plan. For example, Person A and Person B discuss specific measures and create an action plan.

[1286] (Application example 1)

[1287] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1288] There is a need to efficiently manage the content of one-on-one meetings held in the field and link it to specific action items and best practices. However, the current situation is such that the content of meetings is not fully utilized, resulting in a waste of time and resources. In addition, it is a heavy burden for administrators to continue to manage the content of each meeting individually. For this reason, there is a need for a more efficient and automated system.

[1289] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1290] In this invention, the server includes means for starting recording of the meeting, means for transmitting the recorded data to the server, means for the server to recognize the voice of the recorded data and convert it into text, means for the server to summarize the text data, means for the server to analyze the summarized data, means for the server to generate and propose action items, means for automatically generating a proposal list using a generative AI model, and means for notifying the action items and proposal list. This makes it possible to efficiently manage the content of one-on-one meetings and link them to specific actions.

[1291] "Means for starting recording of a meeting" refers to a device or function that allows a user to record the audio of a meeting by pressing a button or other operation.

[1292] "Means for transmitting recorded data to a server" refers to communication functions or programs for uploading recorded audio data to a server in real time or by batch processing.

[1293] "Means for the server to recognize the recorded data and convert it into text" refers to the process of analyzing the voice data on the server and converting it into text data using voice recognition technology.

[1294] "Means for the server to summarize text data" refers to the function of the server using natural language processing (NLP) algorithms, etc. to extract important information from text data and summarize it concisely.

[1295] "Means for the server to analyze the summary data" refers to a data analysis function that analyzes the summary data generated by the server and identifies significant keywords and patterns.

[1296] "Means for the server to generate and suggest action items" refers to the function by which the server automatically generates specific actions and suggestions that the user should take based on the analysis results.

[1297] "Means for automatically generating a suggestion list using a generative AI model" refers to an algorithm or program for automatically creating a suggestion list for a user using a generative AI model.

[1298] "Means for notifying action items and suggestion lists" refers to the system's functionality for notifying the user's terminal of generated action items and suggestion lists.

[1299] The present invention provides a system for efficiently managing the content of one-on-one meetings and linking them to specific actions. The following describes in detail the embodiments of the present invention.

[1300] System Configuration

[1301] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[1302] Voice recording and data transmission

[1303] Yu

[1304] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1305] Step 1:

[1306] When a user starts a one-on-one meeting, they press the recording button on their device, and the device begins capturing audio.

[1307] Input: User presses record button.

[1308] Output: Recorded audio data.

[1309] Specific operation: Captures audio data using the device's microphone function.

[1310] Step 2:

[1311] The device transmits the recorded audio data to the server in real time.

[1312] Input: Pre-recorded audio data.

[1313] Output: The audio data sent to the server.

[1314] Specific operation: Audio data is divided and uploaded to the server via data communication.

[1315] Step 3:

[1316] The server receives the recording and stores it securely.

[1317] Input: Audio data sent from the device.

[1318] Output: Audio data stored on the server.

[1319] What it does: Receives audio data and stores it securely in a database or storage.

[1320] Step 4:

[1321] The server converts the recorded data into text data using voice recognition technology.

[1322] Input: Recorded audio data.

[1323] Output: Text data.

[1324] Specific operation: Uses a speech recognition API to convert voice data into text.

[1325] Step 5:

[1326] The server applies natural language processing (NLP) algorithms to summarize the text data.

[1327] Input: Text data.

[1328] Output: Summarized text data.

[1329] Specific operation: Using NLP algorithms, important keywords and context are extracted from text data and a summary is generated.

[1330] Step 6:

[1331] The server analyzes the summary data and generates action items based thereon.

[1332] Input: Summarized text data.

[1333] Output: Action items.

[1334] Specific Behavior: Analyzes summary data, identifies issues and trends, and generates corresponding action items.

[1335] Step 7:

[1336] The server uses a generative AI model to automatically generate a list of suggestions including relevant know-how and best practices.

[1337] Inputs: Action items and knowledge base data.

[1338] Output: A list of suggestions.

[1339] What it does: It applies generative AI models to extract useful know-how and best practices from a knowledge base and generate a list.

[1340] Step 8:

[1341] The server notifies the user's terminal of the action items and the suggestion list.

[1342] Input: Action items and suggestion lists.

[1343] Output: Notification displayed on the user's device.

[1344] Specific behavior: Using the notification system, the generated action items and suggestion list are sent to the user's device via push notification or email.

[1345] For example, if a factory manager discusses a problem with a new production line during a one-on-one meeting with a field staff member, the server will summarize the content and suggest specific action items such as setting up a trouble-shooting team meeting. In this way, the system can efficiently manage the content of meetings and lead to actual actions.

[1346] For example, an example of a prompt for a generative AI model is: "Please summarize the following text: There is a problem with the new production line. Specifically, there is a problem with the machines frequently stopping. The cause has not yet been identified, but manual intervention may be necessary."

[1347] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1348] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[1349] 1. System Configuration

[1350] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[1351] 2. Description of main functions

[1352] Voice recording and data transmission

[1353] User starts recording:

[1354] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[1355] The device sends the audio data to the server:

[1356] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[1357] Converting audio data to text

[1358] Server receives audio data:

[1359] The server receives the voice data sent from the terminal and stores it securely in data storage.

[1360] Audio to text conversion:

[1361] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[1362] Text summary generation

[1363] The server parses the text data:

[1364] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[1365] Summary generation:

[1366] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[1367] emotion recognition

[1368] The server uses an emotion engine to recognize the user's emotions:

[1369] The server analyzes the user's emotions from the voice and text data, including algorithms that identify emotions from voice tone and text content.

[1370] Data storage and analysis

[1371] Summary and sentiment data storage:

[1372] The server stores the summarized text data and the sentiment analysis results in a database.

[1373] Data Analysis:

[1374] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[1375] Action item generation and suggestions

[1376] Server generates action item:

[1377] The server generates specific action items based on the analysis results and sentiment data, referring to past history and best practices.

[1378] Search our knowledge base:

[1379] The server searches and extracts relevant know-how and best practices from the knowledge base.

[1380] Generate a list of suggestions:

[1381] The server generates specific action suggestions and a list of best practices for the user based on the search results.

[1382] Notification and confirmation

[1383] Send notifications:

[1384] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[1385] User sees action item:

[1386] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[1387] Specific examples

[1388] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[1389] 1. Start recording:

[1390] Person A presses the recording button on the device to start the meeting.

[1391] 2. Sending audio data to the server:

[1392] During the meeting, the device transmits audio data to the server in real time.

[1393] 3. Text:

[1394] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[1395] 4. Summary generation:

[1396] Using an NLP algorithm, summarize it as "Project Y problem."

[1397] 5. Emotion recognition:

[1398] The emotion engine recognizes from the tone of the voice and the content of the text that Person A is feeling stressed about the problem.

[1399] 6. Data Retention:

[1400] The server stores the text data and emotion data in a database.

[1401] 7. Data analysis and action generation:

[1402] The server analyzes similar past data and sentiment data and generates action items such as "set up a meeting to identify problems."

[1403] 8. Know-how proposal:

[1404] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[1405] 9. User confirms:

[1406] Persons A and B check this information on their devices and create an action plan.

[1407] In this way, by combining an emotion recognition engine, it is possible to provide a management system that deepens the content of one-on-one meetings and leads to concrete actions.

[1408] The processing flow will be explained below.

[1409] Step 1:

[1410] User starts recording:

[1411] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[1412] Step 2:

[1413] The device captures audio data:

[1414] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[1415] Step 3:

[1416] The device sends the audio data to the server:

[1417] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[1418] Step 4:

[1419] Server receives audio data:

[1420] The server receives the voice data sent from the device and stores it in a secure data storage.

[1421] Step 5:

[1422] The server converts the audio data to text:

[1423] The server calls the speech recognition API to convert the voice data into text data, which is then stored in a database.

[1424] Step 6:

[1425] The server summarizes the text data:

[1426] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[1427] Step 7:

[1428] The server performs emotion recognition using the emotion engine:

[1429] The server uses an emotion engine to analyze the user's emotions from the voice and text data, using algorithms that identify emotions from voice tone and text content.

[1430] Step 8:

[1431] The server stores the emotion recognition results and summary data:

[1432] The server stores the generated summary data and emotion recognition results in a database for later analysis and statistical processing.

[1433] Step 9:

[1434] The server analyzes the summary data and sentiment data:

[1435] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[1436] Step 10:

[1437] Server generates action item:

[1438] The server generates specific action items based on the analysis results and sentiment data, referencing past history and best practices.

[1439] Step 11:

[1440] Server searches knowledge base:

[1441] The server searches and extracts relevant know-how and best practices from the knowledge base.

[1442] Step 12:

[1443] Server generates suggestion list:

[1444] Based on the extracted information, specific action suggestions and a list of best practices are generated for the user.

[1445] Step 13:

[1446] Server sends notification:

[1447] Generated action items and suggestion lists are notified to the user's device in the form of a popup or email.

[1448] Step 14:

[1449] User sees action item:

[1450] The user uses the device to check the notification and view the generated action items and suggested know-how.

[1451] Step 15:

[1452] User performs an action:

[1453] Based on the action items confirmed by the user, the system performs specific tasks, such as setting up a project meeting or coordinating additional resources.

[1454] Example 2

[1455] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1456] In traditional one-on-one meetings, simply recording the conversation and then manually converting it into text and summarizing it later takes time and effort. Nuances such as the speaker's emotions and tone during the meeting are not recorded, often resulting in insufficient understanding and analysis. Furthermore, opportunities for business improvement tend to be missed because specific action items and effective best practices are not generated.

[1457] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for starting recording of the conference, means for transmitting recorded data to a computer, means for the computer to perform speech recognition on the recorded data and convert it into text data, means for the computer to summarize the text data, means for the computer to analyze the summary data and emotion data, means for the computer to generate and suggest action items, and means for recognizing the user's emotions using an emotion recognition engine. This makes it possible to effectively record and analyze the contents of the conference, obtain a detailed understanding including emotional nuances, and suggest specific improvement actions.

[1458] "Means for starting a meeting recording" refers to a feature that allows users to record one-on-one meetings and other meetings on their devices.

[1459] The "means for transmitting recorded data to a computer" is a function for uploading audio data recorded by a terminal to a server in real time or in batch processing.

[1460] "Means for a computer to recognize the recorded data and convert it into text data" refers to the process by which a server uses a speech recognition API to convert the voice data into text data.

[1461] "Means for computer-generated summarization of text data" refers to the process in which the server uses a natural language processing algorithm to extract important information from the converted text data and summarize it concisely.

[1462] "Means for a computer to analyze summary data and emotion data" refers to a function in which the server uses the stored summary and emotion data to analyze past history and patterns and identify trends and issues.

[1463] "Means for computer-generated and proposed action items" refers to a function in which the server generates specific action items based on the results of data analysis and best practices, and proposes them to the user.

[1464] "Means for recognizing the user's emotions using an emotion recognition engine" refers to the process in which the server analyzes the user's emotions from voice data and text data and records them as emotion data.

[1465] "Storing the summary data and emotion data in an information storage device" refers to the process in which the server safely stores the generated summary text data and emotion data in storage such as a database.

[1466] "Extracting know-how and best practices from the knowledge base and generating a proposal list" refers to the process in which the server searches the knowledge base, extracts relevant information, and creates a specific proposal list for the user.

[1467] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[1468] System Configuration

[1469] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[1470] Main Features

[1471] Voice recording and data transmission

[1472] When a user presses the record button on the device, the device begins capturing audio using the built-in microphone. The device then sends the captured audio data to the server at regular intervals (for example, every 10 seconds). When the device sends the recorded data to the server, it uses Wi-Fi or mobile data communication and sends the audio data via an HTTP POST request.

[1473] Converting audio data to text

[1474] The server receives the voice data sent from the device and stores it in a secure data storage (e.g., Amazon S3).The server then converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).Through this conversion process, all conversation content is obtained as text data.

[1475] Text summary generation

[1476] The server uses natural language processing (NLP) algorithms to extract important keywords and context from the acquired text data and summarize it concisely. For example, the server analyzes the text data using an NLP library such as the Natural Language Toolkit (NLTK) and generates a summary.

[1477] emotion recognition

[1478] The server uses an emotion recognition engine to analyze the user's emotions from the voice data and text data. To determine emotions from the tone of the voice or specific words, the server sends the voice data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to obtain emotion data.

[1479] Data storage and analysis

[1480] The server generates summary text data and stores the sentiment data in a database (e.g., MySQL or PostgreSQL). The server uses the stored data to analyze past history and patterns and identify issues and trends. When analyzing the data, the server uses Structured Query Language (SQL).

[1481] Action item generation and suggestions

[1482] The server generates specific action items based on the analysis results and sentiment data. In addition, the server searches a knowledge base to extract relevant information and generate a list of specific suggestions for the user. The server uses a rule-based engine and machine learning algorithms to generate appropriate action items and match them with the knowledge base.

[1483] Notification and confirmation

[1484] The server notifies the user device of the generated action items and suggestion list. Notification methods include pop-up notifications, emails, and in-app notifications. The server sends real-time notifications to the user's smartphone using services such as Firebase Cloud Messaging (FCM). The user receives the notification on their device, checks the specific action items and suggested know-how, and can then take the next step.

[1485] Specific examples

[1486] For example, consider the case where employee A has a one-on-one meeting with line manager B. A presses the record button on their device to begin the meeting. During the meeting, the device sends audio data to the server in real time. The server uses a speech recognition API to convert the audio data into text: "We're experiencing problems with project Y." The server then uses an NLP algorithm to summarize the text as "Problem with project Y." The server's emotion engine recognizes from the tone of the voice and the content of the text that A is feeling stressed about the problem. The server stores the text data and emotion data in a database, analyzes similar past data and emotion data, and generates action items such as "Arrange a meeting to identify problems." The server extracts and suggests "problem-solving methods that were previously effective on project Y" from its knowledge base. Finally, A and B review this information on their devices and create an action plan.

[1487] Examples of prompt statements

[1488] "Generate a prompt sentence to describe the above 1-on-1 meeting management system."

[1489] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1490] A detailed explanation of the system program processing flow and processing steps

[1491] Step 1:

[1492] User starts recording

[1493] Input: A user starts a 1-on-1 meeting.

[1494] How it works: When a user presses the record button on their device, the device's built-in microphone begins capturing audio.

[1495] Output: Audio data is generated and captured.

[1496] Step 2:

[1497] The device sends the voice data to the server

[1498] Input: Audio data captured by the device.

[1499] How it works: Audio data is sent to the server at regular intervals (e.g., every 10 seconds). The audio data is sent via an HTTP POST request using Wi-Fi or mobile data.

[1500] Output: Audio data uploaded to the server.

[1501] Step 3:

[1502] The server receives and stores the audio data

[1503] Input: The server receives the audio data sent from the device.

[1504] How it works: The server receives the audio data and stores it in secure data storage (e.g. Amazon S3).

[1505] Output: Saved audio data.

[1506] Step 4:

[1507] The server converts the voice data into text

[1508] Input: Stored audio data.

[1509] How it works: The server uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the audio data into text, sends the audio file to the API, and stores the returned text.

[1510] Output: The generated text data.

[1511] Step 5:

[1512] The server summarizes the text data

[1513] Input: Generated text data.

[1514] How it works: The server uses NLP algorithms (e.g., the NLTK library) to analyze the text data, extract important keywords and context, and generate a summary.

[1515] Output: The generated summary.

[1516] Step 6:

[1517] The server recognizes emotions

[1518] Input: Audio and text data.

[1519] How it works: The server uses an emotion recognition API (e.g. IBM Watson Tone Analyzer) to extract emotion data from voice tone and text content.

[1520] Output: Emotion data.

[1521] Step 7:

[1522] The server saves and analyzes the summary and emotion data

[1523] Input: Summary sentences and sentiment data.

[1524] How it works: The server saves the summary sentences and emotion data in a database and analyzes past history and patterns. SQL is used to save the data in the database and perform the analysis.

[1525] Output: Analysis of issues and trends.

[1526] Step 8:

[1527] Server generates and suggests action items

[1528] Input: Analysis results and sentiment data.

[1529] How it works: The server generates specific action items using a rule-based engine and machine learning algorithms, extracts relevant information from a knowledge base, and generates a list of suggestions.

[1530] Output: Generated action items and suggestion list.

[1531] Step 9:

[1532] The server notifies the user

[1533] Input: Generated list of action items and suggestions.

[1534] How it works: The server uses Firebase Cloud Messaging (FCM) or similar to send real-time notifications to the user's smartphone.

[1535] Output: A notification message to the user.

[1536] Step 10:

[1537] User confirms action item

[1538] Input: Notification message.

[1539] Action: The user receives a notification on their device, sees an action item, or suggests a know-how, and takes specific steps based on that.

[1540] Output: User confirmed action items and implementation plans.

[1541] (Application example 2)

[1542] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1543] Current meeting recording and data analysis systems are limited to converting audio data into text, providing a brief summary, and suggesting action items. As a result, they lack insight into the emotional state of participants during meetings and specific, real-time countermeasures based on that information. Furthermore, because they are unable to process data in real time, they lack the ability to respond to emergencies and provide immediate response.

[1544] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1545] In this invention, the server includes means for [analyzing summary data and text data using an emotion recognition engine], means for [acquiring emotion data in real time and proposing specific action items such as setting up an emergency response meeting based on the analysis results], and means for [converting audio data into text in real time while recording a meeting, performing emotion analysis, and immediately reflecting the results]. This makes it possible [to grasp the emotional state of participants in a meeting in real time and immediately propose specific measures].

[1546] The "means for starting recording of a conference" is a function that a user operates to record the audio of a conference or meeting.

[1547] The "means for transmitting recorded data to a server" is a function for uploading recorded voice data to a server in real time or later.

[1548] "Means for the server to recognize the voice of the recorded data and convert it into text" refers to a function that enables the server to use voice recognition technology to convert the recorded voice data into text data.

[1549] The "means for the server to summarize text data" is a function for the server to summarize the text data acquired by the server using a natural language processing algorithm and extract important points.

[1550] The "means for the server to analyze the summarized data and text data using an emotion recognition engine" is a function that enables the server to analyze summarized text data and unsummarized text data using an emotion recognition algorithm to identify the emotional states of conference participants.

[1551] "Means for the server to generate and suggest action items" is a function that allows the server to generate specific action items based on the analysis results and suggest them to the user.

[1552] "A means of acquiring emotional data in real time and proposing specific action items such as setting up emergency response meetings based on the analysis results" is a function in which the server acquires emotional data in real time during a meeting and quickly reflects the analysis results to propose actions such as setting up emergency response meetings.

[1553] "Means for converting audio data into text in real time while recording a meeting, performing sentiment analysis, and immediately displaying the results" refers to a function that enables a robot or system to convert the audio data of a meeting into text in real time while recording it, perform sentiment analysis based on the text data, and immediately present the results to the user.

[1554] The "means for saving summary data and emotion data in a database" is a function that enables the server to safely save summary data and emotion recognition results in a database.

[1555] "Means of inputting a prompt sentence into a generative AI model and proposing optimal best practices from past data" is a function that enables the server to input a specific prompt (input sentence) into the generative AI model, search for optimal best practices from past data, and propose them.

[1556] This invention is a system for making work improvement meetings in factories more efficient, and has the functions of converting voice data into text, recognizing emotions, conducting real-time analysis, and generating and proposing action items.

[1557] System Configuration

[1558] The system includes terminals used in factories, a server for processing recorded data, a database, and an emotion recognition engine. The terminals are primarily used for recording meetings, while the server performs speech recognition and analysis. The database stores text data and emotion data. The emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state.

[1559] Program processing

[1560] The system begins when a user presses a recording button on their device to start a meeting. The recorded audio data is sent in real time to a server, which then converts it into text using speech recognition technology. The server then analyzes the text data using a natural language processing algorithm to generate a summary of the meeting. An emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state and reflects the results in real time.

[1561] Novel Features

[1562] 1. Real-time emotional data acquisition and analysis: The server acquires emotional data in real time during the meeting and, based on the analysis results, proposes specific action items such as setting up an emergency response meeting.

[1563] 2. Instant text conversion and sentiment analysis during meeting recording: Audio data is converted into text in real time during meeting recording, and the text data is instantly analyzed for sentiment.

[1564] 3. Use of generative AI models: By inputting prompt statements into generative AI models, the models can suggest optimal best practices based on past data.

[1565] Specific usage

[1566] The device is equipped with a microphone (e.g., a USB microphone) and captures voice data. The recorded data is sent to a server, where it is converted into text data using the Google Speech Recognition API. The text data is then summarized using a natural language processing algorithm (e.g., BERT), and sentiment analysis is performed using an emotion recognition engine. The analysis results are stored in a database and notified to the user in real time. Examples of prompts include "Please print a summary of this meeting" and "Perform emotion recognition and summarize the issues."

[1567] Specific examples

[1568] For example, when factory worker A holds a one-on-one meeting with manager B, he presses the recording button on his device to start the meeting. During the meeting, audio data is sent to the server in real time and instantly converted into text data. The server analyzes the generated text data, and if worker A says, "There has been an increase in product defects this month," it generates a summary saying, "Investigate the cause of product defects," and recognizes through sentiment analysis that worker A is frustrated. As a result, the server suggests "Set up an emergency meeting to address product defects" and generates specific action items.

[1569] This system makes it possible to identify problems in real time during meetings and respond immediately, which is expected to lead to more efficient work improvements within the factory.

[1570] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1571] Step 1:

[1572] The user presses the record button on the device to start the conference.

[1573] Input: User actions

[1574] How it works: By pressing the record button, the device will begin recording the audio data during the meeting.

[1575] Output: Audio data begins to accumulate on the device.

[1576] Step 2:

[1577] The device sends the recorded data to the server in real time.

[1578] Input: Recorded audio data

[1579] How it works: Recorded data is uploaded to the server in real time, segment by segment.

[1580] Output: Audio data sent to the server

[1581] Step 3:

[1582] The server converts the voice data into text data using a voice recognition API.

[1583] Input: Audio data sent to the server

[1584] How it works: Uses the Google Speech Recognition API to convert audio data into text.

[1585] Output: Text data (e.g., "Product defects are increasing this month.")

[1586] Step 4:

[1587] The server summarizes the text data using natural language processing algorithms.

[1588] Input: Text data converted by speech recognition

[1589] How it works: It uses natural language processing algorithms such as BERT to extract key points from text data and generate summaries.

[1590] Output: Summary data (e.g., "Investigation into the cause of product defects")

[1591] Step 5:

[1592] The server analyzes the summary data and text data using an emotion recognition engine.

[1593] Input: Abstract and text data

[1594] How it works: An emotion recognition engine is used to recognize and analyze user emotions from text data and voice tones.

[1595] Output: Emotion data (e.g., stress and irritation are recognized)

[1596] Step 6:

[1597] The server stores the summary data and the emotion data in a database.

[1598] Input: Summary data and sentiment data

[1599] What it does: Stores data securely in a database.

[1600] Output: Summary data and sentiment data stored in a database

[1601] Step 7:

[1602] The server generates and suggests action items.

[1603] Input: Summary data and sentiment data

[1604] How it works: You input a prompt into the generative AI model, which will then use past data to suggest best practices. For example, "Please provide a summary of this meeting." "Please use emotion recognition to summarize the issues."

[1605] Output: Suggested action items and best practices (e.g., "Schedule an emergency product defect resolution meeting")

[1606] Step 8:

[1607] The user reviews the proposed action items and creates an implementation plan.

[1608] Input: Action items and best practices notified by the server

[1609] How it works: The user sees specific action items on their device and creates an action plan based on them.

[1610] Output: A concrete action plan (e.g., "Schedule an emergency meeting on ____ day to discuss problem-solving methods")

[1611] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1612] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1613] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1614] [Fourth embodiment]

[1615] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1616] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1617] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1618] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1619] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1620] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1621] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1622] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1623] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1624] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1625] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1626] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1627] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1628] The present invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, mainly including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices. The following describes in detail the embodiments of the present invention.

[1629] 1. System Configuration

[1630] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[1631] 2. Description of main functions

[1632] Voice recording and data transmission

[1633] User starts recording:

[1634] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[1635] The device sends the audio data to the server:

[1636] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[1637] Converting audio data to text

[1638] Server receives audio data:

[1639] The server receives the voice data sent from the terminal and stores it securely in data storage.

[1640] Audio to text conversion:

[1641] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[1642] Text summary generation

[1643] The server parses the text data:

[1644] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[1645] Summary generation:

[1646] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[1647] Data analysis and action item generation

[1648] Summary data analysis:

[1649] The server aggregates multiple summaries to identify common issues and trends.

[1650] Generate action items:

[1651] The server uses past meeting data and success stories to generate action items for the issues, such as "reviewing tasks" and "adjusting additional resources."

[1652] Proposal of know-how and best practices

[1653] Search our knowledge base:

[1654] The server searches and extracts relevant know-how and best practices from the knowledge base.

[1655] Generate a list of suggestions:

[1656] The server organizes the search results and creates a specific list of suggestions for the user.

[1657] Notification and confirmation

[1658] Send notifications:

[1659] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[1660] User sees action item:

[1661] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[1662] Specific examples

[1663] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[1664] 1. Start recording:

[1665] Person A presses the recording button on the device to start the meeting.

[1666] 2. Sending audio data to the server:

[1667] During the meeting, the device transmits audio data to the server in real time.

[1668] 3. Text:

[1669] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[1670] 4. Summary generation:

[1671] Using an NLP algorithm, summarize it as "Project Y problem."

[1672] 5. Data analysis and action generation:

[1673] The server analyzes similar past data and generates action items such as "set up a meeting to identify problems."

[1674] 6. Know-how proposal:

[1675] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[1676] 7. User confirms:

[1677] Persons A and B check this information on their devices and create an action plan.

[1678] In this way, the system can efficiently manage the content of one-on-one meetings and link them to concrete actions.

[1679] The processing flow will be explained below.

[1680] Step 1:

[1681] User starts recording:

[1682] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[1683] Step 2:

[1684] The device captures audio data:

[1685] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[1686] Step 3:

[1687] The device sends the audio data to the server:

[1688] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[1689] Step 4:

[1690] Server receives audio data:

[1691] The server receives the voice data sent from the terminal and stores the data in storage.

[1692] Step 5:

[1693] The server converts the audio data to text:

[1694] The server uses a speech recognition API to convert the voice data into text data in real time or in batches, and the converted text data is stored in a database.

[1695] Step 6:

[1696] The server summarizes the text data:

[1697] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[1698] Step 7:

[1699] The server stores the summary data:

[1700] The summarized text data is stored in a database and used for subsequent analysis and statistical processing.

[1701] Step 8:

[1702] The server analyzes the summary data:

[1703] The server aggregates multiple summaries of data and uses an analytics engine to identify issues and trends.

[1704] Step 9:

[1705] Server generates action item:

[1706] Based on the analysis results, the server generates specific action items, referencing past history and best practices.

[1707] Step 10:

[1708] Server searches knowledge base:

[1709] The server searches databases and knowledge bases to extract relevant know-how and best practices.

[1710] Step 11:

[1711] Server generates suggestion list:

[1712] Based on the extracted information, a list of specific action suggestions and best practices is generated to provide to the user.

[1713] Step 12:

[1714] Server sends notification:

[1715] Generated action items and suggestion lists are notified to the user's device, including via pop-up notifications and email notifications.

[1716] Step 13:

[1717] User sees action item:

[1718] The user uses the device to check the notification and view the generated action items and suggested know-how.

[1719] Step 14:

[1720] User performs an action:

[1721] The user then executes specific tasks based on the action items they have confirmed, such as setting up a project meeting or arranging for additional resources.

[1722] Example 1

[1723] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1724] Conventional conference management systems have had difficulty efficiently managing conference content and linking it to specific actions. Furthermore, converting conference recordings into text, summarizing them, and generating action items require a lot of manual work, which takes time and effort. Therefore, there is a demand for a system that can achieve efficient conference management and link it to specific actions.

[1725] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1726] In this invention, the server includes a means for [starting recording of the meeting], a means for [sending the recorded data to the server], a means for [the server converting the recorded data into text data using speech recognition technology], a means for [the server summarizing the text data using a natural language processing algorithm], a means for [the server analyzing the summarized data using a data analysis algorithm], and a means for [the server generating and proposing action items based on past success stories]. This makes it possible to convert the contents of the meeting into text and summarize it in real time, and automatically generate and propose specific action items.

[1727] The "means for starting recording of a conference" refers to the device or software that a user uses to record a conference, specifically a recording button.

[1728] "Means for transmitting recorded data to a server" refers to a function or device that transmits voice data recorded by a user to a server in real time or in batch mode.

[1729] "Means by which the server converts recorded data into text data using speech recognition technology" refers to the functions and processes by which the server converts audio data into text data using a speech recognition API or algorithm.

[1730] "Means by which the server summarizes text data using a natural language processing algorithm" refers to the function or algorithm by which the server uses natural language processing technology to extract important information from text data and summarize it concisely.

[1731] "Means for the server to analyze the summarized data using data analysis algorithms" means the data analysis algorithms used by the server to analyze the summarized text data and identify recurring issues or trends.

[1732] "Means for the server to generate and suggest action items based on past success stories" refers to the process or function in which the server refers to past meeting data and best practices, generates specific action items for specific issues, and suggests them to the user.

[1733] "Means for storing abstract data in a database" refers to the functions and processes for securely storing abstract data generated by the server in a database.

[1734] "Means for the server to extract relevant knowledge and best practices from the knowledge base and generate a list of suggestions" refers to the process or function by which the server searches an existing knowledge base, extracts useful information and best practices that meet the user's needs, and provides them as a list.

[1735] This invention is a system for efficiently managing the content of one-on-one meetings and linking it to specific actions. This system has many functions, including recording, converting audio data into text, generating summaries, analyzing data, generating action items, and proposing know-how and best practices.

[1736] System Configuration

[1737] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[1738] Hardware and Software

[1739] 1. Device: The smartphone or computer used by the user.

[1740] 2. Server: Cloud server that processes data.

[1741] 3. Database: A database for storing meeting data, analysis results, etc.

[1742] 4. Speech recognition APIs: Google Speech-to-Text, Amazon Transcribe, etc.

[1743] 5. Natural Language Processing (NLP) algorithms: Open source libraries such as SpaCy and NLTK.

[1744] System functions and examples

[1745] Voice recording and data transmission

[1746] User starts recording:

[1747] The user opens the dedicated app on their device and presses the record button. By pressing the record button, the app activates the device's microphone and begins capturing audio. For example, employee A uses the smartphone app to record a one-on-one meeting.

[1748] The device sends the audio data to the server:

[1749] The device uploads the captured audio data to the server in regular segments in real time. For example, the audio data is divided into segments every 10 seconds and sent to the server sequentially.

[1750] Converting audio data to text

[1751] Server receives audio data:

[1752] The server receives the voice data sent from the device and stores it securely in data storage, for example, as an audio file in cloud storage.

[1753] The server converts the audio data to text:

[1754] The server converts the received voice data into text using a speech recognition API (e.g., Google Speech-to-Text). For example, a speech saying "Project X's deadline is behind schedule" is obtained as text data.

[1755] Text summary generation

[1756] The server parses the text data:

[1757] The server applies NLP algorithms (e.g., SpaCy) to extract important keywords and context from the text data.

[1758] Server generates summary:

[1759] The server summarizes the original text based on the extracted information. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[1760] Data analysis

[1761] The server aggregates the summary data:

[1762] The server aggregates multiple summaries generated from past meetings and identifies common issues and trends. For example, "late delivery" emerges as a common issue across multiple projects.

[1763] Generate action items

[1764] Server generates action item:

[1765] The server generates specific action items for issues based on past success stories and knowledge bases. For example, it suggests "setting up an emergency meeting" or "arranging additional resources" as a way to address delivery delays.

[1766] Proposal of know-how and best practices

[1767] Server searches for know-how:

[1768] The server searches the knowledge base for relevant know-how and best practices, for example, "project management best practices."

[1769] Server generates suggestion list:

[1770] The server creates a list of suggestions based on the search results and provides it to the user. Specifically, it displays a list of past success stories and procedures based on those stories.

[1771] Notification and confirmation

[1772] Server sends notification:

[1773] The server notifies the user of the generated action items and suggestion list via a pop-up notification on their smartphone or email.

[1774] User sees action item:

[1775] Users check notifications on their devices and view suggested action items and know-how. They then plan specific actions and put them into action. For example, Person A and Person B discuss specific measures based on the information displayed on their devices and create an action plan.

[1776] Prompt Sentence Examples

[1777] "Please summarize the content of the 1-on-1 meeting and propose specific action items and know-how."

[1778] "Convert the recorded audio data into text and extract the key points."

[1779] This system allows you to efficiently manage the content of one-on-one meetings and link them to concrete actions.

[1780] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1781] Step 1:

[1782] User starts recording

[1783] The user opens the dedicated app on their device and presses the record button. The input is the user clicking the record button, and the device starts capturing audio. Specifically, the smartphone's microphone is activated and audio data is captured in real time. The output is the audio data being recorded.

[1784] Step 2:

[1785] The device sends the voice data to the server

[1786] The device uploads captured audio data to the server in regular segments in real time. The input is the audio data being recorded, divided into segments, such as 10 seconds. Specifically, the device sends the audio data segments to the server using HTTP requests. The output is the audio data stored on the server.

[1787] Step 3:

[1788] The server receives the audio data

[1789] The server receives the voice data sent from the terminal and safely stores it in the data storage. The input is the voice data sent from the terminal, and the output is the voice file stored in the data storage. Specifically, the server executes a process to store the voice file in a database.

[1790] Step 4:

[1791] The server converts the voice data into text

[1792] The server uses a speech recognition API (e.g., Google Speech-to-Text) to convert the voice data into text. The input is a saved audio file, and when the server calls the API, the API analyzes the voice data and returns it as text data. The output is text data. For example, the text generated might say, "The delivery date for Project X is behind schedule."

[1793] Step 5:

[1794] The server analyzes the text data

[1795] The server applies an NLP algorithm (e.g., SpaCy) to extract important keywords and context from the text data. The input is text data, which the server analyzes by running the NLP algorithm. The output is important keywords and context data. For example, the important keyword "delayed delivery" is extracted.

[1796] Step 6:

[1797] Server generates summary

[1798] The server summarizes the original text based on the extracted information. The input is important keywords and contextual data, and the server generates a summary by running a summary generation algorithm. The output is summarized text data. For example, the long sentence "Project X is behind schedule, so resources need to be reallocated" is summarized as "Project X is behind schedule."

[1799] Step 7:

[1800] The server aggregates the summary data

[1801] The server aggregates multiple summaries generated from past meetings and identifies frequently occurring issues and trends. The input is multiple summarized text data, which the server analyzes by applying a data aggregation algorithm. The output is the identification of issues and trends. For example, "delayed delivery" is identified as a common issue across multiple projects.

[1802] Step 8:

[1803] Server generates action items

[1804] The server generates specific action items for issues based on past success stories and a knowledge base. The input is the results of identifying issues and trends, and the server generates specific action items by running an action item generation algorithm. The output is a specific action item. For example, it may suggest measures such as "setting up an emergency meeting" or "adjusting additional resources" to address delivery delays.

[1805] Step 9:

[1806] The server searches for know-how

[1807] The server searches the knowledge base to find relevant know-how and best practices. The input is summary data and action items, and the server runs a knowledge base search algorithm to extract relevant information. The output is know-how and best practice information.

[1808] Step 10:

[1809] The server generates a list of suggestions

[1810] The server creates a proposal list based on the search results and provides it to the user. The input is know-how and best practice information, and the server generates the list by running a proposal list generation algorithm. The output is the proposal list provided to the user.

[1811] Step 11:

[1812] The server sends a notification

[1813] The server notifies the user's terminal of the generated action items and suggestion list. The input is the action items and suggestion list, and the server executes the notification process to send the notification to the user. The output is the notification displayed on the user's terminal.

[1814] Step 12:

[1815] User confirms action item

[1816] The user receives a notification on their device and checks the proposed action items and know-how. The input is the notification displayed on the user's device, and the specific action is confirmed by the user checking the operation. The output is the creation of an action plan. For example, Person A and Person B discuss specific measures and create an action plan.

[1817] (Application example 1)

[1818] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1819] There is a need to efficiently manage the content of one-on-one meetings held in the field and link it to specific action items and best practices. However, the current situation is such that the content of meetings is not fully utilized, resulting in a waste of time and resources. In addition, it is a heavy burden for administrators to continue to manage the content of each meeting individually. For this reason, there is a need for a more efficient and automated system.

[1820] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1821] In this invention, the server includes means for starting recording of the meeting, means for transmitting the recorded data to the server, means for the server to recognize the voice of the recorded data and convert it into text, means for the server to summarize the text data, means for the server to analyze the summarized data, means for the server to generate and propose action items, means for automatically generating a proposal list using a generative AI model, and means for notifying the action items and proposal list. This makes it possible to efficiently manage the content of one-on-one meetings and link them to specific actions.

[1822] "Means for starting recording of a meeting" refers to a device or function that allows a user to record the audio of a meeting by pressing a button or other operation.

[1823] "Means for transmitting recorded data to a server" refers to communication functions or programs for uploading recorded audio data to a server in real time or by batch processing.

[1824] "Means for the server to recognize the recorded data and convert it into text" refers to the process of analyzing the voice data on the server and converting it into text data using voice recognition technology.

[1825] "Means for the server to summarize text data" refers to the function of the server using natural language processing (NLP) algorithms, etc. to extract important information from text data and summarize it concisely.

[1826] "Means for the server to analyze the summary data" refers to a data analysis function that analyzes the summary data generated by the server and identifies significant keywords and patterns.

[1827] "Means for the server to generate and suggest action items" refers to the function by which the server automatically generates specific actions and suggestions that the user should take based on the analysis results.

[1828] "Means for automatically generating a suggestion list using a generative AI model" refers to an algorithm or program for automatically creating a suggestion list for a user using a generative AI model.

[1829] "Means for notifying action items and suggestion lists" refers to the system's functionality for notifying the user's terminal of generated action items and suggestion lists.

[1830] The present invention provides a system for efficiently managing the content of one-on-one meetings and linking them to specific actions. The following describes in detail the embodiments of the present invention.

[1831] System Configuration

[1832] This system includes a terminal for users to hold one-on-one meetings, a server for processing data, and a database.

[1833] Voice recording and data transmission

[1834] Yu

[1835] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1836] Step 1:

[1837] When a user starts a one-on-one meeting, they press the recording button on their device, and the device begins capturing audio.

[1838] Input: User presses record button.

[1839] Output: Recorded audio data.

[1840] Specific operation: Captures audio data using the device's microphone function.

[1841] Step 2:

[1842] The device transmits the recorded audio data to the server in real time.

[1843] Input: Pre-recorded audio data.

[1844] Output: The audio data sent to the server.

[1845] Specific operation: Audio data is divided and uploaded to the server via data communication.

[1846] Step 3:

[1847] The server receives the recording and stores it securely.

[1848] Input: Audio data sent from the device.

[1849] Output: Audio data stored on the server.

[1850] What it does: Receives audio data and stores it securely in a database or storage.

[1851] Step 4:

[1852] The server converts the recorded data into text data using voice recognition technology.

[1853] Input: Recorded audio data.

[1854] Output: Text data.

[1855] Specific operation: Uses a speech recognition API to convert voice data into text.

[1856] Step 5:

[1857] The server applies natural language processing (NLP) algorithms to summarize the text data.

[1858] Input: Text data.

[1859] Output: Summarized text data.

[1860] Specific operation: Using NLP algorithms, important keywords and context are extracted from text data and a summary is generated.

[1861] Step 6:

[1862] The server analyzes the summary data and generates action items based thereon.

[1863] Input: Summarized text data.

[1864] Output: Action items.

[1865] Specific Behavior: Analyzes summary data, identifies issues and trends, and generates corresponding action items.

[1866] Step 7:

[1867] The server uses a generative AI model to automatically generate a list of suggestions including relevant know-how and best practices.

[1868] Inputs: Action items and knowledge base data.

[1869] Output: A list of suggestions.

[1870] What it does: It applies generative AI models to extract useful know-how and best practices from a knowledge base and generate a list.

[1871] Step 8:

[1872] The server notifies the user's terminal of the action items and the suggestion list.

[1873] Input: Action items and suggestion lists.

[1874] Output: Notification displayed on the user's device.

[1875] Specific behavior: Using the notification system, the generated action items and suggestion list are sent to the user's device via push notification or email.

[1876] For example, if a factory manager discusses a problem with a new production line during a one-on-one meeting with a field staff member, the server will summarize the content and suggest specific action items such as setting up a trouble-shooting team meeting. In this way, the system can efficiently manage the content of meetings and lead to actual actions.

[1877] For example, an example of a prompt for a generative AI model is: "Please summarize the following text: There is a problem with the new production line. Specifically, there is a problem with the machines frequently stopping. The cause has not yet been identified, but manual intervention may be necessary."

[1878] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1879] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[1880] 1. System Configuration

[1881] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[1882] 2. Description of main functions

[1883] Voice recording and data transmission

[1884] User starts recording:

[1885] A user starts recording the audio of a 1-on-1 meeting using the device. When the user presses the record button, the device begins capturing audio data.

[1886] The device sends the audio data to the server:

[1887] The device captures audio data in real time and uploads it to the server in segments, minimizing the delay before data is processed after the meeting ends.

[1888] Converting audio data to text

[1889] Server receives audio data:

[1890] The server receives the voice data sent from the terminal and stores it securely in data storage.

[1891] Audio to text conversion:

[1892] The server uses a speech recognition API to convert the received voice data into text data. For example, the text data obtained may be something like "Project X's delivery date is behind schedule."

[1893] Text summary generation

[1894] The server parses the text data:

[1895] The server applies natural language processing (NLP) algorithms to extract important keywords and context from the conversational text.

[1896] Summary generation:

[1897] The server uses the extracted information to summarize the original text, for example, into a concise sentence such as "Project X's delivery date is delayed."

[1898] emotion recognition

[1899] The server uses an emotion engine to recognize the user's emotions:

[1900] The server analyzes the user's emotions from the voice and text data, including algorithms that identify emotions from voice tone and text content.

[1901] Data storage and analysis

[1902] Summary and sentiment data storage:

[1903] The server stores the summarized text data and the sentiment analysis results in a database.

[1904] Data Analysis:

[1905] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[1906] Action item generation and suggestions

[1907] Server generates action item:

[1908] The server generates specific action items based on the analysis results and sentiment data, referring to past history and best practices.

[1909] Search our knowledge base:

[1910] The server searches and extracts relevant know-how and best practices from the knowledge base.

[1911] Generate a list of suggestions:

[1912] The server generates specific action suggestions and a list of best practices for the user based on the search results.

[1913] Notification and confirmation

[1914] Send notifications:

[1915] The server notifies the user of the generated action items and suggestion list via a popup notification or email notification.

[1916] User sees action item:

[1917] Users receive notifications and see action items and suggested know-how, allowing them to take concrete steps.

[1918] Specific examples

[1919] For example, consider the case where employee A has a one-on-one meeting with line manager B.

[1920] 1. Start recording:

[1921] Person A presses the recording button on the device to start the meeting.

[1922] 2. Sending audio data to the server:

[1923] During the meeting, the device transmits audio data to the server in real time.

[1924] 3. Text:

[1925] The server uses a speech recognition API to convert the voice data into text that says "We're having problems with project Y."

[1926] 4. Summary generation:

[1927] Using an NLP algorithm, summarize it as "Project Y problem."

[1928] 5. Emotion recognition:

[1929] The emotion engine recognizes from the tone of the voice and the content of the text that Person A is feeling stressed about the problem.

[1930] 6. Data Retention:

[1931] The server stores the text data and emotion data in a database.

[1932] 7. Data analysis and action generation:

[1933] The server analyzes similar past data and sentiment data and generates action items such as "set up a meeting to identify problems."

[1934] 8. Know-how proposal:

[1935] The server extracts and proposes "problem-solving methods that were previously effective in Project Y" from the knowledge base.

[1936] 9. User confirms:

[1937] Persons A and B check this information on their devices and create an action plan.

[1938] In this way, by combining an emotion recognition engine, it is possible to provide a management system that deepens the content of one-on-one meetings and leads to concrete actions.

[1939] The processing flow will be explained below.

[1940] Step 1:

[1941] User starts recording:

[1942] The user opens the dedicated application on their device and presses the start recording button for the 1-on-1 meeting, which starts the recording process.

[1943] Step 2:

[1944] The device captures audio data:

[1945] The device uses its built-in microphone to capture meeting audio in real time and buffers the audio data.

[1946] Step 3:

[1947] The device sends the audio data to the server:

[1948] The device uploads the buffered audio data to the server segment by segment, which allows the data to be stored on the server in real time.

[1949] Step 4:

[1950] Server receives audio data:

[1951] The server receives the voice data sent from the device and stores it in a secure data storage.

[1952] Step 5:

[1953] The server converts the audio data to text:

[1954] The server calls the speech recognition API to convert the voice data into text data, which is then stored in a database.

[1955] Step 6:

[1956] The server summarizes the text data:

[1957] The server uses natural language processing (NLP) algorithms to analyze the text data, extract key points and keywords, and generate a summary.

[1958] Step 7:

[1959] The server performs emotion recognition using the emotion engine:

[1960] The server uses an emotion engine to analyze the user's emotions from the voice and text data, using algorithms that identify emotions from voice tone and text content.

[1961] Step 8:

[1962] The server stores the emotion recognition results and summary data:

[1963] The server stores the generated summary data and emotion recognition results in a database for later analysis and statistical processing.

[1964] Step 9:

[1965] The server analyzes the summary data and sentiment data:

[1966] The server aggregates multiple summary and sentiment data sets and uses an analytics engine to identify issues and trends.

[1967] Step 10:

[1968] Server generates action item:

[1969] The server generates specific action items based on the analysis results and sentiment data, referencing past history and best practices.

[1970] Step 11:

[1971] Server searches knowledge base:

[1972] The server searches and extracts relevant know-how and best practices from the knowledge base.

[1973] Step 12:

[1974] Server generates suggestion list:

[1975] Based on the extracted information, specific action suggestions and a list of best practices are generated for the user.

[1976] Step 13:

[1977] Server sends notification:

[1978] Generated action items and suggestion lists are notified to the user's device in the form of a popup or email.

[1979] Step 14:

[1980] User sees action item:

[1981] The user uses the device to check the notification and view the generated action items and suggested know-how.

[1982] Step 15:

[1983] User performs an action:

[1984] Based on the action items confirmed by the user, the system performs specific tasks, such as setting up a project meeting or coordinating additional resources.

[1985] Example 2

[1986] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1987] In traditional one-on-one meetings, simply recording the conversation and then manually converting it into text and summarizing it later takes time and effort. Nuances such as the speaker's emotions and tone during the meeting are not recorded, often resulting in insufficient understanding and analysis. Furthermore, opportunities for business improvement tend to be missed because specific action items and effective best practices are not generated.

[1988] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for starting recording of the conference, means for transmitting recorded data to a computer, means for the computer to perform speech recognition on the recorded data and convert it into text data, means for the computer to summarize the text data, means for the computer to analyze the summary data and emotion data, means for the computer to generate and suggest action items, and means for recognizing the user's emotions using an emotion recognition engine. This makes it possible to effectively record and analyze the contents of the conference, obtain a detailed understanding including emotional nuances, and suggest specific improvement actions.

[1989] "Means for starting a meeting recording" refers to a feature that allows users to record one-on-one meetings and other meetings on their devices.

[1990] The "means for transmitting recorded data to a computer" is a function for uploading audio data recorded by a terminal to a server in real time or in batch processing.

[1991] "Means for a computer to recognize the recorded data and convert it into text data" refers to the process by which a server uses a speech recognition API to convert the voice data into text data.

[1992] "Means for computer-generated summarization of text data" refers to the process in which the server uses a natural language processing algorithm to extract important information from the converted text data and summarize it concisely.

[1993] "Means for a computer to analyze summary data and emotion data" refers to a function in which the server uses the stored summary and emotion data to analyze past history and patterns and identify trends and issues.

[1994] "Means for computer-generated and proposed action items" refers to a function in which the server generates specific action items based on the results of data analysis and best practices, and proposes them to the user.

[1995] "Means for recognizing the user's emotions using an emotion recognition engine" refers to the process in which the server analyzes the user's emotions from voice data and text data and records them as emotion data.

[1996] "Storing the summary data and emotion data in an information storage device" refers to the process in which the server safely stores the generated summary text data and emotion data in storage such as a database.

[1997] "Extracting know-how and best practices from the knowledge base and generating a proposal list" refers to the process in which the server searches the knowledge base, extracts relevant information, and creates a specific proposal list for the user.

[1998] The present invention relates to a one-on-one meeting management system that incorporates a user emotion recognition engine. This system has functions for recording, converting voice data into text, generating summaries, recognizing emotions, analyzing data, generating action items, and suggesting know-how and best practices. The following describes in detail the embodiments of the present invention.

[1999] System Configuration

[2000] The system includes a terminal for users to conduct one-on-one meetings, a server for processing data, a database, and an emotion recognition engine.

[2001] Main Features

[2002] Voice recording and data transmission

[2003] When a user presses the record button on the device, the device begins capturing audio using the built-in microphone. The device then sends the captured audio data to the server at regular intervals (for example, every 10 seconds). When the device sends the recorded data to the server, it uses Wi-Fi or mobile data communication and sends the audio data via an HTTP POST request.

[2004] Converting audio data to text

[2005] The server receives the voice data sent from the device and stores it in a secure data storage (e.g., Amazon S3).The server then converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).Through this conversion process, all conversation content is obtained as text data.

[2006] Text summary generation

[2007] The server uses natural language processing (NLP) algorithms to extract important keywords and context from the acquired text data and summarize it concisely. For example, the server analyzes the text data using an NLP library such as the Natural Language Toolkit (NLTK) and generates a summary.

[2008] emotion recognition

[2009] The server uses an emotion recognition engine to analyze the user's emotions from the voice data and text data. To determine emotions from the tone of the voice or specific words, the server sends the voice data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to obtain emotion data.

[2010] Data storage and analysis

[2011] The server generates summary text data and stores the sentiment data in a database (e.g., MySQL or PostgreSQL). The server uses the stored data to analyze past history and patterns and identify issues and trends. When analyzing the data, the server uses Structured Query Language (SQL).

[2012] Action item generation and suggestions

[2013] The server generates specific action items based on the analysis results and sentiment data. In addition, the server searches a knowledge base to extract relevant information and generate a list of specific suggestions for the user. The server uses a rule-based engine and machine learning algorithms to generate appropriate action items and match them with the knowledge base.

[2014] Notification and confirmation

[2015] The server notifies the user device of the generated action items and suggestion list. Notification methods include pop-up notifications, emails, and in-app notifications. The server sends real-time notifications to the user's smartphone using services such as Firebase Cloud Messaging (FCM). The user receives the notification on their device, checks the specific action items and suggested know-how, and can then take the next step.

[2016] Specific examples

[2017] For example, consider the case where employee A has a one-on-one meeting with line manager B. A presses the record button on their device to begin the meeting. During the meeting, the device sends audio data to the server in real time. The server uses a speech recognition API to convert the audio data into text: "We're experiencing problems with project Y." The server then uses an NLP algorithm to summarize the text as "Problem with project Y." The server's emotion engine recognizes from the tone of the voice and the content of the text that A is feeling stressed about the problem. The server stores the text data and emotion data in a database, analyzes similar past data and emotion data, and generates action items such as "Arrange a meeting to identify problems." The server extracts and suggests "problem-solving methods that were previously effective on project Y" from its knowledge base. Finally, A and B review this information on their devices and create an action plan.

[2018] Examples of prompt statements

[2019] "Generate a prompt sentence to describe the above 1-on-1 meeting management system."

[2020] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2021] A detailed explanation of the system program processing flow and processing steps

[2022] Step 1:

[2023] User starts recording

[2024] Input: A user starts a 1-on-1 meeting.

[2025] How it works: When a user presses the record button on their device, the device's built-in microphone begins capturing audio.

[2026] Output: Audio data is generated and captured.

[2027] Step 2:

[2028] The device sends the voice data to the server

[2029] Input: Audio data captured by the device.

[2030] How it works: Audio data is sent to the server at regular intervals (e.g., every 10 seconds). The audio data is sent via an HTTP POST request using Wi-Fi or mobile data.

[2031] Output: Audio data uploaded to the server.

[2032] Step 3:

[2033] The server receives and stores the audio data

[2034] Input: The server receives the audio data sent from the device.

[2035] How it works: The server receives the audio data and stores it in secure data storage (e.g. Amazon S3).

[2036] Output: Saved audio data.

[2037] Step 4:

[2038] The server converts the voice data into text

[2039] Input: Stored audio data.

[2040] How it works: The server uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the audio data into text, sends the audio file to the API, and stores the returned text.

[2041] Output: The generated text data.

[2042] Step 5:

[2043] The server summarizes the text data

[2044] Input: Generated text data.

[2045] How it works: The server uses NLP algorithms (e.g., the NLTK library) to analyze the text data, extract important keywords and context, and generate a summary.

[2046] Output: The generated summary.

[2047] Step 6:

[2048] The server recognizes emotions

[2049] Input: Audio and text data.

[2050] How it works: The server uses an emotion recognition API (e.g. IBM Watson Tone Analyzer) to extract emotion data from voice tone and text content.

[2051] Output: Emotion data.

[2052] Step 7:

[2053] The server saves and analyzes the summary and emotion data

[2054] Input: Summary sentences and sentiment data.

[2055] How it works: The server saves the summary sentences and emotion data in a database and analyzes past history and patterns. SQL is used to save the data in the database and perform the analysis.

[2056] Output: Analysis of issues and trends.

[2057] Step 8:

[2058] Server generates and suggests action items

[2059] Input: Analysis results and sentiment data.

[2060] How it works: The server generates specific action items using a rule-based engine and machine learning algorithms, extracts relevant information from a knowledge base, and generates a list of suggestions.

[2061] Output: Generated action items and suggestion list.

[2062] Step 9:

[2063] The server notifies the user

[2064] Input: Generated list of action items and suggestions.

[2065] How it works: The server uses Firebase Cloud Messaging (FCM) or similar to send real-time notifications to the user's smartphone.

[2066] Output: A notification message to the user.

[2067] Step 10:

[2068] User confirms action item

[2069] Input: Notification message.

[2070] Action: The user receives a notification on their device, sees an action item, or suggests a know-how, and takes specific steps based on that.

[2071] Output: User confirmed action items and implementation plans.

[2072] (Application example 2)

[2073] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2074] Current meeting recording and data analysis systems are limited to converting audio data into text, providing a brief summary, and suggesting action items. As a result, they lack insight into the emotional state of participants during meetings and specific, real-time countermeasures based on that information. Furthermore, because they are unable to process data in real time, they lack the ability to respond to emergencies and provide immediate response.

[2075] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2076] In this invention, the server includes means for [analyzing summary data and text data using an emotion recognition engine], means for [acquiring emotion data in real time and proposing specific action items such as setting up an emergency response meeting based on the analysis results], and means for [converting audio data into text in real time while recording a meeting, performing emotion analysis, and immediately reflecting the results]. This makes it possible [to grasp the emotional state of participants in a meeting in real time and immediately propose specific measures].

[2077] The "means for starting recording of a conference" is a function that a user operates to record the audio of a conference or meeting.

[2078] The "means for transmitting recorded data to a server" is a function for uploading recorded voice data to a server in real time or later.

[2079] "Means for the server to recognize the voice of the recorded data and convert it into text" refers to a function that enables the server to use voice recognition technology to convert the recorded voice data into text data.

[2080] The "means for the server to summarize text data" is a function for the server to summarize the text data acquired by the server using a natural language processing algorithm and extract important points.

[2081] The "means for the server to analyze the summarized data and text data using an emotion recognition engine" is a function that enables the server to analyze summarized text data and unsummarized text data using an emotion recognition algorithm to identify the emotional states of conference participants.

[2082] "Means for the server to generate and suggest action items" is a function that allows the server to generate specific action items based on the analysis results and suggest them to the user.

[2083] "A means of acquiring emotional data in real time and proposing specific action items such as setting up emergency response meetings based on the analysis results" is a function in which the server acquires emotional data in real time during a meeting and quickly reflects the analysis results to propose actions such as setting up emergency response meetings.

[2084] "Means for converting audio data into text in real time while recording a meeting, performing sentiment analysis, and immediately displaying the results" refers to a function that enables a robot or system to convert the audio data of a meeting into text in real time while recording it, perform sentiment analysis based on the text data, and immediately present the results to the user.

[2085] The "means for saving summary data and emotion data in a database" is a function that enables the server to safely save summary data and emotion recognition results in a database.

[2086] "Means of inputting a prompt sentence into a generative AI model and proposing optimal best practices from past data" is a function that enables the server to input a specific prompt (input sentence) into the generative AI model, search for optimal best practices from past data, and propose them.

[2087] This invention is a system for making work improvement meetings in factories more efficient, and has the functions of converting voice data into text, recognizing emotions, conducting real-time analysis, and generating and proposing action items.

[2088] System Configuration

[2089] The system includes terminals used in factories, a server for processing recorded data, a database, and an emotion recognition engine. The terminals are primarily used for recording meetings, while the server performs speech recognition and analysis. The database stores text data and emotion data. The emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state.

[2090] Program processing

[2091] The system begins when a user presses a recording button on their device to start a meeting. The recorded audio data is sent in real time to a server, which then converts it into text using speech recognition technology. The server then analyzes the text data using a natural language processing algorithm to generate a summary of the meeting. An emotion recognition engine analyzes the text data and voice tone to identify the user's emotional state and reflects the results in real time.

[2092] Novel Features

[2093] 1. Real-time emotional data acquisition and analysis: The server acquires emotional data in real time during the meeting and, based on the analysis results, proposes specific action items such as setting up an emergency response meeting.

[2094] 2. Instant text conversion and sentiment analysis during meeting recording: Audio data is converted into text in real time during meeting recording, and the text data is instantly analyzed for sentiment.

[2095] 3. Use of generative AI models: By inputting prompt statements into generative AI models, the models can suggest optimal best practices based on past data.

[2096] Specific usage

[2097] The device is equipped with a microphone (e.g., a USB microphone) and captures voice data. The recorded data is sent to a server, where it is converted into text data using the Google Speech Recognition API. The text data is then summarized using a natural language processing algorithm (e.g., BERT), and sentiment analysis is performed using an emotion recognition engine. The analysis results are stored in a database and notified to the user in real time. Examples of prompts include "Please print a summary of this meeting" and "Perform emotion recognition and summarize the issues."

[2098] Specific examples

[2099] For example, when factory worker A holds a one-on-one meeting with manager B, he presses the recording button on his device to start the meeting. During the meeting, audio data is sent to the server in real time and instantly converted into text data. The server analyzes the generated text data, and if worker A says, "There has been an increase in product defects this month," it generates a summary saying, "Investigate the cause of product defects," and recognizes through sentiment analysis that worker A is frustrated. As a result, the server suggests "Set up an emergency meeting to address product defects" and generates specific action items.

[2100] This system makes it possible to identify problems in real time during meetings and respond immediately, which is expected to lead to more efficient work improvements within the factory.

[2101] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2102] Step 1:

[2103] The user presses the record button on the device to start the conference.

[2104] Input: User actions

[2105] How it works: By pressing the record button, the device will begin recording the audio data during the meeting.

[2106] Output: Audio data begins to accumulate on the device.

[2107] Step 2:

[2108] The device sends the recorded data to the server in real time.

[2109] Input: Recorded audio data

[2110] How it works: Recorded data is uploaded to the server in real time, segment by segment.

[2111] Output: Audio data sent to the server

[2112] Step 3:

[2113] The server converts the voice data into text data using a voice recognition API.

[2114] Input: Audio data sent to the server

[2115] How it works: Uses the Google Speech Recognition API to convert audio data into text.

[2116] Output: Text data (e.g., "Product defects are increasing this month.")

[2117] Step 4:

[2118] The server summarizes the text data using natural language processing algorithms.

[2119] Input: Text data converted by speech recognition

[2120] How it works: It uses natural language processing algorithms such as BERT to extract key points from text data and generate summaries.

[2121] Output: Summary data (e.g., "Investigation into the cause of product defects")

[2122] Step 5:

[2123] The server analyzes the summary data and text data using an emotion recognition engine.

[2124] Input: Abstract and text data

[2125] How it works: An emotion recognition engine is used to recognize and analyze user emotions from text data and voice tones.

[2126] Output: Emotion data (e.g., stress and irritation are recognized)

[2127] Step 6:

[2128] The server stores the summary data and the emotion data in a database.

[2129] Input: Summary data and sentiment data

[2130] What it does: Stores data securely in a database.

[2131] Output: Summary data and sentiment data stored in a database

[2132] Step 7:

[2133] The server generates and suggests action items.

[2134] Input: Summary data and sentiment data

[2135] How it works: You input a prompt into the generative AI model, which will then use past data to suggest best practices. For example, "Please provide a summary of this meeting." "Please use emotion recognition to summarize the issues."

[2136] Output: Suggested action items and best practices (e.g., "Schedule an emergency product defect resolution meeting")

[2137] Step 8:

[2138] The user reviews the proposed action items and creates an implementation plan.

[2139] Input: Action items and best practices notified by the server

[2140] How it works: The user sees specific action items on their device and creates an action plan based on them.

[2141] Output: A concrete action plan (e.g., "Schedule an emergency meeting on ____ day to discuss problem-solving methods")

[2142] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2143] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2144] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2145] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2146] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2147] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2148] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2149] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2150] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2151] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2152] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2153] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2154] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2155] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2156] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2157] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2158] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2159] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2160] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2161] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2162] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2163] The following is further disclosed regarding the above embodiment.

[2164] (Claim 1)

[2165] a means for initiating recording of the meeting;

[2166] means for transmitting the recording data to a server;

[2167] A means for the server to recognize the recorded data and convert it into text;

[2168] a means for the server to summarize the text data;

[2169] means for the server to analyze the summary data;

[2170] a means for the server to generate and suggest action items;

[2171] A system including:

[2172] (Claim 2)

[2173] 10. The system of claim 1, further comprising means for storing the summary data in a database.

[2174] (Claim 3)

[2175] 2. The system of claim 1, wherein the server includes means for extracting know-how and best practices from the knowledge base and generating a list of suggestions.

[2176] "Example 1"

[2177] (Claim 1)

[2178] a means for initiating recording of the meeting;

[2179] means for transmitting the recording data to a server;

[2180] A means for the server to convert the recorded data into text data using voice recognition technology;

[2181] A means for the server to summarize the text data using a natural language processing algorithm;

[2182] means for the server to analyze the summarized data using a data analysis algorithm;

[2183] A means for the server to generate and suggest action items based on past success stories;

[2184] A system including:

[2185] (Claim 2)

[2186] 10. The system of claim 1, further comprising means for storing the summary data in a database.

[2187] (Claim 3)

[2188] 10. The system of claim 1, wherein the server includes means for extracting relevant knowledge and best practices from the knowledge base and generating a list of suggestions.

[2189] "Application Example 1"

[2190] (Claim 1)

[2191] a means for initiating recording of the meeting;

[2192] means for transmitting the recording data to a server;

[2193] A means for the server to recognize the recorded data and convert it into text;

[2194] a means for the server to summarize the text data;

[2195] means for the server to analyze the summary data;

[2196] a means for the server to generate and suggest action items;

[2197] A means for automatically generating a list of suggestions using a generative AI model; and

[2198] A means of communicating action items and suggestion lists;

[2199] A system including:

[2200] (Claim 2)

[2201] 10. The system of claim 1, further comprising means for storing the summary data in a database.

[2202] (Claim 3)

[2203] 2. The system of claim 1, wherein the server includes means for extracting know-how and best practices from the knowledge base and generating a list of suggestions, and means for a user to review the action items and suggested know-how.

[2204] "Example 2: Combining Emotion Engines"

[2205] (Claim 1)

[2206] a means for initiating a recording of the meeting;

[2207] means for transmitting the recorded data to a computer;

[2208] A means for a computer to recognize the recorded data by voice and convert it into character data;

[2209] means for summarizing character data by a computer;

[2210] means for a computer to analyze the summary data and the emotion data;

[2211] a computer-generated means for generating and suggesting action items;

[2212] means for recognizing a user's emotion using an emotion recognition engine;

[2213] A system including:

[2214] (Claim 2)

[2215] 10. The system of claim 1, wherein the summary data and the emotion data are stored in an information storage device.

[2216] (Claim 3)

[2217] 2. The system according to claim 1, wherein the computer extracts know-how and best practices from the knowledge base and generates a proposal list.

[2218] "Application example 2 when combining emotion engines"

[2219] (Claim 1)

[2220] a means for initiating recording of the meeting;

[2221] means for transmitting the recording data to a server;

[2222] A means for the server to recognize the recorded data and convert it into text;

[2223] a means for the server to summarize the text data;

[2224] A server analyzes the summary data and the text data using an emotion recognition engine;

[2225] a means for the server to generate and suggest action items;

[2226] A means to obtain emotional data in real time and propose specific action items such as setting up emergency response meetings based on the analysis results, and

[2227] A means to convert audio data into text in real time during recording of a meeting, perform sentiment analysis, and instantly reflect the results.

[2228] A system including:

[2229] (Claim 2)

[2230] 10. The system of claim 1, further comprising means for storing the summary data and the emotion data in a database.

[2231] (Claim 3)

[2232] A means for the server to extract know-how and best practices from the knowledge base and generate a list of suggestions;

[2233] The system of claim 1, further comprising means for inputting prompt statements to the generative AI model to suggest optimal best practices based on past data. [Explanation of symbols]

[2234] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for initiating recording of the meeting; means for transmitting the recording data to a server; A means for the server to recognize the recorded data and convert it into text; a means for the server to summarize the text data; means for the server to analyze the summary data; a means for the server to generate and suggest action items; A system including:

2. 10. The system of claim 1, further comprising means for storing the summary data in a database.

3. 2. The system of claim 1, wherein the server includes means for extracting know-how and best practices from the knowledge base and generating a list of suggestions.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A