system

The system addresses inefficiencies in task management by automatically generating and prioritizing tasks from conversation data, ensuring important tasks are not missed and meeting industry-specific needs, thereby improving productivity.

JP2026070190APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing systems struggle to efficiently manage tasks from conversations without labor-intensive manual input, fail to grasp important information, and lack industry-specific project management capabilities, leading to inefficiencies in task management and prioritization.

Method used

A system that acquires conversation data, analyzes it to extract important information, automatically generates tasks, sets priorities, and provides reminder management, tailored to specific industries using industry-specific models.

Benefits of technology

Enhances task management efficiency by automating task generation and prioritization, ensuring important tasks are not missed, and adapting to the unique needs of various industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070190000001_ABST
    Figure 2026070190000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A data collection method for obtaining conversation data, An information analysis means that analyzes acquired conversation data and extracts important information, A task generation method that generates a task list based on extracted important information, A priority setting mechanism for setting the priority of each task in the task list, A reminder management system that manages reminders based on task progress and deadlines, Industry-specific methods that address specific industries, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a business environment, it is required to manage important tasks without omission from meetings and communications, and to efficiently generate and prioritize tasks. However, there are problems that manual task input and management require labor, and it is difficult to grasp important information without omission from a wide range of discussions. Furthermore, in a specific industry, it is also a problem that a system corresponding to project management peculiar to that industry is required.

Means for Solving the Problems

[0005] This invention solves these problems by providing a means for acquiring conversation data, analyzing it, and extracting important information. Specifically, it provides a system that automatically generates tasks based on the extracted information and uses a means for setting the priority of each task to realize reminder management based on task progress and deadlines. Furthermore, it meets the specific needs of particular industries by providing industry-specific means that can be adapted to specific industries. As a result, users can significantly improve the efficiency of task management and concentrate on essential work.

[0006] "Conversation data" refers to the content of communication recorded in audio or text format.

[0007] "Data collection means" refers to devices or methods equipped with the function of acquiring conversation data and preparing it for processing.

[0008] "Information analysis means" refers to technology equipped with the function of analyzing acquired conversation data and extracting important information from it.

[0009] "Task generation means" refers to a function that automatically creates a task list based on important information extracted by information analysis means.

[0010] "Priority setting means" refers to the techniques and processes used to evaluate the importance and urgency of each task in a generated task list and to rank them accordingly.

[0011] A "reminder management system" refers to a system equipped with functions to appropriately manage the progress and deadlines of tasks and notify users at the necessary time.

[0012] "Industry-specific methods" refer to technologies that utilize industry-specific functions and datasets to address the unique needs and challenges of a particular industry. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention describes a system that acquires conversational data generated during meetings and communications in real time, analyzes it, and automatically generates and manages important tasks. The program processing and specific examples of this system are described below.

[0035] In this system, users participate in meetings and generate conversation data through audio and chat data. The terminal converts this data into text data in real time using speech recognition technology, and also processes related chat logs and email data.

[0036] The server receives text data and related data sent from the terminal. The server then uses natural language processing (NLP) techniques to analyze the data and extract important information from the conversation. Specifically, it employs methods to identify decisions and the next actions to be taken.

[0037] Based on the extracted information, the server automatically generates a task list, listing important action items. The listed tasks are then accompanied by necessary details, referencing the user's past task history and project information.

[0038] For each item in the generated tasks, the server uses machine learning models and rule-based algorithms to set priorities. This makes it easier for users to determine which tasks they should tackle quickly.

[0039] Furthermore, the server sends reminders to the device at appropriate times based on the task's progress and deadline. These reminders prompt the user to take the next action, preventing tasks from being missed or delayed.

[0040] Furthermore, this system can be tailored to specific industries. The server uses pre-configured industry-specific models to perform analyses that address the needs and challenges of those particular industries. These specialized models are trained on industry-specific datasets and contribute to the automation of industry-specific tasks.

[0041] For example, if the user is a project manager in the IT industry, the server analyzes the audio data of the project meeting and automatically generates tasks such as "Prepare for next week's sprint review," setting them to a high priority. Furthermore, it supports smooth project progress by sending reminders to the terminal at appropriate times, such as "The deadline for creating sprint review materials is approaching."

[0042] Thus, the present invention provides users with an efficient task management system, enabling them to streamline business processes and improve productivity.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] Users participate in meetings and engage in conversations and discussions. Communication takes place using voice and chat.

[0046] Step 2:

[0047] The device records audio during meetings in real time and converts it into text data using speech recognition technology. To generate highly accurate text from the recorded audio, it performs ambient noise removal and speaker separation.

[0048] Step 3:

[0049] The terminal sends the generated text data, chat logs used during the meeting, and email data as needed to the server. This transmitted data is then ready for processing by the system.

[0050] Step 4:

[0051] The server analyzes the received text data using a natural language processing (NLP) engine. Through understanding grammatical structure and vocabulary, it extracts important information from the context of the conversation, clarifying decisions and next actions.

[0052] Step 5:

[0053] The server automatically generates a task list based on the extracted information. This list specifically lists the required tasks, and each task includes relevant details. Furthermore, it refers to the user's past work history and supplements task details as needed.

[0054] Step 6:

[0055] The server sets the priority of each task based on its importance and deadline. Using machine learning algorithms, it evaluates the urgency and relative priority of tasks and indicates which tasks the user should tackle first.

[0056] Step 7:

[0057] The server monitors the progress of each task and sends reminders to the user's device based on the deadline. This ensures that users receive timely notifications and prevent tasks from being missed.

[0058] Step 8:

[0059] The servers are specialized for specific industries, employing industry-specific models accordingly. This allows them to address the diverse needs of each industry and achieve specialized task management.

[0060] Step 9:

[0061] Users can view generated task lists and reminders through the device interface and customize the interface to optimize task management to suit their work style. This makes it easier to optimize task placement and management for each industry and team.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] Traditional meeting management systems struggled to extract and properly manage important tasks from conversation content. Furthermore, their lack of flexibility to meet the unique needs of different industries hindered efficient business operations.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes data collection means for acquiring conversation information, information analysis means for analyzing the acquired conversation information and extracting important matters, and task generation means for generating a list of tasks based on the extracted important matters. This makes it possible to immediately identify and manage tasks from conversations in meetings.

[0067] "Conversational information" refers to information generated as audio or text during meetings and communication sessions.

[0068] "Data collection means" refers to a function or device for efficiently acquiring conversational information and incorporating it into a system.

[0069] "Information analysis means" refers to the techniques and methods used to analyze acquired conversational information and extract important information from it.

[0070] "Task generation means" refers to a function for automatically generating tasks based on important information extracted by information analysis means.

[0071] A "prioritization method" refers to a technique for determining which tasks should be prioritized among the generated tasks and setting their order.

[0072] "Notification management means" refers to a function that provides notifications at the appropriate time based on the progress and deadlines of a task, prompting the user to take the next action.

[0073] "Industry-specific methods" refer to technologies and methods for customizing systems to address the specific needs and challenges of a particular industry.

[0074] A "generative AI model" refers to a group of algorithms that use artificial intelligence technology to generate useful information or answers from input data.

[0075] "Prompt generation means" refers to a function or technology for generating natural language questions and commands from conversational information.

[0076] This invention is a system for efficiently processing conversational information generated in meetings and communications, and for achieving automated task generation and management. The system includes, as its components, data collection means, information analysis means, task generation means, priority setting means, notification management means, industry-specific means, generation AI model, and prompt generation means.

[0077] Users participate in meetings and generate conversational information in the form of voice or chat. The device collects this information using a microphone and corresponding software (e.g., a video conferencing application). The collected audio data is converted into text using a common speech recognition API (a comprehensive and widely used API as an example) as a voice identification technology.

[0078] The server analyzes the text information sent from the terminal using a natural language processing library (e.g., SpaCy or OpenNLP). It then extracts important information and uses a generative AI model to generate appropriate prompts. These prompts provide specific, actionable instructions for the generated task list. For example, a possible prompt might be, "List the preparation items for next week's team meeting."

[0079] Furthermore, the server prioritizes the generated tasks through machine learning models and rule-based algorithms. Based on the set priorities, it sends notifications via the terminal at appropriate times according to the progress and deadlines of the work. This prompts users to take the next action, thereby promoting work efficiency.

[0080] Furthermore, the server performs analysis using specialized models for specific industries. This industry-specific approach allows for analysis tailored to the needs of each industry, providing unique tasks and information. Combining these elements makes it possible to achieve efficient task management and optimization of business processes.

[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0082] Step 1:

[0083] Users participate in meetings and generate conversational information in voice and chat formats. This results in voice and text data as input. The device collects this conversational information in real time through microphones and video conferencing software.

[0084] Step 2:

[0085] The device converts collected audio data into text using speech recognition technology. Specifically, it performs the process of converting audio data into text data via a speech recognition API. The input is audio data, and the output is text data. This text data records the content of the conversation as characters.

[0086] Step 3:

[0087] The terminal sends the converted text data and associated chat logs to the server. The server receives this data and prepares it for analysis. The input consists of text data and chat logs, which are used for the subsequent analysis process.

[0088] Step 4:

[0089] The server analyzes the received text data using natural language processing (NLP) techniques. Specifically, it uses an NLP library to extract important information and keywords from the text. The input is text data, and the output is important information as a result of the analysis.

[0090] Step 5:

[0091] Based on the analysis results, the server operates a prompt generation mechanism using a generation AI model. A specific list of tasks is then generated based on the resulting prompt messages. The input is important information, and the output consists of the generated prompt messages and the corresponding task items.

[0092] Step 6:

[0093] The server applies machine learning models and rule-based algorithms to prioritize work items and determine task priorities. The input is work items, and the output is a list of tasks with their priorities set.

[0094] Step 7:

[0095] The server uses a prioritized task list to send notifications to the terminal at appropriate times based on the progress and deadlines of each task. These notifications are intended to prompt the user to take the next action. The input is the task list and the progress of the tasks, and the output is the notification sent to the user.

[0096] (Application Example 1)

[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] In brick-and-mortar retail operations, efficiently managing tasks arising from customer interactions and staff conversations is essential. However, traditional methods require manually recording conversations and manually organizing tasks, hindering operational efficiency. Furthermore, the insufficient use of specific information technology devices for visual task presentation leads to challenges in immediacy and accuracy.

[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0100] In this invention, the server includes information gathering means for acquiring conversation data, data analysis means for analyzing the acquired conversation data and extracting important information, and display means for visually presenting task information to the user using a specific information technology device. This improves customer service and operational efficiency, enabling real-time task management and information presentation.

[0101] "Information gathering means" refers to a system for acquiring information such as conversations and audio data in real time.

[0102] "Data analysis tools" are technologies that analyze acquired information and extract important data and tasks.

[0103] A "task creation method" is a method for automatically generating a list of necessary tasks based on analyzed data.

[0104] A "prioritization mechanism" is a system that determines the priority of generated tasks according to their importance and urgency.

[0105] A "notification management system" is a technology that has the ability to send reminders to users based on the progress and deadlines of tasks.

[0106] "Industry-specific methods" refer to methods for adjusting analysis and task generation functions to suit specific industrial sectors.

[0107] A "display means" is a mechanism that uses a specific information technology device to visually present task information to the user.

[0108] This invention is a system that enables smooth task management by allowing staff to use smart glasses in physical stores to analyze customer interactions and conversations among staff in real time. The server uses a specific voice recognition device as a means of information gathering to collect voice data obtained from the field. Subsequently, the data is analyzed by the server using a data analysis means that employs natural language processing technology to extract important tasks and information. Specifically, using the Hugging Face Transformers library is effective.

[0109] The server also generates a task list based on information extracted by the task creation mechanism and determines the importance of each task using the priority setting mechanism. This information is immediately presented to staff via a specific information technology device, such as smart glasses. This allows users to efficiently manage their work within the physical store.

[0110] Furthermore, reminders are displayed on staff members' smart glasses via notification management tools, based on task progress and deadlines. Industry-specific tools can also be used to provide features optimized for specific store types or industries.

[0111] For example, if a store clerk says to a customer, "This item is sold out. Can you restock it?", a task such as "Restock sold-out item X" will be automatically generated. Furthermore, by utilizing the generative AI model and using a prompt such as, "Analyze the following conversation and extract the important tasks. Conversation: 'This item is sold out. Can you restock it?'", more accurate task analysis becomes possible.

[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0113] Step 1:

[0114] The terminal acquires audio data of staff conversations and customer interactions within the store. This audio data is captured through a built-in speech recognition device in the terminal. The input is conversations within the store, and the output is audio data.

[0115] Step 2:

[0116] The device converts the acquired audio data into text data. This conversion process utilizes speech recognition technology. Specifically, it uses speech recognition software such as Google® Speech API. The input is audio data, and the output is the converted text data.

[0117] Step 3:

[0118] The server receives text data sent from the terminal. During this receiving process, the server uses natural language processing techniques to identify and extract important information. The input is text data, and the output is the analyzed important information. The accuracy of the analysis can be improved by using a generative AI model.

[0119] Step 4:

[0120] The server generates a task list and sets priorities based on the extracted information. This uses machine learning algorithms and also considers the user's past task history. The input is the analyzed key information, and the output is a task list with priorities set.

[0121] Step 5:

[0122] The server sends the generated task list to the terminal, where it is displayed to the staff via smart glasses. This allows users to visually check tasks in real time and respond immediately. The input is a task list with priorities set, and the output is the task list displayed on the smart glasses.

[0123] Step 6:

[0124] The server notifies the user's device of the task progress and deadline at the appropriate time. This prevents the user from missing or delaying tasks. Input is the task list and task management information, and output is the reminder displayed on the device.

[0125] Through the above processing steps, it is possible to improve the efficiency of task management in physical stores.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention combines an emotion engine with a system that acquires conversation data in real time, analyzes the content of that conversation, and generates tasks, thereby enabling task management that responds to the user's emotions. The program processing and specific examples of this system are described below.

[0128] In this invention, users communicate via voice and text during meetings and daily work. The terminal acquires this conversation data in real time and converts it into text data using speech recognition technology.

[0129] The server receives the converted text data and uses natural language processing (NLP) techniques to analyze the conversation. This analysis extracts important information from the conversation and identifies the next necessary actions and decisions.

[0130] Furthermore, the server uses an emotion engine to recognize the user's emotional state from their voice tone and text content. For example, if the user is feeling stressed or exhibiting positive emotions, the emotion engine captures that state as data.

[0131] Based on this emotional data, the server can dynamically adjust task priorities. For example, if the server determines that a user is in a high-stress state, it will prioritize urgent tasks to be handled quickly and adjust task distribution as needed. Conversely, if positive emotions are detected, it can prioritize suggesting challenging and difficult tasks.

[0132] Furthermore, as a reminder management method, adjusting the content and frequency of reminders based on emotions enables task management that takes into account the user's mental state. This allows users to complete tasks within a manageable scope, contributing to maintaining and improving productivity.

[0133] For example, when a user is under high stress as a project deadline approaches, the server will frequently send reminders for important tasks to help the user process them smoothly. Conversely, when a project is progressing successfully and the user is feeling positive, the server will suggest reminders encouraging them to tackle more creative tasks or new projects.

[0134] Thus, the present invention is a system that provides advanced task management that takes into account the emotional state of individual users, thereby not only improving work efficiency but also enabling the optimization of the user's work-life balance.

[0135] The following describes the processing flow.

[0136] Step 1:

[0137] Users participate in meetings and communicate via voice or text. The meetings involve multiple participants, each offering diverse opinions.

[0138] Step 2:

[0139] The device records audio data of conversations in real time and converts that audio data into text using speech recognition technology. If necessary, it performs filtering to remove noise and attempts to identify the speaker.

[0140] Step 3:

[0141] The terminal sends the converted text data to the server. The text data includes decisions and action points discussed at the meeting.

[0142] Step 4:

[0143] The server analyzes the received text data using natural language processing (NLP). This analysis includes contextual understanding and keyword extraction, which extracts important information (e.g., next steps, responsible parties, deadlines).

[0144] Step 5:

[0145] The server utilizes an emotion engine to analyze voice tone and linguistic features in text to identify the user's emotional state. It then quantifies and evaluates the degree of each emotion, such as joy, anger, and stress.

[0146] Step 6:

[0147] Based on the analysis results, the server automatically generates a task list that reflects important information. The generated tasks are then assigned priorities that take into account the results of the emotion engine. Tasks that are highly important under high stress levels are adjusted, such as by lowering their priority.

[0148] Step 7:

[0149] The server manages task reminders, adjusting the content and timing of reminders based on the user's emotional state. If the user is stressed, the frequency of reminders is adjusted accordingly; if they are in a positive mood, the reminders encourage challenging tasks.

[0150] Step 8:

[0151] Users can use their devices to review and edit generated task lists and reminder content. This allows for flexible task management that adapts to the user's own work situation and emotional changes.

[0152] Step 9:

[0153] Throughout the entire process, the server monitors task progress and re-evaluates and adjusts tasks based on newly detected emotional states and environmental changes. This results in optimized task management and increased productivity.

[0154] (Example 2)

[0155] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0156] In today's work environment, efficiently managing busy schedules and multiple tasks is essential. However, traditional task management systems prioritize tasks without considering the user's emotional state, which can cause stress. Furthermore, the inability to manage tasks flexibly based on emotions is a challenge that leads to decreased productivity.

[0157] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0158] In this invention, the server includes an input means for acquiring conversation data, a conversion means for converting the acquired conversation data into text using speech recognition technology, an information analysis means for analyzing the converted text data and extracting important information, an emotion recognition means for determining the user's emotional state based on the analyzed information, a task priority setting means for generating a task list based on the extracted information and emotional state and dynamically setting the priority of each task, and a reminder management means for adjusting the content and frequency of reminders based on the progress of tasks and the user's emotional state. This enables flexible task management that takes into account the user's emotional state, thereby improving work efficiency and user satisfaction.

[0159] "Input means" refers to devices and methods used to acquire conversational data, such as collecting voice and text from users.

[0160] "Conversion means" refers to the technology or process for converting acquired audio data into text data, and utilizes speech recognition technology.

[0161] "Information analysis means" refers to technologies and methods for analyzing converted text data and extracting important information.

[0162] "Emotion recognition means" refers to technologies and devices used to determine a user's emotional state, detecting emotions based on text or audio data.

[0163] "Task priority setting means" refers to methods and techniques for generating a task list based on extracted information and emotional states, and for dynamically setting its priority.

[0164] A "reminder management method" refers to a method or technique for adjusting the content and frequency of reminders based on task progress and the user's emotional state.

[0165] A "generative AI model" refers to a model generated using artificial intelligence technology that outputs results for specific input data.

[0166] A "prompt statement" is a statement used to give instructions to a generative AI model, controlling the model's behavior and output results.

[0167] This invention is a task management system that uses conversational data and can dynamically adjust task priorities while considering the user's emotional state. First, the system acquires information provided by the user through conversation. The terminal uses input devices such as a microphone and keyboard to collect voice and text data in real time.

[0168] The collected conversation data is converted into text format by the device using speech recognition technology. This process utilizes common speech recognition software, such as open-source speech recognition APIs.

[0169] Subsequently, the server receives the converted text data and analyzes it using natural language processing (NLTK) technology. Specifically, it utilizes natural language processing tools such as spaCy and NLTK. Through this analysis, important points are extracted from the conversation, clarifying the actions and decisions that the user should take.

[0170] Next, the server uses emotion recognition technology to determine the user's emotional state. This determination utilizes an emotion engine that evaluates the user's emotions based on the context of their voice and text. The emotion engine detects states such as positive or stressful, and sends this data to the server.

[0171] Based on this information, the server dynamically sets task priorities using a generated AI model. If the user's emotions are stressful, the server will prioritize urgent tasks; if they are in a positive state, it can suggest challenging tasks.

[0172] Furthermore, a reminder function is built in, and the server adjusts the content and frequency of reminders according to the user's emotional state. This makes it possible to process tasks at a manageable pace.

[0173] For example, if it is determined that a user is experiencing stress, the system will automatically re-evaluate tasks by prompting the AI ​​model with the message, "Please reset the priority of the current task according to the stress level." This is expected to improve the user's work efficiency and optimize their work-life balance.

[0174] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0175] Step 1:

[0176] The user engages in conversation, which the device captures in real time as audio and text data. Input is the user's voice and text, and output is digital audio files and text data. The device captures this data using a microphone and keyboard.

[0177] Step 2:

[0178] The device converts the acquired audio data into text using a speech recognition engine. The input is audio data, and the output is the converted text data. The speech recognition software analyzes phonemes and converts them into strings of characters.

[0179] Step 3:

[0180] The server receives text data sent from the terminal. The input is the converted text data, and the output is the text dataset to be analyzed. The server buffers the received text in order to process this data further.

[0181] Step 4:

[0182] The server uses natural language processing techniques to analyze text data and extract important information. The input is a text dataset, and the output is a list of important information. The server performs operations such as tokenization, part-of-speech tagging, and keyword extraction to identify the necessary information.

[0183] Step 5:

[0184] The server analyzes the user's emotional state using an emotion recognition engine. Input consists of text data and speech tone information, while output is user emotional state data. The server performs contextual analysis of the text and speech tone detection to quantify the emotion.

[0185] Step 6:

[0186] The server dynamically sets task priorities using a generative AI model. The input is a list of key information and sentiment data, and the output is a task list with assigned priorities. The server evaluates the extracted information and sentiment data, and automatically sets task priorities based on prompt messages.

[0187] Step 7:

[0188] The server adjusts the content and frequency of reminders based on the user's emotional state. Inputs are a prioritized task list and emotional state data, while output is a customized reminder schedule. The server then develops a reminder delivery plan tailored to the user's situation.

[0189] (Application Example 2)

[0190] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0191] In today's workplace, it is crucial to understand employees' emotional states in real time and implement appropriate task management and security measures accordingly. However, traditional systems struggle to dynamically respond to these emotional changes, often relying on manual adjustments to security protocols and task priorities. This can compromise efficiency and security, posing a significant challenge.

[0192] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0193] In this invention, the server includes data collection means for acquiring conversation data, information analysis means for analyzing the acquired conversation data and extracting important information, and sentiment analysis means for analyzing the user's emotional state and dynamically adjusting security protocols. This enables automatic adjustment of task priorities and dynamic management of security levels according to the emotional state of employees.

[0194] "Data collection means" refers to devices and technologies for acquiring conversation data in real time.

[0195] "Information analysis means" refers to techniques that analyze acquired conversation data and extract important information such as text data.

[0196] A "task generation means" is a device or program that executes the process of generating a task list based on extracted important information.

[0197] A "priority setting method" is a technique for setting priorities for each task in a task list according to its importance and urgency.

[0198] A "reminder management system" is a mechanism that manages the content and sending method of reminders based on the progress and deadline of tasks.

[0199] "Emotional analysis methods" refer to technologies that analyze a user's emotional state and reflect it in the system's behavior and the information it presents.

[0200] "Security management measures" refer to devices or programs that appropriately manage and adjust the security level within a system based on the results of sentiment analysis.

[0201] To realize this invention, the server, terminal, and user collaborate to configure the system. The server is primarily responsible for data analysis and emotion recognition, while the terminal collects and transforms conversational data. Users simply need to perform their daily tasks and communicate naturally.

[0202] The server uses a speech recognition engine (e.g., Google Speech-to-Text) to convert conversational data sent from the terminal into text data. Since the audio data is processed in real time, rapid information analysis is possible. Next, the server uses a natural language processing library (e.g., spaCy) to extract important information from the text data. Based on this information, the task generation mechanism operates and generates a task list for the user.

[0203] Furthermore, the server uses an emotion recognition engine (e.g., IBM Watson® Tone Analyzer) to analyze the user's emotional state. Based on the emotional data, task priorities are dynamically adjusted, security management measures are activated, and security protocols are adjusted as needed.

[0204] For example, if a user is feeling tense or stressed during a meeting, the server can automatically instruct security staff to raise the alert level based on emotion analysis. Conversely, when the mood is relaxed, the system can revert to normal security measures to reduce unnecessary burden. In this way, the goal is to achieve both safety and efficiency in the workplace environment.

[0205] By utilizing generative AI models, it is possible to automatically extract necessary information and emotions from natural user conversations, providing efficient work support. As a result, users can experience improved work efficiency and enhanced safety.

[0206] An example of a prompt that can be input to a generative AI model is, "Help me design a system that captures conversations in real time, analyzes emotional states, and dynamically adjusts office security levels."

[0207] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0208] Step 1:

[0209] The device acquires user conversation data in real time through its built-in microphone. During this process, voice data is input. The device records the voice data in digital format and transmits it to the server via a secure connection.

[0210] Step 2:

[0211] The server converts the received audio data into text data using a speech recognition engine. At this point, the input is audio data and the output is text data. Specifically, the Google Speech-to-Text engine is used to convert the audio into words and add timestamps.

[0212] Step 3:

[0213] The server analyzes the converted text data using spaCy, a natural language processing library. The input is text data, and the output is information extracted with important words and keywords. Specifically, it tags parts of speech such as nouns, verbs, and adjectives, and detects significant topics.

[0214] Step 4:

[0215] The server generates a task list based on the analysis results. The input is a list of important information, and the output is a specific task list. In operation, it automatically assigns tasks related to the user's schedule and projects based on the extracted information.

[0216] Step 5:

[0217] The server uses an emotion recognition engine to analyze the user's emotional state from text data. In this process, the input is text data, and the output is an evaluation of the emotional state. Specifically, it analyzes tone and emotion through IBM Watson Tone Analyzer and classifies the user's emotions as positive, negative, neutral, etc.

[0218] Step 6:

[0219] The server adjusts security levels based on emotional states. The input is an assessment of emotional state, and the output is the adjusted security protocol. Specifically, if tension levels are high, the server may increase the surveillance intensity of security cameras or restrict access privileges.

[0220] Step 7:

[0221] The server generates personalized prompts based on the generated task list and sentiment data. The input is the task list and sentiment data, and the output is the prompt. Specifically, it utilizes a generative AI model to create sentences suggesting what the user should do next. These prompts are then sent to the user's smart device.

[0222] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0223] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0224] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0225] [Second Embodiment]

[0226] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0227] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0228] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0229] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0230] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0231] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0232] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0233] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0234] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0235] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0236] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0237] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0238] This invention describes a system that acquires conversational data generated during meetings and communications in real time, analyzes it, and automatically generates and manages important tasks. The program processing and specific examples of this system are described below.

[0239] In this system, users participate in meetings and generate conversation data through audio and chat data. The terminal converts this data into text data in real time using speech recognition technology, and also processes related chat logs and email data.

[0240] The server receives text data and related data sent from the terminal. The server then uses natural language processing (NLP) techniques to analyze the data and extract important information from the conversation. Specifically, it employs methods to identify decisions and the next actions to be taken.

[0241] Based on the extracted information, the server automatically generates a task list, listing important action items. The listed tasks are then accompanied by necessary details, referencing the user's past task history and project information.

[0242] For each item in the generated tasks, the server uses machine learning models and rule-based algorithms to set priorities. This makes it easier for users to determine which tasks they should tackle quickly.

[0243] Furthermore, the server sends reminders to the device at appropriate times based on the task's progress and deadline. These reminders prompt the user to take the next action, preventing tasks from being missed or delayed.

[0244] Furthermore, this system can be tailored to specific industries. The server uses pre-configured industry-specific models to perform analyses that address the needs and challenges of those particular industries. These specialized models are trained on industry-specific datasets and contribute to the automation of industry-specific tasks.

[0245] For example, if the user is a project manager in the IT industry, the server analyzes the audio data of the project meeting and automatically generates tasks such as "Prepare for next week's sprint review," setting them to a high priority. Furthermore, it supports smooth project progress by sending reminders to the terminal at appropriate times, such as "The deadline for creating sprint review materials is approaching."

[0246] Thus, the present invention provides users with an efficient task management system, enabling them to streamline business processes and improve productivity.

[0247] The following describes the processing flow.

[0248] Step 1:

[0249] Users participate in meetings and engage in conversations and discussions. Communication takes place using voice and chat.

[0250] Step 2:

[0251] The device records audio during meetings in real time and converts it into text data using speech recognition technology. To generate highly accurate text from the recorded audio, it performs ambient noise removal and speaker separation.

[0252] Step 3:

[0253] The terminal sends the generated text data, chat logs used during the meeting, and email data as needed to the server. This transmitted data is then ready for processing by the system.

[0254] Step 4:

[0255] The server analyzes the received text data using a natural language processing (NLP) engine. Through understanding grammatical structure and vocabulary, it extracts important information from the context of the conversation, clarifying decisions and next actions.

[0256] Step 5:

[0257] The server automatically generates a task list based on the extracted information. This list specifically lists the required tasks, and each task includes relevant details. Furthermore, it refers to the user's past work history and supplements task details as needed.

[0258] Step 6:

[0259] The server sets the priority of each task based on its importance and deadline. Using machine learning algorithms, it evaluates the urgency and relative priority of tasks and indicates which tasks the user should tackle first.

[0260] Step 7:

[0261] The server monitors the progress of each task and sends reminders to the user's device based on the deadline. This ensures that users receive timely notifications and prevent tasks from being missed.

[0262] Step 8:

[0263] The servers are specialized for specific industries, employing industry-specific models accordingly. This allows them to address the diverse needs of each industry and achieve specialized task management.

[0264] Step 9:

[0265] Users can view generated task lists and reminders through the device interface and customize the interface to optimize task management to suit their work style. This makes it easier to optimize task placement and management for each industry and team.

[0266] (Example 1)

[0267] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0268] Traditional meeting management systems struggled to extract and properly manage important tasks from conversation content. Furthermore, their lack of flexibility to meet the unique needs of different industries hindered efficient business operations.

[0269] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0270] In this invention, the server includes data collection means for acquiring conversation information, information analysis means for analyzing the acquired conversation information and extracting important matters, and task generation means for generating a list of tasks based on the extracted important matters. This makes it possible to immediately identify and manage tasks from conversations in meetings.

[0271] "Conversational information" refers to information generated as audio or text during meetings and communication sessions.

[0272] "Data collection means" refers to a function or device for efficiently acquiring conversational information and incorporating it into a system.

[0273] "Information analysis means" refers to the techniques and methods used to analyze acquired conversational information and extract important information from it.

[0274] "Task generation means" refers to a function for automatically generating tasks based on important information extracted by information analysis means.

[0275] A "prioritization method" refers to a technique for determining which tasks should be prioritized among the generated tasks and setting their order.

[0276] "Notification management means" refers to a function that provides notifications at the appropriate time based on the progress and deadlines of a task, prompting the user to take the next action.

[0277] "Industry-specific methods" refer to technologies and methods for customizing systems to address the specific needs and challenges of a particular industry.

[0278] A "generative AI model" refers to a group of algorithms that use artificial intelligence technology to generate useful information or answers from input data.

[0279] "Prompt generation means" refers to a function or technology for generating natural language questions and commands from conversational information.

[0280] This invention is a system for efficiently processing conversational information generated in meetings and communications, and for achieving automated task generation and management. The system includes, as its components, data collection means, information analysis means, task generation means, priority setting means, notification management means, industry-specific means, generation AI model, and prompt generation means.

[0281] Users participate in meetings and generate conversational information in the form of voice or chat. The device collects this information using a microphone and corresponding software (e.g., a video conferencing application). The collected audio data is converted into text using a common speech recognition API (a comprehensive and widely used API as an example) as a voice identification technology.

[0282] The server analyzes the text information sent from the terminal using a natural language processing library (e.g., SpaCy or OpenNLP). It then extracts important information and uses a generative AI model to generate appropriate prompts. These prompts provide specific, actionable instructions for the generated task list. For example, a possible prompt might be, "List the preparation items for next week's team meeting."

[0283] Furthermore, the server prioritizes the generated tasks through machine learning models and rule-based algorithms. Based on the set priorities, it sends notifications via the terminal at appropriate times according to the progress and deadlines of the work. This prompts users to take the next action, thereby promoting work efficiency.

[0284] In addition, the server performs analysis using a specialized model for a specific industry. By means of industry specialization, analysis can be performed in accordance with the needs of each industry, and specific tasks and information can be provided. By combining these elements, it is possible to achieve efficient task management and optimization of business processes.

[0285] The flow of the specific process in Example 1 will be described using FIG. 11.

[0286] Step 1:

[0287] The user participates in the meeting and generates conversation information in the form of voice and chat. As a result, voice data and text data are obtained as inputs. The terminal collects this conversation information in real time through a microphone or video conferencing software.

[0288] Step 2:

[0289] The terminal converts the collected voice data into text using voice recognition technology. Specifically, a process of converting voice data into text data is performed through a voice recognition API. The input is voice data, and the output is text data. The content of the conversation is recorded as text in this text data.

[0290] Step 3:

[0291] The terminal sends the converted text data and related chat logs to the server. The server receives these data and prepares for analysis. The inputs are text data and chat logs, which are used in the next analysis process.

[0292] Step 4:

[0293] The server analyzes the received text data using natural language processing technology. Specifically, an NLP library is used to extract important information and keywords from the text. The input is text data, and the output is important information as the analysis result.

[0294] Step 5:

[0295] Based on the analysis results, the server operates a prompt generation mechanism using a generation AI model. A specific list of tasks is then generated based on the resulting prompt messages. The input is important information, and the output consists of the generated prompt messages and the corresponding task items.

[0296] Step 6:

[0297] The server applies machine learning models and rule-based algorithms to prioritize work items and determine task priorities. The input is work items, and the output is a list of tasks with their priorities set.

[0298] Step 7:

[0299] The server uses a prioritized task list to send notifications to the terminal at appropriate times based on the progress and deadlines of each task. These notifications are intended to prompt the user to take the next action. The input is the task list and the progress of the tasks, and the output is the notification sent to the user.

[0300] (Application Example 1)

[0301] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0302] In brick-and-mortar retail operations, efficiently managing tasks arising from customer interactions and staff conversations is essential. However, traditional methods require manually recording conversations and manually organizing tasks, hindering operational efficiency. Furthermore, the insufficient use of specific information technology devices for visual task presentation leads to challenges in immediacy and accuracy.

[0303] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following means.

[0304] In this invention, the server includes an information collection means for acquiring conversation data, a data analysis means for analyzing the acquired conversation data to extract important information, and a display means for visually presenting task information to the user using a specific information technology device. Thereby, it is possible to improve customer service and business efficiency, and enable real-time task management and information presentation.

[0305] The "information collection means" is a mechanism for acquiring information such as conversations and voice data in real time.

[0306] The "data analysis means" is a technology having a function of analyzing the acquired information and extracting important data and tasks.

[0307] The "task creation means" is a method for automatically generating a necessary task list based on the analyzed data.

[0308] The "priority setting means" is a mechanism for determining the priority of the generated tasks according to their importance and urgency.

[0309] The "notification management means" is a technology having the ability to send reminders to the user based on the progress status and deadlines of tasks.

[0310] The "industry specialization means" is a method for adjusting the analysis and task generation functions according to a specific industrial field.

[0311] The "display means" is a mechanism for visually presenting task information to the user using a specific information technology device.

[0312] This invention is a system that enables smooth task management by allowing staff to use smart glasses in physical stores to analyze customer interactions and conversations among staff in real time. The server uses a specific voice recognition device as a means of information gathering to collect voice data obtained from the field. Subsequently, the data is analyzed by the server using a data analysis means that employs natural language processing technology to extract important tasks and information. Specifically, using the Hugging Face Transformers library is effective.

[0313] The server also generates a task list based on information extracted by the task creation mechanism and determines the importance of each task using the priority setting mechanism. This information is immediately presented to staff via a specific information technology device, such as smart glasses. This allows users to efficiently manage their work within the physical store.

[0314] Furthermore, reminders are displayed on staff members' smart glasses via notification management tools, based on task progress and deadlines. Industry-specific tools can also be used to provide features optimized for specific store types or industries.

[0315] For example, if a store clerk says to a customer, "This item is sold out. Can you restock it?", a task such as "Restock sold-out item X" will be automatically generated. Furthermore, by utilizing the generative AI model and using a prompt such as, "Analyze the following conversation and extract the important tasks. Conversation: 'This item is sold out. Can you restock it?'", more accurate task analysis becomes possible.

[0316] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0317] Step 1:

[0318] The terminal acquires audio data of staff conversations and customer interactions within the store. This audio data is captured through a built-in speech recognition device in the terminal. The input is conversations within the store, and the output is audio data.

[0319] Step 2:

[0320] The device converts the acquired audio data into text data. This conversion process utilizes speech recognition technology. Specifically, it uses speech recognition software such as the Google Speech API. The input is audio data, and the output is the converted text data.

[0321] Step 3:

[0322] The server receives text data sent from the terminal. During this receiving process, the server uses natural language processing techniques to identify and extract important information. The input is text data, and the output is the analyzed important information. The accuracy of the analysis can be improved by using a generative AI model.

[0323] Step 4:

[0324] The server generates a task list and sets priorities based on the extracted information. This uses machine learning algorithms and also considers the user's past task history. The input is the analyzed key information, and the output is a task list with priorities set.

[0325] Step 5:

[0326] The server sends the generated task list to the terminal, where it is displayed to the staff via smart glasses. This allows users to visually check tasks in real time and respond immediately. The input is a task list with priorities set, and the output is the task list displayed on the smart glasses.

[0327] Step 6:

[0328] The server notifies the user's device of the task progress and deadline at the appropriate time. This prevents the user from missing or delaying tasks. Input is the task list and task management information, and output is the reminder displayed on the device.

[0329] Through the above processing steps, it is possible to improve the efficiency of task management in physical stores.

[0330] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0331] This invention combines an emotion engine with a system that acquires conversation data in real time, analyzes the content of that conversation, and generates tasks, thereby enabling task management that responds to the user's emotions. The program processing and specific examples of this system are described below.

[0332] In this invention, users communicate via voice and text during meetings and daily work. The terminal acquires this conversation data in real time and converts it into text data using speech recognition technology.

[0333] The server receives the converted text data and uses natural language processing (NLP) techniques to analyze the conversation. This analysis extracts important information from the conversation and identifies the next necessary actions and decisions.

[0334] Furthermore, the server uses an emotion engine to recognize the user's emotional state from their voice tone and text content. For example, if the user is feeling stressed or exhibiting positive emotions, the emotion engine captures that state as data.

[0335] Based on this emotional data, the server can dynamically adjust task priorities. For example, if the server determines that a user is in a high-stress state, it will prioritize urgent tasks to be handled quickly and adjust task distribution as needed. Conversely, if positive emotions are detected, it can prioritize suggesting challenging and difficult tasks.

[0336] Furthermore, as a reminder management method, adjusting the content and frequency of reminders based on emotions enables task management that takes into account the user's mental state. This allows users to complete tasks within a manageable scope, contributing to maintaining and improving productivity.

[0337] For example, when a user is under high stress as a project deadline approaches, the server will frequently send reminders for important tasks to help the user process them smoothly. Conversely, when a project is progressing successfully and the user is feeling positive, the server will suggest reminders encouraging them to tackle more creative tasks or new projects.

[0338] Thus, the present invention is a system that provides advanced task management that takes into account the emotional state of individual users, thereby not only improving work efficiency but also enabling the optimization of the user's work-life balance.

[0339] The following describes the processing flow.

[0340] Step 1:

[0341] Users participate in meetings and communicate via voice or text. The meetings involve multiple participants, each offering diverse opinions.

[0342] Step 2:

[0343] The device records audio data of conversations in real time and converts that audio data into text using speech recognition technology. If necessary, it performs filtering to remove noise and attempts to identify the speaker.

[0344] Step 3:

[0345] The terminal sends the converted text data to the server. The text data includes decisions and action points discussed at the meeting.

[0346] Step 4:

[0347] The server analyzes the received text data using natural language processing (NLP). This analysis includes contextual understanding and keyword extraction, which extracts important information (e.g., next steps, responsible parties, deadlines).

[0348] Step 5:

[0349] The server utilizes an emotion engine to analyze voice tone and linguistic features in text to identify the user's emotional state. It then quantifies and evaluates the degree of each emotion, such as joy, anger, and stress.

[0350] Step 6:

[0351] Based on the analysis results, the server automatically generates a task list that reflects important information. The generated tasks are then assigned priorities that take into account the results of the emotion engine. Tasks that are highly important under high stress levels are adjusted, such as by lowering their priority.

[0352] Step 7:

[0353] The server manages task reminders, adjusting the content and timing of reminders based on the user's emotional state. If the user is stressed, the frequency of reminders is adjusted accordingly; if they are in a positive mood, the reminders encourage challenging tasks.

[0354] Step 8:

[0355] Users can use their devices to review and edit generated task lists and reminder content. This allows for flexible task management that adapts to the user's own work situation and emotional changes.

[0356] Step 9:

[0357] Throughout the entire process, the server monitors task progress and re-evaluates and adjusts tasks based on newly detected emotional states and environmental changes. This results in optimized task management and increased productivity.

[0358] (Example 2)

[0359] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0360] In today's work environment, efficiently managing busy schedules and multiple tasks is essential. However, traditional task management systems prioritize tasks without considering the user's emotional state, which can cause stress. Furthermore, the inability to manage tasks flexibly based on emotions is a challenge that leads to decreased productivity.

[0361] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0362] In this invention, the server includes an input means for acquiring conversation data, a conversion means for converting the acquired conversation data into text using speech recognition technology, an information analysis means for analyzing the converted text data and extracting important information, an emotion recognition means for determining the user's emotional state based on the analyzed information, a task priority setting means for generating a task list based on the extracted information and emotional state and dynamically setting the priority of each task, and a reminder management means for adjusting the content and frequency of reminders based on the progress of tasks and the user's emotional state. This enables flexible task management that takes into account the user's emotional state, thereby improving work efficiency and user satisfaction.

[0363] "Input means" refers to devices and methods used to acquire conversational data, such as collecting voice and text from users.

[0364] "Conversion means" refers to the technology or process for converting acquired audio data into text data, and utilizes speech recognition technology.

[0365] "Information analysis means" refers to technologies and methods for analyzing converted text data and extracting important information.

[0366] "Emotion recognition means" refers to technologies and devices used to determine a user's emotional state, detecting emotions based on text or audio data.

[0367] "Task priority setting means" refers to methods and techniques for generating a task list based on extracted information and emotional states, and for dynamically setting its priority.

[0368] A "reminder management method" refers to a method or technique for adjusting the content and frequency of reminders based on task progress and the user's emotional state.

[0369] A "generative AI model" refers to a model generated using artificial intelligence technology that outputs results for specific input data.

[0370] A "prompt statement" is a statement used to give instructions to a generative AI model, controlling the model's behavior and output results.

[0371] This invention is a task management system that uses conversational data and can dynamically adjust task priorities while considering the user's emotional state. First, the system acquires information provided by the user through conversation. The terminal uses input devices such as a microphone and keyboard to collect voice and text data in real time.

[0372] The collected conversation data is converted into text format by the device using speech recognition technology. This process utilizes common speech recognition software, such as open-source speech recognition APIs.

[0373] Subsequently, the server receives the converted text data and analyzes it using natural language processing (NLTK) technology. Specifically, it utilizes natural language processing tools such as spaCy and NLTK. Through this analysis, important points are extracted from the conversation, clarifying the actions and decisions that the user should take.

[0374] Next, the server uses emotion recognition technology to determine the user's emotional state. This determination utilizes an emotion engine that evaluates the user's emotions based on the context of their voice and text. The emotion engine detects states such as positive or stressful, and sends this data to the server.

[0375] Based on this information, the server dynamically sets task priorities using a generated AI model. If the user's emotions are stressful, the server will prioritize urgent tasks; if they are in a positive state, it can suggest challenging tasks.

[0376] Furthermore, a reminder function is built in, and the server adjusts the content and frequency of reminders according to the user's emotional state. This makes it possible to process tasks at a manageable pace.

[0377] For example, if it is determined that a user is experiencing stress, the system will automatically re-evaluate tasks by prompting the AI ​​model with the message, "Please reset the priority of the current task according to the stress level." This is expected to improve the user's work efficiency and optimize their work-life balance.

[0378] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0379] Step 1:

[0380] The user engages in conversation, which the device captures in real time as audio and text data. Input is the user's voice and text, and output is digital audio files and text data. The device captures this data using a microphone and keyboard.

[0381] Step 2:

[0382] The device converts the acquired audio data into text using a speech recognition engine. The input is audio data, and the output is the converted text data. The speech recognition software analyzes phonemes and converts them into strings of characters.

[0383] Step 3:

[0384] The server receives text data sent from the terminal. The input is the converted text data, and the output is the text dataset to be analyzed. The server buffers the received text in order to process this data further.

[0385] Step 4:

[0386] The server uses natural language processing techniques to analyze text data and extract important information. The input is a text dataset, and the output is a list of important information. The server performs operations such as tokenization, part-of-speech tagging, and keyword extraction to identify the necessary information.

[0387] Step 5:

[0388] The server analyzes the user's emotional state using an emotion recognition engine. Input consists of text data and speech tone information, while output is user emotional state data. The server performs contextual analysis of the text and speech tone detection to quantify the emotion.

[0389] Step 6:

[0390] The server dynamically sets task priorities using a generative AI model. The input is a list of key information and sentiment data, and the output is a task list with assigned priorities. The server evaluates the extracted information and sentiment data, and automatically sets task priorities based on prompt messages.

[0391] Step 7:

[0392] The server adjusts the content and frequency of reminders based on the user's emotional state. Inputs are a prioritized task list and emotional state data, while output is a customized reminder schedule. The server then develops a reminder delivery plan tailored to the user's situation.

[0393] (Application Example 2)

[0394] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0395] In today's workplace, it is crucial to understand employees' emotional states in real time and implement appropriate task management and security measures accordingly. However, traditional systems struggle to dynamically respond to these emotional changes, often relying on manual adjustments to security protocols and task priorities. This can compromise efficiency and security, posing a significant challenge.

[0396] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0397] In this invention, the server includes data collection means for acquiring conversation data, information analysis means for analyzing the acquired conversation data and extracting important information, and sentiment analysis means for analyzing the user's emotional state and dynamically adjusting security protocols. This enables automatic adjustment of task priorities and dynamic management of security levels according to the emotional state of employees.

[0398] "Data collection means" refers to devices and technologies for acquiring conversation data in real time.

[0399] "Information analysis means" refers to techniques that analyze acquired conversation data and extract important information such as text data.

[0400] A "task generation means" is a device or program that executes the process of generating a task list based on extracted important information.

[0401] A "priority setting method" is a technique for setting priorities for each task in a task list according to its importance and urgency.

[0402] A "reminder management system" is a mechanism that manages the content and sending method of reminders based on the progress and deadline of tasks.

[0403] "Emotional analysis methods" refer to technologies that analyze a user's emotional state and reflect it in the system's behavior and the information it presents.

[0404] "Security management measures" refer to devices or programs that appropriately manage and adjust the security level within a system based on the results of sentiment analysis.

[0405] To realize this invention, the server, terminal, and user collaborate to configure the system. The server is primarily responsible for data analysis and emotion recognition, while the terminal collects and transforms conversational data. Users simply need to perform their daily tasks and communicate naturally.

[0406] The server uses a speech recognition engine (e.g., Google Speech-to-Text) to convert conversational data sent from the terminal into text data. Since the audio data is processed in real time, rapid information analysis is possible. Next, the server uses a natural language processing library (e.g., spaCy) to extract important information from the text data. Based on this information, the task generation mechanism operates and generates a task list for the user.

[0407] Furthermore, the server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotional state. Based on the emotional data, task priorities are dynamically adjusted, security management mechanisms are activated, and security protocols are adjusted as needed.

[0408] For example, if a user is feeling tense or stressed during a meeting, the server can automatically instruct security staff to raise the alert level based on emotion analysis. Conversely, when the mood is relaxed, the system can revert to normal security measures to reduce unnecessary burden. In this way, the goal is to achieve both safety and efficiency in the workplace environment.

[0409] By utilizing generative AI models, it is possible to automatically extract necessary information and emotions from natural user conversations, providing efficient work support. As a result, users can experience improved work efficiency and enhanced safety.

[0410] An example of a prompt that can be input to a generative AI model is, "Help me design a system that captures conversations in real time, analyzes emotional states, and dynamically adjusts office security levels."

[0411] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0412] Step 1:

[0413] The device acquires user conversation data in real time through its built-in microphone. During this process, voice data is input. The device records the voice data in digital format and transmits it to the server via a secure connection.

[0414] Step 2:

[0415] The server converts the received audio data into text data using a speech recognition engine. At this point, the input is audio data and the output is text data. Specifically, the Google Speech-to-Text engine is used to convert the audio into words and add timestamps.

[0416] Step 3:

[0417] The server analyzes the converted text data using spaCy, a natural language processing library. The input is text data, and the output is information extracted with important words and keywords. Specifically, it tags parts of speech such as nouns, verbs, and adjectives, and detects significant topics.

[0418] Step 4:

[0419] The server generates a task list based on the analysis results. The input is a list of important information, and the output is a specific task list. In operation, it automatically assigns tasks related to the user's schedule and projects based on the extracted information.

[0420] Step 5:

[0421] The server uses an emotion recognition engine to analyze the user's emotional state from text data. In this process, the input is text data, and the output is an evaluation of the emotional state. Specifically, it analyzes tone and emotion through IBM Watson Tone Analyzer and classifies the user's emotions as positive, negative, neutral, etc.

[0422] Step 6:

[0423] The server adjusts security levels based on emotional states. The input is an assessment of emotional state, and the output is the adjusted security protocol. Specifically, if tension levels are high, the server may increase the surveillance intensity of security cameras or restrict access privileges.

[0424] Step 7:

[0425] The server generates personalized prompts based on the generated task list and sentiment data. The input is the task list and sentiment data, and the output is the prompt. Specifically, it utilizes a generative AI model to create sentences suggesting what the user should do next. These prompts are then sent to the user's smart device.

[0426] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0427] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0428] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0429] [Third Embodiment]

[0430] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0431] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0432] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0433] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0434] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0435] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0436] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0437] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0438] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0439] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0440] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0441] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0442] This invention describes a system that acquires conversational data generated during meetings and communications in real time, analyzes it, and automatically generates and manages important tasks. The program processing and specific examples of this system are described below.

[0443] In this system, users participate in meetings and generate conversation data through audio and chat data. The terminal converts this data into text data in real time using speech recognition technology, and also processes related chat logs and email data.

[0444] The server receives text data and related data sent from the terminal. The server then uses natural language processing (NLP) techniques to analyze the data and extract important information from the conversation. Specifically, it employs methods to identify decisions and the next actions to be taken.

[0445] Based on the extracted information, the server automatically generates a task list, listing important action items. The listed tasks are then accompanied by necessary details, referencing the user's past task history and project information.

[0446] For each item in the generated tasks, the server uses machine learning models and rule-based algorithms to set priorities. This makes it easier for users to determine which tasks they should tackle quickly.

[0447] Furthermore, the server sends reminders to the device at appropriate times based on the task's progress and deadline. These reminders prompt the user to take the next action, preventing tasks from being missed or delayed.

[0448] Furthermore, this system can be tailored to specific industries. The server uses pre-configured industry-specific models to perform analyses that address the needs and challenges of those particular industries. These specialized models are trained on industry-specific datasets and contribute to the automation of industry-specific tasks.

[0449] For example, if the user is a project manager in the IT industry, the server analyzes the audio data of the project meeting and automatically generates tasks such as "Prepare for next week's sprint review," setting them to a high priority. Furthermore, it supports smooth project progress by sending reminders to the terminal at appropriate times, such as "The deadline for creating sprint review materials is approaching."

[0450] Thus, the present invention provides users with an efficient task management system, enabling them to streamline business processes and improve productivity.

[0451] The following describes the processing flow.

[0452] Step 1:

[0453] Users participate in meetings and engage in conversations and discussions. Communication takes place using voice and chat.

[0454] Step 2:

[0455] The device records audio during meetings in real time and converts it into text data using speech recognition technology. To generate highly accurate text from the recorded audio, it performs ambient noise removal and speaker separation.

[0456] Step 3:

[0457] The terminal sends the generated text data, chat logs used during the meeting, and email data as needed to the server. This transmitted data is then ready for processing by the system.

[0458] Step 4:

[0459] The server analyzes the received text data using a natural language processing (NLP) engine. Through understanding grammatical structure and vocabulary, it extracts important information from the context of the conversation, clarifying decisions and next actions.

[0460] Step 5:

[0461] The server automatically generates a task list based on the extracted information. This list specifically lists the required tasks, and each task includes relevant details. Furthermore, it refers to the user's past work history and supplements task details as needed.

[0462] Step 6:

[0463] The server sets the priority of each task based on its importance and deadline. Using machine learning algorithms, it evaluates the urgency and relative priority of tasks and indicates which tasks the user should tackle first.

[0464] Step 7:

[0465] The server monitors the progress of each task and sends reminders to the user's device based on the deadline. This ensures that users receive timely notifications and prevent tasks from being missed.

[0466] Step 8:

[0467] The servers are specialized for specific industries, employing industry-specific models accordingly. This allows them to address the diverse needs of each industry and achieve specialized task management.

[0468] Step 9:

[0469] Users can view generated task lists and reminders through the device interface and customize the interface to optimize task management to suit their work style. This makes it easier to optimize task placement and management for each industry and team.

[0470] (Example 1)

[0471] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0472] Traditional meeting management systems struggled to extract and properly manage important tasks from conversation content. Furthermore, their lack of flexibility to meet the unique needs of different industries hindered efficient business operations.

[0473] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0474] In this invention, the server includes data collection means for acquiring conversation information, information analysis means for analyzing the acquired conversation information and extracting important matters, and task generation means for generating a list of tasks based on the extracted important matters. This makes it possible to immediately identify and manage tasks from conversations in meetings.

[0475] "Conversational information" refers to information generated as audio or text during meetings and communication sessions.

[0476] "Data collection means" refers to a function or device for efficiently acquiring conversational information and incorporating it into a system.

[0477] "Information analysis means" refers to the techniques and methods used to analyze acquired conversational information and extract important information from it.

[0478] "Task generation means" refers to a function for automatically generating tasks based on important information extracted by information analysis means.

[0479] A "prioritization method" refers to a technique for determining which tasks should be prioritized among the generated tasks and setting their order.

[0480] "Notification management means" refers to a function that provides notifications at the appropriate time based on the progress and deadlines of a task, prompting the user to take the next action.

[0481] "Industry-specific methods" refer to technologies and methods for customizing systems to address the specific needs and challenges of a particular industry.

[0482] A "generative AI model" refers to a group of algorithms that use artificial intelligence technology to generate useful information or answers from input data.

[0483] "Prompt generation means" refers to a function or technology for generating natural language questions and commands from conversational information.

[0484] This invention is a system for efficiently processing conversational information generated in meetings and communications, and for achieving automated task generation and management. The system includes, as its components, data collection means, information analysis means, task generation means, priority setting means, notification management means, industry-specific means, generation AI model, and prompt generation means.

[0485] Users participate in meetings and generate conversational information in the form of voice or chat. The device collects this information using a microphone and corresponding software (e.g., a video conferencing application). The collected audio data is converted into text using a common speech recognition API (a comprehensive and widely used API as an example) as a voice identification technology.

[0486] The server analyzes the text information sent from the terminal using a natural language processing library (e.g., SpaCy or OpenNLP). It then extracts important information and uses a generative AI model to generate appropriate prompts. These prompts provide specific, actionable instructions for the generated task list. For example, a possible prompt might be, "List the preparation items for next week's team meeting."

[0487] Furthermore, the server prioritizes the generated tasks through machine learning models and rule-based algorithms. Based on the set priorities, it sends notifications via the terminal at appropriate times according to the progress and deadlines of the work. This prompts users to take the next action, thereby promoting work efficiency.

[0488] Furthermore, the server performs analysis using specialized models for specific industries. This industry-specific approach allows for analysis tailored to the needs of each industry, providing unique tasks and information. Combining these elements makes it possible to achieve efficient task management and optimization of business processes.

[0489] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0490] Step 1:

[0491] Users participate in meetings and generate conversational information in voice and chat formats. This results in voice and text data as input. The device collects this conversational information in real time through microphones and video conferencing software.

[0492] Step 2:

[0493] The device converts collected audio data into text using speech recognition technology. Specifically, it performs the process of converting audio data into text data via a speech recognition API. The input is audio data, and the output is text data. This text data records the content of the conversation as characters.

[0494] Step 3:

[0495] The terminal sends the converted text data and associated chat logs to the server. The server receives this data and prepares it for analysis. The input consists of text data and chat logs, which are used for the subsequent analysis process.

[0496] Step 4:

[0497] The server analyzes the received text data using natural language processing (NLP) techniques. Specifically, it uses an NLP library to extract important information and keywords from the text. The input is text data, and the output is important information as a result of the analysis.

[0498] Step 5:

[0499] Based on the analysis results, the server operates a prompt generation mechanism using a generation AI model. A specific list of tasks is then generated based on the resulting prompt messages. The input is important information, and the output consists of the generated prompt messages and the corresponding task items.

[0500] Step 6:

[0501] The server applies machine learning models and rule-based algorithms to prioritize work items and determine task priorities. The input is work items, and the output is a list of tasks with their priorities set.

[0502] Step 7:

[0503] The server uses a prioritized task list to send notifications to the terminal at appropriate times based on the progress and deadlines of each task. These notifications are intended to prompt the user to take the next action. The input is the task list and the progress of the tasks, and the output is the notification sent to the user.

[0504] (Application Example 1)

[0505] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0506] In brick-and-mortar retail operations, efficiently managing tasks arising from customer interactions and staff conversations is essential. However, traditional methods require manually recording conversations and manually organizing tasks, hindering operational efficiency. Furthermore, the insufficient use of specific information technology devices for visual task presentation leads to challenges in immediacy and accuracy.

[0507] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0508] In this invention, the server includes information gathering means for acquiring conversation data, data analysis means for analyzing the acquired conversation data and extracting important information, and display means for visually presenting task information to the user using a specific information technology device. This improves customer service and operational efficiency, enabling real-time task management and information presentation.

[0509] "Information gathering means" refers to a system for acquiring information such as conversations and audio data in real time.

[0510] "Data analysis tools" are technologies that analyze acquired information and extract important data and tasks.

[0511] A "task creation method" is a method for automatically generating a list of necessary tasks based on analyzed data.

[0512] A "prioritization mechanism" is a system that determines the priority of generated tasks according to their importance and urgency.

[0513] A "notification management system" is a technology that has the ability to send reminders to users based on the progress and deadlines of tasks.

[0514] "Industry-specific methods" refer to methods for adjusting analysis and task generation functions to suit specific industrial sectors.

[0515] A "display means" is a mechanism that uses a specific information technology device to visually present task information to the user.

[0516] This invention is a system that enables smooth task management by allowing staff to use smart glasses in physical stores to analyze customer interactions and conversations among staff in real time. The server uses a specific voice recognition device as a means of information gathering to collect voice data obtained from the field. Subsequently, the data is analyzed by the server using a data analysis means that employs natural language processing technology to extract important tasks and information. Specifically, using the Hugging Face Transformers library is effective.

[0517] The server also generates a task list based on information extracted by the task creation mechanism and determines the importance of each task using the priority setting mechanism. This information is immediately presented to staff via a specific information technology device, such as smart glasses. This allows users to efficiently manage their work within the physical store.

[0518] Furthermore, reminders are displayed on staff members' smart glasses via notification management tools, based on task progress and deadlines. Industry-specific tools can also be used to provide features optimized for specific store types or industries.

[0519] For example, if a store clerk says to a customer, "This item is sold out. Can you restock it?", a task such as "Restock sold-out item X" will be automatically generated. Furthermore, by utilizing the generative AI model and using a prompt such as, "Analyze the following conversation and extract the important tasks. Conversation: 'This item is sold out. Can you restock it?'", more accurate task analysis becomes possible.

[0520] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0521] Step 1:

[0522] The terminal acquires audio data of staff conversations and customer interactions within the store. This audio data is captured through a built-in speech recognition device in the terminal. The input is conversations within the store, and the output is audio data.

[0523] Step 2:

[0524] The device converts the acquired audio data into text data. This conversion process utilizes speech recognition technology. Specifically, it uses speech recognition software such as the Google Speech API. The input is audio data, and the output is the converted text data.

[0525] Step 3:

[0526] The server receives text data sent from the terminal. During this receiving process, the server uses natural language processing techniques to identify and extract important information. The input is text data, and the output is the analyzed important information. The accuracy of the analysis can be improved by using a generative AI model.

[0527] Step 4:

[0528] The server generates a task list and sets priorities based on the extracted information. This uses machine learning algorithms and also considers the user's past task history. The input is the analyzed key information, and the output is a task list with priorities set.

[0529] Step 5:

[0530] The server sends the generated task list to the terminal, where it is displayed to the staff via smart glasses. This allows users to visually check tasks in real time and respond immediately. The input is a task list with priorities set, and the output is the task list displayed on the smart glasses.

[0531] Step 6:

[0532] The server notifies the user's device of the task progress and deadline at the appropriate time. This prevents the user from missing or delaying tasks. Input is the task list and task management information, and output is the reminder displayed on the device.

[0533] Through the above processing steps, it is possible to improve the efficiency of task management in physical stores.

[0534] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0535] This invention combines an emotion engine with a system that acquires conversation data in real time, analyzes the content of that conversation, and generates tasks, thereby enabling task management that responds to the user's emotions. The program processing and specific examples of this system are described below.

[0536] In this invention, users communicate via voice and text during meetings and daily work. The terminal acquires this conversation data in real time and converts it into text data using speech recognition technology.

[0537] The server receives the converted text data and uses natural language processing (NLP) techniques to analyze the conversation. This analysis extracts important information from the conversation and identifies the next necessary actions and decisions.

[0538] Furthermore, the server uses an emotion engine to recognize the user's emotional state from their voice tone and text content. For example, if the user is feeling stressed or exhibiting positive emotions, the emotion engine captures that state as data.

[0539] Based on this emotional data, the server can dynamically adjust task priorities. For example, if the server determines that a user is in a high-stress state, it will prioritize urgent tasks to be handled quickly and adjust task distribution as needed. Conversely, if positive emotions are detected, it can prioritize suggesting challenging and difficult tasks.

[0540] Furthermore, as a reminder management method, adjusting the content and frequency of reminders based on emotions enables task management that takes into account the user's mental state. This allows users to complete tasks within a manageable scope, contributing to maintaining and improving productivity.

[0541] For example, when a user is under high stress as a project deadline approaches, the server will frequently send reminders for important tasks to help the user process them smoothly. Conversely, when a project is progressing successfully and the user is feeling positive, the server will suggest reminders encouraging them to tackle more creative tasks or new projects.

[0542] Thus, the present invention is a system that provides advanced task management that takes into account the emotional state of individual users, thereby not only improving work efficiency but also enabling the optimization of the user's work-life balance.

[0543] The following describes the processing flow.

[0544] Step 1:

[0545] Users participate in meetings and communicate via voice or text. The meetings involve multiple participants, each offering diverse opinions.

[0546] Step 2:

[0547] The device records audio data of conversations in real time and converts that audio data into text using speech recognition technology. If necessary, it performs filtering to remove noise and attempts to identify the speaker.

[0548] Step 3:

[0549] The terminal sends the converted text data to the server. The text data includes decisions and action points discussed at the meeting.

[0550] Step 4:

[0551] The server analyzes the received text data using natural language processing (NLP). This analysis includes contextual understanding and keyword extraction, which extracts important information (e.g., next steps, responsible parties, deadlines).

[0552] Step 5:

[0553] The server utilizes an emotion engine to analyze voice tone and linguistic features in text to identify the user's emotional state. It then quantifies and evaluates the degree of each emotion, such as joy, anger, and stress.

[0554] Step 6:

[0555] Based on the analysis results, the server automatically generates a task list that reflects important information. The generated tasks are then assigned priorities that take into account the results of the emotion engine. Tasks that are highly important under high stress levels are adjusted, such as by lowering their priority.

[0556] Step 7:

[0557] The server manages task reminders, adjusting the content and timing of reminders based on the user's emotional state. If the user is stressed, the frequency of reminders is adjusted accordingly; if they are in a positive mood, the reminders encourage challenging tasks.

[0558] Step 8:

[0559] Users can use their devices to review and edit generated task lists and reminder content. This allows for flexible task management that adapts to the user's own work situation and emotional changes.

[0560] Step 9:

[0561] Throughout the entire process, the server monitors task progress and re-evaluates and adjusts tasks based on newly detected emotional states and environmental changes. This results in optimized task management and increased productivity.

[0562] (Example 2)

[0563] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0564] In today's work environment, efficiently managing busy schedules and multiple tasks is essential. However, traditional task management systems prioritize tasks without considering the user's emotional state, which can cause stress. Furthermore, the inability to manage tasks flexibly based on emotions is a challenge that leads to decreased productivity.

[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0566] In this invention, the server includes an input means for acquiring conversation data, a conversion means for converting the acquired conversation data into text using speech recognition technology, an information analysis means for analyzing the converted text data and extracting important information, an emotion recognition means for determining the user's emotional state based on the analyzed information, a task priority setting means for generating a task list based on the extracted information and emotional state and dynamically setting the priority of each task, and a reminder management means for adjusting the content and frequency of reminders based on the progress of tasks and the user's emotional state. This enables flexible task management that takes into account the user's emotional state, thereby improving work efficiency and user satisfaction.

[0567] "Input means" refers to devices and methods used to acquire conversational data, such as collecting voice and text from users.

[0568] "Conversion means" refers to the technology or process for converting acquired audio data into text data, and utilizes speech recognition technology.

[0569] "Information analysis means" refers to technologies and methods for analyzing converted text data and extracting important information.

[0570] "Emotion recognition means" refers to technologies and devices used to determine a user's emotional state, detecting emotions based on text or audio data.

[0571] "Task priority setting means" refers to methods and techniques for generating a task list based on extracted information and emotional states, and for dynamically setting its priority.

[0572] A "reminder management method" refers to a method or technique for adjusting the content and frequency of reminders based on task progress and the user's emotional state.

[0573] A "generative AI model" refers to a model generated using artificial intelligence technology that outputs results for specific input data.

[0574] A "prompt statement" is a statement used to give instructions to a generative AI model, controlling the model's behavior and output results.

[0575] This invention is a task management system that uses conversational data and can dynamically adjust task priorities while considering the user's emotional state. First, the system acquires information provided by the user through conversation. The terminal uses input devices such as a microphone and keyboard to collect voice and text data in real time.

[0576] The collected conversation data is converted into text format by the device using speech recognition technology. This process utilizes common speech recognition software, such as open-source speech recognition APIs.

[0577] Subsequently, the server receives the converted text data and analyzes it using natural language processing (NLTK) technology. Specifically, it utilizes natural language processing tools such as spaCy and NLTK. Through this analysis, important points are extracted from the conversation, clarifying the actions and decisions that the user should take.

[0578] Next, the server uses emotion recognition technology to determine the user's emotional state. This determination utilizes an emotion engine that evaluates the user's emotions based on the context of their voice and text. The emotion engine detects states such as positive or stressful, and sends this data to the server.

[0579] Based on this information, the server dynamically sets task priorities using a generated AI model. If the user's emotions are stressful, the server will prioritize urgent tasks; if they are in a positive state, it can suggest challenging tasks.

[0580] Furthermore, a reminder function is built in, and the server adjusts the content and frequency of reminders according to the user's emotional state. This makes it possible to process tasks at a manageable pace.

[0581] For example, if it is determined that a user is experiencing stress, the system will automatically re-evaluate tasks by prompting the AI ​​model with the message, "Please reset the priority of the current task according to the stress level." This is expected to improve the user's work efficiency and optimize their work-life balance.

[0582] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0583] Step 1:

[0584] The user engages in conversation, which the device captures in real time as audio and text data. Input is the user's voice and text, and output is digital audio files and text data. The device captures this data using a microphone and keyboard.

[0585] Step 2:

[0586] The device converts the acquired audio data into text using a speech recognition engine. The input is audio data, and the output is the converted text data. The speech recognition software analyzes phonemes and converts them into strings of characters.

[0587] Step 3:

[0588] The server receives text data sent from the terminal. The input is the converted text data, and the output is the text dataset to be analyzed. The server buffers the received text in order to process this data further.

[0589] Step 4:

[0590] The server uses natural language processing techniques to analyze text data and extract important information. The input is a text dataset, and the output is a list of important information. The server performs operations such as tokenization, part-of-speech tagging, and keyword extraction to identify the necessary information.

[0591] Step 5:

[0592] The server analyzes the user's emotional state using an emotion recognition engine. Input consists of text data and speech tone information, while output is user emotional state data. The server performs contextual analysis of the text and speech tone detection to quantify the emotion.

[0593] Step 6:

[0594] The server dynamically sets task priorities using a generative AI model. The input is a list of key information and sentiment data, and the output is a task list with assigned priorities. The server evaluates the extracted information and sentiment data, and automatically sets task priorities based on prompt messages.

[0595] Step 7:

[0596] The server adjusts the content and frequency of reminders based on the user's emotional state. Inputs are a prioritized task list and emotional state data, while output is a customized reminder schedule. The server then develops a reminder delivery plan tailored to the user's situation.

[0597] (Application Example 2)

[0598] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0599] In today's workplace, it is crucial to understand employees' emotional states in real time and implement appropriate task management and security measures accordingly. However, traditional systems struggle to dynamically respond to these emotional changes, often relying on manual adjustments to security protocols and task priorities. This can compromise efficiency and security, posing a significant challenge.

[0600] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0601] In this invention, the server includes data collection means for acquiring conversation data, information analysis means for analyzing the acquired conversation data and extracting important information, and sentiment analysis means for analyzing the user's emotional state and dynamically adjusting security protocols. This enables automatic adjustment of task priorities and dynamic management of security levels according to the emotional state of employees.

[0602] "Data collection means" refers to devices and technologies for acquiring conversation data in real time.

[0603] "Information analysis means" refers to techniques that analyze acquired conversation data and extract important information such as text data.

[0604] A "task generation means" is a device or program that executes the process of generating a task list based on extracted important information.

[0605] A "priority setting method" is a technique for setting priorities for each task in a task list according to its importance and urgency.

[0606] A "reminder management system" is a mechanism that manages the content and sending method of reminders based on the progress and deadline of tasks.

[0607] "Emotional analysis methods" refer to technologies that analyze a user's emotional state and reflect it in the system's behavior and the information it presents.

[0608] "Security management measures" refer to devices or programs that appropriately manage and adjust the security level within a system based on the results of sentiment analysis.

[0609] To realize this invention, the server, terminal, and user collaborate to configure the system. The server is primarily responsible for data analysis and emotion recognition, while the terminal collects and transforms conversational data. Users simply need to perform their daily tasks and communicate naturally.

[0610] The server uses a speech recognition engine (e.g., Google Speech-to-Text) to convert conversational data sent from the terminal into text data. Since the audio data is processed in real time, rapid information analysis is possible. Next, the server uses a natural language processing library (e.g., spaCy) to extract important information from the text data. Based on this information, the task generation mechanism operates and generates a task list for the user.

[0611] Furthermore, the server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotional state. Based on the emotional data, task priorities are dynamically adjusted, security management mechanisms are activated, and security protocols are adjusted as needed.

[0612] For example, if a user is feeling tense or stressed during a meeting, the server can automatically instruct security staff to raise the alert level based on emotion analysis. Conversely, when the mood is relaxed, the system can revert to normal security measures to reduce unnecessary burden. In this way, the goal is to achieve both safety and efficiency in the workplace environment.

[0613] By utilizing generative AI models, it is possible to automatically extract necessary information and emotions from natural user conversations, providing efficient work support. As a result, users can experience improved work efficiency and enhanced safety.

[0614] An example of a prompt that can be input to a generative AI model is, "Help me design a system that captures conversations in real time, analyzes emotional states, and dynamically adjusts office security levels."

[0615] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0616] Step 1:

[0617] The device acquires user conversation data in real time through its built-in microphone. During this process, voice data is input. The device records the voice data in digital format and transmits it to the server via a secure connection.

[0618] Step 2:

[0619] The server converts the received audio data into text data using a speech recognition engine. At this point, the input is audio data and the output is text data. Specifically, the Google Speech-to-Text engine is used to convert the audio into words and add timestamps.

[0620] Step 3:

[0621] The server analyzes the converted text data using spaCy, a natural language processing library. The input is text data, and the output is information extracted with important words and keywords. Specifically, it tags parts of speech such as nouns, verbs, and adjectives, and detects significant topics.

[0622] Step 4:

[0623] The server generates a task list based on the analysis results. The input is a list of important information, and the output is a specific task list. In operation, it automatically assigns tasks related to the user's schedule and projects based on the extracted information.

[0624] Step 5:

[0625] The server uses an emotion recognition engine to analyze the user's emotional state from text data. In this process, the input is text data, and the output is an evaluation of the emotional state. Specifically, it analyzes tone and emotion through IBM Watson Tone Analyzer and classifies the user's emotions as positive, negative, neutral, etc.

[0626] Step 6:

[0627] The server adjusts security levels based on emotional states. The input is an assessment of emotional state, and the output is the adjusted security protocol. Specifically, if tension levels are high, the server may increase the surveillance intensity of security cameras or restrict access privileges.

[0628] Step 7:

[0629] The server generates personalized prompts based on the generated task list and sentiment data. The input is the task list and sentiment data, and the output is the prompt. Specifically, it utilizes a generative AI model to create sentences suggesting what the user should do next. These prompts are then sent to the user's smart device.

[0630] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0631] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0632] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0633] [Fourth Embodiment]

[0634] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0635] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0636] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0637] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0638] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0639] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0640] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0641] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0642] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0643] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0644] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0645] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0646] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0647] This invention describes a system that acquires conversational data generated during meetings and communications in real time, analyzes it, and automatically generates and manages important tasks. The program processing and specific examples of this system are described below.

[0648] In this system, users participate in meetings and generate conversation data through audio and chat data. The terminal converts this data into text data in real time using speech recognition technology, and also processes related chat logs and email data.

[0649] The server receives text data and related data sent from the terminal. The server then uses natural language processing (NLP) techniques to analyze the data and extract important information from the conversation. Specifically, it employs methods to identify decisions and the next actions to be taken.

[0650] Based on the extracted information, the server automatically generates a task list, listing important action items. The listed tasks are then accompanied by necessary details, referencing the user's past task history and project information.

[0651] For each item in the generated tasks, the server uses machine learning models and rule-based algorithms to set priorities. This makes it easier for users to determine which tasks they should tackle quickly.

[0652] Furthermore, the server sends reminders to the device at appropriate times based on the task's progress and deadline. These reminders prompt the user to take the next action, preventing tasks from being missed or delayed.

[0653] Furthermore, this system can be tailored to specific industries. The server uses pre-configured industry-specific models to perform analyses that address the needs and challenges of those particular industries. These specialized models are trained on industry-specific datasets and contribute to the automation of industry-specific tasks.

[0654] For example, if the user is a project manager in the IT industry, the server analyzes the audio data of the project meeting and automatically generates tasks such as "Prepare for next week's sprint review," setting them to a high priority. Furthermore, it supports smooth project progress by sending reminders to the terminal at appropriate times, such as "The deadline for creating sprint review materials is approaching."

[0655] Thus, the present invention provides users with an efficient task management system, enabling them to streamline business processes and improve productivity.

[0656] The following describes the processing flow.

[0657] Step 1:

[0658] Users participate in meetings and engage in conversations and discussions. Communication takes place using voice and chat.

[0659] Step 2:

[0660] The device records audio during meetings in real time and converts it into text data using speech recognition technology. To generate highly accurate text from the recorded audio, it performs ambient noise removal and speaker separation.

[0661] Step 3:

[0662] The terminal sends the generated text data, chat logs used during the meeting, and email data as needed to the server. This transmitted data is then ready for processing by the system.

[0663] Step 4:

[0664] The server analyzes the received text data using a natural language processing (NLP) engine. Through understanding grammatical structure and vocabulary, it extracts important information from the context of the conversation, clarifying decisions and next actions.

[0665] Step 5:

[0666] The server automatically generates a task list based on the extracted information. This list specifically lists the required tasks, and each task includes relevant details. Furthermore, it refers to the user's past work history and supplements task details as needed.

[0667] Step 6:

[0668] The server sets the priority of each task based on its importance and deadline. Using machine learning algorithms, it evaluates the urgency and relative priority of tasks and indicates which tasks the user should tackle first.

[0669] Step 7:

[0670] The server monitors the progress of each task and sends reminders to the user's device based on the deadline. This ensures that users receive timely notifications and prevent tasks from being missed.

[0671] Step 8:

[0672] The servers are specialized for specific industries, employing industry-specific models accordingly. This allows them to address the diverse needs of each industry and achieve specialized task management.

[0673] Step 9:

[0674] Users can view generated task lists and reminders through the device interface and customize the interface to optimize task management to suit their work style. This makes it easier to optimize task placement and management for each industry and team.

[0675] (Example 1)

[0676] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] Traditional meeting management systems struggled to extract and properly manage important tasks from conversation content. Furthermore, their lack of flexibility to meet the unique needs of different industries hindered efficient business operations.

[0678] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0679] In this invention, the server includes data collection means for acquiring conversation information, information analysis means for analyzing the acquired conversation information and extracting important matters, and task generation means for generating a list of tasks based on the extracted important matters. This makes it possible to immediately identify and manage tasks from conversations in meetings.

[0680] "Conversational information" refers to information generated as audio or text during meetings and communication sessions.

[0681] "Data collection means" refers to a function or device for efficiently acquiring conversational information and incorporating it into a system.

[0682] "Information analysis means" refers to the techniques and methods used to analyze acquired conversational information and extract important information from it.

[0683] "Task generation means" refers to a function for automatically generating tasks based on important information extracted by information analysis means.

[0684] A "prioritization method" refers to a technique for determining which tasks should be prioritized among the generated tasks and setting their order.

[0685] "Notification management means" refers to a function that provides notifications at the appropriate time based on the progress and deadlines of a task, prompting the user to take the next action.

[0686] "Industry-specific methods" refer to technologies and methods for customizing systems to address the specific needs and challenges of a particular industry.

[0687] A "generative AI model" refers to a group of algorithms that use artificial intelligence technology to generate useful information or answers from input data.

[0688] "Prompt generation means" refers to a function or technology for generating natural language questions and commands from conversational information.

[0689] This invention is a system for efficiently processing conversational information generated in meetings and communications, and for achieving automated task generation and management. The system includes, as its components, data collection means, information analysis means, task generation means, priority setting means, notification management means, industry-specific means, generation AI model, and prompt generation means.

[0690] Users participate in meetings and generate conversational information in the form of voice or chat. The device collects this information using a microphone and corresponding software (e.g., a video conferencing application). The collected audio data is converted into text using a common speech recognition API (a comprehensive and widely used API as an example) as a voice identification technology.

[0691] The server analyzes the text information sent from the terminal using a natural language processing library (e.g., SpaCy or OpenNLP). It then extracts important information and uses a generative AI model to generate appropriate prompts. These prompts provide specific, actionable instructions for the generated task list. For example, a possible prompt might be, "List the preparation items for next week's team meeting."

[0692] Furthermore, the server prioritizes the generated tasks through machine learning models and rule-based algorithms. Based on the set priorities, it sends notifications via the terminal at appropriate times according to the progress and deadlines of the work. This prompts users to take the next action, thereby promoting work efficiency.

[0693] Furthermore, the server performs analysis using specialized models for specific industries. This industry-specific approach allows for analysis tailored to the needs of each industry, providing unique tasks and information. Combining these elements makes it possible to achieve efficient task management and optimization of business processes.

[0694] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0695] Step 1:

[0696] Users participate in meetings and generate conversational information in voice and chat formats. This results in voice and text data as input. The device collects this conversational information in real time through microphones and video conferencing software.

[0697] Step 2:

[0698] The device converts collected audio data into text using speech recognition technology. Specifically, it performs the process of converting audio data into text data via a speech recognition API. The input is audio data, and the output is text data. This text data records the content of the conversation as characters.

[0699] Step 3:

[0700] The terminal sends the converted text data and associated chat logs to the server. The server receives this data and prepares it for analysis. The input consists of text data and chat logs, which are used for the subsequent analysis process.

[0701] Step 4:

[0702] The server analyzes the received text data using natural language processing (NLP) techniques. Specifically, it uses an NLP library to extract important information and keywords from the text. The input is text data, and the output is important information as a result of the analysis.

[0703] Step 5:

[0704] Based on the analysis results, the server operates a prompt generation mechanism using a generation AI model. A specific list of tasks is then generated based on the resulting prompt messages. The input is important information, and the output consists of the generated prompt messages and the corresponding task items.

[0705] Step 6:

[0706] The server applies machine learning models and rule-based algorithms to prioritize work items and determine task priorities. The input is work items, and the output is a list of tasks with their priorities set.

[0707] Step 7:

[0708] The server uses a prioritized task list to send notifications to the terminal at appropriate times based on the progress and deadlines of each task. These notifications are intended to prompt the user to take the next action. The input is the task list and the progress of the tasks, and the output is the notification sent to the user.

[0709] (Application Example 1)

[0710] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0711] In brick-and-mortar retail operations, efficiently managing tasks arising from customer interactions and staff conversations is essential. However, traditional methods require manually recording conversations and manually organizing tasks, hindering operational efficiency. Furthermore, the insufficient use of specific information technology devices for visual task presentation leads to challenges in immediacy and accuracy.

[0712] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0713] In this invention, the server includes information gathering means for acquiring conversation data, data analysis means for analyzing the acquired conversation data and extracting important information, and display means for visually presenting task information to the user using a specific information technology device. This improves customer service and operational efficiency, enabling real-time task management and information presentation.

[0714] "Information gathering means" refers to a system for acquiring information such as conversations and audio data in real time.

[0715] "Data analysis tools" are technologies that analyze acquired information and extract important data and tasks.

[0716] A "task creation method" is a method for automatically generating a list of necessary tasks based on analyzed data.

[0717] A "prioritization mechanism" is a system that determines the priority of generated tasks according to their importance and urgency.

[0718] A "notification management system" is a technology that has the ability to send reminders to users based on the progress and deadlines of tasks.

[0719] "Industry-specific methods" refer to methods for adjusting analysis and task generation functions to suit specific industrial sectors.

[0720] A "display means" is a mechanism that uses a specific information technology device to visually present task information to the user.

[0721] This invention is a system that enables smooth task management by allowing staff to use smart glasses in physical stores to analyze customer interactions and conversations among staff in real time. The server uses a specific voice recognition device as a means of information gathering to collect voice data obtained from the field. Subsequently, the data is analyzed by the server using a data analysis means that employs natural language processing technology to extract important tasks and information. Specifically, using the Hugging Face Transformers library is effective.

[0722] The server also generates a task list based on information extracted by the task creation mechanism and determines the importance of each task using the priority setting mechanism. This information is immediately presented to staff via a specific information technology device, such as smart glasses. This allows users to efficiently manage their work within the physical store.

[0723] Furthermore, reminders are displayed on staff members' smart glasses via notification management tools, based on task progress and deadlines. Industry-specific tools can also be used to provide features optimized for specific store types or industries.

[0724] For example, if a store clerk says to a customer, "This item is sold out. Can you restock it?", a task such as "Restock sold-out item X" will be automatically generated. Furthermore, by utilizing the generative AI model and using a prompt such as, "Analyze the following conversation and extract the important tasks. Conversation: 'This item is sold out. Can you restock it?'", more accurate task analysis becomes possible.

[0725] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0726] Step 1:

[0727] The terminal acquires audio data of staff conversations and customer interactions within the store. This audio data is captured through a built-in speech recognition device in the terminal. The input is conversations within the store, and the output is audio data.

[0728] Step 2:

[0729] The device converts the acquired audio data into text data. This conversion process utilizes speech recognition technology. Specifically, it uses speech recognition software such as the Google Speech API. The input is audio data, and the output is the converted text data.

[0730] Step 3:

[0731] The server receives text data sent from the terminal. During this receiving process, the server uses natural language processing techniques to identify and extract important information. The input is text data, and the output is the analyzed important information. The accuracy of the analysis can be improved by using a generative AI model.

[0732] Step 4:

[0733] The server generates a task list and sets priorities based on the extracted information. This uses machine learning algorithms and also considers the user's past task history. The input is the analyzed key information, and the output is a task list with priorities set.

[0734] Step 5:

[0735] The server sends the generated task list to the terminal, where it is displayed to the staff via smart glasses. This allows users to visually check tasks in real time and respond immediately. The input is a task list with priorities set, and the output is the task list displayed on the smart glasses.

[0736] Step 6:

[0737] The server notifies the user's device of the task progress and deadline at the appropriate time. This prevents the user from missing or delaying tasks. Input is the task list and task management information, and output is the reminder displayed on the device.

[0738] Through the above processing steps, it is possible to improve the efficiency of task management in physical stores.

[0739] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0740] This invention combines an emotion engine with a system that acquires conversation data in real time, analyzes the content of that conversation, and generates tasks, thereby enabling task management that responds to the user's emotions. The program processing and specific examples of this system are described below.

[0741] In this invention, users communicate via voice and text during meetings and daily work. The terminal acquires this conversation data in real time and converts it into text data using speech recognition technology.

[0742] The server receives the converted text data and uses natural language processing (NLP) techniques to analyze the conversation. This analysis extracts important information from the conversation and identifies the next necessary actions and decisions.

[0743] Furthermore, the server uses an emotion engine to recognize the user's emotional state from their voice tone and text content. For example, if the user is feeling stressed or exhibiting positive emotions, the emotion engine captures that state as data.

[0744] Based on this emotional data, the server can dynamically adjust task priorities. For example, if the server determines that a user is in a high-stress state, it will prioritize urgent tasks to be handled quickly and adjust task distribution as needed. Conversely, if positive emotions are detected, it can prioritize suggesting challenging and difficult tasks.

[0745] Furthermore, as a reminder management method, adjusting the content and frequency of reminders based on emotions enables task management that takes into account the user's mental state. This allows users to complete tasks within a manageable scope, contributing to maintaining and improving productivity.

[0746] For example, when a user is under high stress as a project deadline approaches, the server will frequently send reminders for important tasks to help the user process them smoothly. Conversely, when a project is progressing successfully and the user is feeling positive, the server will suggest reminders encouraging them to tackle more creative tasks or new projects.

[0747] Thus, the present invention is a system that provides advanced task management that takes into account the emotional state of individual users, thereby not only improving work efficiency but also enabling the optimization of the user's work-life balance.

[0748] The following describes the processing flow.

[0749] Step 1:

[0750] Users participate in meetings and communicate via voice or text. The meetings involve multiple participants, each offering diverse opinions.

[0751] Step 2:

[0752] The device records audio data of conversations in real time and converts that audio data into text using speech recognition technology. If necessary, it performs filtering to remove noise and attempts to identify the speaker.

[0753] Step 3:

[0754] The terminal sends the converted text data to the server. The text data includes decisions and action points discussed at the meeting.

[0755] Step 4:

[0756] The server analyzes the received text data using natural language processing (NLP). This analysis includes contextual understanding and keyword extraction, which extracts important information (e.g., next steps, responsible parties, deadlines).

[0757] Step 5:

[0758] The server utilizes an emotion engine to analyze voice tone and linguistic features in text to identify the user's emotional state. It then quantifies and evaluates the degree of each emotion, such as joy, anger, and stress.

[0759] Step 6:

[0760] Based on the analysis results, the server automatically generates a task list that reflects important information. The generated tasks are then assigned priorities that take into account the results of the emotion engine. Tasks that are highly important under high stress levels are adjusted, such as by lowering their priority.

[0761] Step 7:

[0762] The server manages task reminders, adjusting the content and timing of reminders based on the user's emotional state. If the user is stressed, the frequency of reminders is adjusted accordingly; if they are in a positive mood, the reminders encourage challenging tasks.

[0763] Step 8:

[0764] Users can use their devices to review and edit generated task lists and reminder content. This allows for flexible task management that adapts to the user's own work situation and emotional changes.

[0765] Step 9:

[0766] Throughout the entire process, the server monitors task progress and re-evaluates and adjusts tasks based on newly detected emotional states and environmental changes. This results in optimized task management and increased productivity.

[0767] (Example 2)

[0768] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0769] In today's work environment, efficiently managing busy schedules and multiple tasks is essential. However, traditional task management systems prioritize tasks without considering the user's emotional state, which can cause stress. Furthermore, the inability to manage tasks flexibly based on emotions is a challenge that leads to decreased productivity.

[0770] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0771] In this invention, the server includes an input means for acquiring conversation data, a conversion means for converting the acquired conversation data into text using speech recognition technology, an information analysis means for analyzing the converted text data and extracting important information, an emotion recognition means for determining the user's emotional state based on the analyzed information, a task priority setting means for generating a task list based on the extracted information and emotional state and dynamically setting the priority of each task, and a reminder management means for adjusting the content and frequency of reminders based on the progress of tasks and the user's emotional state. This enables flexible task management that takes into account the user's emotional state, thereby improving work efficiency and user satisfaction.

[0772] "Input means" refers to devices and methods used to acquire conversational data, such as collecting voice and text from users.

[0773] "Conversion means" refers to the technology or process for converting acquired audio data into text data, and utilizes speech recognition technology.

[0774] "Information analysis means" refers to technologies and methods for analyzing converted text data and extracting important information.

[0775] "Emotion recognition means" refers to technologies and devices used to determine a user's emotional state, detecting emotions based on text or audio data.

[0776] "Task priority setting means" refers to methods and techniques for generating a task list based on extracted information and emotional states, and for dynamically setting its priority.

[0777] A "reminder management method" refers to a method or technique for adjusting the content and frequency of reminders based on task progress and the user's emotional state.

[0778] A "generative AI model" refers to a model generated using artificial intelligence technology that outputs results for specific input data.

[0779] A "prompt statement" is a statement used to give instructions to a generative AI model, controlling the model's behavior and output results.

[0780] This invention is a task management system that uses conversational data and can dynamically adjust task priorities while considering the user's emotional state. First, the system acquires information provided by the user through conversation. The terminal uses input devices such as a microphone and keyboard to collect voice and text data in real time.

[0781] The collected conversation data is converted into text format by the device using speech recognition technology. This process utilizes common speech recognition software, such as open-source speech recognition APIs.

[0782] Subsequently, the server receives the converted text data and analyzes it using natural language processing (NLTK) technology. Specifically, it utilizes natural language processing tools such as spaCy and NLTK. Through this analysis, important points are extracted from the conversation, clarifying the actions and decisions that the user should take.

[0783] Next, the server uses emotion recognition technology to determine the user's emotional state. This determination utilizes an emotion engine that evaluates the user's emotions based on the context of their voice and text. The emotion engine detects states such as positive or stressful, and sends this data to the server.

[0784] Based on this information, the server dynamically sets task priorities using a generated AI model. If the user's emotions are stressful, the server will prioritize urgent tasks; if they are in a positive state, it can suggest challenging tasks.

[0785] Furthermore, a reminder function is built in, and the server adjusts the content and frequency of reminders according to the user's emotional state. This makes it possible to process tasks at a manageable pace.

[0786] For example, if it is determined that a user is experiencing stress, the system will automatically re-evaluate tasks by prompting the AI ​​model with the message, "Please reset the priority of the current task according to the stress level." This is expected to improve the user's work efficiency and optimize their work-life balance.

[0787] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0788] Step 1:

[0789] The user engages in conversation, which the device captures in real time as audio and text data. Input is the user's voice and text, and output is digital audio files and text data. The device captures this data using a microphone and keyboard.

[0790] Step 2:

[0791] The device converts the acquired audio data into text using a speech recognition engine. The input is audio data, and the output is the converted text data. The speech recognition software analyzes phonemes and converts them into strings of characters.

[0792] Step 3:

[0793] The server receives text data sent from the terminal. The input is the converted text data, and the output is the text dataset to be analyzed. The server buffers the received text in order to process this data further.

[0794] Step 4:

[0795] The server uses natural language processing techniques to analyze text data and extract important information. The input is a text dataset, and the output is a list of important information. The server performs operations such as tokenization, part-of-speech tagging, and keyword extraction to identify the necessary information.

[0796] Step 5:

[0797] The server analyzes the user's emotional state using an emotion recognition engine. Input consists of text data and speech tone information, while output is user emotional state data. The server performs contextual analysis of the text and speech tone detection to quantify the emotion.

[0798] Step 6:

[0799] The server dynamically sets task priorities using a generative AI model. The input is a list of key information and sentiment data, and the output is a task list with assigned priorities. The server evaluates the extracted information and sentiment data, and automatically sets task priorities based on prompt messages.

[0800] Step 7:

[0801] The server adjusts the content and frequency of reminders based on the user's emotional state. Inputs are a prioritized task list and emotional state data, while output is a customized reminder schedule. The server then develops a reminder delivery plan tailored to the user's situation.

[0802] (Application Example 2)

[0803] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0804] In today's workplace, it is crucial to understand employees' emotional states in real time and implement appropriate task management and security measures accordingly. However, traditional systems struggle to dynamically respond to these emotional changes, often relying on manual adjustments to security protocols and task priorities. This can compromise efficiency and security, posing a significant challenge.

[0805] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0806] In this invention, the server includes data collection means for acquiring conversation data, information analysis means for analyzing the acquired conversation data and extracting important information, and sentiment analysis means for analyzing the user's emotional state and dynamically adjusting security protocols. This enables automatic adjustment of task priorities and dynamic management of security levels according to the emotional state of employees.

[0807] "Data collection means" refers to devices and technologies for acquiring conversation data in real time.

[0808] "Information analysis means" refers to techniques that analyze acquired conversation data and extract important information such as text data.

[0809] A "task generation means" is a device or program that executes the process of generating a task list based on extracted important information.

[0810] A "priority setting method" is a technique for setting priorities for each task in a task list according to its importance and urgency.

[0811] A "reminder management system" is a mechanism that manages the content and sending method of reminders based on the progress and deadline of tasks.

[0812] "Emotional analysis methods" refer to technologies that analyze a user's emotional state and reflect it in the system's behavior and the information it presents.

[0813] "Security management measures" refer to devices or programs that appropriately manage and adjust the security level within a system based on the results of sentiment analysis.

[0814] To realize this invention, the server, terminal, and user collaborate to configure the system. The server is primarily responsible for data analysis and emotion recognition, while the terminal collects and transforms conversational data. Users simply need to perform their daily tasks and communicate naturally.

[0815] The server uses a speech recognition engine (e.g., Google Speech-to-Text) to convert conversational data sent from the terminal into text data. Since the audio data is processed in real time, rapid information analysis is possible. Next, the server uses a natural language processing library (e.g., spaCy) to extract important information from the text data. Based on this information, the task generation mechanism operates and generates a task list for the user.

[0816] Furthermore, the server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotional state. Based on the emotional data, task priorities are dynamically adjusted, security management mechanisms are activated, and security protocols are adjusted as needed.

[0817] For example, if a user is feeling tense or stressed during a meeting, the server can automatically instruct security staff to raise the alert level based on emotion analysis. Conversely, when the mood is relaxed, the system can revert to normal security measures to reduce unnecessary burden. In this way, the goal is to achieve both safety and efficiency in the workplace environment.

[0818] By utilizing generative AI models, it is possible to automatically extract necessary information and emotions from natural user conversations, providing efficient work support. As a result, users can experience improved work efficiency and enhanced safety.

[0819] An example of a prompt that can be input to a generative AI model is, "Help me design a system that captures conversations in real time, analyzes emotional states, and dynamically adjusts office security levels."

[0820] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0821] Step 1:

[0822] The device acquires user conversation data in real time through its built-in microphone. During this process, voice data is input. The device records the voice data in digital format and transmits it to the server via a secure connection.

[0823] Step 2:

[0824] The server converts the received audio data into text data using a speech recognition engine. At this point, the input is audio data and the output is text data. Specifically, the Google Speech-to-Text engine is used to convert the audio into words and add timestamps.

[0825] Step 3:

[0826] The server analyzes the converted text data using spaCy, a natural language processing library. The input is text data, and the output is information extracted with important words and keywords. Specifically, it tags parts of speech such as nouns, verbs, and adjectives, and detects significant topics.

[0827] Step 4:

[0828] The server generates a task list based on the analysis results. The input is a list of important information, and the output is a specific task list. In operation, it automatically assigns tasks related to the user's schedule and projects based on the extracted information.

[0829] Step 5:

[0830] The server uses an emotion recognition engine to analyze the user's emotional state from text data. In this process, the input is text data, and the output is an evaluation of the emotional state. Specifically, it analyzes tone and emotion through IBM Watson Tone Analyzer and classifies the user's emotions as positive, negative, neutral, etc.

[0831] Step 6:

[0832] The server adjusts security levels based on emotional states. The input is an assessment of emotional state, and the output is the adjusted security protocol. Specifically, if tension levels are high, the server may increase the surveillance intensity of security cameras or restrict access privileges.

[0833] Step 7:

[0834] The server generates personalized prompts based on the generated task list and sentiment data. The input is the task list and sentiment data, and the output is the prompt. Specifically, it utilizes a generative AI model to create sentences suggesting what the user should do next. These prompts are then sent to the user's smart device.

[0835] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0836] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0837] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0838] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0839] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0840] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0841] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0842] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0843] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0844] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0845] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0846] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0847] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0848] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0849] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0850] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0851] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0852] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0853] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0854] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0855] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0856] The following is further disclosed regarding the embodiments described above.

[0857] (Claim 1)

[0858] A data collection method for obtaining conversation data,

[0859] An information analysis means that analyzes acquired conversation data and extracts important information,

[0860] A task generation method that generates a task list based on extracted important information,

[0861] A priority setting mechanism for setting the priority of each task in the task list,

[0862] A reminder management system that manages reminders based on task progress and deadlines,

[0863] Industry-specific methods that address specific industries,

[0864] A system that includes this.

[0865] (Claim 2)

[0866] The system according to claim 1, which converts acquired conversation data into text using speech recognition technology.

[0867] (Claim 3)

[0868] The system according to claim 1, wherein the task generation means adds detailed information to the task list by referring to the user's past task history.

[0869] "Example 1"

[0870] (Claim 1)

[0871] A data collection method for obtaining conversation information,

[0872] An information analysis means that analyzes acquired conversation information and extracts important information,

[0873] A task generation means that generates a list of tasks based on the extracted important items,

[0874] A priority setting mechanism for setting the priority of each task in the task list,

[0875] A notification management system that manages notifications based on the progress and deadlines of the work,

[0876] Industry-specific measures that address specific industries,

[0877] A prompt generation means that generates prompt sentences from conversational information using a generative AI model,

[0878] A system that includes this.

[0879] (Claim 2)

[0880] The system according to claim 1, which converts acquired conversation information into text information using speech recognition technology.

[0881] (Claim 3)

[0882] The system according to claim 1, wherein the task generation means adds detailed information to the task list by referring to the user's past task history.

[0883] "Application Example 1"

[0884] (Claim 1)

[0885] Information gathering methods for obtaining conversation data,

[0886] A data analysis method that analyzes acquired conversation data to extract important information,

[0887] A task creation method that creates a task list based on extracted important information,

[0888] A means for setting the priority of each task in a task list,

[0889] A notification management system that manages reminders based on task progress and deadlines,

[0890] Industry-specific measures that address specific industries,

[0891] A display means that visually presents task information to the user using a specific information technology device,

[0892] A system that includes this.

[0893] (Claim 2)

[0894] The system according to claim 1, which converts acquired conversation data into text using speech recognition technology.

[0895] (Claim 3)

[0896] The system according to claim 1, wherein the task creation means adds detailed information to the task list by referring to the user's past task history.

[0897] "Example 2 of combining an emotion engine"

[0898] (Claim 1)

[0899] An input method for acquiring conversation data,

[0900] A conversion means that converts acquired conversation data into text using speech recognition technology,

[0901] An information analysis means for analyzing converted text data and extracting important information,

[0902] An emotion recognition means that determines the user's emotional state based on the analyzed information,

[0903] A task priority setting means generates a task list based on extracted information and emotional states, and dynamically sets the priority of each task.

[0904] A reminder management system that adjusts the content and frequency of reminders based on task progress and the user's emotional state,

[0905] A system that includes this.

[0906] (Claim 2)

[0907] The system according to claim 1, which adjusts the priority of generated tasks using a generation AI model.

[0908] (Claim 3)

[0909] The system according to claim 1, wherein the emotion recognition means evaluates the user's emotions using a prompt statement.

[0910] "Application example 2 when combining with an emotional engine"

[0911] (Claim 1)

[0912] A data collection method for obtaining conversation data,

[0913] An information analysis means that analyzes acquired conversation data and extracts important information,

[0914] A task generation method that generates a task list based on extracted important information,

[0915] A priority setting mechanism for setting the priority of each task in the task list,

[0916] A reminder management system that manages reminders based on task progress and deadlines,

[0917] A sentiment analysis method that analyzes the user's emotional state and dynamically adjusts security protocols,

[0918] Security management means for managing the security level of the system,

[0919] A system that includes this.

[0920] (Claim 2)

[0921] The system according to claim 1, which converts acquired conversation data into text using speech recognition technology.

[0922] (Claim 3)

[0923] The system according to claim 1, wherein the task generation means adds detailed information to the task list by referring to the user's past task history. [Explanation of Symbols]

[0924] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A data collection method for obtaining conversation data, An information analysis means that analyzes acquired conversation data and extracts important information, A task generation method that generates a task list based on extracted important information, A priority setting mechanism for setting the priority of each task in the task list, A reminder management system that manages reminders based on task progress and deadlines, Industry-specific methods that address specific industries, A system that includes this.

2. The system according to claim 1, which converts acquired conversation data into text using speech recognition technology.

3. The system according to claim 1, wherein the task generation means adds detailed information to the task list by referring to the user's past task history.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A