system

A voice-activated system with a generative model automates task management and scheduling, addressing inefficiencies and errors in manual business operations by optimizing tasks and adapting to user emotions.

JP2026073337APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In modern business environments, managing multiple operations and tasks simultaneously is complex and time-consuming, particularly for small and medium-sized enterprises, leading to inefficiencies, human errors, and information leakage due to manual task management and scheduling.

Method used

A system utilizing a generative model on a server to convert voice data into text, automate meeting scheduling, and prioritize tasks, incorporating a terminal for voice input and notification, thereby optimizing task management and reducing user burden.

Benefits of technology

The system enhances work efficiency and accuracy by automating task management and scheduling, minimizing human error and information leakage, and providing flexible, emotion-aware responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073337000001_ABST
    Figure 2026073337000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means equipped with a terminal for receiving voice input, A means of processing audio data on a server to convert it to text, A method for optimizing a task using a generative model, A means of notifying the device of optimized tasks, A means of receiving and processing additional voice commands from the user. A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a modern business environment, many operations and tasks occur simultaneously, and it is complex and time-consuming to manage them appropriately. Especially for managers and team leaders of small and medium-sized enterprises, scheduling daily meetings and prioritizing tasks may hinder efficient business operations. Furthermore, since these tasks are performed manually, there is a high risk of human error and information leakage. There is a need to solve such problems.

Means for Solving the Problems

[0005] This invention provides a system that optimizes tasks using a generative model, via a server that converts voice data into text, and includes a terminal for receiving voice input. This system notifies the terminal of the optimized tasks and also accepts and processes additional voice instructions from the user, thereby improving work efficiency and accurate task management. Furthermore, it includes means for automating meeting scheduling and task prioritization, aiming to reduce the burden on users and streamline business processes.

[0006] "Voice input" is a method of transmitting what a user says as sound to a digital device to give instructions for operation or data processing.

[0007] A "terminal" is a device used by users to input and receive information, and in this system, it refers to a device used for voice capture and notification of operation results.

[0008] "Audio data" refers to data that digitally represents the waveform of sound acquired through audio input, enabling the transmission and analysis of information.

[0009] A "server" is a computer system that interacts with multiple terminals via a network to process and store data and provide services.

[0010] "Converting to text" is the process of converting data such as audio and images into a string of natural language characters, and is usually done using speech recognition technology.

[0011] A "generative model" is an algorithm that uses machine learning and artificial intelligence technologies to generate rational predictions and results from input data.

[0012] "Task optimization" is the process of efficiently managing various tasks in business operations and work, and making adjustments and improvements to enhance overall productivity.

[0013] "Notification" refers to a system communicating important information or results to a user through methods such as audio or visuals.

[0014] "Voice commands" refer to a means by which a user communicates specific operations or instructions to the system by voice, and these commands serve as the basis for the system's operation.

[0015] "Schedule setting" is the act of registering events or tasks for a specified date and time, and managing the plan for their execution.

[0016] Prioritization is the process of reviewing the importance of ongoing or planned tasks and replanning them to maximize the effectiveness of resource allocation. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0019] First, the language used in the following description will be explained.

[0020] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] The system of this invention primarily consists of a terminal that accepts voice input and a server that processes the input voice. The user inputs work-related instructions by voice through a microphone built into the terminal. This voice input includes scheduling meetings and changing task priorities.

[0039] The terminal receives voice input and sends it to the server as digital voice data. The server converts the received voice data into text using speech recognition technology. By analyzing this textual data, it is possible to clearly understand the user's instructions.

[0040] The server optimizes the assigned tasks using a generative model. In this process, it determines the optimal schedule and adjusts task priorities based on factors such as the company's calendar, existing appointments, and resource availability.

[0041] Once optimization is complete, the server notifies the terminal of optimized task and schedule information. The terminal receives the information from the server and notifies the user visually or audibly. This allows the user to stay informed of the latest schedule status and task priorities.

[0042] For example, if a user commands, "Schedule a project meeting for next Tuesday," the device sends voice data to the server. The server converts the voice data to text, determines the optimal schedule for the project meeting through a generative model, and returns the result to the device. The device then notifies the user, "We have scheduled a project meeting for next Tuesday at 2 PM." This process ensures that meeting schedules are set efficiently and automatically, eliminating the need for manual adjustments by the user.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] Users give voice instructions to the device's microphone to change meeting schedules or tasks. These voice instructions are recorded by the device as digital audio data.

[0046] Step 2:

[0047] The device uses a network connection to upload the recorded audio data to the server in order to send it to the server.

[0048] Step 3:

[0049] The server converts the received audio data into text using speech recognition technology. This allows the content of the voice instructions to be saved as digital text.

[0050] Step 4:

[0051] The server analyzes the transcribed instructions and identifies the task processing requirements based on their content. The results of this analysis determine the necessary business processes.

[0052] Step 5:

[0053] The server uses a generative model to optimize tasks based on the assigned work content. This optimization includes a process of scheduling based on specific conditions.

[0054] Step 6:

[0055] The server sends optimized schedule and task information to the device, including in the form of text and notification data.

[0056] Step 7:

[0057] The terminal analyzes the information received from the server and notifies the user of its contents visually or audibly. This allows the user to confirm the results.

[0058] Step 8:

[0059] The user reviews the received notification and provides additional voice instructions as needed. For example, they can perform instructions to change or cancel something.

[0060] Step 9:

[0061] The terminal sends additional user instructions to the server again, and the same process is restarted.

[0062] This series of steps allows the system to efficiently automate business tasks based on user voice commands, making management easier.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] There is a need to streamline the processing of work instructions based on voice input and to automatically optimize schedules and task management. Conventional systems require considerable effort and time to process voice data and optimize tasks, and also require user interaction, so new technologies are needed to solve this problem.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes means for processing voice information with computing resources to convert it into text information, means for optimizing the work using a generation algorithm, and means for notifying the device of the optimized work. This enables the automatic processing of work instructions via voice input, and allows for the rapid and efficient execution of optimized schedules and task management.

[0068] "Voice input" is a method by which digital devices receive instructions and information spoken by the user.

[0069] A "device" is a collection of hardware and software designed to perform a specific function.

[0070] "Voice information" refers to digital data acquired through voice input.

[0071] "Text information" refers to digital data converted into a string format, a format that can be processed more easily by computers.

[0072] "Computational resources" refer to the processing power and storage capacity provided by servers and computers.

[0073] A "generative algorithm" is a set of steps for calculating new outputs or results based on existing data.

[0074] "Work optimization" is the process of efficiently and effectively adjusting tasks and schedules set to achieve a specific goal.

[0075] "Notification" refers to sending a message to inform a user of information.

[0076] "Meeting planning" refers to the process of making the necessary preparations and setting schedules for holding a meeting or event.

[0077] "Priority" is an indicator that shows the order and importance required when performing tasks or processes.

[0078] The system of this invention efficiently processes work instructions based on voice input and automatically optimizes schedules and task management. Specific embodiments for carrying out the invention are shown below.

[0079] First, the user uses a device equipped with a voice input device. This device includes a microphone and is a digital device such as a smartphone or personal computer. The user verbally inputs instructions into this device, and the device captures the voice as digital audio data.

[0080] Next, the device transmits the acquired audio data to a server via the internet connection. The server, equipped with powerful computing resources, uses speech recognition technology to convert the audio data into text data. This process utilizes commercial speech recognition services or open-source speech processing libraries.

[0081] After the voice data is converted to text data, the server uses a generative AI model (for example, a generative model specializing in natural language processing) to analyze the user's instructions and extract specific task information. For example, if the user instructs, "Schedule a project meeting for next Tuesday," this command is understood as a request to schedule a meeting. The server then refers to the company's calendar information, resource availability, and other appointments to determine the optimal meeting time. The generative model uses prompt statements to break down the task into smaller parts and generate an optimized schedule. An example of a prompt statement might be, "If the user wants to schedule a project meeting, please suggest the factors to consider and the optimal scheduling time."

[0082] Ultimately, the server sends optimized schedule information and task priorities back to the terminal. Upon receiving this information, the terminal prepares to notify the user visually or audibly. This allows the user to receive information such as, "We have scheduled a project meeting for next Tuesday at 2 PM." This entire process enables users to efficiently manage their schedules and tasks without any hassle.

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] The user verbally inputs work instructions using a voice input device. The terminal captures this audio through the microphone and temporarily stores it as digital audio data. At this point, the input is the user's voice instructions, and the output is digital audio data.

[0086] Step 2:

[0087] The terminal sends audio data to the server. The audio data is transferred to the server via network communication. Specifically, this involves packetization of digital audio data and data transfer based on a protocol. The input is the audio data stored on the terminal, and the output is the audio data delivered to the server.

[0088] Step 3:

[0089] The server uses speech recognition technology to convert audio data into text data. Specifically, the speech recognition engine applies speech signal processing algorithms to analyze phonemes and convert them into text based on a language model. The input is audio data, and the output is text data.

[0090] Step 4:

[0091] The server uses a generative AI model to analyze text data and understand user instructions. At this stage, natural language processing is performed, and the instructions are extracted using a parsing algorithm. The generative AI model then uses the prompt text to perform additional inference and optimization. The input is text data, and the output is the parsed instructions.

[0092] Step 5:

[0093] The server optimizes tasks based on user instructions. Here, the optimal schedule is calculated by referencing the company's calendar and resource availability. Generative model prompts are used to refine the calculation results. The input is the parsed instructions, and the output is the optimized schedule.

[0094] Step 6:

[0095] The server sends optimized schedule information to the terminal. The calculated schedule is sent back to the terminal via network communication. The input is the optimized schedule, and the output is the schedule information transmitted to the terminal.

[0096] Step 7:

[0097] The device notifies the user of the schedule information it has received. Specifically, an application on the device conveys the information to the user visually or audibly. The notification app displays or plays an audio message, and the user receives a notification such as, "A project meeting has been scheduled for next Tuesday at 2 PM." The input is the schedule information received from the server, and the output is the notification result sent to the user.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] In factory manufacturing environments, automated voice control is needed to quickly and efficiently provide work instructions and adjust schedules. However, current technology does not adequately optimize the work process after receiving voice instructions, resulting in reduced work efficiency.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes means for processing voice information with a data processing device for converting it into text information, means for optimizing work using a generative model, and means for controlling work processes in the manufacturing industry with voice instructions. This enables smooth reception of work instructions by voice at the factory floor and the generation and execution of an optimal work schedule.

[0103] A "device for receiving voice input" is a device that detects voice commands from a user and acquires that voice information.

[0104] A "data processing device for converting audio information into text information" is a device that analyzes acquired audio information and performs the process of converting it into text information.

[0105] "Methods for optimizing work using generative models" refer to methods that utilize generative models to plan efficient work schedules based on work instruction data.

[0106] "Means for notifying the device of optimized work" refers to means for informing the corresponding device of optimized work schedules and instructions.

[0107] "Means for receiving and processing additional voice instructions from the user" refers to means that, when a user gives a new instruction, the system receives that instruction and processes the necessary actions.

[0108] "Means for controlling work processes in manufacturing using voice commands" refers to means for receiving voice commands for various tasks in a manufacturing process and controlling the content of those tasks accordingly.

[0109] To implement this invention, a device for receiving voice input and a server for processing it are required. The user gives work instructions by voice using a device such as a smartphone or tablet. This device is equipped with a high-performance microphone that can reliably capture the voice instructions spoken by the user.

[0110] The server functions as a data processing unit for converting speech information into text information. This process utilizes the Google® Cloud Speech-to-Text API as the speech recognition technology. This API enables high-precision speech recognition and is used to analyze the resulting text information.

[0111] By utilizing a generative model, the server optimizes the work process based on work instructions from the user. For example, it uses the Python library scipy.optimize to generate the optimal work schedule within the given resources and constraints. This optimized schedule information is then notified to the device.

[0112] If the user provides additional voice instructions, these instructions will also be received and processed in real time by the server, which will readjust the work schedule as needed. This will improve productivity on the factory floor and enable more efficient work progress.

[0113] For example, if a factory line operator gives a voice command such as, "Finish the next 100 products within 3 hours," this voice command is acquired by the device and converted into text data by a server. Then, an optimal production schedule is formulated using a generative AI model, and the robots begin operation at the appropriate time.

[0114] An example of a prompt message might be, "Generate the optimal work schedule based on this voice command and send the operation instructions to the designated robot." This would enable the automation and efficiency of tasks based on voice commands.

[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0116] Step 1:

[0117] The user inputs work instructions by voice through the terminal. The input voice is captured as digital audio data by the microphone built into the terminal. This data is input as an audio signal.

[0118] Step 2:

[0119] The device sends the acquired digital audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert this audio data into text. The output is the text data converted from the audio.

[0120] Step 3:

[0121] The server analyzes text data obtained using a generative model and optimizes the work schedule. Here, the Python library `scipy.optimize` is used to calculate the optimal schedule based on the input work instruction data. The optimized schedule information is then output.

[0122] Step 4:

[0123] The server notifies the terminal of optimized schedule information. The terminal presents this information to the user visually or audibly. As a result of this output, the user can check the latest schedule.

[0124] Step 5:

[0125] If the user provides additional voice instructions, the process from step 1 is repeated. This allows for continuous real-time adjustment of the work schedule. The output is the updated schedule information.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention combines a system comprising a terminal that accepts voice input, a server that converts voice data into text, and a generative model that optimizes tasks with an emotion engine that recognizes the user's emotions. The user can use the terminal to give voice instructions for scheduling meetings or changing task priorities. The terminal transmits these voice instructions to the server as digital voice data.

[0128] The server converts the received audio data into text and optimizes the task using a generative model. During this process, an emotion engine analyzes the user's emotions from the audio data and adjusts the optimization parameters accordingly. For example, if the user is stressed, the system can ease the schedule and postpone less important tasks.

[0129] Furthermore, based on the analysis results of the emotion engine, the device can change the way notifications are presented. For example, it can switch to concise and direct notifications when the user is busy or irritated, providing an interface that is considerate of the user's emotions.

[0130] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the emotion engine recognizes this emotion and adjusts the system's response accordingly. The server quickly processes the meeting cancellation, and the terminal sends a concise notification to the user stating, "Tomorrow's meeting has been canceled." This enables flexible responses that take emotions into consideration, reducing the user's workload.

[0131] In this way, this system, which incorporates an emotion engine, brings voice-input-based task management and automation closer to a more human-like response, providing a more user-friendly tool.

[0132] The following describes the processing flow.

[0133] Step 1:

[0134] The user gives work-related instructions by voice into the device's microphone. For example, they might say, "Cancel this Friday's meeting." This voice is captured by the device as digital audio data.

[0135] Step 2:

[0136] The device uses a network connection to upload the captured audio data to the server. Appropriate security protocols are used during this process to ensure data integrity.

[0137] Step 3:

[0138] The server receives the uploaded audio data and converts it into digital text using a speech recognition engine. Accuracy is crucial for this text conversion, as it forms the basis for accurately understanding the instructions.

[0139] Step 4:

[0140] The server analyzes the converted text and uses an emotion engine to identify emotions from the user's voice. This emotion information is used to identify a variety of emotional states, such as stress, anger, and joy.

[0141] Step 5:

[0142] The generative model calculates optimized task and schedule suggestions based on text and sentiment information. For example, if it determines that a user is stressed, it will suggest prioritizing the cancellation of low-priority meetings.

[0143] Step 6:

[0144] The server sends task and schedule information to the device, adjusted according to the user's emotions. At this point, the server adjusts the priority of information and the wording of notifications to match the user's emotions.

[0145] Step 7:

[0146] The device analyzes the received information and notifies the user in a way that is appropriate to their emotions. To encourage user focus, notifications may use concise and direct language.

[0147] Step 8:

[0148] The user reviews the notification, provides further voice instructions if necessary, and the process restarts. Emotion-responsive processing is a crucial element in improving the user experience.

[0149] (Example 2)

[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0151] Conventional voice input systems simply convert speech to text, making it difficult to dynamically optimize tasks and adjust interfaces while considering the user's emotions and state. Therefore, there was a need for flexible and human-like responses to reduce user stress and frustration.

[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0153] In this invention, the server includes means for processing voice information with communication equipment for converting it into text information, means for performing emotion analysis and dynamically adjusting task priorities based on the analysis results, and means for changing the interface representation according to the user's emotional state. This enables task optimization and notification adjustment that are adapted to the user's emotions and state.

[0154] A "device for receiving voice input" is a device that obtains voice instructions from a user and processes them as digital data.

[0155] "Communication equipment for converting audio information into text information" refers to a server or network-based device that converts received audio data into text data.

[0156] "Optimizing tasks using generative algorithms" is an AI technology that aims to improve the efficiency and priority of tasks based on the data received.

[0157] "Means for notifying the device of optimized work" refers to a function that sends optimization information processed by the server to the user's device.

[0158] "Means for receiving and processing additional voice instructions from users" refers to a system that allows users to give additional voice instructions after the initial input, and that continuously receives and processes these instructions.

[0159] "A means of performing emotion analysis and dynamically adjusting task priorities" refers to a system that analyzes emotions from the user's voice and adjusts task priorities in real time based on the results.

[0160] "Means of changing interface expression according to the user's emotional state" refers to a function that adapts the display method and notification content of the device to the situation based on the emotions the user is feeling.

[0161] This invention provides a mechanism for a task management system that uses voice input to achieve more flexible responses that are considerate of human emotions. The system includes a device for receiving voice input, a communication device for converting voice information into text information, and a configuration for optimizing tasks using a generation algorithm.

[0162] The user first provides voice instructions for the task to a voice input device. These voice instructions are captured using a terminal device equipped with a standard microphone. The terminal sends the captured voice to a server, which uses speech recognition technology to convert the voice data into text data. Google Cloud Speech-to-Text or similar speech recognition APIs can be used for this process.

[0163] The server uses a generative AI model to optimize tasks based on the converted text data. This process employs generative algorithms such as OpenAI's GPT series. Furthermore, the server utilizes an emotion engine to analyze the user's emotions from their voice, dynamically adjusting task and schedule priorities. As a result, if the user is experiencing stress, the system can automatically ease the schedule or postpone less important tasks.

[0164] Furthermore, based on sentiment analysis, the device modifies the content and wording of notifications to the user, providing an interface that is appropriate to the user's current emotional state. This feature allows for, for example, a concise and direct notification such as "I have canceled tomorrow's meeting" if the device determines that the user is in a bad mood.

[0165] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the system's emotion engine will recognize that emotion and quickly process the meeting cancellation. At this time, a notification will appear on the device stating, "Tomorrow's meeting has been canceled."

[0166] An example of a prompt for a generative AI model might be, "How would the system respond if a user, in an unhappy voice, requested that the next meeting be postponed?"

[0167] This configuration enables intuitive and effective task management and notifications that respond to the user's emotions and state.

[0168] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0169] Step 1:

[0170] The user provides task instructions verbally through a voice input device. Specifically, they might give instructions such as, "Cancel tomorrow's meeting." The input data is voice data. The terminal acquires this as an audio signal and stores it as digital data. This voice data forms the basis for the next processing step.

[0171] Step 2:

[0172] The terminal sends the acquired voice data to the server. At this time, the voice data is received by communication equipment via the network. The input received by the server is digital voice data, which is converted into text information using a speech recognition API (e.g., Google Cloud Speech-to-Text). The output generated is text information of the voice command.

[0173] Step 3:

[0174] The server receives the text information converted from the speech and uses a generative AI model (e.g., OpenAI's GPT series) to optimize the task. The input used here is the converted text information and previously analyzed emotion data. The emotion engine operates here, analyzing the user's emotions from the speech. Based on the emotion analysis results, the generative model dynamically adjusts the task priority and schedule to output optimized task information.

[0175] Step 4:

[0176] Upon receiving optimized task information from the server, the terminal adjusts the notification content to the user based on sentiment analysis. For example, if the user is in a bad mood, it generates a concise and direct notification such as, "I have canceled tomorrow's meeting." The inputs here are the optimized task information and the user's emotional state, and the output is the adjusted notification message.

[0177] Step 5:

[0178] The device presents the user with optimized notification messages. Notifications can be delivered visually or audibly. This allows users to receive appropriate notifications that take into account their emotions and task priorities, enabling effective task management. Output is delivered in the form of direct notifications to the user.

[0179] (Application Example 2)

[0180] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0181] In modern society, users may experience stress and dissatisfaction when using food delivery services, particularly regarding ordering and delivery. Insufficient voice instructions and emotionally responsive service are major factors that detract from the overall service experience. Furthermore, inadequate order management and delivery planning during delivery can reduce efficiency and negatively impact the user experience.

[0182] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0183] In this invention, the server includes means for processing audio data with an information processing device for converting it into text information, means for optimizing operations using a generative model, and means including an emotion analysis engine for analyzing the user's emotions. This enables the optimization of order management and delivery plans in accordance with the user's emotions, making it possible to provide a more personalized food delivery service.

[0184] "Voice input" refers to a means of input that allows users to convey instructions to devices and systems through their voice.

[0185] An "input / output device" is hardware that receives audio data and transmits the processing results to the user.

[0186] An "information processing device" is a computer system used to convert audio data into text information or to perform calculations using generative models.

[0187] "Textual information" refers to text data obtained by converting audio data into a format that can be processed by a computer.

[0188] A "generative model" is an artificial intelligence algorithm designed to optimize a specific task.

[0189] "Work optimization" refers to methods for maximizing the efficiency and effectiveness of work under specific conditions.

[0190] A "sentiment analysis engine" is software that analyzes and determines a user's emotions from their voice or text.

[0191] "Notification content" refers to messages or alerts sent to inform users of processing results or other information.

[0192] "Order management" refers to the entire process of receiving customer orders and processing and tracking them appropriately.

[0193] A "delivery plan" is a plan of schedules and routes to efficiently deliver ordered goods to the customer.

[0194] This invention realizes an order management system for food delivery services that takes into account voice commands and emotions. The system is configured as follows:

[0195] First, the user places an order by voice using an input / output device that accepts voice input (e.g., a smartphone or smart device). The input / output device sends this voice input as digital voice data to the server. The server uses an information processing device to convert the voice data into text. Specifically, it utilizes speech recognition software such as the Google Cloud Speech-to-Text API to convert the voice data into text information.

[0196] The server then uses a generative model to optimize order management and delivery planning. Possible generative models used here include OpenAI's GPT and deep learning models. This is used to calculate the shortest delivery route and the optimal delivery time.

[0197] Furthermore, an emotion analysis engine analyzes the user's emotions from text data. Emotion recognition tools such as IBM Watson® are used for this analysis. This analysis reflects the user's emotions in delivery plans and notification content. For example, if the system detects that the user is in a hurry, it will arrange for the shortest possible delivery time, resulting in faster delivery.

[0198] Notification content is adjusted based on analyzed emotions and sent to the user via input / output devices. Because notifications are concise and relevant to the user's needs, a stress-free and comfortable interface is provided.

[0199] For example, if a user requests "I want my pasta delivered immediately!", the server recognizes this "urgency" and generates a top-priority delivery plan tailored to that emotion. It then concisely informs the user of the estimated delivery time.

[0200] The following is an example of a possible prompt: "Write a program that uses an emotion engine to recognize the user's emotions and generate the optimal delivery route and time."

[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0202] Step 1:

[0203] The user provides voice input through an input / output device. The voice data is captured as input and converted into a format for transmission to the speech recognition system. This allows the terminal to collect voice information as digital data and prepare the voice data necessary for the next processing step.

[0204] Step 2:

[0205] The server uses speech recognition software to convert audio data into text. This process takes audio data as input and outputs text. The Google Cloud Speech-to-Text API is used to convert speech into text format based on acoustic features. This makes the content of the speech available as text.

[0206] Step 3:

[0207] The server analyzes the user's emotions from text data using an emotion analysis engine. Based on the text data as input, it performs data calculations to identify emotions and outputs emotion tags (e.g., anxiety, joy, frustration, etc.) as the analysis result. IBM Watson is a possible emotion analysis engine to be used. Specifically, it analyzes the frequency of occurrence of certain patterns and keywords in the text to infer the emotional state.

[0208] Step 4:

[0209] The server uses a generative model to optimize order management and delivery planning. It processes sentiment analysis results and text data as input to optimize delivery schedules and routes. In this process, it utilizes a generative AI model (e.g., OpenAI GPT) to apply processing based on sentiment. The optimized delivery plan is then output.

[0210] Step 5:

[0211] The server processes the optimization results into notification content and sends it to the terminal. Here, it considers the optimization results and sentiment data as input, transforming them into a notification message that is easy for the user to understand. The output is a concise yet detailed notification for the user, allowing them to grasp the situation in real time.

[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0215] [Second Embodiment]

[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0228] The system of this invention primarily consists of a terminal that accepts voice input and a server that processes the input voice. The user inputs work-related instructions by voice through a microphone built into the terminal. This voice input includes scheduling meetings and changing task priorities.

[0229] The terminal receives voice input and sends it to the server as digital voice data. The server converts the received voice data into text using speech recognition technology. By analyzing this textual data, it is possible to clearly understand the user's instructions.

[0230] The server optimizes the assigned tasks using a generative model. In this process, it determines the optimal schedule and adjusts task priorities based on factors such as the company's calendar, existing appointments, and resource availability.

[0231] Once optimization is complete, the server notifies the terminal of optimized task and schedule information. The terminal receives the information from the server and notifies the user visually or audibly. This allows the user to stay informed of the latest schedule status and task priorities.

[0232] For example, if a user commands, "Schedule a project meeting for next Tuesday," the device sends voice data to the server. The server converts the voice data to text, determines the optimal schedule for the project meeting through a generative model, and returns the result to the device. The device then notifies the user, "We have scheduled a project meeting for next Tuesday at 2 PM." This process ensures that meeting schedules are set efficiently and automatically, eliminating the need for manual adjustments by the user.

[0233] The following describes the processing flow.

[0234] Step 1:

[0235] Users give voice instructions to the device's microphone to change meeting schedules or tasks. These voice instructions are recorded by the device as digital audio data.

[0236] Step 2:

[0237] The device uses a network connection to upload the recorded audio data to the server in order to send it to the server.

[0238] Step 3:

[0239] The server converts the received audio data into text using speech recognition technology. This allows the content of the voice instructions to be saved as digital text.

[0240] Step 4:

[0241] The server analyzes the transcribed instructions and identifies the task processing requirements based on their content. The results of this analysis determine the necessary business processes.

[0242] Step 5:

[0243] The server uses a generative model to optimize tasks based on the assigned work content. This optimization includes a process of scheduling based on specific conditions.

[0244] Step 6:

[0245] The server sends optimized schedule and task information to the device, including in the form of text and notification data.

[0246] Step 7:

[0247] The terminal analyzes the information received from the server and notifies the user of its contents visually or audibly. This allows the user to confirm the results.

[0248] Step 8:

[0249] The user reviews the received notification and provides additional voice instructions as needed. For example, they can perform instructions to change or cancel something.

[0250] Step 9:

[0251] The terminal sends additional user instructions to the server again, and the same process is restarted.

[0252] This series of steps allows the system to efficiently automate business tasks based on user voice commands, making management easier.

[0253] (Example 1)

[0254] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0255] There is a need to streamline the processing of work instructions based on voice input and to automatically optimize schedules and task management. Conventional systems require considerable effort and time to process voice data and optimize tasks, and also require user interaction, so new technologies are needed to solve this problem.

[0256] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0257] In this invention, the server includes means for processing voice information with computing resources to convert it into text information, means for optimizing the work using a generation algorithm, and means for notifying the device of the optimized work. This enables the automatic processing of work instructions via voice input, and allows for the rapid and efficient execution of optimized schedules and task management.

[0258] "Voice input" is a method by which digital devices receive instructions and information spoken by the user.

[0259] A "device" is a collection of hardware and software designed to perform a specific function.

[0260] "Voice information" refers to digital data acquired through voice input.

[0261] "Text information" refers to digital data converted into a string format, a format that can be processed more easily by computers.

[0262] "Computational resources" refer to the processing power and storage capacity provided by servers and computers.

[0263] A "generative algorithm" is a set of steps for calculating new outputs or results based on existing data.

[0264] "Work optimization" is the process of efficiently and effectively adjusting tasks and schedules set to achieve a specific goal.

[0265] "Notification" refers to sending a message to inform a user of information.

[0266] "Meeting planning" refers to the process of making the necessary preparations and setting schedules for holding a meeting or event.

[0267] "Priority" is an indicator that shows the order and importance required when performing tasks or processes.

[0268] The system of this invention efficiently processes work instructions based on voice input and automatically optimizes schedules and task management. Specific embodiments for carrying out the invention are shown below.

[0269] First, the user uses a device equipped with a voice input device. This device includes a microphone and is a digital device such as a smartphone or personal computer. The user verbally inputs instructions into this device, and the device captures the voice as digital audio data.

[0270] Next, the device transmits the acquired audio data to a server via the internet connection. The server, equipped with powerful computing resources, uses speech recognition technology to convert the audio data into text data. This process utilizes commercial speech recognition services or open-source speech processing libraries.

[0271] After the voice data is converted to text data, the server uses a generative AI model (for example, a generative model specializing in natural language processing) to analyze the user's instructions and extract specific task information. For example, if the user instructs, "Schedule a project meeting for next Tuesday," this command is understood as a request to schedule a meeting. The server then refers to the company's calendar information, resource availability, and other appointments to determine the optimal meeting time. The generative model uses prompt statements to break down the task into smaller parts and generate an optimized schedule. An example of a prompt statement might be, "If the user wants to schedule a project meeting, please suggest the factors to consider and the optimal scheduling time."

[0272] Ultimately, the server sends optimized schedule information and task priorities back to the terminal. Upon receiving this information, the terminal prepares to notify the user visually or audibly. This allows the user to receive information such as, "We have scheduled a project meeting for next Tuesday at 2 PM." This entire process enables users to efficiently manage their schedules and tasks without any hassle.

[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0274] Step 1:

[0275] The user verbally inputs work instructions using a voice input device. The terminal captures this audio through the microphone and temporarily stores it as digital audio data. At this point, the input is the user's voice instructions, and the output is digital audio data.

[0276] Step 2:

[0277] The terminal sends voice data to the server. The voice data is transferred to the server via network communication. As specific operations here, packetization of digital voice data and data transfer based on protocols are performed. The input is the voice data stored in the terminal, and the output is the voice data delivered to the server.

[0278] Step 3:

[0279] The server uses voice recognition technology to convert the voice data into text data. Specifically, the voice recognition engine applies voice signal processing algorithms to analyze phonemes and convert them into text based on a language model. The input is the voice data, and the output is the text data.

[0280] Step 4:

[0281] The server uses a generative AI model to analyze the text data and understand the content of the user's instructions. At this stage, natural language processing is performed, and the instruction content is extracted by an analysis algorithm. The generative AI model uses a prompt sentence to perform additional inferences and optimizations. The input is the text data, and the output is the analyzed instruction content.

[0282] Step 5:

[0283] The server optimizes the task based on the user's instructions. Here, while referring to the company's calendar and the availability of resources, an optimal schedule is calculated. A prompt from the generative model is used to review the calculation results. The input is the analyzed instruction content, and the output is the optimized schedule.

[0284] Step 6:

[0285] The server sends the optimized schedule information to the terminal. The calculated schedule is sent back to the terminal via network communication. The input is the optimized schedule, and the output is the schedule information transmitted to the terminal.

[0286] Step 7:

[0287] The terminal notifies the user of the received schedule information. Specifically, the application on the terminal conveys the information to the user visually or aurally. The notification application performs display and audio playback, and the user receives a notification such as "A project meeting has been set for 2:00 PM next Tuesday." The input is the schedule information received from the server, and the output is the notification result to the user.

[0288] (Application Example 1)

[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server," and the smart glasses 214 are referred to as the "terminal."

[0290] In a factory's manufacturing site, in order to quickly and efficiently give work instructions and adjust schedules, automation by voice instructions is required. However, with current technology, there is a problem that the optimization of the work process after receiving a voice instruction is not sufficiently performed, resulting in a decrease in work efficiency.

[0291] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0292] In this invention, the server includes means for performing processing with a data processing device for converting voice information into character information, means for optimizing work using a generation model, and means for controlling work processes in manufacturing by voice instructions. Thereby, work instructions at the factory site can be smoothly received by voice, enabling the generation and execution of an optimal work schedule.

[0293] The "device for receiving voice input" is a device for detecting a voice instruction from a user and acquiring the voice information.

[0294] A "data processing device for converting audio information into text information" is a device that analyzes acquired audio information and performs the process of converting it into text information.

[0295] "Methods for optimizing work using generative models" refer to methods that utilize generative models to plan efficient work schedules based on work instruction data.

[0296] "Means for notifying the device of optimized work" refers to means for informing the corresponding device of optimized work schedules and instructions.

[0297] "Means for receiving and processing additional voice instructions from the user" refers to means that, when a user gives a new instruction, the system receives that instruction and processes the necessary actions.

[0298] "Means for controlling work processes in manufacturing using voice commands" refers to means for receiving voice commands for various tasks in a manufacturing process and controlling the content of those tasks accordingly.

[0299] To implement this invention, a device for receiving voice input and a server for processing it are required. The user gives work instructions by voice using a device such as a smartphone or tablet. This device is equipped with a high-performance microphone that can reliably capture the voice instructions spoken by the user.

[0300] The server functions as a data processing unit for converting speech information into text information. This process utilizes the Google Cloud Speech-to-Text API as the speech recognition technology. This API enables high-accuracy speech recognition and is used to analyze the resulting text information.

[0301] The server utilizes the generation model to optimize the work process based on the work instructions from the user. For example, using the scipy.optimize library in Python, an optimal work schedule is generated within the given resources and constraints. This optimized schedule information is notified to the device.

[0302] When the user gives additional voice instructions, the server newly accepts the instructions, processes them in real time, and readjusts the work schedule as necessary. This improves productivity at the factory site and enables efficient progress of work.

[0303] As a specific example, when a factory line operator gives a voice instruction such as "Finish the next 100 products within 3 hours", this voice is acquired by the device and converted into character data by the server. Then, an optimal production schedule is formulated using the generation AI model, and the robot starts operating at an appropriate timing.

[0304] As an example of the prompt text, content such as "Generate an optimal work schedule based on this voice instruction and send an operation instruction to the specified robot." can be considered. This enables automation and efficiency improvement of work based on voice instructions. <0,000,961> The flow of the specific process in Application Example 1 will be described using FIG. 12. [[ID=十七]]

[0306] Step 1:

[0307] The user inputs a work instruction by voice through the terminal. The input voice is acquired as digital voice data by the microphone installed on the terminal. This data is an input as an audio signal.

[0308] Step 2:

[0309] The device sends the acquired digital audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert this audio data into text. The output is the text data converted from the audio.

[0310] Step 3:

[0311] The server analyzes text data obtained using a generative model and optimizes the work schedule. Here, the Python library `scipy.optimize` is used to calculate the optimal schedule based on the input work instruction data. The optimized schedule information is then output.

[0312] Step 4:

[0313] The server notifies the terminal of optimized schedule information. The terminal presents this information to the user visually or audibly. As a result of this output, the user can check the latest schedule.

[0314] Step 5:

[0315] If the user provides additional voice instructions, the process from step 1 is repeated. This allows for continuous real-time adjustment of the work schedule. The output is the updated schedule information.

[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0317] This invention combines a system comprising a terminal that accepts voice input, a server that converts voice data into text, and a generative model that optimizes tasks with an emotion engine that recognizes the user's emotions. The user can use the terminal to give voice instructions for scheduling meetings or changing task priorities. The terminal transmits these voice instructions to the server as digital voice data.

[0318] The server converts the received audio data into text and optimizes the task using a generative model. During this process, an emotion engine analyzes the user's emotions from the audio data and adjusts the optimization parameters accordingly. For example, if the user is stressed, the system can ease the schedule and postpone less important tasks.

[0319] Furthermore, based on the analysis results of the emotion engine, the device can change the way notifications are presented. For example, it can switch to concise and direct notifications when the user is busy or irritated, providing an interface that is considerate of the user's emotions.

[0320] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the emotion engine recognizes this emotion and adjusts the system's response accordingly. The server quickly processes the meeting cancellation, and the terminal sends a concise notification to the user stating, "Tomorrow's meeting has been canceled." This enables flexible responses that take emotions into consideration, reducing the user's workload.

[0321] In this way, this system, which incorporates an emotion engine, brings voice-input-based task management and automation closer to a more human-like response, providing a more user-friendly tool.

[0322] The following describes the processing flow.

[0323] Step 1:

[0324] The user gives work-related instructions by voice into the device's microphone. For example, they might say, "Cancel this Friday's meeting." This voice is captured by the device as digital audio data.

[0325] Step 2:

[0326] The device uses a network connection to upload the captured audio data to the server. Appropriate security protocols are used during this process to ensure data integrity.

[0327] Step 3:

[0328] The server receives the uploaded audio data and converts it into digital text using a speech recognition engine. Accuracy is crucial for this text conversion, as it forms the basis for accurately understanding the instructions.

[0329] Step 4:

[0330] The server analyzes the converted text and uses an emotion engine to identify emotions from the user's voice. This emotion information is used to identify a variety of emotional states, such as stress, anger, and joy.

[0331] Step 5:

[0332] The generative model calculates optimized task and schedule suggestions based on text and sentiment information. For example, if it determines that a user is stressed, it will suggest prioritizing the cancellation of low-priority meetings.

[0333] Step 6:

[0334] The server sends task and schedule information to the device, adjusted according to the user's emotions. At this point, the server adjusts the priority of information and the wording of notifications to match the user's emotions.

[0335] Step 7:

[0336] The device analyzes the received information and notifies the user in a way that is appropriate to their emotions. To encourage user focus, notifications may use concise and direct language.

[0337] Step 8:

[0338] The user reviews the notification, provides further voice instructions if necessary, and the process restarts. Emotion-responsive processing is a crucial element in improving the user experience.

[0339] (Example 2)

[0340] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0341] Conventional voice input systems simply convert speech to text, making it difficult to dynamically optimize tasks and adjust interfaces while considering the user's emotions and state. Therefore, there was a need for flexible and human-like responses to reduce user stress and frustration.

[0342] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0343] In this invention, the server includes means for processing voice information with communication equipment for converting it into text information, means for performing emotion analysis and dynamically adjusting task priorities based on the analysis results, and means for changing the interface representation according to the user's emotional state. This enables task optimization and notification adjustment that are adapted to the user's emotions and state.

[0344] A "device for receiving voice input" is a device that obtains voice instructions from a user and processes them as digital data.

[0345] "Communication equipment for converting audio information into text information" refers to a server or network-based device that converts received audio data into text data.

[0346] "Optimizing tasks using generative algorithms" is an AI technology that aims to improve the efficiency and priority of tasks based on the data received.

[0347] "Means for notifying the device of optimized work" refers to a function that sends optimization information processed by the server to the user's device.

[0348] "Means for receiving and processing additional voice instructions from users" refers to a system that allows users to give additional voice instructions after the initial input, and that continuously receives and processes these instructions.

[0349] "A means of performing emotion analysis and dynamically adjusting task priorities" refers to a system that analyzes emotions from the user's voice and adjusts task priorities in real time based on the results.

[0350] "Means of changing interface expression according to the user's emotional state" refers to a function that adapts the display method and notification content of the device to the situation based on the emotions the user is feeling.

[0351] This invention provides a mechanism for a task management system that uses voice input to achieve more flexible responses that are considerate of human emotions. The system includes a device for receiving voice input, a communication device for converting voice information into text information, and a configuration for optimizing tasks using a generation algorithm.

[0352] The user first provides voice instructions for the task to a voice input device. These voice instructions are captured using a terminal device equipped with a standard microphone. The terminal sends the captured voice to a server, which uses speech recognition technology to convert the voice data into text data. Google Cloud Speech-to-Text or similar speech recognition APIs can be used for this process.

[0353] The server uses a generative AI model to optimize tasks based on the converted text data. This process employs generative algorithms such as OpenAI's GPT series. Furthermore, the server utilizes an emotion engine to analyze the user's emotions from their voice, dynamically adjusting task and schedule priorities. As a result, if the user is experiencing stress, the system can automatically ease the schedule or postpone less important tasks.

[0354] Furthermore, based on sentiment analysis, the device modifies the content and wording of notifications to the user, providing an interface that is appropriate to the user's current emotional state. This feature allows for, for example, a concise and direct notification such as "I have canceled tomorrow's meeting" if the device determines that the user is in a bad mood.

[0355] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the system's emotion engine will recognize that emotion and quickly process the meeting cancellation. At this time, a notification will appear on the device stating, "Tomorrow's meeting has been canceled."

[0356] An example of a prompt for a generative AI model might be, "How would the system respond if a user, in an unhappy voice, requested that the next meeting be postponed?"

[0357] This configuration enables intuitive and effective task management and notifications that respond to the user's emotions and state.

[0358] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0359] Step 1:

[0360] The user provides task instructions verbally through a voice input device. Specifically, they might give instructions such as, "Cancel tomorrow's meeting." The input data is voice data. The terminal acquires this as an audio signal and stores it as digital data. This voice data forms the basis for the next processing step.

[0361] Step 2:

[0362] The terminal sends the acquired voice data to the server. At this time, the voice data is received by communication equipment via the network. The input received by the server is digital voice data, which is converted into text information using a speech recognition API (e.g., Google Cloud Speech-to-Text). The output generated is text information of the voice command.

[0363] Step 3:

[0364] The server receives the text information converted from the speech and uses a generative AI model (e.g., OpenAI's GPT series) to optimize the task. The input used here is the converted text information and previously analyzed emotion data. The emotion engine operates here, analyzing the user's emotions from the speech. Based on the emotion analysis results, the generative model dynamically adjusts the task priority and schedule to output optimized task information.

[0365] Step 4:

[0366] Upon receiving optimized task information from the server, the terminal adjusts the notification content to the user based on sentiment analysis. For example, if the user is in a bad mood, it generates a concise and direct notification such as, "I have canceled tomorrow's meeting." The inputs here are the optimized task information and the user's emotional state, and the output is the adjusted notification message.

[0367] Step 5:

[0368] The device presents the user with optimized notification messages. Notifications can be delivered visually or audibly. This allows users to receive appropriate notifications that take into account their emotions and task priorities, enabling effective task management. Output is delivered in the form of direct notifications to the user.

[0369] (Application Example 2)

[0370] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0371] In modern society, users may experience stress and dissatisfaction when using food delivery services, particularly regarding ordering and delivery. Insufficient voice instructions and emotionally responsive service are major factors that detract from the overall service experience. Furthermore, inadequate order management and delivery planning during delivery can reduce efficiency and negatively impact the user experience.

[0372] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0373] In this invention, the server includes means for processing audio data with an information processing device for converting it into text information, means for optimizing operations using a generative model, and means including an emotion analysis engine for analyzing the user's emotions. This enables the optimization of order management and delivery plans in accordance with the user's emotions, making it possible to provide a more personalized food delivery service.

[0374] "Voice input" refers to a means of input that allows users to convey instructions to devices and systems through their voice.

[0375] An "input / output device" is hardware that receives audio data and transmits the processing results to the user.

[0376] An "information processing device" is a computer system used to convert audio data into text information or to perform calculations using generative models.

[0377] "Textual information" refers to text data obtained by converting audio data into a format that can be processed by a computer.

[0378] A "generative model" is an artificial intelligence algorithm designed to optimize a specific task.

[0379] "Work optimization" refers to methods for maximizing the efficiency and effectiveness of work under specific conditions.

[0380] A "sentiment analysis engine" is software that analyzes and determines a user's emotions from their voice or text.

[0381] "Notification content" refers to messages or alerts sent to inform users of processing results or other information.

[0382] "Order management" refers to the entire process of receiving customer orders and processing and tracking them appropriately.

[0383] A "delivery plan" is a plan of schedules and routes to efficiently deliver ordered goods to the customer.

[0384] This invention realizes an order management system for food delivery services that takes into account voice commands and emotions. The system is configured as follows:

[0385] First, the user places an order by voice using an input / output device that accepts voice input (e.g., a smartphone or smart device). The input / output device sends this voice input as digital voice data to the server. The server uses an information processing device to convert the voice data into text. Specifically, it utilizes speech recognition software such as the Google Cloud Speech-to-Text API to convert the voice data into text information.

[0386] The server then uses a generative model to optimize order management and delivery planning. Possible generative models used here include OpenAI's GPT and deep learning models. This is used to calculate the shortest delivery route and the optimal delivery time.

[0387] Furthermore, an emotion analysis engine analyzes the user's emotions from text data. Emotion recognition tools such as IBM Watson are used for this analysis. This analysis reflects the user's emotions in delivery plans and notification content. For example, if the system detects that the user is in a hurry, it will arrange for the shortest possible delivery time, resulting in faster delivery.

[0388] Notification content is adjusted based on analyzed emotions and sent to the user via input / output devices. Because notifications are concise and relevant to the user's needs, a stress-free and comfortable interface is provided.

[0389] For example, if a user requests "I want my pasta delivered immediately!", the server recognizes this "urgency" and generates a top-priority delivery plan tailored to that emotion. It then concisely informs the user of the estimated delivery time.

[0390] The following is an example of a possible prompt: "Write a program that uses an emotion engine to recognize the user's emotions and generate the optimal delivery route and time."

[0391] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0392] Step 1:

[0393] The user provides voice input through an input / output device. The voice data is captured as input and converted into a format for transmission to the speech recognition system. This allows the terminal to collect voice information as digital data and prepare the voice data necessary for the next processing step.

[0394] Step 2:

[0395] The server uses speech recognition software to convert audio data into text. This process takes audio data as input and outputs text. The Google Cloud Speech-to-Text API is used to convert speech into text format based on acoustic features. This makes the content of the speech available as text.

[0396] Step 3:

[0397] The server analyzes the user's emotions from text data using an emotion analysis engine. Based on the text data as input, it performs data calculations to identify emotions and outputs emotion tags (e.g., anxiety, joy, frustration, etc.) as the analysis result. IBM Watson is a possible emotion analysis engine to be used. Specifically, it analyzes the frequency of occurrence of certain patterns and keywords in the text to infer the emotional state.

[0398] Step 4:

[0399] The server uses a generative model to optimize order management and delivery planning. It processes sentiment analysis results and text data as input to optimize delivery schedules and routes. In this process, it utilizes a generative AI model (e.g., OpenAI GPT) to apply processing based on sentiment. The optimized delivery plan is then output.

[0400] Step 5:

[0401] The server processes the optimization results into notification content and sends it to the terminal. Here, it considers the optimization results and sentiment data as input, transforming them into a notification message that is easy for the user to understand. The output is a concise yet detailed notification for the user, allowing them to grasp the situation in real time.

[0402] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0403] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0404] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0405] [Third Embodiment]

[0406] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0407] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0408] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0409] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0410] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0411] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0412] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0413] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0414] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0415] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0416] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0417] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0418] The system of this invention primarily consists of a terminal that accepts voice input and a server that processes the input voice. The user inputs work-related instructions by voice through a microphone built into the terminal. This voice input includes scheduling meetings and changing task priorities.

[0419] The terminal receives voice input and sends it to the server as digital voice data. The server converts the received voice data into text using speech recognition technology. By analyzing this textual data, it is possible to clearly understand the user's instructions.

[0420] The server optimizes the assigned tasks using a generative model. In this process, it determines the optimal schedule and adjusts task priorities based on factors such as the company's calendar, existing appointments, and resource availability.

[0421] Once optimization is complete, the server notifies the terminal of optimized task and schedule information. The terminal receives the information from the server and notifies the user visually or audibly. This allows the user to stay informed of the latest schedule status and task priorities.

[0422] For example, if a user commands, "Schedule a project meeting for next Tuesday," the device sends voice data to the server. The server converts the voice data to text, determines the optimal schedule for the project meeting through a generative model, and returns the result to the device. The device then notifies the user, "We have scheduled a project meeting for next Tuesday at 2 PM." This process ensures that meeting schedules are set efficiently and automatically, eliminating the need for manual adjustments by the user.

[0423] The following describes the processing flow.

[0424] Step 1:

[0425] Users give voice instructions to the device's microphone to change meeting schedules or tasks. These voice instructions are recorded by the device as digital audio data.

[0426] Step 2:

[0427] The device uses a network connection to upload the recorded audio data to the server in order to send it to the server.

[0428] Step 3:

[0429] The server converts the received audio data into text using speech recognition technology. This allows the content of the voice instructions to be saved as digital text.

[0430] Step 4:

[0431] The server analyzes the transcribed instructions and identifies the task processing requirements based on their content. The results of this analysis determine the necessary business processes.

[0432] Step 5:

[0433] The server uses a generative model to optimize tasks based on the assigned work content. This optimization includes a process of scheduling based on specific conditions.

[0434] Step 6:

[0435] The server sends optimized schedule and task information to the device, including in the form of text and notification data.

[0436] Step 7:

[0437] The terminal analyzes the information received from the server and notifies the user of its contents visually or audibly. This allows the user to confirm the results.

[0438] Step 8:

[0439] The user reviews the received notification and provides additional voice instructions as needed. For example, they can perform instructions to change or cancel something.

[0440] Step 9:

[0441] The terminal sends additional user instructions to the server again, and the same process is restarted.

[0442] This series of steps allows the system to efficiently automate business tasks based on user voice commands, making management easier.

[0443] (Example 1)

[0444] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0445] There is a need to streamline the processing of work instructions based on voice input and to automatically optimize schedules and task management. Conventional systems require considerable effort and time to process voice data and optimize tasks, and also require user interaction, so new technologies are needed to solve this problem.

[0446] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0447] In this invention, the server includes means for processing voice information with computing resources to convert it into text information, means for optimizing the work using a generation algorithm, and means for notifying the device of the optimized work. This enables the automatic processing of work instructions via voice input, and allows for the rapid and efficient execution of optimized schedules and task management.

[0448] "Voice input" is a method by which digital devices receive instructions and information spoken by the user.

[0449] A "device" is a collection of hardware and software designed to perform a specific function.

[0450] "Voice information" refers to digital data acquired through voice input.

[0451] "Text information" refers to digital data converted into a string format, a format that can be processed more easily by computers.

[0452] "Computational resources" refer to the processing power and storage capacity provided by servers and computers.

[0453] A "generative algorithm" is a set of steps for calculating new outputs or results based on existing data.

[0454] "Work optimization" is the process of efficiently and effectively adjusting tasks and schedules set to achieve a specific goal.

[0455] "Notification" refers to sending a message to inform a user of information.

[0456] "Meeting planning" refers to the process of making the necessary preparations and setting schedules for holding a meeting or event.

[0457] "Priority" is an indicator that shows the order and importance required when performing tasks or processes.

[0458] The system of this invention efficiently processes work instructions based on voice input and automatically optimizes schedules and task management. Specific embodiments for carrying out the invention are shown below.

[0459] First, the user uses a device equipped with a voice input device. This device includes a microphone and is a digital device such as a smartphone or personal computer. The user verbally inputs instructions into this device, and the device captures the voice as digital audio data.

[0460] Next, the device transmits the acquired audio data to a server via the internet connection. The server, equipped with powerful computing resources, uses speech recognition technology to convert the audio data into text data. This process utilizes commercial speech recognition services or open-source speech processing libraries.

[0461] After the voice data is converted to text data, the server uses a generative AI model (for example, a generative model specializing in natural language processing) to analyze the user's instructions and extract specific task information. For example, if the user instructs, "Schedule a project meeting for next Tuesday," this command is understood as a request to schedule a meeting. The server then refers to the company's calendar information, resource availability, and other appointments to determine the optimal meeting time. The generative model uses prompt statements to break down the task into smaller parts and generate an optimized schedule. An example of a prompt statement might be, "If the user wants to schedule a project meeting, please suggest the factors to consider and the optimal scheduling time."

[0462] Ultimately, the server sends optimized schedule information and task priorities back to the terminal. Upon receiving this information, the terminal prepares to notify the user visually or audibly. This allows the user to receive information such as, "We have scheduled a project meeting for next Tuesday at 2 PM." This entire process enables users to efficiently manage their schedules and tasks without any hassle.

[0463] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0464] Step 1:

[0465] The user verbally inputs work instructions using a voice input device. The terminal captures this audio through the microphone and temporarily stores it as digital audio data. At this point, the input is the user's voice instructions, and the output is digital audio data.

[0466] Step 2:

[0467] The terminal sends audio data to the server. The audio data is transferred to the server via network communication. Specifically, this involves packetization of digital audio data and data transfer based on a protocol. The input is the audio data stored on the terminal, and the output is the audio data delivered to the server.

[0468] Step 3:

[0469] The server uses speech recognition technology to convert audio data into text data. Specifically, the speech recognition engine applies speech signal processing algorithms to analyze phonemes and convert them into text based on a language model. The input is audio data, and the output is text data.

[0470] Step 4:

[0471] The server uses a generative AI model to analyze text data and understand user instructions. At this stage, natural language processing is performed, and the instructions are extracted using a parsing algorithm. The generative AI model then uses the prompt text to perform additional inference and optimization. The input is text data, and the output is the parsed instructions.

[0472] Step 5:

[0473] The server optimizes tasks based on user instructions. Here, the optimal schedule is calculated by referencing the company's calendar and resource availability. Generative model prompts are used to refine the calculation results. The input is the parsed instructions, and the output is the optimized schedule.

[0474] Step 6:

[0475] The server sends optimized schedule information to the terminal. The calculated schedule is sent back to the terminal via network communication. The input is the optimized schedule, and the output is the schedule information transmitted to the terminal.

[0476] Step 7:

[0477] The device notifies the user of the schedule information it has received. Specifically, an application on the device conveys the information to the user visually or audibly. The notification app displays or plays an audio message, and the user receives a notification such as, "A project meeting has been scheduled for next Tuesday at 2 PM." The input is the schedule information received from the server, and the output is the notification result sent to the user.

[0478] (Application Example 1)

[0479] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0480] In factory manufacturing environments, automation using voice commands is needed to quickly and efficiently give work instructions and adjust schedules. However, current technology does not adequately optimize the work process after receiving voice commands, resulting in reduced work efficiency.

[0481] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0482] In this invention, the server includes means for processing voice information with a data processing device for converting it into text information, means for optimizing work using a generative model, and means for controlling work processes in the manufacturing industry with voice instructions. This enables smooth reception of work instructions by voice at the factory floor and the generation and execution of an optimal work schedule.

[0483] A "device for receiving voice input" is a device that detects voice commands from a user and acquires that voice information.

[0484] A "data processing device for converting audio information into text information" is a device that analyzes acquired audio information and performs the process of converting it into text information.

[0485] "Methods for optimizing work using generative models" refer to methods that utilize generative models to plan efficient work schedules based on work instruction data.

[0486] "Means for notifying the device of optimized work" refers to means for informing the corresponding device of optimized work schedules and instructions.

[0487] "Means for receiving and processing additional voice instructions from the user" refers to means that, when a user gives a new instruction, the system receives that instruction and processes the necessary actions.

[0488] "Means for controlling work processes in manufacturing using voice commands" refers to means for receiving voice commands for various tasks in a manufacturing process and controlling the content of those tasks accordingly.

[0489] To implement this invention, a device for receiving voice input and a server for processing it are required. The user gives work instructions by voice using a device such as a smartphone or tablet. This device is equipped with a high-performance microphone that can reliably capture the voice instructions spoken by the user.

[0490] The server functions as a data processing unit for converting speech information into text information. This process utilizes the Google Cloud Speech-to-Text API as the speech recognition technology. This API enables high-accuracy speech recognition and is used to analyze the resulting text information.

[0491] By utilizing a generative model, the server optimizes the work process based on work instructions from the user. For example, it uses the Python library scipy.optimize to generate the optimal work schedule within the given resources and constraints. This optimized schedule information is then notified to the device.

[0492] If the user provides additional voice instructions, these instructions will also be accepted and processed in real time by the server, which will readjust the work schedule as needed. This will improve productivity on the factory floor and enable more efficient work progress.

[0493] For example, if a factory line operator gives a voice command such as, "Finish the next 100 products within 3 hours," this voice command is acquired by the device and converted into text data by a server. Then, an optimal production schedule is formulated using a generative AI model, and the robots begin operation at the appropriate time.

[0494] An example of a prompt message might be, "Generate the optimal work schedule based on this voice command and send the operation instructions to the designated robot." This would enable the automation and efficiency of tasks based on voice commands.

[0495] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0496] Step 1:

[0497] The user inputs work instructions by voice through the terminal. The input voice is captured as digital audio data by the microphone built into the terminal. This data is input as an audio signal.

[0498] Step 2:

[0499] The device sends the acquired digital audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert this audio data into text. The output is the text data converted from the audio.

[0500] Step 3:

[0501] The server analyzes text data obtained using a generative model and optimizes the work schedule. Here, the Python library scipy.optimize is used to calculate the optimal schedule based on the input work instruction data. The optimized schedule information is then output.

[0502] Step 4:

[0503] The server notifies the terminal of optimized schedule information. The terminal presents this information to the user visually or audibly. As a result of this output, the user can check the latest schedule.

[0504] Step 5:

[0505] If the user provides additional voice instructions, the process from step 1 is repeated. This allows for continuous real-time adjustment of the work schedule. The output is the updated schedule information.

[0506] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0507] This invention combines a system comprising a terminal that accepts voice input, a server that converts voice data into text, and a generative model that optimizes tasks with an emotion engine that recognizes the user's emotions. The user can use the terminal to give voice instructions for scheduling meetings or changing task priorities. The terminal transmits these voice instructions to the server as digital voice data.

[0508] The server converts the received audio data into text and optimizes the task using a generative model. During this process, an emotion engine analyzes the user's emotions from the audio data and adjusts the optimization parameters accordingly. For example, if the user is stressed, the system can ease the schedule and postpone less important tasks.

[0509] Furthermore, based on the analysis results of the emotion engine, the device can change the way notifications are presented. For example, it can switch to concise and direct notifications when the user is busy or irritated, providing an interface that is considerate of the user's emotions.

[0510] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the emotion engine recognizes this emotion and adjusts the system's response accordingly. The server quickly processes the meeting cancellation, and the terminal sends a concise notification to the user stating, "Tomorrow's meeting has been canceled." This enables flexible responses that take emotions into consideration, reducing the user's workload.

[0511] In this way, this system, which incorporates an emotion engine, brings voice-input-based task management and automation closer to a more human-like response, providing a more user-friendly tool.

[0512] The following describes the processing flow.

[0513] Step 1:

[0514] The user gives work-related instructions by voice into the device's microphone. For example, they might say, "Cancel this Friday's meeting." This voice is captured by the device as digital audio data.

[0515] Step 2:

[0516] The device uses a network connection to upload the captured audio data to the server. Appropriate security protocols are used during this process to ensure data integrity.

[0517] Step 3:

[0518] The server receives the uploaded audio data and converts it into digital text using a speech recognition engine. Accuracy is crucial for this text conversion, as it forms the basis for accurately understanding the instructions.

[0519] Step 4:

[0520] The server analyzes the converted text and uses an emotion engine to identify emotions from the user's voice. This emotion information is used to identify a variety of emotional states, such as stress, anger, and joy.

[0521] Step 5:

[0522] The generative model calculates optimized task and schedule suggestions based on text and sentiment information. For example, if it determines that a user is stressed, it will suggest prioritizing the cancellation of low-priority meetings.

[0523] Step 6:

[0524] The server sends task and schedule information to the device, adjusted according to the user's emotions. At this point, the server adjusts the priority of information and the wording of notifications to match the user's emotions.

[0525] Step 7:

[0526] The device analyzes the received information and notifies the user in a way that is appropriate to their emotions. To encourage user focus, notifications may use concise and direct language.

[0527] Step 8:

[0528] The user reviews the notification, provides further voice instructions if necessary, and the process restarts. Emotion-responsive processing is a crucial element in improving the user experience.

[0529] (Example 2)

[0530] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0531] Conventional voice input systems simply convert speech to text, making it difficult to dynamically optimize tasks and adjust interfaces while considering the user's emotions and state of mind. Therefore, there was a need for flexible and human-like responses to reduce user stress and frustration.

[0532] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0533] In this invention, the server includes means for processing voice information with a communication device for converting it into text information, means for performing emotion analysis and dynamically adjusting task priorities based on the analysis results, and means for changing the interface representation according to the user's emotional state. This enables task optimization and notification adjustment that are adapted to the user's emotions and state.

[0534] A "device for receiving voice input" is a device that obtains voice instructions from a user and processes them as digital data.

[0535] "Communication equipment for converting audio information into text information" refers to a server or network-based device that converts received audio data into text data.

[0536] "Optimizing tasks using generative algorithms" is an AI technology that aims to improve the efficiency and priority of tasks based on the data received.

[0537] "Means for notifying the device of optimized work" refers to a function that sends optimization information processed by the server to the user's device.

[0538] "Means for receiving and processing additional voice instructions from users" refers to a system that allows users to give additional voice instructions after the initial input, and that continuously receives and processes these instructions.

[0539] "A means of performing emotion analysis and dynamically adjusting task priorities" refers to a system that analyzes emotions from the user's voice and adjusts task priorities in real time based on the results.

[0540] "Means of changing interface expression according to the user's emotional state" refers to a function that adapts the display method and notification content of the device to the situation based on the emotions the user is feeling.

[0541] This invention provides a mechanism for a task management system that uses voice input to achieve more flexible responses that are considerate of human emotions. The system includes a device for receiving voice input, a communication device for converting voice information into text information, and a configuration for optimizing tasks using a generation algorithm.

[0542] The user first provides voice instructions for the task to a voice input device. These voice instructions are captured using a terminal device equipped with a standard microphone. The terminal sends the captured voice to a server, which uses speech recognition technology to convert the voice data into text data. Google Cloud Speech-to-Text or similar speech recognition APIs can be used for this process.

[0543] The server uses a generative AI model to optimize tasks based on the converted text data. This process employs generative algorithms such as OpenAI's GPT series. Furthermore, the server utilizes an emotion engine to analyze the user's emotions from their voice, dynamically adjusting task and schedule priorities. As a result, if the user is experiencing stress, the system can automatically ease the schedule or postpone less important tasks.

[0544] Furthermore, based on sentiment analysis, the device modifies the content and wording of notifications to the user, providing an interface that is appropriate to the user's current emotional state. This feature allows for, for example, a concise and direct notification such as "I have canceled tomorrow's meeting" if the device determines that the user is in a bad mood.

[0545] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the system's emotion engine will recognize that emotion and quickly process the meeting cancellation. At this time, a notification will appear on the device stating, "Tomorrow's meeting has been canceled."

[0546] An example of a prompt for a generative AI model might be, "How would the system respond if a user, in an unhappy voice, requested that the next meeting be postponed?"

[0547] This configuration enables intuitive and effective task management and notifications that respond to the user's emotions and state.

[0548] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0549] Step 1:

[0550] The user provides task instructions verbally through a voice input device. Specifically, they might give instructions such as, "Cancel tomorrow's meeting." The input data is voice data. The terminal acquires this as an audio signal and stores it as digital data. This voice data forms the basis for the next processing step.

[0551] Step 2:

[0552] The terminal sends the acquired voice data to the server. At this time, the voice data is received by communication equipment via the network. The input received by the server is digital voice data, which is converted into text information using a speech recognition API (e.g., Google Cloud Speech-to-Text). The output generated is text information of the voice command.

[0553] Step 3:

[0554] The server receives the text information converted from the speech and uses a generative AI model (e.g., OpenAI's GPT series) to optimize the task. The input used here is the converted text information and previously analyzed emotion data. The emotion engine operates here, analyzing the user's emotions from the speech. Based on the emotion analysis results, the generative model dynamically adjusts the task priority and schedule to output optimized task information.

[0555] Step 4:

[0556] Upon receiving optimized task information from the server, the terminal adjusts the notification content to the user based on sentiment analysis. For example, if the user is in a bad mood, it generates a concise and direct notification such as, "I have canceled tomorrow's meeting." The inputs here are the optimized task information and the user's emotional state, and the output is the adjusted notification message.

[0557] Step 5:

[0558] The device presents the user with optimized notification messages. Notifications can be delivered visually or audibly. This allows users to receive appropriate notifications that take into account their emotions and task priorities, enabling effective task management. Output is delivered in the form of direct notifications to the user.

[0559] (Application Example 2)

[0560] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0561] In modern society, users may experience stress and dissatisfaction when using food delivery services, particularly regarding ordering and delivery. Insufficient voice instructions and emotionally responsive service are major factors that detract from the overall service experience. Furthermore, inadequate order management and delivery planning during delivery can reduce efficiency and negatively impact the user experience.

[0562] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0563] In this invention, the server includes means for processing audio data with an information processing device for converting it into text information, means for optimizing operations using a generative model, and means including an emotion analysis engine for analyzing the user's emotions. This enables the optimization of order management and delivery plans in accordance with the user's emotions, making it possible to provide a more personalized food delivery service.

[0564] "Voice input" refers to a means of input that allows users to convey instructions to devices and systems through their voice.

[0565] An "input / output device" is hardware that receives audio data and transmits the processing results to the user.

[0566] An "information processing device" is a computer system used to convert audio data into text information or to perform calculations using generative models.

[0567] "Textual information" refers to text data obtained by converting audio data into a format that can be processed by a computer.

[0568] A "generative model" is an artificial intelligence algorithm designed to optimize a specific task.

[0569] "Work optimization" refers to methods for maximizing the efficiency and effectiveness of work under specific conditions.

[0570] A "sentiment analysis engine" is software that analyzes and determines a user's emotions from their voice or text.

[0571] "Notification content" refers to messages or alerts sent to inform users of processing results or other information.

[0572] "Order management" refers to the entire process of receiving customer orders and processing and tracking them appropriately.

[0573] A "delivery plan" is a plan of schedules and routes to efficiently deliver ordered goods to the customer.

[0574] This invention realizes an order management system for food delivery services that takes into account voice commands and emotions. The system is configured as follows:

[0575] First, the user places an order by voice using an input / output device that accepts voice input (e.g., a smartphone or smart device). The input / output device sends this voice input as digital voice data to the server. The server uses an information processing device to convert the voice data into text. Specifically, it utilizes speech recognition software such as the Google Cloud Speech-to-Text API to convert the voice data into text information.

[0576] The server then uses a generative model to optimize order management and delivery planning. Possible generative models used here include OpenAI's GPT and deep learning models. This is used to calculate the shortest delivery route and the optimal delivery time.

[0577] Furthermore, an emotion analysis engine analyzes the user's emotions from text data. Emotion recognition tools such as IBM Watson are used for this analysis. This analysis reflects the user's emotions in delivery plans and notification content. For example, if the system detects that the user is in a hurry, it will arrange for the shortest possible delivery time, resulting in faster delivery.

[0578] Notification content is adjusted based on analyzed emotions and sent to the user via input / output devices. Because notifications are concise and relevant to the user's needs, a stress-free and comfortable interface is provided.

[0579] For example, if a user requests "I want my pasta delivered immediately!", the server recognizes this "urgency" and generates a top-priority delivery plan tailored to that emotion. It then concisely informs the user of the estimated delivery time.

[0580] The following is an example of a possible prompt: "Write a program that uses an emotion engine to recognize the user's emotions and generate the optimal delivery route and time."

[0581] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0582] Step 1:

[0583] The user provides voice input through an input / output device. The voice data is captured as input and converted into a format for transmission to the speech recognition system. This allows the terminal to collect voice information as digital data and prepare the voice data necessary for the next processing step.

[0584] Step 2:

[0585] The server uses speech recognition software to convert audio data into text. This process takes audio data as input and outputs text. The Google Cloud Speech-to-Text API is used to convert speech into text format based on acoustic features. This makes the content of the speech available as text.

[0586] Step 3:

[0587] The server analyzes the user's emotions from text data using an emotion analysis engine. Based on the text data as input, it performs data calculations to identify emotions and outputs emotion tags (e.g., anxiety, joy, frustration, etc.) as the analysis result. IBM Watson is a possible emotion analysis engine to be used. Specifically, it analyzes the frequency of occurrence of certain patterns and keywords in the text to infer the emotional state.

[0588] Step 4:

[0589] The server uses a generative model to optimize order management and delivery planning. It processes sentiment analysis results and text data as input to optimize delivery schedules and routes. In this process, it utilizes a generative AI model (e.g., OpenAI GPT) to apply processing based on sentiment. The optimized delivery plan is then output.

[0590] Step 5:

[0591] The server processes the optimization results into notification content and sends it to the terminal. Here, it considers the optimization results and sentiment data as input, transforming them into a notification message that is easy for the user to understand. The output is a concise yet detailed notification for the user, allowing them to grasp the situation in real time.

[0592] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0593] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0594] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0595] [Fourth Embodiment]

[0596] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0597] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0598] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0599] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0600] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0601] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0602] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0603] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0604] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0605] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0606] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0607] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0608] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0609] The system of this invention primarily consists of a terminal that accepts voice input and a server that processes the input voice. The user inputs work-related instructions by voice through a microphone built into the terminal. This voice input includes scheduling meetings and changing task priorities.

[0610] The terminal receives voice input and sends it to the server as digital voice data. The server converts the received voice data into text using speech recognition technology. By analyzing this textual data, it is possible to clearly understand the user's instructions.

[0611] The server optimizes the assigned tasks using a generative model. In this process, it determines the optimal schedule and adjusts task priorities based on factors such as the company's calendar, existing appointments, and resource availability.

[0612] Once optimization is complete, the server notifies the terminal of optimized task and schedule information. The terminal receives the information from the server and notifies the user visually or audibly. This allows the user to stay informed of the latest schedule status and task priorities.

[0613] For example, if a user commands, "Schedule a project meeting for next Tuesday," the device sends voice data to the server. The server converts the voice data to text, determines the optimal schedule for the project meeting through a generative model, and returns the result to the device. The device then notifies the user, "We have scheduled a project meeting for next Tuesday at 2 PM." This process ensures that meeting schedules are set efficiently and automatically, eliminating the need for manual adjustments by the user.

[0614] The following describes the processing flow.

[0615] Step 1:

[0616] Users give voice instructions to the device's microphone to change meeting schedules or tasks. These voice instructions are recorded by the device as digital audio data.

[0617] Step 2:

[0618] The device uses a network connection to upload the recorded audio data to the server in order to send it to the server.

[0619] Step 3:

[0620] The server converts the received audio data into text using speech recognition technology. This allows the content of the voice instructions to be saved as digital text.

[0621] Step 4:

[0622] The server analyzes the transcribed instructions and identifies the task processing requirements based on their content. The results of this analysis determine the necessary business processes.

[0623] Step 5:

[0624] The server uses a generative model to optimize tasks based on the assigned work content. This optimization includes a process of scheduling based on specific conditions.

[0625] Step 6:

[0626] The server sends optimized schedule and task information to the device, including in the form of text and notification data.

[0627] Step 7:

[0628] The terminal analyzes the information received from the server and notifies the user of its contents visually or audibly. This allows the user to confirm the results.

[0629] Step 8:

[0630] The user reviews the received notification and provides additional voice instructions as needed. For example, they can perform instructions to change or cancel something.

[0631] Step 9:

[0632] The terminal sends additional user instructions to the server again, and the same process is restarted.

[0633] This series of steps allows the system to efficiently automate business tasks based on user voice commands, making management easier.

[0634] (Example 1)

[0635] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0636] There is a need to streamline the processing of work instructions based on voice input and to automatically optimize schedules and task management. Conventional systems require considerable effort and time to process voice data and optimize tasks, and also require user interaction, so new technologies are needed to solve this problem.

[0637] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0638] In this invention, the server includes means for processing voice information with computing resources to convert it into text information, means for optimizing the work using a generation algorithm, and means for notifying the device of the optimized work. This enables the automatic processing of work instructions via voice input, and allows for the rapid and efficient execution of optimized schedules and task management.

[0639] "Voice input" is a method by which digital devices receive instructions and information spoken by the user.

[0640] A "device" is a collection of hardware and software designed to perform a specific function.

[0641] "Voice information" refers to digital data acquired through voice input.

[0642] "Text information" refers to digital data converted into a string format, a format that can be processed more easily by computers.

[0643] "Computational resources" refer to the processing power and storage capacity provided by servers and computers.

[0644] A "generative algorithm" is a set of steps for calculating new outputs or results based on existing data.

[0645] "Work optimization" is the process of efficiently and effectively adjusting tasks and schedules set to achieve a specific goal.

[0646] "Notification" refers to sending a message to inform a user of information.

[0647] "Meeting planning" refers to the process of making the necessary preparations and setting schedules for holding a meeting or event.

[0648] "Priority" is an indicator that shows the order and importance required when performing tasks or processes.

[0649] The system of this invention efficiently processes work instructions based on voice input and automatically optimizes schedules and task management. Specific embodiments for carrying out the invention are shown below.

[0650] First, the user uses a device equipped with a voice input device. This device includes a microphone and is a digital device such as a smartphone or personal computer. The user verbally inputs instructions into this device, and the device captures the voice as digital audio data.

[0651] Next, the device transmits the acquired audio data to a server via the internet connection. The server, equipped with powerful computing resources, uses speech recognition technology to convert the audio data into text data. This process utilizes commercial speech recognition services or open-source speech processing libraries.

[0652] After the voice data is converted to text data, the server uses a generative AI model (for example, a generative model specializing in natural language processing) to analyze the user's instructions and extract specific task information. For example, if the user instructs, "Schedule a project meeting for next Tuesday," this command is understood as a request to schedule a meeting. The server then refers to the company's calendar information, resource availability, and other appointments to determine the optimal meeting time. The generative model uses prompt statements to break down the task into smaller parts and generate an optimized schedule. An example of a prompt statement might be, "If the user wants to schedule a project meeting, please suggest the factors to consider and the optimal scheduling time."

[0653] Ultimately, the server sends optimized schedule information and task priorities back to the terminal. Upon receiving this information, the terminal prepares to notify the user visually or audibly. This allows the user to receive information such as, "We have scheduled a project meeting for next Tuesday at 2 PM." This entire process enables users to efficiently manage their schedules and tasks without any hassle.

[0654] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0655] Step 1:

[0656] The user verbally inputs work instructions using a voice input device. The terminal captures this audio through the microphone and temporarily stores it as digital audio data. At this point, the input is the user's voice instructions, and the output is digital audio data.

[0657] Step 2:

[0658] The terminal sends audio data to the server. The audio data is transferred to the server via network communication. Specifically, this involves packetization of digital audio data and data transfer based on a protocol. The input is the audio data stored on the terminal, and the output is the audio data delivered to the server.

[0659] Step 3:

[0660] The server uses speech recognition technology to convert audio data into text data. Specifically, the speech recognition engine applies speech signal processing algorithms to analyze phonemes and convert them into text based on a language model. The input is audio data, and the output is text data.

[0661] Step 4:

[0662] The server uses a generative AI model to analyze text data and understand user instructions. At this stage, natural language processing is performed, and the instructions are extracted using a parsing algorithm. The generative AI model then uses the prompt text to perform additional inference and optimization. The input is text data, and the output is the parsed instructions.

[0663] Step 5:

[0664] The server optimizes tasks based on user instructions. Here, the optimal schedule is calculated by referencing the company's calendar and resource availability. Generative model prompts are used to refine the calculation results. The input is the parsed instructions, and the output is the optimized schedule.

[0665] Step 6:

[0666] The server sends optimized schedule information to the terminal. The calculated schedule is sent back to the terminal via network communication. The input is the optimized schedule, and the output is the schedule information transmitted to the terminal.

[0667] Step 7:

[0668] The device notifies the user of the schedule information it has received. Specifically, an application on the device conveys the information to the user visually or audibly. The notification app displays or plays an audio message, and the user receives a notification such as, "A project meeting has been scheduled for next Tuesday at 2 PM." The input is the schedule information received from the server, and the output is the notification result sent to the user.

[0669] (Application Example 1)

[0670] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] In factory manufacturing environments, automation using voice commands is needed to quickly and efficiently give work instructions and adjust schedules. However, current technology does not adequately optimize the work process after receiving voice commands, resulting in reduced work efficiency.

[0672] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0673] In this invention, the server includes means for processing voice information with a data processing device for converting it into text information, means for optimizing work using a generative model, and means for controlling work processes in the manufacturing industry with voice instructions. This enables smooth reception of work instructions by voice at the factory floor and the generation and execution of an optimal work schedule.

[0674] A "device for receiving voice input" is a device that detects voice commands from a user and acquires that voice information.

[0675] A "data processing device for converting audio information into text information" is a device that analyzes acquired audio information and performs the process of converting it into text information.

[0676] "Methods for optimizing work using generative models" refer to methods that utilize generative models to plan efficient work schedules based on work instruction data.

[0677] "Means for notifying the device of optimized work" refers to means for informing the corresponding device of optimized work schedules and instructions.

[0678] "Means for receiving and processing additional voice instructions from the user" refers to means that, when a user gives a new instruction, the system receives that instruction and processes the necessary actions.

[0679] "Means for controlling work processes in manufacturing using voice commands" refers to means for receiving voice commands for various tasks in a manufacturing process and controlling the content of those tasks accordingly.

[0680] To implement this invention, a device for receiving voice input and a server for processing it are required. The user gives work instructions by voice using a device such as a smartphone or tablet. This device is equipped with a high-performance microphone that can reliably capture the voice instructions spoken by the user.

[0681] The server functions as a data processing unit for converting speech information into text information. This process utilizes the Google Cloud Speech-to-Text API as the speech recognition technology. This API enables high-accuracy speech recognition and is used to analyze the resulting text information.

[0682] By utilizing a generative model, the server optimizes the work process based on work instructions from the user. For example, it uses the Python library scipy.optimize to generate the optimal work schedule within the given resources and constraints. This optimized schedule information is then notified to the device.

[0683] If the user provides additional voice instructions, these instructions will also be accepted and processed in real time by the server, which will readjust the work schedule as needed. This will improve productivity on the factory floor and enable more efficient work progress.

[0684] For example, if a factory line operator gives a voice command such as, "Finish the next 100 products within 3 hours," this voice command is acquired by the device and converted into text data by a server. Then, an optimal production schedule is formulated using a generative AI model, and the robots begin operation at the appropriate time.

[0685] An example of a prompt message might be, "Generate the optimal work schedule based on this voice command and send the operation instructions to the designated robot." This would enable the automation and efficiency of tasks based on voice commands.

[0686] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0687] Step 1:

[0688] The user inputs work instructions by voice through the terminal. The input voice is captured as digital audio data by the microphone built into the terminal. This data is input as an audio signal.

[0689] Step 2:

[0690] The device sends the acquired digital audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert this audio data into text. The output is the text data converted from the audio.

[0691] Step 3:

[0692] The server analyzes text data obtained using a generative model and optimizes the work schedule. Here, the Python library scipy.optimize is used to calculate the optimal schedule based on the input work instruction data. The optimized schedule information is then output.

[0693] Step 4:

[0694] The server notifies the terminal of optimized schedule information. The terminal presents this information to the user visually or audibly. As a result of this output, the user can check the latest schedule.

[0695] Step 5:

[0696] If the user provides additional voice instructions, the process from step 1 is repeated. This allows for continuous real-time adjustment of the work schedule. The output is the updated schedule information.

[0697] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0698] This invention combines a system comprising a terminal that accepts voice input, a server that converts voice data into text, and a generative model that optimizes tasks with an emotion engine that recognizes the user's emotions. The user can use the terminal to give voice instructions for scheduling meetings or changing task priorities. The terminal transmits these voice instructions to the server as digital voice data.

[0699] The server converts the received audio data into text and optimizes the task using a generative model. During this process, an emotion engine analyzes the user's emotions from the audio data and adjusts the optimization parameters accordingly. For example, if the user is stressed, the system can ease the schedule and postpone less important tasks.

[0700] Furthermore, based on the analysis results of the emotion engine, the device can change the way notifications are presented. For example, it can switch to concise and direct notifications when the user is busy or irritated, providing an interface that is considerate of the user's emotions.

[0701] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the emotion engine recognizes this emotion and adjusts the system's response accordingly. The server quickly processes the meeting cancellation, and the terminal sends a concise notification to the user stating, "Tomorrow's meeting has been canceled." This enables flexible responses that take emotions into consideration, reducing the user's workload.

[0702] In this way, this system, which incorporates an emotion engine, brings voice-input-based task management and automation closer to a more human-like response, providing a more user-friendly tool.

[0703] The following describes the processing flow.

[0704] Step 1:

[0705] The user gives work-related instructions by voice into the device's microphone. For example, they might say, "Cancel this Friday's meeting." This voice is captured by the device as digital audio data.

[0706] Step 2:

[0707] The device uses a network connection to upload the captured audio data to the server. Appropriate security protocols are used during this process to ensure data integrity.

[0708] Step 3:

[0709] The server receives the uploaded audio data and converts it into digital text using a speech recognition engine. Accuracy is crucial for this text conversion, as it forms the basis for accurately understanding the instructions.

[0710] Step 4:

[0711] The server analyzes the converted text and uses an emotion engine to identify emotions from the user's voice. This emotion information is used to identify a variety of emotional states, such as stress, anger, and joy.

[0712] Step 5:

[0713] The generative model calculates optimized task and schedule suggestions based on text and sentiment information. For example, if it determines that a user is stressed, it will suggest prioritizing the cancellation of low-priority meetings.

[0714] Step 6:

[0715] The server sends task and schedule information to the device, adjusted according to the user's emotions. At this point, the server adjusts the priority of information and the wording of notifications to match the user's emotions.

[0716] Step 7:

[0717] The device analyzes the received information and notifies the user in a way that is appropriate to their emotions. To encourage user focus, notifications may use concise and direct language.

[0718] Step 8:

[0719] The user reviews the notification, provides further voice instructions if necessary, and the process restarts. Emotion-responsive processing is a crucial element in improving the user experience.

[0720] (Example 2)

[0721] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0722] Conventional voice input systems simply convert speech to text, making it difficult to dynamically optimize tasks and adjust interfaces while considering the user's emotions and state of mind. Therefore, there was a need for flexible and human-like responses to reduce user stress and frustration.

[0723] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0724] In this invention, the server includes means for processing voice information with a communication device for converting it into text information, means for performing emotion analysis and dynamically adjusting task priorities based on the analysis results, and means for changing the interface representation according to the user's emotional state. This enables task optimization and notification adjustment that are adapted to the user's emotions and state.

[0725] A "device for receiving voice input" is a device that obtains voice instructions from a user and processes them as digital data.

[0726] "Communication equipment for converting audio information into text information" refers to a server or network-based device that converts received audio data into text data.

[0727] "Optimizing tasks using generative algorithms" is an AI technology that aims to improve the efficiency and priority of tasks based on the data received.

[0728] "Means for notifying the device of optimized work" refers to a function that sends optimization information processed by the server to the user's device.

[0729] "Means for receiving and processing additional voice instructions from users" refers to a system that allows users to give additional voice instructions after the initial input, and that continuously receives and processes these instructions.

[0730] "A means of performing emotion analysis and dynamically adjusting task priorities" refers to a system that analyzes emotions from the user's voice and adjusts task priorities in real time based on the results.

[0731] "Means of changing interface expression according to the user's emotional state" refers to a function that adapts the display method and notification content of the device to the situation based on the emotions the user is feeling.

[0732] This invention provides a mechanism for a task management system that uses voice input to achieve more flexible responses that are considerate of human emotions. The system includes a device for receiving voice input, a communication device for converting voice information into text information, and a configuration for optimizing tasks using a generation algorithm.

[0733] The user first provides voice instructions for the task to a voice input device. These voice instructions are captured using a terminal device equipped with a standard microphone. The terminal sends the captured voice to a server, which uses speech recognition technology to convert the voice data into text data. Google Cloud Speech-to-Text or similar speech recognition APIs can be used for this process.

[0734] The server uses a generative AI model to optimize tasks based on the converted text data. This process employs generative algorithms such as OpenAI's GPT series. Furthermore, the server utilizes an emotion engine to analyze the user's emotions from their voice, dynamically adjusting task and schedule priorities. As a result, if the user is experiencing stress, the system can automatically ease the schedule or postpone less important tasks.

[0735] Furthermore, based on sentiment analysis, the device modifies the content and wording of notifications to the user, providing an interface that is appropriate to the user's current emotional state. This feature allows for, for example, a concise and direct notification such as "I have canceled tomorrow's meeting" if the device determines that the user is in a bad mood.

[0736] For example, if a user gives an unhappy instruction such as "Cancel tomorrow's meeting," the system's emotion engine will recognize that emotion and quickly process the meeting cancellation. At this time, a notification will appear on the device stating, "Tomorrow's meeting has been canceled."

[0737] An example of a prompt for a generative AI model might be, "How would the system respond if a user, in an unhappy voice, requested that the next meeting be postponed?"

[0738] This configuration enables intuitive and effective task management and notifications that respond to the user's emotions and state.

[0739] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0740] Step 1:

[0741] The user provides task instructions verbally through a voice input device. Specifically, they might give instructions such as, "Cancel tomorrow's meeting." The input data is voice data. The terminal acquires this as an audio signal and stores it as digital data. This voice data forms the basis for the next processing step.

[0742] Step 2:

[0743] The terminal sends the acquired voice data to the server. At this time, the voice data is received by communication equipment via the network. The input received by the server is digital voice data, which is converted into text information using a speech recognition API (e.g., Google Cloud Speech-to-Text). The output generated is text information of the voice command.

[0744] Step 3:

[0745] The server receives the text information converted from the speech and uses a generative AI model (e.g., OpenAI's GPT series) to optimize the task. The input used here is the converted text information and previously analyzed emotion data. The emotion engine operates here, analyzing the user's emotions from the speech. Based on the emotion analysis results, the generative model dynamically adjusts the task priority and schedule to output optimized task information.

[0746] Step 4:

[0747] Upon receiving optimized task information from the server, the terminal adjusts the notification content to the user based on sentiment analysis. For example, if the user is in a bad mood, it generates a concise and direct notification such as, "I have canceled tomorrow's meeting." The inputs here are the optimized task information and the user's emotional state, and the output is the adjusted notification message.

[0748] Step 5:

[0749] The device presents the user with optimized notification messages. Notifications can be delivered visually or audibly. This allows users to receive appropriate notifications that take into account their emotions and task priorities, enabling effective task management. Output is delivered in the form of direct notifications to the user.

[0750] (Application Example 2)

[0751] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0752] In modern society, users may experience stress and dissatisfaction when using food delivery services, particularly regarding ordering and delivery. Insufficient voice instructions and emotionally responsive service are major factors that detract from the overall service experience. Furthermore, inadequate order management and delivery planning during delivery can reduce efficiency and negatively impact the user experience.

[0753] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0754] In this invention, the server includes means for processing audio data with an information processing device for converting it into text information, means for optimizing operations using a generative model, and means including an emotion analysis engine for analyzing the user's emotions. This enables the optimization of order management and delivery plans in accordance with the user's emotions, making it possible to provide a more personalized food delivery service.

[0755] "Voice input" refers to a means of input that allows users to convey instructions to devices and systems through their voice.

[0756] An "input / output device" is hardware that receives audio data and transmits the processing results to the user.

[0757] An "information processing device" is a computer system used to convert audio data into text information or to perform calculations using generative models.

[0758] "Textual information" refers to text data obtained by converting audio data into a format that can be processed by a computer.

[0759] A "generative model" is an artificial intelligence algorithm designed to optimize a specific task.

[0760] "Work optimization" refers to methods for maximizing the efficiency and effectiveness of work under specific conditions.

[0761] A "sentiment analysis engine" is software that analyzes and determines a user's emotions from their voice or text.

[0762] "Notification content" refers to messages or alerts sent to inform users of processing results or other information.

[0763] "Order management" refers to the entire process of receiving customer orders and processing and tracking them appropriately.

[0764] A "delivery plan" is a plan of schedules and routes to efficiently deliver ordered goods to the customer.

[0765] This invention realizes an order management system for food delivery services that takes into account voice commands and emotions. The system is configured as follows:

[0766] First, the user places an order by voice using an input / output device that accepts voice input (e.g., a smartphone or smart device). The input / output device sends this voice input as digital voice data to the server. The server uses an information processing device to convert the voice data into text. Specifically, it utilizes speech recognition software such as the Google Cloud Speech-to-Text API to convert the voice data into text information.

[0767] The server then uses a generative model to optimize order management and delivery planning. Possible generative models used here include OpenAI's GPT and deep learning models. This is used to calculate the shortest delivery route and the optimal delivery time.

[0768] Furthermore, an emotion analysis engine analyzes the user's emotions from text data. Emotion recognition tools such as IBM Watson are used for this analysis. This analysis reflects the user's emotions in delivery plans and notification content. For example, if the system detects that the user is in a hurry, it will arrange for the shortest possible delivery time, resulting in faster delivery.

[0769] Notification content is adjusted based on analyzed emotions and sent to the user via input / output devices. Because notifications are concise and relevant to the user's needs, a stress-free and comfortable interface is provided.

[0770] For example, if a user requests "I want my pasta delivered immediately!", the server recognizes this "urgency" and generates a top-priority delivery plan tailored to that emotion. It then concisely informs the user of the estimated delivery time.

[0771] The following is an example of a possible prompt: "Write a program that uses an emotion engine to recognize the user's emotions and generate the optimal delivery route and time."

[0772] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0773] Step 1:

[0774] The user provides voice input through an input / output device. The voice data is captured as input and converted into a format for transmission to the speech recognition system. This allows the terminal to collect voice information as digital data and prepare the voice data necessary for the next processing step.

[0775] Step 2:

[0776] The server uses speech recognition software to convert audio data into text. This process takes audio data as input and outputs text. The Google Cloud Speech-to-Text API is used to convert speech into text format based on acoustic features. This makes the content of the speech available as text.

[0777] Step 3:

[0778] The server analyzes the user's emotions from text data using an emotion analysis engine. Based on the text data as input, it performs data calculations to identify emotions and outputs emotion tags (e.g., anxiety, joy, frustration, etc.) as the analysis result. IBM Watson is a possible emotion analysis engine to be used. Specifically, it analyzes the frequency of occurrence of certain patterns and keywords in the text to infer the emotional state.

[0779] Step 4:

[0780] The server uses a generative model to optimize order management and delivery planning. It processes sentiment analysis results and text data as input to optimize delivery schedules and routes. In this process, it utilizes a generative AI model (e.g., OpenAI GPT) to apply processing based on sentiment. The optimized delivery plan is then output.

[0781] Step 5:

[0782] The server processes the optimization results into notification content and sends it to the terminal. Here, it considers the optimization results and sentiment data as input, transforming them into a notification message that is easy for the user to understand. The output is a concise yet detailed notification for the user, allowing them to grasp the situation in real time.

[0783] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0784] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0785] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0786] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0787] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0788] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0789] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0790] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0791] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0792] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0793] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0794] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0795] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0796] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0797] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0798] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0799] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0800] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0801] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0802] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0803] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0804] The following is further disclosed regarding the embodiments described above.

[0805] (Claim 1)

[0806] A means equipped with a terminal for receiving voice input,

[0807] A means of processing audio data on a server to convert it to text,

[0808] A method for optimizing a task using a generative model,

[0809] A means of notifying the device of optimized tasks,

[0810] A means of receiving and processing additional voice commands from the user.

[0811] A system that includes this.

[0812] (Claim 2)

[0813] The system according to claim 1, further comprising means for automating the scheduling of meetings.

[0814] (Claim 3)

[0815] The system according to claim 1, further comprising means for processing user instructions to change the priority of tasks.

[0816] "Example 1"

[0817] (Claim 1)

[0818] A means equipped with a device for acquiring voice input,

[0819] A means of processing audio information using computing resources to convert it into text information,

[0820] Means for optimizing the work using a generation algorithm,

[0821] Means for notifying the device of optimized work,

[0822] A means for receiving and processing additional voice commands from the user.

[0823] A system that includes this.

[0824] (Claim 2)

[0825] The system according to claim 1, further comprising means for automating the planning and setting of meetings.

[0826] (Claim 3)

[0827] The system according to claim 1, further comprising means for processing user instructions to change the priority of tasks.

[0828] "Application Example 1"

[0829] (Claim 1)

[0830] A means equipped with a device for receiving voice input,

[0831] A means for processing audio information using a data processing device for converting audio information into text information,

[0832] A means of optimizing work using a generative model,

[0833] Means for notifying the device of optimized work,

[0834] A means for receiving and processing additional voice instructions from the user,

[0835] A means of controlling work processes in the manufacturing industry using voice commands,

[0836] ...

[0837] A system that includes this.

[0838] (Claim 2)

[0839] The system according to claim 1, further comprising means for automating meeting scheduling.

[0840] (Claim 3)

[0841] The system according to claim 1, further comprising means for processing user instructions to change the priority of tasks.

[0842] "Example 2 of combining an emotion engine"

[0843] (Claim 1)

[0844] A means equipped with a device for receiving voice input,

[0845] A means of processing audio information using a communication device to convert it into text information,

[0846] A means of optimizing the work using a generation algorithm,

[0847] Means for notifying the device of optimized work,

[0848] A means for receiving and processing additional voice instructions from the user,

[0849] A means of performing emotion analysis and dynamically adjusting task priorities based on the analysis results,

[0850] A means of changing the interface expression according to the user's emotional state.

[0851] A system that includes this.

[0852] (Claim 2)

[0853] The system according to claim 1, further comprising means for automating the scheduling of meetings and adjusting them based on sentiment analysis results.

[0854] (Claim 3)

[0855] The system according to claim 1, further comprising means for processing user instructions to change the priority of tasks and adjusting them based on the emotional state.

[0856] "Application example 2 when combining with an emotional engine"

[0857] (Claim 1)

[0858] A means equipped with an input / output device for receiving audio input,

[0859] A means for processing audio data using an information processing device for converting it into text information,

[0860] A means of optimizing work using a generative model,

[0861] Means for notifying the device of optimized work,

[0862] A means including an emotion analysis engine for analyzing the emotions of users,

[0863] A means of adjusting work optimization based on emotions,

[0864] A means of adjusting notification content according to the user's emotions.

[0865] A system that includes this.

[0866] (Claim 2)

[0867] The system according to claim 1, further comprising means for automating order management.

[0868] (Claim 3)

[0869] The system according to claim 1, further comprising means for optimizing the delivery plan. [Explanation of Symbols]

[0870] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means equipped with a terminal for receiving voice input, A means of processing audio data on a server to convert it to text, A method for optimizing a task using a generative model, A means of notifying the device of optimized tasks, A means of receiving and processing additional voice commands from the user. A system that includes this.

2. The system according to claim 1, further comprising means for automating the scheduling of meetings.

3. The system according to claim 1, further comprising means for processing user instructions to change the priority of tasks.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A