system
The system enhances business automation by converting voice commands to text, analyzing and automating tasks, and providing emotional state-aware notifications, addressing inefficiencies in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Existing business automation systems lack flexibility and accuracy in processing voice commands, particularly in handling complex instructions and diverse tasks, leading to inefficiencies in productivity and business operations.
A system that captures voice instructions, converts them into text data, analyzes the content, and automates tasks such as schedule management and email creation, with real-time notification of completion status, utilizing speech recognition and natural language processing.
Improves work efficiency by enabling seamless task execution through voice commands, ensuring accurate and timely completion notifications, and adapting to user emotional states for personalized task management.
Smart Images

Figure 2026069146000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] [[ID=3�]]In a modern business environment, employees need to efficiently handle a variety of tasks. In particular, email correspondence and schedule management consume a great deal of time, which hinders productivity improvement. Also, in many cases, differences in IT literacy may affect the speed of business operations, which has an adverse impact on the overall business efficiency of the company. There is a need to solve the above problems and provide a mechanism that allows anyone to easily automate and streamline their work.
Means for Solving the Problems
[0005] This invention captures voice instructions from a user using a terminal means for acquiring voice data. The acquired voice data is converted into text data by a conversion means, and the content of the instructions is interpreted by an analysis means. Furthermore, based on the analyzed instructions, a business automation means works in conjunction to automatically perform various tasks such as schedule management, email creation, and data management, thereby improving work efficiency. In addition, the completion status of the tasks is notified to the user in a timely manner via a notification means, allowing the user to quickly check the results. As a result, it becomes possible to smoothly carry out tasks using only voice instructions, thereby achieving improved work efficiency.
[0006] "Voice data" refers to data that represents instructions spoken by a user in a digital format.
[0007] "Terminal means" refers to a device or apparatus that captures audio data and transmits it to a server or other processing means.
[0008] "Conversion means" refers to a technology or device that has the function of converting audio data acquired by a terminal means into text format.
[0009] "Analysis means" refers to a technology or device that has the function of understanding the content of text data obtained by the conversion means and interpreting work instructions.
[0010] "Business automation means" refers to technology or equipment that has the function of automatically performing business operations based on instructions interpreted by analysis means.
[0011] "Notification means" refers to technology or devices that communicate with users to inform them of the completion status of a task. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] The present invention relates to a system for automating tasks based on voice commands, and specific embodiments thereof are shown below.
[0034] First, the user gives work instructions using voice. For example, they might say a specific instruction such as, "Add a project meeting to the schedule for 10 AM tomorrow."
[0035] The device captures the user's voice and generates audio data. This audio data is then transmitted to a server via the network.
[0036] When the server receives audio data, it uses a speech recognition engine to convert it into text data. This text data is then interpreted by an analysis tool to identify the specific instructions required for the business automation tool.
[0037] Based on the analyzed instructions, the server executes automated tasks. This includes various actions in accordance with user instructions, such as schedule management, email creation, and data manipulation. For example, when automatically adding a meeting schedule using the Calendar API, the server accurately reflects the date, time, and content of the meeting.
[0038] Furthermore, when a task is completed, the server sends a completion notification to the terminal via a notification system. The user can then check the notification displayed on the terminal screen and confirm that the task has been performed correctly.
[0039] Thus, the system in this invention receives voice input from the user and provides a series of steps to perform the tasks desired by the user through automated functions, thereby improving work efficiency.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The user speaks work instructions into their terminal. These instructions should be clear and specific, so that the system can easily understand them.
[0043] Step 2:
[0044] The device uses a built-in or connected microphone to capture the user's voice in real time. This audio is temporarily stored as digital audio data.
[0045] Step 3:
[0046] The device sends the captured audio data to the server. This communication is typically conducted through a secure protocol to maintain data integrity and confidentiality.
[0047] Step 4:
[0048] The server inputs the received audio data into a speech recognition engine and converts it into text data. This conversion utilizes speech recognition algorithms and applies language models, among other things.
[0049] Step 5:
[0050] The server passes the converted text data to a natural language processing module, which analyzes the instructions. Here, the instructions are broken down into specific tasks, such as changing schedules or composing emails.
[0051] Step 6:
[0052] Based on the analysis results, the server selects the appropriate automation method and executes the necessary tasks. For example, it might access the calendar API to add a meeting at a specified date and time.
[0053] Step 7:
[0054] The server provides feedback to the user's terminal once the task is completed. This feedback includes a notification indicating that the task was successful.
[0055] Step 8:
[0056] Users can check notifications from their devices and confirm that the tasks they instructed have been carried out correctly. Users can also issue further voice instructions as needed.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] Conventional business automation systems lack the flexibility and accuracy to effectively process voice commands, making it difficult to properly interpret and execute user instructions. In particular, there is a need for a means to achieve rapid and accurate automation processes when handling complex instructions and diverse tasks.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes an information processing device for acquiring voice information, a conversion device for converting the acquired voice information into text information, and an analysis device for analyzing the text information obtained by the conversion device and interpreting the instruction content. This makes it possible to execute automated processes quickly and accurately for complex instructions and diverse tasks.
[0062] "Voice information" refers to instructions and information obtained from users via voice.
[0063] An "information processing device" is a device that has the function of capturing and digitizing audio information.
[0064] A "conversion device" is a device used to convert acquired audio information into text information.
[0065] "Textual information" refers to data that represents audio information as a string of characters.
[0066] An "analysis device" is a device that analyzes converted character information and interprets the specific instructions it contains.
[0067] An "automation device" is a device that automates tasks based on analyzed instructions.
[0068] A "notification device" is a device or function that informs the user of the work status.
[0069] "Other systems" refer to external systems that work in conjunction with each other to perform tasks.
[0070] "Schedule management" refers to the task of managing information and events related to schedules and calendars.
[0071] "Communication creation" refers to the task of creating and sending emails and messages.
[0072] "Information management" refers to the systematic collection, storage, and use of various types of data.
[0073] This invention is a system that automates tasks based on voice commands. This system is implemented through a series of processes involving the user, terminal, and server.
[0074] Users input work requests and instructions into the terminal as voice commands. For example, they might issue a voice command such as, "Please add a team meeting tomorrow at 3 PM." The terminal is equipped with an information processing device that converts the voice received from the user into digital voice data. This device typically includes a microphone and software for voice digitization.
[0075] The terminal sends the generated voice data to the server using a communication protocol. The server has a conversion device that uses speech recognition technology to convert the voice data into text information. Generally, speech recognition APIs or services (e.g., speech recognition cloud services) are used.
[0076] The server further uses natural language processing technology to analyze textual information and interpret user instructions. The analysis device performs this role. This analysis identifies specific automation tasks, such as schedule management and communication creation.
[0077] The server then works with external integrated systems (e.g., scheduling APIs and database management systems) to execute automated tasks based on the analyzed instructions. For example, when automatically adding a meeting to a schedule, it accurately reflects the date, time, and content of that meeting.
[0078] Once the task is complete, the server generates a notification and sends it to the terminal. The user receives the notification through the terminal and can confirm that the instructions were carried out correctly.
[0079] A concrete example is when a user gives a voice command saying, "Please cancel the meeting scheduled for 2 PM." The terminal captures this as voice data and sends it to the server. The server converts the voice data to text and cancels the meeting through analysis. After the task is completed, the terminal receives a notification and informs the user. In this way, an automated system based on voice commands streamlines business processes.
[0080] An example of a prompt message is, "Please tell me the procedure for a user to give work instructions by voice."
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The user provides work instructions via voice. The input consists of the user's voice instructions, specifically something like, "Please add a meeting at 3 PM." The user's spoken words are the central data for this step.
[0084] Step 2:
[0085] The terminal captures the user's voice using a microphone and converts it into digital audio data. The input for this step is the user's voice, and the output is digital audio data. Since audio is an analog signal, it undergoes digital processing and is stored as binary data. This process includes audio signal processing.
[0086] Step 3:
[0087] The terminal sends the generated digital audio data to the server. The input is the digital audio data within the terminal, and the output is the audio data sent to the server over the network. In this step, the data is transferred using an appropriate communication protocol.
[0088] Step 4:
[0089] The server converts the received audio data into text information using a speech recognition engine. The input is the audio data passed to the server, and the output is the instructions in text format. This conversion process uses a speech recognition algorithm, which is part of a generative AI model.
[0090] Step 5:
[0091] The server analyzes the converted text information and interprets the content of the instructions. The input is text information, and the output is the interpreted specific task or instruction. In this step, natural language processing is performed, and the parser plays a role in analyzing the content in detail.
[0092] Step 6:
[0093] The server executes the appropriate automated task based on the instructions. The input is the interpreted instructions, and the output is the result of the automated work. For example, this could include calling a scheduling API to add a meeting to the schedule.
[0094] Step 7:
[0095] The server generates a message to notify the terminal of the completion status of the task and sends it. The input is the completion status of the task, and the output is the notification message sent to the user. In this step, a notification device is used to generate a message based on an example prompt.
[0096] Step 8:
[0097] The user receives notifications on their device to confirm that the task was performed correctly. The input is the notification sent from the server, and the output is the user's understanding and confirmation action. The user can view the notification through their device's display.
[0098] (Application Example 1)
[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] In autonomous vehicles, it is necessary for the user to safely and efficiently control the vehicle's functions and respond appropriately to the situation while driving. However, operating the vehicle using hands or eyes while driving carries a risk of compromising safety. This invention aims to solve this problem by providing a system that allows the vehicle to be operated automatically by voice commands.
[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0102] In this invention, the server includes a device means for acquiring voice information, a conversion means for converting the acquired voice information into text information, an analysis means for analyzing the text information obtained by the conversion means and interpreting instructions, and means for supporting tasks related to the operation of the vehicle. This makes it possible for the user to safely and efficiently control the functions of the vehicle by voice while driving.
[0103] "Voice information" refers to data acquired based on the voice spoken by the user, and is information input as a voice signal.
[0104] "Device means" is a general term for equipment and devices for acquiring audio information, and in this invention, it specifically refers to devices that have the function of capturing audio data.
[0105] A "conversion means" is a means for performing the process of converting audio information into text information, and it generates text data using speech recognition technology.
[0106] "Textual information" refers to data in text format obtained as a result of converting audio information using a conversion means, and is the subject of interpretation by an analysis means.
[0107] "Analysis means" refers to a process or technology that interprets user instructions based on textual information obtained by the conversion means and derives specific work instructions.
[0108] "Instructions" refer to specific requests or commands given by the user as voice information, and are information that is interpreted by the analysis tool.
[0109] "Means to support tasks related to vehicle operation" refer to means for controlling and adjusting functions within the vehicle based on voice commands, thereby improving convenience and safety in autonomous vehicles.
[0110] In a system implementing the present invention, a device for processing voice information and a series of related software operate in cooperation. When a user inputs specific instructions by voice, that voice information is acquired by the terminal's device. Examples of such devices include smartphones and in-vehicle information systems.
[0111] The acquired audio information is sent from the terminal to the server via a speech recognition API (e.g., Google® Cloud Speech-to-Text API). On the server, the audio information is converted into text information using a conversion means. This text information is interpreted by an analysis means to identify the user's instructions. A natural language processing model (e.g., a generative AI model) is used as the analysis means.
[0112] Next, based on the analyzed instructions, the automated system performs operations to control the vehicle's functions. In-vehicle APIs (e.g., CarPlay® and ANDROID® Auto) are used to support tasks related to vehicle operation. Specifically, this includes adjusting the air conditioning temperature and providing automatic route guidance via navigation.
[0113] Once the task is complete, the server sends a completion notification to the terminal using a notification method, and the user can check the result on the terminal's display screen. For example, if a user says, "Tell me about nearby restaurants," the voice is converted to text, and based on the analysis, directions to the nearest restaurant are displayed on the vehicle's navigation system.
[0114] An example of a prompt for a generative AI model could be: "Please provide brief steps to guide the user to the nearest destination based on voice commands."
[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0116] Step 1:
[0117] The device acquires the user's voice commands. Here, a smartphone or in-vehicle system captures voice data using a microphone. The input is the user's voice itself, and the output is voice data converted into a digital format.
[0118] Step 2:
[0119] The terminal sends the acquired audio data to the server. A communication module is used to deliver the audio data to the server over the network. The input here is digitized audio data, and the output is the audio data as it has been transferred to the server's designated data receiving endpoint.
[0120] Step 3:
[0121] The server receives audio data and uses a conversion mechanism to convert it into text data via a speech recognition API. The speech recognition model analyzes the audio signal and converts it into text. The input is the audio data received by the server, and the output is text data in text format.
[0122] Step 4:
[0123] The server uses a generative AI model to analyze text data and identify instructions. A natural language processing model takes text data as input and analyzes it to identify the user's instructions. The input is text data, and the output is specific instructions.
[0124] Step 5:
[0125] The server automates vehicle operation using means to assist with vehicle operation based on identified instructions. It calls in-vehicle APIs to control specific vehicle functions such as adjusting the air conditioning temperature and setting the navigation system. The input is the instructions obtained through analysis, and the output is the automated functional state within the vehicle.
[0126] Step 6:
[0127] After completing the task, the server notifies the terminal of the result using a notification method. A completion message is sent to the terminal via the network. The input is the status information indicating task completion, and the output is the notification message displayed on the terminal.
[0128] Step 7:
[0129] The user checks the notification on their device and confirms that the task has been completed correctly. The user can see the completion notification message on the device's display. The input is the notification message received from the server, and the output is the completion status information that the user confirms.
[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0131] This invention is a system that automates and further improves the efficiency of tasks using voice commands and emotion recognition.
[0132] The user inputs work-related instructions by voice into the terminal. During this process, information about the user's emotional state is also captured as voice data.
[0133] The device sends voice data to the server. The server converts the received voice data into text data using a speech recognition engine and simultaneously analyzes the user's emotional state using an emotion engine. The emotion engine determines whether the user is experiencing emotions such as "anger," "joy," or "sadness" based on features such as the tone and rhythm of the voice.
[0134] The analysis method analyzes the textualized instructions and also takes into account emotional information obtained from the emotion engine. Based on this emotional information, for example, if it is determined that the user is experiencing stress, the system determines which tasks should be prioritized.
[0135] The business automation system executes tasks by coordinating with necessary external systems based on analysis results. For example, if it determines that a user is in a "high-stress state," it can prioritize scheduling tasks and suggest rescheduling meetings.
[0136] Once a task is completed, a notification system will inform the device of the result. It's also possible to adjust the notification's wording based on emotional information. For example, if the user is feeling "joy," the notification will use positive language.
[0137] Thus, the system of the present invention provides a more user-friendly interface and enables smooth business operations by comprehensively utilizing user voice commands and emotional information.
[0138] The following describes the processing flow.
[0139] Step 1:
[0140] The user speaks in a natural voice along with work instructions. This includes tone, speed, and emphasis, and also conveys emotional information.
[0141] Step 2:
[0142] The device captures the user's voice through the microphone. The captured voice is temporarily stored as digital audio data containing the user's instructions and emotional information.
[0143] Step 3:
[0144] The device sends the captured audio data to the server. This communication is conducted using a secure protocol.
[0145] Step 4:
[0146] The server converts the audio data into text data using a speech recognition engine. During this process, the content of the speech is faithfully reproduced as text.
[0147] Step 5:
[0148] The server analyzes the converted text data using a natural language processing module. Here, the user's instructions are clearly interpreted, and specific business tasks are extracted.
[0149] Step 6:
[0150] Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone and volume of the voice to determine emotions such as "joy," "anger," and "sadness."
[0151] Step 7:
[0152] The server selects the most appropriate automation method for tasks based on the analyzed instructions and emotional information. For example, if it determines that a user is experiencing stress, it prioritizes tasks that require assistance.
[0153] Step 8:
[0154] The server interacts with other systems as needed to perform specific tasks. Data exchange with external systems allows for optimization of schedules and reallocation of resources, for example.
[0155] Step 9:
[0156] After completing a task, the server uses a notification system to report the results to the terminal. The notification reflects emotional information and is presented in a format and content appropriate to the user.
[0157] Step 10:
[0158] Users can check notifications from their devices and evaluate whether the assigned tasks were performed as expected. Additional instructions can be given via voice if necessary.
[0159] (Example 2)
[0160] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0161] Existing business automation systems often process tasks uniformly without considering the user's emotional state, which can increase user stress and burden. Furthermore, notifications may not be sensitive to the user's psychological state, hindering efficient communication. This can negatively impact the user experience.
[0162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0163] In this invention, the server includes means for acquiring voice data and capturing information including the user's emotional state; means for converting the acquired voice data into text data using a speech recognition engine; and means for analyzing the user's emotional state from the text data obtained by the conversion means and the tone and rhythm of the voice. This enables user-friendly system operation by adjusting the content of automated tasks and notifications according to the user's emotional state.
[0164] "Voice data" refers to information recorded in digital or analog format as audio signals, and includes user voice commands and emotional expressions.
[0165] "Terminal means" refers to a device or apparatus that acquires voice data from a user and transmits that information to a server.
[0166] "Conversion means" refers to software or algorithms used to convert acquired audio data into text data.
[0167] A "speech recognition engine" is a technology or program that extracts linguistic information from speech data and converts its content into text format.
[0168] "Analysis means" refers to a device or software that has the function of analyzing the content of instructions converted from voice data into text and the user's emotional state to determine the content and priority of tasks.
[0169] An "emotion engine" is a component of a system that determines and analyzes a user's emotional state based on characteristics such as tone and rhythm of their voice.
[0170] "Automation methods" are those that efficiently process tasks based on analyzed instructions and emotional information, and have the functionality to link with external systems as needed.
[0171] A "notification means" is a method or device for informing a user of the results of a task, and it is possible to adjust the content and expression of the notification according to the user's emotional state.
[0172] This invention is a system that automates and streamlines tasks through voice commands and emotion recognition. The user inputs task-related instructions via voice into a terminal. The terminal transmits the voice data to a server. This voice data also includes the user's emotional information.
[0173] The server uses a speech recognition engine to convert received audio data into text data. It is possible to customize and use commonly available APIs as the speech recognition engine. Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the tone and rhythmic characteristics within the audio data. The emotion engine uses music analysis technology and machine learning-based emotion analysis algorithms to determine whether the user is experiencing emotions such as "joy," "anger," or "sadness."
[0174] Based on analyzed instructions and emotional information, the server automates tasks. Depending on the type of task, it integrates with schedule management and data management systems to perform the task. For example, if a user requests that a meeting be rescheduled and also indicates stress, the server will prioritize schedule management and suggest rescheduling the meeting.
[0175] Once a task is completed, the server notifies the terminal of the result via a notification system. The notification's wording can be adjusted based on emotional information. For example, if the user is feeling "joy," the notification will be expressed with a positive message.
[0176] For example, if a user gives a voice command saying, "Remind me of the project meeting tomorrow at 10 AM," the server will acknowledge the command, add the event to the calendar if necessary, and send a notification at the appropriate time.
[0177] An example of a prompt sentence to input into the generating AI model is: "Please describe the flow of a system that efficiently automates tasks instructed by the user via voice and prioritizes them according to their emotional state."
[0178] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0179] Step 1:
[0180] The user inputs voice instructions related to their work into the terminal. This voice input includes the user's spoken content, tone, and rhythm. The terminal captures this voice data using a microphone. The input data is raw audio waveform data.
[0181] Step 2:
[0182] The terminal converts the captured audio data into a digital format and sends it to the server over the network. At this point, the input is digital audio data, and the output is the data packets sent to the server.
[0183] Step 3:
[0184] The server inputs the received digital audio data into the speech recognition engine. The engine analyzes the audio signal and converts it into corresponding text data. Specifically, it uses a feature extraction algorithm to identify phonemes and generates words based on them. Here, the input is audio signal data, and the output is text data.
[0185] Step 4:
[0186] The server simultaneously uses an emotion engine to analyze the tone and rhythm of the voice data to identify the emotional state. A machine learning model is used for the analysis, calculating emotions such as "joy" or "anger" from the voice patterns. The input is voice characteristic data, and the output is an emotional state (e.g., joy, anger, sadness).
[0187] Step 5:
[0188] The server analyzes the instructions based on the converted text data and analyzed emotional state. The analysis extracts keywords and phrases from the text and associates them with the work content. The input is text data and emotional information, and the output is a profile of the work instructions.
[0189] Step 6:
[0190] The server integrates with appropriate external systems and automates necessary tasks based on the work order profile and sentiment information. Specifically, it accesses a scheduling management system to coordinate meeting schedules. The input is the work order profile, and the output is the result of the completed tasks.
[0191] Step 7:
[0192] Once a task is completed, the server notifies the terminal of the result via a notification system. This notification is crafted with sentiment information in mind, using positive or mitigating language. The input is the result of the task execution, and the output is the adjusted notification message.
[0193] (Application Example 2)
[0194] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0195] In brick-and-mortar stores where customers interact with staff, employees are required to appropriately recognize each customer's emotional state and respond quickly and appropriately based on that understanding. However, manual emotional recognition and response make it difficult to respond quickly to the diverse emotions and situations of customers. Furthermore, providing personalized service tailored to each customer's emotions is necessary to improve customer satisfaction.
[0196] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0197] In this invention, the server includes a device for acquiring audio signals, a processing means for converting the acquired audio signals into text information, an analysis means for analyzing the text information obtained by the processing means and interpreting the information content, an automation means for automating tasks based on the analyzed information content and emotional information, and a notification means for notifying the user of the completion status of the tasks. This enables efficient and accurate responses that take customer emotions into consideration, as well as the provision of personalized services.
[0198] An "audio signal" refers to a signal that represents audio information in digital or analog format.
[0199] "Device" refers to hardware that has a mechanical or electronic mechanism for achieving a specific purpose.
[0200] A "processing means" is a means for receiving input data and converting it into a desired format or information.
[0201] "Textual information" refers to information composed of characters in natural language that possesses meaning.
[0202] "Analysis means" refers to the means used to analyze input data and interpret specific information or commands.
[0203] "Information content" refers to the actual content of knowledge or instructions transmitted through data or communication.
[0204] "Emotional information" refers to information about emotional states detected from the characteristics of speech and text.
[0205] "Automation methods" are technical means to perform specific tasks or processes without human intervention.
[0206] "Notification means" refers to methods for informing users or systems of status or results.
[0207] "User" refers to a person who operates a system or device, or an end-user.
[0208] To implement this invention, a system utilizing voice commands and emotion recognition will be constructed. This system will primarily aim to improve the efficiency and personalization of customer service in retail stores.
[0209] The server receives audio signals from voice acquisition devices such as smartphones. The hardware used is expected to include smart devices with microphones and network-connected communication devices. The audio signals are converted into text information by the Google Cloud Speech-to-Text API. Then, IBM Watson® Tone Analyzer analyzes sentiment information based on the text information extracted from the audio.
[0210] If the user is a store employee, the analyzed information is used to determine the appropriate response to a customer. For example, if the emotional information indicates that a customer is making a complaint, the system generates an alert to prioritize that customer. This alert is sent to the employee's mobile device to prompt a quick response.
[0211] Furthermore, as a concrete example of improving the efficiency of customer service based on emotion recognition, the system supports prompt phrases such as "Customer feedback: My order is wrong, what are you going to do about it? (Emotion: Anger)." This prompt allows staff to immediately consider countermeasures and provide appropriate service.
[0212] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0213] Step 1:
[0214] The user (store staff) uses the voice input function of their smartphone to capture the customer's voice on the device. This input is acquired as a real-time audio signal.
[0215] Step 2:
[0216] The device sends the acquired audio signal to the Google Cloud Speech-to-Text API and receives the resulting text information. Here, the audio signal is converted into text data through language analysis.
[0217] Step 3:
[0218] The server sends text information to IBM Watson's Tone Analyzer, which analyzes and receives customer sentiment information. In this step, the tone and atmosphere from the text information are analyzed to identify emotional states such as "anger" or "joy."
[0219] Step 4:
[0220] The server analyzes a combination of text and sentiment information to determine the priority of customer responses and the necessary actions. This analysis determines that a customer needs to be prioritized, especially if they are experiencing significant stress.
[0221] Step 5:
[0222] The server sends a notification to the terminal along with the analysis results. The user (store staff) receives specific response strategies and suggestions based on the emotional information, and responds to the customer accordingly. At this stage, the system presents examples of how to provide quick and appropriate customer service.
[0223] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0224] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0225] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0226] [Second Embodiment]
[0227] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0228] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0229] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0230] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0231] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0232] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0233] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0234] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0235] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0236] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0237] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0238] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0239] The present invention relates to a system for automating tasks based on voice commands, and specific embodiments thereof are shown below.
[0240] First, the user gives work instructions using voice. For example, they might say a specific instruction such as, "Add a project meeting to the schedule for 10 AM tomorrow."
[0241] The device captures the user's voice and generates audio data. This audio data is then transmitted to a server via the network.
[0242] When the server receives audio data, it uses a speech recognition engine to convert it into text data. This text data is then interpreted by an analysis tool to identify the specific instructions required for the business automation tool.
[0243] Based on the analyzed instructions, the server executes automated tasks. This includes various actions in accordance with user instructions, such as schedule management, email creation, and data manipulation. For example, when automatically adding a meeting schedule using the Calendar API, the server accurately reflects the date, time, and content of the meeting.
[0244] Furthermore, when a task is completed, the server sends a completion notification to the terminal via a notification system. The user can then check the notification displayed on the terminal screen and confirm that the task has been performed correctly.
[0245] Thus, the system in this invention receives voice input from the user and provides a series of steps to perform the tasks desired by the user through automated functions, thereby improving work efficiency.
[0246] The following describes the processing flow.
[0247] Step 1:
[0248] The user speaks work instructions into their terminal. These instructions should be clear and specific, so that the system can easily understand them.
[0249] Step 2:
[0250] The device uses a built-in or connected microphone to capture the user's voice in real time. This audio is temporarily stored as digital audio data.
[0251] Step 3:
[0252] The device sends the captured audio data to the server. This communication is typically conducted through a secure protocol to maintain data integrity and confidentiality.
[0253] Step 4:
[0254] The server inputs the received audio data into a speech recognition engine and converts it into text data. This conversion utilizes speech recognition algorithms and applies language models, among other things.
[0255] Step 5:
[0256] The server passes the converted text data to a natural language processing module, which analyzes the instructions. Here, the instructions are broken down into specific tasks, such as changing schedules or composing emails.
[0257] Step 6:
[0258] Based on the analysis results, the server selects the appropriate automation method and executes the necessary tasks. For example, it might access the calendar API to add a meeting at a specified date and time.
[0259] Step 7:
[0260] The server provides feedback to the user's terminal once the task is completed. This feedback includes a notification indicating that the task was successful.
[0261] Step 8:
[0262] Users can check notifications from their devices and confirm that the tasks they instructed have been carried out correctly. Users can also issue further voice instructions as needed.
[0263] (Example 1)
[0264] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0265] Conventional business automation systems lack the flexibility and accuracy to effectively process voice commands, making it difficult to properly interpret and execute user instructions. In particular, there is a need for a means to achieve rapid and accurate automation processes when handling complex instructions and diverse tasks.
[0266] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0267] In this invention, the server includes an information processing device for acquiring voice information, a conversion device for converting the acquired voice information into text information, and an analysis device for analyzing the text information obtained by the conversion device and interpreting the instruction content. This makes it possible to execute automated processes quickly and accurately for complex instructions and diverse tasks.
[0268] "Voice information" refers to instructions and information obtained from users via voice.
[0269] An "information processing device" is a device that has the function of capturing and digitizing audio information.
[0270] A "conversion device" is a device used to convert acquired audio information into text information.
[0271] "Textual information" refers to data that represents audio information as a string of characters.
[0272] An "analysis device" is a device that analyzes converted character information and interprets the specific instructions it contains.
[0273] An "automation device" is a device that automates tasks based on analyzed instructions.
[0274] A "notification device" is a device or function that informs the user of the work status.
[0275] "Other systems" refer to external systems that work in conjunction with each other to perform tasks.
[0276] "Schedule management" refers to the task of managing information and events related to schedules and calendars.
[0277] "Communication creation" refers to the task of creating and sending emails and messages.
[0278] "Information management" refers to the systematic collection, storage, and use of various types of data.
[0279] This invention is a system that automates tasks based on voice commands. This system is implemented through a series of processes involving the user, terminal, and server.
[0280] Users input work requests and instructions into the terminal as voice commands. For example, they might issue a voice command such as, "Please add a team meeting tomorrow at 3 PM." The terminal is equipped with an information processing device that converts the voice received from the user into digital voice data. This device typically includes a microphone and software for voice digitization.
[0281] The terminal sends the generated voice data to the server using a communication protocol. The server has a conversion device that uses speech recognition technology to convert the voice data into text information. Generally, speech recognition APIs or services (e.g., speech recognition cloud services) are used.
[0282] The server further uses natural language processing technology to analyze textual information and interpret user instructions. The analysis device performs this role. This analysis identifies specific automation tasks, such as schedule management and communication creation.
[0283] After that, the server executes automated tasks based on the analyzed instructions by collaborating with external integration systems (e.g., schedule management APIs and database management systems). For example, when automatically adding a meeting to the schedule, it accurately reflects the date, time, and content.
[0284] When the task is completed, the server generates a notification and sends it to the terminal. The user can receive the notification through the terminal and confirm that the instructed content has been accurately executed.
[0285] As a specific example, it may be the case where the user instructs verbally "Please cancel the meeting scheduled for 2 PM". The terminal captures this as voice data and sends it to the server. The server converts the voice data into text and cancels the meeting through analysis. After the task is completed, the terminal receives the notification and informs the user. In this way, the automated system based on voice instructions improves the business process.
[0286] An example of a prompt sentence is "Please teach me the procedure when the user gives a business instruction verbally."
[0287] The flow of the specific process in Example 1 will be described using FIG. 11.
[0288] Step 1:
[0289] The user gives a business instruction verbally. The input is the user's verbal instruction, specifically, something like "Please add a meeting at 3 PM". What the user says is the central data for this step.
[0290] Step 2:
[0291] The terminal captures the user's voice using a microphone and converts it into digital audio data. The input for this step is the user's voice, and the output is digital audio data. Since audio is an analog signal, it undergoes digital processing and is stored as binary data. This process includes audio signal processing.
[0292] Step 3:
[0293] The terminal sends the generated digital audio data to the server. The input is the digital audio data within the terminal, and the output is the audio data sent to the server over the network. In this step, the data is transferred using an appropriate communication protocol.
[0294] Step 4:
[0295] The server converts the received audio data into text information using a speech recognition engine. The input is the audio data passed to the server, and the output is the instructions in text format. This conversion process uses a speech recognition algorithm, which is part of a generative AI model.
[0296] Step 5:
[0297] The server analyzes the converted text information and interprets the content of the instructions. The input is text information, and the output is the interpreted specific task or instruction. In this step, natural language processing is performed, and the parser plays a role in analyzing the content in detail.
[0298] Step 6:
[0299] The server executes the appropriate automated task based on the instructions. The input is the interpreted instructions, and the output is the result of the automated work. For example, this could include calling a scheduling API to add a meeting to the schedule.
[0300] Step 7:
[0301] The server generates a message for notifying the completion status of a task and sends it to the terminal. The input is the completion status of the work, and the output is the notification message sent to the user. In this step, a notification device is used to generate a message based on an example of a prompt sentence.
[0302] Step 8:
[0303] The user receives a notification from the terminal and confirms that the task has been correctly performed. The input is the notification sent from the server, and the output is the user's understanding and confirmation action. The user can confirm the notification through the display of the terminal.
[0304] (Application Example 1)
[0305] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0306] In an autonomous vehicle, it is required that the user safely and efficiently control the functions of the vehicle during driving and take appropriate actions according to the situation. However, operating the vehicle using hands or line of sight during driving poses a risk of compromising safety. The present invention aims to solve this problem by providing a system that can automatically operate the vehicle by voice instructions.
[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0308] In this invention, the server includes a device means for acquiring voice information, a conversion means for converting the acquired voice information into text information, an analysis means for analyzing the text information obtained by the conversion means to interpret the instruction items, and a means for assisting the work related to the operation of the vehicle. Thereby, it becomes possible for the user to safely and efficiently control the functions of the vehicle by voice during driving.
[0309] "Voice information" refers to data acquired based on the voice spoken by the user, and is information input as a voice signal.
[0310] "Device means" is a general term for equipment and devices for acquiring audio information, and in this invention, it specifically refers to devices that have the function of capturing audio data.
[0311] A "conversion means" is a means for performing the process of converting audio information into text information, and it generates text data using speech recognition technology.
[0312] "Textual information" refers to data in text format obtained as a result of converting audio information using a conversion means, and is the subject of interpretation by an analysis means.
[0313] "Analysis means" refers to a process or technology that interprets user instructions based on textual information obtained by the conversion means and derives specific work instructions.
[0314] "Instructions" refer to specific requests or commands given by the user as voice information, and are information that is interpreted by the analysis tool.
[0315] "Means to support tasks related to vehicle operation" refer to means for controlling and adjusting functions within the vehicle based on voice commands, thereby improving convenience and safety in autonomous vehicles.
[0316] In a system implementing the present invention, a device for processing voice information and a series of related software operate in cooperation. When a user inputs specific instructions by voice, that voice information is acquired by the terminal's device. Examples of such devices include smartphones and in-vehicle information systems.
[0317] The acquired audio information is sent from the terminal to the server via a speech recognition API (e.g., Google Cloud Speech-to-Text API). On the server, the audio information is converted into text information using a conversion means. This text information is interpreted by an analysis means to identify the user's instructions. A natural language processing model (e.g., a generative AI model) is used as the analysis means.
[0318] Next, based on the analyzed instructions, the automated system performs operations to control the vehicle's functions. In-vehicle APIs (e.g., CarPlay and Android Auto) are used to support tasks related to vehicle operation. Specifically, this includes adjusting the air conditioning temperature and providing automatic route guidance via navigation.
[0319] Once the task is complete, the server sends a completion notification to the terminal using a notification method, and the user can check the result on the terminal's display screen. For example, if a user says, "Tell me about nearby restaurants," the voice is converted to text, and based on the analysis, directions to the nearest restaurant are displayed on the vehicle's navigation system.
[0320] An example of a prompt for a generative AI model could be: "Please provide brief steps to guide the user to the nearest destination based on voice commands."
[0321] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0322] Step 1:
[0323] The device acquires the user's voice commands. Here, a smartphone or in-vehicle system captures voice data using a microphone. The input is the user's voice itself, and the output is voice data converted into a digital format.
[0324] Step 2:
[0325] The terminal sends the acquired audio data to the server. A communication module is used to deliver the audio data to the server over the network. The input here is digitized audio data, and the output is the audio data as it has been transferred to the server's designated data receiving endpoint.
[0326] Step 3:
[0327] The server receives audio data and uses a conversion mechanism to convert it into text data via a speech recognition API. The speech recognition model analyzes the audio signal and converts it into text. The input is the audio data received by the server, and the output is text data in text format.
[0328] Step 4:
[0329] The server uses a generative AI model to analyze text data and identify instructions. A natural language processing model takes text data as input and analyzes it to identify the user's instructions. The input is text data, and the output is specific instructions.
[0330] Step 5:
[0331] The server automates vehicle operation using means to assist with vehicle operation based on identified instructions. It calls in-vehicle APIs to control specific vehicle functions such as adjusting the air conditioning temperature and setting the navigation system. The input is the instructions obtained through analysis, and the output is the automated functional state within the vehicle.
[0332] Step 6:
[0333] After completing the task, the server notifies the terminal of the result using a notification method. A completion message is sent to the terminal via the network. The input is the status information indicating task completion, and the output is the notification message displayed on the terminal.
[0334] Step 7:
[0335] The user checks the notification on their device and confirms that the task has been completed correctly. The user can see the completion notification message on the device's display. The input is the notification message received from the server, and the output is the completion status information that the user confirms.
[0336] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0337] This invention is a system that automates and further improves the efficiency of tasks using voice commands and emotion recognition.
[0338] The user inputs work-related instructions by voice into the terminal. During this process, information about the user's emotional state is also captured as voice data.
[0339] The device sends voice data to the server. The server converts the received voice data into text data using a speech recognition engine and simultaneously analyzes the user's emotional state using an emotion engine. The emotion engine determines whether the user is experiencing emotions such as "anger," "joy," or "sadness" based on features such as the tone and rhythm of the voice.
[0340] The analysis method analyzes the textualized instructions and also takes into account emotional information obtained from the emotion engine. Based on this emotional information, for example, if it is determined that the user is experiencing stress, the system determines which tasks should be prioritized.
[0341] The business automation system executes tasks by coordinating with necessary external systems based on analysis results. For example, if it determines that a user is in a "high-stress state," it can prioritize scheduling tasks and suggest rescheduling meetings.
[0342] Once a task is completed, a notification system will inform the device of the result. It's also possible to adjust the notification's wording based on emotional information. For example, if the user is feeling "joy," the notification will use positive language.
[0343] Thus, the system of the present invention provides a more user-friendly interface and enables smooth business operations by comprehensively utilizing user voice commands and emotional information.
[0344] The following describes the processing flow.
[0345] Step 1:
[0346] The user speaks in a natural voice along with work instructions. This includes tone, speed, and emphasis, and also conveys emotional information.
[0347] Step 2:
[0348] The device captures the user's voice through the microphone. The captured voice is temporarily stored as digital audio data containing the user's instructions and emotional information.
[0349] Step 3:
[0350] The device sends the captured audio data to the server. This communication is conducted using a secure protocol.
[0351] Step 4:
[0352] The server converts the audio data into text data using a speech recognition engine. During this process, the content of the speech is faithfully reproduced as text.
[0353] Step 5:
[0354] The server analyzes the converted text data using a natural language processing module. Here, the user's instructions are clearly interpreted, and specific business tasks are extracted.
[0355] Step 6:
[0356] Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone and volume of the voice to determine emotions such as "joy," "anger," and "sadness."
[0357] Step 7:
[0358] The server selects the most appropriate automation method for tasks based on the analyzed instructions and emotional information. For example, if it determines that a user is experiencing stress, it prioritizes tasks that require assistance.
[0359] Step 8:
[0360] The server interacts with other systems as needed to perform specific tasks. Data exchange with external systems allows for optimization of schedules and reallocation of resources, for example.
[0361] Step 9:
[0362] After completing a task, the server uses a notification system to report the results to the terminal. The notification reflects emotional information and is presented in a format and content appropriate to the user.
[0363] Step 10:
[0364] Users can check notifications from their devices and evaluate whether the assigned tasks were performed as expected. Additional instructions can be given via voice if necessary.
[0365] (Example 2)
[0366] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0367] Existing business automation systems often process tasks uniformly without considering the user's emotional state, which can increase user stress and burden. Furthermore, notifications may not be sensitive to the user's psychological state, hindering efficient communication. This can negatively impact the user experience.
[0368] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0369] In this invention, the server includes means for acquiring voice data and capturing information including the user's emotional state; means for converting the acquired voice data into text data using a speech recognition engine; and means for analyzing the user's emotional state from the text data obtained by the conversion means and the tone and rhythm of the voice. This enables user-friendly system operation by adjusting the content of automated tasks and notifications according to the user's emotional state.
[0370] "Voice data" refers to information recorded in digital or analog format as audio signals, and includes user voice commands and emotional expressions.
[0371] "Terminal means" refers to a device or apparatus that acquires voice data from a user and transmits that information to a server.
[0372] "Conversion means" refers to software or algorithms used to convert acquired audio data into text data.
[0373] A "speech recognition engine" is a technology or program that extracts linguistic information from speech data and converts its content into text format.
[0374] "Analysis means" refers to a device or software that has the function of analyzing the content of instructions converted from voice data into text and the user's emotional state to determine the content and priority of tasks.
[0375] An "emotion engine" is a component of a system that determines and analyzes a user's emotional state based on characteristics such as tone and rhythm of their voice.
[0376] "Automation methods" are those that efficiently process tasks based on analyzed instructions and emotional information, and have the functionality to link with external systems as needed.
[0377] A "notification means" is a method or device for informing a user of the results of a task, and it is possible to adjust the content and expression of the notification according to the user's emotional state.
[0378] This invention is a system that automates and streamlines tasks through voice commands and emotion recognition. The user inputs task-related instructions via voice into a terminal. The terminal transmits the voice data to a server. This voice data also includes the user's emotional information.
[0379] The server uses a speech recognition engine to convert received audio data into text data. It is possible to customize and use commonly available APIs as the speech recognition engine. Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the tone and rhythmic characteristics within the audio data. The emotion engine uses music analysis technology and machine learning-based emotion analysis algorithms to determine whether the user is experiencing emotions such as "joy," "anger," or "sadness."
[0380] Based on analyzed instructions and emotional information, the server automates tasks. Depending on the type of task, it integrates with schedule management and data management systems to perform the task. For example, if a user requests that a meeting be rescheduled and also indicates stress, the server will prioritize schedule management and suggest rescheduling the meeting.
[0381] Once a task is completed, the server notifies the terminal of the result via a notification system. The notification's wording can be adjusted based on emotional information. For example, if the user is feeling "joy," the notification will be expressed with a positive message.
[0382] For example, if a user gives a voice command saying, "Remind me of the project meeting tomorrow at 10 AM," the server will acknowledge the command, add the event to the calendar if necessary, and send a notification at the appropriate time.
[0383] An example of a prompt sentence to input into the generating AI model is: "Please describe the flow of a system that efficiently automates tasks instructed by the user via voice and prioritizes them according to their emotional state."
[0384] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0385] Step 1:
[0386] The user inputs voice instructions related to their work into the terminal. This voice input includes the user's spoken content, tone, and rhythm. The terminal captures this voice data using a microphone. The input data is raw audio waveform data.
[0387] Step 2:
[0388] The terminal converts the captured audio data into a digital format and sends it to the server over the network. At this point, the input is digital audio data, and the output is the data packets sent to the server.
[0389] Step 3:
[0390] The server inputs the received digital audio data into the speech recognition engine. The engine analyzes the audio signal and converts it into corresponding text data. Specifically, it uses a feature extraction algorithm to identify phonemes and generates words based on them. Here, the input is audio signal data, and the output is text data.
[0391] Step 4:
[0392] The server simultaneously uses an emotion engine to analyze the tone and rhythm of the voice data to identify the emotional state. A machine learning model is used for the analysis, calculating emotions such as "joy" or "anger" from the voice patterns. The input is voice characteristic data, and the output is an emotional state (e.g., joy, anger, sadness).
[0393] Step 5:
[0394] The server analyzes the instructions based on the converted text data and analyzed emotional state. The analysis extracts keywords and phrases from the text and associates them with the work content. The input is text data and emotional information, and the output is a profile of the work instructions.
[0395] Step 6:
[0396] The server integrates with appropriate external systems and automates necessary tasks based on the work order profile and sentiment information. Specifically, it accesses a scheduling management system to coordinate meeting schedules. The input is the work order profile, and the output is the result of the completed tasks.
[0397] Step 7:
[0398] Once a task is completed, the server notifies the terminal of the result via a notification system. This notification is crafted with sentiment information in mind, using positive or mitigating language. The input is the result of the task execution, and the output is the adjusted notification message.
[0399] (Application Example 2)
[0400] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0401] In brick-and-mortar stores where customers interact with staff, employees are required to appropriately recognize each customer's emotional state and respond quickly and appropriately based on that understanding. However, manual emotional recognition and response make it difficult to respond quickly to the diverse emotions and situations of customers. Furthermore, providing personalized service tailored to each customer's emotions is necessary to improve customer satisfaction.
[0402] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0403] In this invention, the server includes a device for acquiring audio signals, a processing means for converting the acquired audio signals into text information, an analysis means for analyzing the text information obtained by the processing means and interpreting the information content, an automation means for automating tasks based on the analyzed information content and emotional information, and a notification means for notifying the user of the completion status of the tasks. This enables efficient and accurate responses that take customer emotions into consideration, as well as the provision of personalized services.
[0404] An "audio signal" refers to a signal that represents audio information in digital or analog format.
[0405] "Device" refers to hardware that has a mechanical or electronic mechanism for achieving a specific purpose.
[0406] A "processing means" is a means for receiving input data and converting it into a desired format or information.
[0407] "Textual information" refers to information composed of characters in natural language that possesses meaning.
[0408] "Analysis means" refers to the means used to analyze input data and interpret specific information or commands.
[0409] "Information content" refers to the actual content of knowledge or instructions transmitted through data or communication.
[0410] "Emotional information" refers to information about emotional states detected from the characteristics of speech and text.
[0411] "Automation methods" are technical means to perform specific tasks or processes without human intervention.
[0412] "Notification means" refers to methods for informing users or systems of status or results.
[0413] "User" refers to a person who operates a system or device, or an end-user.
[0414] To implement this invention, a system utilizing voice commands and emotion recognition will be constructed. This system will primarily aim to improve the efficiency and personalization of customer service in retail stores.
[0415] The server receives audio signals from voice acquisition devices such as smartphones. The hardware used is expected to include smart devices with microphones and network-connected communication devices. The audio signals are converted into text information by the Google Cloud Speech-to-Text API. Then, IBM Watson's Tone Analyzer analyzes sentiment information based on the text information extracted from the audio.
[0416] If the user is a store employee, the analyzed information is used to determine the appropriate response to a customer. For example, if the emotional information indicates that a customer is making a complaint, the system generates an alert to prioritize that customer. This alert is sent to the employee's mobile device to prompt a quick response.
[0417] Furthermore, as a concrete example of improving the efficiency of customer service based on emotion recognition, the system supports prompt phrases such as "Customer feedback: My order is wrong, what are you going to do about it? (Emotion: Anger)." This prompt allows staff to immediately consider countermeasures and provide appropriate service.
[0418] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0419] Step 1:
[0420] The user (store staff) uses the voice input function of their smartphone to capture the customer's voice on the device. This input is acquired as a real-time audio signal.
[0421] Step 2:
[0422] The device sends the acquired audio signal to the Google Cloud Speech-to-Text API and receives the resulting text information. Here, the audio signal is converted into text data through language analysis.
[0423] Step 3:
[0424] The server sends text information to IBM Watson's Tone Analyzer, which analyzes and receives customer sentiment information. In this step, the tone and atmosphere from the text information are analyzed to identify emotional states such as "anger" or "joy."
[0425] Step 4:
[0426] The server analyzes a combination of text and sentiment information to determine the priority of customer responses and the necessary actions. This analysis determines that a customer needs to be prioritized, especially if they are experiencing significant stress.
[0427] Step 5:
[0428] The server sends a notification to the terminal along with the analysis results. The user (store staff) receives specific response strategies and suggestions based on the emotional information, and responds to the customer accordingly. At this stage, the system presents examples of how to provide quick and appropriate customer service.
[0429] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0430] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0431] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0432] [Third Embodiment]
[0433] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0434] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0435] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0436] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0437] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0438] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0439] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0440] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0441] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0442] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0443] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0444] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0445] The present invention relates to a system for automating tasks based on voice commands, and specific embodiments thereof are shown below.
[0446] First, the user gives work instructions using voice. For example, they might say a specific instruction such as, "Add a project meeting to the schedule for 10 AM tomorrow."
[0447] The device captures the user's voice and generates audio data. This audio data is then transmitted to a server via the network.
[0448] When the server receives audio data, it uses a speech recognition engine to convert it into text data. This text data is then interpreted by an analysis tool to identify the specific instructions required for the business automation tool.
[0449] Based on the analyzed instructions, the server executes automated tasks. This includes various actions in accordance with user instructions, such as schedule management, email creation, and data manipulation. For example, when automatically adding a meeting schedule using the Calendar API, the server accurately reflects the date, time, and content of the meeting.
[0450] Furthermore, when a task is completed, the server sends a completion notification to the terminal via a notification system. The user can then check the notification displayed on the terminal screen and confirm that the task has been performed correctly.
[0451] Thus, the system in this invention receives voice input from the user and provides a series of steps to perform the tasks desired by the user through automated functions, thereby improving work efficiency.
[0452] The following describes the processing flow.
[0453] Step 1:
[0454] The user speaks work instructions into their terminal. These instructions should be clear and specific, so that the system can easily understand them.
[0455] Step 2:
[0456] The device uses a built-in or connected microphone to capture the user's voice in real time. This audio is temporarily stored as digital audio data.
[0457] Step 3:
[0458] The device sends the captured audio data to the server. This communication is typically conducted through a secure protocol to maintain data integrity and confidentiality.
[0459] Step 4:
[0460] The server inputs the received audio data into a speech recognition engine and converts it into text data. This conversion utilizes speech recognition algorithms and applies language models, among other things.
[0461] Step 5:
[0462] The server passes the converted text data to a natural language processing module, which analyzes the instructions. Here, the instructions are broken down into specific tasks, such as changing schedules or composing emails.
[0463] Step 6:
[0464] Based on the analysis results, the server selects the appropriate automation method and executes the necessary tasks. For example, it might access the calendar API to add a meeting at a specified date and time.
[0465] Step 7:
[0466] The server provides feedback to the user's terminal once the task is completed. This feedback includes a notification indicating that the task was successful.
[0467] Step 8:
[0468] Users can check notifications from their devices and confirm that the tasks they instructed have been carried out correctly. Users can also issue further voice instructions as needed.
[0469] (Example 1)
[0470] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0471] Conventional business automation systems lack the flexibility and accuracy to effectively process voice commands, making it difficult to properly interpret and execute user instructions. In particular, there is a need for a means to achieve rapid and accurate automation processes when handling complex instructions and diverse tasks.
[0472] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0473] In this invention, the server includes an information processing device for acquiring voice information, a conversion device for converting the acquired voice information into text information, and an analysis device for analyzing the text information obtained by the conversion device and interpreting the instruction content. This makes it possible to execute automated processes quickly and accurately for complex instructions and diverse tasks.
[0474] "Voice information" refers to instructions and information obtained from users via voice.
[0475] An "information processing device" is a device that has the function of capturing and digitizing audio information.
[0476] A "conversion device" is a device used to convert acquired audio information into text information.
[0477] "Textual information" refers to data that represents audio information as a string of characters.
[0478] An "analysis device" is a device that analyzes converted character information and interprets the specific instructions it contains.
[0479] An "automation device" is a device that automates tasks based on analyzed instructions.
[0480] A "notification device" is a device or function that informs the user of the work status.
[0481] "Other systems" refer to external systems that work in conjunction with each other to perform tasks.
[0482] "Schedule management" refers to the task of managing information and events related to schedules and calendars.
[0483] "Communication creation" refers to the task of creating and sending emails and messages.
[0484] "Information management" refers to the systematic collection, storage, and use of various types of data.
[0485] This invention is a system that automates tasks based on voice commands. This system is implemented through a series of processes involving the user, terminal, and server.
[0486] Users input work requests and instructions into the terminal as voice commands. For example, they might issue a voice command such as, "Please add a team meeting tomorrow at 3 PM." The terminal is equipped with an information processing device that converts the voice received from the user into digital voice data. This device typically includes a microphone and software for voice digitization.
[0487] The terminal sends the generated voice data to the server using a communication protocol. The server has a conversion device that uses speech recognition technology to convert the voice data into text information. Generally, speech recognition APIs or services (e.g., speech recognition cloud services) are used.
[0488] The server further uses natural language processing technology to analyze textual information and interpret user instructions. The analysis device performs this role. This analysis identifies specific automation tasks, such as schedule management and communication creation.
[0489] The server then works with external integrated systems (e.g., scheduling APIs and database management systems) to execute automated tasks based on the analyzed instructions. For example, when automatically adding a meeting to a schedule, it accurately reflects the date, time, and content of that meeting.
[0490] Once the task is complete, the server generates a notification and sends it to the terminal. The user receives the notification through the terminal and can confirm that the instructions were carried out correctly.
[0491] A concrete example is when a user gives a voice command saying, "Please cancel the meeting scheduled for 2 PM." The terminal captures this as voice data and sends it to the server. The server converts the voice data to text and cancels the meeting through analysis. After the task is completed, the terminal receives a notification and informs the user. In this way, an automated system based on voice commands streamlines business processes.
[0492] An example of a prompt message is, "Please tell me the procedure for a user to give work instructions by voice."
[0493] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0494] Step 1:
[0495] The user provides work instructions via voice. The input consists of the user's voice instructions, specifically something like, "Please add a meeting at 3 PM." The user's spoken words are the central data for this step.
[0496] Step 2:
[0497] The terminal captures the user's voice using a microphone and converts it into digital audio data. The input for this step is the user's voice, and the output is digital audio data. Since audio is an analog signal, it undergoes digital processing and is stored as binary data. This process includes audio signal processing.
[0498] Step 3:
[0499] The terminal sends the generated digital audio data to the server. The input is the digital audio data within the terminal, and the output is the audio data sent to the server over the network. In this step, the data is transferred using an appropriate communication protocol.
[0500] Step 4:
[0501] The server converts the received audio data into text information using a speech recognition engine. The input is the audio data passed to the server, and the output is the instructions in text format. This conversion process uses a speech recognition algorithm, which is part of a generative AI model.
[0502] Step 5:
[0503] The server analyzes the converted text information and interprets the content of the instructions. The input is text information, and the output is the interpreted specific task or instruction. In this step, natural language processing is performed, and the parser plays a role in analyzing the content in detail.
[0504] Step 6:
[0505] The server executes the appropriate automated task based on the instructions. The input is the interpreted instructions, and the output is the result of the automated work. For example, this could include calling a scheduling API to add a meeting to the schedule.
[0506] Step 7:
[0507] The server generates a message to notify the terminal of the completion status of the task and sends it. The input is the completion status of the task, and the output is the notification message sent to the user. In this step, a notification device is used to generate a message based on an example prompt.
[0508] Step 8:
[0509] The user receives notifications on their device to confirm that the task was performed correctly. The input is the notification sent from the server, and the output is the user's understanding and confirmation action. The user can view the notification through their device's display.
[0510] (Application Example 1)
[0511] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0512] In autonomous vehicles, it is necessary for the user to safely and efficiently control the vehicle's functions and respond appropriately to the situation while driving. However, operating the vehicle using hands or eyes while driving carries a risk of compromising safety. This invention aims to solve this problem by providing a system that allows the vehicle to be operated automatically by voice commands.
[0513] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0514] In this invention, the server includes a device means for acquiring voice information, a conversion means for converting the acquired voice information into text information, an analysis means for analyzing the text information obtained by the conversion means and interpreting instructions, and means for supporting tasks related to the operation of the vehicle. This makes it possible for the user to safely and efficiently control the functions of the vehicle by voice while driving.
[0515] "Voice information" refers to data acquired based on the voice spoken by the user, and is information input as a voice signal.
[0516] "Device means" is a general term for equipment and devices for acquiring audio information, and in this invention, it specifically refers to devices that have the function of capturing audio data.
[0517] A "conversion means" is a means for performing the process of converting audio information into text information, and it generates text data using speech recognition technology.
[0518] "Textual information" refers to data in text format obtained as a result of converting audio information using a conversion means, and is the subject of interpretation by an analysis means.
[0519] "Analysis means" refers to a process or technology that interprets user instructions based on textual information obtained by the conversion means and derives specific work instructions.
[0520] "Instructions" refer to specific requests or commands given by the user as voice information, and are information that is interpreted by the analysis tool.
[0521] "Means to support tasks related to vehicle operation" refer to means for controlling and adjusting functions within the vehicle based on voice commands, thereby improving convenience and safety in autonomous vehicles.
[0522] In a system implementing the present invention, a device for processing voice information and a series of related software operate in cooperation. When a user inputs specific instructions by voice, that voice information is acquired by the terminal's device. Examples of such devices include smartphones and in-vehicle information systems.
[0523] The acquired audio information is sent from the terminal to the server via a speech recognition API (e.g., Google Cloud Speech-to-Text API). On the server, the audio information is converted into text information using a conversion means. This text information is interpreted by an analysis means to identify the user's instructions. A natural language processing model (e.g., a generative AI model) is used as the analysis means.
[0524] Next, based on the analyzed instructions, the automated system performs operations to control the vehicle's functions. In-vehicle APIs (e.g., CarPlay and Android Auto) are used to support tasks related to vehicle operation. Specifically, this includes adjusting the air conditioning temperature and providing automatic route guidance via navigation.
[0525] Once the task is complete, the server sends a completion notification to the terminal using a notification method, and the user can check the result on the terminal's display screen. For example, if a user says, "Tell me about nearby restaurants," the voice is converted to text, and based on the analysis, directions to the nearest restaurant are displayed on the vehicle's navigation system.
[0526] An example of a prompt for a generative AI model could be: "Please provide brief steps to guide the user to the nearest destination based on voice commands."
[0527] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0528] Step 1:
[0529] The device acquires the user's voice commands. Here, a smartphone or in-vehicle system captures voice data using a microphone. The input is the user's voice itself, and the output is voice data converted into a digital format.
[0530] Step 2:
[0531] The terminal sends the acquired audio data to the server. A communication module is used to deliver the audio data to the server over the network. The input here is digitized audio data, and the output is the audio data as it has been transferred to the server's designated data receiving endpoint.
[0532] Step 3:
[0533] The server receives audio data and uses a conversion mechanism to convert it into text data via a speech recognition API. The speech recognition model analyzes the audio signal and converts it into text. The input is the audio data received by the server, and the output is text data in text format.
[0534] Step 4:
[0535] The server uses a generative AI model to analyze text data and identify instructions. A natural language processing model takes text data as input and analyzes it to identify the user's instructions. The input is text data, and the output is specific instructions.
[0536] Step 5:
[0537] The server automates vehicle operation using means to assist with vehicle operation based on identified instructions. It calls in-vehicle APIs to control specific vehicle functions such as adjusting the air conditioning temperature and setting the navigation system. The input is the instructions obtained through analysis, and the output is the automated functional state within the vehicle.
[0538] Step 6:
[0539] After completing the task, the server notifies the terminal of the result using a notification method. A completion message is sent to the terminal via the network. The input is the status information indicating task completion, and the output is the notification message displayed on the terminal.
[0540] Step 7:
[0541] The user checks the notification on their device and confirms that the task has been completed correctly. The user can see the completion notification message on the device's display. The input is the notification message received from the server, and the output is the completion status information that the user confirms.
[0542] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0543] This invention is a system that automates and further improves the efficiency of tasks using voice commands and emotion recognition.
[0544] The user inputs work-related instructions by voice into the terminal. During this process, information about the user's emotional state is also captured as voice data.
[0545] The device sends voice data to the server. The server converts the received voice data into text data using a speech recognition engine and simultaneously analyzes the user's emotional state using an emotion engine. The emotion engine determines whether the user is experiencing emotions such as "anger," "joy," or "sadness" based on features such as the tone and rhythm of the voice.
[0546] The analysis method analyzes the textualized instructions and also takes into account emotional information obtained from the emotion engine. Based on this emotional information, for example, if it is determined that the user is experiencing stress, the system determines which tasks should be prioritized.
[0547] The business automation system executes tasks by coordinating with necessary external systems based on analysis results. For example, if it determines that a user is in a "high-stress state," it can prioritize scheduling tasks and suggest rescheduling meetings.
[0548] Once a task is completed, a notification system will inform the device of the result. It's also possible to adjust the notification's wording based on emotional information. For example, if the user is feeling "joy," the notification will use positive language.
[0549] Thus, the system of the present invention provides a more user-friendly interface and enables smooth business operations by comprehensively utilizing user voice commands and emotional information.
[0550] The following describes the processing flow.
[0551] Step 1:
[0552] The user speaks in a natural voice along with work instructions. This includes tone, speed, and emphasis, and also conveys emotional information.
[0553] Step 2:
[0554] The device captures the user's voice through the microphone. The captured voice is temporarily stored as digital audio data containing the user's instructions and emotional information.
[0555] Step 3:
[0556] The device sends the captured audio data to the server. This communication is conducted using a secure protocol.
[0557] Step 4:
[0558] The server converts the audio data into text data using a speech recognition engine. During this process, the content of the speech is faithfully reproduced as text.
[0559] Step 5:
[0560] The server analyzes the converted text data using a natural language processing module. Here, the user's instructions are clearly interpreted, and specific business tasks are extracted.
[0561] Step 6:
[0562] Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone and volume of the voice to determine emotions such as "joy," "anger," and "sadness."
[0563] Step 7:
[0564] The server selects the most appropriate automation method for tasks based on the analyzed instructions and emotional information. For example, if it determines that a user is experiencing stress, it prioritizes tasks that require assistance.
[0565] Step 8:
[0566] The server interacts with other systems as needed to perform specific tasks. Data exchange with external systems allows for optimization of schedules and reallocation of resources, for example.
[0567] Step 9:
[0568] After completing a task, the server uses a notification system to report the results to the terminal. The notification reflects emotional information and is presented in a format and content appropriate to the user.
[0569] Step 10:
[0570] Users can check notifications from their devices and evaluate whether the assigned tasks were performed as expected. Additional instructions can be given via voice if necessary.
[0571] (Example 2)
[0572] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0573] Existing business automation systems often process tasks uniformly without considering the user's emotional state, which can increase user stress and burden. Furthermore, notifications may not be sensitive to the user's psychological state, hindering efficient communication. This can negatively impact the user experience.
[0574] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0575] In this invention, the server includes means for acquiring voice data and capturing information including the user's emotional state; means for converting the acquired voice data into text data using a speech recognition engine; and means for analyzing the user's emotional state from the text data obtained by the conversion means and the tone and rhythm of the voice. This enables user-friendly system operation by adjusting the content of automated tasks and notifications according to the user's emotional state.
[0576] "Voice data" refers to information recorded in digital or analog format as audio signals, and includes user voice commands and emotional expressions.
[0577] "Terminal means" refers to a device or apparatus that acquires voice data from a user and transmits that information to a server.
[0578] "Conversion means" refers to software or algorithms used to convert acquired audio data into text data.
[0579] A "speech recognition engine" is a technology or program that extracts linguistic information from speech data and converts its content into text format.
[0580] "Analysis means" refers to a device or software that has the function of analyzing the content of instructions converted from voice data into text and the user's emotional state to determine the content and priority of tasks.
[0581] An "emotion engine" is a component of a system that determines and analyzes a user's emotional state based on characteristics such as tone and rhythm of their voice.
[0582] "Automation methods" are those that efficiently process tasks based on analyzed instructions and emotional information, and have the functionality to link with external systems as needed.
[0583] A "notification means" is a method or device for informing a user of the results of a task, and it is possible to adjust the content and expression of the notification according to the user's emotional state.
[0584] This invention is a system that automates and streamlines tasks through voice commands and emotion recognition. The user inputs task-related instructions via voice into a terminal. The terminal transmits the voice data to a server. This voice data also includes the user's emotional information.
[0585] The server uses a speech recognition engine to convert received audio data into text data. It is possible to customize and use commonly available APIs as the speech recognition engine. Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the tone and rhythmic characteristics within the audio data. The emotion engine uses music analysis technology and machine learning-based emotion analysis algorithms to determine whether the user is experiencing emotions such as "joy," "anger," or "sadness."
[0586] Based on analyzed instructions and emotional information, the server automates tasks. Depending on the type of task, it integrates with schedule management and data management systems to perform the task. For example, if a user requests that a meeting be rescheduled and also indicates stress, the server will prioritize schedule management and suggest rescheduling the meeting.
[0587] Once a task is completed, the server notifies the terminal of the result via a notification system. The notification's wording can be adjusted based on emotional information. For example, if the user is feeling "joy," the notification will be expressed with a positive message.
[0588] For example, if a user gives a voice command saying, "Remind me of the project meeting tomorrow at 10 AM," the server will acknowledge the command, add the event to the calendar if necessary, and send a notification at the appropriate time.
[0589] An example of a prompt sentence to input into the generating AI model is: "Please describe the flow of a system that efficiently automates tasks instructed by the user via voice and prioritizes them according to their emotional state."
[0590] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0591] Step 1:
[0592] The user inputs voice instructions related to their work into the terminal. This voice input includes the user's spoken content, tone, and rhythm. The terminal captures this voice data using a microphone. The input data is raw audio waveform data.
[0593] Step 2:
[0594] The terminal converts the captured audio data into a digital format and sends it to the server over the network. At this point, the input is digital audio data, and the output is the data packets sent to the server.
[0595] Step 3:
[0596] The server inputs the received digital audio data into the speech recognition engine. The engine analyzes the audio signal and converts it into corresponding text data. Specifically, it uses a feature extraction algorithm to identify phonemes and generates words based on them. Here, the input is audio signal data, and the output is text data.
[0597] Step 4:
[0598] The server simultaneously uses an emotion engine to analyze the tone and rhythm of the voice data to identify the emotional state. A machine learning model is used for the analysis, calculating emotions such as "joy" or "anger" from the voice patterns. The input is voice characteristic data, and the output is an emotional state (e.g., joy, anger, sadness).
[0599] Step 5:
[0600] The server analyzes the instructions based on the converted text data and analyzed emotional state. The analysis extracts keywords and phrases from the text and associates them with the work content. The input is text data and emotional information, and the output is a profile of the work instructions.
[0601] Step 6:
[0602] The server integrates with appropriate external systems and automates necessary tasks based on the work order profile and sentiment information. Specifically, it accesses a scheduling management system to coordinate meeting schedules. The input is the work order profile, and the output is the result of the completed tasks.
[0603] Step 7:
[0604] Once a task is completed, the server notifies the terminal of the result via a notification system. This notification is crafted with sentiment information in mind, using positive or mitigating language. The input is the result of the task execution, and the output is the adjusted notification message.
[0605] (Application Example 2)
[0606] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0607] In brick-and-mortar stores where customers interact with staff, employees are required to appropriately recognize each customer's emotional state and respond quickly and appropriately based on that understanding. However, manual emotional recognition and response make it difficult to respond quickly to the diverse emotions and situations of customers. Furthermore, providing personalized service tailored to each customer's emotions is necessary to improve customer satisfaction.
[0608] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0609] In this invention, the server includes a device for acquiring audio signals, a processing means for converting the acquired audio signals into text information, an analysis means for analyzing the text information obtained by the processing means and interpreting the information content, an automation means for automating tasks based on the analyzed information content and emotional information, and a notification means for notifying the user of the completion status of the tasks. This enables efficient and accurate responses that take customer emotions into consideration, as well as the provision of personalized services.
[0610] An "audio signal" refers to a signal that represents audio information in digital or analog format.
[0611] "Device" refers to hardware that has a mechanical or electronic mechanism for achieving a specific purpose.
[0612] A "processing means" is a means for receiving input data and converting it into a desired format or information.
[0613] "Textual information" refers to information composed of characters in natural language that possesses meaning.
[0614] "Analysis means" refers to the means used to analyze input data and interpret specific information or commands.
[0615] "Information content" refers to the actual content of knowledge or instructions transmitted through data or communication.
[0616] "Emotional information" refers to information about emotional states detected from the characteristics of speech and text.
[0617] "Automation methods" are technical means to perform specific tasks or processes without human intervention.
[0618] "Notification means" refers to methods for informing users or systems of status or results.
[0619] "User" refers to a person who operates a system or device, or an end-user.
[0620] To implement this invention, a system utilizing voice commands and emotion recognition will be constructed. This system will primarily aim to improve the efficiency and personalization of customer service in retail stores.
[0621] The server receives audio signals from voice acquisition devices such as smartphones. The hardware used is expected to include smart devices with microphones and network-connected communication devices. The audio signals are converted into text information by the Google Cloud Speech-to-Text API. Then, IBM Watson's Tone Analyzer analyzes sentiment information based on the text information extracted from the audio.
[0622] If the user is a store employee, the analyzed information is used to determine the appropriate response to a customer. For example, if the emotional information indicates that a customer is making a complaint, the system generates an alert to prioritize that customer. This alert is sent to the employee's mobile device to prompt a quick response.
[0623] Furthermore, as a concrete example of improving the efficiency of customer service based on emotion recognition, the system supports prompt phrases such as "Customer feedback: My order is wrong, what are you going to do about it? (Emotion: Anger)." This prompt allows staff to immediately consider countermeasures and provide appropriate service.
[0624] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0625] Step 1:
[0626] The user (store staff) uses the voice input function of their smartphone to capture the customer's voice on the device. This input is acquired as a real-time audio signal.
[0627] Step 2:
[0628] The device sends the acquired audio signal to the Google Cloud Speech-to-Text API and receives the resulting text information. Here, the audio signal is converted into text data through language analysis.
[0629] Step 3:
[0630] The server sends text information to IBM Watson's Tone Analyzer, which analyzes and receives customer sentiment information. In this step, the tone and atmosphere from the text information are analyzed to identify emotional states such as "anger" or "joy."
[0631] Step 4:
[0632] The server analyzes a combination of text and sentiment information to determine the priority of customer responses and the necessary actions. This analysis determines that a customer needs to be prioritized, especially if they are experiencing significant stress.
[0633] Step 5:
[0634] The server sends a notification to the terminal along with the analysis results. The user (store staff) receives specific response strategies and suggestions based on the emotional information, and responds to the customer accordingly. At this stage, the system presents examples of how to provide quick and appropriate customer service.
[0635] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0636] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0637] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0638] [Fourth Embodiment]
[0639] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0640] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0641] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0642] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0643] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0644] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0645] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0646] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0647] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0648] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0649] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0650] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0651] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0652] The present invention relates to a system for automating tasks based on voice commands, and specific embodiments thereof are shown below.
[0653] First, the user gives work instructions using voice. For example, they might say a specific instruction such as, "Add a project meeting to the schedule for 10 AM tomorrow."
[0654] The device captures the user's voice and generates audio data. This audio data is then transmitted to a server via the network.
[0655] When the server receives audio data, it uses a speech recognition engine to convert it into text data. This text data is then interpreted by an analysis tool to identify the specific instructions required for the business automation tool.
[0656] Based on the analyzed instructions, the server executes automated tasks. This includes various actions in accordance with user instructions, such as schedule management, email creation, and data manipulation. For example, when automatically adding a meeting schedule using the Calendar API, the server accurately reflects the date, time, and content of the meeting.
[0657] Furthermore, when a task is completed, the server sends a completion notification to the terminal via a notification system. The user can then check the notification displayed on the terminal screen and confirm that the task has been performed correctly.
[0658] Thus, the system in this invention receives voice input from the user and provides a series of steps to perform the tasks desired by the user through automated functions, thereby improving work efficiency.
[0659] The following describes the processing flow.
[0660] Step 1:
[0661] The user speaks work instructions into their terminal. These instructions should be clear and specific, so that the system can easily understand them.
[0662] Step 2:
[0663] The device uses a built-in or connected microphone to capture the user's voice in real time. This audio is temporarily stored as digital audio data.
[0664] Step 3:
[0665] The device sends the captured audio data to the server. This communication is typically conducted through a secure protocol to maintain data integrity and confidentiality.
[0666] Step 4:
[0667] The server inputs the received audio data into a speech recognition engine and converts it into text data. This conversion utilizes speech recognition algorithms and applies language models, among other things.
[0668] Step 5:
[0669] The server passes the converted text data to a natural language processing module, which analyzes the instructions. Here, the instructions are broken down into specific tasks, such as changing schedules or composing emails.
[0670] Step 6:
[0671] Based on the analysis results, the server selects the appropriate automation method and executes the necessary tasks. For example, it might access the calendar API to add a meeting at a specified date and time.
[0672] Step 7:
[0673] The server provides feedback to the user's terminal once the task is completed. This feedback includes a notification indicating that the task was successful.
[0674] Step 8:
[0675] Users can check notifications from their devices and confirm that the tasks they instructed have been carried out correctly. Users can also issue further voice instructions as needed.
[0676] (Example 1)
[0677] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0678] Conventional business automation systems lack the flexibility and accuracy to effectively process voice commands, making it difficult to properly interpret and execute user instructions. In particular, there is a need for a means to achieve rapid and accurate automation processes when handling complex instructions and diverse tasks.
[0679] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0680] In this invention, the server includes an information processing device for acquiring voice information, a conversion device for converting the acquired voice information into text information, and an analysis device for analyzing the text information obtained by the conversion device and interpreting the instruction content. This makes it possible to execute automated processes quickly and accurately for complex instructions and diverse tasks.
[0681] "Voice information" refers to instructions and information obtained from users via voice.
[0682] An "information processing device" is a device that has the function of capturing and digitizing audio information.
[0683] A "conversion device" is a device used to convert acquired audio information into text information.
[0684] "Textual information" refers to data that represents audio information as a string of characters.
[0685] An "analysis device" is a device that analyzes converted character information and interprets the specific instructions it contains.
[0686] An "automation device" is a device that automates tasks based on analyzed instructions.
[0687] A "notification device" is a device or function that informs the user of the work status.
[0688] "Other systems" refer to external systems that work in conjunction with each other to perform tasks.
[0689] "Schedule management" refers to the task of managing information and events related to schedules and calendars.
[0690] "Communication creation" refers to the task of creating and sending emails and messages.
[0691] "Information management" refers to the systematic collection, storage, and use of various types of data.
[0692] This invention is a system that automates tasks based on voice commands. This system is implemented through a series of processes involving the user, terminal, and server.
[0693] Users input work requests and instructions into the terminal as voice commands. For example, they might issue a voice command such as, "Please add a team meeting tomorrow at 3 PM." The terminal is equipped with an information processing device that converts the voice received from the user into digital voice data. This device typically includes a microphone and software for voice digitization.
[0694] The terminal sends the generated voice data to the server using a communication protocol. The server has a conversion device that uses speech recognition technology to convert the voice data into text information. Generally, speech recognition APIs or services (e.g., speech recognition cloud services) are used.
[0695] The server further uses natural language processing technology to analyze textual information and interpret user instructions. The analysis device performs this role. This analysis identifies specific automation tasks, such as schedule management and communication creation.
[0696] The server then works with external integrated systems (e.g., scheduling APIs and database management systems) to execute automated tasks based on the analyzed instructions. For example, when automatically adding a meeting to a schedule, it accurately reflects the date, time, and content of that meeting.
[0697] Once the task is complete, the server generates a notification and sends it to the terminal. The user receives the notification through the terminal and can confirm that the instructions were carried out correctly.
[0698] A concrete example is when a user gives a voice command saying, "Please cancel the meeting scheduled for 2 PM." The terminal captures this as voice data and sends it to the server. The server converts the voice data to text and cancels the meeting through analysis. After the task is completed, the terminal receives a notification and informs the user. In this way, an automated system based on voice commands streamlines business processes.
[0699] An example of a prompt message is, "Please tell me the procedure for a user to give work instructions by voice."
[0700] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0701] Step 1:
[0702] The user provides work instructions via voice. The input consists of the user's voice instructions, specifically something like, "Please add a meeting at 3 PM." The user's spoken words are the central data for this step.
[0703] Step 2:
[0704] The terminal captures the user's voice using a microphone and converts it into digital audio data. The input for this step is the user's voice, and the output is digital audio data. Since audio is an analog signal, it undergoes digital processing and is stored as binary data. This process includes audio signal processing.
[0705] Step 3:
[0706] The terminal sends the generated digital audio data to the server. The input is the digital audio data within the terminal, and the output is the audio data sent to the server over the network. In this step, the data is transferred using an appropriate communication protocol.
[0707] Step 4:
[0708] The server converts the received audio data into text information using a speech recognition engine. The input is the audio data passed to the server, and the output is the instructions in text format. This conversion process uses a speech recognition algorithm, which is part of a generative AI model.
[0709] Step 5:
[0710] The server analyzes the converted text information and interprets the content of the instructions. The input is text information, and the output is the interpreted specific task or instruction. In this step, natural language processing is performed, and the parser plays a role in analyzing the content in detail.
[0711] Step 6:
[0712] The server executes the appropriate automated task based on the instructions. The input is the interpreted instructions, and the output is the result of the automated work. For example, this could include calling a scheduling API to add a meeting to the schedule.
[0713] Step 7:
[0714] The server generates a message to notify the terminal of the completion status of the task and sends it. The input is the completion status of the task, and the output is the notification message sent to the user. In this step, a notification device is used to generate a message based on an example prompt.
[0715] Step 8:
[0716] The user receives notifications on their device to confirm that the task was performed correctly. The input is the notification sent from the server, and the output is the user's understanding and confirmation action. The user can view the notification through their device's display.
[0717] (Application Example 1)
[0718] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0719] In autonomous vehicles, it is necessary for the user to safely and efficiently control the vehicle's functions and respond appropriately to the situation while driving. However, operating the vehicle using hands or eyes while driving carries a risk of compromising safety. This invention aims to solve this problem by providing a system that allows the vehicle to be operated automatically by voice commands.
[0720] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0721] In this invention, the server includes a device means for acquiring voice information, a conversion means for converting the acquired voice information into text information, an analysis means for analyzing the text information obtained by the conversion means and interpreting instructions, and means for supporting tasks related to the operation of the vehicle. This makes it possible for the user to safely and efficiently control the functions of the vehicle by voice while driving.
[0722] "Voice information" refers to data acquired based on the voice spoken by the user, and is information input as a voice signal.
[0723] "Device means" is a general term for equipment and devices for acquiring audio information, and in this invention, it specifically refers to devices that have the function of capturing audio data.
[0724] A "conversion means" is a means for performing the process of converting audio information into text information, and it generates text data using speech recognition technology.
[0725] "Textual information" refers to data in text format obtained as a result of converting audio information using a conversion means, and is the subject of interpretation by an analysis means.
[0726] "Analysis means" refers to a process or technology that interprets user instructions based on textual information obtained by the conversion means and derives specific work instructions.
[0727] "Instructions" refer to specific requests or commands given by the user as voice information, and are information that is interpreted by the analysis tool.
[0728] "Means to support tasks related to vehicle operation" refer to means for controlling and adjusting functions within the vehicle based on voice commands, thereby improving convenience and safety in autonomous vehicles.
[0729] In a system implementing the present invention, a device for processing voice information and a series of related software operate in cooperation. When a user inputs specific instructions by voice, that voice information is acquired by the terminal's device. Examples of such devices include smartphones and in-vehicle information systems.
[0730] The acquired audio information is sent from the terminal to the server via a speech recognition API (e.g., Google Cloud Speech-to-Text API). On the server, the audio information is converted into text information using a conversion means. This text information is interpreted by an analysis means to identify the user's instructions. A natural language processing model (e.g., a generative AI model) is used as the analysis means.
[0731] Next, based on the analyzed instructions, the automated system performs operations to control the vehicle's functions. In-vehicle APIs (e.g., CarPlay and Android Auto) are used to support tasks related to vehicle operation. Specifically, this includes adjusting the air conditioning temperature and providing automatic route guidance via navigation.
[0732] Once the task is complete, the server sends a completion notification to the terminal using a notification method, and the user can check the result on the terminal's display screen. For example, if a user says, "Tell me about nearby restaurants," the voice is converted to text, and based on the analysis, directions to the nearest restaurant are displayed on the vehicle's navigation system.
[0733] An example of a prompt for a generative AI model could be: "Please provide brief steps to guide the user to the nearest destination based on voice commands."
[0734] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0735] Step 1:
[0736] The device acquires the user's voice commands. Here, a smartphone or in-vehicle system captures voice data using a microphone. The input is the user's voice itself, and the output is voice data converted into a digital format.
[0737] Step 2:
[0738] The terminal sends the acquired audio data to the server. A communication module is used to deliver the audio data to the server over the network. The input here is digitized audio data, and the output is the audio data as it has been transferred to the server's designated data receiving endpoint.
[0739] Step 3:
[0740] The server receives audio data and uses a conversion mechanism to convert it into text data via a speech recognition API. The speech recognition model analyzes the audio signal and converts it into text. The input is the audio data received by the server, and the output is text data in text format.
[0741] Step 4:
[0742] The server uses a generative AI model to analyze text data and identify instructions. A natural language processing model takes text data as input and analyzes it to identify the user's instructions. The input is text data, and the output is specific instructions.
[0743] Step 5:
[0744] The server automates vehicle operation using means to assist with vehicle operation based on identified instructions. It calls in-vehicle APIs to control specific vehicle functions such as adjusting the air conditioning temperature and setting the navigation system. The input is the instructions obtained through analysis, and the output is the automated functional state within the vehicle.
[0745] Step 6:
[0746] After completing the task, the server notifies the terminal of the result using a notification method. A completion message is sent to the terminal via the network. The input is the status information indicating task completion, and the output is the notification message displayed on the terminal.
[0747] Step 7:
[0748] The user checks the notification on their device and confirms that the task has been completed correctly. The user can see the completion notification message on the device's display. The input is the notification message received from the server, and the output is the completion status information that the user confirms.
[0749] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0750] This invention is a system that automates and further improves the efficiency of tasks using voice commands and emotion recognition.
[0751] The user inputs work-related instructions by voice into the terminal. During this process, information about the user's emotional state is also captured as voice data.
[0752] The device sends voice data to the server. The server converts the received voice data into text data using a speech recognition engine and simultaneously analyzes the user's emotional state using an emotion engine. The emotion engine determines whether the user is experiencing emotions such as "anger," "joy," or "sadness" based on features such as the tone and rhythm of the voice.
[0753] The analysis method analyzes the textualized instructions and also takes into account emotional information obtained from the emotion engine. Based on this emotional information, for example, if it is determined that the user is experiencing stress, the system determines which tasks should be prioritized.
[0754] The business automation system executes tasks by coordinating with necessary external systems based on analysis results. For example, if it determines that a user is in a "high-stress state," it can prioritize scheduling tasks and suggest rescheduling meetings.
[0755] Once a task is completed, a notification system will inform the device of the result. It's also possible to adjust the notification's wording based on emotional information. For example, if the user is feeling "joy," the notification will use positive language.
[0756] Thus, the system of the present invention provides a more user-friendly interface and enables smooth business operations by comprehensively utilizing user voice commands and emotional information.
[0757] The following describes the processing flow.
[0758] Step 1:
[0759] The user speaks in a natural voice along with work instructions. This includes tone, speed, and emphasis, and also conveys emotional information.
[0760] Step 2:
[0761] The device captures the user's voice through the microphone. The captured voice is temporarily stored as digital audio data containing the user's instructions and emotional information.
[0762] Step 3:
[0763] The device sends the captured audio data to the server. This communication is conducted using a secure protocol.
[0764] Step 4:
[0765] The server converts the audio data into text data using a speech recognition engine. During this process, the content of the speech is faithfully reproduced as text.
[0766] Step 5:
[0767] The server analyzes the converted text data using a natural language processing module. Here, the user's instructions are clearly interpreted, and specific business tasks are extracted.
[0768] Step 6:
[0769] Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone and volume of the voice to determine emotions such as "joy," "anger," and "sadness."
[0770] Step 7:
[0771] The server selects the most appropriate automation method for tasks based on the analyzed instructions and emotional information. For example, if it determines that a user is experiencing stress, it prioritizes tasks that require assistance.
[0772] Step 8:
[0773] The server interacts with other systems as needed to perform specific tasks. Data exchange with external systems allows for optimization of schedules and reallocation of resources, for example.
[0774] Step 9:
[0775] After completing a task, the server uses a notification system to report the results to the terminal. The notification reflects emotional information and is presented in a format and content appropriate to the user.
[0776] Step 10:
[0777] Users can check notifications from their devices and evaluate whether the assigned tasks were performed as expected. Additional instructions can be given via voice if necessary.
[0778] (Example 2)
[0779] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0780] Existing business automation systems often process tasks uniformly without considering the user's emotional state, which can increase user stress and burden. Furthermore, notifications may not be sensitive to the user's psychological state, hindering efficient communication. This can negatively impact the user experience.
[0781] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0782] In this invention, the server includes means for acquiring voice data and capturing information including the user's emotional state; means for converting the acquired voice data into text data using a speech recognition engine; and means for analyzing the user's emotional state from the text data obtained by the conversion means and the tone and rhythm of the voice. This enables user-friendly system operation by adjusting the content of automated tasks and notifications according to the user's emotional state.
[0783] "Voice data" refers to information recorded in digital or analog format as audio signals, and includes user voice commands and emotional expressions.
[0784] "Terminal means" refers to a device or apparatus that acquires voice data from a user and transmits that information to a server.
[0785] "Conversion means" refers to software or algorithms used to convert acquired audio data into text data.
[0786] A "speech recognition engine" is a technology or program that extracts linguistic information from speech data and converts its content into text format.
[0787] "Analysis means" refers to a device or software that has the function of analyzing the content of instructions converted from voice data into text and the user's emotional state to determine the content and priority of tasks.
[0788] An "emotion engine" is a component of a system that determines and analyzes a user's emotional state based on characteristics such as tone and rhythm of their voice.
[0789] "Automation methods" are those that efficiently process tasks based on analyzed instructions and emotional information, and have the functionality to link with external systems as needed.
[0790] A "notification means" is a method or device for informing a user of the results of a task, and it is possible to adjust the content and expression of the notification according to the user's emotional state.
[0791] This invention is a system that automates and streamlines tasks through voice commands and emotion recognition. The user inputs task-related instructions via voice into a terminal. The terminal transmits the voice data to a server. This voice data also includes the user's emotional information.
[0792] The server uses a speech recognition engine to convert received audio data into text data. It is possible to customize and use commonly available APIs as the speech recognition engine. Simultaneously, the server uses an emotion engine to analyze the user's emotional state from the tone and rhythmic characteristics within the audio data. The emotion engine uses music analysis technology and machine learning-based emotion analysis algorithms to determine whether the user is experiencing emotions such as "joy," "anger," or "sadness."
[0793] Based on analyzed instructions and emotional information, the server automates tasks. Depending on the type of task, it integrates with schedule management and data management systems to perform the task. For example, if a user requests that a meeting be rescheduled and also indicates stress, the server will prioritize schedule management and suggest rescheduling the meeting.
[0794] Once a task is completed, the server notifies the terminal of the result via a notification system. The notification's wording can be adjusted based on emotional information. For example, if the user is feeling "joy," the notification will be expressed with a positive message.
[0795] For example, if a user gives a voice command saying, "Remind me of the project meeting tomorrow at 10 AM," the server will acknowledge the command, add the event to the calendar if necessary, and send a notification at the appropriate time.
[0796] An example of a prompt sentence to input into the generating AI model is: "Please describe the flow of a system that efficiently automates tasks instructed by the user via voice and prioritizes them according to their emotional state."
[0797] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0798] Step 1:
[0799] The user inputs voice instructions related to their work into the terminal. This voice input includes the user's spoken content, tone, and rhythm. The terminal captures this voice data using a microphone. The input data is raw audio waveform data.
[0800] Step 2:
[0801] The terminal converts the captured audio data into a digital format and sends it to the server over the network. At this point, the input is digital audio data, and the output is the data packets sent to the server.
[0802] Step 3:
[0803] The server inputs the received digital audio data into the speech recognition engine. The engine analyzes the audio signal and converts it into corresponding text data. Specifically, it uses a feature extraction algorithm to identify phonemes and generates words based on them. Here, the input is audio signal data, and the output is text data.
[0804] Step 4:
[0805] The server simultaneously uses an emotion engine to analyze the tone and rhythm of the voice data to identify the emotional state. A machine learning model is used for the analysis, calculating emotions such as "joy" or "anger" from the voice patterns. The input is voice characteristic data, and the output is an emotional state (e.g., joy, anger, sadness).
[0806] Step 5:
[0807] The server analyzes the instructions based on the converted text data and analyzed emotional state. The analysis extracts keywords and phrases from the text and associates them with the work content. The input is text data and emotional information, and the output is a profile of the work instructions.
[0808] Step 6:
[0809] The server integrates with appropriate external systems and automates necessary tasks based on the work order profile and sentiment information. Specifically, it accesses a scheduling management system to coordinate meeting schedules. The input is the work order profile, and the output is the result of the completed tasks.
[0810] Step 7:
[0811] Once a task is completed, the server notifies the terminal of the result via a notification system. This notification is crafted with sentiment information in mind, using positive or mitigating language. The input is the result of the task execution, and the output is the adjusted notification message.
[0812] (Application Example 2)
[0813] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0814] In brick-and-mortar stores where customers interact with staff, employees are required to appropriately recognize each customer's emotional state and respond quickly and appropriately based on that understanding. However, manual emotional recognition and response make it difficult to respond quickly to the diverse emotions and situations of customers. Furthermore, providing personalized service tailored to each customer's emotions is necessary to improve customer satisfaction.
[0815] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0816] In this invention, the server includes a device for acquiring audio signals, a processing means for converting the acquired audio signals into text information, an analysis means for analyzing the text information obtained by the processing means and interpreting the information content, an automation means for automating tasks based on the analyzed information content and emotional information, and a notification means for notifying the user of the completion status of the tasks. This enables efficient and accurate responses that take customer emotions into consideration, as well as the provision of personalized services.
[0817] An "audio signal" refers to a signal that represents audio information in digital or analog format.
[0818] "Device" refers to hardware that has a mechanical or electronic mechanism for achieving a specific purpose.
[0819] A "processing means" is a means for receiving input data and converting it into a desired format or information.
[0820] "Textual information" refers to information composed of characters in natural language that possesses meaning.
[0821] "Analysis means" refers to the means used to analyze input data and interpret specific information or commands.
[0822] "Information content" refers to the actual content of knowledge or instructions transmitted through data or communication.
[0823] "Emotional information" refers to information about emotional states detected from the characteristics of speech and text.
[0824] "Automation methods" are technical means to perform specific tasks or processes without human intervention.
[0825] "Notification means" refers to methods for informing users or systems of status or results.
[0826] "User" refers to a person who operates a system or device, or an end-user.
[0827] To implement this invention, a system utilizing voice commands and emotion recognition will be constructed. This system will primarily aim to improve the efficiency and personalization of customer service in retail stores.
[0828] The server receives audio signals from voice acquisition devices such as smartphones. The hardware used is expected to include smart devices with microphones and network-connected communication devices. The audio signals are converted into text information by the Google Cloud Speech-to-Text API. Then, IBM Watson's Tone Analyzer analyzes sentiment information based on the text information extracted from the audio.
[0829] If the user is a store employee, the analyzed information is used to determine the appropriate response to a customer. For example, if the emotional information indicates that a customer is making a complaint, the system generates an alert to prioritize that customer. This alert is sent to the employee's mobile device to prompt a quick response.
[0830] Furthermore, as a concrete example of improving the efficiency of customer service based on emotion recognition, the system supports prompt phrases such as "Customer feedback: My order is wrong, what are you going to do about it? (Emotion: Anger)." This prompt allows staff to immediately consider countermeasures and provide appropriate service.
[0831] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0832] Step 1:
[0833] The user (store staff) uses the voice input function of their smartphone to capture the customer's voice on the device. This input is acquired as a real-time audio signal.
[0834] Step 2:
[0835] The device sends the acquired audio signal to the Google Cloud Speech-to-Text API and receives the resulting text information. Here, the audio signal is converted into text data through language analysis.
[0836] Step 3:
[0837] The server sends text information to IBM Watson's Tone Analyzer, which analyzes and receives customer sentiment information. In this step, the tone and atmosphere from the text information are analyzed to identify emotional states such as "anger" or "joy."
[0838] Step 4:
[0839] The server analyzes a combination of text and sentiment information to determine the priority of customer responses and the necessary actions. This analysis determines that a customer needs to be prioritized, especially if they are experiencing significant stress.
[0840] Step 5:
[0841] The server sends a notification to the terminal along with the analysis results. The user (store staff) receives specific response strategies and suggestions based on the emotional information, and responds to the customer accordingly. At this stage, the system presents examples of how to provide quick and appropriate customer service.
[0842] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0843] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0844] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0845] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0846] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0847] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0848] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0849] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0850] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0851] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0852] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0853] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0854] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0855] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0856] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0857] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0858] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0859] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0860] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0861] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0862] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0863] The following is further disclosed regarding the embodiments described above.
[0864] (Claim 1)
[0865] A terminal means for acquiring audio data,
[0866] A conversion means for converting acquired audio data into text data,
[0867] An analysis means that analyzes the text data obtained by the conversion means and interprets the instructions,
[0868] An automation method that automates tasks based on the analyzed instructions,
[0869] A notification method for informing users of the completion status of tasks,
[0870] A system that includes this.
[0871] (Claim 2)
[0872] The system according to claim 1, wherein the analysis means determines whether the task is schedule management, email creation, or data management.
[0873] (Claim 3)
[0874] The system according to claim 1, wherein the automation means performs work in cooperation with an external system based on the judgment result of the analysis means.
[0875] "Example 1"
[0876] (Claim 1)
[0877] An information processing device that acquires audio information,
[0878] A conversion device that converts acquired audio information into text information,
[0879] An analysis device that analyzes the character information obtained by the conversion device and interprets the instructions,
[0880] An automated device that automates tasks based on the analyzed instructions,
[0881] A notification device that informs the user of the completion status of the work,
[0882] A system that includes this.
[0883] (Claim 2)
[0884] The system according to claim 1, wherein the analysis device determines the content of one of the following tasks: schedule management, communication creation, or information management.
[0885] (Claim 3)
[0886] The system according to claim 1, wherein the automated device performs tasks in cooperation with other systems based on the judgment results of the analysis device.
[0887] "Application Example 1"
[0888] (Claim 1)
[0889] A device means for acquiring audio information,
[0890] A conversion means for converting acquired audio information into text information,
[0891] An analysis means that analyzes the textual information obtained by the conversion means and interprets the instructions,
[0892] An automation means that automates tasks based on the analyzed instructions,
[0893] Means to support tasks related to vehicle operation,
[0894] A notification method to inform the user of the completion status of the work,
[0895] A system that includes this.
[0896] (Claim 2)
[0897] The system according to claim 1, wherein the analysis means determines the content of any of the following tasks: planning management, communication creation, or information management, and interprets instructions relating to environmental adjustment or route guidance within the vehicle.
[0898] (Claim 3)
[0899] The system according to claim 1, wherein the automation means performs work in cooperation with an external device based on the judgment result of the analysis means and assists in the functional control of the vehicle.
[0900] "Example 2 of combining an emotion engine"
[0901] (Claim 1)
[0902] A terminal means for acquiring audio data and capturing information including the user's emotional state,
[0903] A conversion means for converting acquired audio data into text data using a speech recognition engine,
[0904] An analysis means that analyzes the user's emotional state from the text data obtained by the conversion means and the tone and rhythm of the voice,
[0905] An automation method that automates tasks based on analyzed instructions and emotional information,
[0906] A notification method that informs the user of the completion status of a task using expressions adjusted based on emotional information,
[0907] A system that includes this.
[0908] (Claim 2)
[0909] The system according to claim 1, wherein the analysis means determines the content of one of the tasks: schedule management, email creation, or data management, and further takes into account the user's emotional state.
[0910] (Claim 3)
[0911] The system according to claim 1, wherein the automation means cooperates with an external system based on the judgment results of the analysis means, and performs the work after adjusting priorities based on emotional states.
[0912] "Application example 2 of combining emotional engines"
[0913] (Claim 1)
[0914] A device for acquiring audio signals,
[0915] A processing means for converting acquired audio signals into text information,
[0916] An analysis means that analyzes the text information obtained by the processing means and interprets the information content,
[0917] An automation means that automates tasks based on the analyzed information content and emotional information,
[0918] A notification method for informing the user of the completion status of the work,
[0919] A system that includes this.
[0920] (Claim 2)
[0921] The system according to claim 1, wherein the analysis means determines the content of any of the following tasks: time management, electronic message creation, or information management, and adjusts the priority of the tasks based on emotion recognition.
[0922] (Claim 3)
[0923] The system according to claim 1, wherein the automation means performs tasks in cooperation with an external device based on the judgment results of the analysis means and emotional information, and provides emotionally appropriate feedback to the user. [Explanation of Symbols]
[0924] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A terminal means for acquiring audio data, A conversion means for converting acquired audio data into text data, An analysis means that analyzes the text data obtained by the conversion means and interprets the instructions, An automation method that automates tasks based on the analyzed instructions, A notification method for informing users of the completion status of tasks, A system that includes this.
2. The system according to claim 1, wherein the analysis means determines whether the task is schedule management, email creation, or data management.
3. The system according to claim 1, wherein the automation means performs work in cooperation with an external system based on the judgment result of the analysis means.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A