system

A system using natural language processing to convert voice inputs into text and control smart devices addresses the challenge of senior users operating complex technology, enhancing learning efficiency and reducing digital divide.

JP2026104599APending Publication Date: 2026-06-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-13
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Senior users face challenges in operating complex smart devices due to unfamiliarity with digital technology, leading to inefficiencies in learning and digital divide, and existing educational methods fail to provide personalized guidance.

Method used

A system that utilizes natural language processing to convert voice input into text, analyze user requests, control smart devices, and provide voice feedback, creating an interactive learning experience.

Benefits of technology

Enables senior users to easily operate smart devices by converting voice commands into actionable instructions, providing intuitive guidance and reducing anxiety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026104599000001_ABST
    Figure 2026104599000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A speech recognition means that receives voice input from a user and converts it into text data, A natural language processing means that analyzes the aforementioned text data and determines the user's request, A control means for operating the user's information processing device based on the determined request, A feedback mechanism that provides voice feedback to the user regarding the results of the operation, A means of executing a time management function based on voice requests, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For the senior layer, the operation of smart devices is complex and diverse, so there are significant hurdles to its use. Their unfamiliarity with digital devices limits daily communication and information access, contributing to the digital divide. With previous educational methods, it has been difficult to provide guidance according to individual needs, making learning inefficient. Therefore, there is a need for a system that can provide an interactive and personalized learning experience.

Means for Solving the Problems

[0005] To solve this problem, the present invention provides a natural language processing means that converts voice input from the user into text data and analyzes the text data to determine the user's request. Furthermore, it includes a control means that operates the user's smart device based on that determination, and provides voice feedback to the user on the results of the operation, thereby creating an interactive system in which the user can learn the actual procedure while operating. This makes it possible for senior users to easily use smart devices and enjoy their benefits.

[0006] "User" refers to an individual who uses the system, and specifically to senior citizens who are unfamiliar with operating digital devices.

[0007] "Voice input" is a method of sending voice information spoken by a user to a system as digital data.

[0008] "Text data" refers to digital character information converted from speech using speech recognition technology.

[0009] "Voice recognition means" refers to an element that has the function of analyzing voice input from a user and converting it into text data.

[0010] "Natural language processing means" refers to elements that analyze text data obtained by speech recognition means and perform processing to understand the user's intentions and requests.

[0011] "Control means" refers to elements for operating applications and functions on a smart device based on user requests determined by natural language processing means.

[0012] A "feedback mechanism" is an element that communicates the results of smart device operations and the next steps to the user via voice, thereby assisting the user's learning.

[0013] A "smart device" refers to a digital device such as a smartphone or tablet that is operated by a system and used by the user in their daily life. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention relates to an interactive system that assists in the operation of smart devices based on voice input from the user. In particular, it provides a technology that combines voice recognition, natural language processing, and device control to enable senior users to easily operate digital devices.

[0036] First, the user requests a specific action by voice to their smart device. This voice input is captured by the device and converted into text data using speech recognition software within the device. This text data is then sent from the device to a server to understand the user's request.

[0037] The server analyzes the received text data using a natural language processing algorithm. This analysis determines what the user wants and generates a set of instructions to address that request. The server then sends these instructions to the terminal, providing specific control commands for the smart device.

[0038] Based on instructions from the server, the terminal launches the appropriate application or performs specific device functions. The terminal also uses feedback mechanisms to provide the user with voice guidance regarding the results of the operation and the next steps. This allows the user to use the device with confidence while confirming the operating procedures.

[0039] As a concrete example, consider the case where a user voice-inputs "I want to take a picture." The device converts this voice into text data and sends it to the server. The server recognizes the "take a picture" request and sends a command to launch the camera app to the device. The device opens the camera app and supports the user's operation by providing voice guidance such as, "The camera has been launched. Please press the shutter button."

[0040] This system will make it easier for senior users who were previously apprehensive about using digital devices to use them.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user provides specific operational instructions to the smart device via voice. For example, they might say, "I want to take a picture."

[0044] Step 2:

[0045] The terminal receives voice input from the user and uses its built-in speech recognition software to convert this voice into text data. It then sends this text data to the server.

[0046] Step 3:

[0047] The server uses natural language processing algorithms to analyze the user's request in order to parse the received text data. It identifies the instruction "Take a picture" and generates the corresponding procedure.

[0048] Step 4:

[0049] The server sends the generated operating procedure to the terminal, instructing it on the specific commands it should execute.

[0050] Step 5:

[0051] The device follows instructions from the server and launches the camera app. It also provides voice guidance to the user, informing them that the camera has been launched and prompting them to "Press the shutter button."

[0052] Step 6:

[0053] The user follows the instructions on their device and takes a photo by pressing the shutter button in the camera app.

[0054] Step 7:

[0055] The device detects when a photo has been taken and saves the photo to the gallery. It then provides voice feedback to the user when the saving process is complete.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] Currently, many users, especially the elderly, experience confusion and anxiety when using electronic devices. This limits the use of digital devices and diminishes user convenience. In addition, there are challenges in terms of accuracy and ease of use when controlling devices using voice input.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes a voice processing means that receives voice input and converts it into text data, a language analysis means that analyzes the text data to determine the user's request, and a control mechanism that operates the user's electronic device. This enables intuitive and simple operation of the electronic device via voice input.

[0061] "Voice processing means" refers to a device or software that has the function of receiving voice input from a user and converting it into analyzable text data.

[0062] "Language analysis means" refers to an element that has the function of analyzing text data and understanding and judging the content of the user's request, and includes algorithms for natural language processing.

[0063] A "control mechanism" is a system for specifically operating the applications and functions of an electronic device based on user requirements.

[0064] A "response-providing means" is a device or software that has the function of providing feedback to the user, such as the result of an operation or the next step in an operation, via voice or other means.

[0065] "Electronic devices" refer to devices capable of performing digital processing, such as smartphones, tablets, and smart speakers.

[0066] This invention is a system that improves the convenience of users operating electronic devices by voice. Designed to allow a wide range of users, including seniors, to intuitively operate the device, it is achieved through the following main components:

[0067] First, the user gives instructions to the device using voice. For example, they can issue specific commands by voice, such as "Play music" or "I want to take a picture." The voice is captured by the microphone built into the smart device.

[0068] Next, the device is responsible for converting voice input into text data. This uses high-precision speech recognition software (for example, commonly used speech recognition APIs) to accurately generate text from speech. Noise reduction and other processes are performed in the background during this process to improve the accuracy of the conversion.

[0069] Furthermore, the terminal sends the converted text data to the server. The data is transmitted securely and then sent to the server for analysis. The server analyzes the received text data using natural language processing techniques (e.g., generative AI models). Here, the user's true intentions and requests are determined, and appropriate control procedures are generated.

[0070] The server then returns appropriately generated operating instructions to the terminal, which in turn controls the corresponding functions of the electronic device. For example, when playing music, the music app is automatically launched and the song is played. Simultaneously with the control, voice feedback about the operation result is also provided. An announcement such as "Music playback app launched. Playing the current playlist" is made, allowing the user to confirm that the operation is complete.

[0071] For example, if a user gives a voice command such as "Play music," the device converts the voice into text data, which is then analyzed by the server. Based on the analysis results, the server issues a command to launch the music app, and the device follows suit, launching the music app and starting playback.

[0072] An example of a prompt for a generative AI model would be: "Explain how the system should process a user's request to 'play music.'"

[0073] This system enables intuitive and effective operation of electronic devices using voice commands, significantly improving user convenience.

[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0075] Step 1:

[0076] The user inputs specific commands to the smart device using voice. For example, they might say, "Play music." This voice input is the starting point of the program, and the smart device's built-in microphone captures the voice. The input data is the user's voice.

[0077] Step 2:

[0078] The device converts acquired voice input into text data using speech recognition software. Specifically, a high-precision speech recognition engine analyzes the voice signal and generates the corresponding text. The input is the user's voice data, and the output is text data.

[0079] Step 3:

[0080] The terminal sends the converted text data to the server. The data is encrypted and transferred securely, thus protecting user privacy. The input is text data, and the output is the completion of the data transmission to the server.

[0081] Step 4:

[0082] The server analyzes the received text data using natural language processing techniques. It uses a generative AI model to understand the intent of the text and generate specific control procedures related to the user's request. The input is text data, and the output is the generated operation procedure.

[0083] Step 5:

[0084] The server sends the generated operation procedure to the terminal. This prepares the terminal to perform the control requested by the user. The input is the operation procedure data, and the output is the completion of sending the instruction to the terminal.

[0085] Step 6:

[0086] The terminal performs operations on the user's electronic device based on the received instructions. Specifically, for example, a music playback application is automatically launched, and the specified operation (e.g., playing music) is performed. The input is instruction data from the server, and the output is the operating state of the device based on those instructions.

[0087] Step 7:

[0088] The device provides the user with voice feedback regarding the results of the operation and information about the next steps. This feedback allows the user to confirm that the operation was successful and to know what to do next. The input is data of the operation result, and the output is voice feedback.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] The challenge is to provide a system that supports elderly and other users unfamiliar with operating information devices, enabling them to easily operate these digital devices using voice commands and efficiently handle tasks, particularly those related to time management.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes speech recognition means that receive voice input from the user and convert it into text data, natural language processing means that analyze the text data and determine the user's request, and control means that operate the user's information processing device based on the determined request. This enables elderly people to manage their time, such as setting reminders by voice.

[0094] A "speech recognition means" is a device that receives speech input from a user and converts that speech into text data.

[0095] "Natural language processing means" includes technologies for analyzing text data converted by speech recognition means and determining user requests.

[0096] "Control means" refers to a device that has the function of operating the user's information processing device based on requests determined by natural language processing means.

[0097] A "feedback mechanism" is a function that notifies the user of the results of an operation via voice.

[0098] "Means for executing time management functions based on voice requests" refers to technology for executing time management functions, such as setting reminders, in response to a user's voice requests.

[0099] This invention includes an implementation of a system that combines speech recognition, natural language processing, and device control technologies. For example, Google® Speech-to-Text API can be used as speech recognition software. This converts voice input from the user into text data. This text data is sent to a server in the cloud, where it is analyzed using natural language processing algorithms such as IBM Watson® Natural Language Understanding. Based on the analysis results, generated instructions are sent to the user's information processing device via control means, and the device is operated appropriately.

[0100] For example, if a user says to their smartphone, "Set a reminder to take my medicine at 8 AM tomorrow," the system will process this voice and set a reminder. Once this operation is complete, a voice confirmation message saying, "Reminder set," will be provided via a feedback mechanism.

[0101] The advantage of this system lies in its ability to allow elderly users to intuitively operate digital devices through voice commands. This makes it easy for even users who are not tech-savvy to utilize time management functions.

[0102] Examples of prompt statements include the following:

[0103] "Please explain how you would interpret a user's voice message, 'Remind me to take my medicine at 7 AM tomorrow,' and how you would set a reminder."

[0104] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0105] Step 1:

[0106] The user makes a voice input to their smart device. Here, the user gives a voice command such as, "Set up a reminder to take my medicine tomorrow at 8 AM." The input is voice data.

[0107] Step 2:

[0108] The device receives voice input and converts the voice data into text data using the Google Speech-to-Text API. The output is text data. In this process, the voice waveform is mapped to text.

[0109] Step 3:

[0110] The terminal sends text data to a server in the cloud. The input is text data, and the output is the transmission of data to the server. At this stage, data transfer takes place over the network.

[0111] Step 4:

[0112] The server analyzes the received text data using IBM Watson Natural Language Understanding. The input is text data, and the output is the analysis result of the user request. In this process, natural language processing algorithms are used to understand the meaning of the text.

[0113] Step 5:

[0114] Based on the analysis results, the server determines the time management requirement and generates instructions for setting reminders. The input is the analysis results, and the output is instructions for the device. During this process, control commands for specific tasks are generated.

[0115] Step 6:

[0116] The server sends control commands to the terminal. The input is the control commands, and the output is the command sent to the terminal. Even at this stage, data is exchanged via the communication network.

[0117] Step 7:

[0118] Based on the commands received by the device, it executes time management functions and sets reminders. The input is a control command, and the output is the completion of the reminder setting. The app launches on the device and the settings are saved.

[0119] Step 8:

[0120] The device uses a feedback mechanism to notify the user via voice, "Reminder set." The input is the completion information for setting the reminder, and the output is the voice feedback. In this final step, the user is informed of the result of the operation.

[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0122] This invention is an interactive system for operating a smart device based on voice input from a user, and further possesses the function of recognizing the user's emotional state and adjusting the device operation and feedback accordingly. The specific implementation method of this system will be described below.

[0123] First, the user inputs operating instructions by voice into the smart device. The device captures this voice, converts it into text data using voice recognition, and simultaneously uses an emotion engine to identify the user's emotions from the voice. The converted text data and identified emotion data are sent from the device to the server.

[0124] The server analyzes text data using natural language processing to determine user requests. It also takes into account emotional information obtained by the emotion engine to determine the optimal response and action based on the user's current emotional state. Based on this determination, the server uses control mechanisms to instruct specific operations on the smart device and sets the method of providing feedback to the user. In particular, it adjusts the tone and content of the feedback based on emotional information to ensure the user feels as comfortable using the device as possible.

[0125] For example, if a user impatiently says "I want to take a picture" using voice input, the emotion engine will recognize that impatience. The server will recognize the request to "take a picture" and, taking into account the user's impatient emotional state, instruct the device to provide gentle feedback such as, "It's okay to calm down. I'll launch the camera now." This provides an emotionally empathetic response, reducing the stress the user experiences during the operation.

[0126] In this way, by incorporating user emotional states into the interaction, the system becomes more intuitive and user-friendly, providing a particularly accessible learning and operating environment for senior users.

[0127] The following describes the processing flow.

[0128] Step 1:

[0129] The user provides voice input to a smart device, giving instructions for specific actions. This voice input may contain emotions.

[0130] Step 2:

[0131] The device receives the user's voice and converts it into text data using speech recognition technology. Simultaneously, it uses an emotion engine to identify the user's emotional state from the voice.

[0132] Step 3:

[0133] The terminal sends the converted text data and sentiment data to the server. This prepares the server to consider the user's request and sentiment simultaneously.

[0134] Step 4:

[0135] The server uses natural language processing to analyze text data and identify the user's specific requests. It determines actions such as "take a picture" and analyzes emotional information obtained from voice to understand the user's emotional state.

[0136] Step 5:

[0137] Based on the acquired requests and emotions, the server sends instructions to the terminal through control mechanisms. These instructions include specific operations on the smart device (e.g., launching the camera app) and emotionally sensitive feedback.

[0138] Step 6:

[0139] The device follows instructions from the server, performing specified device operations and providing emotionally responsive voice feedback to the user. For users experiencing frustration, it provides guidance in a calmer tone.

[0140] Step 7:

[0141] The user follows the instructions on the device and completes the necessary device operations. The device notifies the server of the results and provides guidance on the next steps as needed. By adjusting the guidance method according to the user's mood, a stress-free experience is provided.

[0142] (Example 2)

[0143] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0144] Conventional voice control systems process information and provide feedback without considering the user's emotional state, resulting in a uniform user experience and a particular difficulty in appropriately responding to emotionally charged voice input. Therefore, there is an urgent need to realize an interactive system that can provide appropriate operation and feedback in accordance with the user's emotions.

[0145] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0146] In this invention, the server includes speech recognition means for receiving voice input and converting it into text data, emotion recognition means for identifying the user's emotional state from the voice input, natural language processing means for analyzing the text data and determining the user's request, and feedback adjustment means for adjusting the tone and content of the feedback. This enables adaptive operation and feedback that takes the user's emotional state into consideration.

[0147] "Voice recognition means" refers to technology that digitizes voice input from a user and converts that voice information into text data. It includes devices and programs that analyze voice patterns and convert them into corresponding strings of characters.

[0148] "Emotion recognition means" refers to technology that identifies a user's emotions from voice input. This includes algorithms and devices for analyzing voice characteristics and inferring emotional states.

[0149] "Natural language processing means" refers to technologies for analyzing character data and understanding user requests and intentions. This includes methods for appropriately interpreting input strings and guiding necessary responses and actions.

[0150] "Feedback adjustment techniques" refer to technologies that adjust the tone and content of feedback based on the user's emotional state. They are means of improving the user experience by generating appropriate responses and taking emotions into consideration.

[0151] "Control means" refers to technology that operates the user's electronic devices based on determined requests. This includes functions that instruct appropriate actions based on data processing.

[0152] This system is an interactive system that performs emotionally responsive device operation based on user voice input. When the user gives voice commands, the terminal captures the voice. The device's microphone is used for voice capture, and the voice signal is converted into digital data.

[0153] The device converts this audio data into text data using speech recognition software (e.g., a general-purpose speech recognition API). Simultaneously, it uses an emotion engine (e.g., a common emotion analysis API) to identify the user's emotions from the audio. The resulting text data and emotion data are then sent from the device to the server.

[0154] The server uses natural language processing techniques (e.g., general-purpose natural language processing models) to analyze the received text data and understand the user's request. It also considers sentiment data and determines the optimal response or action based on the user's emotional state. In particular, it adjusts the content and tone of feedback based on sentiment information to ensure that what is communicated to the user is appropriate.

[0155] For example, if a user angrily commands "Play some music," the server will take that anger into consideration and instruct the device to provide considerate feedback such as "Playing some relaxing music."

[0156] An example of a prompt that utilizes a generative AI model is an instruction such as, "When the user becomes excited and gives a voice command, generate feedback to soothe their emotions." This system configuration enables device operation that is sensitive to the user's emotions.

[0157] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0158] Step 1:

[0159] The user directs their voice input towards the smart device. The device captures the voice through its microphone and receives it as digital audio data. This digital audio data becomes the input.

[0160] Step 2:

[0161] The terminal uses speech recognition software to convert digital speech data into text data. The input is digital speech data, and the output is text data. The software analyzes the speech waveform and generates the corresponding string of characters.

[0162] Step 3:

[0163] The device uses an emotion engine to analyze the user's emotions from their voice. The input is digital voice data, and the output is identified emotion data. This process estimates the user's emotions by evaluating characteristics such as voice tone and speed.

[0164] Step 4:

[0165] The terminal sends converted text data and sentiment data to the server. The input is text data and sentiment data, and the output is the data sent to the server. The data is securely transmitted to the server over the network.

[0166] Step 5:

[0167] The server uses natural language processing technology to analyze text data and determine user requests. The input is text data, and the output is the interpreted user request. It analyzes the structure and meaning of the text to identify the request.

[0168] Step 6:

[0169] The server considers emotional data and uses feedback adjustment mechanisms to determine the most appropriate response. The input is emotional data and the interpreted request, while the output is the adjusted feedback content. It selects a tone and content appropriate to the emotional state.

[0170] Step 7:

[0171] The server sends the determined feedback content to the terminal, which then provides feedback to the user via audio or screen. The input is the feedback content, and the output is the display of feedback to the user or audio playback. This ensures that an appropriate response is given to the user's input.

[0172] (Application Example 2)

[0173] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0174] The aim is to create a more natural and intuitive interface by providing device responses that take into account the user's emotional state when they operate the device via voice input. In particular, in information processing devices used in the home, there is a challenge in reducing user stress and improving the user experience by enabling feedback that responds to the user's emotions.

[0175] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0176] In this invention, the server includes speech recognition means, natural language processing means, and emotion recognition means. This makes it possible to convert speech input into text data and simultaneously identify requests and emotions.

[0177] "Voice recognition means" refers to a device or technology that receives voice input from a user and converts it into text data.

[0178] "Natural language processing means" refers to technologies that have the ability to analyze text data and determine user requests.

[0179] "Emotion recognition means" refers to technologies and devices that identify a user's emotional state from voice input.

[0180] An "information processing device" refers to an electronic device that controls its operation based on the user's requests and emotional state.

[0181] "Control means" refers to means for operating an information processing device based on judged requests and identified emotional states.

[0182] A "feedback mechanism" is a technology that provides users with information in voice that corresponds to the results of their actions and their emotions.

[0183] This invention relates to a method for interactively operating an information processing device using speech recognition means, natural language processing means, emotion recognition means, and feedback means. The system is intended for use in the home and aims to provide personalized feedback based on the user's emotional state.

[0184] First, the user inputs voice data into the information processing device via a microphone. The device then converts this voice into text data using speech recognition software such as the Google Speech API.

[0185] Next, the converted text is analyzed by natural language processing tools to understand the user's request. This process extracts specific operation instructions from the text data.

[0186] In parallel, voice input is analyzed through emotion recognition mechanisms to identify the user's emotional state. This emotion analysis utilizes emotion recognition models built with TENSORFLOW® or PyTorch.

[0187] The server controls the operation of the information processing device by considering requests obtained through natural language processing and emotions obtained through emotion recognition. In particular, it can adjust the tone and content of responses provided through feedback mechanisms according to the user's emotional state.

[0188] For example, if a user asks in a tense voice, "What's the weather forecast?", the system will generate a gentle and reassuring response such as, "It's going to be sunny today. It'll be warm during the day, so please relax."

[0189] An example of a prompt for a generative AI model might be, "Please think of a reassuring response for a user who is feeling anxious." In this way, the system supports the user's actions by being sensitive to their emotions.

[0190] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0191] Step 1:

[0192] The terminal receives voice input from the user through the microphone and acquires that voice data. It obtains an audio signal as input and proceeds to process it as digital data.

[0193] Step 2:

[0194] The device processes the acquired audio data using speech recognition technology and converts it into text data. Specifically, it uses APIs such as Google Speech to convert digital audio signals into natural language text. The output is text data indicating the request.

[0195] Step 3:

[0196] The device analyzes voice data using emotion recognition technology to identify the user's emotional state based on their voice tone, pitch, and speed. It utilizes emotion recognition models built with TensorFlow or PyTorch for analysis. The output is data including emotion labels.

[0197] Step 4:

[0198] The server analyzes text data using natural language processing to understand user requests. The input is text data converted by the terminal, and the server analyzes the request content to identify operation commands. The output is information containing these operation commands.

[0199] Step 5:

[0200] The server determines the operation of the information processing device based on the acquired request and sentiment information. It combines the obtained operation commands and sentiment data to determine the optimal operation and sends a control signal to the terminal.

[0201] Step 6:

[0202] The terminal receives control signals from the server and operates the information processing device. Specific actions include launching applications and executing specific functions.

[0203] Step 7:

[0204] The device generates voice feedback for the user, responding with a tone and content that matches the user's emotional state. For example, if the user is feeling anxious, it will play reassuring content in a gentle voice through the speaker. The output is a voice response that is attentive to the user's emotions.

[0205] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0206] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0207] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0208] [Second Embodiment]

[0209] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0210] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0211] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0212] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0213] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0214] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0215] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0216] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0217] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0218] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0219] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0220] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0221] This invention relates to an interactive system that assists in the operation of smart devices based on voice input from the user. In particular, it provides a technology that combines voice recognition, natural language processing, and device control to enable senior users to easily operate digital devices.

[0222] First, the user requests a specific action by voice to their smart device. This voice input is captured by the device and converted into text data using speech recognition software within the device. This text data is then sent from the device to a server to understand the user's request.

[0223] The server analyzes the received text data using a natural language processing algorithm. This analysis determines what the user wants and generates a set of instructions to address that request. The server then sends these instructions to the terminal, providing specific control commands for the smart device.

[0224] Based on instructions from the server, the terminal launches the appropriate application or performs specific device functions. The terminal also uses feedback mechanisms to provide the user with voice guidance regarding the results of the operation and the next steps. This allows the user to use the device with confidence while confirming the operating procedures.

[0225] As a concrete example, consider the case where a user voice-inputs "I want to take a picture." The device converts this voice into text data and sends it to the server. The server recognizes the "take a picture" request and sends a command to launch the camera app to the device. The device opens the camera app and supports the user's operation by providing voice guidance such as, "The camera has been launched. Please press the shutter button."

[0226] This system will make it easier for senior users who were previously apprehensive about using digital devices to use them.

[0227] The following describes the processing flow.

[0228] Step 1:

[0229] The user provides specific operational instructions to the smart device via voice. For example, they might say, "I want to take a picture."

[0230] Step 2:

[0231] The terminal receives voice input from the user and uses its built-in speech recognition software to convert this voice into text data. It then sends this text data to the server.

[0232] Step 3:

[0233] The server uses natural language processing algorithms to analyze the user's request in order to parse the received text data. It identifies the instruction "Take a picture" and generates the corresponding procedure.

[0234] Step 4:

[0235] The server sends the generated operating procedure to the terminal, instructing it on the specific commands it should execute.

[0236] Step 5:

[0237] The device follows instructions from the server and launches the camera app. It also provides voice guidance to the user, informing them that the camera has been launched and prompting them to "Press the shutter button."

[0238] Step 6:

[0239] The user follows the instructions on their device and takes a photo by pressing the shutter button in the camera app.

[0240] Step 7:

[0241] The device detects when a photo has been taken and saves the photo to the gallery. It then provides voice feedback to the user when the saving process is complete.

[0242] (Example 1)

[0243] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0244] Currently, many users, especially the elderly, experience confusion and anxiety when using electronic devices. This limits the use of digital devices and diminishes user convenience. In addition, there are challenges in terms of accuracy and ease of use when controlling devices using voice input.

[0245] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0246] In this invention, the server includes a voice processing means that receives voice input and converts it into text data, a language analysis means that analyzes the text data to determine the user's request, and a control mechanism that operates the user's electronic device. This enables intuitive and simple operation of the electronic device via voice input.

[0247] "Voice processing means" refers to a device or software that has the function of receiving voice input from a user and converting it into analyzable text data.

[0248] "Language analysis means" refers to an element that has the function of analyzing text data and understanding and judging the content of the user's request, and includes algorithms for natural language processing.

[0249] A "control mechanism" is a system for specifically operating the applications and functions of an electronic device based on user requirements.

[0250] A "response-providing means" is a device or software that has the function of providing feedback to the user, such as the result of an operation or the next step in an operation, via voice or other means.

[0251] "Electronic devices" refer to devices capable of performing digital processing, such as smartphones, tablets, and smart speakers.

[0252] This invention is a system that improves the convenience of users operating electronic devices by voice. Designed to allow a wide range of users, including seniors, to intuitively operate the device, it is achieved through the following main components:

[0253] First, the user gives instructions to the device using voice commands. For example, they can issue specific commands by voice, such as "Play music" or "I want to take a picture." The voice is captured by the microphone built into the smart device.

[0254] Next, the device is responsible for converting voice input into text data. This uses high-precision speech recognition software (for example, commonly used speech recognition APIs) to accurately generate text from speech. Noise reduction and other processes are performed in the background during this process to improve the accuracy of the conversion.

[0255] Furthermore, the terminal sends the converted text data to the server. The data is transmitted securely and then sent to the server for analysis. The server analyzes the received text data using natural language processing techniques (e.g., generative AI models). Here, the user's true intentions and requests are determined, and appropriate control procedures are generated.

[0256] The server then returns appropriately generated operating instructions to the terminal, which in turn controls the corresponding functions of the electronic device. For example, when playing music, the music app is automatically launched and the song is played. Simultaneously with the control, voice feedback about the operation result is also provided. An announcement such as "Music playback app launched. Playing the current playlist" is made, allowing the user to confirm that the operation is complete.

[0257] For example, if a user gives a voice command such as "Play music," the device converts the voice into text data, which is then analyzed by the server. Based on the analysis results, the server issues a command to launch the music app, and the device follows suit, launching the music app and starting playback.

[0258] An example of a prompt for a generative AI model would be: "Explain how the system should process a user's request to 'play music.'"

[0259] This system enables intuitive and effective operation of electronic devices using voice commands, significantly improving user convenience.

[0260] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0261] Step 1:

[0262] The user inputs specific commands to the smart device using voice. For example, they might say, "Play music." This voice input is the starting point of the program, and the smart device's built-in microphone captures the voice. The input data is the user's voice.

[0263] Step 2:

[0264] The device converts acquired voice input into text data using speech recognition software. Specifically, a high-precision speech recognition engine analyzes the voice signal and generates the corresponding text. The input is the user's voice data, and the output is text data.

[0265] Step 3:

[0266] The terminal sends the converted text data to the server. The data is encrypted and transferred securely, thus protecting user privacy. The input is text data, and the output is the completion of the data transmission to the server.

[0267] Step 4:

[0268] The server analyzes the received text data using natural language processing techniques. It uses a generative AI model to understand the intent of the text and generate specific control procedures related to the user's request. The input is text data, and the output is the generated operation procedure.

[0269] Step 5:

[0270] The server sends the generated operation procedure to the terminal. This prepares the terminal to perform the control requested by the user. The input is the operation procedure data, and the output is the completion of sending the instruction to the terminal.

[0271] Step 6:

[0272] The terminal performs operations on the user's electronic device based on the received instructions. Specifically, for example, a music playback application is automatically launched, and the specified operation (e.g., playing music) is performed. The input is instruction data from the server, and the output is the operating state of the device based on those instructions.

[0273] Step 7:

[0274] The device provides the user with voice feedback regarding the results of the operation and information about the next steps. This feedback allows the user to confirm that the operation was successful and to know what to do next. The input is data of the operation result, and the output is voice feedback.

[0275] (Application Example 1)

[0276] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0277] The challenge is to provide a system that supports elderly and other users unfamiliar with operating information devices, enabling them to easily operate these digital devices using voice commands and efficiently handle tasks, particularly those related to time management.

[0278] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following respective means.

[0279] In this invention, the server includes voice recognition means for receiving voice input from a user and converting it into text data, natural language processing means for analyzing the text data and determining the user's request, and control means for operating the user's information processing apparatus based on the determined request. Thereby, time management such as setting a reminder by voice by an elderly person becomes possible.

[0280] "Voice recognition means" has a function of receiving voice input from a user and converting the voice into text data.

[0281] "Natural language processing means" includes technologies for analyzing the text data converted by the voice recognition means and determining the user's request.

[0282] "Control means" has a function of operating the user's information processing apparatus based on the request determined by the natural language processing means.

[0283] "Feedback means" is provided with a function of notifying the user of the operation result by voice.

[0284] "Means for executing a time management function based on a voice request" is a technology for executing functions related to time management such as setting a reminder in response to a user's voice request.

[0285] In this invention, the system includes an implementation that combines speech recognition, natural language processing, and device control technologies. As speech recognition software, for example, the Google Speech-to-Text API can be used. Thereby, the voice input from the user is converted into text data. This text data is sent to a server on the cloud and analyzed on the server side using natural language processing algorithms such as IBM Watson Natural Language Understanding. Based on the analysis results, the generated instructions are sent to the user's information processing device through the control means, and the device is appropriately operated.

[0286] For example, when the user instructs the smartphone "Set a reminder to take medicine at 8 am tomorrow", the system processes this voice and sets a reminder. When this operation is completed, voice feedback "The reminder has been set" is provided by the feedback means.

[0287] The advantage of this system is that the elderly can intuitively operate digital devices through voice. Therefore, even users who are not good at operating can easily use the time management function.

[0288] Specific examples of the prompt sentences include the following.

[0289] [[ID=!15]] "When the user says in voice 'Tell me to take medicine at 7 am tomorrow', please explain how to interpret it and set a reminder."

[0290] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0291] Step 1:

[0292] The user makes a voice input towards the smart device. Here, the user gives an instruction in voice "Set a reminder to take medicine at 8 am tomorrow". The input is voice data.

[0293] Note: There seems to be a mistake in the original text where it says "7 am" in the prompt sentence example in , which is likely a typo and should probably be "7 am". The translation is done accordingly. Also, there was a `!` added in front of in the translation as it seems to be an error in the original numbering system, but the tag is preserved as per the instruction. Step 2:

[0294] The device receives voice input and converts the voice data into text data using the Google Speech-to-Text API. The output is text data. In this process, the voice waveform is mapped to text.

[0295] Step 3:

[0296] The terminal sends text data to a server in the cloud. The input is text data, and the output is the transmission of data to the server. At this stage, data transfer takes place over the network.

[0297] Step 4:

[0298] The server analyzes the received text data using IBM Watson Natural Language Understanding. The input is text data, and the output is the analysis result of the user request. In this process, natural language processing algorithms are used to understand the meaning of the text.

[0299] Step 5:

[0300] Based on the analysis results, the server determines the time management requirement and generates instructions for setting reminders. The input is the analysis results, and the output is instructions for the device. During this process, control commands for specific tasks are generated.

[0301] Step 6:

[0302] The server sends control commands to the terminal. The input is the control commands, and the output is the command sent to the terminal. Even at this stage, data is exchanged via the communication network.

[0303] Step 7:

[0304] Based on the command received by the terminal, execute the time management function and set the reminder. The input is the control command, and the output is the completion of the reminder setting. The app is launched within the terminal, and the set content is saved.

[0305] Step 8:

[0306] The terminal uses the feedback means to notify the user audibly that "the reminder has been set". The input is the completion information of the reminder setting, and the output is the audio feedback. The operation result is conveyed to the user in this final step.

[0307] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.

[0308] The present invention is an interactive system for operating a smart device based on voice input from a user, and further has a function of recognizing the user's emotional state and adjusting the operation and feedback of the device accordingly. Hereinafter, a specific implementation method of this system will be described.

[0309] First, the user inputs an operation instruction audibly towards the smart device. The terminal captures this voice, converts it into text data using the voice recognition means, and simultaneously uses the emotion engine to identify the user's emotion from the voice. The converted text data and the identified emotion data are transmitted from the terminal to the server.

[0310] The server analyzes text data using natural language processing to determine user requests. It also takes into account emotional information obtained by the emotion engine to determine the optimal response and action based on the user's current emotional state. Based on this determination, the server uses control mechanisms to instruct specific operations on the smart device and sets the method of providing feedback to the user. In particular, it adjusts the tone and content of the feedback based on emotional information to ensure the user feels as comfortable using the device as possible.

[0311] For example, if a user impatiently says "I want to take a picture" using voice input, the emotion engine will recognize that impatience. The server will recognize the request to "take a picture" and, taking into account the user's impatient emotional state, instruct the device to provide gentle feedback such as, "It's okay to calm down. I'll launch the camera now." This provides an emotionally empathetic response, reducing the stress the user experiences during the operation.

[0312] By incorporating user emotional states into the interaction in this way, the system becomes more intuitive and user-friendly, providing a particularly accessible learning and operating environment for senior users.

[0313] The following describes the processing flow.

[0314] Step 1:

[0315] The user provides voice input to a smart device, giving instructions for specific actions. This voice input may contain emotions.

[0316] Step 2:

[0317] The device receives the user's voice and converts it into text data using speech recognition technology. Simultaneously, it uses an emotion engine to identify the user's emotional state from the voice.

[0318] Step 3:

[0319] The terminal sends the converted text data and sentiment data to the server. This prepares the server to consider the user's request and sentiment simultaneously.

[0320] Step 4:

[0321] The server uses natural language processing to analyze text data and identify the user's specific requests. It determines actions such as "take a picture" and analyzes emotional information obtained from voice to understand the user's emotional state.

[0322] Step 5:

[0323] Based on the acquired requests and emotions, the server sends instructions to the terminal through control mechanisms. These instructions include specific operations on the smart device (e.g., launching the camera app) and emotionally sensitive feedback.

[0324] Step 6:

[0325] The device follows instructions from the server, performing specified device operations and providing emotionally responsive voice feedback to the user. For users experiencing frustration, it provides guidance in a calmer tone.

[0326] Step 7:

[0327] The user follows the instructions on the device and completes the necessary device operations. The device notifies the server of the results and provides guidance on the next steps as needed. By adjusting the guidance method according to the user's mood, a stress-free experience is provided.

[0328] (Example 2)

[0329] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0330] Conventional voice control systems process information and provide feedback without considering the user's emotional state, resulting in a uniform user experience and a particular difficulty in appropriately responding to emotionally charged voice input. Therefore, there is an urgent need to realize an interactive system that can provide appropriate operation and feedback in accordance with the user's emotions.

[0331] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0332] In this invention, the server includes speech recognition means for receiving voice input and converting it into text data, emotion recognition means for identifying the user's emotional state from the voice input, natural language processing means for analyzing the text data and determining the user's request, and feedback adjustment means for adjusting the tone and content of the feedback. This enables adaptive operation and feedback that takes the user's emotional state into consideration.

[0333] "Voice recognition means" refers to technology that digitizes voice input from a user and converts that voice information into text data. It includes devices and programs that analyze voice patterns and convert them into corresponding strings of characters.

[0334] "Emotion recognition means" refers to technology that identifies a user's emotions from voice input. This includes algorithms and devices for analyzing voice characteristics and inferring emotional states.

[0335] "Natural language processing means" refers to technologies for analyzing character data and understanding user requests and intentions. This includes methods for appropriately interpreting input strings and guiding necessary responses and actions.

[0336] "Feedback adjustment techniques" refer to technologies that adjust the tone and content of feedback based on the user's emotional state. They are means of improving the user experience by generating appropriate responses and taking emotions into consideration.

[0337] "Control means" refers to technology that operates the user's electronic devices based on determined requests. This includes functions that instruct appropriate actions based on data processing.

[0338] This system is an interactive system that performs device operations that take emotions into consideration based on the user's voice input. When the user gives operation instructions by voice, the terminal captures the voice. The device's microphone is used for voice capture, and the voice signal is converted into digital data.

[0339] The device converts this audio data into text data using speech recognition software (e.g., a general-purpose speech recognition API). Simultaneously, it uses an emotion engine (e.g., a common emotion analysis API) to identify the user's emotions from the audio. The resulting text data and emotion data are then sent from the device to the server.

[0340] The server uses natural language processing techniques (e.g., general-purpose natural language processing models) to analyze the received text data and understand the user's request. It also considers sentiment data and determines the optimal response or action based on the user's emotional state. In particular, it adjusts the content and tone of feedback based on sentiment information to ensure that what is communicated to the user is appropriate.

[0341] For example, if a user angrily commands "Play some music," the server will take that anger into consideration and instruct the device to provide considerate feedback such as "Playing some relaxing music."

[0342] An example of a prompt that utilizes a generative AI model is an instruction such as, "When the user becomes excited and gives a voice command, generate feedback to soothe their emotions." This system configuration enables device operation that is sensitive to the user's emotions.

[0343] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0344] Step 1:

[0345] The user directs their voice input towards the smart device. The device captures the voice through its microphone and receives it as digital audio data. This digital audio data becomes the input.

[0346] Step 2:

[0347] The terminal uses speech recognition software to convert digital speech data into text data. The input is digital speech data, and the output is text data. The software analyzes the speech waveform and generates the corresponding string of characters.

[0348] Step 3:

[0349] The device uses an emotion engine to analyze the user's emotions from their voice. The input is digital voice data, and the output is identified emotion data. This process estimates the user's emotions by evaluating characteristics such as voice tone and speed.

[0350] Step 4:

[0351] The terminal sends converted text data and sentiment data to the server. The input is text data and sentiment data, and the output is the data sent to the server. The data is securely transmitted to the server over the network.

[0352] Step 5:

[0353] The server uses natural language processing technology to analyze text data and determine user requests. The input is text data, and the output is the interpreted user request. It analyzes the structure and meaning of the text to identify the request.

[0354] Step 6:

[0355] The server considers emotional data and uses feedback adjustment mechanisms to determine the most appropriate response. The input is emotional data and the interpreted request, while the output is the adjusted feedback content. It selects a tone and content appropriate to the emotional state.

[0356] Step 7:

[0357] The server sends the determined feedback content to the terminal, which then provides feedback to the user via audio or screen. The input is the feedback content, and the output is the display of feedback to the user or audio playback. This ensures that an appropriate response is given to the user's input.

[0358] (Application Example 2)

[0359] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0360] The aim is to create a more natural and intuitive interface by providing device responses that take into account the user's emotional state when they operate the device via voice input. In particular, in information processing devices used in the home, there is a challenge in reducing user stress and improving the user experience by enabling feedback that responds to the user's emotions.

[0361] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0362] In this invention, the server includes speech recognition means, natural language processing means, and emotion recognition means. This makes it possible to convert speech input into text data and simultaneously identify requests and emotions.

[0363] "Voice recognition means" refers to a device or technology that receives voice input from a user and converts it into text data.

[0364] "Natural language processing means" refers to technologies that have the ability to analyze text data and determine user requests.

[0365] "Emotion recognition means" refers to technologies and devices that identify a user's emotional state from voice input.

[0366] An "information processing device" refers to an electronic device that controls its operation based on the user's requests and emotional state.

[0367] "Control means" refers to means for operating an information processing device based on judged requests and identified emotional states.

[0368] A "feedback mechanism" is a technology that provides users with information in voice that corresponds to the results of their actions and their emotions.

[0369] This invention relates to a method for interactively operating an information processing device using speech recognition means, natural language processing means, emotion recognition means, and feedback means. The system is intended for use in the home and aims to provide personalized feedback based on the user's emotional state.

[0370] First, the user inputs voice data into the information processing device via a microphone. The device then converts this voice into text data using speech recognition software such as the Google Speech API.

[0371] Next, the converted text is analyzed by natural language processing tools to understand the user's request. This process extracts specific operation instructions from the text data.

[0372] In parallel, voice input is analyzed through emotion recognition mechanisms to identify the user's emotional state. This emotion analysis utilizes emotion recognition models built with TensorFlow or PyTorch.

[0373] The server controls the operation of the information processing device by considering requests obtained through natural language processing and emotions obtained through emotion recognition. In particular, it can adjust the tone and content of responses provided through feedback mechanisms according to the user's emotional state.

[0374] For example, if a user asks in a tense voice, "What's the weather forecast?", the system will generate a gentle and reassuring response such as, "It's going to be sunny today. It'll be warm during the day, so please relax."

[0375] An example of a prompt for a generative AI model might be, "Please think of a reassuring response for a user who is feeling anxious." In this way, the system supports the user's actions by being sensitive to their emotions.

[0376] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0377] Step 1:

[0378] The terminal receives voice input from the user through the microphone and acquires that voice data. It obtains an audio signal as input and proceeds to process it as digital data.

[0379] Step 2:

[0380] The device processes the acquired audio data using speech recognition technology and converts it into text data. Specifically, it uses APIs such as Google Speech to convert digital audio signals into natural language text. The output is text data indicating the request.

[0381] Step 3:

[0382] The device analyzes voice data using emotion recognition technology to identify the user's emotional state based on their voice tone, pitch, and speed. It utilizes emotion recognition models built with TensorFlow or PyTorch for analysis. The output is data including emotion labels.

[0383] Step 4:

[0384] The server analyzes text data using natural language processing to understand user requests. The input is text data converted by the terminal, and the server analyzes the request content to identify operation commands. The output is information containing these operation commands.

[0385] Step 5:

[0386] The server determines the operation of the information processing device based on the acquired request and sentiment information. It combines the obtained operation commands and sentiment data to determine the optimal operation and sends a control signal to the terminal.

[0387] Step 6:

[0388] The terminal receives control signals from the server and operates the information processing device. Specific actions include launching applications and executing specific functions.

[0389] Step 7:

[0390] The device generates voice feedback for the user, responding with a tone and content that matches the user's emotional state. For example, if the user is feeling anxious, it will play reassuring content in a gentle voice through the speaker. The output is a voice response that is attentive to the user's emotions.

[0391] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0392] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0393] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0394] [Third Embodiment]

[0395] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0396] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0397] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0398] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0399] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0400] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0401] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0402] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0403] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0404] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0405] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0406] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0407] This invention relates to an interactive system that assists in the operation of smart devices based on voice input from the user. In particular, it provides a technology that combines voice recognition, natural language processing, and device control to enable senior users to easily operate digital devices.

[0408] First, the user requests a specific action by voice to their smart device. This voice input is captured by the device and converted into text data using speech recognition software within the device. This text data is then sent from the device to a server to understand the user's request.

[0409] The server analyzes the received text data using a natural language processing algorithm. This analysis determines what the user wants and generates a set of instructions to address that request. The server then sends these instructions to the terminal, providing specific control commands for the smart device.

[0410] Based on instructions from the server, the terminal launches the appropriate application or performs specific device functions. The terminal also uses feedback mechanisms to provide the user with voice guidance regarding the results of the operation and the next steps. This allows the user to use the device with confidence while confirming the operating procedures.

[0411] As a concrete example, consider the case where a user voice-inputs "I want to take a picture." The device converts this voice into text data and sends it to the server. The server recognizes the "take a picture" request and sends a command to launch the camera app to the device. The device opens the camera app and supports the user's operation by providing voice guidance such as, "The camera has been launched. Please press the shutter button."

[0412] This system will make it easier for senior users who were previously apprehensive about using digital devices to use them.

[0413] The following describes the processing flow.

[0414] Step 1:

[0415] The user provides specific operational instructions to the smart device via voice. For example, they might say, "I want to take a picture."

[0416] Step 2:

[0417] The terminal receives voice input from the user and uses its built-in speech recognition software to convert this voice into text data. It then sends this text data to the server.

[0418] Step 3:

[0419] The server uses natural language processing algorithms to analyze the user's request in order to parse the received text data. It identifies the instruction "Take a picture" and generates the corresponding procedure.

[0420] Step 4:

[0421] The server sends the generated operating procedure to the terminal, instructing it on the specific commands it should execute.

[0422] Step 5:

[0423] The device follows instructions from the server and launches the camera app. It also provides voice guidance to the user, informing them that the camera has been launched and prompting them to "Press the shutter button."

[0424] Step 6:

[0425] The user follows the instructions on their device and takes a photo by pressing the shutter button in the camera app.

[0426] Step 7:

[0427] The device detects when a photo has been taken and saves the photo to the gallery. It then provides voice feedback to the user when the saving process is complete.

[0428] (Example 1)

[0429] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0430] Currently, many users, especially the elderly, experience confusion and anxiety when using electronic devices. This limits the use of digital devices and diminishes user convenience. In addition, there are challenges in terms of accuracy and ease of use when controlling devices using voice input.

[0431] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0432] In this invention, the server includes a voice processing means that receives voice input and converts it into text data, a language analysis means that analyzes the text data to determine the user's request, and a control mechanism that operates the user's electronic device. This enables intuitive and simple operation of the electronic device via voice input.

[0433] "Voice processing means" refers to a device or software that has the function of receiving voice input from a user and converting it into analyzable text data.

[0434] "Language analysis means" refers to an element that has the function of analyzing text data and understanding and judging the content of the user's request, and includes algorithms for natural language processing.

[0435] A "control mechanism" is a system for specifically operating the applications and functions of an electronic device based on user requirements.

[0436] A "response-providing means" is a device or software that has the function of providing feedback to the user, such as the result of an operation or the next step in an operation, via voice or other means.

[0437] "Electronic devices" refer to devices capable of performing digital processing, such as smartphones, tablets, and smart speakers.

[0438] This invention is a system that improves the convenience of users operating electronic devices by voice. Designed to allow a wide range of users, including seniors, to intuitively operate the device, it is achieved through the following main components:

[0439] First, the user gives instructions to the device using voice commands. For example, they can issue specific commands by voice, such as "Play music" or "I want to take a picture." The voice is captured by the microphone built into the smart device.

[0440] Next, the device is responsible for converting voice input into text data. This uses high-precision speech recognition software (for example, commonly used speech recognition APIs) to accurately generate text from speech. Noise reduction and other processes are performed in the background during this process to improve the accuracy of the conversion.

[0441] Furthermore, the terminal sends the converted text data to the server. The data is transmitted securely and then sent to the server for analysis. The server analyzes the received text data using natural language processing techniques (e.g., generative AI models). Here, the user's true intentions and requests are determined, and appropriate control procedures are generated.

[0442] The server then returns appropriately generated operating instructions to the terminal, which in turn controls the corresponding functions of the electronic device. For example, when playing music, the music app is automatically launched and the song is played. Simultaneously with the control, voice feedback about the operation result is also provided. An announcement such as "Music playback app launched. Playing the current playlist" is made, allowing the user to confirm that the operation is complete.

[0443] For example, if a user gives a voice command such as "Play music," the device converts the voice into text data, which is then analyzed by the server. Based on the analysis results, the server issues a command to launch the music app, and the device follows suit, launching the music app and starting playback.

[0444] An example of a prompt for a generative AI model would be: "Explain how the system should process a user's request to 'play music.'"

[0445] This system enables intuitive and effective operation of electronic devices using voice commands, significantly improving user convenience.

[0446] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0447] Step 1:

[0448] The user inputs specific commands to the smart device using voice. For example, they might say, "Play music." This voice input is the starting point of the program, and the smart device's built-in microphone captures the voice. The input data is the user's voice.

[0449] Step 2:

[0450] The device converts acquired voice input into text data using speech recognition software. Specifically, a high-precision speech recognition engine analyzes the voice signal and generates the corresponding text. The input is the user's voice data, and the output is text data.

[0451] Step 3:

[0452] The terminal sends the converted text data to the server. The data is encrypted and transferred securely, thus protecting user privacy. The input is text data, and the output is the completion of the data transmission to the server.

[0453] Step 4:

[0454] The server analyzes the received text data using natural language processing techniques. It uses a generative AI model to understand the intent of the text and generate specific control procedures related to the user's request. The input is text data, and the output is the generated operation procedure.

[0455] Step 5:

[0456] The server sends the generated operation procedure to the terminal. This prepares the terminal to perform the control requested by the user. The input is the operation procedure data, and the output is the completion of sending the instruction to the terminal.

[0457] Step 6:

[0458] The terminal performs operations on the user's electronic device based on the received instructions. Specifically, for example, a music playback application is automatically launched, and the specified operation (e.g., playing music) is performed. The input is instruction data from the server, and the output is the operating state of the device based on those instructions.

[0459] Step 7:

[0460] The device provides the user with voice feedback regarding the results of the operation and information about the next steps. This feedback allows the user to confirm that the operation was successful and to know what to do next. The input is data of the operation result, and the output is voice feedback.

[0461] (Application Example 1)

[0462] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0463] The challenge is to provide a system that supports elderly and other users unfamiliar with operating information devices, enabling them to easily operate these digital devices using voice commands and efficiently handle tasks, particularly those related to time management.

[0464] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0465] In this invention, the server includes speech recognition means that receive voice input from the user and convert it into text data, natural language processing means that analyze the text data and determine the user's request, and control means that operate the user's information processing device based on the determined request. This enables elderly people to manage their time, such as setting reminders by voice.

[0466] A "speech recognition means" is a device that receives speech input from a user and converts that speech into text data.

[0467] "Natural language processing means" includes technologies for analyzing text data converted by speech recognition means and determining user requests.

[0468] "Control means" refers to a device that has the function of operating the user's information processing device based on requests determined by natural language processing means.

[0469] A "feedback mechanism" is a function that notifies the user of the results of an operation via voice.

[0470] "Means for executing time management functions based on voice requests" refers to technology for executing time management functions, such as setting reminders, in response to a user's voice requests.

[0471] This invention includes an implementation of a system that combines speech recognition, natural language processing, and device control technologies. For example, the Google Speech-to-Text API can be used as speech recognition software. This converts voice input from the user into text data. This text data is sent to a server in the cloud, where it is analyzed using natural language processing algorithms such as IBM Watson Natural Language Understanding. Based on the analysis results, generated instructions are sent to the user's information processing device via a control means, and the device is operated appropriately.

[0472] For example, if a user says to their smartphone, "Set a reminder to take my medicine at 8 AM tomorrow," the system will process this voice and set a reminder. Once this operation is complete, a voice confirmation message saying, "Reminder set," will be provided via a feedback mechanism.

[0473] The advantage of this system lies in its ability to allow elderly users to intuitively operate digital devices through voice commands. This makes it easy for even users who are not tech-savvy to utilize time management functions.

[0474] Examples of prompt statements include the following:

[0475] "Please explain how you would interpret a user's voice message, 'Remind me to take my medicine at 7 AM tomorrow,' and how you would set a reminder."

[0476] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0477] Step 1:

[0478] The user makes a voice input to their smart device. Here, the user gives a voice command such as, "Set up a reminder to take my medicine tomorrow at 8 AM." The input is voice data.

[0479] Step 2:

[0480] The device receives voice input and converts the voice data into text data using the Google Speech-to-Text API. The output is text data. In this process, the voice waveform is mapped to text.

[0481] Step 3:

[0482] The terminal sends text data to a server in the cloud. The input is text data, and the output is the transmission of data to the server. At this stage, data transfer takes place over the network.

[0483] Step 4:

[0484] The server analyzes the received text data using IBM Watson Natural Language Understanding. The input is text data, and the output is the analysis result of the user request. In this process, natural language processing algorithms are used to understand the meaning of the text.

[0485] Step 5:

[0486] Based on the analysis results, the server determines the time management requirement and generates instructions for setting reminders. The input is the analysis results, and the output is instructions for the device. During this process, control commands for specific tasks are generated.

[0487] Step 6:

[0488] The server sends control commands to the terminal. The input is the control commands, and the output is the command sent to the terminal. Even at this stage, data is exchanged via the communication network.

[0489] Step 7:

[0490] Based on the commands received by the device, it executes time management functions and sets reminders. The input is a control command, and the output is the completion of the reminder setting. The app launches on the device and the settings are saved.

[0491] Step 8:

[0492] The device uses a feedback mechanism to notify the user via voice, "Reminder set." The input is the completion information for setting the reminder, and the output is the voice feedback. In this final step, the user is informed of the result of the operation.

[0493] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0494] This invention is an interactive system for operating a smart device based on voice input from a user, and further possesses the function of recognizing the user's emotional state and adjusting the device operation and feedback accordingly. The specific implementation method of this system will be described below.

[0495] First, the user inputs operating instructions by voice into the smart device. The device captures this voice, converts it into text data using voice recognition, and simultaneously uses an emotion engine to identify the user's emotions from the voice. The converted text data and identified emotion data are sent from the device to the server.

[0496] The server analyzes text data using natural language processing to determine user requests. It also takes into account emotional information obtained by the emotion engine to determine the optimal response and action based on the user's current emotional state. Based on this determination, the server uses control mechanisms to instruct specific operations on the smart device and sets the method of providing feedback to the user. In particular, it adjusts the tone and content of the feedback based on emotional information to ensure the user feels as comfortable using the device as possible.

[0497] For example, if a user impatiently says "I want to take a picture" using voice input, the emotion engine will recognize that impatience. The server will recognize the request to "take a picture" and, taking into account the user's impatient emotional state, instruct the device to provide gentle feedback such as, "It's okay to calm down. I'll launch the camera now." This provides an emotionally empathetic response, reducing the stress the user experiences during the operation.

[0498] By incorporating user emotional states into the interaction in this way, the system becomes more intuitive and user-friendly, providing a particularly accessible learning and operating environment for senior users.

[0499] The following describes the processing flow.

[0500] Step 1:

[0501] The user provides voice input to a smart device, giving instructions for specific actions. This voice input may contain emotions.

[0502] Step 2:

[0503] The device receives the user's voice and converts it into text data using speech recognition technology. Simultaneously, it uses an emotion engine to identify the user's emotional state from the voice.

[0504] Step 3:

[0505] The terminal sends the converted text data and sentiment data to the server. This prepares the server to consider the user's request and sentiment simultaneously.

[0506] Step 4:

[0507] The server uses natural language processing to analyze text data and identify the user's specific requests. It determines actions such as "take a picture" and analyzes emotional information obtained from voice to understand the user's emotional state.

[0508] Step 5:

[0509] Based on the acquired requests and emotions, the server sends instructions to the terminal through control mechanisms. These instructions include specific operations on the smart device (e.g., launching the camera app) and emotionally sensitive feedback.

[0510] Step 6:

[0511] The device follows instructions from the server, performing specified device operations and providing emotionally responsive voice feedback to the user. For users experiencing frustration, it provides guidance in a calmer tone.

[0512] Step 7:

[0513] The user follows the instructions on the device and completes the necessary device operations. The device notifies the server of the results and provides guidance on the next steps as needed. By adjusting the guidance method according to the user's mood, a stress-free experience is provided.

[0514] (Example 2)

[0515] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0516] Conventional voice control systems process information and provide feedback without considering the user's emotional state, resulting in a uniform user experience and a particular difficulty in appropriately responding to emotionally charged voice input. Therefore, there is an urgent need to realize an interactive system that can provide appropriate operation and feedback in accordance with the user's emotions.

[0517] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0518] In this invention, the server includes speech recognition means for receiving voice input and converting it into text data, emotion recognition means for identifying the user's emotional state from the voice input, natural language processing means for analyzing the text data and determining the user's request, and feedback adjustment means for adjusting the tone and content of the feedback. This enables adaptive operation and feedback that takes the user's emotional state into consideration.

[0519] "Voice recognition means" refers to technology that digitizes voice input from a user and converts that voice information into text data. It includes devices and programs that analyze voice patterns and convert them into corresponding strings of characters.

[0520] "Emotion recognition means" refers to technology that identifies a user's emotions from voice input. This includes algorithms and devices for analyzing voice characteristics and inferring emotional states.

[0521] "Natural language processing means" refers to technologies for analyzing character data and understanding user requests and intentions. This includes methods for appropriately interpreting input strings and guiding necessary responses and actions.

[0522] "Feedback adjustment techniques" refer to technologies that adjust the tone and content of feedback based on the user's emotional state. They are means of improving the user experience by generating appropriate responses and taking emotions into consideration.

[0523] "Control means" refers to technology that operates the user's electronic devices based on determined requests. This includes functions that instruct appropriate actions based on data processing.

[0524] This system is an interactive system that performs emotionally responsive device operation based on user voice input. When the user gives voice commands, the terminal captures the voice. The device's microphone is used for voice capture, and the voice signal is converted into digital data.

[0525] The device converts this audio data into text data using speech recognition software (e.g., a general-purpose speech recognition API). Simultaneously, it uses an emotion engine (e.g., a common emotion analysis API) to identify the user's emotions from the audio. The resulting text data and emotion data are then sent from the device to the server.

[0526] The server uses natural language processing techniques (e.g., general-purpose natural language processing models) to analyze the received text data and understand the user's request. It also considers sentiment data and determines the optimal response or action based on the user's emotional state. In particular, it adjusts the content and tone of feedback based on sentiment information to ensure that what is communicated to the user is appropriate.

[0527] For example, if a user angrily commands "Play some music," the server will take that anger into consideration and instruct the device to provide considerate feedback such as "Playing some relaxing music."

[0528] An example of a prompt that utilizes a generative AI model is an instruction such as, "When the user becomes excited and gives a voice command, generate feedback to soothe their emotions." This system configuration enables device operation that is sensitive to the user's emotions.

[0529] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0530] Step 1:

[0531] The user directs their voice input towards the smart device. The device captures the voice through its microphone and receives it as digital audio data. This digital audio data becomes the input.

[0532] Step 2:

[0533] The terminal uses speech recognition software to convert digital speech data into text data. The input is digital speech data, and the output is text data. The software analyzes the speech waveform and generates the corresponding string of characters.

[0534] Step 3:

[0535] The device uses an emotion engine to analyze the user's emotions from their voice. The input is digital voice data, and the output is identified emotion data. This process estimates the user's emotions by evaluating characteristics such as voice tone and speed.

[0536] Step 4:

[0537] The terminal sends converted text data and sentiment data to the server. The input is text data and sentiment data, and the output is the data sent to the server. The data is securely transmitted to the server over the network.

[0538] Step 5:

[0539] The server uses natural language processing technology to analyze text data and determine user requests. The input is text data, and the output is the interpreted user request. It analyzes the structure and meaning of the text to identify the request.

[0540] Step 6:

[0541] The server considers emotional data and uses feedback adjustment mechanisms to determine the most appropriate response. The input is emotional data and the interpreted request, while the output is the adjusted feedback content. It selects a tone and content appropriate to the emotional state.

[0542] Step 7:

[0543] The server sends the determined feedback content to the terminal, which then provides feedback to the user via audio or screen. The input is the feedback content, and the output is the display of feedback to the user or audio playback. This ensures that an appropriate response is given to the user's input.

[0544] (Application Example 2)

[0545] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0546] The aim is to create a more natural and intuitive interface by providing device responses that take into account the user's emotional state when they operate the device via voice input. In particular, in information processing devices used in the home, there is a challenge in reducing user stress and improving the user experience by enabling feedback that responds to the user's emotions.

[0547] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0548] In this invention, the server includes speech recognition means, natural language processing means, and emotion recognition means. This makes it possible to convert speech input into text data and simultaneously identify requests and emotions.

[0549] "Voice recognition means" refers to a device or technology that receives voice input from a user and converts it into text data.

[0550] "Natural language processing means" refers to technologies that have the ability to analyze text data and determine user requests.

[0551] "Emotion recognition means" refers to technologies and devices that identify a user's emotional state from voice input.

[0552] An "information processing device" refers to an electronic device that controls its operation based on the user's requests and emotional state.

[0553] "Control means" refers to means for operating an information processing device based on judged requests and identified emotional states.

[0554] A "feedback mechanism" is a technology that provides users with information in voice that corresponds to the results of their actions and their emotions.

[0555] This invention relates to a method for interactively operating an information processing device using speech recognition means, natural language processing means, emotion recognition means, and feedback means. The system is intended for use in the home and aims to provide personalized feedback based on the user's emotional state.

[0556] First, the user inputs voice data into the information processing device via a microphone. The device then converts this voice into text data using speech recognition software such as the Google Speech API.

[0557] Next, the converted text is analyzed by natural language processing tools to understand the user's request. This process extracts specific operation instructions from the text data.

[0558] In parallel, voice input is analyzed through emotion recognition mechanisms to identify the user's emotional state. This emotion analysis utilizes emotion recognition models built with TensorFlow or PyTorch.

[0559] The server controls the operation of the information processing device by considering requests obtained through natural language processing and emotions obtained through emotion recognition. In particular, it can adjust the tone and content of responses provided through feedback mechanisms according to the user's emotional state.

[0560] For example, if a user asks in a tense voice, "What's the weather forecast?", the system will generate a gentle and reassuring response such as, "It's going to be sunny today. It'll be warm during the day, so please relax."

[0561] An example of a prompt for a generative AI model might be, "Please think of a reassuring response for a user who is feeling anxious." In this way, the system supports the user's actions by being sensitive to their emotions.

[0562] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0563] Step 1:

[0564] The terminal receives voice input from the user through the microphone and acquires that voice data. It obtains an audio signal as input and proceeds to process it as digital data.

[0565] Step 2:

[0566] The device processes the acquired audio data using speech recognition technology and converts it into text data. Specifically, it uses APIs such as Google Speech to convert digital audio signals into natural language text. The output is text data indicating the request.

[0567] Step 3:

[0568] The device analyzes voice data using emotion recognition technology to identify the user's emotional state based on their voice tone, pitch, and speed. It utilizes emotion recognition models built with TensorFlow or PyTorch for analysis. The output is data including emotion labels.

[0569] Step 4:

[0570] The server analyzes text data using natural language processing to understand user requests. The input is text data converted by the terminal, and the server analyzes the request content to identify operation commands. The output is information containing these operation commands.

[0571] Step 5:

[0572] The server determines the operation of the information processing device based on the acquired request and sentiment information. It combines the obtained operation commands and sentiment data to determine the optimal operation and sends a control signal to the terminal.

[0573] Step 6:

[0574] The terminal receives control signals from the server and operates the information processing device. Specific actions include launching applications and executing specific functions.

[0575] Step 7:

[0576] The device generates voice feedback for the user, responding with a tone and content that matches the user's emotional state. For example, if the user is feeling anxious, it will play reassuring content in a gentle voice through the speaker. The output is a voice response that is attentive to the user's emotions.

[0577] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0578] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0579] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0580] [Fourth Embodiment]

[0581] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0582] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0583] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0584] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0585] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0586] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0587] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0588] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0589] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0590] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0591] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0592] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0593] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0594] This invention relates to an interactive system that assists in the operation of smart devices based on voice input from the user. In particular, it provides a technology that combines voice recognition, natural language processing, and device control to enable senior users to easily operate digital devices.

[0595] First, the user requests a specific action by voice to their smart device. This voice input is captured by the device and converted into text data using speech recognition software within the device. This text data is then sent from the device to a server to understand the user's request.

[0596] The server analyzes the received text data using a natural language processing algorithm. This analysis determines what the user wants and generates a set of instructions to address that request. The server then sends these instructions to the terminal, providing specific control commands for the smart device.

[0597] Based on instructions from the server, the terminal launches the appropriate application or performs specific device functions. The terminal also uses feedback mechanisms to provide the user with voice guidance regarding the results of the operation and the next steps. This allows the user to use the device with confidence while confirming the operating procedures.

[0598] As a concrete example, consider the case where a user voice-inputs "I want to take a picture." The device converts this voice into text data and sends it to the server. The server recognizes the "take a picture" request and sends a command to launch the camera app to the device. The device opens the camera app and supports the user's operation by providing voice guidance such as, "The camera has been launched. Please press the shutter button."

[0599] This system will make it easier for senior users who were previously apprehensive about using digital devices to use them.

[0600] The following describes the processing flow.

[0601] Step 1:

[0602] The user provides specific operational instructions to the smart device via voice. For example, they might say, "I want to take a picture."

[0603] Step 2:

[0604] The terminal receives voice input from the user and uses its built-in speech recognition software to convert this voice into text data. It then sends this text data to the server.

[0605] Step 3:

[0606] The server uses natural language processing algorithms to analyze the user's request in order to parse the received text data. It identifies the instruction "Take a picture" and generates the corresponding procedure.

[0607] Step 4:

[0608] The server sends the generated operating procedure to the terminal, instructing it on the specific commands it should execute.

[0609] Step 5:

[0610] The device follows instructions from the server and launches the camera app. It also provides voice guidance to the user, informing them that the camera has been launched and prompting them to "Press the shutter button."

[0611] Step 6:

[0612] The user follows the instructions on their device and takes a photo by pressing the shutter button in the camera app.

[0613] Step 7:

[0614] The device detects when a photo has been taken and saves the photo to the gallery. It then provides voice feedback to the user when the saving process is complete.

[0615] (Example 1)

[0616] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0617] Currently, many users, especially the elderly, experience confusion and anxiety when using electronic devices. This limits the use of digital devices and diminishes user convenience. In addition, there are challenges in terms of accuracy and ease of use when controlling devices using voice input.

[0618] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0619] In this invention, the server includes a voice processing means that receives voice input and converts it into text data, a language analysis means that analyzes the text data to determine the user's request, and a control mechanism that operates the user's electronic device. This enables intuitive and simple operation of the electronic device via voice input.

[0620] "Voice processing means" refers to a device or software that has the function of receiving voice input from a user and converting it into analyzable text data.

[0621] "Language analysis means" refers to an element that has the function of analyzing text data and understanding and judging the content of the user's request, and includes algorithms for natural language processing.

[0622] A "control mechanism" is a system for specifically operating the applications and functions of an electronic device based on user requirements.

[0623] A "response-providing means" is a device or software that has the function of providing feedback to the user, such as the result of an operation or the next step in an operation, via voice or other means.

[0624] "Electronic devices" refer to devices capable of performing digital processing, such as smartphones, tablets, and smart speakers.

[0625] This invention is a system that improves the convenience of users operating electronic devices by voice. Designed to allow a wide range of users, including seniors, to intuitively operate the device, it is achieved through the following main components:

[0626] First, the user gives instructions to the device using voice commands. For example, they can issue specific commands by voice, such as "Play music" or "I want to take a picture." The voice is captured by the microphone built into the smart device.

[0627] Next, the device is responsible for converting voice input into text data. This uses high-precision speech recognition software (for example, commonly used speech recognition APIs) to accurately generate text from speech. Noise reduction and other processes are performed in the background during this process to improve the accuracy of the conversion.

[0628] Furthermore, the terminal sends the converted text data to the server. The data is transmitted securely and then sent to the server for analysis. The server analyzes the received text data using natural language processing techniques (e.g., generative AI models). Here, the user's true intentions and requests are determined, and appropriate control procedures are generated.

[0629] The server then returns appropriately generated operating instructions to the terminal, which in turn controls the corresponding functions of the electronic device. For example, when playing music, the music app is automatically launched and the song is played. Simultaneously with the control, voice feedback about the operation result is also provided. An announcement such as "Music playback app launched. Playing the current playlist" is made, allowing the user to confirm that the operation is complete.

[0630] For example, if a user gives a voice command such as "Play music," the device converts the voice into text data, which is then analyzed by the server. Based on the analysis results, the server issues a command to launch the music app, and the device follows suit, launching the music app and starting playback.

[0631] An example of a prompt for a generative AI model would be: "Explain how the system should process a user's request to 'play music.'"

[0632] This system enables intuitive and effective operation of electronic devices using voice commands, significantly improving user convenience.

[0633] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0634] Step 1:

[0635] The user inputs specific commands to the smart device using voice. For example, they might say, "Play music." This voice input is the starting point of the program, and the smart device's built-in microphone captures the voice. The input data is the user's voice.

[0636] Step 2:

[0637] The device converts acquired voice input into text data using speech recognition software. Specifically, a high-precision speech recognition engine analyzes the voice signal and generates the corresponding text. The input is the user's voice data, and the output is text data.

[0638] Step 3:

[0639] The terminal sends the converted text data to the server. The data is encrypted and transferred securely, thus protecting user privacy. The input is text data, and the output is the completion of the data transmission to the server.

[0640] Step 4:

[0641] The server analyzes the received text data using natural language processing techniques. It uses a generative AI model to understand the intent of the text and generate specific control procedures related to the user's request. The input is text data, and the output is the generated operation procedure.

[0642] Step 5:

[0643] The server sends the generated operation procedure to the terminal. This prepares the terminal to perform the control requested by the user. The input is the operation procedure data, and the output is the completion of sending the instruction to the terminal.

[0644] Step 6:

[0645] The terminal performs operations on the user's electronic device based on the received instructions. Specifically, for example, a music playback application is automatically launched, and the specified operation (e.g., playing music) is performed. The input is instruction data from the server, and the output is the operating state of the device based on those instructions.

[0646] Step 7:

[0647] The device provides the user with voice feedback regarding the results of the operation and information about the next steps. This feedback allows the user to confirm that the operation was successful and to know what to do next. The input is data of the operation result, and the output is voice feedback.

[0648] (Application Example 1)

[0649] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0650] The challenge is to provide a system that supports elderly and other users unfamiliar with operating information devices, enabling them to easily operate these digital devices using voice commands and efficiently handle tasks, particularly those related to time management.

[0651] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0652] In this invention, the server includes speech recognition means that receive voice input from the user and convert it into text data, natural language processing means that analyze the text data and determine the user's request, and control means that operate the user's information processing device based on the determined request. This enables elderly people to manage their time, such as setting reminders by voice.

[0653] A "speech recognition means" is a device that receives speech input from a user and converts that speech into text data.

[0654] "Natural language processing means" includes technologies for analyzing text data converted by speech recognition means and determining user requests.

[0655] "Control means" refers to a device that has the function of operating the user's information processing device based on requests determined by natural language processing means.

[0656] A "feedback mechanism" is a function that notifies the user of the results of an operation via voice.

[0657] "Means for executing time management functions based on voice requests" refers to technology for executing time management functions, such as setting reminders, in response to a user's voice requests.

[0658] This invention includes an implementation of a system that combines speech recognition, natural language processing, and device control technologies. For example, the Google Speech-to-Text API can be used as speech recognition software. This converts voice input from the user into text data. This text data is sent to a server in the cloud, where it is analyzed using natural language processing algorithms such as IBM Watson Natural Language Understanding. Based on the analysis results, generated instructions are sent to the user's information processing device via a control means, and the device is operated appropriately.

[0659] For example, if a user says to their smartphone, "Set a reminder to take my medicine at 8 AM tomorrow," the system will process this voice and set a reminder. Once this operation is complete, a voice confirmation message saying, "Reminder set," will be provided via a feedback mechanism.

[0660] The advantage of this system lies in its ability to allow elderly users to intuitively operate digital devices through voice commands. This makes it easy for even users who are not tech-savvy to utilize time management functions.

[0661] Examples of prompt statements include the following:

[0662] "Please explain how you would interpret a user's voice message, 'Remind me to take my medicine at 7 AM tomorrow,' and how you would set a reminder."

[0663] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0664] Step 1:

[0665] The user makes a voice input to their smart device. Here, the user gives a voice command such as, "Set up a reminder to take my medicine tomorrow at 8 AM." The input is voice data.

[0666] Step 2:

[0667] The device receives voice input and converts the voice data into text data using the Google Speech-to-Text API. The output is text data. In this process, the voice waveform is mapped to text.

[0668] Step 3:

[0669] The terminal sends text data to a server in the cloud. The input is text data, and the output is the transmission of data to the server. At this stage, data transfer takes place over the network.

[0670] Step 4:

[0671] The server analyzes the received text data using IBM Watson Natural Language Understanding. The input is text data, and the output is the analysis result of the user request. In this process, natural language processing algorithms are used to understand the meaning of the text.

[0672] Step 5:

[0673] Based on the analysis results, the server determines the time management requirement and generates instructions for setting reminders. The input is the analysis results, and the output is instructions for the device. During this process, control commands for specific tasks are generated.

[0674] Step 6:

[0675] The server sends control commands to the terminal. The input is the control commands, and the output is the command sent to the terminal. Even at this stage, data is exchanged via the communication network.

[0676] Step 7:

[0677] Based on the commands received by the device, it executes time management functions and sets reminders. The input is a control command, and the output is the completion of the reminder setting. The app launches on the device and the settings are saved.

[0678] Step 8:

[0679] The device uses a feedback mechanism to notify the user via voice, "Reminder set." The input is the completion information for setting the reminder, and the output is the voice feedback. In this final step, the user is informed of the result of the operation.

[0680] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0681] This invention is an interactive system for operating a smart device based on voice input from a user, and further possesses the function of recognizing the user's emotional state and adjusting the device operation and feedback accordingly. The specific implementation method of this system will be described below.

[0682] First, the user inputs operating instructions by voice into the smart device. The device captures this voice, converts it into text data using voice recognition, and simultaneously uses an emotion engine to identify the user's emotions from the voice. The converted text data and identified emotion data are sent from the device to the server.

[0683] The server analyzes text data using natural language processing to determine user requests. It also takes into account emotional information obtained by the emotion engine to determine the optimal response and action based on the user's current emotional state. Based on this determination, the server uses control mechanisms to instruct specific operations on the smart device and sets the method of providing feedback to the user. In particular, it adjusts the tone and content of the feedback based on emotional information to ensure the user feels as comfortable using the device as possible.

[0684] For example, if a user impatiently says "I want to take a picture" using voice input, the emotion engine will recognize that impatience. The server will recognize the request to "take a picture" and, taking into account the user's impatient emotional state, instruct the device to provide gentle feedback such as, "It's okay to calm down. I'll launch the camera now." This provides an emotionally empathetic response, reducing the stress the user experiences during the operation.

[0685] By incorporating user emotional states into the interaction in this way, the system becomes more intuitive and user-friendly, providing a particularly accessible learning and operating environment for senior users.

[0686] The following describes the processing flow.

[0687] Step 1:

[0688] The user provides voice input to a smart device, giving instructions for specific actions. This voice input may contain emotions.

[0689] Step 2:

[0690] The device receives the user's voice and converts it into text data using speech recognition technology. Simultaneously, it uses an emotion engine to identify the user's emotional state from the voice.

[0691] Step 3:

[0692] The terminal sends the converted text data and sentiment data to the server. This prepares the server to consider the user's request and sentiment simultaneously.

[0693] Step 4:

[0694] The server uses natural language processing to analyze text data and identify the user's specific requests. It determines actions such as "take a picture" and analyzes emotional information obtained from voice to understand the user's emotional state.

[0695] Step 5:

[0696] Based on the acquired requests and emotions, the server sends instructions to the terminal through control mechanisms. These instructions include specific operations on the smart device (e.g., launching the camera app) and emotionally sensitive feedback.

[0697] Step 6:

[0698] The device follows instructions from the server, performing specified device operations and providing emotionally responsive voice feedback to the user. For users experiencing frustration, it provides guidance in a calmer tone.

[0699] Step 7:

[0700] The user follows the instructions on the device and completes the necessary device operations. The device notifies the server of the results and provides guidance on the next steps as needed. By adjusting the guidance method according to the user's mood, a stress-free experience is provided.

[0701] (Example 2)

[0702] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0703] Conventional voice control systems process information and provide feedback without considering the user's emotional state, resulting in a uniform user experience and a particular difficulty in appropriately responding to emotionally charged voice input. Therefore, there is an urgent need to realize an interactive system that can provide appropriate operation and feedback in accordance with the user's emotions.

[0704] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0705] In this invention, the server includes speech recognition means for receiving voice input and converting it into text data, emotion recognition means for identifying the user's emotional state from the voice input, natural language processing means for analyzing the text data and determining the user's request, and feedback adjustment means for adjusting the tone and content of the feedback. This enables adaptive operation and feedback that takes the user's emotional state into consideration.

[0706] "Voice recognition means" refers to technology that digitizes voice input from a user and converts that voice information into text data. It includes devices and programs that analyze voice patterns and convert them into corresponding strings of characters.

[0707] "Emotion recognition means" refers to technology that identifies a user's emotions from voice input. This includes algorithms and devices for analyzing voice characteristics and inferring emotional states.

[0708] "Natural language processing means" refers to technologies for analyzing character data and understanding user requests and intentions. This includes methods for appropriately interpreting input strings and guiding necessary responses and actions.

[0709] "Feedback adjustment techniques" refer to technologies that adjust the tone and content of feedback based on the user's emotional state. They are means of improving the user experience by generating appropriate responses and taking emotions into consideration.

[0710] "Control means" refers to technology that operates the user's electronic devices based on determined requests. This includes functions that instruct appropriate actions based on data processing.

[0711] This system is an interactive system that performs emotionally responsive device operation based on user voice input. When the user gives voice commands, the terminal captures the voice. The device's microphone is used for voice capture, and the voice signal is converted into digital data.

[0712] The device converts this audio data into text data using speech recognition software (e.g., a general-purpose speech recognition API). Simultaneously, it uses an emotion engine (e.g., a common emotion analysis API) to identify the user's emotions from the audio. The resulting text data and emotion data are then sent from the device to the server.

[0713] The server uses natural language processing techniques (e.g., general-purpose natural language processing models) to analyze the received text data and understand the user's request. It also considers sentiment data and determines the optimal response or action based on the user's emotional state. In particular, it adjusts the content and tone of feedback based on sentiment information to ensure that what is communicated to the user is appropriate.

[0714] For example, if a user angrily commands "Play some music," the server will take that anger into consideration and instruct the device to provide considerate feedback such as "Playing some relaxing music."

[0715] An example of a prompt that utilizes a generative AI model is an instruction such as, "When the user becomes excited and gives a voice command, generate feedback to soothe their emotions." This system configuration enables device operation that is sensitive to the user's emotions.

[0716] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0717] Step 1:

[0718] The user directs their voice input towards the smart device. The device captures the voice through its microphone and receives it as digital audio data. This digital audio data becomes the input.

[0719] Step 2:

[0720] The terminal uses speech recognition software to convert digital speech data into text data. The input is digital speech data, and the output is text data. The software analyzes the speech waveform and generates the corresponding string of characters.

[0721] Step 3:

[0722] The device uses an emotion engine to analyze the user's emotions from their voice. The input is digital voice data, and the output is identified emotion data. This process estimates the user's emotions by evaluating characteristics such as voice tone and speed.

[0723] Step 4:

[0724] The terminal sends converted text data and sentiment data to the server. The input is text data and sentiment data, and the output is the data sent to the server. The data is securely transmitted to the server over the network.

[0725] Step 5:

[0726] The server uses natural language processing technology to analyze text data and determine user requests. The input is text data, and the output is the interpreted user request. It analyzes the structure and meaning of the text to identify the request.

[0727] Step 6:

[0728] The server considers emotional data and uses feedback adjustment mechanisms to determine the most appropriate response. The input is emotional data and the interpreted request, while the output is the adjusted feedback content. It selects a tone and content appropriate to the emotional state.

[0729] Step 7:

[0730] The server sends the determined feedback content to the terminal, which then provides feedback to the user via audio or screen. The input is the feedback content, and the output is the display of feedback to the user or audio playback. This ensures that an appropriate response is given to the user's input.

[0731] (Application Example 2)

[0732] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0733] The aim is to create a more natural and intuitive interface by providing device responses that take into account the user's emotional state when they operate the device via voice input. In particular, in information processing devices used in the home, there is a challenge in reducing user stress and improving the user experience by enabling feedback that responds to the user's emotions.

[0734] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0735] In this invention, the server includes speech recognition means, natural language processing means, and emotion recognition means. This makes it possible to convert speech input into text data and simultaneously identify requests and emotions.

[0736] "Voice recognition means" refers to a device or technology that receives voice input from a user and converts it into text data.

[0737] "Natural language processing means" refers to technologies that have the ability to analyze text data and determine user requests.

[0738] "Emotion recognition means" refers to technologies and devices that identify a user's emotional state from voice input.

[0739] An "information processing device" refers to an electronic device that controls its operation based on the user's requests and emotional state.

[0740] "Control means" refers to means for operating an information processing device based on judged requests and identified emotional states.

[0741] A "feedback mechanism" is a technology that provides users with information in voice that corresponds to the results of their actions and their emotions.

[0742] This invention relates to a method for interactively operating an information processing device using speech recognition means, natural language processing means, emotion recognition means, and feedback means. The system is intended for use in the home and aims to provide personalized feedback based on the user's emotional state.

[0743] First, the user inputs voice data into the information processing device via a microphone. The device then converts this voice into text data using speech recognition software such as the Google Speech API.

[0744] Next, the converted text is analyzed by natural language processing tools to understand the user's request. This process extracts specific operation instructions from the text data.

[0745] In parallel, voice input is analyzed through emotion recognition mechanisms to identify the user's emotional state. This emotion analysis utilizes emotion recognition models built with TensorFlow or PyTorch.

[0746] The server controls the operation of the information processing device by considering requests obtained through natural language processing and emotions obtained through emotion recognition. In particular, it can adjust the tone and content of responses provided through feedback mechanisms according to the user's emotional state.

[0747] For example, if a user asks in a tense voice, "What's the weather forecast?", the system will generate a gentle and reassuring response such as, "It's going to be sunny today. It'll be warm during the day, so please relax."

[0748] An example of a prompt for a generative AI model might be, "Please think of a reassuring response for a user who is feeling anxious." In this way, the system supports the user's actions by being sensitive to their emotions.

[0749] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0750] Step 1:

[0751] The terminal receives voice input from the user through the microphone and acquires that voice data. It obtains an audio signal as input and proceeds to process it as digital data.

[0752] Step 2:

[0753] The device processes the acquired audio data using speech recognition technology and converts it into text data. Specifically, it uses APIs such as Google Speech to convert digital audio signals into natural language text. The output is text data indicating the request.

[0754] Step 3:

[0755] The device analyzes voice data using emotion recognition technology to identify the user's emotional state based on their voice tone, pitch, and speed. It utilizes emotion recognition models built with TensorFlow or PyTorch for analysis. The output is data including emotion labels.

[0756] Step 4:

[0757] The server analyzes text data using natural language processing to understand user requests. The input is text data converted by the terminal, and the server analyzes the request content to identify operation commands. The output is information containing these operation commands.

[0758] Step 5:

[0759] The server determines the operation of the information processing device based on the acquired request and sentiment information. It combines the obtained operation commands and sentiment data to determine the optimal operation and sends a control signal to the terminal.

[0760] Step 6:

[0761] The terminal receives control signals from the server and operates the information processing device. Specific actions include launching applications and executing specific functions.

[0762] Step 7:

[0763] The device generates voice feedback for the user, responding with a tone and content that matches the user's emotional state. For example, if the user is feeling anxious, it will play reassuring content in a gentle voice through the speaker. The output is a voice response that is attentive to the user's emotions.

[0764] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0765] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0766] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0767] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0768] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0769] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0770] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0771] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0772] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0773] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0774] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0775] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0776] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0777] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0778] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0779] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0780] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0781] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0782] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0783] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0784] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0785] The following is further disclosed regarding the embodiments described above.

[0786] (Claim 1)

[0787] A speech recognition means that receives voice input from a user and converts it into text data,

[0788] A natural language processing means that analyzes the aforementioned text data and determines the user's request,

[0789] A control means for operating the user's smart device based on the determined request,

[0790] A feedback mechanism that provides voice feedback to the user regarding the results of the operation,

[0791] A system that includes this.

[0792] (Claim 2)

[0793] The system according to claim 1, characterized in that the control means has a function to automatically launch an application.

[0794] (Claim 3)

[0795] The system according to claim 1, characterized in that the feedback means has a function to provide step-by-step guidance for each stage of the operation.

[0796] "Example 1"

[0797] (Claim 1)

[0798] A voice processing means that receives voice input from a user and converts it into text data,

[0799] A language analysis means that analyzes the aforementioned text data and determines the user's request,

[0800] A control mechanism that operates the user's electronic device based on the determined request,

[0801] A means of providing feedback to the user by voice regarding the results of the operation,

[0802] A system that includes this.

[0803] (Claim 2)

[0804] The system according to claim 1, characterized in that the control mechanism has a function to automatically start software.

[0805] (Claim 3)

[0806] The system according to claim 1, characterized in that the reaction providing means has a function to provide sequential guidance at each stage of the operation.

[0807] "Application Example 1"

[0808] (Claim 1)

[0809] A speech recognition means that receives voice input from a user and converts it into text data,

[0810] A natural language processing means that analyzes the aforementioned text data and determines the user's request,

[0811] A control means for operating the user's information processing device based on the determined request,

[0812] A feedback mechanism that provides voice feedback to the user regarding the results of the operation,

[0813] A means of executing a time management function based on voice requests,

[0814] A system that includes this.

[0815] (Claim 2)

[0816] The system according to claim 1, characterized in that the control means has a function to automatically start an application program.

[0817] (Claim 3)

[0818] The system according to claim 1, characterized in that the feedback means has a function to provide step-by-step guidance at each stage of the operation.

[0819] "Example 2 of combining an emotion engine"

[0820] (Claim 1)

[0821] A speech recognition means that receives voice input from a user and converts it into text data,

[0822] An emotion recognition means for identifying the user's emotional state from the aforementioned voice input,

[0823] A natural language processing means that analyzes the aforementioned character data and determines the user's request,

[0824] A feedback adjustment means that adjusts the tone and content of the feedback based on the aforementioned emotional state,

[0825] A control means for operating the user's electronic device based on the determined request,

[0826] A system that includes this.

[0827] (Claim 2)

[0828] The system according to claim 1, characterized in that the control means has a function to automatically start a program.

[0829] (Claim 3)

[0830] The system according to claim 1, characterized in that the feedback adjustment means has a function to provide feedback according to the user's emotional state.

[0831] "Application example 2 when combining with an emotional engine"

[0832] (Claim 1)

[0833] A speech recognition means that receives voice input from a user and converts it into text data,

[0834] A natural language processing means that analyzes the aforementioned text data and determines the user's request,

[0835] An emotion recognition means that identifies the user's emotional state from voice input,

[0836] A control means for operating a user's information processing device based on judged requests and identified emotional states,

[0837] A feedback mechanism that provides users with voice feedback that corresponds to the results of their actions and their emotions,

[0838] A system that includes this.

[0839] (Claim 2)

[0840] The system according to claim 1, characterized in that the control means has a function to generate a personalized response according to the emotional state.

[0841] (Claim 3)

[0842] The system according to claim 1, characterized in that the feedback means has a function to adjust the tone and content of the feedback according to the emotional state. [Explanation of Symbols]

[0843] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A speech recognition means that receives voice input from a user and converts it into text data, A natural language processing means that analyzes the aforementioned text data and determines the user's request, A control means for operating the user's information processing device based on the determined request, A feedback mechanism that provides voice feedback to the user regarding the results of the operation, A means of executing a time management function based on voice requests, A system that includes this.

2. The system according to claim 1, characterized in that the control means has a function to automatically start an application program.

3. The system according to claim 1, characterized in that the feedback means has a function to provide step-by-step guidance at each stage of the operation.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A