system
A voice-based system converts user input to text, analyzes screen information, simulates input device operations, and provides feedback, addressing the inconvenience of physical input devices and enabling adaptation to new applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-10
- Publication Date
- 2026-06-22
Smart Images

Figure 2026101188000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional computer operations require direct use of input devices such as keyboards and mice, which is inconvenient for users who have difficulty with these operations, especially users with physical limitations. Also, every time new functions or applications are added to a PC, new learning is often required to operate them. There is a need to provide a system that overcomes such problems and allows users to efficiently operate a PC with a natural interface.
Means for Solving the Problems
[0005] This invention solves the problem by providing a means to receive user instructions using voice input, convert the voice into text, analyze screen information output to a display device, and identify executable actions corresponding to the user's instructions. Furthermore, it includes means to simulate the operation of the input device based on the identified actions, and feeds back the results of the operation using speech synthesis technology. In addition, by providing operations that can be adapted to new functions or applications based on the analyzed screen information, it realizes simple and versatile operation.
[0006] "Voice input" is a method by which users communicate operational instructions to a system through their voice.
[0007] A "user" refers to a person who operates and gives instructions to a system.
[0008] "Converting to text" refers to the process of converting audio data into written information.
[0009] A "display device" is a device that provides visual information to a user, such as a computer screen.
[0010] "Screen information" refers to visual information such as windows, text, and icons currently displayed on a display device.
[0011] "Analysis" is the process of processing given information and extracting meaningful data and structures from it.
[0012] An "action" refers to a specific operation that a system performs based on certain instructions.
[0013] "Simulation" is a technique that mimics and executes actual keyboard and mouse operations.
[0014] "Feedback" refers to a response from the system that informs the user of the results of an operation performed by the system.
[0015] "Voice synthesis technology" refers to the technology of converting character information generated by a computer into voice and outputting it.
[0016] "Adaptation to new functions or applications" refers to the ability to smoothly handle software and operations that did not exist before.
Brief Explanation of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] The system of the present invention enables the user to operate a PC through voice input. Its embodiments are described in detail below.
[0039] The system begins with the user giving a voice command into the microphone. For example, consider the command, "Open a web browser and search for news." The terminal receives the voice input and converts it into text using Speech-to-Text technology. The converted text is then used for subsequent processing.
[0040] The device captures the user's PC screen as an image and analyzes the displayed windows, icons, and text fields. This analysis uses generative AI technology. Through this analysis, the system identifies how the user's voice commands are translated into screen operations.
[0041] For example, if the device receives the voice command, "Open a web browser and search for news," it will first display an existing browser window or launch a new browser, and then enter the keyword "news" into the URL bar. It will also automatically simulate pressing the Enter key to perform the search, if necessary.
[0042] The device uses speech synthesis technology to provide feedback to the user about the operations performed. A voice message such as "News search performed" is output, allowing the user to confirm that the instructions were carried out correctly.
[0043] Furthermore, based on the screen information it analyzes, the device provides flexible voice control for newly installed applications and added features. This versatility allows users to continue using their PC without having to learn new operating methods.
[0044] This system allows users to operate their PCs through natural voice interaction without using physical input devices. This is particularly useful for users who have difficulty with manual operation, and also provides a comfortable and efficient operating environment for general users.
[0045] The following describes the processing flow.
[0046] Step 1:
[0047] The user gives voice commands into the PC's microphone. For example, they might say, "Open a text editor and create a new document."
[0048] Step 2:
[0049] The device captures voice input from the microphone and immediately performs Speech-to-Text processing using its speech recognition engine. In this process, the voice data is converted into text data.
[0050] Step 3:
[0051] The device captures the screen and uses AI to analyze information about currently displayed windows and icons. This allows it to identify which applications and windows are open.
[0052] Step 4:
[0053] The device determines what action needs to be performed based on the analyzed screen information and text converted from the audio. If the user instructs it to "open a text editor," it checks if an editor is already open and launches a new one if it's not.
[0054] Step 5:
[0055] The terminal simulates and performs keyboard input and mouse operations based on the specified action. For example, if the instruction is "Create a new document," it selects the "New" menu in the text editor and places the cursor where to input text in the document.
[0056] Step 6:
[0057] The device uses speech synthesis technology to provide feedback to the user on whether the operation was performed successfully. It outputs a voice message such as "A new document has been created" to inform the user of the result.
[0058] Step 7:
[0059] Based on the feedback received, the user can either give further instructions or terminate the process. By referring to the feedback, the user can verify that the instructions were followed correctly.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] In voice-based operation assistance technologies, it is crucial to provide an environment where users can easily perform complex tasks. However, conventional technologies have limitations in terms of voice input accuracy and responsiveness, making it particularly difficult to flexibly adapt to new functions and applications. Furthermore, it is difficult to accurately simulate on-screen actions, resulting in a lack of quick and clear feedback to users. There is a need to solve these problems and provide an efficient and natural voice-based interface.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes means for acquiring user instructions using acoustic signals, means for converting the acquired acoustic signals into text, and means for analyzing video information output to a display means and determining executable actions in accordance with the user's instructions. This allows the user to operate the computer naturally through voice, receive accurate feedback while flexibly adapting to new functions and applications.
[0065] An "acoustic signal" is an electrical signal that represents air vibrations, including the voice of a user.
[0066] "User" refers to an individual or entity that operates the system through voice input.
[0067] "Display means" refers to devices or media used to provide information visually, and is usually a computer monitor or screen.
[0068] "Visual information" refers to digital data or screen content that is output to a display device and is visually recognizable.
[0069] "Converting to text" is the process of analyzing an audio signal and converting it into text data that corresponds to the speech.
[0070] "Analysis" refers to the process of investigating video information and data to understand its structure and content.
[0071] "Executable actions" refer to the specific operations and processes that a computer can perform based on user instructions.
[0072] "Simulating" is the process of imitating the operation of actual input devices in order to obtain equivalent effects.
[0073] "Responding" refers to the act of conveying results or information in response to a user's instructions, and is often provided as audio or visual feedback.
[0074] The present invention provides a system that assists user operation through voice input. Specifically, the user first emits an acoustic signal, i.e., a voice command, into the microphone of the terminal. The terminal receives this acoustic signal and converts the voice into text using voice recognition software technology. The technology used at this stage includes general voice recognition APIs.
[0075] Next, the terminal acquires the video information displayed on the display device and uses a generative AI model to analyze it. The generative AI model includes a system known as a common natural language processing technique, which interprets and converts user instructions into executable actions. This is achieved by analyzing the arrangement of icons and the state of windows on the screen.
[0076] For example, based on a prompt message such as "Open your email and compose a new message," the system will automatically launch the email application and simulate opening a new message window. During this process, necessary mouse clicks and keyboard inputs will also be simulated by the terminal.
[0077] Once processing is complete, the terminal provides voice feedback to the user. This is done using synthesized speech technology, and the user receives a voice message such as, "New message creation complete." In this way, the user's voice commands are executed through the system, resulting in efficient and natural operation.
[0078] Furthermore, this system can flexibly adapt to newly installed programs and functions. Therefore, users do not need to relearn how to operate new applications when they are added. In this way, users can enjoy an intuitive and user-friendly computing environment that allows them to control a variety of programs using only their voice.
[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0080] Step 1:
[0081] Acquiring voice input
[0082] The user gives clear voice instructions into the microphone. For example, they might say, "Open the document and add a new page."
[0083] The input is an acoustic signal from the user.
[0084] The output is raw audio data received by the terminal as an acoustic signal.
[0085] Step 2:
[0086] Speech-to-text conversion
[0087] The terminal converts the received audio data into text data using speech recognition software (such as a speech recognition API).
[0088] The input is raw audio data as an acoustic signal.
[0089] As part of the data processing, the acoustic signal is analyzed and the audio pattern is converted into a corresponding string of characters.
[0090] The output is text data corresponding to the voice instructions.
[0091] Step 3:
[0092] Acquisition and analysis of video information
[0093] The terminal captures screen information from the display device and performs analysis using a generated AI model.
[0094] The input consists of captured video data and converted text data.
[0095] As a data calculation tool, it analyzes the position and state of elements on the screen and generates data to determine actions that correspond to user instructions.
[0096] The output consists of specific operating instructions based on the analysis results.
[0097] Step 4:
[0098] Execution of specific operations
[0099] The terminal performs necessary operations based on the analysis results and simulates the operation of input devices. For example, it opens the appropriate application and performs the necessary key inputs and mouse operations to use the corresponding functions.
[0100] The input consists of specific instructions for operation.
[0101] The output is the result of the user's intended operation being executed on the device.
[0102] Step 5:
[0103] Provide feedback
[0104] The device uses voice generation technology to communicate to the user that the operation is complete via an acoustic signal.
[0105] The input is a message indicating the completion of the operation.
[0106] The output is an acoustic signal that is reproduced as audio feedback.
[0107] Step 6:
[0108] Adapting to new features
[0109] The device also analyzes voice commands for newly installed functions and applications, enabling them to be operated.
[0110] The input is new video information to be analyzed and adapted to.
[0111] The output is an operable environment for new functions.
[0112] (Application Example 1)
[0113] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0114] In modern information technology, many users experience difficulties accessing digital services and applications due to manual interface operation. In particular, for users with physical limitations or those unfamiliar with manual operation, a natural and intuitive voice-based interface is useful, but current technology offers limited, versatile, and flexible solutions. This invention aims to solve these problems and provide a system that allows users to efficiently operate digital services via voice.
[0115] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0116] In this invention, the system includes means for acquiring user instructions using voice input, means for converting the acquired voice into text, and means for analyzing visual information output to a display means to determine executable processing based on the user's instructions. This enables users to intuitively operate digital services using only their voice.
[0117] "Voice input" is a method by which a system can receive instructions and information spoken by the user.
[0118] "User instructions" refer to the specific content that a user issues to request a particular operation or process from the system.
[0119] "Acquired audio" refers to the audio signal received through audio input.
[0120] "Converting to a string" means processing an audio signal and converting it into data in a corresponding text format.
[0121] "Display means" refers to a device or function for visually presenting information, and includes monitors and displays.
[0122] "Visual information" refers to data such as images and text that are presented visually through a display means.
[0123] "Analyzing" means scrutinizing data in order to understand the information and recognize the necessary structure.
[0124] "Executable processes" refer to operations and actions that a system can perform based on user instructions.
[0125] "Reproducing the operation of a device" means simulating the operations performed by an actual device within a system.
[0126] "Reporting" means that the system communicates the results or status of its actions to the user.
[0127] "Speech conversion technology" is a technology that converts text information into a speech format.
[0128] "Communicating through sound" means providing information to users aurally.
[0129] "New features and programs" refer to new operations or uses added to the system.
[0130] The system that realizes this application will enable users to intuitively manipulate digital information through a series of processes including speech recognition, data analysis, and feedback provision. Specifically, it will utilize smart devices such as smart glasses, smartphones, or other computing devices.
[0131] The device first receives the user's voice input via the microphone and converts it into text using a speech recognition API (e.g., Google® Cloud Speech-to-Text). Next, it uses a generative AI model (e.g., OpenAI® GPT series) to identify the user's instructions from the converted text and analyzes the visual information based on those instructions.
[0132] Based on the analyzed visual information, the device determines and reproduces the executable actions required in response to the user's request. After the determined actions are performed, the results are reported to the user via voice feedback. This voice feedback is generated using an API that employs speech conversion technology (e.g., Amazon Polly).
[0133] For example, if a user says, "I want to contact my family," the device will launch a video call app and automatically make a call to a registered contact. Also, if the user says, "I want to check my next hospital appointment," the device will open a calendar app and read out the date.
[0134] Examples of prompt messages include, "Please display the next hospital appointment," and "Please make a video call to family member [name]."
[0135] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0136] Step 1:
[0137] The terminal acquires the user's voice input through a microphone device. This input is an audio signal for capturing the user's verbal instructions in real time. The format of the voice input is digitized audio data.
[0138] Step 2:
[0139] The device converts the acquired audio signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). This API analyzes the audio waveform and sequentially converts human speech into corresponding text. The input is audio data, and the output is text data in which the audio has been converted into a string.
[0140] Step 3:
[0141] The server utilizes a generative AI model (e.g., OpenAI's GPT series) to analyze the converted text data and identify the user's intended instructions. This analysis uses natural language processing techniques to understand the context and determine the necessary actions. The input is text data, and the output is executable processing information based on the analyzed instructions.
[0142] Step 4:
[0143] The device executes visual information and action simulations based on the specified instructions. For example, if told to "make a video call," it will launch the relevant application and automatically select the appropriate contact. Input is executable processing information, and output is the system's execution.
[0144] Step 5:
[0145] The device generates audio feedback to report the processing results to the user using speech conversion technology. An API (e.g., Amazon Polly) synthesizes text information into speech, generating sounds to communicate the current process and results to the user. The input is text information of the processing results, and the output is audio feedback.
[0146] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0147] The system of the present invention aims to enable users to operate their PCs more naturally and interactively using voice input and emotion recognition technology. Embodiments are described in detail below.
[0148] The user first speaks a voice command into the PC's microphone. For example, they might say, "Open the document, then start rehearsing the presentation." Once this voice is captured by the device, Speech-to-Text technology converts the speech into text. This ensures that the user's instructions are clearly understood.
[0149] The terminal captures the screen currently displayed on the display device and analyzes that information. It collects information about the displayed applications and windows and matches it with the user's voice commands. This identifies the target for performing the operation instructed by the user.
[0150] Furthermore, this system is equipped with an emotion engine that analyzes the user's emotions from their voice. The emotion engine analyzes, for example, the tone, speed, and volume of the voice to determine whether the user is excited, calm, or dissatisfied. As a result, the system adjusts its response according to the user's emotional state.
[0151] For example, if a user gives an instruction in a voice that conveys urgency, such as "Do something about this document right now!", the device will recognize the user's emotion as "urgency," execute the instruction more quickly than usual, and provide proactive feedback using speech synthesis technology, such as "I've opened the document immediately, what do you want to do next?"
[0152] Once an operation is performed, the terminal provides feedback to the user. Using speech synthesis technology, an easy-to-understand voice message is generated for the user, providing confirmation of the operation's success and advice on the next steps.
[0153] This system allows users not only to control their PCs with their voice, but also to experience customized interactions tailored to their emotional state, enabling them to enjoy a more comfortable and efficient operating environment.
[0154] The following describes the processing flow.
[0155] Step 1:
[0156] The user gives voice commands into the PC's microphone. For example, they might say, "Open my email and check for the latest incoming emails."
[0157] Step 2:
[0158] The device captures audio through its microphone, inputs the audio data into a Speech-to-Text engine in real time, and converts it into text. This text conversion allows for clear analysis of the user's instructions.
[0159] Step 3:
[0160] During the process of processing voice input, the device uses its built-in emotion engine to analyze the tone, speed, and volume of the voice to identify the user's emotional state. For example, states such as excitement, relaxation, and dissatisfaction are analyzed.
[0161] Step 4:
[0162] The device captures the current screen information and analyzes it using AI technology. Based on the analysis results, actions related to voice commands are identified. For example, it determines whether the email app is already open or if it needs to be opened again.
[0163] Step 5:
[0164] The device performs actions determined by the analyzed voice commands and screen information. It simulates operations such as opening the email app and selecting the most recent received email from the list.
[0165] Step 6:
[0166] The device utilizes the results of the emotion engine to adjust feedback according to the user's emotional state. For example, if the user is expressing dissatisfaction, it will generate a response in a gentle tone such as, "I've seen your email. Is there anything I can help you with?"
[0167] Step 7:
[0168] The device uses speech synthesis technology to provide voice feedback on the results of the operations performed. The feedback is generated in real time, informing the user of the success of the operation and prompting them to take the next action.
[0169] This processing flow allows the system to quickly and accurately perform PC operations based on voice input and provide the user with appropriate responses that reflect their emotions.
[0170] (Example 2)
[0171] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0172] Traditional computer systems made it difficult for users to operate them using natural voice commands and failed to provide responses that took user emotions into consideration. This resulted in users experiencing stress during voice input and decreased operational efficiency. Furthermore, there was a lack of methods to flexibly apply operations based on information displayed on the screen.
[0173] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0174] In this invention, the server includes means for receiving user instructions using voice input, means for using a natural language processing device to convert the received voice into text, means for analyzing image information acquired from a display device to identify the target to be executed based on the user's instructions, and means for analyzing the user's emotions from the voice and adjusting the response based on those emotions. This enables the user to effectively operate the computer with a natural voice and receive a customized response that is tailored to their emotions.
[0175] "Voice input" is a method by which users give voice commands directly to a computer through a microphone.
[0176] A "natural language processing device" is a device equipped with technology to convert speech data into text, and has the function of making user voice instructions into a format that a computer can understand.
[0177] A "display device" is a device that displays information visually, such as a computer screen or monitor.
[0178] "Analyzing image information" is the process of analyzing screen captures and image data to understand what is displayed within them and extract specific data.
[0179] "Analyzing user emotions from voice" refers to a technology that analyzes the tone, speed, volume, and other elements contained in voice data to determine the user's emotional state.
[0180] A "generative AI technology model" is a model equipped with an algorithm that uses artificial intelligence to automatically generate content and feedback based on user instructions.
[0181] "Speech synthesis technology" is a technology that converts text data into speech, providing users with computer responses as voice.
[0182] This invention is a system that allows users to operate a computer using natural voice, enabling more human-like interaction. The system combines voice input, image information analysis, generative AI technology models, and speech synthesis technology.
[0183] First, the user gives voice instructions through the device's microphone, and the device receives this voice data. The voice is then converted into text using a natural language processing unit. In this process, common technologies can be used as natural language processing techniques, for example.
[0184] Next, the terminal acquires image information from the display device and analyzes its contents. For example, image processing software is used to identify the appropriate action to take based on the user's instructions. The terminal also incorporates an emotion analyzer that analyzes the user's emotions from voice data. For instance, it can determine whether the user is feeling anxious or calm.
[0185] By utilizing generative AI technology models, the device generates responses that correspond to the user's instructions and emotions. This generative AI model can be modified using common artificial intelligence algorithms. For example, if a user instructs, "Can you open my presentation materials and help me with something?", the device might generate a response such as, "I'll open the materials now, and then shall we practice the slides together?"
[0186] Finally, the device uses speech synthesis technology to provide the generated response to the user as audio. This process utilizes speech synthesis technology to provide the user with natural and easy-to-understand voice messages.
[0187] This system allows users to operate computers through voice commands and emotionally responsive interactions, resulting in improved operational efficiency and a richer user experience.
[0188] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0189] Step 1:
[0190] The user gives voice commands to the device. For example, they might say, "Open the presentation materials." The device receives this voice data as input through its built-in microphone. The voice data is then converted from a sound waveform into digital audio data.
[0191] Step 2:
[0192] The terminal transmits the received audio data to a natural language processing unit, which converts the audio into text. A speech recognition engine is used here to extract word sequences from the audio waveform as output. This converted text serves as the basic data for understanding the user's instructions.
[0193] Step 3:
[0194] The terminal acquires image information from the display device and analyzes it to associate it with user instructions. For example, image processing software performs a screen capture and extracts information about the active window and application. This analysis identifies actions that can be performed based on the user's instructions.
[0195] Step 4:
[0196] The device uses a voice analysis engine to evaluate the user's emotional state from voice data. Parameters such as voice tone, speed, and volume are used as input data, and the emotional state is output. For example, it can determine whether the user is speaking in an expectant voice or in an anxious voice.
[0197] Step 5:
[0198] The device utilizes a generative AI technology model to generate responses that correspond to the user's emotional state and instructions. Input includes text data and emotional information obtained in the previous step, and the output is an appropriate response. For example, it might generate a message like, "I've opened the document. What should I do next?"
[0199] Step 6:
[0200] The device converts text responses generated using speech synthesis technology into speech. The synthesized speech data is provided as output for communication with the user and played back from the device to the user as voice feedback. This process ensures that the user receives voice feedback appropriate to the instructions given.
[0201] (Application Example 2)
[0202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0203] Current consumer robots have limitations in that voice control is restricted to simple commands, and they cannot provide flexible responses or adjust operations based on the user's emotional state. Furthermore, while users desire more natural communication, existing systems are unable to meet this need.
[0204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0205] In this invention, the server includes means for receiving operation instructions using voice input, means for converting the received voice into text data, and means for analyzing video information displayed on a display device to identify executable actions corresponding to the instructions. This enables operation and feedback in accordance with the user's voice instructions and emotional state.
[0206] "Voice input" is a method for capturing voice information spoken by the user.
[0207] "Operation instructions" refer to the instructions a user gives to a robot to perform a specific action.
[0208] "Converting to text data" refers to the process of analyzing audio information after it has been input and converting it into a corresponding string of characters.
[0209] "Display devices" refer to devices that output information visually, such as screens and displays.
[0210] "Analyzing video information" is the process of analyzing the displayed visual information to obtain the necessary data.
[0211] "Identifying feasible actions" means identifying specific tasks that should be performed based on the analyzed information and instructions.
[0212] An "input device" is a device used to receive data or instructions, and includes keyboards and mice.
[0213] "To mimic an operation" means to reproduce an actual mechanical or electronic action based on an identified movement.
[0214] "Analyzing emotional state" means analyzing the tone and speed of voice from audio data to identify the user's emotions.
[0215] "Adjusting the speed of actions and response content" refers to the process of changing the speed of actions performed and the content of voice feedback according to the user's emotions.
[0216] This invention aims to achieve advanced interaction in consumer robots by combining voice input and emotion recognition technology. This system receives voice commands from the user, converts them into text data, and processes them in the robot's control unit. The details are described below.
[0217] The server first receives operation instructions from the user via a voice input device. The voice data is processed in real time and converted into text data using the Google Speech-to-Text API. This conversion process ensures that each instruction is interpreted as a clear digital command.
[0218] Furthermore, the server executes a Python script that implements an emotion recognition algorithm to analyze the user's emotional state along with the text data. Specifically, it analyzes acoustic features such as voice tone, speed, and volume to identify whether the user is in a particular emotional state, such as excitement, relaxation, or anxiety. This allows the server to adjust the operation speed and feedback content as needed.
[0219] Based on this information, the robot quickly identifies the specified action to analyze the displayed video information and generates an action scenario. The server sends commands to the robot's control unit, which then performs physical actions such as household chores according to the user's instructions. Simultaneously, it generates audio feedback using Amazon Polly to notify the user of the results and the next steps.
[0220] For example, if a user instructs the robot to "clean the kitchen and then prepare dinner," the robot will convert the voice into text, perform the kitchen cleaning task, check the ingredients in the refrigerator, and begin preparing the necessary ingredients. During the process, the robot will provide feedback such as, "I've started cleaning the kitchen. What can I help you with next?"
[0221] Examples of prompt statements that can be used as input to a generative AI model include the following:
[0222] "When a user gives a voice command with emotion, please convert that command into text and explain the action that corresponds to that emotion."
[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0224] Step 1:
[0225] The terminal receives voice commands spoken by the user via the microphone. The input is an audio signal, and the output is captured by the terminal as a digital signal. This is the process of acquiring audio data.
[0226] Step 2:
[0227] The device converts the received audio signal into text data using the Google Speech-to-Text API. The input is an audio signal, and the output is the converted string. This conversion process involves phonological analysis and formatting of specific instructions.
[0228] Step 3:
[0229] The server analyzes the acquired text data to identify the specific tasks instructed by the user. The input is text data, and the output is a task list. This includes interpreting the meaning of the instructions using natural language processing techniques.
[0230] Step 4:
[0231] The device applies an emotion recognition algorithm to voice data to analyze the user's emotional state. The input is voice data, and the output is information about the emotional state. It performs specific actions to identify the user's emotions by analyzing the tone, speed, and intensity of the voice.
[0232] Step 5:
[0233] For each task, it generates instructions to adjust the execution speed and priority based on the user's emotional state. The input is a task list and emotional information, and the output is an adjusted task list. This includes actions to determine timing and order according to the user's emotions.
[0234] Step 6:
[0235] The terminal sends these instructions to the robot's control unit. The input is a coordinated task list, and the output is specific control signals to the robot. It is a process of instructing physical actions using a communication protocol.
[0236] Step 7:
[0237] The robot initiates a specified action based on a control signal, and its progress is monitored by a server during the operation. The input is the control signal, and the output is progress data. This includes performing actual physical tasks and tracking the results in real time.
[0238] Step 8:
[0239] The server receives progress data and provides voice feedback to the user regarding the progress and results via Amazon Polly. The input is progress data, and the output is a generated voice message. It is a voice generation process designed to communicate results in an accessible way.
[0240] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0241] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0242] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0243] [Second Embodiment]
[0244] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0245] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0246] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0247] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0248] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0249] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0250] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0251] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0252] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0253] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0254] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0255] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0256] The system of the present invention enables the user to operate a PC through voice input. Its embodiments are described in detail below.
[0257] The system begins with the user giving a voice command into the microphone. For example, consider the command, "Open a web browser and search for news." The terminal receives the voice input and converts it into text using Speech-to-Text technology. The converted text is then used for subsequent processing.
[0258] The device captures the user's PC screen as an image and analyzes the displayed windows, icons, and text fields. This analysis uses generative AI technology. Through this analysis, the system identifies how the user's voice commands are translated into screen operations.
[0259] For example, if the device receives the voice command, "Open a web browser and search for news," it will first display an existing browser window or launch a new browser, and then enter the keyword "news" into the URL bar. It will also automatically simulate pressing the Enter key to perform the search, if necessary.
[0260] The device uses speech synthesis technology to provide feedback to the user about the operations performed. A voice message such as "News search performed" is output, allowing the user to confirm that the instructions were carried out correctly.
[0261] Furthermore, based on the screen information it analyzes, the device provides flexible voice control for newly installed applications and added features. This versatility allows users to continue using their PC without having to learn new operating methods.
[0262] This system allows users to operate their PCs through natural voice interaction without using physical input devices. This is particularly useful for users who have difficulty with manual operation, and also provides a comfortable and efficient operating environment for general users.
[0263] The following describes the processing flow.
[0264] Step 1:
[0265] The user gives voice commands into the PC's microphone. For example, they might say, "Open a text editor and create a new document."
[0266] Step 2:
[0267] The device captures voice input from the microphone and immediately performs Speech-to-Text processing using its speech recognition engine. In this process, the voice data is converted into text data.
[0268] Step 3:
[0269] The device captures the screen and uses AI to analyze information about currently displayed windows and icons. This allows it to identify which applications and windows are open.
[0270] Step 4:
[0271] The device determines what action needs to be performed based on the analyzed screen information and text converted from the audio. If the user instructs it to "open a text editor," it checks if an editor is already open and launches a new one if it's not.
[0272] Step 5:
[0273] The terminal simulates and performs keyboard input and mouse operations based on the specified action. For example, if the instruction is "Create a new document," it selects the "New" menu in the text editor and places the cursor where to input text in the document.
[0274] Step 6:
[0275] The device uses speech synthesis technology to provide feedback to the user on whether the operation was performed successfully. It outputs a voice message such as "A new document has been created" to inform the user of the result.
[0276] Step 7:
[0277] Based on the received feedback, the user issues further instructions or ends the process. By referring to the feedback, the user can confirm that the instructions are being correctly executed.
[0278] (Example 1)
[0279] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0280] In an operation support technology using voice input, it is important to provide an environment in which users can perform complex tasks simply. However, conventional technologies have problems with the accuracy and responsiveness of voice input, especially the difficulty of flexibly adapting to new functions and applications. Also, it is difficult to appropriately simulate operations on the screen, and there is a lack of prompt and clear feedback to the user. There is a need to solve such problems and provide an efficient and natural interface by voice.
[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0282] In this invention, the server includes means for acquiring a user's instruction using an acoustic signal, means for converting the acquired acoustic signal into characters, and means for analyzing video information output to a display means and determining an executable operation according to the user's instruction. Thereby, the user can naturally operate the computer through voice, and while flexibly adapting to new functions and applications, can receive accurate feedback.
[0283] An "acoustic signal" is what represents the vibration of air as an electrical signal, including the user's voice, etc.
[0284] A "user" refers to an individual or entity that operates the system through voice input.
[0285] "Display means" refers to a device or medium for visually providing information, usually a computer monitor or screen.
[0286] "Video information" refers to digital data or screen content that is output to the display means and is visually recognizable.
[0287] "Convert to text" is a process of analyzing an acoustic signal and converting it into text data corresponding to the voice.
[0288] "Analysis" refers to the work of investigating video information or data and understanding its structure and content.
[0289] "Executable operation" refers to specific operations or processes that a computer can perform based on the instructions of the user.
[0290] "Simulate" is a process of imitating the operation of the actual input means to obtain an equivalent effect.
[0291] "Respond" refers to the act of conveying the result or information in response to the user's instruction, and in many cases, it is provided as voice or visual feedback.
[0292] The system of the present invention is a system that supports the user's operation through voice input. Specifically, first, the user emits an acoustic signal, that is, a voice instruction, towards the microphone of the terminal. The terminal receives this acoustic signal and uses software technology for voice recognition to convert the voice into text. The technology used at this stage includes general voice recognition APIs.
[0293] Next, the terminal acquires the video information presented on the display means and uses a generative AI model to analyze it. The generative AI model includes a system known as general natural language processing technology, which interprets and converts the user's instruction into an executable operation. This is achieved by analyzing the arrangement of icons on the screen and the state of windows.
[0294] For example, based on a prompt message such as "Open your email and compose a new message," the system will automatically launch the email application and simulate opening a new message window. During this process, necessary mouse clicks and keyboard inputs will also be simulated by the terminal.
[0295] Once processing is complete, the terminal provides voice feedback to the user. This is done using synthesized speech technology, and the user receives a voice message such as, "New message creation complete." In this way, the user's voice commands are executed through the system, resulting in efficient and natural operation.
[0296] Furthermore, this system can flexibly adapt to newly installed programs and functions. Therefore, users do not need to relearn how to operate new applications when they are added. In this way, users can enjoy an intuitive and user-friendly computing environment that allows them to control a variety of programs using only their voice.
[0297] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0298] Step 1:
[0299] Acquiring voice input
[0300] The user gives clear voice instructions into the microphone. For example, they might say, "Open the document and add a new page."
[0301] The input is an acoustic signal from the user.
[0302] The output is raw audio data received by the terminal as an acoustic signal.
[0303] Step 2:
[0304] Conversion from Voice to Text
[0305] The terminal converts the received voice data into character data using voice recognition software (such as a voice recognition API).
[0306] The input is raw voice data as an acoustic signal.
[0307] As data processing, the acoustic signal is analyzed and the voice pattern is converted into a corresponding character string.
[0308] The output is text data corresponding to the voice instruction.
[0309] Step 3:
[0310] Acquisition and Analysis of Video Information
[0311] The terminal captures the screen information of the display means and performs analysis using the generated AI model.
[0312] The input is the captured video data and the converted text data.
[0313] As data calculation, the position and state of the elements on the screen are analyzed, and data for determining the operation corresponding to the user instruction is generated.
[0314] The output is a specific operation instruction as the analysis result.
[0315] Step 4:
[0316] Execution of Specific Operations
[0317] The terminal executes the necessary operations based on the analysis result and simulates the operation of the input device. For example, it opens an appropriate application and performs key inputs and mouse operations necessary to use the corresponding function.
[0318] The input is a specific operation instruction.
[0319] The output is the result of the user's intended operation being executed on the device.
[0320] Step 5:
[0321] Provide feedback
[0322] The device uses voice generation technology to communicate to the user that the operation is complete via an acoustic signal.
[0323] The input is a message indicating the completion of the operation.
[0324] The output is an acoustic signal that is reproduced as audio feedback.
[0325] Step 6:
[0326] Adapting to new features
[0327] The device also analyzes voice commands for newly installed functions and applications, enabling them to be operated.
[0328] The input is new video information to be analyzed and adapted to.
[0329] The output is an operable environment for new functions.
[0330] (Application Example 1)
[0331] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0332] In modern information technology, many users experience difficulties accessing digital services and applications due to manual interface operation. In particular, for users with physical limitations or those unfamiliar with manual operation, a natural and intuitive voice-based interface is useful, but current technology offers limited, versatile, and flexible solutions. This invention aims to solve these problems and provide a system that allows users to efficiently operate digital services via voice.
[0333] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0334] In this invention, the system includes means for acquiring user instructions using voice input, means for converting the acquired voice into text, and means for analyzing visual information output to a display means to determine executable processing based on the user's instructions. This enables users to intuitively operate digital services using only their voice.
[0335] "Voice input" is a method by which a system can receive instructions and information spoken by the user.
[0336] "User instructions" refer to the specific content that a user issues to request a particular operation or process from the system.
[0337] "Acquired audio" refers to the audio signal received through audio input.
[0338] "Converting to a string" means processing an audio signal and converting it into data in a corresponding text format.
[0339] "Display means" refers to a device or function for visually presenting information, and includes monitors and displays.
[0340] "Visual information" refers to data such as images and text that are presented visually through a display means.
[0341] "Analyzing" means scrutinizing data in order to understand the information and recognize the necessary structure.
[0342] "Executable processes" refer to operations and actions that a system can perform based on user instructions.
[0343] "Reproducing the operation of a device" means simulating the operations performed by an actual device within a system.
[0344] "Reporting" means that the system communicates the results or status of its actions to the user.
[0345] "Speech conversion technology" is a technology that converts text information into a speech format.
[0346] "Communicating through sound" means providing information to users aurally.
[0347] "New features and programs" refer to new operations or uses added to the system.
[0348] The system that realizes this application will enable users to intuitively manipulate digital information through a series of processes including speech recognition, data analysis, and feedback provision. Specifically, it will utilize smart devices such as smart glasses, smartphones, or other computing devices.
[0349] The device first receives the user's voice input via the microphone and converts it into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). Next, it uses a generative AI model (e.g., OpenAI's GPT series) to identify the user's instructions from the converted text and analyzes the visual information based on those instructions.
[0350] Based on the analyzed visual information, the device determines and reproduces the executable actions required in response to the user's request. After the determined actions are performed, the results are reported to the user via voice feedback. This voice feedback is generated using an API that employs speech conversion technology (e.g., Amazon Polly).
[0351] For example, if a user says, "I want to contact my family," the device will launch a video call app and automatically make a call to a registered contact. Also, if the user says, "I want to check my next hospital appointment," the device will open a calendar app and read out the date.
[0352] Examples of prompt messages include, "Please display the next hospital appointment," and "Please make a video call to family member [name]."
[0353] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0354] Step 1:
[0355] The terminal acquires the user's voice input through a microphone device. This input is an audio signal for capturing the user's verbal instructions in real time. The format of the voice input is digitized audio data.
[0356] Step 2:
[0357] The device converts the acquired audio signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). This API analyzes the audio waveform and sequentially converts human speech into corresponding text. The input is audio data, and the output is text data in which the audio has been converted into a string.
[0358] Step 3:
[0359] The server utilizes a generative AI model (e.g., OpenAI's GPT series) to analyze the converted text data and identify the user's intended instructions. This analysis uses natural language processing techniques to understand the context and determine the necessary actions. The input is text data, and the output is executable processing information based on the analyzed instructions.
[0360] Step 4:
[0361] The device executes visual information and action simulations based on the specified instructions. For example, if told to "make a video call," it will launch the relevant application and automatically select the appropriate contact. Input is executable processing information, and output is the system's execution.
[0362] Step 5:
[0363] The device generates audio feedback to report the processing results to the user using speech conversion technology. An API (e.g., Amazon Polly) synthesizes text information into speech, generating sounds to communicate the current process and results to the user. The input is text information of the processing results, and the output is audio feedback.
[0364] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0365] The system of the present invention aims to enable users to operate their PCs more naturally and interactively using voice input and emotion recognition technology. Embodiments are described in detail below.
[0366] The user first speaks a voice command into the PC's microphone. For example, they might say, "Open the document, then start rehearsing the presentation." Once this voice is captured by the device, Speech-to-Text technology converts the speech into text. This ensures that the user's instructions are clearly understood.
[0367] The terminal captures the screen currently displayed on the display device and analyzes that information. It collects information about the displayed applications and windows and matches it with the user's voice commands. This identifies the target for performing the operation instructed by the user.
[0368] Furthermore, this system is equipped with an emotion engine that analyzes the user's emotions from their voice. The emotion engine analyzes, for example, the tone, speed, and volume of the voice to determine whether the user is excited, calm, or dissatisfied. As a result, the system adjusts its response according to the user's emotional state.
[0369] For example, if a user gives an instruction in a voice that conveys urgency, such as "Do something about this document right now!", the device will recognize the user's emotion as "urgency," execute the instruction more quickly than usual, and provide proactive feedback using speech synthesis technology, such as "I've opened the document immediately, what do you want to do next?"
[0370] Once an operation is performed, the terminal provides feedback to the user. Using speech synthesis technology, an easy-to-understand voice message is generated for the user, providing confirmation of the operation's success and advice on the next steps.
[0371] This system allows users not only to control their PCs with their voice, but also to experience customized interactions tailored to their emotional state, enabling them to enjoy a more comfortable and efficient operating environment.
[0372] The following describes the processing flow.
[0373] Step 1:
[0374] The user gives voice commands into the PC's microphone. For example, they might say, "Open my email and check for the latest incoming emails."
[0375] Step 2:
[0376] The device captures audio through its microphone, inputs the audio data into a Speech-to-Text engine in real time, and converts it into text. This text conversion allows for clear analysis of the user's instructions.
[0377] Step 3:
[0378] During the process of processing voice input, the device uses its built-in emotion engine to analyze the tone, speed, and volume of the voice to identify the user's emotional state. For example, states such as excitement, relaxation, and dissatisfaction are analyzed.
[0379] Step 4:
[0380] The device captures the current screen information and analyzes it using AI technology. Based on the analysis results, actions related to voice commands are identified. For example, it determines whether the email app is already open or if it needs to be opened again.
[0381] Step 5:
[0382] The device performs actions determined by the analyzed voice commands and screen information. It simulates operations such as opening the email app and selecting the most recent received email from the list.
[0383] Step 6:
[0384] The device utilizes the results of the emotion engine to adjust feedback according to the user's emotional state. For example, if the user is expressing dissatisfaction, it will generate a response in a gentle tone such as, "I've seen your email. Is there anything I can help you with?"
[0385] Step 7:
[0386] The device uses speech synthesis technology to provide voice feedback on the results of the operations performed. The feedback is generated in real time, informing the user of the success of the operation and prompting them to take the next action.
[0387] This processing flow allows the system to quickly and accurately perform PC operations based on voice input and provide the user with appropriate responses that reflect their emotions.
[0388] (Example 2)
[0389] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0390] Traditional computer systems made it difficult for users to operate them using natural voice commands and failed to provide responses that took user emotions into consideration. This resulted in users experiencing stress during voice input and decreased operational efficiency. Furthermore, there was a lack of methods to flexibly apply operations based on information displayed on the screen.
[0391] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0392] In this invention, the server includes means for receiving user instructions using voice input, means for using a natural language processing device to convert the received voice into text, means for analyzing image information acquired from a display device to identify the target to be executed based on the user's instructions, and means for analyzing the user's emotions from the voice and adjusting the response based on those emotions. This enables the user to effectively operate the computer with a natural voice and receive a customized response that is tailored to their emotions.
[0393] "Voice input" is a method by which users give voice commands directly to a computer through a microphone.
[0394] A "natural language processing device" is a device equipped with technology to convert speech data into text, and has the function of making user voice instructions into a format that a computer can understand.
[0395] A "display device" is a device that displays information visually, such as a computer screen or monitor.
[0396] "Analyzing image information" is the process of analyzing screen captures and image data to understand what is displayed within them and extract specific data.
[0397] "Analyzing user emotions from voice" refers to a technology that analyzes the tone, speed, volume, and other elements contained in voice data to determine the user's emotional state.
[0398] A "generative AI technology model" is a model equipped with an algorithm that uses artificial intelligence to automatically generate content and feedback based on user instructions.
[0399] "Speech synthesis technology" is a technology that converts text data into speech, providing users with computer responses as voice.
[0400] This invention is a system that allows users to operate a computer using natural voice, enabling more human-like interaction. The system combines voice input, image information analysis, generative AI technology models, and speech synthesis technology.
[0401] First, the user gives voice instructions through the device's microphone, and the device receives this voice data. The voice is then converted into text using a natural language processing unit. In this process, common technologies can be used as natural language processing techniques, for example.
[0402] Next, the terminal acquires image information from the display device and analyzes its contents. For example, image processing software is used to identify the appropriate action to take based on the user's instructions. The terminal also incorporates an emotion analyzer that analyzes the user's emotions from voice data. For instance, it can determine whether the user is feeling anxious or calm.
[0403] By utilizing generative AI technology models, the device generates responses that correspond to the user's instructions and emotions. This generative AI model can be modified using common artificial intelligence algorithms. For example, if a user instructs, "Can you open my presentation materials and help me with something?", the device might generate a response such as, "I'll open the materials now, and then shall we practice the slides together?"
[0404] Finally, the device uses speech synthesis technology to provide the generated response to the user as audio. This process utilizes speech synthesis technology to provide the user with natural and easy-to-understand voice messages.
[0405] This system allows users to operate computers through voice commands and emotionally responsive interactions, resulting in improved operational efficiency and a richer user experience.
[0406] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0407] Step 1:
[0408] The user gives voice commands to the device. For example, they might say, "Open the presentation materials." The device receives this voice data as input through its built-in microphone. The voice data is then converted from a sound waveform into digital audio data.
[0409] Step 2:
[0410] The terminal transmits the received audio data to a natural language processing unit, which converts the audio into text. A speech recognition engine is used here to extract word sequences from the audio waveform as output. This converted text serves as the basic data for understanding the user's instructions.
[0411] Step 3:
[0412] The terminal acquires image information from the display device and analyzes it to associate it with user instructions. For example, image processing software performs a screen capture and extracts information about the active window and application. This analysis identifies actions that can be performed based on the user's instructions.
[0413] Step 4:
[0414] The device uses a voice analysis engine to evaluate the user's emotional state from voice data. Parameters such as voice tone, speed, and volume are used as input data, and the emotional state is output. For example, it can determine whether the user is speaking in an expectant voice or in an anxious voice.
[0415] Step 5:
[0416] The device utilizes a generative AI technology model to generate responses that correspond to the user's emotional state and instructions. Input includes text data and emotional information obtained in the previous step, and the output is an appropriate response. For example, it might generate a message like, "I've opened the document. What should I do next?"
[0417] Step 6:
[0418] The device converts text responses generated using speech synthesis technology into speech. The synthesized speech data is provided as output for communication with the user and played back from the device to the user as voice feedback. This process ensures that the user receives voice feedback appropriate to the instructions given.
[0419] (Application Example 2)
[0420] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0421] Current consumer robots have limitations in that voice control is restricted to simple commands, and they cannot provide flexible responses or adjust operations based on the user's emotional state. Furthermore, while users desire more natural communication, existing systems are unable to meet this need.
[0422] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0423] In this invention, the server includes means for receiving operation instructions using voice input, means for converting the received voice into text data, and means for analyzing video information displayed on a display device to identify executable actions corresponding to the instructions. This enables operation and feedback in accordance with the user's voice instructions and emotional state.
[0424] "Voice input" is a method for capturing voice information spoken by the user.
[0425] "Operation instructions" refer to the instructions a user gives to a robot to perform a specific action.
[0426] "Converting to text data" refers to the process of analyzing audio information after it has been input and converting it into a corresponding string of characters.
[0427] "Display devices" refer to devices that output information visually, such as screens and displays.
[0428] "Analyzing video information" is the process of analyzing the displayed visual information to obtain the necessary data.
[0429] "Identifying feasible actions" means identifying specific tasks that should be performed based on the analyzed information and instructions.
[0430] An "input device" is a device used to receive data or instructions, and includes keyboards and mice.
[0431] "To mimic an operation" means to reproduce an actual mechanical or electronic action based on an identified movement.
[0432] "Analyzing emotional state" means analyzing the tone and speed of voice from audio data to identify the user's emotions.
[0433] "Adjusting the speed of actions and response content" refers to the process of changing the speed of actions performed and the content of voice feedback according to the user's emotions.
[0434] This invention aims to achieve advanced interaction in consumer robots by combining voice input and emotion recognition technology. This system receives voice commands from the user, converts them into text data, and processes them in the robot's control unit. The details are described below.
[0435] The server first receives operation instructions from the user via a voice input device. The voice data is processed in real time and converted into text data using the Google Speech-to-Text API. This conversion process ensures that each instruction is interpreted as a clear digital command.
[0436] Furthermore, the server executes a Python script that implements an emotion recognition algorithm to analyze the user's emotional state along with the text data. Specifically, it analyzes acoustic features such as voice tone, speed, and volume to identify whether the user is in a particular emotional state, such as excitement, relaxation, or anxiety. This allows the server to adjust the operation speed and feedback content as needed.
[0437] Based on this information, the robot quickly identifies the specified action to analyze the displayed video information and generates an action scenario. The server sends commands to the robot's control unit, which then performs physical actions such as household chores according to the user's instructions. Simultaneously, it generates audio feedback using Amazon Polly to notify the user of the results and the next steps.
[0438] For example, if a user instructs the robot to "clean the kitchen and then prepare dinner," the robot will convert the voice into text, perform the kitchen cleaning task, check the ingredients in the refrigerator, and begin preparing the necessary ingredients. During the process, the robot will provide feedback such as, "I've started cleaning the kitchen. What can I help you with next?"
[0439] Examples of prompt statements that can be used as input to a generative AI model include the following:
[0440] "When a user gives a voice command with emotion, please convert that command into text and explain the action that corresponds to that emotion."
[0441] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0442] Step 1:
[0443] The terminal receives voice commands spoken by the user via the microphone. The input is an audio signal, and the output is captured by the terminal as a digital signal. This is the process of acquiring audio data.
[0444] Step 2:
[0445] The device converts the received audio signal into text data using the Google Speech-to-Text API. The input is an audio signal, and the output is the converted string. This conversion process involves phonological analysis and formatting of specific instructions.
[0446] Step 3:
[0447] The server analyzes the acquired text data to identify the specific tasks instructed by the user. The input is text data, and the output is a task list. This includes interpreting the meaning of the instructions using natural language processing techniques.
[0448] Step 4:
[0449] The device applies an emotion recognition algorithm to voice data to analyze the user's emotional state. The input is voice data, and the output is information about the emotional state. It performs specific actions to identify the user's emotions by analyzing the tone, speed, and intensity of the voice.
[0450] Step 5:
[0451] For each task, it generates instructions to adjust the execution speed and priority based on the user's emotional state. The input is a task list and emotional information, and the output is an adjusted task list. This includes actions to determine timing and order according to the user's emotions.
[0452] Step 6:
[0453] The terminal sends these instructions to the robot's control unit. The input is a coordinated task list, and the output is specific control signals to the robot. It is a process of instructing physical actions using a communication protocol.
[0454] Step 7:
[0455] The robot initiates a specified action based on a control signal, and its progress is monitored by a server during the operation. The input is the control signal, and the output is progress data. This includes performing actual physical tasks and tracking the results in real time.
[0456] Step 8:
[0457] The server receives progress data and provides voice feedback to the user regarding the progress and results via Amazon Polly. The input is progress data, and the output is a generated voice message. It is a voice generation process designed to communicate results in an accessible way.
[0458] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0459] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0460] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0461] [Third Embodiment]
[0462] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0463] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0464] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0465] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0466] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0467] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0468] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0469] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0470] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0471] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0472] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0473] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0474] The system of the present invention enables the user to operate a PC through voice input. Its embodiments are described in detail below.
[0475] The system begins with the user giving a voice command into the microphone. For example, consider the command, "Open a web browser and search for news." The terminal receives the voice input and converts it into text using Speech-to-Text technology. The converted text is then used for subsequent processing.
[0476] The device captures the user's PC screen as an image and analyzes the displayed windows, icons, and text fields. This analysis uses generative AI technology. Through this analysis, the system identifies how the user's voice commands are translated into screen operations.
[0477] For example, if the device receives the voice command, "Open a web browser and search for news," it will first display an existing browser window or launch a new browser, and then enter the keyword "news" into the URL bar. It will also automatically simulate pressing the Enter key to perform the search, if necessary.
[0478] The device uses speech synthesis technology to provide feedback to the user about the operations performed. A voice message such as "News search performed" is output, allowing the user to confirm that the instructions were carried out correctly.
[0479] Furthermore, based on the screen information it analyzes, the device provides flexible voice control for newly installed applications and added features. This versatility allows users to continue using their PC without having to learn new operating methods.
[0480] This system allows users to operate their PCs through natural voice interaction without using physical input devices. This is particularly useful for users who have difficulty with manual operation, and also provides a comfortable and efficient operating environment for general users.
[0481] The following describes the processing flow.
[0482] Step 1:
[0483] The user gives voice commands into the PC's microphone. For example, they might say, "Open a text editor and create a new document."
[0484] Step 2:
[0485] The device captures voice input from the microphone and immediately performs Speech-to-Text processing using its speech recognition engine. In this process, the voice data is converted into text data.
[0486] Step 3:
[0487] The device captures the screen and uses AI to analyze information about currently displayed windows and icons. This allows it to identify which applications and windows are open.
[0488] Step 4:
[0489] The device determines what action needs to be performed based on the analyzed screen information and text converted from the audio. If the user instructs it to "open a text editor," it checks if an editor is already open and launches a new one if it's not.
[0490] Step 5:
[0491] The terminal simulates and performs keyboard input and mouse operations based on the specified action. For example, if the instruction is "Create a new document," it selects the "New" menu in the text editor and places the cursor where to input text in the document.
[0492] Step 6:
[0493] The device uses speech synthesis technology to provide feedback to the user on whether the operation was performed successfully. It outputs a voice message such as "A new document has been created" to inform the user of the result.
[0494] Step 7:
[0495] Based on the feedback received, the user can either give further instructions or terminate the process. By referring to the feedback, the user can verify that the instructions were followed correctly.
[0496] (Example 1)
[0497] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0498] In voice-based operation assistance technologies, it is crucial to provide an environment where users can easily perform complex tasks. However, conventional technologies have limitations in terms of voice input accuracy and responsiveness, making it particularly difficult to flexibly adapt to new functions and applications. Furthermore, it is difficult to accurately simulate on-screen actions, resulting in a lack of quick and clear feedback to users. There is a need to solve these problems and provide an efficient and natural voice-based interface.
[0499] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0500] In this invention, the server includes means for acquiring user instructions using acoustic signals, means for converting the acquired acoustic signals into text, and means for analyzing video information output to a display means and determining executable actions in accordance with the user's instructions. This allows the user to operate the computer naturally through voice, receive accurate feedback while flexibly adapting to new functions and applications.
[0501] An "acoustic signal" is an electrical signal that represents air vibrations, including the voice of a user.
[0502] "User" refers to an individual or entity that operates the system through voice input.
[0503] "Display means" refers to devices or media used to provide information visually, and is usually a computer monitor or screen.
[0504] "Visual information" refers to digital data or screen content that is output to a display device and is visually recognizable.
[0505] "Converting to text" is the process of analyzing an audio signal and converting it into text data that corresponds to the speech.
[0506] "Analysis" refers to the process of investigating video information and data to understand its structure and content.
[0507] "Executable actions" refer to the specific operations and processes that a computer can perform based on user instructions.
[0508] "Simulating" is the process of imitating the operation of actual input devices in order to obtain equivalent effects.
[0509] "Responding" refers to the act of conveying results or information in response to a user's instructions, and is often provided as audio or visual feedback.
[0510] The present invention provides a system that assists user operation through voice input. Specifically, the user first emits an acoustic signal, i.e., a voice command, into the microphone of the terminal. The terminal receives this acoustic signal and converts the voice into text using voice recognition software technology. The technology used at this stage includes general voice recognition APIs.
[0511] Next, the terminal acquires the video information displayed on the display device and uses a generative AI model to analyze it. The generative AI model includes a system known as a common natural language processing technique, which interprets and converts user instructions into executable actions. This is achieved by analyzing the arrangement of icons and the state of windows on the screen.
[0512] For example, based on a prompt message such as "Open your email and compose a new message," the system will automatically launch the email application and simulate opening a new message window. During this process, necessary mouse clicks and keyboard inputs will also be simulated by the terminal.
[0513] Once processing is complete, the terminal provides voice feedback to the user. This is done using synthesized speech technology, and the user receives a voice message such as, "New message creation complete." In this way, the user's voice commands are executed through the system, resulting in efficient and natural operation.
[0514] Furthermore, this system can flexibly adapt to newly installed programs and functions. Therefore, users do not need to relearn how to operate new applications when they are added. In this way, users can enjoy an intuitive and user-friendly computing environment that allows them to control a variety of programs using only their voice.
[0515] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0516] Step 1:
[0517] Acquiring voice input
[0518] The user gives clear voice instructions into the microphone. For example, they might say, "Open the document and add a new page."
[0519] The input is an acoustic signal from the user.
[0520] The output is raw audio data received by the terminal as an acoustic signal.
[0521] Step 2:
[0522] Speech-to-text conversion
[0523] The terminal converts the received audio data into text data using speech recognition software (such as a speech recognition API).
[0524] The input is raw audio data as an acoustic signal.
[0525] As part of the data processing, the acoustic signal is analyzed and the audio pattern is converted into a corresponding string of characters.
[0526] The output is text data corresponding to the voice instructions.
[0527] Step 3:
[0528] Acquisition and analysis of video information
[0529] The terminal captures screen information from the display device and performs analysis using a generated AI model.
[0530] The input consists of captured video data and converted text data.
[0531] As a data calculation tool, it analyzes the position and state of elements on the screen and generates data to determine actions that correspond to user instructions.
[0532] The output consists of specific operating instructions based on the analysis results.
[0533] Step 4:
[0534] Execution of specific operations
[0535] The terminal performs necessary operations based on the analysis results and simulates the operation of input devices. For example, it opens the appropriate application and performs the necessary key inputs and mouse operations to use the corresponding functions.
[0536] The input consists of specific instructions for operation.
[0537] The output is the result of the user's intended operation being executed on the device.
[0538] Step 5:
[0539] Provide feedback
[0540] The device uses voice generation technology to communicate to the user that the operation is complete via an acoustic signal.
[0541] The input is a message indicating the completion of the operation.
[0542] The output is an acoustic signal that is reproduced as audio feedback.
[0543] Step 6:
[0544] Adapting to new features
[0545] The device also analyzes voice commands for newly installed functions and applications, enabling them to be operated.
[0546] The input is new video information to be analyzed and adapted to.
[0547] The output is an operable environment for new functions.
[0548] (Application Example 1)
[0549] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0550] In modern information technology, many users experience difficulties accessing digital services and applications due to manual interface operation. In particular, for users with physical limitations or those unfamiliar with manual operation, a natural and intuitive voice-based interface is useful, but current technology offers limited, versatile, and flexible solutions. This invention aims to solve these problems and provide a system that allows users to efficiently operate digital services via voice.
[0551] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0552] In this invention, the system includes means for acquiring user instructions using voice input, means for converting the acquired voice into text, and means for analyzing visual information output to a display means to determine executable processing based on the user's instructions. This enables users to intuitively operate digital services using only their voice.
[0553] "Voice input" is a method by which a system can receive instructions and information spoken by the user.
[0554] "User instructions" refer to the specific content that a user issues to request a particular operation or process from the system.
[0555] "Acquired audio" refers to the audio signal received through audio input.
[0556] "Converting to a string" means processing an audio signal and converting it into data in a corresponding text format.
[0557] "Display means" refers to a device or function for visually presenting information, and includes monitors and displays.
[0558] "Visual information" refers to data such as images and text that are presented visually through a display means.
[0559] "Analyzing" means scrutinizing data in order to understand the information and recognize the necessary structure.
[0560] "Executable processes" refer to operations and actions that a system can perform based on user instructions.
[0561] "Reproducing the operation of a device" means simulating the operations performed by an actual device within a system.
[0562] "Reporting" means that the system communicates the results or status of its actions to the user.
[0563] "Speech conversion technology" is a technology that converts text information into a speech format.
[0564] "Communicating through sound" means providing information to users aurally.
[0565] "New features and programs" refer to new operations or uses added to the system.
[0566] The system that realizes this application will enable users to intuitively manipulate digital information through a series of processes including speech recognition, data analysis, and feedback provision. Specifically, it will utilize smart devices such as smart glasses, smartphones, or other computing devices.
[0567] The device first receives the user's voice input via the microphone and converts it into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). Next, it uses a generative AI model (e.g., OpenAI's GPT series) to identify the user's instructions from the converted text and analyzes the visual information based on those instructions.
[0568] Based on the analyzed visual information, the device determines and reproduces the executable actions required in response to the user's request. After the determined actions are performed, the results are reported to the user via voice feedback. This voice feedback is generated using an API that employs speech conversion technology (e.g., Amazon Polly).
[0569] For example, if a user says, "I want to contact my family," the device will launch a video call app and automatically make a call to a registered contact. Also, if the user says, "I want to check my next hospital appointment," the device will open a calendar app and read out the date.
[0570] Examples of prompt messages include, "Please display the next hospital appointment," and "Please make a video call to family member [name]."
[0571] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0572] Step 1:
[0573] The terminal acquires the user's voice input through a microphone device. This input is an audio signal for capturing the user's verbal instructions in real time. The format of the voice input is digitized audio data.
[0574] Step 2:
[0575] The device converts the acquired audio signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). This API analyzes the audio waveform and sequentially converts human speech into corresponding text. The input is audio data, and the output is text data in which the audio has been converted into a string.
[0576] Step 3:
[0577] The server utilizes a generative AI model (e.g., OpenAI's GPT series) to analyze the converted text data and identify the user's intended instructions. This analysis uses natural language processing techniques to understand the context and determine the necessary actions. The input is text data, and the output is executable processing information based on the analyzed instructions.
[0578] Step 4:
[0579] The device executes visual information and action simulations based on the specified instructions. For example, if told to "make a video call," it will launch the relevant application and automatically select the appropriate contact. Input is executable processing information, and output is the system's execution.
[0580] Step 5:
[0581] The device generates audio feedback to report the processing results to the user using speech conversion technology. An API (e.g., Amazon Polly) synthesizes text information into speech, generating sounds to communicate the current process and results to the user. The input is text information of the processing results, and the output is audio feedback.
[0582] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0583] The system of the present invention aims to enable users to operate their PCs more naturally and interactively using voice input and emotion recognition technology. Embodiments are described in detail below.
[0584] The user first speaks a voice command into the PC's microphone. For example, they might say, "Open the document, then start rehearsing the presentation." Once this voice is captured by the device, Speech-to-Text technology converts the speech into text. This ensures that the user's instructions are clearly understood.
[0585] The terminal captures the screen currently displayed on the display device and analyzes that information. It collects information about the displayed applications and windows and matches it with the user's voice commands. This identifies the target for performing the operation instructed by the user.
[0586] Furthermore, this system is equipped with an emotion engine that analyzes the user's emotions from their voice. The emotion engine analyzes, for example, the tone, speed, and volume of the voice to determine whether the user is excited, calm, or dissatisfied. As a result, the system adjusts its response according to the user's emotional state.
[0587] For example, if a user gives an instruction in a voice that conveys urgency, such as "Do something about this document right now!", the device will recognize the user's emotion as "urgency," execute the instruction more quickly than usual, and provide proactive feedback using speech synthesis technology, such as "I've opened the document immediately, what do you want to do next?"
[0588] Once an operation is performed, the terminal provides feedback to the user. Using speech synthesis technology, an easy-to-understand voice message is generated for the user, providing confirmation of the operation's success and advice on the next steps.
[0589] This system allows users not only to control their PCs with their voice, but also to experience customized interactions tailored to their emotional state, enabling them to enjoy a more comfortable and efficient operating environment.
[0590] The following describes the processing flow.
[0591] Step 1:
[0592] The user gives voice commands into the PC's microphone. For example, they might say, "Open my email and check for the latest incoming emails."
[0593] Step 2:
[0594] The device captures audio through its microphone, inputs the audio data into a Speech-to-Text engine in real time, and converts it into text. This text conversion allows for clear analysis of the user's instructions.
[0595] Step 3:
[0596] During the process of processing voice input, the device uses its built-in emotion engine to analyze the tone, speed, and volume of the voice to identify the user's emotional state. For example, states such as excitement, relaxation, and dissatisfaction are analyzed.
[0597] Step 4:
[0598] The device captures the current screen information and analyzes it using AI technology. Based on the analysis results, actions related to voice commands are identified. For example, it determines whether the email app is already open or if it needs to be opened again.
[0599] Step 5:
[0600] The device performs actions determined by the analyzed voice commands and screen information. It simulates operations such as opening the email app and selecting the most recent received email from the list.
[0601] Step 6:
[0602] The device utilizes the results of the emotion engine to adjust feedback according to the user's emotional state. For example, if the user is expressing dissatisfaction, it will generate a response in a gentle tone such as, "I've seen your email. Is there anything I can help you with?"
[0603] Step 7:
[0604] The device uses speech synthesis technology to provide voice feedback on the results of the operations performed. The feedback is generated in real time, informing the user of the success of the operation and prompting them to take the next action.
[0605] This processing flow allows the system to quickly and accurately perform PC operations based on voice input and provide the user with appropriate responses that reflect their emotions.
[0606] (Example 2)
[0607] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0608] Traditional computer systems made it difficult for users to operate them using natural voice commands and failed to provide responses that took user emotions into consideration. This resulted in users experiencing stress during voice input and decreased operational efficiency. Furthermore, there was a lack of methods to flexibly apply operations based on information displayed on the screen.
[0609] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0610] In this invention, the server includes means for receiving user instructions using voice input, means for using a natural language processing device to convert the received voice into text, means for analyzing image information acquired from a display device to identify the target to be executed based on the user's instructions, and means for analyzing the user's emotions from the voice and adjusting the response based on those emotions. This enables the user to effectively operate the computer with a natural voice and receive a customized response that is tailored to their emotions.
[0611] "Voice input" is a method by which users give voice commands directly to a computer through a microphone.
[0612] A "natural language processing device" is a device equipped with technology to convert speech data into text, and has the function of making user voice instructions into a format that a computer can understand.
[0613] A "display device" is a device that displays information visually, such as a computer screen or monitor.
[0614] "Analyzing image information" is the process of analyzing screen captures and image data to understand what is displayed within them and extract specific data.
[0615] "Analyzing user emotions from voice" refers to a technology that analyzes the tone, speed, volume, and other elements contained in voice data to determine the user's emotional state.
[0616] A "generative AI technology model" is a model equipped with an algorithm that uses artificial intelligence to automatically generate content and feedback based on user instructions.
[0617] "Speech synthesis technology" is a technology that converts text data into speech, providing users with computer responses as voice.
[0618] This invention is a system that allows users to operate a computer using natural voice, enabling more human-like interaction. The system combines voice input, image information analysis, generative AI technology models, and speech synthesis technology.
[0619] First, the user gives voice instructions through the device's microphone, and the device receives this voice data. The voice is then converted into text using a natural language processing unit. In this process, common technologies can be used as natural language processing techniques, for example.
[0620] Next, the terminal acquires image information from the display device and analyzes its contents. For example, image processing software is used to identify the appropriate action to take based on the user's instructions. The terminal also incorporates an emotion analyzer that analyzes the user's emotions from voice data. For instance, it can determine whether the user is feeling anxious or calm.
[0621] By utilizing generative AI technology models, the device generates responses that correspond to the user's instructions and emotions. This generative AI model can be modified using common artificial intelligence algorithms. For example, if a user instructs, "Can you open my presentation materials and help me with something?", the device might generate a response such as, "I'll open the materials now, and then shall we practice the slides together?"
[0622] Finally, the device uses speech synthesis technology to provide the generated response to the user as audio. This process utilizes speech synthesis technology to provide the user with natural and easy-to-understand voice messages.
[0623] This system allows users to operate computers through voice commands and emotionally responsive interactions, resulting in improved operational efficiency and a richer user experience.
[0624] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0625] Step 1:
[0626] The user gives voice commands to the device. For example, they might say, "Open the presentation materials." The device receives this voice data as input through its built-in microphone. The voice data is then converted from a sound waveform into digital audio data.
[0627] Step 2:
[0628] The terminal transmits the received audio data to a natural language processing unit, which converts the audio into text. A speech recognition engine is used here to extract word sequences from the audio waveform as output. This converted text serves as the basic data for understanding the user's instructions.
[0629] Step 3:
[0630] The terminal acquires image information from the display device and analyzes it to associate it with user instructions. For example, image processing software performs a screen capture and extracts information about the active window and application. This analysis identifies actions that can be performed based on the user's instructions.
[0631] Step 4:
[0632] The device uses a voice analysis engine to evaluate the user's emotional state from voice data. Parameters such as voice tone, speed, and volume are used as input data, and the emotional state is output. For example, it can determine whether the user is speaking in an expectant voice or in an anxious voice.
[0633] Step 5:
[0634] The device utilizes a generative AI technology model to generate responses that correspond to the user's emotional state and instructions. Input includes text data and emotional information obtained in the previous step, and the output is an appropriate response. For example, it might generate a message like, "I've opened the document. What should I do next?"
[0635] Step 6:
[0636] The device converts text responses generated using speech synthesis technology into speech. The synthesized speech data is provided as output for communication with the user and played back from the device to the user as voice feedback. This process ensures that the user receives voice feedback appropriate to the instructions given.
[0637] (Application Example 2)
[0638] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0639] Current consumer robots have limitations in that voice control is restricted to simple commands, and they cannot provide flexible responses or adjust operations based on the user's emotional state. Furthermore, while users desire more natural communication, existing systems are unable to meet this need.
[0640] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0641] In this invention, the server includes means for receiving operation instructions using voice input, means for converting the received voice into text data, and means for analyzing video information displayed on a display device to identify executable actions corresponding to the instructions. This enables operation and feedback in accordance with the user's voice instructions and emotional state.
[0642] "Voice input" is a method for capturing voice information spoken by the user.
[0643] "Operation instructions" refer to the instructions a user gives to a robot to perform a specific action.
[0644] "Converting to text data" refers to the process of analyzing audio information after it has been input and converting it into a corresponding string of characters.
[0645] "Display devices" refer to devices that output information visually, such as screens and displays.
[0646] "Analyzing video information" is the process of analyzing the displayed visual information to obtain the necessary data.
[0647] "Identifying feasible actions" means identifying specific tasks that should be performed based on the analyzed information and instructions.
[0648] An "input device" is a device used to receive data or instructions, and includes keyboards and mice.
[0649] "To mimic an operation" means to reproduce an actual mechanical or electronic action based on an identified movement.
[0650] "Analyzing emotional state" means analyzing the tone and speed of voice from audio data to identify the user's emotions.
[0651] "Adjusting the speed of actions and response content" refers to the process of changing the speed of actions performed and the content of voice feedback according to the user's emotions.
[0652] This invention aims to achieve advanced interaction in consumer robots by combining voice input and emotion recognition technology. This system receives voice commands from the user, converts them into text data, and processes them in the robot's control unit. The details are described below.
[0653] The server first receives operation instructions from the user via a voice input device. The voice data is processed in real time and converted into text data using the Google Speech-to-Text API. This conversion process ensures that each instruction is interpreted as a clear digital command.
[0654] Furthermore, the server executes a Python script that implements an emotion recognition algorithm to analyze the user's emotional state along with the text data. Specifically, it analyzes acoustic features such as voice tone, speed, and volume to identify whether the user is in a particular emotional state, such as excitement, relaxation, or anxiety. This allows the server to adjust the operation speed and feedback content as needed.
[0655] Based on this information, the robot quickly identifies the specified action to analyze the displayed video information and generates an action scenario. The server sends commands to the robot's control unit, which then performs physical actions such as household chores according to the user's instructions. Simultaneously, it generates audio feedback using Amazon Polly to notify the user of the results and the next steps.
[0656] For example, if a user instructs the robot to "clean the kitchen and then prepare dinner," the robot will convert the voice into text, perform the kitchen cleaning task, check the ingredients in the refrigerator, and begin preparing the necessary ingredients. During the process, the robot will provide feedback such as, "I've started cleaning the kitchen. What can I help you with next?"
[0657] Examples of prompt statements that can be used as input to a generative AI model include the following:
[0658] "When a user gives a voice command with emotion, please convert that command into text and explain the action that corresponds to that emotion."
[0659] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0660] Step 1:
[0661] The terminal receives voice commands spoken by the user via the microphone. The input is an audio signal, and the output is captured by the terminal as a digital signal. This is the process of acquiring audio data.
[0662] Step 2:
[0663] The device converts the received audio signal into text data using the Google Speech-to-Text API. The input is an audio signal, and the output is the converted string. This conversion process involves phonological analysis and formatting of specific instructions.
[0664] Step 3:
[0665] The server analyzes the acquired text data to identify the specific tasks instructed by the user. The input is text data, and the output is a task list. This includes interpreting the meaning of the instructions using natural language processing techniques.
[0666] Step 4:
[0667] The device applies an emotion recognition algorithm to voice data to analyze the user's emotional state. The input is voice data, and the output is information about the emotional state. It performs specific actions to identify the user's emotions by analyzing the tone, speed, and intensity of the voice.
[0668] Step 5:
[0669] For each task, it generates instructions to adjust the execution speed and priority based on the user's emotional state. The input is a task list and emotional information, and the output is an adjusted task list. This includes actions to determine timing and order according to the user's emotions.
[0670] Step 6:
[0671] The terminal sends these instructions to the robot's control unit. The input is a coordinated task list, and the output is specific control signals to the robot. It is a process of instructing physical actions using a communication protocol.
[0672] Step 7:
[0673] The robot initiates a specified action based on a control signal, and its progress is monitored by a server during the operation. The input is the control signal, and the output is progress data. This includes performing actual physical tasks and tracking the results in real time.
[0674] Step 8:
[0675] The server receives progress data and provides voice feedback to the user regarding the progress and results via Amazon Polly. The input is progress data, and the output is a generated voice message. It is a voice generation process designed to communicate results in an accessible way.
[0676] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0677] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0678] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0679] [Fourth Embodiment]
[0680] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0681] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0682] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0683] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0684] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0685] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0686] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0687] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0688] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0689] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0690] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0691] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0692] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0693] The system of the present invention enables the user to operate a PC through voice input. Its embodiments are described in detail below.
[0694] The system begins with the user giving a voice command into the microphone. For example, consider the command, "Open a web browser and search for news." The terminal receives the voice input and converts it into text using Speech-to-Text technology. The converted text is then used for subsequent processing.
[0695] The device captures the user's PC screen as an image and analyzes the displayed windows, icons, and text fields. This analysis uses generative AI technology. Through this analysis, the system identifies how the user's voice commands are translated into screen operations.
[0696] For example, if the device receives the voice command, "Open a web browser and search for news," it will first display an existing browser window or launch a new browser, and then enter the keyword "news" into the URL bar. It will also automatically simulate pressing the Enter key to perform the search, if necessary.
[0697] The device uses speech synthesis technology to provide feedback to the user about the operations performed. A voice message such as "News search performed" is output, allowing the user to confirm that the instructions were carried out correctly.
[0698] Furthermore, based on the screen information it analyzes, the device provides flexible voice control for newly installed applications and added features. This versatility allows users to continue using their PC without having to learn new operating methods.
[0699] This system allows users to operate their PCs through natural voice interaction without using physical input devices. This is particularly useful for users who have difficulty with manual operation, and also provides a comfortable and efficient operating environment for general users.
[0700] The following describes the processing flow.
[0701] Step 1:
[0702] The user gives voice commands into the PC's microphone. For example, they might say, "Open a text editor and create a new document."
[0703] Step 2:
[0704] The device captures voice input from the microphone and immediately performs Speech-to-Text processing using its speech recognition engine. In this process, the voice data is converted into text data.
[0705] Step 3:
[0706] The device captures the screen and uses AI to analyze information about currently displayed windows and icons. This allows it to identify which applications and windows are open.
[0707] Step 4:
[0708] The device determines what action needs to be performed based on the analyzed screen information and text converted from the audio. If the user instructs it to "open a text editor," it checks if an editor is already open and launches a new one if it's not.
[0709] Step 5:
[0710] The terminal simulates and performs keyboard input and mouse operations based on the specified action. For example, if the instruction is "Create a new document," it selects the "New" menu in the text editor and places the cursor where to input text in the document.
[0711] Step 6:
[0712] The device uses speech synthesis technology to provide feedback to the user on whether the operation was performed successfully. It outputs a voice message such as "A new document has been created" to inform the user of the result.
[0713] Step 7:
[0714] Based on the feedback received, the user can either give further instructions or terminate the process. By referring to the feedback, the user can verify that the instructions were followed correctly.
[0715] (Example 1)
[0716] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0717] In voice-based operation assistance technologies, it is crucial to provide an environment where users can easily perform complex tasks. However, conventional technologies have limitations in terms of voice input accuracy and responsiveness, making it particularly difficult to flexibly adapt to new functions and applications. Furthermore, it is difficult to accurately simulate on-screen actions, resulting in a lack of quick and clear feedback to users. There is a need to solve these problems and provide an efficient and natural voice-based interface.
[0718] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0719] In this invention, the server includes means for acquiring user instructions using acoustic signals, means for converting the acquired acoustic signals into text, and means for analyzing video information output to a display means and determining executable actions in accordance with the user's instructions. This allows the user to operate the computer naturally through voice, receive accurate feedback while flexibly adapting to new functions and applications.
[0720] An "acoustic signal" is an electrical signal that represents air vibrations, including the voice of a user.
[0721] "User" refers to an individual or entity that operates the system through voice input.
[0722] "Display means" refers to devices or media used to provide information visually, and is usually a computer monitor or screen.
[0723] "Visual information" refers to digital data or screen content that is output to a display device and is visually recognizable.
[0724] "Converting to text" is the process of analyzing an audio signal and converting it into text data that corresponds to the speech.
[0725] "Analysis" refers to the process of investigating video information and data to understand its structure and content.
[0726] "Executable actions" refer to the specific operations and processes that a computer can perform based on user instructions.
[0727] "Simulating" is the process of imitating the operation of actual input devices in order to obtain equivalent effects.
[0728] "Responding" refers to the act of conveying results or information in response to a user's instructions, and is often provided as audio or visual feedback.
[0729] The present invention provides a system that assists user operation through voice input. Specifically, the user first emits an acoustic signal, i.e., a voice command, into the microphone of the terminal. The terminal receives this acoustic signal and converts the voice into text using voice recognition software technology. The technology used at this stage includes general voice recognition APIs.
[0730] Next, the terminal acquires the video information displayed on the display device and uses a generative AI model to analyze it. The generative AI model includes a system known as a common natural language processing technique, which interprets and converts user instructions into executable actions. This is achieved by analyzing the arrangement of icons and the state of windows on the screen.
[0731] For example, based on a prompt message such as "Open your email and compose a new message," the system will automatically launch the email application and simulate opening a new message window. During this process, necessary mouse clicks and keyboard inputs will also be simulated by the terminal.
[0732] Once processing is complete, the terminal provides voice feedback to the user. This is done using synthesized speech technology, and the user receives a voice message such as, "New message creation complete." In this way, the user's voice commands are executed through the system, resulting in efficient and natural operation.
[0733] Furthermore, this system can flexibly adapt to newly installed programs and functions. Therefore, users do not need to relearn how to operate new applications when they are added. In this way, users can enjoy an intuitive and user-friendly computing environment that allows them to control a variety of programs using only their voice.
[0734] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0735] Step 1:
[0736] Acquiring voice input
[0737] The user gives clear voice instructions into the microphone. For example, they might say, "Open the document and add a new page."
[0738] The input is an acoustic signal from the user.
[0739] The output is raw audio data received by the terminal as an acoustic signal.
[0740] Step 2:
[0741] Speech-to-text conversion
[0742] The terminal converts the received audio data into text data using speech recognition software (such as a speech recognition API).
[0743] The input is raw audio data as an acoustic signal.
[0744] As part of the data processing, the acoustic signal is analyzed and the audio pattern is converted into a corresponding string of characters.
[0745] The output is text data corresponding to the voice instructions.
[0746] Step 3:
[0747] Acquisition and analysis of video information
[0748] The terminal captures screen information from the display device and performs analysis using a generated AI model.
[0749] The input consists of captured video data and converted text data.
[0750] As a data calculation tool, it analyzes the position and state of elements on the screen and generates data to determine actions that correspond to user instructions.
[0751] The output consists of specific operating instructions based on the analysis results.
[0752] Step 4:
[0753] Execution of specific operations
[0754] The terminal performs necessary operations based on the analysis results and simulates the operation of input devices. For example, it opens the appropriate application and performs the necessary key inputs and mouse operations to use the corresponding functions.
[0755] The input consists of specific instructions for operation.
[0756] The output is the result of the user's intended operation being executed on the device.
[0757] Step 5:
[0758] Provide feedback
[0759] The device uses voice generation technology to communicate to the user that the operation is complete via an acoustic signal.
[0760] The input is a message indicating the completion of the operation.
[0761] The output is an acoustic signal that is reproduced as audio feedback.
[0762] Step 6:
[0763] Adapting to new features
[0764] The device also analyzes voice commands for newly installed functions and applications, enabling them to be operated.
[0765] The input is new video information to be analyzed and adapted to.
[0766] The output is an operable environment for new functions.
[0767] (Application Example 1)
[0768] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0769] In modern information technology, many users experience difficulties accessing digital services and applications due to manual interface operation. In particular, for users with physical limitations or those unfamiliar with manual operation, a natural and intuitive voice-based interface is useful, but current technology offers limited, versatile, and flexible solutions. This invention aims to solve these problems and provide a system that allows users to efficiently operate digital services via voice.
[0770] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0771] In this invention, the system includes means for acquiring user instructions using voice input, means for converting the acquired voice into text, and means for analyzing visual information output to a display means to determine executable processing based on the user's instructions. This enables users to intuitively operate digital services using only their voice.
[0772] "Voice input" is a method by which a system can receive instructions and information spoken by the user.
[0773] "User instructions" refer to the specific content that a user issues to request a particular operation or process from the system.
[0774] "Acquired audio" refers to the audio signal received through audio input.
[0775] "Converting to a string" means processing an audio signal and converting it into data in a corresponding text format.
[0776] "Display means" refers to a device or function for visually presenting information, and includes monitors and displays.
[0777] "Visual information" refers to data such as images and text that are presented visually through a display means.
[0778] "Analyzing" means scrutinizing data in order to understand the information and recognize the necessary structure.
[0779] "Executable processes" refer to operations and actions that a system can perform based on user instructions.
[0780] "Reproducing the operation of a device" means simulating the operations performed by an actual device within a system.
[0781] "Reporting" means that the system communicates the results or status of its actions to the user.
[0782] "Speech conversion technology" is a technology that converts text information into a speech format.
[0783] "Communicating through sound" means providing information to users aurally.
[0784] "New features and programs" refer to new operations or uses added to the system.
[0785] The system that realizes this application will enable users to intuitively manipulate digital information through a series of processes including speech recognition, data analysis, and feedback provision. Specifically, it will utilize smart devices such as smart glasses, smartphones, or other computing devices.
[0786] The device first receives the user's voice input via the microphone and converts it into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). Next, it uses a generative AI model (e.g., OpenAI's GPT series) to identify the user's instructions from the converted text and analyzes the visual information based on those instructions.
[0787] Based on the analyzed visual information, the device determines and reproduces the executable actions required in response to the user's request. After the determined actions are performed, the results are reported to the user via voice feedback. This voice feedback is generated using an API that employs speech conversion technology (e.g., Amazon Polly).
[0788] For example, if a user says, "I want to contact my family," the device will launch a video call app and automatically make a call to a registered contact. Also, if the user says, "I want to check my next hospital appointment," the device will open a calendar app and read out the date.
[0789] Examples of prompt messages include, "Please display the next hospital appointment," and "Please make a video call to family member [name]."
[0790] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0791] Step 1:
[0792] The terminal acquires the user's voice input through a microphone device. This input is an audio signal for capturing the user's verbal instructions in real time. The format of the voice input is digitized audio data.
[0793] Step 2:
[0794] The device converts the acquired audio signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). This API analyzes the audio waveform and sequentially converts human speech into corresponding text. The input is audio data, and the output is text data in which the audio has been converted into a string.
[0795] Step 3:
[0796] The server utilizes a generative AI model (e.g., OpenAI's GPT series) to analyze the converted text data and identify the user's intended instructions. This analysis uses natural language processing techniques to understand the context and determine the necessary actions. The input is text data, and the output is executable processing information based on the analyzed instructions.
[0797] Step 4:
[0798] The device executes visual information and action simulations based on the specified instructions. For example, if told to "make a video call," it will launch the relevant application and automatically select the appropriate contact. Input is executable processing information, and output is the system's execution.
[0799] Step 5:
[0800] The device generates audio feedback to report the processing results to the user using speech conversion technology. An API (e.g., Amazon Polly) synthesizes text information into speech, generating sounds to communicate the current process and results to the user. The input is text information of the processing results, and the output is audio feedback.
[0801] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0802] The system of the present invention aims to enable users to operate their PCs more naturally and interactively using voice input and emotion recognition technology. Embodiments are described in detail below.
[0803] The user first speaks a voice command into the PC's microphone. For example, they might say, "Open the document, then start rehearsing the presentation." Once this voice is captured by the device, Speech-to-Text technology converts the speech into text. This ensures that the user's instructions are clearly understood.
[0804] The terminal captures the screen currently displayed on the display device and analyzes that information. It collects information about the displayed applications and windows and matches it with the user's voice commands. This identifies the target for performing the operation instructed by the user.
[0805] Furthermore, this system is equipped with an emotion engine that analyzes the user's emotions from their voice. The emotion engine analyzes, for example, the tone, speed, and volume of the voice to determine whether the user is excited, calm, or dissatisfied. As a result, the system adjusts its response according to the user's emotional state.
[0806] For example, if a user gives an instruction in a voice that conveys urgency, such as "Do something about this document right now!", the device will recognize the user's emotion as "urgency," execute the instruction more quickly than usual, and provide proactive feedback using speech synthesis technology, such as "I've opened the document immediately, what do you want to do next?"
[0807] Once an operation is performed, the terminal provides feedback to the user. Using speech synthesis technology, an easy-to-understand voice message is generated for the user, providing confirmation of the operation's success and advice on the next steps.
[0808] This system allows users not only to control their PCs with their voice, but also to experience customized interactions tailored to their emotional state, enabling them to enjoy a more comfortable and efficient operating environment.
[0809] The following describes the processing flow.
[0810] Step 1:
[0811] The user gives voice commands into the PC's microphone. For example, they might say, "Open my email and check for the latest incoming emails."
[0812] Step 2:
[0813] The device captures audio through its microphone, inputs the audio data into a Speech-to-Text engine in real time, and converts it into text. This text conversion allows for clear analysis of the user's instructions.
[0814] Step 3:
[0815] During the process of processing voice input, the device uses its built-in emotion engine to analyze the tone, speed, and volume of the voice to identify the user's emotional state. For example, states such as excitement, relaxation, and dissatisfaction are analyzed.
[0816] Step 4:
[0817] The device captures the current screen information and analyzes it using AI technology. Based on the analysis results, actions related to voice commands are identified. For example, it determines whether the email app is already open or if it needs to be opened again.
[0818] Step 5:
[0819] The device performs actions determined by the analyzed voice commands and screen information. It simulates operations such as opening the email app and selecting the most recent received email from the list.
[0820] Step 6:
[0821] The device utilizes the results of the emotion engine to adjust feedback according to the user's emotional state. For example, if the user is expressing dissatisfaction, it will generate a response in a gentle tone such as, "I've seen your email. Is there anything I can help you with?"
[0822] Step 7:
[0823] The device uses speech synthesis technology to provide voice feedback on the results of the operations performed. The feedback is generated in real time, informing the user of the success of the operation and prompting them to take the next action.
[0824] This processing flow allows the system to quickly and accurately perform PC operations based on voice input and provide the user with appropriate responses that reflect their emotions.
[0825] (Example 2)
[0826] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0827] Traditional computer systems made it difficult for users to operate them using natural voice commands and failed to provide responses that took user emotions into consideration. This resulted in users experiencing stress during voice input and decreased operational efficiency. Furthermore, there was a lack of methods to flexibly apply operations based on information displayed on the screen.
[0828] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0829] In this invention, the server includes means for receiving user instructions using voice input, means for using a natural language processing device to convert the received voice into text, means for analyzing image information acquired from a display device to identify the target to be executed based on the user's instructions, and means for analyzing the user's emotions from the voice and adjusting the response based on those emotions. This enables the user to effectively operate the computer with a natural voice and receive a customized response that is tailored to their emotions.
[0830] "Voice input" is a method by which users give voice commands directly to a computer through a microphone.
[0831] A "natural language processing device" is a device equipped with technology to convert speech data into text, and has the function of making user voice instructions into a format that a computer can understand.
[0832] A "display device" is a device that displays information visually, such as a computer screen or monitor.
[0833] "Analyzing image information" is the process of analyzing screen captures and image data to understand what is displayed within them and extract specific data.
[0834] "Analyzing user emotions from voice" refers to a technology that analyzes the tone, speed, volume, and other elements contained in voice data to determine the user's emotional state.
[0835] A "generative AI technology model" is a model equipped with an algorithm that uses artificial intelligence to automatically generate content and feedback based on user instructions.
[0836] "Speech synthesis technology" is a technology that converts text data into speech, providing users with computer responses as voice.
[0837] This invention is a system that allows users to operate a computer using natural voice, enabling more human-like interaction. The system combines voice input, image information analysis, generative AI technology models, and speech synthesis technology.
[0838] First, the user gives voice instructions through the device's microphone, and the device receives this voice data. The voice is then converted into text using a natural language processing unit. In this process, common technologies can be used as natural language processing techniques, for example.
[0839] Next, the terminal acquires image information from the display device and analyzes its contents. For example, image processing software is used to identify the appropriate action to take based on the user's instructions. The terminal also incorporates an emotion analyzer that analyzes the user's emotions from voice data. For instance, it can determine whether the user is feeling anxious or calm.
[0840] By utilizing generative AI technology models, the device generates responses that correspond to the user's instructions and emotions. This generative AI model can be modified using common artificial intelligence algorithms. For example, if a user instructs, "Can you open my presentation materials and help me with something?", the device might generate a response such as, "I'll open the materials now, and then shall we practice the slides together?"
[0841] Finally, the device uses speech synthesis technology to provide the generated response to the user as audio. This process utilizes speech synthesis technology to provide the user with natural and easy-to-understand voice messages.
[0842] This system allows users to operate computers through voice commands and emotionally responsive interactions, resulting in improved operational efficiency and a richer user experience.
[0843] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0844] Step 1:
[0845] The user gives voice commands to the device. For example, they might say, "Open the presentation materials." The device receives this voice data as input through its built-in microphone. The voice data is then converted from a sound waveform into digital audio data.
[0846] Step 2:
[0847] The terminal transmits the received audio data to a natural language processing unit, which converts the audio into text. A speech recognition engine is used here to extract word sequences from the audio waveform as output. This converted text serves as the basic data for understanding the user's instructions.
[0848] Step 3:
[0849] The terminal acquires image information from the display device and analyzes it to associate it with user instructions. For example, image processing software performs a screen capture and extracts information about the active window and application. This analysis identifies actions that can be performed based on the user's instructions.
[0850] Step 4:
[0851] The device uses a voice analysis engine to evaluate the user's emotional state from voice data. Parameters such as voice tone, speed, and volume are used as input data, and the emotional state is output. For example, it can determine whether the user is speaking in an expectant voice or in an anxious voice.
[0852] Step 5:
[0853] The device utilizes a generative AI technology model to generate responses that correspond to the user's emotional state and instructions. Input includes text data and emotional information obtained in the previous step, and the output is an appropriate response. For example, it might generate a message like, "I've opened the document. What should I do next?"
[0854] Step 6:
[0855] The device converts text responses generated using speech synthesis technology into speech. The synthesized speech data is provided as output for communication with the user and played back from the device to the user as voice feedback. This process ensures that the user receives voice feedback appropriate to the instructions given.
[0856] (Application Example 2)
[0857] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0858] Current consumer robots have limitations in that voice control is restricted to simple commands, and they cannot provide flexible responses or adjust operations based on the user's emotional state. Furthermore, while users desire more natural communication, existing systems are unable to meet this need.
[0859] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0860] In this invention, the server includes means for receiving operation instructions using voice input, means for converting the received voice into text data, and means for analyzing video information displayed on a display device to identify executable actions corresponding to the instructions. This enables operation and feedback in accordance with the user's voice instructions and emotional state.
[0861] "Voice input" is a method for capturing voice information spoken by the user.
[0862] "Operation instructions" refer to the instructions a user gives to a robot to perform a specific action.
[0863] "Converting to text data" refers to the process of analyzing audio information after it has been input and converting it into a corresponding string of characters.
[0864] "Display devices" refer to devices that output information visually, such as screens and displays.
[0865] "Analyzing video information" is the process of analyzing the displayed visual information to obtain the necessary data.
[0866] "Identifying feasible actions" means identifying specific tasks that should be performed based on the analyzed information and instructions.
[0867] An "input device" is a device used to receive data or instructions, and includes keyboards and mice.
[0868] "To mimic an operation" means to reproduce an actual mechanical or electronic action based on an identified movement.
[0869] "Analyzing emotional state" means analyzing the tone and speed of voice from audio data to identify the user's emotions.
[0870] "Adjusting the speed of actions and response content" refers to the process of changing the speed of actions performed and the content of voice feedback according to the user's emotions.
[0871] This invention aims to achieve advanced interaction in consumer robots by combining voice input and emotion recognition technology. This system receives voice commands from the user, converts them into text data, and processes them in the robot's control unit. The details are described below.
[0872] The server first receives operation instructions from the user via a voice input device. The voice data is processed in real time and converted into text data using the Google Speech-to-Text API. This conversion process ensures that each instruction is interpreted as a clear digital command.
[0873] Furthermore, the server executes a Python script that implements an emotion recognition algorithm to analyze the user's emotional state along with the text data. Specifically, it analyzes acoustic features such as voice tone, speed, and volume to identify whether the user is in a particular emotional state, such as excitement, relaxation, or anxiety. This allows the server to adjust the operation speed and feedback content as needed.
[0874] Based on this information, the robot quickly identifies the specified action to analyze the displayed video information and generates an action scenario. The server sends commands to the robot's control unit, which then performs physical actions such as household chores according to the user's instructions. Simultaneously, it generates audio feedback using Amazon Polly to notify the user of the results and the next steps.
[0875] For example, if a user instructs the robot to "clean the kitchen and then prepare dinner," the robot will convert the voice into text, perform the kitchen cleaning task, check the ingredients in the refrigerator, and begin preparing the necessary ingredients. During the process, the robot will provide feedback such as, "I've started cleaning the kitchen. What can I help you with next?"
[0876] Examples of prompt statements that can be used as input to a generative AI model include the following:
[0877] "When a user gives a voice command with emotion, please convert that command into text and explain the action that corresponds to that emotion."
[0878] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0879] Step 1:
[0880] The terminal receives voice commands spoken by the user via the microphone. The input is an audio signal, and the output is captured by the terminal as a digital signal. This is the process of acquiring audio data.
[0881] Step 2:
[0882] The device converts the received audio signal into text data using the Google Speech-to-Text API. The input is an audio signal, and the output is the converted string. This conversion process involves phonological analysis and formatting of specific instructions.
[0883] Step 3:
[0884] The server analyzes the acquired text data to identify the specific tasks instructed by the user. The input is text data, and the output is a task list. This includes interpreting the meaning of the instructions using natural language processing techniques.
[0885] Step 4:
[0886] The device applies an emotion recognition algorithm to voice data to analyze the user's emotional state. The input is voice data, and the output is information about the emotional state. It performs specific actions to identify the user's emotions by analyzing the tone, speed, and intensity of the voice.
[0887] Step 5:
[0888] For each task, it generates instructions to adjust the execution speed and priority based on the user's emotional state. The input is a task list and emotional information, and the output is an adjusted task list. This includes actions to determine timing and order according to the user's emotions.
[0889] Step 6:
[0890] The terminal sends these instructions to the robot's control unit. The input is a coordinated task list, and the output is specific control signals to the robot. It is a process of instructing physical actions using a communication protocol.
[0891] Step 7:
[0892] The robot initiates a specified action based on a control signal, and its progress is monitored by a server during the operation. The input is the control signal, and the output is progress data. This includes performing actual physical tasks and tracking the results in real time.
[0893] Step 8:
[0894] The server receives progress data and provides voice feedback to the user regarding the progress and results via Amazon Polly. The input is progress data, and the output is a generated voice message. It is a voice generation process designed to communicate results in an accessible way.
[0895] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0896] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0897] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0898] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0899] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0900] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0901] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0902] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0903] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0904] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0905] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0906] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0907] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0908] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0909] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0910] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0911] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0912] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0913] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0914] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0915] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0916] The following is further disclosed regarding the embodiments described above.
[0917] (Claim 1)
[0918] A means of receiving user instructions using voice input,
[0919] A means of converting received audio into text,
[0920] A means for analyzing screen information output to a display device to identify executable actions in response to user instructions,
[0921] A means for simulating the operation of an input device based on a specified action,
[0922] A system that includes means for providing feedback to the user on the results of an operation.
[0923] (Claim 2)
[0924] The system according to claim 1, which uses speech synthesis technology to convey feedback in voice.
[0925] (Claim 3)
[0926] The system according to claim 1, which provides operations that can be adapted to new functions or applications based on the analyzed screen information.
[0927] "Example 1"
[0928] (Claim 1)
[0929] A means of obtaining user instructions using acoustic signals,
[0930] A means for converting acquired acoustic signals into text,
[0931] A means for analyzing video information output to a display means and determining executable actions in response to user instructions,
[0932] A means for simulating the operation of an input means based on the determined action,
[0933] A system that includes a means of providing the user with the results of an operation.
[0934] (Claim 2)
[0935] The system according to claim 1, which uses speech generation technology to transmit a response as an acoustic signal.
[0936] (Claim 3)
[0937] The system according to claim 1, which provides operations that can accommodate new functions or applications based on analyzed video information.
[0938] "Application Example 1"
[0939] (Claim 1)
[0940] A means of obtaining user instructions using voice input,
[0941] A means of converting acquired audio into text,
[0942] A means for analyzing visual information output to a display means and determining executable processing based on user instructions,
[0943] A means for reproducing the operation of the equipment based on the determined process,
[0944] A means of reporting the results of the operation to the user,
[0945] A means of transmitting reports by sound using voice conversion technology,
[0946] A system that includes means for providing new functions and operations applicable to programs based on analyzed visual information.
[0947] (Claim 2)
[0948] The system according to claim 1, which enables users to operate digital services through a natural voice interface.
[0949] (Claim 3)
[0950] The system according to claim 1, which enables communication processing in response to user instructions on a smart device.
[0951] "Example 2 of combining an emotion engine"
[0952] (Claim 1)
[0953] A means of receiving user instructions using voice input,
[0954] A means of using a natural language processing device that converts received audio into text,
[0955] A means for analyzing image information acquired from a display device and identifying the target to be executed based on user instructions,
[0956] A means for analyzing the user's emotions from their voice and adjusting the response based on those emotions,
[0957] A means for simulating the operation of an input device based on the identified target of execution,
[0958] A means of creating feedback tailored to user instructions using a generative AI technology model,
[0959] A system that includes a means of communicating the results of an operation to the user using speech synthesis technology.
[0960] (Claim 2)
[0961] The system according to claim 1, which uses an emotion analysis device to evaluate the user's emotional state and provides corresponding voice feedback.
[0962] (Claim 3)
[0963] The system according to claim 1, which provides a function that can be adapted to other functions or application fields based on analyzed image information and user emotions.
[0964] "Application example 2 when combining with an emotional engine"
[0965] (Claim 1)
[0966] A means of receiving operation instructions using voice input,
[0967] A means of converting received audio into text data,
[0968] A means for analyzing video information displayed on a display device to identify executable actions in response to instructions,
[0969] Means for mimicking the operation of an input device based on identified actions,
[0970] A means of returning the results after the operation to the user,
[0971] A system that analyzes emotional states from voice input and includes means to adjust the speed of actions and the content of responses according to those emotions.
[0972] (Claim 2)
[0973] The system according to claim 1, which uses speech synthesis technology to communicate results by voice.
[0974] (Claim 3)
[0975] The system according to claim 1, which performs operations that can flexibly respond to new functions or application software based on the analyzed video information. [Explanation of Symbols]
[0976] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining user instructions using voice input, A means of converting acquired audio into text, A means for analyzing visual information output to a display means and determining executable processing based on user instructions, A means for reproducing the operation of the equipment based on the determined process, A means of reporting the results of the operation to the user, A means of transmitting reports by sound using voice conversion technology, A system that includes means for providing new functions and operations applicable to programs based on analyzed visual information.
2. The system according to claim 1, which enables users to operate digital services through a natural voice interface.
3. The system according to claim 1, which enables communication processing in response to user instructions on a smart device.