System
The integration of gesture detection, eye tracking, and voice command analysis in smart home systems allows for intuitive control of home appliances through a virtual interface, addressing the limitations of conventional voice-only systems.
Patent Information
- Application Number
- JP2024120596
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional smart home systems rely solely on voice commands, lacking intuitive control methods such as gestures or eye tracking, resulting in a limited user experience and inefficient interaction with digital interfaces.
A system that integrates gesture detection, eye tracking, and voice command reception, with a server analyzing these inputs to generate a virtual interface displayed on the device, allowing users to control home appliances intuitively through this interface.
Enables users to efficiently and intuitively operate smart home appliances using gestures, gaze, and voice commands, enhancing the user experience by providing a more interactive and efficient control method.
Smart Images

Figure 2026019187000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional smart home systems have limitations in user interaction and rely solely on voice commands, resulting in a limited user experience. They also lack intuitive control via gestures or eye tracking, resulting in insufficient integration with the home's digital interface. This makes it difficult for users to enjoy an efficient and entertaining home experience. Furthermore, there is a lack of technology to complement the visual and interactive experience. [Means for solving the problem]
[0005] The present invention provides a system that includes means for detecting gestures, tracking gaze, and receiving voice commands, and also provides a means for analyzing data obtained from these input means. It also includes a means for generating a virtual interface based on the analysis results and displaying it in a home space. Controlling home appliances through the virtual interface enables intuitive and efficient operation. Specifically, gestures, gaze, and voice are acquired using the device's built-in sensors, camera, and microphone, and this data is sent to an analysis server. The server analyzes the data and generates a virtual interface based on the user's intentions, which is displayed on the device. Home appliances can be controlled through this interface, providing a new smart home experience that goes beyond conventional limitations.
[0006] A "gesture detection means" is a device or technology that detects a user's physical movements and recognizes them as digital signals.
[0007] "Eye tracking means" refers to devices or technologies that monitor a user's eye movements in real time and detect the direction and focus of their gaze.
[0008] A "means for receiving voice commands" is any device or technology that captures a user's voice and recognizes the voice instructions as digital signals.
[0009] "Means for analyzing data" refers to devices or technologies that process acquired gesture, gaze, and voice data and interpret the user's intentions.
[0010] "Means for generating a virtual interface" refers to devices or technologies that create interfaces such as virtual operation panels and buttons that can be operated by users based on the analysis results.
[0011] The "means for displaying a virtual interface" refers to a device or technology for visually displaying the generated virtual interface to the user.
[0012] The "means for controlling a home appliance" refers to a device or technology for executing a user's instruction obtained through a virtual interface on an actual home appliance.
[0013] A "terminal" is an electronic device that interacts with the user and is equipped with sensors, cameras, microphones, etc.
[0014] A "server" is a computer system that receives data sent from a terminal and analyzes and processes it. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention relates to a system that enables a user to intuitively operate various home appliances in a smart home, which includes a gesture detection unit, an eye tracking unit, a voice command receiving unit, a data analysis unit, a virtual interface generation unit, a virtual interface display unit, and a home appliance control unit.
[0037] Overall system picture
[0038] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user can then control the home appliances through the displayed virtual interface.
[0039] Device Role
[0040] The terminal has the following roles:
[0041] 1. Gesture and gaze capture:
[0042] Uses built-in sensors and cameras to capture user gestures and eye movements.
[0043] 2. Receiving voice commands:
[0044] A microphone is used to receive voice commands from the user.
[0045] 3. Data transmission:
[0046] The acquired gesture, gaze, and voice data is sent to the server.
[0047] 4. View virtual interfaces:
[0048] The virtual interface received from the server is displayed in the home space using AR.
[0049] Server Roles
[0050] The server has the following roles:
[0051] 1. Receiving data:
[0052] Receives gesture, gaze, and voice data sent from the device.
[0053] 2. Data Analysis:
[0054] The received data is analyzed to understand the user's intent. For example, it can recognize from gesture data that the user has performed the "turn on" action to turn on a light, or from voice data that the user has said "Turn on the light."
[0055] 3. Create a virtual interface:
[0056] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated to assist the user in their operations.
[0057] 4. Send virtual interface:
[0058] The generated virtual interface data is sent to the terminal.
[0059] User interaction
[0060] Users can issue voice commands to their devices using gestures or gaze. For example, if a user is in the living room and says, "Turn on the light," the device captures the voice and sends it to the server. The server analyzes the voice data and generates a control signal to turn on the light based on the user's intention. The device then sends the received control signal to the smart light, and the light turns on.
[0061] Specific examples
[0062] Lighting Control
[0063] 1. User says "Turn on the light":
[0064] The device's microphone captures this audio and sends it to the server.
[0065] 2. The server analyzes the audio data:
[0066] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[0067] 3. The terminal will display the virtual interface:
[0068] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[0069] 4. User's gaze lands on the "ON" button:
[0070] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[0071] 5. The light comes on:
[0072] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[0073] Playing music
[0074] 1. User controls the music player using gaze and gestures:
[0075] Select the music player icon with your gaze and press the play button with a gesture.
[0076] 2. The device sends the operation data to the server:
[0077] The terminal transmits gaze and gesture data to a server.
[0078] 3. The server analyzes the data and generates a playback signal:
[0079] The server understands the user's intention to play music and generates a play command.
[0080] 4. The device sends a signal to the music player:
[0081] The device sends a playback signal to the music player, and the music is played.
[0082] In this way, the system of the present invention can provide users with a highly intuitive and efficient smart home experience.
[0083] The processing flow will be explained below.
[0084] Step 1:
[0085] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[0086] Step 2:
[0087] The device monitors the user's movements and captures gesture, gaze, and voice data, for example, when the user moves their hand, looks in a particular direction, or issues a voice command.
[0088] Step 3:
[0089] The device first stores the captured data internally and then prepares it for transmission to the server, which includes packetizing the data and implementing security measures.
[0090] Step 4:
[0091] The device sends gesture data, gaze data, and voice data to the server, where the data is encrypted and sent in real time.
[0092] Step 5:
[0093] The server receives the data sent from the device and performs checks to ensure the data is consistent and complete.
[0094] Step 6:
[0095] The server analyzes the received data and uses gesture recognition, eye tracking, and voice analysis algorithms to understand the user's intent. For example, it analyzes the "raise your hand" gesture and the "turn on the light" voice command.
[0096] Step 7:
[0097] Based on the analysis results, the server generates a virtual interface according to the user's intentions, such as creating an "ON / OFF" button for a light or a control panel for playing music.
[0098] Step 8:
[0099] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[0100] Step 9:
[0101] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze and gestures.
[0102] Step 10:
[0103] The user operates the home appliance using the displayed virtual interface, for example, by selecting the "ON" button with their gaze and clicking the button with a gesture.
[0104] Step 11:
[0105] The device recaptures the user's gaze and gestures and recognizes that a part of the virtual interface has been manipulated.
[0106] Step 12:
[0107] The terminal sends this information to the server, which generates specific control signals for the home appliances.
[0108] Step 13:
[0109] The server sends the generated control signal back to the device, which then transmits the control signal to the actual home appliance, causing the light to turn on or the music to play.
[0110] This series of steps creates a smart home experience where users can intuitively control home appliances through gaze, gestures, and voice.
[0111] Example 1
[0112] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0113] In recent years, smart homes have become increasingly popular, and users have an increasing need to efficiently control a variety of home appliances. However, operation using conventional remote controls or smartphones is not intuitive, making it difficult for elderly people and those less familiar with technology to use. Furthermore, there are few systems that combine multiple input methods, such as gaze, gestures, and voice, to control home appliances, making it difficult to accurately understand the user's intentions and control them.
[0114] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0115] In this invention, the server includes means for analyzing user gesture, gaze, and voice data, means for generating an appropriate virtual interface based on the analysis results, and means for transmitting the generated virtual interface data to the terminal and displaying it in AR, thereby enabling users to intuitively and efficiently operate home appliances using multiple input means.
[0116] A "gesture detection means" is a device or technology that uses a sensor or camera to recognize the hand or body movements made by a user.
[0117] "Eye tracking means" refers to a device or technology that tracks the user's eye movements using a camera, infrared sensor, etc., to determine their gaze.
[0118] The "means for receiving voice commands" refers to a device or technology that receives the voice uttered by the user using a microphone or the like and analyzes the content of the voice.
[0119] "Means for analyzing data" refers to devices or technologies that process input data such as gestures, gaze, and voice to understand the user's intent.
[0120] "Means for generating a virtual interface" refers to a device or technology that creates a virtual interface that can be operated by the user based on the results of the analysis.
[0121] A "means for displaying a virtual interface" is a device or technique for visually presenting the generated virtual interface to a user.
[0122] A "means for controlling a home appliance" is a device or technology that sends appropriate control signals to a home appliance that is operated by a user through a virtual interface.
[0123] The "means for transmitting data to a server" refers to a device or technology that transmits captured data such as gestures, gazes, and voices to a server via a network.
[0124] "Means for the server to analyze data" refers to a device or technology that processes the gesture, gaze, and voice data received by the server to understand the user's intention.
[0125] "Means for generating an appropriate virtual interface" refers to a device or technology that allows the server to create a virtual interface for the user to operate based on the results of data analysis.
[0126] "Means for transmitting to a terminal and displaying in AR within the home" refers to a device or technology that transmits virtual interface data generated from a server to a terminal, which then displays it in AR (augmented reality) within the home.
[0127] The present invention relates to a system that allows users to intuitively operate smart home appliances. This system detects gestures, gaze, and voice inputs from the user and analyzes the data on a server to understand the user's intentions and generate and display an appropriate virtual interface. Furthermore, the virtual interface can be used to control the home appliances.
[0128] Hardware and Software Examples
[0129] Device: Built-in sensors, camera, microphone
[0130] The device captures gestures, gaze, and voice and sends this data to a server.
[0131] Server: Data analysis device
[0132] The server analyzes the received data and generates a virtual interface according to the user's intentions.
[0133] Virtual interface generation and display devices: AR devices, smartphones
[0134] The generated virtual interface is displayed in AR, providing the user with a visual means of operation.
[0135] Program processing
[0136] 1. Role of the device
[0137] The device uses its built-in sensors and camera to capture user gestures and eye movements, and a microphone to receive voice commands, which are then sent to a server.
[0138] 2. Server Roles
[0139] The server receives gesture, gaze, and voice data sent from the device and analyzes them to understand the user's intentions. Based on the analysis results, it generates an appropriate virtual interface and sends it to the device.
[0140] 3. Viewing Virtual Interfaces
[0141] The device receives the virtual interface from the server and displays it in the home using AR, allowing users to operate the virtual interface with their gaze and gestures to control home appliances.
[0142] Specific examples
[0143] Lighting Control
[0144] 1. The user says "Turn on the light"
[0145] The device's microphone captures this audio and sends it to the server.
[0146] 2. The server analyzes the audio data
[0147] The server understands the intent "turn on the light" and creates an "ON" interface for the light.
[0148] 3. The device displays the virtual interface
[0149] Based on the data received by the device, the light's "ON / OFF" button is displayed inside the home using AR.
[0150] 4. Users look at the "ON" button
[0151] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[0152] 5. The lights come on
[0153] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[0154] Playing music
[0155] 1. Users control the music player using gaze and gestures
[0156] Select the music player icon with your gaze and press the play button with a gesture.
[0157] 2. The device sends the operation data to the server.
[0158] The terminal transmits gaze and gesture data to a server.
[0159] 3. The server analyzes the data and generates a playback signal
[0160] The server understands the user's intention to play music and generates a play command.
[0161] 4. The device sends a signal to the music player
[0162] The device sends a playback signal to the music player, and the music is played.
[0163] Prompt Sentence Examples
[0164] A detailed description of the system can be created by inputting the following prompts into the generative AI model:
[0165] Example prompt:
[0166] "Please explain how this system works when a user wants to control a smart home light. For example, please provide a detailed description of the steps that occur when a user issues a voice command such as 'Turn on the light.'"
[0167] As described above, the system of the present invention enables users to intuitively and efficiently operate smart home appliances.
[0168] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0169] Step 1:
[0170] The terminal captures the user's input
[0171] The user issues the voice command "Turn on the lights."
[0172] Input: Voice commands, gestures, gaze
[0173] How it works: The device's microphone receives voice commands. The device's camera and eye-tracking sensors capture the user's gestures (pointing) and gaze.
[0174] Output: Voice data, gesture data, gaze data
[0175] Step 2:
[0176] The device sends the data to the server
[0177] The device transmits the captured voice data, gesture data, and gaze data to the server.
[0178] Input: Voice data, gesture data, gaze data
[0179] How it works: The device sends data to the server using Wi-Fi, Bluetooth, etc.
[0180] Output: Data sent
[0181] Step 3:
[0182] The server receives and analyzes the data
[0183] The server receives the voice data, gesture data, and gaze data transmitted from the terminal.
[0184] Input: Sent voice data, gesture data, and gaze data
[0185] How it works: The server uses a speech analysis algorithm to analyze the user's intent of "turn on the light." It understands from gesture data that the user is pointing at the light. It also checks from gaze data where the user is looking.
[0186] Output: User intent and target as analysis results
[0187] Step 4:
[0188] The server creates a virtual interface
[0189] Based on the user's intent analyzed by the server, an appropriate virtual interface (e.g., an "ON / OFF" button for a light) is generated.
[0190] Input: User intent, object
[0191] Operation: Based on the user's intention, the server designs a virtual interface including an "ON / OFF" button for the light and generates it as data.
[0192] Output: Generated virtual interface data
[0193] Step 5:
[0194] The server sends the virtual interface data to the terminal.
[0195] The server transmits the generated virtual interface data to the terminal.
[0196] Input: Virtual interface data
[0197] How it works: The server sends virtual interface data to the device via Wi-Fi or the internet.
[0198] Output: Virtual interface data sent
[0199] Step 6:
[0200] The device displays the virtual interface in AR
[0201] The device displays the virtual interface received from the server in the home space using AR.
[0202] Input: Virtual interface data
[0203] Operation: Uses the device's AR function to display the light's "ON / OFF" button at the specified location.
[0204] Output: AR displayed virtual interface
[0205] Step 7:
[0206] The user operates the virtual interface
[0207] The user selects the "ON" button with their gaze and then presses the "ON" button with a gesture.
[0208] Input: AR displayed virtual interface
[0209] Actions: The device captures the user's gaze and gestures.
[0210] Output: User operation data
[0211] Step 8:
[0212] The device sends signals to control home appliances.
[0213] Based on the user's operation, the terminal generates a control signal for the light (e.g., "turn on") and sends it to the smart light.
[0214] Input: User operation data
[0215] How it works: The device wirelessly sends the appropriate control signal to the smart light.
[0216] Output: Transmitted control signal
[0217] Step 9:
[0218] Home appliances work
[0219] The smart light receives a control signal from the device and performs the specified action (e.g., turning on the light).
[0220] Input: Control signal
[0221] Action: The smart light follows the control signal and turns on.
[0222] Output: The result is that the light turns on.
[0223] (Application example 1)
[0224] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0225] Conventional in-vehicle device operation involves numerous buttons and complex control panels, which can be inconvenient for drivers while driving and potentially compromise safety. In particular, operating entertainment and air conditioning systems requires intuitive methods that allow drivers to keep their eyes on the road. Furthermore, there is a lack of intuitive interfaces that integrate gesture, gaze, and voice commands, so these issues must be addressed simultaneously.
[0226] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0227] In this invention, the server includes a means for detecting gestures, a means for tracking gaze, and a means for receiving voice commands. This allows intuitive operation of in-car devices using gestures, gaze, and voice commands. Specifically, the driver can safely and efficiently control the entertainment system and air conditioning system in the car using gaze, gestures, and voice commands. This minimizes eye movement while driving, enabling intuitive in-car operation, improving convenience and safety for the driver.
[0228] A "gesture" is a way for a user to perform a specific action that is recognized by a camera or sensor and generates a corresponding signal.
[0229] "Eye tracking" is a method of tracking the movement of a user's eyes, recognizing the direction and focus of their gaze, and generating a signal.
[0230] "Voice command" is a means of receiving voice uttered by a user, analyzing the content of the voice, and generating a corresponding signal.
[0231] "Data analytics" is the means of understanding user intent by analyzing information obtained from gestures, gaze, and voice commands.
[0232] A "virtual interface" is a user interface that is generated based on analyzed information and is displayed through a digital device rather than a physical interface.
[0233] "Device control" is a means for users to operate and control specific devices through the generated virtual interface.
[0234] A "device" is a device that captures gestures and eye contact and receives voice commands, including a camera, built-in sensors, and microphone.
[0235] An "image capture device" is a device such as a camera that captures visual information and is used to detect the user's gaze and gestures.
[0236] "Built-in detection devices" refer to sensors built into the device, which are used to capture user behavior and environmental information.
[0237] An "entertainment system" is a system installed in a vehicle that provides entertainment functions such as music playback and video viewing.
[0238] The system that realizes this application example is designed for intuitive in-car operation, integrating gestures, gaze, and voice commands to allow users to control various in-car devices.
[0239] Overall system picture
[0240] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user controls the in-vehicle devices through the displayed virtual interface.
[0241] Device Role
[0242] The terminal has the following roles:
[0243] Gesture and gaze capture: Captures user gestures and gaze movements using built-in sensors and cameras.
[0244] Receiving voice commands: Uses a microphone to receive the user's voice commands.
[0245] Data transmission: The acquired gesture, gaze, and voice data are sent to the server.
[0246] Displaying the virtual interface: The virtual interface received from the server is displayed on the in-car display or head-mounted display.
[0247] Server Roles
[0248] The server has the following roles:
[0249] Data reception: Receives gesture, gaze, and voice data sent from the device.
[0250] Data analysis: Analyzes the received data and understands the user's intent. For example, it recognizes from gesture data that the user has turned up the volume, or from voice data that the user said "Turn up the volume."
[0251] Virtual interface generation: Based on the analysis results, a virtual interface (e.g., the "play / stop" button on a music player) is generated to assist the user in their operations.
[0252] Send virtual interface: Send the created virtual interface to the terminal.
[0253] Hardware and software used
[0254] Hardware: Camera, microphone, built-in sensors
[0255] Software: OpenCV, SpeechRecognition module, GazeTracker, GestureRecognizer
[0256] Processing Overview
[0257] The server analyzes the data sent from the device and generates a virtual interface based on the analysis results. This virtual interface is displayed on the device, and the user can operate the interface using gaze and gestures. For example, if a user says "Turn up the volume" in a car, the server captures and analyzes the voice command. Then, it generates a volume adjustment button on the virtual interface and displays it on the device. This allows the user to intuitively adjust the volume of the entertainment system.
[0258] Specific examples
[0259] To change the volume, the user can say "Turn up the volume." This voice command is captured by the device's microphone and sent to the server. The server analyzes the voice data and understands the intent of the volume adjustment. As a result, a virtual interface for adjusting the volume is displayed on the device. The user can use their gaze or gestures to operate the interface and adjust the volume.
[0260] An example of a prompt for a generative AI model is as follows:
[0261] Write the code to do the camera capture. Use OpenCV.
[0262]
[0263] Write code to recognize speech. Use the SpeechRecognition module.
[0264]
[0265] Write code to track your gaze. Use the GazeTracker module.
[0266]
[0267] Write code to recognize gestures. Use the GestureRecognizer module.
[0268] In this way, an interactive in-car operation system that combines gestures, gaze, and voice can be realized.
[0269] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0270] Step 1:
[0271] Users input gestures, gaze, and voice commands. The device captures these inputs using its built-in camera, sensors, and microphone. Specifically, the camera captures gestures and gaze movements as video, and the microphone captures voice commands as audio data.
[0272] Input: User gestures, eye movements, and voice commands
[0273] Output: Captured video and audio data
[0274] Step 2:
[0275] The device sends the captured data to the server, which includes gesture and gaze information as video data and voice commands as audio data.
[0276] Input: Captured video and audio data
[0277] Output: Data sent to the server
[0278] Step 3:
[0279] The server analyzes gesture and gaze data from the received video data and voice commands from the audio data. Specifically, the gesture recognition module recognizes gestures from the video data, and the eye-tracking module identifies the direction and focus of gaze. The voice recognition module analyzes the audio data and converts the user's intentions into text.
[0280] Input: Data sent to the server
[0281] Output: Analyzed gesture data, gaze data, voice commands
[0282] Step 4:
[0283] The server understands the user's intention based on the analyzed gesture data, gaze data, and voice command. For example, it recognizes that the user turned up the volume from the gesture data, that the user looked at a specific button from the gaze data, and that the user said "Turn up the volume" from the voice command.
[0284] Input: Analyzed gesture data, gaze data, voice commands
[0285] Output: User intent
[0286] Step 5:
[0287] The server generates a virtual interface based on the user's intentions, such as buttons for adjusting the volume or an operation panel for an entertainment system.
[0288] Input: User intent
[0289] Output: Generated virtual interface data
[0290] Step 6:
[0291] The server transmits the generated virtual interface data to the terminal, which receives the data and displays the virtual interface on an in-car display or a head-mounted display.
[0292] Input: Generated virtual interface data
[0293] Output: Virtual interface data sent to the terminal
[0294] Step 7:
[0295] The user operates the displayed virtual interface using gaze, gestures, and voice commands. For example, if the user gazes at the volume control button and gestures to increase the volume, the corresponding operation is executed. The device then transmits the newly acquired user input data back to the server.
[0296] Input: New user input to the virtual interface
[0297] Output: Control signal to in-vehicle equipment
[0298] Step 8:
[0299] The device then sends control signals to the vehicle's equipment based on the analysis results, allowing the equipment to be operated as intended by the user, for example by adjusting the volume.
[0300] Input: Control signal to in-vehicle equipment
[0301] Output: Device operation results
[0302] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0303] This invention relates to a system that controls various smart home appliances based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[0304] Overall system picture
[0305] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user can then operate the home appliance through the displayed virtual interface.
[0306] Device Role
[0307] The terminal has the following roles:
[0308] 1. Gesture and gaze capture:
[0309] Uses built-in sensors and cameras to capture user gestures and eye movements.
[0310] 2. Receiving voice commands:
[0311] A microphone is used to receive voice commands from the user.
[0312] 3. Detecting the user's emotional state:
[0313] It uses a camera and microphone to capture the user's facial expressions and voice tone, and an emotion engine to detect emotions.
[0314] 4. Data transmission:
[0315] The acquired gesture, gaze, voice, and emotion data is sent to the server.
[0316] 5. View virtual interfaces:
[0317] The virtual interface received from the server is displayed in the home space using AR.
[0318] Server Roles
[0319] The server has the following roles:
[0320] 1. Receiving data:
[0321] Receive gesture, gaze, voice, and emotion data sent from the terminal.
[0322] 2. Data Analysis:
[0323] The received data is analyzed to understand the user's intentions and emotional state. For example, hand movements are analyzed from gesture data, gaze data from points of gaze, command content from voice data, and user emotions from emotional data.
[0324] 3. Create a virtual interface:
[0325] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated according to the user's intentions and emotional state.
[0326] 4. Send virtual interface:
[0327] The generated virtual interface data is sent to the terminal.
[0328] User interaction
[0329] Users can use gestures, gaze, and voice commands to communicate their emotional state through facial expressions and tone of voice. For example, if the emotion engine detects that the user is tired, the system can dim the lights or play relaxing music.
[0330] Specific examples
[0331] Lighting Control
[0332] 1. User says "Turn on the light":
[0333] The device's microphone captures this audio and sends it to the server.
[0334] 2. The server analyzes the audio data:
[0335] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[0336] 3. The terminal will display the virtual interface:
[0337] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[0338] 4. User's gaze lands on the "ON" button:
[0339] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[0340] 5. The light comes on:
[0341] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[0342] Playing music
[0343] 1. User controls the music player using gaze and gestures:
[0344] Select the music player icon with your gaze and press the play button with a gesture.
[0345] 2. The device sends the operation data to the server:
[0346] The terminal transmits gaze and gesture data to a server.
[0347] 3. The server analyzes the data and generates a playback signal:
[0348] The server understands the user's intention to play music and generates a play command.
[0349] 4. The device sends a signal to the music player:
[0350] The device sends a playback signal to the music player, and the music is played.
[0351] Emotion-based preferences
[0352] 1. If the user is tired:
[0353] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[0354] 2. The device sends the emotion data to the server:
[0355] The emotion engine transmits the detected emotion data to the server.
[0356] 3. The server analyzes the emotion data:
[0357] The server analyzes that the user is tired and generates a virtual interface for relaxation mode.
[0358] 4. The device will display the virtual interface in relaxed mode:
[0359] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[0360] 5. The environment changes:
[0361] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[0362] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[0363] The processing flow will be explained below.
[0364] Step 1:
[0365] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[0366] Step 2:
[0367] The device monitors the user's movements and captures gesture, gaze, and voice data, such as when the user moves their hand, looks in a particular direction, or makes a sound.
[0368] Step 3:
[0369] The device captures the user's facial expressions and voice tone, which are then analyzed by the emotion engine to generate the user's emotional data, including the user's smiling or serious face and voice tone.
[0370] Step 4:
[0371] The device captures gesture, gaze, voice, and emotion data, stores it internally, and then prepares it for transmission to the server, including packetizing the data and encrypting it for privacy purposes.
[0372] Step 5:
[0373] The device transmits gesture data, gaze data, voice data, and emotion data to the server, where the data is encrypted and transmitted in real time.
[0374] Step 6:
[0375] The server receives the data sent from the terminal, checks the consistency and integrity of the received data, and performs any pre-processing necessary for analysis.
[0376] Step 7:
[0377] The server analyzes the received data, using gesture recognition algorithms, eye tracking algorithms, and voice analysis algorithms to understand the user's intention, and then uses an emotion engine to analyze the user's emotional state. For example, a "raise your hand" gesture, a "turn on the light" voice command, or a smile can be analyzed.
[0378] Step 8:
[0379] Based on the analysis results, the server generates a virtual interface that corresponds to the user's intentions and emotions. For example, if the user is tired, it generates a relaxation mode interface.
[0380] Step 9:
[0381] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[0382] Step 10:
[0383] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze, gestures, and emotional state.
[0384] Step 11:
[0385] The user operates the home appliance using the displayed virtual interface. For example, they can select the "ON" button with their gaze and click the button with their gesture. The interface and behavior are adjusted based on the user's emotions.
[0386] Step 12:
[0387] The device recaptures the user's gaze and gestures to recognize when parts of the virtual interface are manipulated, and emotional state is also continuously monitored.
[0388] Step 13:
[0389] The terminal transmits the operation data to the server, which generates specific control signals for the home appliances. If the emotion data changes, the control signals are adjusted accordingly.
[0390] Step 14:
[0391] The server sends the generated control signal back to the terminal, which then transmits the control signal to the actual home appliance, for example, turning on a light, playing music, or changing the environment setting to relaxation mode.
[0392] This series of steps allows users to enjoy an intuitive and personalized smart home experience through gaze, gesture, voice, and emotion.
[0393] Example 2
[0394] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0395] Current smart home systems provide operations based on user gestures, gaze, and voice commands, but they do not adequately consider the user's emotional state when controlling the interface or operation of home appliances. As a result, they are unable to flexibly respond to the user's diverse needs and momentary emotional state, making it difficult to provide an intuitive and personalized experience.
[0396] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for detecting gestures, means for tracking gaze, means for receiving voice commands, means for analyzing data obtained from the gesture means, the gaze tracking means, and the voice command means, means for generating a virtual interface based on the analysis result, means for displaying the virtual interface, means for controlling home appliances through the virtual interface, means for analyzing the emotional state of the user, and means for adjusting the operation of the virtual interface and the home appliances based on the emotional state. This enables an intuitive and personalized smart home experience based on the user's intentions and emotions.
[0397] "Means for detecting gestures" refers to technology that uses sensors to recognize the movements of a user's hands or body and collects them as data.
[0398] "Means for tracking gaze" refers to technology that uses a camera or sensor to capture the movement of a user's gaze and analyze the position and direction of that gaze.
[0399] "Means for receiving voice commands" refers to technology that uses a microphone to pick up the voices emitted by the user and collects them as data.
[0400] "Means for analyzing data" refers to technology for analyzing collected data such as gestures, gaze, and voice commands to understand the user's intentions and state.
[0401] "Means for generating a virtual interface" refers to a technology for virtually creating an interface that can be operated by a user based on the analysis results.
[0402] "Means for displaying a virtual interface" refers to a technique for displaying the generated virtual interface in a location where the user can see or operate it.
[0403] "Means for controlling home appliances" refers to technology that operates electronic devices and equipment in the home based on user instructions through a virtual interface.
[0404] "Means for analyzing the user's emotional state" refers to technology for determining the user's current emotions from their facial expressions, tone of voice, etc.
[0405] "Means for adjusting the operation of a virtual interface or a home appliance based on the emotional state" refers to a technology for setting the optimal operating mode of a virtual interface or a home appliance based on the analyzed emotional state.
[0406] This system controls various electronic devices in a smart home based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[0407] Overall system picture
[0408] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user then operates the home appliance through the displayed virtual interface.
[0409] Hardware and software used
[0410] Device: A typical smartphone or tablet device, using built-in sensors, cameras, and microphones to capture gestures, gaze, and voice.
[0411] Server: Uses a high-performance cloud-based computer to analyze data, generate virtual interfaces, and manipulate emotional states.
[0412] software:
[0413] Speech recognition engine: Converts voice commands into text.
[0414] AI model: Analyzes gesture and gaze data to understand user intent.
[0415] Emotion engine: Analyzes user emotions from facial expressions and voice tone.
[0416] AR module: displays a virtual interface.
[0417] Specific operation example
[0418] Lighting Control
[0419] 1. The user says "Turn on the light":
[0420] The device's microphone captures this audio and sends it to the server.
[0421] 2. The server analyzes the audio data:
[0422] The server uses a speech recognition engine to convert the command "turn on the lights" into text and understand its intent.
[0423] 3. The terminal will display the virtual interface:
[0424] The server generates an "ON / OFF" interface for the light and sends it to the device, which displays the "ON / OFF" button in AR.
[0425] 4. User's gaze focuses on the "ON" button:
[0426] The device's eye-tracking sensor determines where the user is looking and confirms that they are looking at the "ON" button.
[0427] 5. The light comes on:
[0428] The device sends an "ON" signal to the light, and the light turns on.
[0429] Emotion-based preferences
[0430] 1. If the user is tired:
[0431] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[0432] 2. The device sends the emotion data to the server:
[0433] The emotion engine transmits the detected emotion data to the server.
[0434] 3. The server analyzes the emotion data:
[0435] The server analyzes that the user is tired and generates a virtual interface in relaxation mode.
[0436] 4. The device will display the virtual interface in relaxed mode:
[0437] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[0438] 5. The environment changes:
[0439] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[0440] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[0441] Prompt Sentence Examples
[0442] An example of a prompt is as follows:
[0443] "Turn on the lights"
[0444] "Play music."
[0445] "Put the room in relaxation mode"
[0446] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0447] Step 1:
[0448] The user provides input. Specifically, the user makes gestures toward the device, directs their gaze toward a specific object, or issues a voice command. These inputs are captured by the device. Input data includes gestures, gaze, and voice commands.
[0449] Step 2:
[0450] The device captures data: its built-in sensors capture user gestures, its camera tracks gaze, and its microphone records voice commands. The input data is gestures, gaze, and voice, and the output data is the raw captured data.
[0451] Step 3:
[0452] The device sends the captured data to the server. The device's communication module assembles gesture, gaze, and voice data into packets and transmits them to the server. The input data is the captured raw data, and the output data is the transmitted packets.
[0453] Step 4:
[0454] The server receives the data. The server's network interface receives the data packets sent from the terminal and stores them in memory for analysis. The input data is the data packets from the terminal, and the output data is the data for analysis stored in memory.
[0455] Step 5:
[0456] The server analyzes the data. The server's AI model analyzes the gesture data to identify hand movements and gaze data to identify the user's point of gaze. The voice recognition engine converts voice commands into text and understands the content. Furthermore, the emotion engine analyzes the user's emotional state from their facial expressions and tone of voice. The input data is stored data for analysis, and the output data is the analysis results that indicate the user's intentions and emotional state.
[0457] Examples:
[0458] Recognizes the "waving" gesture from gesture data.
[0459] Eye gaze data determined that the person was "watching TV."
[0460] The command "Turn on the lights" is converted from voice data into text.
[0461] Emotional data was analyzed to determine that the user was relaxed.
[0462] Step 6:
[0463] The server generates a virtual interface based on the analysis results. The server designs a virtual interface according to the user's intentions, such as an "ON / OFF" button on a light or a "PLAY / STOP" button on a music player. The input data is the analysis results, and the output data is the design of the virtual interface.
[0464] Step 7:
[0465] The server sends the virtual interface to the terminal. The generated virtual interface data is packaged into packets and sent to the terminal. The input data is the virtual interface design, and the output data is the packets sent to the terminal.
[0466] Step 8:
[0467] The device displays a virtual interface. The device's AR module renders the virtual interface in space and displays it in the user's field of view. The input data are packets sent from the server, and the output data is the displayed virtual interface.
[0468] Step 9:
[0469] The user operates the virtual interface by looking at or tapping buttons on the displayed virtual interface. The input data is the displayed virtual interface, and the output is the user's operation.
[0470] Step 10:
[0471] The terminal transmits a signal to the home appliance. The communication module of the terminal transmits a control signal to the home appliance in response to a user operation. The input data is the user's operation, and the output data is the control signal transmitted to the home appliance.
[0472] Examples:
[0473] When you look at the "ON" button, the device sends a signal to the light to "turn it on."
[0474] Tapping the "Play" button sends a "Play" signal to the music player.
[0475] This allows the system to deliver an intuitive and personalized smart home experience based on the user's intent and emotional state.
[0476] (Application example 2)
[0477] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0478] While conventional home appliance control systems simplify the operation of individual devices, they are not sufficient to improve the user's in-store experience. Furthermore, the information provided and service suggestions provided in-store are standardized, making it difficult to personalize them to meet the needs and emotional state of each individual user. Furthermore, conventional systems have the problem of being unable to suggest services that take the user's emotions into account.
[0479] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing gesture data, means for analyzing gaze data, means for analyzing voice data, means for analyzing emotion data, means for generating a virtual interface based on the analysis results, means for providing information and services through the virtual interface, and means for analyzing the user's emotional state and proposing optimal services. This enables users to intuitively interact in a physical store using gestures, gaze, and voice commands, and to be provided with personalized services according to their emotional state.
[0480] "Gesture sensing means" refers to a device or sensor used to sense the movement of a user's hands or body.
[0481] "Eye tracking means" means a camera or other device used to track a user's eye movements and determine the direction of gaze.
[0482] "Means for receiving voice commands" refers to a device that captures the voice uttered by the user using a microphone or the like and analyzes it.
[0483] "Means for analyzing" refers to the software or hardware used to process data obtained from gestures, gaze, voice commands, etc., to understand the user's intent and emotional state.
[0484] "Means for generating a virtual interface" refers to a device or software for designing and displaying a virtual interface that can be operated by a user based on the analysis results.
[0485] The term "means for displaying a virtual interface" refers to a device, such as a display or projector, for visually presenting the generated virtual interface to a user.
[0486] "Means for controlling a home appliance" refers to a device or system that controls a home appliance through a virtual interface.
[0487] "Information and service provision means for physical stores" refers to technologies and devices used to provide product information and services to users within a store.
[0488] "Means for analyzing the user's emotional state and suggesting optimal services" refers to software or devices that use cameras or microphones to sense the user's emotional state and suggest appropriate services or information based on that state.
[0489] That's all.
[0490] This invention is a system that receives user gestures, gaze, and voice commands as inputs, analyzes them, and controls home appliances. Furthermore, this system can recognize the user's emotional state and provide services according to that state. Specific embodiments for implementing this invention are described below.
[0491] 1. System Configuration
[0492] Hardware Configuration
[0493] 1. Device:
[0494] Gesture detection: Built-in sensors (e.g. accelerometer, infrared sensor)
[0495] Eye tracking: Camera (e.g. webcam, infrared camera)
[0496] Voice command reception: Microphone
[0497] Virtual interface display: displays, AR headsets
[0498] 2. Server:
[0499] Data analysis: high-performance processors, databases
[0500] Emotion Engine: Emotion Recognition Algorithm
[0501] Software Configuration
[0502] 1. On the device:
[0503] Gesture recognition program: OpenCV
[0504] Eye-tracking software: Dedicated eye-tracking library
[0505] Speech recognition software: Libraries like SpeechRecognition
[0506] Emotion Recognizer: EmotionRecognizer
[0507] 2. Server side:
[0508] Data Analysis Program
[0509] Virtual Interface Generator
[0510] Emotion Engine
[0511] 2. Program processing content
[0512] Gesture Analysis
[0513] The device uses built-in sensors and a camera to capture the user's hand and body movements and detect gestures, and the data is sent to a server where an analysis program interprets the user's intentions.
[0514] Eye tracking
[0515] The device's camera is used to track the user's eye movements and determine the direction of their gaze. The gaze data is sent to a server, where an analysis program identifies the point of gaze.
[0516] Voice command analysis
[0517] The device's microphone receives voice commands, which are interpreted by speech recognition software, and the results are sent to a server, which analyzes the intent of the command.
[0518] Emotion analysis
[0519] The device's camera and microphone are used to capture the user's facial expressions and tone of voice, and this data is analyzed by an emotion recognition program to determine the user's emotional state.
[0520] Virtual Interface Creation
[0521] The server generates a virtual interface based on the analysis of gesture, gaze, voice, and emotion data. The generated interface is sent to the terminal and presented to the user through a display device (such as a monitor or AR headset).
[0522] Information provision and service proposals
[0523] The server then provides optimal information and services based on the results of the user's data analysis. For example, if the user's gaze is directed at a particular product, it will present information about that product. If it detects a tired emotional state, it will suggest a relaxing environment.
[0524] 3. Examples and prompts
[0525] As a concrete example, imagine a user visits a store, looks at a product on the shelf, and detailed information about that product is displayed on smart glasses. Here is an example prompt:
[0526] When a user looks at a product on the shelf, the smart glasses will display detailed information about that product. If the user raises their hand to make a gesture, more information and reviews will be displayed.
[0527] In this way, the present invention can provide a more intuitive and personalized experience based on the user's intent and emotional state.
[0528] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0529] Step 1:
[0530] When a user enters the device's visual range, the device's camera captures the user's gaze and hand gestures. The input is image data from the camera, and the output is gesture and gaze data. The device does this in real time.
[0531] Step 2:
[0532] The device's microphone receives the user's voice commands. The input is the voice data obtained from the microphone, and the output is the interpreted voice command. The voice data is analyzed using the SpeechRecognition library.
[0533] Step 3:
[0534] The captured gesture, gaze, and voice data is sent from the device to the server. The input is the raw data sent from the device, and the output is the raw data sent to the server, which the device transmits via wireless communication.
[0535] Step 4:
[0536] The server analyzes the received gesture, gaze, and voice data. The input is the raw data sent from the device, and the output is the analysis result. The server analyzes the data using OpenCV, a dedicated gaze tracking library, and emotion recognition algorithms.
[0537] Step 5:
[0538] The server uses emotion recognition algorithms to analyze the user's emotional state from their facial expressions and tone of voice. The input is video and audio data from the camera and microphone, and the output is the analyzed emotional state. This is done by the emotion engine.
[0539] Step 6:
[0540] The server generates a virtual interface based on the analysis results. The input is the analyzed gesture, gaze, voice, and emotion data, and the output is the virtual interface. This is done by a virtual interface generation program.
[0541] Step 7:
[0542] The generated virtual interface is sent from the server to the device and displayed on the device's display or AR headset. The input is the virtual interface data sent from the server, and the output is the interface displayed on the device.
[0543] Step 8:
[0544] Users control home appliances using the displayed virtual interface. For example, users select products with their gaze and request more information with hand gestures. The input is the user's gaze and gesture data, and the output is the control command.
[0545] Step 9:
[0546] The server provides information and services based on the user's emotional state. For example, if it detects that the user is tired, it suggests relaxing environments and products. The input is the analyzed emotional data, and the output is the suggested information and services.
[0547] That's all.
[0548] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0549] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0550] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0551] [Second embodiment]
[0552] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0553] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0554] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0555] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0556] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0557] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0558] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0559] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0560] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0561] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0562] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0563] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0564] The present invention relates to a system that enables a user to intuitively operate various home appliances in a smart home, which includes a gesture detection unit, an eye tracking unit, a voice command receiving unit, a data analysis unit, a virtual interface generation unit, a virtual interface display unit, and a home appliance control unit.
[0565] Overall system picture
[0566] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user can then control the home appliances through the displayed virtual interface.
[0567] Device Role
[0568] The terminal has the following roles:
[0569] 1. Gesture and gaze capture:
[0570] Uses built-in sensors and cameras to capture user gestures and eye movements.
[0571] 2. Receiving voice commands:
[0572] A microphone is used to receive voice commands from the user.
[0573] 3. Data transmission:
[0574] The acquired gesture, gaze, and voice data is sent to the server.
[0575] 4. View virtual interfaces:
[0576] The virtual interface received from the server is displayed in the home space using AR.
[0577] Server Roles
[0578] The server has the following roles:
[0579] 1. Receiving data:
[0580] Receives gesture, gaze, and voice data sent from the device.
[0581] 2. Data Analysis:
[0582] The received data is analyzed to understand the user's intent. For example, it can recognize from gesture data that the user has performed the "turn on" action to turn on a light, or from voice data that the user has said "Turn on the light."
[0583] 3. Create a virtual interface:
[0584] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated to assist the user in their operations.
[0585] 4. Send virtual interface:
[0586] The generated virtual interface data is sent to the terminal.
[0587] User interaction
[0588] Users can issue voice commands to their devices using gestures or gaze. For example, if a user is in the living room and says, "Turn on the light," the device captures the voice and sends it to the server. The server analyzes the voice data and generates a control signal to turn on the light based on the user's intention. The device then sends the received control signal to the smart light, and the light turns on.
[0589] Specific examples
[0590] Lighting Control
[0591] 1. User says "Turn on the light":
[0592] The device's microphone captures this audio and sends it to the server.
[0593] 2. The server analyzes the audio data:
[0594] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[0595] 3. The terminal will display the virtual interface:
[0596] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[0597] 4. User's gaze lands on the "ON" button:
[0598] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[0599] 5. The light comes on:
[0600] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[0601] Playing music
[0602] 1. User controls the music player using gaze and gestures:
[0603] Select the music player icon with your gaze and press the play button with a gesture.
[0604] 2. The device sends the operation data to the server:
[0605] The terminal transmits gaze and gesture data to a server.
[0606] 3. The server analyzes the data and generates a playback signal:
[0607] The server understands the user's intention to play music and generates a play command.
[0608] 4. The device sends a signal to the music player:
[0609] The device sends a playback signal to the music player, and the music is played.
[0610] In this way, the system of the present invention can provide users with a highly intuitive and efficient smart home experience.
[0611] The processing flow will be explained below.
[0612] Step 1:
[0613] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[0614] Step 2:
[0615] The device monitors the user's movements and captures gesture, gaze, and voice data, for example, when the user moves their hand, looks in a particular direction, or issues a voice command.
[0616] Step 3:
[0617] The device first stores the captured data internally and then prepares it for transmission to the server, which includes packetizing the data and implementing security measures.
[0618] Step 4:
[0619] The device sends gesture data, gaze data, and voice data to the server, where the data is encrypted and sent in real time.
[0620] Step 5:
[0621] The server receives the data sent from the device and performs checks to ensure the data is consistent and complete.
[0622] Step 6:
[0623] The server analyzes the received data and uses gesture recognition, eye tracking, and voice analysis algorithms to understand the user's intent. For example, it analyzes the "raise your hand" gesture and the "turn on the light" voice command.
[0624] Step 7:
[0625] Based on the analysis results, the server generates a virtual interface according to the user's intentions, such as creating an "ON / OFF" button for a light or a control panel for playing music.
[0626] Step 8:
[0627] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[0628] Step 9:
[0629] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze and gestures.
[0630] Step 10:
[0631] The user operates the home appliance using the displayed virtual interface, for example, by selecting the "ON" button with their gaze and clicking the button with a gesture.
[0632] Step 11:
[0633] The device recaptures the user's gaze and gestures and recognizes that a part of the virtual interface has been manipulated.
[0634] Step 12:
[0635] The terminal sends this information to the server, which generates specific control signals for the home appliances.
[0636] Step 13:
[0637] The server sends the generated control signal back to the device, which then transmits the control signal to the actual home appliance, causing the light to turn on or the music to play.
[0638] This series of steps creates a smart home experience where users can intuitively control home appliances through gaze, gestures, and voice.
[0639] Example 1
[0640] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0641] In recent years, smart homes have become increasingly popular, and users have an increasing need to efficiently control a variety of home appliances. However, operation using conventional remote controls or smartphones is not intuitive, making it difficult for elderly people and those less familiar with technology to use. Furthermore, there are few systems that combine multiple input methods, such as gaze, gestures, and voice, to control home appliances, making it difficult to accurately understand the user's intentions and control them.
[0642] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0643] In this invention, the server includes means for analyzing user gesture, gaze, and voice data, means for generating an appropriate virtual interface based on the analysis results, and means for transmitting the generated virtual interface data to the terminal and displaying it in AR, thereby enabling users to intuitively and efficiently operate home appliances using multiple input means.
[0644] A "gesture detection means" is a device or technology that uses a sensor or camera to recognize the hand or body movements made by a user.
[0645] "Eye tracking means" refers to a device or technology that tracks the user's eye movements using a camera, infrared sensor, etc., to determine their gaze.
[0646] The "means for receiving voice commands" refers to a device or technology that receives the voice uttered by the user using a microphone or the like and analyzes the content of the voice.
[0647] "Means for analyzing data" refers to devices or technologies that process input data such as gestures, gaze, and voice to understand the user's intent.
[0648] "Means for generating a virtual interface" refers to a device or technology that creates a virtual interface that can be operated by the user based on the results of the analysis.
[0649] A "means for displaying a virtual interface" is a device or technique for visually presenting the generated virtual interface to a user.
[0650] A "means for controlling a home appliance" is a device or technology that sends appropriate control signals to a home appliance that is operated by a user through a virtual interface.
[0651] The "means for transmitting data to a server" refers to a device or technology that transmits captured data such as gestures, gazes, and voices to a server via a network.
[0652] "Means for the server to analyze data" refers to a device or technology that processes the gesture, gaze, and voice data received by the server to understand the user's intention.
[0653] "Means for generating an appropriate virtual interface" refers to a device or technology that allows the server to create a virtual interface for the user to operate based on the results of data analysis.
[0654] "Means for transmitting to a terminal and displaying in AR within the home" refers to a device or technology that transmits virtual interface data generated from a server to a terminal, which then displays it in AR (augmented reality) within the home.
[0655] The present invention relates to a system that allows users to intuitively operate smart home appliances. This system detects gestures, gaze, and voice inputs from the user and analyzes the data on a server to understand the user's intentions and generate and display an appropriate virtual interface. Furthermore, the virtual interface can be used to control the home appliances.
[0656] Hardware and Software Examples
[0657] Device: Built-in sensors, camera, microphone
[0658] The device captures gestures, gaze, and voice and sends this data to a server.
[0659] Server: Data analysis device
[0660] The server analyzes the received data and generates a virtual interface according to the user's intentions.
[0661] Virtual interface generation and display devices: AR devices, smartphones
[0662] The generated virtual interface is displayed in AR, providing the user with a visual means of operation.
[0663] Program processing
[0664] 1. Role of the device
[0665] The device uses its built-in sensors and camera to capture user gestures and eye movements, and a microphone to receive voice commands, which are then sent to a server.
[0666] 2. Server Roles
[0667] The server receives gesture, gaze, and voice data sent from the device and analyzes them to understand the user's intentions. Based on the analysis results, it generates an appropriate virtual interface and sends it to the device.
[0668] 3. Viewing Virtual Interfaces
[0669] The device receives the virtual interface from the server and displays it in the home using AR, allowing users to operate the virtual interface with their gaze and gestures to control home appliances.
[0670] Specific examples
[0671] Lighting Control
[0672] 1. The user says "Turn on the light"
[0673] The device's microphone captures this audio and sends it to the server.
[0674] 2. The server analyzes the audio data
[0675] The server understands the intent "turn on the light" and creates an "ON" interface for the light.
[0676] 3. The device displays the virtual interface
[0677] Based on the data received by the device, the light's "ON / OFF" button is displayed inside the home using AR.
[0678] 4. Users look at the "ON" button
[0679] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[0680] 5. The lights come on
[0681] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[0682] Playing music
[0683] 1. Users control the music player using gaze and gestures
[0684] Select the music player icon with your gaze and press the play button with a gesture.
[0685] 2. The device sends the operation data to the server.
[0686] The terminal transmits gaze and gesture data to a server.
[0687] 3. The server analyzes the data and generates a playback signal
[0688] The server understands the user's intention to play music and generates a play command.
[0689] 4. The device sends a signal to the music player
[0690] The device sends a playback signal to the music player, and the music is played.
[0691] Prompt Sentence Examples
[0692] A detailed description of the system can be created by inputting the following prompts into the generative AI model:
[0693] Example prompt:
[0694] "Please explain how this system works when a user wants to control a smart home light. For example, please provide a detailed description of the steps that occur when a user issues a voice command such as 'Turn on the light.'"
[0695] As described above, the system of the present invention enables users to intuitively and efficiently operate smart home appliances.
[0696] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0697] Step 1:
[0698] The terminal captures the user's input
[0699] The user issues the voice command "Turn on the lights."
[0700] Input: Voice commands, gestures, gaze
[0701] How it works: The device's microphone receives voice commands. The device's camera and eye-tracking sensors capture the user's gestures (pointing) and gaze.
[0702] Output: Voice data, gesture data, gaze data
[0703] Step 2:
[0704] The device sends the data to the server
[0705] The device transmits the captured voice data, gesture data, and gaze data to the server.
[0706] Input: Voice data, gesture data, gaze data
[0707] How it works: The device sends data to the server using Wi-Fi, Bluetooth, etc.
[0708] Output: Data sent
[0709] Step 3:
[0710] The server receives and analyzes the data
[0711] The server receives the voice data, gesture data, and gaze data transmitted from the terminal.
[0712] Input: Sent voice data, gesture data, and gaze data
[0713] How it works: The server uses a speech analysis algorithm to analyze the user's intent of "turn on the light." It understands from gesture data that the user is pointing at the light. It also checks from gaze data where the user is looking.
[0714] Output: User intent and target as analysis results
[0715] Step 4:
[0716] The server creates a virtual interface
[0717] Based on the user's intent analyzed by the server, an appropriate virtual interface (e.g., an "ON / OFF" button for a light) is generated.
[0718] Input: User intent, object
[0719] Operation: Based on the user's intention, the server designs a virtual interface including an "ON / OFF" button for the light and generates it as data.
[0720] Output: Generated virtual interface data
[0721] Step 5:
[0722] The server sends the virtual interface data to the terminal.
[0723] The server transmits the generated virtual interface data to the terminal.
[0724] Input: Virtual interface data
[0725] How it works: The server sends virtual interface data to the device via Wi-Fi or the internet.
[0726] Output: Virtual interface data sent
[0727] Step 6:
[0728] The device displays the virtual interface in AR
[0729] The device displays the virtual interface received from the server in the home space using AR.
[0730] Input: Virtual interface data
[0731] Operation: Uses the device's AR function to display the light's "ON / OFF" button at the specified location.
[0732] Output: AR displayed virtual interface
[0733] Step 7:
[0734] The user operates the virtual interface
[0735] The user selects the "ON" button with their gaze and then presses the "ON" button with a gesture.
[0736] Input: AR displayed virtual interface
[0737] Actions: The device captures the user's gaze and gestures.
[0738] Output: User operation data
[0739] Step 8:
[0740] The device sends signals to control home appliances.
[0741] Based on the user's operation, the terminal generates a control signal for the light (e.g., "turn on") and sends it to the smart light.
[0742] Input: User operation data
[0743] How it works: The device wirelessly sends the appropriate control signal to the smart light.
[0744] Output: Transmitted control signal
[0745] Step 9:
[0746] Home appliances work
[0747] The smart light receives a control signal from the device and performs the specified action (e.g., turning on the light).
[0748] Input: Control signal
[0749] Action: The smart light follows the control signal and turns on.
[0750] Output: The result is that the light turns on.
[0751] (Application example 1)
[0752] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0753] Conventional in-vehicle device operation involves numerous buttons and complex control panels, which can be inconvenient for drivers while driving and potentially compromise safety. In particular, operating entertainment and air conditioning systems requires intuitive methods that allow drivers to keep their eyes on the road. Furthermore, there is a lack of intuitive interfaces that integrate gesture, gaze, and voice commands, so these issues must be addressed simultaneously.
[0754] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0755] In this invention, the server includes a means for detecting gestures, a means for tracking gaze, and a means for receiving voice commands. This allows intuitive operation of in-car devices using gestures, gaze, and voice commands. Specifically, the driver can safely and efficiently control the entertainment system and air conditioning system in the car using gaze, gestures, and voice commands. This minimizes eye movement while driving, enabling intuitive in-car operation, improving convenience and safety for the driver.
[0756] A "gesture" is a way for a user to perform a specific action that is recognized by a camera or sensor and generates a corresponding signal.
[0757] "Eye tracking" is a method of tracking the movement of a user's eyes, recognizing the direction and focus of their gaze, and generating a signal.
[0758] "Voice command" is a means of receiving voice uttered by a user, analyzing the content of the voice, and generating a corresponding signal.
[0759] "Data analytics" is the means of understanding user intent by analyzing information obtained from gestures, gaze, and voice commands.
[0760] A "virtual interface" is a user interface that is generated based on analyzed information and is displayed through a digital device rather than a physical interface.
[0761] "Device control" is a means for users to operate and control specific devices through the generated virtual interface.
[0762] A "device" is a device that captures gestures and eye contact and receives voice commands, including a camera, built-in sensors, and microphone.
[0763] An "image capture device" is a device such as a camera that captures visual information and is used to detect the user's gaze and gestures.
[0764] "Built-in detection devices" refer to sensors built into the device, which are used to capture user behavior and environmental information.
[0765] An "entertainment system" is a system installed in a vehicle that provides entertainment functions such as music playback and video viewing.
[0766] The system that realizes this application example is designed for intuitive in-car operation, integrating gestures, gaze, and voice commands to allow users to control various in-car devices.
[0767] Overall system picture
[0768] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user controls the in-vehicle devices through the displayed virtual interface.
[0769] Device Role
[0770] The terminal has the following roles:
[0771] Gesture and gaze capture: Captures user gestures and gaze movements using built-in sensors and cameras.
[0772] Receiving voice commands: Uses a microphone to receive the user's voice commands.
[0773] Data transmission: The acquired gesture, gaze, and voice data are sent to the server.
[0774] Displaying the virtual interface: The virtual interface received from the server is displayed on the in-car display or head-mounted display.
[0775] Server Roles
[0776] The server has the following roles:
[0777] Data reception: Receives gesture, gaze, and voice data sent from the device.
[0778] Data analysis: Analyzes the received data and understands the user's intent. For example, it recognizes from gesture data that the user has turned up the volume, or from voice data that the user said "Turn up the volume."
[0779] Virtual interface generation: Based on the analysis results, a virtual interface (e.g., the "play / stop" button on a music player) is generated to assist the user in their operations.
[0780] Send virtual interface: Send the created virtual interface to the terminal.
[0781] Hardware and software used
[0782] Hardware: Camera, microphone, built-in sensors
[0783] Software: OpenCV, SpeechRecognition module, GazeTracker, GestureRecognizer
[0784] Processing Overview
[0785] The server analyzes the data sent from the device and generates a virtual interface based on the analysis results. This virtual interface is displayed on the device, and the user can operate the interface using gaze and gestures. For example, if a user says "Turn up the volume" in a car, the server captures and analyzes the voice command. Then, it generates a volume adjustment button on the virtual interface and displays it on the device. This allows the user to intuitively adjust the volume of the entertainment system.
[0786] Specific examples
[0787] To change the volume, the user can say "Turn up the volume." This voice command is captured by the device's microphone and sent to the server. The server analyzes the voice data and understands the intent of the volume adjustment. As a result, a virtual interface for adjusting the volume is displayed on the device. The user can use their gaze or gestures to operate the interface and adjust the volume.
[0788] An example of a prompt for a generative AI model is as follows:
[0789] Write the code to do the camera capture. Use OpenCV.
[0790]
[0791] Write code to recognize speech. Use the SpeechRecognition module.
[0792]
[0793] Write code to track your gaze. Use the GazeTracker module.
[0794]
[0795] Write code to recognize gestures. Use the GestureRecognizer module.
[0796] In this way, an interactive in-car operation system that combines gestures, gaze, and voice can be realized.
[0797] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0798] Step 1:
[0799] Users input gestures, gaze, and voice commands. The device captures these inputs using its built-in camera, sensors, and microphone. Specifically, the camera captures gestures and gaze movements as video, and the microphone captures voice commands as audio data.
[0800] Input: User gestures, eye movements, and voice commands
[0801] Output: Captured video and audio data
[0802] Step 2:
[0803] The device sends the captured data to the server, which includes gesture and gaze information as video data and voice commands as audio data.
[0804] Input: Captured video and audio data
[0805] Output: Data sent to the server
[0806] Step 3:
[0807] The server analyzes gesture and gaze data from the received video data and voice commands from the audio data. Specifically, the gesture recognition module recognizes gestures from the video data, and the eye-tracking module identifies the direction and focus of gaze. The voice recognition module analyzes the audio data and converts the user's intentions into text.
[0808] Input: Data sent to the server
[0809] Output: Analyzed gesture data, gaze data, voice commands
[0810] Step 4:
[0811] The server understands the user's intention based on the analyzed gesture data, gaze data, and voice command. For example, it recognizes that the user turned up the volume from the gesture data, that the user looked at a specific button from the gaze data, and that the user said "Turn up the volume" from the voice command.
[0812] Input: Analyzed gesture data, gaze data, voice commands
[0813] Output: User intent
[0814] Step 5:
[0815] The server generates a virtual interface based on the user's intentions, such as buttons for adjusting the volume or an operation panel for an entertainment system.
[0816] Input: User intent
[0817] Output: Generated virtual interface data
[0818] Step 6:
[0819] The server transmits the generated virtual interface data to the terminal, which receives the data and displays the virtual interface on an in-car display or a head-mounted display.
[0820] Input: Generated virtual interface data
[0821] Output: Virtual interface data sent to the terminal
[0822] Step 7:
[0823] The user operates the displayed virtual interface using gaze, gestures, and voice commands. For example, if the user gazes at the volume control button and gestures to increase the volume, the corresponding operation is executed. The device then transmits the newly acquired user input data back to the server.
[0824] Input: New user input to the virtual interface
[0825] Output: Control signal to in-vehicle equipment
[0826] Step 8:
[0827] The device then sends control signals to the vehicle's equipment based on the analysis results, allowing the equipment to be operated as intended by the user, for example by adjusting the volume.
[0828] Input: Control signal to in-vehicle equipment
[0829] Output: Device operation results
[0830] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0831] This invention relates to a system that controls various smart home appliances based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[0832] Overall system picture
[0833] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user can then operate the home appliance through the displayed virtual interface.
[0834] Device Role
[0835] The terminal has the following roles:
[0836] 1. Gesture and gaze capture:
[0837] Uses built-in sensors and cameras to capture user gestures and eye movements.
[0838] 2. Receiving voice commands:
[0839] A microphone is used to receive voice commands from the user.
[0840] 3. Detecting the user's emotional state:
[0841] It uses a camera and microphone to capture the user's facial expressions and voice tone, and an emotion engine to detect emotions.
[0842] 4. Data transmission:
[0843] The acquired gesture, gaze, voice, and emotion data is sent to the server.
[0844] 5. View virtual interfaces:
[0845] The virtual interface received from the server is displayed in the home space using AR.
[0846] Server Roles
[0847] The server has the following roles:
[0848] 1. Receiving data:
[0849] Receive gesture, gaze, voice, and emotion data sent from the terminal.
[0850] 2. Data Analysis:
[0851] The received data is analyzed to understand the user's intentions and emotional state. For example, hand movements are analyzed from gesture data, gaze data from points of gaze, command content from voice data, and user emotions from emotional data.
[0852] 3. Create a virtual interface:
[0853] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated according to the user's intentions and emotional state.
[0854] 4. Send virtual interface:
[0855] The generated virtual interface data is sent to the terminal.
[0856] User interaction
[0857] Users can use gestures, gaze, and voice commands to communicate their emotional state through facial expressions and tone of voice. For example, if the emotion engine detects that the user is tired, the system can dim the lights or play relaxing music.
[0858] Specific examples
[0859] Lighting Control
[0860] 1. User says "Turn on the light":
[0861] The device's microphone captures this audio and sends it to the server.
[0862] 2. The server analyzes the audio data:
[0863] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[0864] 3. The terminal will display the virtual interface:
[0865] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[0866] 4. User's gaze lands on the "ON" button:
[0867] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[0868] 5. The light comes on:
[0869] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[0870] Playing music
[0871] 1. User controls the music player using gaze and gestures:
[0872] Select the music player icon with your gaze and press the play button with a gesture.
[0873] 2. The device sends the operation data to the server:
[0874] The terminal transmits gaze and gesture data to a server.
[0875] 3. The server analyzes the data and generates a playback signal:
[0876] The server understands the user's intention to play music and generates a play command.
[0877] 4. The device sends a signal to the music player:
[0878] The device sends a playback signal to the music player, and the music is played.
[0879] Emotion-based preferences
[0880] 1. If the user is tired:
[0881] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[0882] 2. The device sends the emotion data to the server:
[0883] The emotion engine transmits the detected emotion data to the server.
[0884] 3. The server analyzes the emotion data:
[0885] The server analyzes that the user is tired and generates a virtual interface for relaxation mode.
[0886] 4. The device will display the virtual interface in relaxed mode:
[0887] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[0888] 5. The environment changes:
[0889] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[0890] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[0891] The processing flow will be explained below.
[0892] Step 1:
[0893] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[0894] Step 2:
[0895] The device monitors the user's movements and captures gesture, gaze, and voice data, such as when the user moves their hand, looks in a particular direction, or makes a sound.
[0896] Step 3:
[0897] The device captures the user's facial expressions and voice tone, which are then analyzed by the emotion engine to generate the user's emotional data, including the user's smiling or serious face and voice tone.
[0898] Step 4:
[0899] The device captures gesture, gaze, voice, and emotion data, stores it internally, and then prepares it for transmission to the server, including packetizing the data and encrypting it for privacy purposes.
[0900] Step 5:
[0901] The device transmits gesture data, gaze data, voice data, and emotion data to the server, where the data is encrypted and transmitted in real time.
[0902] Step 6:
[0903] The server receives the data sent from the terminal, checks the consistency and integrity of the received data, and performs any pre-processing necessary for analysis.
[0904] Step 7:
[0905] The server analyzes the received data, using gesture recognition algorithms, eye tracking algorithms, and voice analysis algorithms to understand the user's intention, and then uses an emotion engine to analyze the user's emotional state. For example, a "raise your hand" gesture, a "turn on the light" voice command, or a smile can be analyzed.
[0906] Step 8:
[0907] Based on the analysis results, the server generates a virtual interface that corresponds to the user's intentions and emotions. For example, if the user is tired, it generates a relaxation mode interface.
[0908] Step 9:
[0909] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[0910] Step 10:
[0911] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze, gestures, and emotional state.
[0912] Step 11:
[0913] The user operates the home appliance using the displayed virtual interface. For example, they can select the "ON" button with their gaze and click the button with their gesture. The interface and behavior are adjusted based on the user's emotions.
[0914] Step 12:
[0915] The device recaptures the user's gaze and gestures to recognize when parts of the virtual interface are manipulated, and emotional state is also continuously monitored.
[0916] Step 13:
[0917] The terminal transmits the operation data to the server, which generates specific control signals for the home appliances. If the emotion data changes, the control signals are adjusted accordingly.
[0918] Step 14:
[0919] The server sends the generated control signal back to the terminal, which then transmits the control signal to the actual home appliance, for example, turning on a light, playing music, or changing the environment setting to relaxation mode.
[0920] This series of steps allows users to enjoy an intuitive and personalized smart home experience through gaze, gesture, voice, and emotion.
[0921] Example 2
[0922] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0923] Current smart home systems provide operations based on user gestures, gaze, and voice commands, but they do not adequately consider the user's emotional state when controlling the interface or operation of home appliances. As a result, they are unable to flexibly respond to the user's diverse needs and momentary emotional state, making it difficult to provide an intuitive and personalized experience.
[0924] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for detecting gestures, means for tracking gaze, means for receiving voice commands, means for analyzing data obtained from the gesture means, the gaze tracking means, and the voice command means, means for generating a virtual interface based on the analysis result, means for displaying the virtual interface, means for controlling home appliances through the virtual interface, means for analyzing the emotional state of the user, and means for adjusting the operation of the virtual interface and the home appliances based on the emotional state. This enables an intuitive and personalized smart home experience based on the user's intentions and emotions.
[0925] "Means for detecting gestures" refers to technology that uses sensors to recognize the movements of a user's hands or body and collects them as data.
[0926] "Means for tracking gaze" refers to technology that uses a camera or sensor to capture the movement of a user's gaze and analyze the position and direction of that gaze.
[0927] "Means for receiving voice commands" refers to technology that uses a microphone to pick up the voices emitted by the user and collects them as data.
[0928] "Means for analyzing data" refers to technology for analyzing collected data such as gestures, gaze, and voice commands to understand the user's intentions and state.
[0929] "Means for generating a virtual interface" refers to a technology for virtually creating an interface that can be operated by a user based on the analysis results.
[0930] "Means for displaying a virtual interface" refers to a technique for displaying the generated virtual interface in a location where the user can see or operate it.
[0931] "Means for controlling home appliances" refers to technology that operates electronic devices and equipment in the home based on user instructions through a virtual interface.
[0932] "Means for analyzing the user's emotional state" refers to technology for determining the user's current emotions from their facial expressions, tone of voice, etc.
[0933] "Means for adjusting the operation of a virtual interface or a home appliance based on the emotional state" refers to a technology for setting the optimal operating mode of a virtual interface or a home appliance based on the analyzed emotional state.
[0934] This system controls various electronic devices in a smart home based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[0935] Overall system picture
[0936] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user then operates the home appliance through the displayed virtual interface.
[0937] Hardware and software used
[0938] Device: A typical smartphone or tablet device, using built-in sensors, cameras, and microphones to capture gestures, gaze, and voice.
[0939] Server: Uses a high-performance cloud-based computer to analyze data, generate virtual interfaces, and manipulate emotional states.
[0940] software:
[0941] Speech recognition engine: Converts voice commands into text.
[0942] AI model: Analyzes gesture and gaze data to understand user intent.
[0943] Emotion engine: Analyzes user emotions from facial expressions and voice tone.
[0944] AR module: displays a virtual interface.
[0945] Specific operation example
[0946] Lighting Control
[0947] 1. The user says "Turn on the light":
[0948] The device's microphone captures this audio and sends it to the server.
[0949] 2. The server analyzes the audio data:
[0950] The server uses a speech recognition engine to convert the command "turn on the lights" into text and understand its intent.
[0951] 3. The terminal will display the virtual interface:
[0952] The server generates an "ON / OFF" interface for the light and sends it to the device, which displays the "ON / OFF" button in AR.
[0953] 4. User's gaze focuses on the "ON" button:
[0954] The device's eye-tracking sensor determines where the user is looking and confirms that they are looking at the "ON" button.
[0955] 5. The light comes on:
[0956] The device sends an "ON" signal to the light, and the light turns on.
[0957] Emotion-based preferences
[0958] 1. If the user is tired:
[0959] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[0960] 2. The device sends the emotion data to the server:
[0961] The emotion engine transmits the detected emotion data to the server.
[0962] 3. The server analyzes the emotion data:
[0963] The server analyzes that the user is tired and generates a virtual interface in relaxation mode.
[0964] 4. The device will display the virtual interface in relaxed mode:
[0965] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[0966] 5. The environment changes:
[0967] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[0968] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[0969] Prompt Sentence Examples
[0970] An example of a prompt is as follows:
[0971] "Turn on the lights"
[0972] "Play music."
[0973] "Put the room in relaxation mode"
[0974] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0975] Step 1:
[0976] The user provides input. Specifically, the user makes gestures toward the device, directs their gaze toward a specific object, or issues a voice command. These inputs are captured by the device. Input data includes gestures, gaze, and voice commands.
[0977] Step 2:
[0978] The device captures data: its built-in sensors capture user gestures, its camera tracks gaze, and its microphone records voice commands. The input data is gestures, gaze, and voice, and the output data is the raw captured data.
[0979] Step 3:
[0980] The device sends the captured data to the server. The device's communication module assembles gesture, gaze, and voice data into packets and transmits them to the server. The input data is the captured raw data, and the output data is the transmitted packets.
[0981] Step 4:
[0982] The server receives the data. The server's network interface receives the data packets sent from the terminal and stores them in memory for analysis. The input data is the data packets from the terminal, and the output data is the data for analysis stored in memory.
[0983] Step 5:
[0984] The server analyzes the data. The server's AI model analyzes the gesture data to identify hand movements and gaze data to identify the user's point of gaze. The voice recognition engine converts voice commands into text and understands the content. Furthermore, the emotion engine analyzes the user's emotional state from their facial expressions and tone of voice. The input data is stored data for analysis, and the output data is the analysis results that indicate the user's intentions and emotional state.
[0985] Examples:
[0986] Recognizes the "waving" gesture from gesture data.
[0987] Eye gaze data determined that the person was "watching TV."
[0988] The command "Turn on the lights" is converted from voice data into text.
[0989] Emotional data was analyzed to determine that the user was relaxed.
[0990] Step 6:
[0991] The server generates a virtual interface based on the analysis results. The server designs a virtual interface according to the user's intentions, such as an "ON / OFF" button on a light or a "PLAY / STOP" button on a music player. The input data is the analysis results, and the output data is the design of the virtual interface.
[0992] Step 7:
[0993] The server sends the virtual interface to the terminal. The generated virtual interface data is packaged into packets and sent to the terminal. The input data is the virtual interface design, and the output data is the packets sent to the terminal.
[0994] Step 8:
[0995] The device displays a virtual interface. The device's AR module renders the virtual interface in space and displays it in the user's field of view. The input data are packets sent from the server, and the output data is the displayed virtual interface.
[0996] Step 9:
[0997] The user operates the virtual interface by looking at or tapping buttons on the displayed virtual interface. The input data is the displayed virtual interface, and the output is the user's operation.
[0998] Step 10:
[0999] The terminal transmits a signal to the home appliance. The communication module of the terminal transmits a control signal to the home appliance in response to a user operation. The input data is the user's operation, and the output data is the control signal transmitted to the home appliance.
[1000] Examples:
[1001] When you look at the "ON" button, the device sends a signal to the light to "turn it on."
[1002] Tapping the "Play" button sends a "Play" signal to the music player.
[1003] This allows the system to deliver an intuitive and personalized smart home experience based on the user's intent and emotional state.
[1004] (Application example 2)
[1005] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1006] While conventional home appliance control systems simplify the operation of individual devices, they are not sufficient to improve the user's in-store experience. Furthermore, the information provided and service suggestions provided in-store are standardized, making it difficult to personalize them to meet the needs and emotional state of each individual user. Furthermore, conventional systems have the problem of being unable to suggest services that take the user's emotions into account.
[1007] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing gesture data, means for analyzing gaze data, means for analyzing voice data, means for analyzing emotion data, means for generating a virtual interface based on the analysis results, means for providing information and services through the virtual interface, and means for analyzing the user's emotional state and proposing optimal services. This enables users to intuitively interact in a physical store using gestures, gaze, and voice commands, and to be provided with personalized services according to their emotional state.
[1008] "Gesture sensing means" refers to a device or sensor used to sense the movement of a user's hands or body.
[1009] "Eye tracking means" means a camera or other device used to track a user's eye movements and determine the direction of gaze.
[1010] "Means for receiving voice commands" refers to a device that captures the voice uttered by the user using a microphone or the like and analyzes it.
[1011] "Means for analyzing" refers to the software or hardware used to process data obtained from gestures, gaze, voice commands, etc., to understand the user's intent and emotional state.
[1012] "Means for generating a virtual interface" refers to a device or software for designing and displaying a virtual interface that can be operated by a user based on the analysis results.
[1013] The term "means for displaying a virtual interface" refers to a device, such as a display or projector, for visually presenting the generated virtual interface to a user.
[1014] "Means for controlling a home appliance" refers to a device or system that controls a home appliance through a virtual interface.
[1015] "Information and service provision means for physical stores" refers to technologies and devices used to provide product information and services to users within a store.
[1016] "Means for analyzing the user's emotional state and suggesting optimal services" refers to software or devices that use cameras or microphones to sense the user's emotional state and suggest appropriate services or information based on that state.
[1017] That's all.
[1018] This invention is a system that receives user gestures, gaze, and voice commands as inputs, analyzes them, and controls home appliances. Furthermore, this system can recognize the user's emotional state and provide services according to that state. Specific embodiments for implementing this invention are described below.
[1019] 1. System Configuration
[1020] Hardware Configuration
[1021] 1. Device:
[1022] Gesture detection: Built-in sensors (e.g. accelerometer, infrared sensor)
[1023] Eye tracking: Camera (e.g. webcam, infrared camera)
[1024] Voice command reception: Microphone
[1025] Virtual interface display: displays, AR headsets
[1026] 2. Server:
[1027] Data analysis: high-performance processors, databases
[1028] Emotion Engine: Emotion Recognition Algorithm
[1029] Software Configuration
[1030] 1. On the device:
[1031] Gesture recognition program: OpenCV
[1032] Eye-tracking software: Dedicated eye-tracking library
[1033] Speech recognition software: Libraries like SpeechRecognition
[1034] Emotion Recognizer: EmotionRecognizer
[1035] 2. Server side:
[1036] Data Analysis Program
[1037] Virtual Interface Generator
[1038] Emotion Engine
[1039] 2. Program processing content
[1040] Gesture Analysis
[1041] The device uses built-in sensors and a camera to capture the user's hand and body movements and detect gestures, and the data is sent to a server where an analysis program interprets the user's intentions.
[1042] Eye tracking
[1043] The device's camera is used to track the user's eye movements and determine the direction of their gaze. The gaze data is sent to a server, where an analysis program identifies the point of gaze.
[1044] Voice command analysis
[1045] The device's microphone receives voice commands, which are interpreted by speech recognition software, and the results are sent to a server, which analyzes the intent of the command.
[1046] Emotion analysis
[1047] The device's camera and microphone are used to capture the user's facial expressions and tone of voice, and this data is analyzed by an emotion recognition program to determine the user's emotional state.
[1048] Virtual Interface Creation
[1049] The server generates a virtual interface based on the analysis of gesture, gaze, voice, and emotion data. The generated interface is sent to the terminal and presented to the user through a display device (such as a monitor or AR headset).
[1050] Information provision and service proposals
[1051] The server then provides optimal information and services based on the results of the user's data analysis. For example, if the user's gaze is directed at a particular product, it will present information about that product. If it detects a tired emotional state, it will suggest a relaxing environment.
[1052] 3. Examples and prompts
[1053] As a concrete example, imagine a user visits a store, looks at a product on the shelf, and detailed information about that product is displayed on smart glasses. Here is an example prompt:
[1054] When a user looks at a product on the shelf, the smart glasses will display detailed information about that product. If the user raises their hand to make a gesture, more information and reviews will be displayed.
[1055] In this way, the present invention can provide a more intuitive and personalized experience based on the user's intent and emotional state.
[1056] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1057] Step 1:
[1058] When a user enters the device's visual range, the device's camera captures the user's gaze and hand gestures. The input is image data from the camera, and the output is gesture and gaze data. The device does this in real time.
[1059] Step 2:
[1060] The device's microphone receives the user's voice commands. The input is the voice data obtained from the microphone, and the output is the interpreted voice command. The voice data is analyzed using the SpeechRecognition library.
[1061] Step 3:
[1062] The captured gesture, gaze, and voice data is sent from the device to the server. The input is the raw data sent from the device, and the output is the raw data sent to the server, which the device transmits via wireless communication.
[1063] Step 4:
[1064] The server analyzes the received gesture, gaze, and voice data. The input is the raw data sent from the device, and the output is the analysis result. The server analyzes the data using OpenCV, a dedicated gaze tracking library, and emotion recognition algorithms.
[1065] Step 5:
[1066] The server uses emotion recognition algorithms to analyze the user's emotional state from their facial expressions and tone of voice. The input is video and audio data from the camera and microphone, and the output is the analyzed emotional state. This is done by the emotion engine.
[1067] Step 6:
[1068] The server generates a virtual interface based on the analysis results. The input is the analyzed gesture, gaze, voice, and emotion data, and the output is the virtual interface. This is done by a virtual interface generation program.
[1069] Step 7:
[1070] The generated virtual interface is sent from the server to the device and displayed on the device's display or AR headset. The input is the virtual interface data sent from the server, and the output is the interface displayed on the device.
[1071] Step 8:
[1072] Users control home appliances using the displayed virtual interface. For example, users select products with their gaze and request more information with hand gestures. The input is the user's gaze and gesture data, and the output is the control command.
[1073] Step 9:
[1074] The server provides information and services based on the user's emotional state. For example, if it detects that the user is tired, it suggests relaxing environments and products. The input is the analyzed emotional data, and the output is the suggested information and services.
[1075] That's all.
[1076] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1077] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1078] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1079] [Third embodiment]
[1080] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1081] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1082] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1083] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1084] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1085] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1086] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1087] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1088] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1089] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1090] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1091] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1092] The present invention relates to a system that enables a user to intuitively operate various home appliances in a smart home, which includes a gesture detection unit, an eye tracking unit, a voice command receiving unit, a data analysis unit, a virtual interface generation unit, a virtual interface display unit, and a home appliance control unit.
[1093] Overall system picture
[1094] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user can then control the home appliances through the displayed virtual interface.
[1095] Device Role
[1096] The terminal has the following roles:
[1097] 1. Gesture and gaze capture:
[1098] Uses built-in sensors and cameras to capture user gestures and eye movements.
[1099] 2. Receiving voice commands:
[1100] A microphone is used to receive voice commands from the user.
[1101] 3. Data transmission:
[1102] The acquired gesture, gaze, and voice data is sent to the server.
[1103] 4. View virtual interfaces:
[1104] The virtual interface received from the server is displayed in the home space using AR.
[1105] Server Roles
[1106] The server has the following roles:
[1107] 1. Receiving data:
[1108] Receives gesture, gaze, and voice data sent from the device.
[1109] 2. Data Analysis:
[1110] The received data is analyzed to understand the user's intent. For example, it can recognize from gesture data that the user has performed the "turn on" action to turn on a light, or from voice data that the user has said "Turn on the light."
[1111] 3. Create a virtual interface:
[1112] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated to assist the user in their operations.
[1113] 4. Send virtual interface:
[1114] The generated virtual interface data is sent to the terminal.
[1115] User interaction
[1116] Users can issue voice commands to their devices using gestures or gaze. For example, if a user is in the living room and says, "Turn on the light," the device captures the voice and sends it to the server. The server analyzes the voice data and generates a control signal to turn on the light based on the user's intention. The device then sends the received control signal to the smart light, and the light turns on.
[1117] Specific examples
[1118] Lighting Control
[1119] 1. User says "Turn on the light":
[1120] The device's microphone captures this audio and sends it to the server.
[1121] 2. The server analyzes the audio data:
[1122] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[1123] 3. The terminal will display the virtual interface:
[1124] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[1125] 4. User's gaze lands on the "ON" button:
[1126] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[1127] 5. The light comes on:
[1128] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[1129] Playing music
[1130] 1. User controls the music player using gaze and gestures:
[1131] Select the music player icon with your gaze and press the play button with a gesture.
[1132] 2. The device sends the operation data to the server:
[1133] The terminal transmits gaze and gesture data to a server.
[1134] 3. The server analyzes the data and generates a playback signal:
[1135] The server understands the user's intention to play music and generates a play command.
[1136] 4. The device sends a signal to the music player:
[1137] The device sends a playback signal to the music player, and the music is played.
[1138] In this way, the system of the present invention can provide users with a highly intuitive and efficient smart home experience.
[1139] The processing flow will be explained below.
[1140] Step 1:
[1141] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[1142] Step 2:
[1143] The device monitors the user's movements and captures gesture, gaze, and voice data, for example, when the user moves their hand, looks in a particular direction, or issues a voice command.
[1144] Step 3:
[1145] The device first stores the captured data internally and then prepares it for transmission to the server, which includes packetizing the data and implementing security measures.
[1146] Step 4:
[1147] The device sends gesture data, gaze data, and voice data to the server, where the data is encrypted and sent in real time.
[1148] Step 5:
[1149] The server receives the data sent from the device and performs checks to ensure the data is consistent and complete.
[1150] Step 6:
[1151] The server analyzes the received data and uses gesture recognition, eye tracking, and voice analysis algorithms to understand the user's intent. For example, it analyzes the "raise your hand" gesture and the "turn on the light" voice command.
[1152] Step 7:
[1153] Based on the analysis results, the server generates a virtual interface according to the user's intentions, such as creating an "ON / OFF" button for a light or a control panel for playing music.
[1154] Step 8:
[1155] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[1156] Step 9:
[1157] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze and gestures.
[1158] Step 10:
[1159] The user operates the home appliance using the displayed virtual interface, for example, by selecting the "ON" button with their gaze and clicking the button with a gesture.
[1160] Step 11:
[1161] The device recaptures the user's gaze and gestures and recognizes that a part of the virtual interface has been manipulated.
[1162] Step 12:
[1163] The terminal sends this information to the server, which generates specific control signals for the home appliances.
[1164] Step 13:
[1165] The server sends the generated control signal back to the device, which then transmits the control signal to the actual home appliance, causing the light to turn on or the music to play.
[1166] This series of steps creates a smart home experience where users can intuitively control home appliances through gaze, gestures, and voice.
[1167] Example 1
[1168] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1169] In recent years, smart homes have become increasingly popular, and users have an increasing need to efficiently control a variety of home appliances. However, operation using conventional remote controls or smartphones is not intuitive, making it difficult for elderly people and those less familiar with technology to use. Furthermore, there are few systems that combine multiple input methods, such as gaze, gestures, and voice, to control home appliances, making it difficult to accurately understand the user's intentions and control them.
[1170] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1171] In this invention, the server includes means for analyzing user gesture, gaze, and voice data, means for generating an appropriate virtual interface based on the analysis results, and means for transmitting the generated virtual interface data to the terminal and displaying it in AR, thereby enabling users to intuitively and efficiently operate home appliances using multiple input means.
[1172] A "gesture detection means" is a device or technology that uses a sensor or camera to recognize the hand or body movements made by a user.
[1173] "Eye tracking means" refers to a device or technology that tracks the user's eye movements using a camera, infrared sensor, etc., to determine their gaze.
[1174] The "means for receiving voice commands" refers to a device or technology that receives the voice uttered by the user using a microphone or the like and analyzes the content of the voice.
[1175] "Means for analyzing data" refers to devices or technologies that process input data such as gestures, gaze, and voice to understand the user's intent.
[1176] "Means for generating a virtual interface" refers to a device or technology that creates a virtual interface that can be operated by the user based on the results of the analysis.
[1177] A "means for displaying a virtual interface" is a device or technique for visually presenting the generated virtual interface to a user.
[1178] A "means for controlling a home appliance" is a device or technology that sends appropriate control signals to a home appliance that is operated by a user through a virtual interface.
[1179] The "means for transmitting data to a server" refers to a device or technology that transmits captured data such as gestures, gazes, and voices to a server via a network.
[1180] "Means for the server to analyze data" refers to a device or technology that processes the gesture, gaze, and voice data received by the server to understand the user's intention.
[1181] "Means for generating an appropriate virtual interface" refers to a device or technology that allows the server to create a virtual interface for the user to operate based on the results of data analysis.
[1182] "Means for transmitting to a terminal and displaying in AR within the home" refers to a device or technology that transmits virtual interface data generated from a server to a terminal, which then displays it in AR (augmented reality) within the home.
[1183] The present invention relates to a system that allows users to intuitively operate smart home appliances. This system detects gestures, gaze, and voice inputs from the user and analyzes the data on a server to understand the user's intentions and generate and display an appropriate virtual interface. Furthermore, the virtual interface can be used to control the home appliances.
[1184] Hardware and Software Examples
[1185] Device: Built-in sensors, camera, microphone
[1186] The device captures gestures, gaze, and voice and sends this data to a server.
[1187] Server: Data analysis device
[1188] The server analyzes the received data and generates a virtual interface according to the user's intentions.
[1189] Virtual interface generation and display devices: AR devices, smartphones
[1190] The generated virtual interface is displayed in AR, providing the user with a visual means of operation.
[1191] Program processing
[1192] 1. Role of the device
[1193] The device uses its built-in sensors and camera to capture user gestures and eye movements, and a microphone to receive voice commands, which are then sent to a server.
[1194] 2. Server Roles
[1195] The server receives gesture, gaze, and voice data sent from the device and analyzes them to understand the user's intentions. Based on the analysis results, it generates an appropriate virtual interface and sends it to the device.
[1196] 3. Viewing Virtual Interfaces
[1197] The device receives the virtual interface from the server and displays it in the home using AR, allowing users to operate the virtual interface with their gaze and gestures to control home appliances.
[1198] Specific examples
[1199] Lighting Control
[1200] 1. The user says "Turn on the light"
[1201] The device's microphone captures this audio and sends it to the server.
[1202] 2. The server analyzes the audio data
[1203] The server understands the intent "turn on the light" and creates an "ON" interface for the light.
[1204] 3. The device displays the virtual interface
[1205] Based on the data received by the device, the light's "ON / OFF" button is displayed inside the home using AR.
[1206] 4. Users look at the "ON" button
[1207] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[1208] 5. The lights come on
[1209] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[1210] Playing music
[1211] 1. Users control the music player using gaze and gestures
[1212] Select the music player icon with your gaze and press the play button with a gesture.
[1213] 2. The device sends the operation data to the server.
[1214] The terminal transmits gaze and gesture data to a server.
[1215] 3. The server analyzes the data and generates a playback signal
[1216] The server understands the user's intention to play music and generates a play command.
[1217] 4. The device sends a signal to the music player
[1218] The device sends a playback signal to the music player, and the music is played.
[1219] Prompt Sentence Examples
[1220] A detailed description of the system can be created by inputting the following prompts into the generative AI model:
[1221] Example prompt:
[1222] "Please explain how this system works when a user wants to control a smart home light. For example, please provide a detailed description of the steps that occur when a user issues a voice command such as 'Turn on the light.'"
[1223] As described above, the system of the present invention enables users to intuitively and efficiently operate smart home appliances.
[1224] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1225] Step 1:
[1226] The terminal captures the user's input
[1227] The user issues the voice command "Turn on the lights."
[1228] Input: Voice commands, gestures, gaze
[1229] How it works: The device's microphone receives voice commands. The device's camera and eye-tracking sensors capture the user's gestures (pointing) and gaze.
[1230] Output: Voice data, gesture data, gaze data
[1231] Step 2:
[1232] The device sends the data to the server
[1233] The device transmits the captured voice data, gesture data, and gaze data to the server.
[1234] Input: Voice data, gesture data, gaze data
[1235] How it works: The device sends data to the server using Wi-Fi, Bluetooth, etc.
[1236] Output: Data sent
[1237] Step 3:
[1238] The server receives and analyzes the data
[1239] The server receives the voice data, gesture data, and gaze data transmitted from the terminal.
[1240] Input: Sent voice data, gesture data, and gaze data
[1241] How it works: The server uses a speech analysis algorithm to analyze the user's intent of "turn on the light." It understands from gesture data that the user is pointing at the light. It also checks from gaze data where the user is looking.
[1242] Output: User intent and target as analysis results
[1243] Step 4:
[1244] The server creates a virtual interface
[1245] Based on the user's intent analyzed by the server, an appropriate virtual interface (e.g., an "ON / OFF" button for a light) is generated.
[1246] Input: User intent, object
[1247] Operation: Based on the user's intention, the server designs a virtual interface including an "ON / OFF" button for the light and generates it as data.
[1248] Output: Generated virtual interface data
[1249] Step 5:
[1250] The server sends the virtual interface data to the terminal.
[1251] The server transmits the generated virtual interface data to the terminal.
[1252] Input: Virtual interface data
[1253] How it works: The server sends virtual interface data to the device via Wi-Fi or the internet.
[1254] Output: Virtual interface data sent
[1255] Step 6:
[1256] The device displays the virtual interface in AR
[1257] The device displays the virtual interface received from the server in the home space using AR.
[1258] Input: Virtual interface data
[1259] Operation: Uses the device's AR function to display the light's "ON / OFF" button at the specified location.
[1260] Output: AR displayed virtual interface
[1261] Step 7:
[1262] The user operates the virtual interface
[1263] The user selects the "ON" button with their gaze and then presses the "ON" button with a gesture.
[1264] Input: AR displayed virtual interface
[1265] Actions: The device captures the user's gaze and gestures.
[1266] Output: User operation data
[1267] Step 8:
[1268] The device sends signals to control home appliances.
[1269] Based on the user's operation, the terminal generates a control signal for the light (e.g., "turn on") and sends it to the smart light.
[1270] Input: User operation data
[1271] How it works: The device wirelessly sends the appropriate control signal to the smart light.
[1272] Output: Transmitted control signal
[1273] Step 9:
[1274] Home appliances work
[1275] The smart light receives a control signal from the device and performs the specified action (e.g., turning on the light).
[1276] Input: Control signal
[1277] Action: The smart light follows the control signal and turns on.
[1278] Output: The result is that the light turns on.
[1279] (Application example 1)
[1280] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1281] Conventional in-vehicle device operation involves numerous buttons and complex control panels, which can be inconvenient for drivers while driving and potentially compromise safety. In particular, operating entertainment and air conditioning systems requires intuitive methods that allow drivers to keep their eyes on the road. Furthermore, there is a lack of intuitive interfaces that integrate gesture, gaze, and voice commands, so these issues must be addressed simultaneously.
[1282] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1283] In this invention, the server includes a means for detecting gestures, a means for tracking gaze, and a means for receiving voice commands. This allows intuitive operation of in-car devices using gestures, gaze, and voice commands. Specifically, the driver can safely and efficiently control the entertainment system and air conditioning system in the car using gaze, gestures, and voice commands. This minimizes eye movement while driving, enabling intuitive in-car operation, improving convenience and safety for the driver.
[1284] A "gesture" is a way for a user to perform a specific action that is recognized by a camera or sensor and generates a corresponding signal.
[1285] "Eye tracking" is a method of tracking the movement of a user's eyes, recognizing the direction and focus of their gaze, and generating a signal.
[1286] "Voice command" is a means of receiving voice uttered by a user, analyzing the content of the voice, and generating a corresponding signal.
[1287] "Data analytics" is the means of understanding user intent by analyzing information obtained from gestures, gaze, and voice commands.
[1288] A "virtual interface" is a user interface that is generated based on analyzed information and is displayed through a digital device rather than a physical interface.
[1289] "Device control" is a means for users to operate and control specific devices through the generated virtual interface.
[1290] A "device" is a device that captures gestures and eye contact and receives voice commands, including a camera, built-in sensors, and microphone.
[1291] An "image capture device" is a device such as a camera that captures visual information and is used to detect the user's gaze and gestures.
[1292] "Built-in detection devices" refer to sensors built into the device, which are used to capture user behavior and environmental information.
[1293] An "entertainment system" is a system installed in a vehicle that provides entertainment functions such as music playback and video viewing.
[1294] The system that realizes this application example is designed for intuitive in-car operation, integrating gestures, gaze, and voice commands to allow users to control various in-car devices.
[1295] Overall system picture
[1296] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user controls the in-vehicle devices through the displayed virtual interface.
[1297] Device Role
[1298] The terminal has the following roles:
[1299] Gesture and gaze capture: Captures user gestures and gaze movements using built-in sensors and cameras.
[1300] Receiving voice commands: Uses a microphone to receive the user's voice commands.
[1301] Data transmission: The acquired gesture, gaze, and voice data are sent to the server.
[1302] Displaying the virtual interface: The virtual interface received from the server is displayed on the in-car display or head-mounted display.
[1303] Server Roles
[1304] The server has the following roles:
[1305] Data reception: Receives gesture, gaze, and voice data sent from the device.
[1306] Data analysis: Analyzes the received data and understands the user's intent. For example, it recognizes from gesture data that the user has turned up the volume, or from voice data that the user said "Turn up the volume."
[1307] Virtual interface generation: Based on the analysis results, a virtual interface (e.g., the "play / stop" button on a music player) is generated to assist the user in their operations.
[1308] Send virtual interface: Send the created virtual interface to the terminal.
[1309] Hardware and software used
[1310] Hardware: Camera, microphone, built-in sensors
[1311] Software: OpenCV, SpeechRecognition module, GazeTracker, GestureRecognizer
[1312] Processing Overview
[1313] The server analyzes the data sent from the device and generates a virtual interface based on the analysis results. This virtual interface is displayed on the device, and the user can operate the interface using gaze and gestures. For example, if a user says "Turn up the volume" in a car, the server captures and analyzes the voice command. Then, it generates a volume adjustment button on the virtual interface and displays it on the device. This allows the user to intuitively adjust the volume of the entertainment system.
[1314] Specific examples
[1315] To change the volume, the user can say "Turn up the volume." This voice command is captured by the device's microphone and sent to the server. The server analyzes the voice data and understands the intent of the volume adjustment. As a result, a virtual interface for adjusting the volume is displayed on the device. The user can use their gaze or gestures to operate the interface and adjust the volume.
[1316] An example of a prompt for a generative AI model is as follows:
[1317] Write the code to do the camera capture. Use OpenCV.
[1318]
[1319] Write code to recognize speech. Use the SpeechRecognition module.
[1320]
[1321] Write code to track your gaze. Use the GazeTracker module.
[1322]
[1323] Write code to recognize gestures. Use the GestureRecognizer module.
[1324] In this way, an interactive in-car operation system that combines gestures, gaze, and voice can be realized.
[1325] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1326] Step 1:
[1327] Users input gestures, gaze, and voice commands. The device captures these inputs using its built-in camera, sensors, and microphone. Specifically, the camera captures gestures and gaze movements as video, and the microphone captures voice commands as audio data.
[1328] Input: User gestures, eye movements, and voice commands
[1329] Output: Captured video and audio data
[1330] Step 2:
[1331] The device sends the captured data to the server, which includes gesture and gaze information as video data and voice commands as audio data.
[1332] Input: Captured video and audio data
[1333] Output: Data sent to the server
[1334] Step 3:
[1335] The server analyzes gesture and gaze data from the received video data and voice commands from the audio data. Specifically, the gesture recognition module recognizes gestures from the video data, and the eye-tracking module identifies the direction and focus of gaze. The voice recognition module analyzes the audio data and converts the user's intentions into text.
[1336] Input: Data sent to the server
[1337] Output: Analyzed gesture data, gaze data, voice commands
[1338] Step 4:
[1339] The server understands the user's intention based on the analyzed gesture data, gaze data, and voice command. For example, it recognizes that the user turned up the volume from the gesture data, that the user looked at a specific button from the gaze data, and that the user said "Turn up the volume" from the voice command.
[1340] Input: Analyzed gesture data, gaze data, voice commands
[1341] Output: User intent
[1342] Step 5:
[1343] The server generates a virtual interface based on the user's intentions, such as buttons for adjusting the volume or an operation panel for an entertainment system.
[1344] Input: User intent
[1345] Output: Generated virtual interface data
[1346] Step 6:
[1347] The server transmits the generated virtual interface data to the terminal, which receives the data and displays the virtual interface on an in-car display or a head-mounted display.
[1348] Input: Generated virtual interface data
[1349] Output: Virtual interface data sent to the terminal
[1350] Step 7:
[1351] The user operates the displayed virtual interface using gaze, gestures, and voice commands. For example, if the user gazes at the volume control button and gestures to increase the volume, the corresponding operation is executed. The device then transmits the newly acquired user input data back to the server.
[1352] Input: New user input to the virtual interface
[1353] Output: Control signal to in-vehicle equipment
[1354] Step 8:
[1355] The device then sends control signals to the vehicle's equipment based on the analysis results, allowing the equipment to be operated as intended by the user, for example by adjusting the volume.
[1356] Input: Control signal to in-vehicle equipment
[1357] Output: Device operation results
[1358] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1359] This invention relates to a system that controls various smart home appliances based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[1360] Overall system picture
[1361] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user can then operate the home appliance through the displayed virtual interface.
[1362] Device Role
[1363] The terminal has the following roles:
[1364] 1. Gesture and gaze capture:
[1365] Uses built-in sensors and cameras to capture user gestures and eye movements.
[1366] 2. Receiving voice commands:
[1367] A microphone is used to receive voice commands from the user.
[1368] 3. Detecting the user's emotional state:
[1369] It uses a camera and microphone to capture the user's facial expressions and voice tone, and an emotion engine to detect emotions.
[1370] 4. Data transmission:
[1371] The acquired gesture, gaze, voice, and emotion data is sent to the server.
[1372] 5. View virtual interfaces:
[1373] The virtual interface received from the server is displayed in the home space using AR.
[1374] Server Roles
[1375] The server has the following roles:
[1376] 1. Receiving data:
[1377] Receive gesture, gaze, voice, and emotion data sent from the terminal.
[1378] 2. Data Analysis:
[1379] The received data is analyzed to understand the user's intentions and emotional state. For example, hand movements are analyzed from gesture data, gaze data from points of gaze, command content from voice data, and user emotions from emotional data.
[1380] 3. Create a virtual interface:
[1381] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated according to the user's intentions and emotional state.
[1382] 4. Send virtual interface:
[1383] The generated virtual interface data is sent to the terminal.
[1384] User interaction
[1385] Users can use gestures, gaze, and voice commands to communicate their emotional state through facial expressions and tone of voice. For example, if the emotion engine detects that the user is tired, the system can dim the lights or play relaxing music.
[1386] Specific examples
[1387] Lighting Control
[1388] 1. User says "Turn on the light":
[1389] The device's microphone captures this audio and sends it to the server.
[1390] 2. The server analyzes the audio data:
[1391] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[1392] 3. The terminal will display the virtual interface:
[1393] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[1394] 4. User's gaze lands on the "ON" button:
[1395] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[1396] 5. The light comes on:
[1397] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[1398] Playing music
[1399] 1. User controls the music player using gaze and gestures:
[1400] Select the music player icon with your gaze and press the play button with a gesture.
[1401] 2. The device sends the operation data to the server:
[1402] The terminal transmits gaze and gesture data to a server.
[1403] 3. The server analyzes the data and generates a playback signal:
[1404] The server understands the user's intention to play music and generates a play command.
[1405] 4. The device sends a signal to the music player:
[1406] The device sends a playback signal to the music player, and the music is played.
[1407] Emotion-based preferences
[1408] 1. If the user is tired:
[1409] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[1410] 2. The device sends the emotion data to the server:
[1411] The emotion engine transmits the detected emotion data to the server.
[1412] 3. The server analyzes the emotion data:
[1413] The server analyzes that the user is tired and generates a virtual interface for relaxation mode.
[1414] 4. The device will display the virtual interface in relaxed mode:
[1415] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[1416] 5. The environment changes:
[1417] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[1418] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[1419] The processing flow will be explained below.
[1420] Step 1:
[1421] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[1422] Step 2:
[1423] The device monitors the user's movements and captures gesture, gaze, and voice data, such as when the user moves their hand, looks in a particular direction, or makes a sound.
[1424] Step 3:
[1425] The device captures the user's facial expressions and voice tone, which are then analyzed by the emotion engine to generate the user's emotional data, including the user's smiling or serious face and voice tone.
[1426] Step 4:
[1427] The device captures gesture, gaze, voice, and emotion data, stores it internally, and then prepares it for transmission to the server, including packetizing the data and encrypting it for privacy purposes.
[1428] Step 5:
[1429] The device transmits gesture data, gaze data, voice data, and emotion data to the server, where the data is encrypted and transmitted in real time.
[1430] Step 6:
[1431] The server receives the data sent from the terminal, checks the consistency and integrity of the received data, and performs any pre-processing necessary for analysis.
[1432] Step 7:
[1433] The server analyzes the received data, using gesture recognition algorithms, eye tracking algorithms, and voice analysis algorithms to understand the user's intention, and then uses an emotion engine to analyze the user's emotional state. For example, a "raise your hand" gesture, a "turn on the light" voice command, or a smile can be analyzed.
[1434] Step 8:
[1435] Based on the analysis results, the server generates a virtual interface that corresponds to the user's intentions and emotions. For example, if the user is tired, it generates a relaxation mode interface.
[1436] Step 9:
[1437] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[1438] Step 10:
[1439] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze, gestures, and emotional state.
[1440] Step 11:
[1441] The user operates the home appliance using the displayed virtual interface. For example, they can select the "ON" button with their gaze and click the button with their gesture. The interface and behavior are adjusted based on the user's emotions.
[1442] Step 12:
[1443] The device recaptures the user's gaze and gestures to recognize when parts of the virtual interface are manipulated, and emotional state is also continuously monitored.
[1444] Step 13:
[1445] The terminal transmits the operation data to the server, which generates specific control signals for the home appliances. If the emotion data changes, the control signals are adjusted accordingly.
[1446] Step 14:
[1447] The server sends the generated control signal back to the terminal, which then transmits the control signal to the actual home appliance, for example, turning on a light, playing music, or changing the environment setting to relaxation mode.
[1448] This series of steps allows users to enjoy an intuitive and personalized smart home experience through gaze, gesture, voice, and emotion.
[1449] Example 2
[1450] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1451] Current smart home systems provide operations based on user gestures, gaze, and voice commands, but they do not adequately consider the user's emotional state when controlling the interface or operation of home appliances. As a result, they are unable to flexibly respond to the user's diverse needs and momentary emotional state, making it difficult to provide an intuitive and personalized experience.
[1452] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for detecting gestures, means for tracking gaze, means for receiving voice commands, means for analyzing data obtained from the gesture means, the gaze tracking means, and the voice command means, means for generating a virtual interface based on the analysis result, means for displaying the virtual interface, means for controlling home appliances through the virtual interface, means for analyzing the emotional state of the user, and means for adjusting the operation of the virtual interface and the home appliances based on the emotional state. This enables an intuitive and personalized smart home experience based on the user's intentions and emotions.
[1453] "Means for detecting gestures" refers to technology that uses sensors to recognize the movements of a user's hands or body and collects them as data.
[1454] "Means for tracking gaze" refers to technology that uses a camera or sensor to capture the movement of a user's gaze and analyze the position and direction of that gaze.
[1455] "Means for receiving voice commands" refers to technology that uses a microphone to pick up the voices emitted by the user and collects them as data.
[1456] "Means for analyzing data" refers to technology for analyzing collected data such as gestures, gaze, and voice commands to understand the user's intentions and state.
[1457] "Means for generating a virtual interface" refers to a technology for virtually creating an interface that can be operated by a user based on the analysis results.
[1458] "Means for displaying a virtual interface" refers to a technique for displaying the generated virtual interface in a location where the user can see or operate it.
[1459] "Means for controlling home appliances" refers to technology that operates electronic devices and equipment in the home based on user instructions through a virtual interface.
[1460] "Means for analyzing the user's emotional state" refers to technology for determining the user's current emotions from their facial expressions, tone of voice, etc.
[1461] "Means for adjusting the operation of a virtual interface or a home appliance based on the emotional state" refers to a technology for setting the optimal operating mode of a virtual interface or a home appliance based on the analyzed emotional state.
[1462] This system controls various electronic devices in a smart home based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[1463] Overall system picture
[1464] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user then operates the home appliance through the displayed virtual interface.
[1465] Hardware and software used
[1466] Device: A typical smartphone or tablet device, using built-in sensors, cameras, and microphones to capture gestures, gaze, and voice.
[1467] Server: Uses a high-performance cloud-based computer to analyze data, generate virtual interfaces, and manipulate emotional states.
[1468] software:
[1469] Speech recognition engine: Converts voice commands into text.
[1470] AI model: Analyzes gesture and gaze data to understand user intent.
[1471] Emotion engine: Analyzes user emotions from facial expressions and voice tone.
[1472] AR module: displays a virtual interface.
[1473] Specific operation example
[1474] Lighting Control
[1475] 1. The user says "Turn on the light":
[1476] The device's microphone captures this audio and sends it to the server.
[1477] 2. The server analyzes the audio data:
[1478] The server uses a speech recognition engine to convert the command "turn on the lights" into text and understand its intent.
[1479] 3. The terminal will display the virtual interface:
[1480] The server generates an "ON / OFF" interface for the light and sends it to the device, which displays the "ON / OFF" button in AR.
[1481] 4. User's gaze focuses on the "ON" button:
[1482] The device's eye-tracking sensor determines where the user is looking and confirms that they are looking at the "ON" button.
[1483] 5. The light comes on:
[1484] The device sends an "ON" signal to the light, and the light turns on.
[1485] Emotion-based preferences
[1486] 1. If the user is tired:
[1487] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[1488] 2. The device sends the emotion data to the server:
[1489] The emotion engine transmits the detected emotion data to the server.
[1490] 3. The server analyzes the emotion data:
[1491] The server analyzes that the user is tired and generates a virtual interface in relaxation mode.
[1492] 4. The device will display the virtual interface in relaxed mode:
[1493] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[1494] 5. The environment changes:
[1495] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[1496] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[1497] Prompt Sentence Examples
[1498] An example of a prompt is as follows:
[1499] "Turn on the lights"
[1500] "Play music."
[1501] "Put the room in relaxation mode"
[1502] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1503] Step 1:
[1504] The user provides input. Specifically, the user makes gestures toward the device, directs their gaze toward a specific object, or issues a voice command. These inputs are captured by the device. Input data includes gestures, gaze, and voice commands.
[1505] Step 2:
[1506] The device captures data: its built-in sensors capture user gestures, its camera tracks gaze, and its microphone records voice commands. The input data is gestures, gaze, and voice, and the output data is the raw captured data.
[1507] Step 3:
[1508] The device sends the captured data to the server. The device's communication module assembles gesture, gaze, and voice data into packets and transmits them to the server. The input data is the captured raw data, and the output data is the transmitted packets.
[1509] Step 4:
[1510] The server receives the data. The server's network interface receives the data packets sent from the terminal and stores them in memory for analysis. The input data is the data packets from the terminal, and the output data is the data for analysis stored in memory.
[1511] Step 5:
[1512] The server analyzes the data. The server's AI model analyzes the gesture data to identify hand movements and gaze data to identify the user's point of gaze. The voice recognition engine converts voice commands into text and understands the content. Furthermore, the emotion engine analyzes the user's emotional state from their facial expressions and tone of voice. The input data is stored data for analysis, and the output data is the analysis results that indicate the user's intentions and emotional state.
[1513] Examples:
[1514] Recognizes the "waving" gesture from gesture data.
[1515] Eye gaze data determined that the person was "watching TV."
[1516] The command "Turn on the lights" is converted from voice data into text.
[1517] Emotional data was analyzed to determine that the user was relaxed.
[1518] Step 6:
[1519] The server generates a virtual interface based on the analysis results. The server designs a virtual interface according to the user's intentions, such as an "ON / OFF" button on a light or a "PLAY / STOP" button on a music player. The input data is the analysis results, and the output data is the design of the virtual interface.
[1520] Step 7:
[1521] The server sends the virtual interface to the terminal. The generated virtual interface data is packaged into packets and sent to the terminal. The input data is the virtual interface design, and the output data is the packets sent to the terminal.
[1522] Step 8:
[1523] The device displays a virtual interface. The device's AR module renders the virtual interface in space and displays it in the user's field of view. The input data are packets sent from the server, and the output data is the displayed virtual interface.
[1524] Step 9:
[1525] The user operates the virtual interface by looking at or tapping buttons on the displayed virtual interface. The input data is the displayed virtual interface, and the output is the user's operation.
[1526] Step 10:
[1527] The terminal transmits a signal to the home appliance. The communication module of the terminal transmits a control signal to the home appliance in response to a user operation. The input data is the user's operation, and the output data is the control signal transmitted to the home appliance.
[1528] Examples:
[1529] When you look at the "ON" button, the device sends a signal to the light to "turn it on."
[1530] Tapping the "Play" button sends a "Play" signal to the music player.
[1531] This allows the system to deliver an intuitive and personalized smart home experience based on the user's intent and emotional state.
[1532] (Application example 2)
[1533] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1534] While conventional home appliance control systems simplify the operation of individual devices, they are not sufficient to improve the user's in-store experience. Furthermore, the information provided and service suggestions provided in-store are standardized, making it difficult to personalize them to meet the needs and emotional state of each individual user. Furthermore, conventional systems have the problem of being unable to suggest services that take the user's emotions into account.
[1535] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing gesture data, means for analyzing gaze data, means for analyzing voice data, means for analyzing emotion data, means for generating a virtual interface based on the analysis results, means for providing information and services through the virtual interface, and means for analyzing the user's emotional state and proposing optimal services. This enables users to intuitively interact in a physical store using gestures, gaze, and voice commands, and to be provided with personalized services according to their emotional state.
[1536] "Gesture sensing means" refers to a device or sensor used to sense the movement of a user's hands or body.
[1537] "Eye tracking means" means a camera or other device used to track a user's eye movements and determine the direction of gaze.
[1538] "Means for receiving voice commands" refers to a device that captures the voice uttered by the user using a microphone or the like and analyzes it.
[1539] "Means for analyzing" refers to the software or hardware used to process data obtained from gestures, gaze, voice commands, etc., to understand the user's intent and emotional state.
[1540] "Means for generating a virtual interface" refers to a device or software for designing and displaying a virtual interface that can be operated by a user based on the analysis results.
[1541] The term "means for displaying a virtual interface" refers to a device, such as a display or projector, for visually presenting the generated virtual interface to a user.
[1542] "Means for controlling a home appliance" refers to a device or system that controls a home appliance through a virtual interface.
[1543] "Information and service provision means for physical stores" refers to technologies and devices used to provide product information and services to users within a store.
[1544] "Means for analyzing the user's emotional state and suggesting optimal services" refers to software or devices that use cameras or microphones to sense the user's emotional state and suggest appropriate services or information based on that state.
[1545] That's all.
[1546] This invention is a system that receives user gestures, gaze, and voice commands as inputs, analyzes them, and controls home appliances. Furthermore, this system can recognize the user's emotional state and provide services according to that state. Specific embodiments for implementing this invention are described below.
[1547] 1. System Configuration
[1548] Hardware Configuration
[1549] 1. Device:
[1550] Gesture detection: Built-in sensors (e.g. accelerometer, infrared sensor)
[1551] Eye tracking: Camera (e.g. webcam, infrared camera)
[1552] Voice command reception: Microphone
[1553] Virtual interface display: displays, AR headsets
[1554] 2. Server:
[1555] Data analysis: high-performance processors, databases
[1556] Emotion Engine: Emotion Recognition Algorithm
[1557] Software Configuration
[1558] 1. On the device:
[1559] Gesture recognition program: OpenCV
[1560] Eye-tracking software: Dedicated eye-tracking library
[1561] Speech recognition software: Libraries like SpeechRecognition
[1562] Emotion Recognizer: EmotionRecognizer
[1563] 2. Server side:
[1564] Data Analysis Program
[1565] Virtual Interface Generator
[1566] Emotion Engine
[1567] 2. Program processing content
[1568] Gesture Analysis
[1569] The device uses built-in sensors and a camera to capture the user's hand and body movements and detect gestures, and the data is sent to a server where an analysis program interprets the user's intentions.
[1570] Eye tracking
[1571] The device's camera is used to track the user's eye movements and determine the direction of their gaze. The gaze data is sent to a server, where an analysis program identifies the point of gaze.
[1572] Voice command analysis
[1573] The device's microphone receives voice commands, which are interpreted by speech recognition software, and the results are sent to a server, which analyzes the intent of the command.
[1574] Emotion analysis
[1575] The device's camera and microphone are used to capture the user's facial expressions and tone of voice, and this data is analyzed by an emotion recognition program to determine the user's emotional state.
[1576] Virtual Interface Creation
[1577] The server generates a virtual interface based on the analysis of gesture, gaze, voice, and emotion data. The generated interface is sent to the terminal and presented to the user through a display device (such as a monitor or AR headset).
[1578] Information provision and service proposals
[1579] The server then provides optimal information and services based on the results of the user's data analysis. For example, if the user's gaze is directed at a particular product, it will present information about that product. If it detects a tired emotional state, it will suggest a relaxing environment.
[1580] 3. Examples and prompts
[1581] As a concrete example, imagine a user visits a store, looks at a product on the shelf, and detailed information about that product is displayed on smart glasses. Here is an example prompt:
[1582] When a user looks at a product on the shelf, the smart glasses will display detailed information about that product. If the user raises their hand to make a gesture, more information and reviews will be displayed.
[1583] In this way, the present invention can provide a more intuitive and personalized experience based on the user's intent and emotional state.
[1584] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1585] Step 1:
[1586] When a user enters the device's visual range, the device's camera captures the user's gaze and hand gestures. The input is image data from the camera, and the output is gesture and gaze data. The device does this in real time.
[1587] Step 2:
[1588] The device's microphone receives the user's voice commands. The input is the voice data obtained from the microphone, and the output is the interpreted voice command. The voice data is analyzed using the SpeechRecognition library.
[1589] Step 3:
[1590] The captured gesture, gaze, and voice data is sent from the device to the server. The input is the raw data sent from the device, and the output is the raw data sent to the server. The device transmits this via wireless communication.
[1591] Step 4:
[1592] The server analyzes the received gesture, gaze, and voice data. The input is the raw data sent from the device, and the output is the analysis result. The server analyzes the data using OpenCV, a dedicated gaze tracking library, and emotion recognition algorithms.
[1593] Step 5:
[1594] The server uses emotion recognition algorithms to analyze the user's emotional state from their facial expressions and tone of voice. The input is video and audio data from the camera and microphone, and the output is the analyzed emotional state. This is done by the emotion engine.
[1595] Step 6:
[1596] The server generates a virtual interface based on the analysis results. The input is the analyzed gesture, gaze, voice, and emotion data, and the output is the virtual interface. This is done by a virtual interface generation program.
[1597] Step 7:
[1598] The generated virtual interface is sent from the server to the device and displayed on the device's display or AR headset. The input is the virtual interface data sent from the server, and the output is the interface displayed on the device.
[1599] Step 8:
[1600] Users control home appliances using the displayed virtual interface. For example, users select products with their gaze and request more information with hand gestures. The input is the user's gaze and gesture data, and the output is the control command.
[1601] Step 9:
[1602] The server provides information and services based on the user's emotional state. For example, if it detects that the user is tired, it suggests relaxing environments and products. The input is the analyzed emotional data, and the output is the suggested information and services.
[1603] That's all.
[1604] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1605] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1606] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1607] [Fourth embodiment]
[1608] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1609] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1610] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1611] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1612] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1613] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1614] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1615] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1616] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1617] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1618] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1619] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1620] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1621] The present invention relates to a system that enables a user to intuitively operate various home appliances in a smart home, which includes a gesture detection unit, an eye tracking unit, a voice command receiving unit, a data analysis unit, a virtual interface generation unit, a virtual interface display unit, and a home appliance control unit.
[1622] Overall system picture
[1623] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user can then control the home appliances through the displayed virtual interface.
[1624] Device Role
[1625] The terminal has the following roles:
[1626] 1. Gesture and gaze capture:
[1627] Uses built-in sensors and cameras to capture user gestures and eye movements.
[1628] 2. Receiving voice commands:
[1629] A microphone is used to receive voice commands from the user.
[1630] 3. Data transmission:
[1631] The acquired gesture, gaze, and voice data is sent to the server.
[1632] 4. View virtual interfaces:
[1633] The virtual interface received from the server is displayed in the home space using AR.
[1634] Server Roles
[1635] The server has the following roles:
[1636] 1. Receiving data:
[1637] Receives gesture, gaze, and voice data sent from the device.
[1638] 2. Data Analysis:
[1639] The received data is analyzed to understand the user's intent. For example, it can recognize from gesture data that the user has performed the "turn on" action to turn on a light, or from voice data that the user has said "Turn on the light."
[1640] 3. Create a virtual interface:
[1641] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated to assist the user in their operations.
[1642] 4. Send virtual interface:
[1643] The generated virtual interface data is sent to the terminal.
[1644] User interaction
[1645] Users can issue voice commands to their devices using gestures or gaze. For example, if a user is in the living room and says, "Turn on the light," the device captures the voice and sends it to the server. The server analyzes the voice data and generates a control signal to turn on the light based on the user's intention. The device then sends the received control signal to the smart light, and the light turns on.
[1646] Specific examples
[1647] Lighting Control
[1648] 1. User says "Turn on the light":
[1649] The device's microphone captures this audio and sends it to the server.
[1650] 2. The server analyzes the audio data:
[1651] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[1652] 3. The terminal will display the virtual interface:
[1653] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[1654] 4. User's gaze lands on the "ON" button:
[1655] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[1656] 5. The light comes on:
[1657] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[1658] Playing music
[1659] 1. User controls the music player using gaze and gestures:
[1660] Select the music player icon with your gaze and press the play button with a gesture.
[1661] 2. The device sends the operation data to the server:
[1662] The terminal transmits gaze and gesture data to a server.
[1663] 3. The server analyzes the data and generates a playback signal:
[1664] The server understands the user's intention to play music and generates a play command.
[1665] 4. The device sends a signal to the music player:
[1666] The device sends a playback signal to the music player, and the music is played.
[1667] In this way, the system of the present invention can provide users with a highly intuitive and efficient smart home experience.
[1668] The processing flow will be explained below.
[1669] Step 1:
[1670] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[1671] Step 2:
[1672] The device monitors the user's movements and captures gesture, gaze, and voice data, for example, when the user moves their hand, looks in a particular direction, or issues a voice command.
[1673] Step 3:
[1674] The device first stores the captured data internally and then prepares it for transmission to the server, which includes packetizing the data and implementing security measures.
[1675] Step 4:
[1676] The device sends gesture data, gaze data, and voice data to the server, where the data is encrypted and sent in real time.
[1677] Step 5:
[1678] The server receives the data sent from the device and performs checks to ensure the data is consistent and complete.
[1679] Step 6:
[1680] The server analyzes the received data and uses gesture recognition, eye tracking, and voice analysis algorithms to understand the user's intent. For example, it analyzes the "raise your hand" gesture and the "turn on the light" voice command.
[1681] Step 7:
[1682] Based on the analysis results, the server generates a virtual interface according to the user's intentions, such as creating an "ON / OFF" button for a light or a control panel for playing music.
[1683] Step 8:
[1684] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[1685] Step 9:
[1686] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze and gestures.
[1687] Step 10:
[1688] The user operates the home appliance using the displayed virtual interface, for example, by selecting the "ON" button with their gaze and clicking the button with a gesture.
[1689] Step 11:
[1690] The device recaptures the user's gaze and gestures and recognizes that a part of the virtual interface has been manipulated.
[1691] Step 12:
[1692] The terminal sends this information to the server, which generates specific control signals for the home appliances.
[1693] Step 13:
[1694] The server sends the generated control signal back to the device, which then transmits the control signal to the actual home appliance, causing the light to turn on or the music to play.
[1695] This series of steps creates a smart home experience where users can intuitively control home appliances through gaze, gestures, and voice.
[1696] Example 1
[1697] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1698] In recent years, smart homes have become increasingly popular, and users have an increasing need to efficiently control a variety of home appliances. However, operation using conventional remote controls or smartphones is not intuitive, making it difficult for elderly people and those less familiar with technology to use. Furthermore, there are few systems that combine multiple input methods, such as gaze, gestures, and voice, to control home appliances, making it difficult to accurately understand the user's intentions and control them.
[1699] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1700] In this invention, the server includes means for analyzing user gesture, gaze, and voice data, means for generating an appropriate virtual interface based on the analysis results, and means for transmitting the generated virtual interface data to the terminal and displaying it in AR, thereby enabling users to intuitively and efficiently operate home appliances using multiple input means.
[1701] A "gesture detection means" is a device or technology that uses a sensor or camera to recognize the hand or body movements made by a user.
[1702] "Eye tracking means" refers to a device or technology that tracks the user's eye movements using a camera, infrared sensor, etc., to determine their gaze.
[1703] The "means for receiving voice commands" refers to a device or technology that receives the voice uttered by the user using a microphone or the like and analyzes the content of the voice.
[1704] "Means for analyzing data" refers to devices or technologies that process input data such as gestures, gaze, and voice to understand the user's intent.
[1705] "Means for generating a virtual interface" refers to a device or technology that creates a virtual interface that can be operated by the user based on the results of the analysis.
[1706] A "means for displaying a virtual interface" is a device or technique for visually presenting the generated virtual interface to a user.
[1707] A "means for controlling a home appliance" is a device or technology that sends appropriate control signals to a home appliance that is operated by a user through a virtual interface.
[1708] The "means for transmitting data to a server" refers to a device or technology that transmits captured data such as gestures, gazes, and voices to a server via a network.
[1709] "Means for the server to analyze data" refers to a device or technology that processes the gesture, gaze, and voice data received by the server to understand the user's intention.
[1710] "Means for generating an appropriate virtual interface" refers to a device or technology that allows the server to create a virtual interface for the user to operate based on the results of data analysis.
[1711] "Means for transmitting to a terminal and displaying in AR within the home" refers to a device or technology that transmits virtual interface data generated from a server to a terminal, which then displays it in AR (augmented reality) within the home.
[1712] The present invention relates to a system that allows users to intuitively operate smart home appliances. This system detects gestures, gaze, and voice inputs from the user and analyzes the data on a server to understand the user's intentions and generate and display an appropriate virtual interface. Furthermore, the virtual interface can be used to control the home appliances.
[1713] Hardware and Software Examples
[1714] Device: Built-in sensors, camera, microphone
[1715] The device captures gestures, gaze, and voice and sends this data to a server.
[1716] Server: Data analysis device
[1717] The server analyzes the received data and generates a virtual interface according to the user's intentions.
[1718] Virtual interface generation and display devices: AR devices, smartphones
[1719] The generated virtual interface is displayed in AR, providing the user with a visual means of operation.
[1720] Program processing
[1721] 1. Role of the device
[1722] The device uses its built-in sensors and camera to capture user gestures and eye movements, and a microphone to receive voice commands, which are then sent to a server.
[1723] 2. Server Roles
[1724] The server receives gesture, gaze, and voice data sent from the device and analyzes them to understand the user's intentions. Based on the analysis results, it generates an appropriate virtual interface and sends it to the device.
[1725] 3. Viewing Virtual Interfaces
[1726] The device receives the virtual interface from the server and displays it in the home using AR, allowing users to operate the virtual interface with their gaze and gestures to control home appliances.
[1727] Specific examples
[1728] Lighting Control
[1729] 1. The user says "Turn on the light"
[1730] The device's microphone captures this audio and sends it to the server.
[1731] 2. The server analyzes the audio data
[1732] The server understands the intent "turn on the light" and creates an "ON" interface for the light.
[1733] 3. The device displays the virtual interface
[1734] Based on the data received by the device, the light's "ON / OFF" button is displayed inside the home using AR.
[1735] 4. Users look at the "ON" button
[1736] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[1737] 5. The lights come on
[1738] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[1739] Playing music
[1740] 1. Users control the music player using gaze and gestures
[1741] Select the music player icon with your gaze and press the play button with a gesture.
[1742] 2. The device sends the operation data to the server.
[1743] The terminal transmits gaze and gesture data to a server.
[1744] 3. The server analyzes the data and generates a playback signal
[1745] The server understands the user's intention to play music and generates a play command.
[1746] 4. The device sends a signal to the music player
[1747] The device sends a playback signal to the music player, and the music is played.
[1748] Prompt Sentence Examples
[1749] A detailed description of the system can be created by inputting the following prompts into the generative AI model:
[1750] Example prompt:
[1751] "Please explain how this system works when a user wants to control a smart home light. For example, please provide a detailed description of the steps that occur when a user issues a voice command such as 'Turn on the light.'"
[1752] As described above, the system of the present invention enables users to intuitively and efficiently operate smart home appliances.
[1753] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1754] Step 1:
[1755] The terminal captures the user's input
[1756] The user issues the voice command "Turn on the lights."
[1757] Input: Voice commands, gestures, gaze
[1758] How it works: The device's microphone receives voice commands. The device's camera and eye-tracking sensors capture the user's gestures (pointing) and gaze.
[1759] Output: Voice data, gesture data, gaze data
[1760] Step 2:
[1761] The device sends the data to the server
[1762] The device transmits the captured voice data, gesture data, and gaze data to the server.
[1763] Input: Voice data, gesture data, gaze data
[1764] How it works: The device sends data to the server using Wi-Fi, Bluetooth, etc.
[1765] Output: Data sent
[1766] Step 3:
[1767] The server receives and analyzes the data
[1768] The server receives the voice data, gesture data, and gaze data transmitted from the terminal.
[1769] Input: Sent voice data, gesture data, and gaze data
[1770] How it works: The server uses a speech analysis algorithm to analyze the user's intent of "turn on the light." It understands from gesture data that the user is pointing at the light. It also checks from gaze data where the user is looking.
[1771] Output: User intent and target as analysis results
[1772] Step 4:
[1773] The server creates a virtual interface
[1774] Based on the user's intent analyzed by the server, an appropriate virtual interface (e.g., an "ON / OFF" button for a light) is generated.
[1775] Input: User intent, object
[1776] Operation: Based on the user's intention, the server designs a virtual interface including an "ON / OFF" button for the light and generates it as data.
[1777] Output: Generated virtual interface data
[1778] Step 5:
[1779] The server sends the virtual interface data to the terminal.
[1780] The server transmits the generated virtual interface data to the terminal.
[1781] Input: Virtual interface data
[1782] How it works: The server sends virtual interface data to the device via Wi-Fi or the internet.
[1783] Output: Virtual interface data sent
[1784] Step 6:
[1785] The device displays the virtual interface in AR
[1786] The device displays the virtual interface received from the server in the home space using AR.
[1787] Input: Virtual interface data
[1788] Operation: Uses the device's AR function to display the light's "ON / OFF" button at the specified location.
[1789] Output: AR displayed virtual interface
[1790] Step 7:
[1791] The user operates the virtual interface
[1792] The user selects the "ON" button with their gaze and then presses the "ON" button with a gesture.
[1793] Input: AR displayed virtual interface
[1794] Actions: The device captures the user's gaze and gestures.
[1795] Output: User operation data
[1796] Step 8:
[1797] The device sends signals to control home appliances.
[1798] Based on the user's operation, the terminal generates a control signal for the light (e.g., "turn on") and sends it to the smart light.
[1799] Input: User operation data
[1800] How it works: The device wirelessly sends the appropriate control signal to the smart light.
[1801] Output: Transmitted control signal
[1802] Step 9:
[1803] Home appliances work
[1804] The smart light receives a control signal from the device and performs the specified action (e.g., turning on the light).
[1805] Input: Control signal
[1806] Action: The smart light follows the control signal and turns on.
[1807] Output: The result is that the light turns on.
[1808] (Application example 1)
[1809] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1810] Conventional in-vehicle device operation involves numerous buttons and complex control panels, which can be inconvenient for drivers while driving and potentially compromise safety. In particular, operating entertainment and air conditioning systems requires intuitive methods that allow drivers to keep their eyes on the road. Furthermore, there is a lack of intuitive interfaces that integrate gesture, gaze, and voice commands, so these issues must be addressed simultaneously.
[1811] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1812] In this invention, the server includes a means for detecting gestures, a means for tracking gaze, and a means for receiving voice commands. This allows intuitive operation of in-car devices using gestures, gaze, and voice commands. Specifically, the driver can safely and efficiently control the entertainment system and air conditioning system in the car using gaze, gestures, and voice commands. This minimizes eye movement while driving, enabling intuitive in-car operation, improving convenience and safety for the driver.
[1813] A "gesture" is a way for a user to perform a specific action that is recognized by a camera or sensor and generates a corresponding signal.
[1814] "Eye tracking" is a method of tracking the movement of a user's eyes, recognizing the direction and focus of their gaze, and generating a signal.
[1815] "Voice command" is a means of receiving voice uttered by a user, analyzing the content of the voice, and generating a corresponding signal.
[1816] "Data analytics" is the means of understanding user intent by analyzing information obtained from gestures, gaze, and voice commands.
[1817] A "virtual interface" is a user interface that is generated based on analyzed information and is displayed through a digital device rather than a physical interface.
[1818] "Device control" is a means for users to operate and control specific devices through the generated virtual interface.
[1819] A "device" is a device that captures gestures and eye contact and receives voice commands, including a camera, built-in sensors, and microphone.
[1820] An "image capture device" is a device such as a camera that captures visual information and is used to detect the user's gaze and gestures.
[1821] "Built-in detection devices" refer to sensors built into the device, which are used to capture user behavior and environmental information.
[1822] An "entertainment system" is a system installed in a vehicle that provides entertainment functions such as music playback and video viewing.
[1823] The system that realizes this application example is designed for intuitive in-car operation, integrating gestures, gaze, and voice commands to allow users to control various in-car devices.
[1824] Overall system picture
[1825] The user inputs gestures, gaze, and voice through the device. The device captures these inputs and sends them to the server. The server analyzes the transmitted data and understands the user's intent. The server then generates an appropriate virtual interface and displays it on the device. The user controls the in-vehicle devices through the displayed virtual interface.
[1826] Device Role
[1827] The terminal has the following roles:
[1828] Gesture and gaze capture: Captures user gestures and gaze movements using built-in sensors and cameras.
[1829] Receiving voice commands: Uses a microphone to receive the user's voice commands.
[1830] Data transmission: The acquired gesture, gaze, and voice data are sent to the server.
[1831] Displaying the virtual interface: The virtual interface received from the server is displayed on the in-car display or head-mounted display.
[1832] Server Roles
[1833] The server has the following roles:
[1834] Data reception: Receives gesture, gaze, and voice data sent from the device.
[1835] Data analysis: Analyzes the received data and understands the user's intent. For example, it recognizes from gesture data that the user has turned up the volume, or from voice data that the user said "Turn up the volume."
[1836] Virtual interface generation: Based on the analysis results, a virtual interface (e.g., the "play / stop" button on a music player) is generated to assist the user in their operations.
[1837] Send virtual interface: Send the created virtual interface to the terminal.
[1838] Hardware and software used
[1839] Hardware: Camera, microphone, built-in sensors
[1840] Software: OpenCV, SpeechRecognition module, GazeTracker, GestureRecognizer
[1841] Processing Overview
[1842] The server analyzes the data sent from the device and generates a virtual interface based on the analysis results. This virtual interface is displayed on the device, and the user can operate the interface using gaze and gestures. For example, if a user says "Turn up the volume" in a car, the server captures and analyzes the voice command. Then, it generates a volume adjustment button on the virtual interface and displays it on the device. This allows the user to intuitively adjust the volume of the entertainment system.
[1843] Specific examples
[1844] To change the volume, the user can say "Turn up the volume." This voice command is captured by the device's microphone and sent to the server. The server analyzes the voice data and understands the intent of the volume adjustment. As a result, a virtual interface for adjusting the volume is displayed on the device. The user can use their gaze or gestures to operate the interface and adjust the volume.
[1845] An example of a prompt for a generative AI model is as follows:
[1846] Write the code to do the camera capture. Use OpenCV.
[1847]
[1848] Write code to recognize speech. Use the SpeechRecognition module.
[1849]
[1850] Write code to track your gaze. Use the GazeTracker module.
[1851]
[1852] Write code to recognize gestures. Use the GestureRecognizer module.
[1853] In this way, an interactive in-car operation system that combines gestures, gaze, and voice can be realized.
[1854] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1855] Step 1:
[1856] Users input gestures, gaze, and voice commands. The device captures these inputs using its built-in camera, sensors, and microphone. Specifically, the camera captures gestures and gaze movements as video, and the microphone captures voice commands as audio data.
[1857] Input: User gestures, eye movements, and voice commands
[1858] Output: Captured video and audio data
[1859] Step 2:
[1860] The device sends the captured data to the server, which includes gesture and gaze information as video data and voice commands as audio data.
[1861] Input: Captured video and audio data
[1862] Output: Data sent to the server
[1863] Step 3:
[1864] The server analyzes gesture and gaze data from the received video data and voice commands from the audio data. Specifically, the gesture recognition module recognizes gestures from the video data, and the eye-tracking module identifies the direction and focus of gaze. The voice recognition module analyzes the audio data and converts the user's intentions into text.
[1865] Input: Data sent to the server
[1866] Output: Analyzed gesture data, gaze data, voice commands
[1867] Step 4:
[1868] The server understands the user's intention based on the analyzed gesture data, gaze data, and voice command. For example, it recognizes that the user turned up the volume from the gesture data, that the user looked at a specific button from the gaze data, and that the user said "Turn up the volume" from the voice command.
[1869] Input: Analyzed gesture data, gaze data, voice commands
[1870] Output: User intent
[1871] Step 5:
[1872] The server generates a virtual interface based on the user's intentions, such as buttons for adjusting the volume or an operation panel for an entertainment system.
[1873] Input: User intent
[1874] Output: Generated virtual interface data
[1875] Step 6:
[1876] The server transmits the generated virtual interface data to the terminal, which receives the data and displays the virtual interface on an in-car display or a head-mounted display.
[1877] Input: Generated virtual interface data
[1878] Output: Virtual interface data sent to the terminal
[1879] Step 7:
[1880] The user operates the displayed virtual interface using gaze, gestures, and voice commands. For example, if the user gazes at the volume control button and gestures to increase the volume, the corresponding operation is executed. The device then transmits the newly acquired user input data back to the server.
[1881] Input: New user input to the virtual interface
[1882] Output: Control signal to in-vehicle equipment
[1883] Step 8:
[1884] The device then sends control signals to the vehicle's equipment based on the analysis results, allowing the equipment to be operated as intended by the user, for example by adjusting the volume.
[1885] Input: Control signal to in-vehicle equipment
[1886] Output: Device operation results
[1887] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1888] This invention relates to a system that controls various smart home appliances based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[1889] Overall system picture
[1890] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user can then operate the home appliance through the displayed virtual interface.
[1891] Device Role
[1892] The terminal has the following roles:
[1893] 1. Gesture and gaze capture:
[1894] Uses built-in sensors and cameras to capture user gestures and eye movements.
[1895] 2. Receiving voice commands:
[1896] A microphone is used to receive voice commands from the user.
[1897] 3. Detecting the user's emotional state:
[1898] It uses a camera and microphone to capture the user's facial expressions and voice tone, and an emotion engine to detect emotions.
[1899] 4. Data transmission:
[1900] The acquired gesture, gaze, voice, and emotion data is sent to the server.
[1901] 5. View virtual interfaces:
[1902] The virtual interface received from the server is displayed in the home space using AR.
[1903] Server Roles
[1904] The server has the following roles:
[1905] 1. Receiving data:
[1906] Receive gesture, gaze, voice, and emotion data sent from the terminal.
[1907] 2. Data Analysis:
[1908] The received data is analyzed to understand the user's intentions and emotional state. For example, hand movements are analyzed from gesture data, gaze data from points of gaze, command content from voice data, and user emotions from emotional data.
[1909] 3. Create a virtual interface:
[1910] Based on the analysis results, a virtual interface (e.g., an "ON / OFF" button for a light) is generated according to the user's intentions and emotional state.
[1911] 4. Send virtual interface:
[1912] The generated virtual interface data is sent to the terminal.
[1913] User interaction
[1914] Users can use gestures, gaze, and voice commands to communicate their emotional state through facial expressions and tone of voice. For example, if the emotion engine detects that the user is tired, the system can dim the lights or play relaxing music.
[1915] Specific examples
[1916] Lighting Control
[1917] 1. User says "Turn on the light":
[1918] The device's microphone captures this audio and sends it to the server.
[1919] 2. The server analyzes the audio data:
[1920] The server understands the intent of "turn on the light" and creates an "ON" interface for the light.
[1921] 3. The terminal will display the virtual interface:
[1922] Based on the data received, the device displays the light's "ON / OFF" button in the home using AR.
[1923] 4. User's gaze lands on the "ON" button:
[1924] The device tracks your gaze and recognizes when you are looking at the "ON" button.
[1925] 5. The light comes on:
[1926] Based on the recognition result, the device sends a "turn on" signal to the smart light, and the light turns on.
[1927] Playing music
[1928] 1. User controls the music player using gaze and gestures:
[1929] Select the music player icon with your gaze and press the play button with a gesture.
[1930] 2. The device sends the operation data to the server:
[1931] The terminal transmits gaze and gesture data to a server.
[1932] 3. The server analyzes the data and generates a playback signal:
[1933] The server understands the user's intention to play music and generates a play command.
[1934] 4. The device sends a signal to the music player:
[1935] The device sends a playback signal to the music player, and the music is played.
[1936] Emotion-based preferences
[1937] 1. If the user is tired:
[1938] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[1939] 2. The device sends the emotion data to the server:
[1940] The emotion engine transmits the detected emotion data to the server.
[1941] 3. The server analyzes the emotion data:
[1942] The server analyzes that the user is tired and generates a virtual interface for relaxation mode.
[1943] 4. The device will display the virtual interface in relaxed mode:
[1944] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[1945] 5. The environment changes:
[1946] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[1947] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[1948] The processing flow will be explained below.
[1949] Step 1:
[1950] The device wakes up its sensors and prepares them to capture gesture, gaze, and voice data in real time, including initializing the camera and microphone and calibrating the sensors.
[1951] Step 2:
[1952] The device monitors the user's movements and captures gesture, gaze, and voice data, such as when the user moves their hand, looks in a particular direction, or makes a sound.
[1953] Step 3:
[1954] The device captures the user's facial expressions and voice tone, which are then analyzed by the emotion engine to generate the user's emotional data, including the user's smiling or serious face and voice tone.
[1955] Step 4:
[1956] The device captures gesture, gaze, voice, and emotion data, stores it internally, and then prepares it for transmission to the server, including packetizing the data and encrypting it for privacy purposes.
[1957] Step 5:
[1958] The device transmits gesture data, gaze data, voice data, and emotion data to the server, where the data is encrypted and transmitted in real time.
[1959] Step 6:
[1960] The server receives the data sent from the terminal, checks the consistency and integrity of the received data, and performs any pre-processing necessary for analysis.
[1961] Step 7:
[1962] The server analyzes the received data, using gesture recognition algorithms, eye tracking algorithms, and voice analysis algorithms to understand the user's intention, and then uses an emotion engine to analyze the user's emotional state. For example, a "raise your hand" gesture, a "turn on the light" voice command, or a smile can be analyzed.
[1963] Step 8:
[1964] Based on the analysis results, the server generates a virtual interface that corresponds to the user's intentions and emotions. For example, if the user is tired, it generates a relaxation mode interface.
[1965] Step 9:
[1966] The server sends the generated virtual interface data to the device, including interface location information and user interaction tracking information.
[1967] Step 10:
[1968] Based on the virtual interface data received by the device, an AR interface is displayed in the home space, and this interface changes dynamically according to the user's gaze, gestures, and emotional state.
[1969] Step 11:
[1970] The user operates the home appliance using the displayed virtual interface. For example, they can select the "ON" button with their gaze and click the button with their gesture. The interface and behavior are adjusted based on the user's emotions.
[1971] Step 12:
[1972] The device recaptures the user's gaze and gestures to recognize when parts of the virtual interface are manipulated, and emotional state is also continuously monitored.
[1973] Step 13:
[1974] The terminal transmits the operation data to the server, which generates specific control signals for the home appliances. If the emotion data changes, the control signals are adjusted accordingly.
[1975] Step 14:
[1976] The server sends the generated control signal back to the terminal, which then transmits the control signal to the actual home appliance, for example, turning on a light, playing music, or changing the environment setting to relaxation mode.
[1977] This series of steps allows users to enjoy an intuitive and personalized smart home experience through gaze, gesture, voice, and emotion.
[1978] Example 2
[1979] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1980] Current smart home systems provide operations based on user gestures, gaze, and voice commands, but they do not adequately consider the user's emotional state when controlling the interface or operation of home appliances. As a result, they are unable to flexibly respond to the user's diverse needs and momentary emotional state, making it difficult to provide an intuitive and personalized experience.
[1981] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for detecting gestures, means for tracking gaze, means for receiving voice commands, means for analyzing data obtained from the gesture means, the gaze tracking means, and the voice command means, means for generating a virtual interface based on the analysis result, means for displaying the virtual interface, means for controlling home appliances through the virtual interface, means for analyzing the emotional state of the user, and means for adjusting the operation of the virtual interface and the home appliances based on the emotional state. This enables an intuitive and personalized smart home experience based on the user's intentions and emotions.
[1982] "Means for detecting gestures" refers to technology that uses sensors to recognize the movements of a user's hands or body and collects them as data.
[1983] "Means for tracking gaze" refers to technology that uses a camera or sensor to capture the movement of a user's gaze and analyze the position and direction of that gaze.
[1984] "Means for receiving voice commands" refers to technology that uses a microphone to pick up the voices emitted by the user and collects them as data.
[1985] "Means for analyzing data" refers to technology for analyzing collected data such as gestures, gaze, and voice commands to understand the user's intentions and state.
[1986] "Means for generating a virtual interface" refers to a technology for virtually creating an interface that can be operated by a user based on the analysis results.
[1987] "Means for displaying a virtual interface" refers to a technique for displaying the generated virtual interface in a location where the user can see or operate it.
[1988] "Means for controlling home appliances" refers to technology that operates electronic devices and equipment in the home based on user instructions through a virtual interface.
[1989] "Means for analyzing the user's emotional state" refers to technology for determining the user's current emotions from their facial expressions, tone of voice, etc.
[1990] "Means for adjusting the operation of a virtual interface or a home appliance based on the emotional state" refers to a technology for setting the optimal operating mode of a virtual interface or a home appliance based on the analyzed emotional state.
[1991] This system controls various electronic devices in a smart home based on the input and analysis of user gestures, gaze, and voice commands. Furthermore, by incorporating an emotion engine, it recognizes the user's emotional state and adjusts the interface and behavior of the appliances accordingly.
[1992] Overall system picture
[1993] The user uses the device to input gestures, gaze, and voice. The device captures these inputs and sends the data to the server. The server analyzes the data to understand the user's intentions and emotional state. The server then generates a virtual interface based on the analysis results and displays it on the device. The user then operates the home appliance through the displayed virtual interface.
[1994] Hardware and software used
[1995] Device: A typical smartphone or tablet device, using built-in sensors, cameras, and microphones to capture gestures, gaze, and voice.
[1996] Server: Uses a high-performance cloud-based computer to analyze data, generate virtual interfaces, and manipulate emotional states.
[1997] software:
[1998] Speech recognition engine: Converts voice commands into text.
[1999] AI model: Analyzes gesture and gaze data to understand user intent.
[2000] Emotion engine: Analyzes user emotions from facial expressions and voice tone.
[2001] AR module: displays a virtual interface.
[2002] Specific operation example
[2003] Lighting Control
[2004] 1. The user says "Turn on the light":
[2005] The device's microphone captures this audio and sends it to the server.
[2006] 2. The server analyzes the audio data:
[2007] The server uses a speech recognition engine to convert the command "turn on the lights" into text and understand its intent.
[2008] 3. The terminal will display the virtual interface:
[2009] The server generates an "ON / OFF" interface for the light and sends it to the device, which displays the "ON / OFF" button in AR.
[2010] 4. User's gaze focuses on the "ON" button:
[2011] The device's eye-tracking sensor determines where the user is looking and confirms that they are looking at the "ON" button.
[2012] 5. The light comes on:
[2013] The device sends an "ON" signal to the light, and the light turns on.
[2014] Emotion-based preferences
[2015] 1. If the user is tired:
[2016] The device's camera and microphone detect fatigue from the user's facial expressions and tone of voice.
[2017] 2. The device sends the emotion data to the server:
[2018] The emotion engine transmits the detected emotion data to the server.
[2019] 3. The server analyzes the emotion data:
[2020] The server analyzes that the user is tired and generates a virtual interface in relaxation mode.
[2021] 4. The device will display the virtual interface in relaxed mode:
[2022] The device offers a relaxation mode, for example by dimming the lights and playing relaxation music.
[2023] 5. The environment changes:
[2024] The device dims the smart lights and sends a signal to the music player to play relaxation music.
[2025] In this way, the system of the present invention provides a more intuitive and personalized smart home experience based on the user's intentions and emotions.
[2026] Prompt Sentence Examples
[2027] An example of a prompt is as follows:
[2028] "Turn on the lights"
[2029] "Play music."
[2030] "Put the room in relaxation mode"
[2031] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2032] Step 1:
[2033] The user provides input. Specifically, the user makes gestures toward the device, directs their gaze toward a specific object, or issues a voice command. These inputs are captured by the device. Input data includes gestures, gaze, and voice commands.
[2034] Step 2:
[2035] The device captures data: its built-in sensors capture user gestures, its camera tracks gaze, and its microphone records voice commands. The input data is gestures, gaze, and voice, and the output data is the raw captured data.
[2036] Step 3:
[2037] The device sends the captured data to the server. The device's communication module assembles gesture, gaze, and voice data into packets and transmits them to the server. The input data is the captured raw data, and the output data is the transmitted packets.
[2038] Step 4:
[2039] The server receives the data. The server's network interface receives the data packets sent from the terminal and stores them in memory for analysis. The input data is the data packets from the terminal, and the output data is the data for analysis stored in memory.
[2040] Step 5:
[2041] The server analyzes the data. The server's AI model analyzes the gesture data to identify hand movements and gaze data to identify the user's point of gaze. The voice recognition engine converts voice commands into text and understands the content. Furthermore, the emotion engine analyzes the user's emotional state from their facial expressions and tone of voice. The input data is stored data for analysis, and the output data is the analysis results that indicate the user's intentions and emotional state.
[2042] Examples:
[2043] Recognizes the "waving" gesture from gesture data.
[2044] Eye gaze data determined that the person was "watching TV."
[2045] The command "Turn on the lights" is converted from voice data into text.
[2046] Emotional data was analyzed to determine that the user was relaxed.
[2047] Step 6:
[2048] The server generates a virtual interface based on the analysis results. The server designs a virtual interface according to the user's intentions, such as an "ON / OFF" button on a light or a "PLAY / STOP" button on a music player. The input data is the analysis results, and the output data is the design of the virtual interface.
[2049] Step 7:
[2050] The server sends the virtual interface to the terminal. The generated virtual interface data is packaged into packets and sent to the terminal. The input data is the virtual interface design, and the output data is the packets sent to the terminal.
[2051] Step 8:
[2052] The device displays a virtual interface. The device's AR module renders the virtual interface in space and displays it in the user's field of view. The input data are packets sent from the server, and the output data is the displayed virtual interface.
[2053] Step 9:
[2054] The user operates the virtual interface by looking at or tapping buttons on the displayed virtual interface. The input data is the displayed virtual interface, and the output is the user's operation.
[2055] Step 10:
[2056] The terminal transmits a signal to the home appliance. The communication module of the terminal transmits a control signal to the home appliance in response to a user operation. The input data is the user's operation, and the output data is the control signal transmitted to the home appliance.
[2057] Examples:
[2058] When you look at the "ON" button, the device sends a signal to the light to "turn it on."
[2059] Tapping the "Play" button sends a "Play" signal to the music player.
[2060] This allows the system to deliver an intuitive and personalized smart home experience based on the user's intent and emotional state.
[2061] (Application example 2)
[2062] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2063] While conventional home appliance control systems simplify the operation of individual devices, they are not sufficient to improve the user's in-store experience. Furthermore, the information provided and service suggestions provided in-store are standardized, making it difficult to personalize them to meet the needs and emotional state of each individual user. Furthermore, conventional systems have the problem of being unable to suggest services that take the user's emotions into account.
[2064] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing gesture data, means for analyzing gaze data, means for analyzing voice data, means for analyzing emotion data, means for generating a virtual interface based on the analysis results, means for providing information and services through the virtual interface, and means for analyzing the user's emotional state and proposing optimal services. This enables users to intuitively interact in a physical store using gestures, gaze, and voice commands, and to be provided with personalized services according to their emotional state.
[2065] "Gesture sensing means" refers to a device or sensor used to sense the movement of a user's hands or body.
[2066] "Eye tracking means" means a camera or other device used to track a user's eye movements and determine the direction of gaze.
[2067] "Means for receiving voice commands" refers to a device that captures the voice uttered by the user using a microphone or the like and analyzes it.
[2068] "Means for analyzing" refers to the software or hardware used to process data obtained from gestures, gaze, voice commands, etc., to understand the user's intent and emotional state.
[2069] "Means for generating a virtual interface" refers to a device or software for designing and displaying a virtual interface that can be operated by a user based on the analysis results.
[2070] The term "means for displaying a virtual interface" refers to a device, such as a display or projector, for visually presenting the generated virtual interface to a user.
[2071] "Means for controlling a home appliance" refers to a device or system that controls a home appliance through a virtual interface.
[2072] "Information and service provision means for physical stores" refers to technologies and devices used to provide product information and services to users within a store.
[2073] "Means for analyzing the user's emotional state and suggesting optimal services" refers to software or devices that use cameras or microphones to sense the user's emotional state and suggest appropriate services or information based on that state.
[2074] That's all.
[2075] This invention is a system that receives user gestures, gaze, and voice commands as inputs, analyzes them, and controls home appliances. Furthermore, this system can recognize the user's emotional state and provide services according to that state. Specific embodiments for implementing this invention are described below.
[2076] 1. System Configuration
[2077] Hardware Configuration
[2078] 1. Device:
[2079] Gesture detection: Built-in sensors (e.g. accelerometer, infrared sensor)
[2080] Eye tracking: Camera (e.g. webcam, infrared camera)
[2081] Voice command reception: Microphone
[2082] Virtual interface display: displays, AR headsets
[2083] 2. Server:
[2084] Data analysis: high-performance processors, databases
[2085] Emotion Engine: Emotion Recognition Algorithm
[2086] Software Configuration
[2087] 1. On the device:
[2088] Gesture recognition program: OpenCV
[2089] Eye-tracking software: Dedicated eye-tracking library
[2090] Speech recognition software: Libraries like SpeechRecognition
[2091] Emotion Recognizer: EmotionRecognizer
[2092] 2. Server side:
[2093] Data Analysis Program
[2094] Virtual Interface Generator
[2095] Emotion Engine
[2096] 2. Program processing content
[2097] Gesture Analysis
[2098] The device uses built-in sensors and a camera to capture the user's hand and body movements and detect gestures, and the data is sent to a server where an analysis program interprets the user's intentions.
[2099] Eye tracking
[2100] The device's camera is used to track the user's eye movements and determine the direction of their gaze. The gaze data is sent to a server, where an analysis program identifies the point of gaze.
[2101] Voice command analysis
[2102] The device's microphone receives voice commands, which are interpreted by speech recognition software, and the results are sent to a server, which analyzes the intent of the command.
[2103] Emotion analysis
[2104] The device's camera and microphone are used to capture the user's facial expressions and tone of voice, and this data is analyzed by an emotion recognition program to determine the user's emotional state.
[2105] Virtual Interface Creation
[2106] The server generates a virtual interface based on the analysis of gesture, gaze, voice, and emotion data. The generated interface is sent to the terminal and presented to the user through a display device (such as a monitor or AR headset).
[2107] Information provision and service proposals
[2108] The server then provides optimal information and services based on the results of the user's data analysis. For example, if the user's gaze is directed at a particular product, it will present information about that product. If it detects a tired emotional state, it will suggest a relaxing environment.
[2109] 3. Examples and prompts
[2110] As a concrete example, imagine a user visits a store, looks at a product on the shelf, and detailed information about that product is displayed on smart glasses. Here is an example prompt:
[2111] When a user looks at a product on the shelf, the smart glasses will display detailed information about that product. If the user raises their hand to make a gesture, more information and reviews will be displayed.
[2112] In this way, the present invention can provide a more intuitive and personalized experience based on the user's intent and emotional state.
[2113] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2114] Step 1:
[2115] When a user enters the device's visual range, the device's camera captures the user's gaze and hand gestures. The input is image data from the camera, and the output is gesture and gaze data. The device does this in real time.
[2116] Step 2:
[2117] The device's microphone receives the user's voice commands. The input is the voice data obtained from the microphone, and the output is the interpreted voice command. The voice data is analyzed using the SpeechRecognition library.
[2118] Step 3:
[2119] The captured gesture, gaze, and voice data is sent from the device to the server. The input is the raw data sent from the device, and the output is the raw data sent to the server, which the device transmits via wireless communication.
[2120] Step 4:
[2121] The server analyzes the received gesture, gaze, and voice data. The input is the raw data sent from the device, and the output is the analysis result. The server analyzes the data using OpenCV, a dedicated gaze tracking library, and emotion recognition algorithms.
[2122] Step 5:
[2123] The server uses emotion recognition algorithms to analyze the user's emotional state from their facial expressions and tone of voice. The input is video and audio data from the camera and microphone, and the output is the analyzed emotional state. This is done by the emotion engine.
[2124] Step 6:
[2125] The server generates a virtual interface based on the analysis results. The input is the analyzed gesture, gaze, voice, and emotion data, and the output is the virtual interface. This is done by a virtual interface generation program.
[2126] Step 7:
[2127] The generated virtual interface is sent from the server to the device and displayed on the device's display or AR headset. The input is the virtual interface data sent from the server, and the output is the interface displayed on the device.
[2128] Step 8:
[2129] Users control home appliances using the displayed virtual interface. For example, users select products with their gaze and request more information with hand gestures. The input is the user's gaze and gesture data, and the output is the control command.
[2130] Step 9:
[2131] The server provides information and services based on the user's emotional state. For example, if it detects that the user is tired, it suggests relaxing environments and products. The input is the analyzed emotional data, and the output is the suggested information and services.
[2132] That's all.
[2133] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2134] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2135] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2136] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2137] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2138] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2139] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2140] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2141] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2142] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2143] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2144] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2145] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2146] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2147] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2148] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2149] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2150] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2151] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2152] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2153] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2154] The following is further disclosed regarding the above embodiment.
[2155] (Claim 1)
[2156] a means for detecting a gesture;
[2157] a means for tracking gaze; and
[2158] means for receiving voice commands;
[2159] means for analyzing data obtained from said gesture means, eye tracking means and voice command means;
[2160] means for generating a virtual interface based on the analysis results;
[2161] means for displaying the virtual interface;
[2162] A system including means for controlling a home appliance through said virtual interface.
[2163] (Claim 2)
[2164] 2. The system according to claim 1, wherein the means for detecting the gesture uses a built-in sensor of the terminal.
[2165] (Claim 3)
[2166] 2. The system according to claim 1, wherein the means for tracking the gaze uses a camera of the terminal.
[2167] (Claim 4)
[2168] 2. The system of claim 1, wherein the means for receiving voice commands uses a microphone of the terminal.
[2169] (Claim 5)
[2170] 10. The system of claim 1, wherein the means for displaying the virtual interface displays the AR data in a space within the home.
[2171] (Claim 6)
[2172] 2. The system according to claim 1, wherein the means for controlling the home appliance uses a control signal generated based on the analysis result.
[2173] "Example 1"
[2174] (Claim 1)
[2175] a means for detecting a gesture;
[2176] a means for tracking gaze; and
[2177] means for receiving voice commands;
[2178] means for analyzing data obtained from said gesture means, eye tracking means and voice command means;
[2179] means for generating a virtual interface based on the analysis results;
[2180] means for displaying the virtual interface;
[2181] means for controlling a home appliance through the virtual interface;
[2182] A means for transmitting the acquired user gesture, gaze, and voice data to a server;
[2183] A means for the server to analyze the received data, understand the user's intent, and generate an appropriate virtual interface;
[2184] The system includes a means for transmitting virtual interface data from a server to a terminal and displaying it in AR within the home.
[2185] (Claim 2)
[2186] 2. The system according to claim 1, wherein the means for detecting the gesture uses a built-in sensor of the terminal.
[2187] (Claim 3)
[2188] 2. The system according to claim 1, wherein the means for tracking the gaze uses a camera of the terminal.
[2189] "Application Example 1"
[2190] (Claim 1)
[2191] a means for detecting a gesture;
[2192] a means for tracking gaze; and
[2193] means for receiving voice commands;
[2194] means for analyzing data obtained from said gesture means, eye tracking means and voice command means;
[2195] means for generating a virtual interface based on the analysis results;
[2196] means for displaying the virtual interface;
[2197] A system including means for controlling an appliance through said virtual interface.
[2198] (Claim 2)
[2199] 2. The system of claim 1, wherein the means for detecting gestures uses a built-in detection device of the terminal.
[2200] (Claim 3)
[2201] 2. The system according to claim 1, wherein the means for tracking the gaze uses an image capture device of the terminal.
[2202] (Claim 4)
[2203] 2. The system according to claim 1, further comprising means for intuitively operating devices in the vehicle.
[2204] (Claim 5)
[2205] 10. The system of claim 1, further comprising means for controlling an in-vehicle entertainment system based on gaze, gesture, and voice commands.
[2206] "Example 2: Combining Emotion Engines"
[2207] (Claim 1)
[2208] a means for detecting a gesture;
[2209] a means for tracking gaze; and
[2210] means for receiving voice commands;
[2211] means for analyzing data obtained from said gesture means, eye tracking means and voice command means;
[2212] means for generating a virtual interface based on the analysis results;
[2213] means for displaying the virtual interface;
[2214] means for controlling a home appliance through the virtual interface;
[2215] a means for analyzing the emotional state of a user;
[2216] means for adjusting the behavior of a virtual interface or a home appliance based on said emotional state;
[2217] A system including:
[2218] (Claim 2)
[2219] 2. The system according to claim 1, wherein the means for detecting the gesture uses a built-in sensor of the terminal.
[2220] (Claim 3)
[2221] 2. The system according to claim 1, wherein the means for tracking the gaze uses a camera of the terminal.
[2222] "Application example 2 when combining emotion engines"
[2223] (Claim 1)
[2224] a means for detecting a gesture;
[2225] a means for tracking gaze; and
[2226] means for receiving voice commands;
[2227] means for analyzing data obtained from said gesture means, eye tracking means and voice command means;
[2228] means for generating a virtual interface based on the analysis result;
[2229] means for displaying the virtual interface;
[2230] means for controlling a home appliance through the virtual interface;
[2231] A means for brick-and-mortar stores to provide information and services through a virtual interface;
[2232] A system that includes a means for analyzing a user's emotional state and suggesting optimal services.
[2233] (Claim 2)
[2234] 2. The system according to claim 1, wherein the means for detecting the gesture uses a built-in sensor of the terminal.
[2235] (Claim 3)
[2236] 2. The system according to claim 1, wherein the means for tracking the gaze uses a camera of the terminal.
[2237] That's all. [Explanation of symbols]
[2238] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for detecting a gesture; a means for tracking gaze; means for receiving voice commands; means for analyzing data obtained from said gesture means, eye tracking means and voice command means; means for generating a virtual interface based on the analysis results; means for displaying the virtual interface; A system including means for controlling a home appliance through said virtual interface.
2. 2. The system according to claim 1, wherein the means for detecting the gesture uses a built-in sensor of the terminal.
3. 2. The system of claim 1, wherein the means for tracking the gaze uses a camera of the terminal.
4. 2. The system of claim 1, wherein the means for receiving voice commands uses a microphone of the terminal.
5. 2. The system of claim 1, wherein the means for displaying the virtual interface displays AR data in a space within the home.
6. 2. The system according to claim 1, wherein the means for controlling the home appliance uses a control signal generated based on the analysis result.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A